---
title: "AI Citations for Law Firms Start in Your Intake Data - Carlos Arias"
description: "AI citations for law firms start with language, not schema. Here's how to mine intake calls, site search and forms for the phrasing models answer."
url: "https://carlosarias.com/blog/law-firm-marketing/law-firm-intake-language-ai-citations"
---

[Law Firm Marketing](/blog/categories/law-firm-marketing)

# AI Citations for Law Firms Start in Your Intake Data

AI citations for law firms start with language, not schema. Here's how to mine intake calls, site search and forms for the phrasing models answer.

  [Carlos Arias](/blog/authors/carlos-arias) · September 19, 2026  · 8 min read

![A black ink brush character on cream paper, its last stroke trailing into a small circle above a red seal.](/_astro/cover.Avuy-Rg4_Z2434WC.webp)

*A black ink brush character on cream paper, its last stroke trailing into a small circle above a red seal. AI-generated illustration by Carlos Arias .*

If you want AI citations for law firms to start landing on your pages, stop rewriting title tags and go read your call transcripts. Answer engines reward content phrased the way people actually ask. Your intake recordings and that contact-form free-text box already hold the phrasing, captured from people who were minutes away from hiring a lawyer. Your site search log holds the rest.

Most firms are sitting on that data. Almost none have opened it.

That’s the method. The rest of this piece is how to run it without ever publishing something a client told you in confidence.

## Schema is not the lever

The shallow version of answer engine optimization for legal goes like this: bolt FAQPage markup onto the practice pages and wait for ChatGPT to notice. It doesn’t work, and we now have a measurement of how much it doesn’t work.

Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, then matched each one against control URLs on other domains with similar pre-period citation levels. The result was noise: +2.4% in Google AI Mode, +2.2% in ChatGPT, and a small but statistically significant 4.6% decline in AI Overviews. The correlation everyone quotes in pitch decks, that cited pages are almost three times more likely to carry JSON-LD than uncited ones, held up fine. The causation didn’t.

Structured data still earns its keep for entity clarity and rich results. We ship it. We just won’t sell it to you as the reason a model quotes your firm, because the evidence says it isn’t.

## Why AI citations for law firms no longer follow your rankings

Here’s the part that reorders the whole budget conversation.

Ahrefs also ran 863,000 keywords against 4 million AI Overview URLs and found that only 38% of cited pages ranked in the top 10 for the query that produced the citation. In their July 2025 analysis, that figure was 76%. It halved in about seven months. The rest of the citations split almost evenly between positions 11 to 100 (31.2%) and beyond position 100 (31.0%).

Ahrefs attributes the shift to query fan-out. When an AI Overview fires, Google decomposes the original query into a spray of related sub-questions, then cites the pages that keep showing up across the whole set. Your position for the head term is one input among dozens. Coverage of the surrounding questions is the real qualifier.

Now put that next to how the questions are worded. Semrush ran 17 months of ChatGPT clickstream data against its 27-billion-keyword database and clocked the average Google query at 3.4 words against a 23-word ChatGPT prompt. Switch ChatGPT’s search mode on and the prompt collapses back to 4.2 words. Left in chat mode, 65% to 85% of those prompts matched no keyword in the database at all.

Your keyword research file is three words wide. The question that triggers the answer is twenty-three. That’s the gap, and no amount of markup closes it.

## Step 1: pull the four places client language already lives

Every firm we look at has at least two of these running and unread. Nobody assigned them an owner, so they fill up and get purged on a retention schedule.

### The four sources worth opening

- Call tracking transcripts, first 40 seconds only. The opening of an intake call is the caller explaining their own situation before anyone coaches them into legal vocabulary. That’s the raw material. After intake starts asking structured questions, the language is yours, not theirs.
- Internal site search. Usually a neglected log in your analytics property. It’s short and high-intent, and it tells you which questions your navigation failed to answer.
- Contact-form free text. The “briefly describe your situation” box. This is the single densest source of natural phrasing you own, because people type it unsupervised at 11pm.
- Voicemail and after-hours chat. Same content, worse capture. If those inquiries are dying in a mailbox, you have a bigger problem than AEO, and we’ve written about where law firm leads actually die.

### Strip it before it touches a content brief

We are not flexible on this constraint. You are mining phrasing patterns, not matters. Strip names, dates, carriers, venues, injuries, dollar figures and anything else that identifies a human before the text goes anywhere near a content brief. No client’s words get republished. No matter facts become marketing. What survives the strip is grammar and vocabulary: how a worried person describes a problem out loud. That’s all you needed.

## Step 2: convert the strip into question strings, not keywords

A keyword is a noun phrase. A question string is a full sentence with the caller’s anxiety still attached, and that’s the format the models are matching against.

### What the translation actually looks like

A keyword tool hands you “boca raton car accident lawyer.” What people actually say, once you’ve read fifty openings, is closer to: the other driver’s insurance already called me and wants a recorded statement, do I need a lawyer or am I overreacting. Same practice area. Completely different retrieval target.

### Cluster until four or five shapes are left

Write the strings down verbatim in shape, sanitized in content. Then cluster them. You’ll typically find that a single practice area produces somewhere between eight and twenty recurring question shapes, and that four or five of them account for most of the volume. Those four or five are your page.

Notice what happened. You didn’t do keyword research. You did transcription, and the research was already finished by the people who called you.

## Step 3: rewrite the practice page so the answer is the first thing on it

This is where a specific rewrite shape matters more than another paragraph of theory.

### Brochure page against retrieval page

A standard car accident page, the kind three quarters of firms run, is built as a brochure:

Standard page
Rewritten for retrieval

H1: Car Accident Attorney in [City]
H1 unchanged, then a 45-word plain answer to the top question string before any other element

H2: Why Choose Our Firm
H2: Should you give the insurance adjuster a recorded statement?

H2: Types of Car Accident Cases We Handle
H2: What happens if the other driver’s insurance calls before you hire anyone

H2: Our Results
H2: How long you have to file in Florida, and what pauses that clock

H2: Contact Us Today
H2: What to bring to a first consultation

The left column answers “who are you.” The right column answers “what do I do right now.” Only one of those is a question anybody types into ChatGPT at 11pm.

### Three rules govern the rewrite

Each H2 has to be a question a real caller asked, in their words, not the practice-area label. Each section has to open with the answer in one or two sentences, then support it, because a section that buries its conclusion under three paragraphs of context can’t be lifted cleanly into a generated answer. And every claim in it has to survive a bar-counsel read, which means no outcome promises and no comparative superlatives dressed up as helpfulness.

What you end up with is a page that covers the fan-out surface instead of a single head term. That’s the mechanism the Ahrefs data is describing. AI citations for law firms land on the page that keeps appearing across the sub-queries, not on the page holding position three.

### Do not spin this across forty city pages

One caution before you scale. Do not run the same question set across forty near-identical city pages. That’s low-value programmatic SEO with a new coat of paint, and the failure mode is identical.

## Step 4: fix the entity, then accept what you can’t see

The pages are now phrased correctly. Whether a model says your name out loud is a separate problem, and it is only half yours to solve.

### Two numbers from 126 million prompts

Two facts from Semrush’s 2026 AI Visibility Index, which analyzed 126 million US AI search prompts between January and April 2026 across ChatGPT, Gemini, AI Mode and AI Overviews.

First, ChatGPT cites an average of 15 sources per response while Gemini cites three. Being the fourth-best source on a topic still gets you into one of those and not the other. Second, 61.7% of citations are ghost citations: the source link is there, the brand name never appears in the answer text. Your content can be doing the work and staying invisible in the output.

### AI citations for law firms need a resolvable entity

Which is exactly why the firm name, the attorney names, the office address and the practice descriptions have to be byte-identical across the site, the Google Business Profile, the bar directories and every citation source underneath them. When a model does decide to name someone, it names the entity it can resolve confidently. That work overlaps almost entirely with the local foundation we described in SEO for law firms in 2026.

## Measuring this without lying to yourself

Semrush’s same index reports that 45% of marketing leaders can’t accurately measure brand visibility in AI answers, and only 9% have tooling that tracks every metric they care about. Believe that number. It matches what we see.

So here’s what we actually instrument. Take the specific question strings you built pages around and check the four surfaces by hand on a fixed cadence. Log every check in a sheet with the date against it. Branded search volume goes next to it, because assisted AI research usually surfaces a day later as someone Googling your firm name. Then referral traffic from chatgpt.com and perplexity.ai, with the volume reading far smaller than the influence. It always does.

What we won’t do is hand you a GEO score. Nobody has a defensible one. The honest report says “you appeared for six of the eleven strings we targeted, here are the six, here’s what changed.”

We have been building software professionally since 1999, and all in on agentic AI since 2022. Different decades, same habit. Read what the person actually said before you design the thing that answers them. The people who win a new retrieval system are the ones who wrote that down, while everyone else optimized the container.

## What we refuse to sell you here

No guarantee that a model will name your firm. Nobody controls model output, and anyone quoting you a citation count is quoting a number they invented.

No AI-written practice pages built from these transcripts. The whole advantage of this method is that the language is real. Running it through a generator to produce forty variants destroys the only asset you had.

And no chat widget as the answer to any of this. That’s a capture tool at best, and it’s not what earns a citation.

## If this is your problem

If your content keeps getting ignored by the models and you’ve already paid someone for schema and a content calendar, the diagnosis is usually language, not markup. Your call recordings know the questions. The pages don’t ask them yet.

The mining pass is a few hours of reading, and you can run it yourself with what’s described above. If you’d rather have it done by the operator who would actually build the pages, write us at hi@carlosarias.com and tell us which practice area leaks the most.

    Tags [#AI Search](/tags/ai-search/)
#Answer Engine Optimization [#Law Firm SEO](/tags/law-firm-seo/)[#Content Strategy](/tags/content-strategy/)[#Intake Automation](/tags/intake-automation/)   Share        Written by [Carlos Arias](/blog/authors/carlos-arias)

Marketing Engineer for law firms. I combine digital marketing, software, data, automation and AI to improve the whole system — from first click to signed case.

         On this page

- Schema is not the lever
- Why AI citations for law firms no longer follow your rankings
- Step 1: pull the four places client language already lives
- The four sources worth opening
- Strip it before it touches a content brief
- Step 2: convert the strip into question strings, not keywords
- What the translation actually looks like
- Cluster until four or five shapes are left
- Step 3: rewrite the practice page so the answer is the first thing on it
- Brochure page against retrieval page
- Three rules govern the rewrite
- Do not spin this across forty city pages
- Step 4: fix the entity, then accept what you can’t see
- Two numbers from 126 million prompts
- AI citations for law firms need a resolvable entity
- Measuring this without lying to yourself
- What we refuse to sell you here
- If this is your problem

## Continue reading

      [Law Firm Marketing](/blog/categories/law-firm-marketing) · September 11, 2026  [### How Law Firm Marketing Actually Works: The Systems Guide (2026)](/blog/law-firm-marketing/law-firm-marketing-systems-guide-2026/)

Law firm marketing is not posts and ads — it is the system from search to signed case. SEO, local, paid, website, intake and measurement, with the places firms lose cases they already paid for.

  Carlos Arias · 18 min
      [Law Firm Marketing](/blog/categories/law-firm-marketing) · September 11, 2026  [### SEO for Law Firms in 2026: Local First, Then Practice Pages, Then AI Citations](/blog/law-firm-marketing/seo-for-law-firms-2026/)

Law firm SEO in 2026 starts with the map pack, then practice-area pages that convert, then eligibility for AI citations. Timelines, intake, and what agencies get wrong — with no ranking guarantees.

  Carlos Arias · 14 min
      [Law Firm Marketing](/blog/categories/law-firm-marketing) · September 18, 2026  [### Law Firm Website AI Visibility: It's an Architecture Problem](/blog/law-firm-marketing/machine-first-architecture-law-firm-website/)

Law firm website AI visibility is an architecture problem, not a content problem. The four pillars, what breaks on a practice-area page, and the fix.

  Carlos Arias · 8 min

## Stay in the loop.

One email when it’s worth it — new posts and updates, no spam.

Thanks — check your inbox to confirm.

Free. Unsubscribe in one click.
