New · A dedicated AI SEO channel: see conversions from ChatGPT, Perplexity, Claude & Gemini →
← Blog
Part of: How to Track ChatGPT, Perplexity & Gemini Conversions, Not Just Mentions →

LLM Tracking: How It Works, What to Measure and How to Choose a Tool

Portrait of Samy ThuillierBy ··12 min read
LLM tracking dashboard comparing brand mentions across AI assistants with AI referral conversions by landing page

LLM tracking is monitoring how AI assistants such as ChatGPT, Gemini, Claude, Perplexity and Google’s AI Overviews mention, cite and describe your brand. A tracker sends the same set of prompts to each assistant on a schedule, checks every answer for your brand, competitors, links and tone, and reports how often you appear over time. It measures presence in answers, so you pair it with referral traffic and conversions to see what that presence is worth.

This guide covers how the tools collect data, which metrics hold up, how many prompts you need before a change is real, what GA4 can and cannot show, how to choose a tool and how to set tracking up. It also covers the parts most tool roundups skip: the other meaning of the term, the sample size math, debugging bad data, and a worked example that runs from visibility to revenue.

What is LLM tracking?

A large language model (LLM) is the model that writes the answer. ChatGPT, Gemini, Claude and Perplexity are apps built on top of such models, with web search and other features added. When people say “LLM tracking” in marketing, they mean tracking those apps’ answers about a category: does the assistant name you, link to you, describe you correctly, and recommend you over a rival?

It exists because classic tools do not see these answers. A rank tracker records where your URL sits for a keyword on a results page. An assistant does not show a fixed list. It writes a new answer each time, and that answer can differ by wording, location, account and model version. Your page can rank first on Google and still be missing from the AI answer to the same question.

Rank trackingSocial listeningLLM tracking
WatchesSearch results pagesPosts on social networks, news, forumsAnswers generated by AI assistants
Unit trackedKeywordMention in a postPrompt (or a keyword expanded into prompts)
Core metricPosition of your URLMention volume and sentimentShare of answers that mention or cite you
StabilityFairly stable day to dayDepends on what people postVaries from run to run, so needs repeated sampling
Tells you about traffic?Indirectly, with click-through estimatesNoNo. Traffic needs analytics

Two different things are called LLM tracking

Before you compare tools, check which job you have. Google’s own AI Overview for this query asks whether you want brand visibility or cost and token tracking, and several vendor pages blur the two.

Brand visibility tracking (this guide)LLM observability (developer tracking)
Who uses itMarketing, SEO, PR, foundersEngineers building an app on a language model
Whose modelPublic assistants: ChatGPT, Gemini, Claude, Perplexity, AI OverviewsYour own app and the model API it calls
What gets loggedAnswers to prompts you chose: mentions, citations, position, sentimentEvery request: prompt, response, tokens, cost, latency, errors
Typical question“Does ChatGPT recommend us for our category?”“Why did our support bot get slower and pricier this week?”
Example toolsSemrush AI toolkit, SE Ranking, Profound, Peec AI, Otterly, LLMrefsLangfuse, LangSmith, Helicone

If you landed here to monitor an app you built, you want an observability tool. The rest of this guide is about the marketing meaning.

How LLM tracking tools work

Every tracker follows the same loop, with different choices at each step:

  1. Choose inputs. Most tools ask you for prompts. Some take your SEO keywords and expand each one into several conversational prompts (LLMrefs works this way). Others start from buyer personas and generate the prompts those people would type (Gumshoe’s approach).
  2. Run them. The tool sends each prompt to each assistant, in each chosen country and language, on a daily or weekly schedule.
  3. Parse the answers. It detects brand and competitor names, linked citations and their domains, where in a list you appear, and the tone of the sentence around your name.
  4. Aggregate. It turns many answers into rates, share of voice and trend lines, often with a proprietary score on top.

API collection vs interface collection

This is the method question to ask every vendor. Results can be pulled through a model’s API or by loading the consumer app the way a person would.

Through the APIThrough the app interface
Speed and costFast and cheap at scaleSlower and more expensive to run
Closeness to what users seeCan differ: the app adds search, formatting and features the raw model lacksCloser to a real user’s answer
Rich elements (shopping cards, maps, tables)Usually missingCaptured
Good forLarge prompt sets, model behavior researchReporting what buyers actually see

Neither method sees a logged-in user’s memory or past chats, so no tool reproduces an individual buyer’s answer exactly. Treat every tracker as a sample of typical answers, not a copy of what each person saw.

The metrics that matter

MetricHow it is calculatedUse it for
Mention rate (visibility %)Answers that name your brand ÷ all answers checkedYour headline number, per engine and per prompt group
Citation rateAnswers that link to your domain ÷ all answersWhether you are a source, which is what can send clicks
Share of voiceYour mentions ÷ mentions of you plus tracked competitorsComparing against named rivals
Average positionMean rank of your brand inside list-style answersA secondary signal only (see below)
SentimentTone of the text around your nameSpotting negative framing early
AccuracyAnswers with wrong facts about you ÷ branded answersFixing outdated pricing, features or claims
Cited source domainsWhich sites the assistant cites for your promptsFinding the pages and publications to earn a place on

Track mention and citation rate separately. An assistant can name you without linking, or cite your page while recommending someone else. For the scoring math behind vendor dashboards, see our breakdown of the AI visibility score.

Where the top guides disagree

Is “position” in an AI answer meaningful?

Some trackers sell average position as a core metric, like a rank. Other guides argue that because assistants reorder their lists from one run to the next, a position is close to meaningless and only the share of answers that include you counts. Both have a point. Position from a single run is noise. Position averaged over many runs of the same prompt is a usable secondary signal, as long as the mention rate is your primary one. Never report a single “we rank #2 in ChatGPT” screenshot as a result.

What does each tool cost?

Review articles in the current results list different prices, prompt limits and included engines for the same products. Vendors in this category change plans often, and many charge per prompt, per answer credit or per extra engine. Use list articles to build a shortlist, then read the price and the date on the vendor’s own pricing page. Check three things: how many answers per month your prompt set needs, which engines are add-ons, and how much history you keep.

How many prompts you need before a change is real

Guides say you need a “statistically meaningful” sample. Here is what that means in numbers. A mention rate is a proportion, so its rough 95% margin of error follows the standard formula:

Margin of error ≈ 1.96 × √( p × (1 − p) ÷ n )

Here p is your mention rate and n is the number of answers behind it. The example below is illustrative, using a mention rate of 30%.

Answers checked (n)Margin of errorYour 30% really means
25±18 pointsAnywhere from about 12% to 48%
100±9 pointsAbout 21% to 39%
400±4.5 pointsAbout 25.5% to 34.5%

Comparing two periods is harder, because both carry error. With 400 answers in each month and a rate near 30%, the change has to be about 6 points or more (1.96 × √(2 × 0.21 ÷ 400) ≈ 0.064) before you can call it real. With 100 answers per month, a jump from 30% to 36% tells you nothing.

Then work out how many answers your setup produces. Again illustrative:

Illustrative monthly answer volume
40 prompts x 4 engines x 3 runs a week x 4 weeks = 1,920 answers a month
Per engine: 40 x 3 x 4 = 480 answers, about ±4 points at 30%
Per prompt group of 10 prompts on one engine: 120 answers, about ±8 points

Two caveats. Repeated runs of the same prompt are not fully independent, so treat these margins as the best case. And the sample you need depends on the level you report at: a brand-wide number settles quickly, while a single prompt on a single engine never will. Report at the level your sample supports.

The three layers of LLM tracking

Most tools cover only the first layer. A complete setup watches all three, because each answers a different question.

LayerWhat you seeWhere the data comes fromBlind spot
1. AnswersMentions, citations, sentiment, competitors in AI answersPrompt trackersSays nothing about clicks or buyers
2. AI crawlersWhich AI bots fetch which of your pages, and how oftenServer or CDN logs, bot analytics featuresA fetch is not a recommendation
3. Visits and conversionsAI referral sessions by landing page, conversions, valueYour analytics or conversion trackerCannot see the prompt; visits with no referrer look Direct

For layer 2, the assistants publish their user agents. OpenAI’s crawler documentation separates OAI-SearchBot (used to show sites in ChatGPT search), GPTBot (used for model training) and ChatGPT-User (fetches a page when a user asks, and robots.txt rules may not apply to it). Perplexity’s crawler documentation makes the same split between PerplexityBot and Perplexity-User. Search your logs for these names: a page the search bots never fetch is unlikely to be cited.

What GA4 can and cannot tell you about LLM traffic

A common complaint from marketers: GA4 shows visits from ChatGPT or Perplexity, but not what people asked. That is how it works, and no setting changes it.

GA4 can showGA4 cannot show
Sessions whose referrer is an assistant domain (chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com and others)The prompt or conversation that led to the click
The landing page of each of those sessionsAnswers that mentioned you but were never clicked
Conversions and their value from those sessionsAI visits that arrive with no referrer, which are counted as Direct
Any utm_source tag present on the arriving linkWhich competitor the assistant recommended alongside you

So the landing page is your best clue to intent. A visit that lands on your comparison page came from a comparison question. A visit to a deep how-to article came from a problem question. Build a custom channel group or segment for assistant referrers, then report by landing page. Our guide to tracking conversions from ChatGPT, Perplexity and Gemini walks through the setup. If you want that layer without building it, SEOConversion identifies AI assistant visits from referrers and supported UTM signals and reports their conversions and value by landing page. It does not track prompts or what assistants say, so it complements a prompt tracker rather than replacing one, and visits with no referrer stay Direct.

Worked example: from visibility to revenue by landing page

All numbers here are illustrative, for a made-up B2B software company. Its prompt tracker shows a 22% mention rate overall: strong on “best tool for” prompts, weak on “X vs Y” comparisons. Is the comparison gap worth fixing? Visibility alone cannot say. Here are its AI referral visits for one quarter, with each demo request valued at $500 and each newsletter signup at $20 (see how to calculate conversion value for setting your own figures).

Landing pageAI visitsConversionsValueValue per visit
/pricing1809 demos$4,500$25.00
/compare/us-vs-rival2406 demos$3,000$12.50
/integrations/crm603 demos$1,500$25.00
/blog/how-to-guide5201 demo + 4 signups$580$1.12
Total1,000$9,580$9.58

The blog post gets over half the AI visits but under 7% of the value ($580 of $9,580). The comparison page earns $12.50 a visit even though the brand is weak on comparison prompts. That points the work at comparison prompts: if better coverage there added another 240 visits at the same rate, that would be about $3,000 a quarter. Raising the overall mention rate by pushing more how-to citations would mostly add $1-a-visit traffic. The prompt tracker tells you where you are missing. The landing page report tells you which gaps pay.

Decision rule

Rank prompt groups by the value per AI visit of the page that answers them, not by how far your mention rate trails competitors. Fix the gap where a visit is worth the most first.

LLM tracking tools by category

The tools named most often across this year’s roundups fall into five groups. Within a group, the differences are mostly engine coverage, prompt limits, collection method and price.

CategoryExamplesBest fit
SEO suites with an AI moduleSemrush AI Visibility Toolkit, SE Ranking, Ahrefs Brand Radar, Nightwatch, SurferTeams already paying for the suite who want AI and classic rankings side by side
Dedicated AI visibility trackersProfound, Peec AI, Otterly, LLMrefs, ZipTie, Scrunch, AthenaHQ, GumshoeTeams that need prompt-level detail, more engines or more countries
GEO and content platforms with trackingWritesonic, AirOps, Goodie, AIclicksTeams that want tracking and content production in one place
Media and PR monitoringMeltwater GenAI LensComms teams that already monitor news and social
Enterprise search intelligenceGrowByDataLarge retail and multi-market brands buying custom reporting

How to choose an LLM tracking tool

If this is true for youPrioritizeAsk the vendor
You already use an SEO suite dailyIts AI module firstWhich engines are included, and is collection via the app or the API?
Buyers use several assistantsEngine coverage on your planWhich engines cost extra, and is data broken out per engine?
You sell in several countries or languagesLocation and language targetingCan I run the same prompt per country and compare?
You manage many brands or clientsMulti-project accounts and exportsIs pricing per domain, per prompt or per credit?
You must prove revenueExports or API to join with analyticsCan I export answers with dates to match my landing page data?
Budget is tightA small frozen prompt set, or manual trackingHow many answers a month does my set use?

Run two shortlisted tools on the same frozen prompt set for a month, and spot check 10 answers by hand in each. Keep the tool whose numbers match what you see, not the one with the best-looking score.

How to set up LLM tracking in six steps

  1. Decide the question. Category visibility, answer accuracy, competitor movement or revenue from AI visits. One primary goal decides prompts, engines and tools.
  2. Build and freeze a prompt set. Write 30 to 50 prompts across four intents: category (“best X for Y”), comparison (“X vs Y”), problem (“how do I fix Z”) and branded (“is X good for Y”, for accuracy). Source them from Search Console queries, sales calls, support tickets and forum threads. Save the list with a version date.
  3. Choose engines and locations. Track the assistants and countries your buyers use, and read each engine on its own, because each cites different sources.
  4. Add competitors and set the cadence. Add three to five rivals. Run each prompt several times per period so your rates rest on enough answers (see the sample size section).
  5. Connect traffic and conversions. Report AI referral visits by landing page, with conversions and value. Without this layer you can grow visibility and never know whether it paid. The SEO conversion tracking guide covers the same landing-page method for organic search.
  6. Review monthly, refresh deliberately. Act on gaps and wrong facts each month. Change the prompt set only on a schedule, such as quarterly, and note the new version so trends stay comparable.

The free route uses the same steps with a spreadsheet: one row per prompt, engine and run, with columns for mentioned, linked, position, tone and cited domains. Use a fresh or logged-out session each time.

When LLM tracking data looks wrong: debugging

SymptomLikely causeWhat to do
Your manual check disagrees with the toolYour account’s memory, location or model version differs from the tool’sRecheck logged out, in the tool’s country. Compare rates over many runs, not one answer
Every brand dropped at onceModel update, or the vendor changed its collection methodCheck whether competitors moved too and read the vendor changelog before touching content
Big swings week to weekToo few answers behind the numberAdd runs, or report at a higher level (all prompts, not one)
Mentions you know are false positivesBrand name is a common word or shared with another companyAdd exclusion terms or track the exact product name
Visibility up, AI referral visits flatMentions without links, or visits arriving with no referrerCheck citation rate. Watch Direct visits to deep pages that rarely get Direct traffic
Trend broke after a quarterPrompt set changed, or you switched toolsRe-baseline. Never splice two tools or two prompt versions into one line
Never cited, even for prompts you should winAI search bots blocked by robots.txt, a firewall or the CDNSearch logs for OAI-SearchBot and PerplexityBot, then allow them

Acting on what you find

  • Missing from category prompts: look at the cited source domains. The review sites, comparison articles and forum threads assistants cite are where you need a presence.
  • Wrong facts in branded answers: update the pages that state them, then the third-party pages the assistant cites for those claims.
  • Cited but not named: make your brand and product name explicit in the passages that get quoted, not only in the logo and header.
  • Never cited: confirm the search bots can fetch your pages. Allowing OAI-SearchBot while blocking GPTBot is a valid choice: per OpenAI’s documentation the two settings are independent.

For the content side of earning mentions, see how to rank in ChatGPT.

If you meant AI tracking you

Some people search this term because they want the opposite: to limit what AI services record about them. That is a privacy question. In each assistant you use, open the privacy or data settings and turn off use of your chats for model training, use temporary chats for sensitive topics, and delete history you do not need. On your phone, review which apps can access your microphone, location and contacts. Brand LLM tracking tools do not see individual users’ chats. They only see answers to prompts the tool sends itself.

Frequently asked questions

What is LLM tracking and how does it work?

LLM tracking is monitoring how AI assistants such as ChatGPT, Gemini, Claude, Perplexity and Google’s AI Overviews mention, cite and describe your brand. A tool sends a fixed set of prompts to each assistant on a schedule, parses every answer for your brand, competitors, links and tone, and reports the results as rates over time. The same term is also used by developers for monitoring their own LLM apps, which is a different job.

What are the best LLM visibility trackers?

There is no single best one, because the tools solve different problems. If you already pay for an SEO suite, start with its AI module (Semrush, SE Ranking, Ahrefs Brand Radar or Nightwatch). If you need broad engine coverage and prompt-level detail, look at dedicated trackers such as Profound, Peec AI, Otterly or LLMrefs. Shortlist two, run the same frozen prompt set in both for a month and keep the one whose data matches your manual spot checks.

Is ChatGPT an LLM?

ChatGPT is a chatbot app built on large language models made by OpenAI, so in everyday speech people call it an LLM. Strictly, the LLM is the model underneath and ChatGPT is the product around it, which adds web search, memory and other features. That difference matters for tracking: answers in the ChatGPT app can differ from answers the same model gives through the API.

Is there a free way to do LLM tracking?

Yes, if your prompt set is small. Write 20 to 30 buyer questions, run each one in a logged-out or fresh session on each assistant, and log in a spreadsheet whether you were mentioned, linked, where you appeared and which sources were cited. Repeat on a fixed schedule. It is slow, and small samples swing a lot, but it teaches you how assistants treat your category before you pay for a tool.

Can Google Analytics show which prompts sent traffic to my site?

No. GA4 can show sessions that arrive with a referrer such as chatgpt.com or perplexity.ai, plus the landing page and any conversions, but assistants do not pass the user’s prompt. Visits that arrive with no referrer are counted as Direct. The landing page is the best available clue to what the person asked.

How do I stop AI from tracking me?

That is a different question from brand LLM tracking. To limit what AI services keep about you, open the privacy or data settings in each assistant you use and turn off use of your chats for model training, use temporary or incognito chats for sensitive topics, and review which apps on your phone can use your microphone, location and contacts.

Prompt trackers show where you appear. Conversions show what it pays.

SEOConversion identifies visits from ChatGPT, Perplexity, Claude, Gemini and Copilot by referrer and reports their conversions and value by landing page, with one cookieless script.

Start free