To track brand mentions in ChatGPT, run a fixed set of real buyer questions under documented conditions and save each answer over time. Record whether your brand is absent, mentioned, cited, or recommended; which competitors appear; which sources are shown; and what changed since the previous check. Compare like with like instead of treating one favorable answer as a trend.
This method measures a controlled sample. It does not reveal every answer ChatGPT gives, estimate a global mention total, or provide access to other users’ private conversations. OpenAI says even ChatGPT Business workspace members and admins do not automatically see other members’ private chat history. See OpenAI’s Business privacy and sharing guidance.
A useful tracker is a repeatable experiment, not a listening device. It shows how your chosen buyer questions were answered in your recorded tests.
What “tracking ChatGPT mentions” actually means
Traditional media monitoring can search a public stream of posts, pages, or broadcasts. ChatGPT conversations are not a public stream. A brand cannot search all private answers for its name.
Practical ChatGPT tracking therefore uses controlled prompt monitoring:
- Choose a stable set of buyer questions.
- Define the ChatGPT mode, language, market, and test conditions.
- Run the questions on a schedule.
- Save the complete answers and visible sources.
- Classify mentions, citations, recommendations, and competitor appearances.
- Compare the same questions across dates.
- Connect movement to a change log without claiming unproven causation.
Monitoring software generally automates some or all of that workflow by running and storing its own checks. It does not gain access to unrelated users’ chats.
This distinction matters when you report results. “Our brand appeared in 6 of 10 tracked prompts this week” is defensible. “ChatGPT mentioned us 60% of the time everywhere” is not.
The sample should represent decisions buyers genuinely make. Branded prompts such as “Tell me about Acme” can test factual accuracy, but they do not show whether Acme enters an unprompted shortlist. Keep branded accuracy checks separate from unbranded visibility checks.
Track mentions, citations, and recommendations separately
Do not collapse every appearance into one “AI visibility” event. A mention, a citation, and a recommendation answer different questions.
| Outcome | Classification rule | Example | What it tells you |
|---|---|---|---|
| Mention | The answer names the correct brand or product entity | “Options include Acme, Northstar, and Delta.” | The brand entered the answer, but the answer may not endorse or source it |
| Citation | A visible source links to the brand’s owned domain and supports the answer | The answer cites acme.com/pricing for plan facts | ChatGPT used or surfaced an owned source in that answer |
| Recommendation | The answer presents the brand as a suitable choice or shortlist option for the stated need | “Acme is worth considering if you need…” | The answer connects the brand to buyer fit, not merely awareness |
Record these as separate yes/no fields. A brand can be:
- Mentioned but not cited.
- Cited for a fact without being recommended.
- Recommended using third-party sources rather than its own site.
- Mentioned only as an option that does not fit the buyer.
For recommendation, save the reason and any limitation the answer gives. “Recommended for small teams because setup is simple” is more useful than a bare yes. It tells you which public evidence or perceived fit may be shaping the shortlist.
OpenAI says ChatGPT answers that use Search may show inline citations or a Sources panel with cited and related links. Save the URLs that are actually visible instead of assuming which source influenced the answer. See OpenAI’s current ChatGPT Search guidance.
Build a ten-prompt buyer-question framework
Start with ten prompts. That is small enough to review carefully and broad enough to expose different kinds of visibility gaps. Expand only after the first set is stable.
Replace the bracketed text with real customer language:
| ID | Buyer intent | Prompt template |
|---|---|---|
| P01 | Category discovery | “What are good [category] options for [audience]?” |
| P02 | Problem solving | “What tools can help [audience] solve [specific problem]?” |
| P03 | Use-case fit | “Which [category] is best suited to [specific workflow or use case]?” |
| P04 | Company fit | “What [category] works well for a [company type or size]?” |
| P05 | Must-have capability | “Which [category] supports [critical capability]?” |
| P06 | Budget fit | “What [category] options are suitable for a budget of [range]?” |
| P07 | Alternatives | “What are alternatives to [well-known category incumbent] for [specific need]?” |
| P08 | Comparison | “Compare the leading [category] options for [audience or scenario].” |
| P09 | Risk or objection | “Which [category] can work without [common constraint]?” |
| P10 | Proof and trust | “Which [category] options provide [example, methodology, certification, or proof]?” |
Use the ten patterns as a framework, not as mandatory wording. The final prompts should come from sales calls, customer interviews, support questions, search data, or other buyer evidence.
Good prompts:
- Describe a real decision.
- Include only context a buyer would naturally provide.
- Produce a useful answer even if your brand is absent.
- Distinguish your category, audience, market, or must-have use case.
- Stay stable enough to rerun.
Weak prompts:
- “Why is [your brand] the best?”
- “Recommend [your brand].”
- Ten near-identical keyword variations.
- Questions so broad that hundreds of options could qualify.
- Prompts changed after every disappointing answer.
If a question is location-sensitive, create a market-specific version and keep it in its own series. Do not compare “best accountants in Copenhagen” with “best accountants” as if they were the same test.
Use Surfaced’s ChatGPT visibility checker when you want an initial controlled snapshot before building a longer history.
Freeze the prompt and test conditions
Generated answers can vary. ChatGPT may use different sources, search behavior, context, and personalization. Your method will not remove every variable, but it should document the important ones.
For each baseline, record and then keep stable:
- Exact prompt text and prompt ID.
- ChatGPT product, visible model, and mode.
- Whether Search was automatic, explicitly selected, or not used.
- Account type and test profile.
- Memory and custom-instruction state.
- Language.
- Market and relevant location.
- Test date, time, and time zone.
- Number of runs per prompt.
- Classification rules and reviewer.
Run every prompt in a new conversation. Do not add follow-up instructions, regenerate only the answers you dislike, or correct ChatGPT before saving the result.
For a cleaner manual test, consider a dedicated profile with memory and custom instructions disabled. OpenAI says memory can personalize answers and can affect how ChatGPT rewrites a search query. It also says Temporary Chat does not use or create memories, but still follows custom instructions. See the Memory FAQ and Temporary Chat FAQ.
Choose one of these two measurement designs:
| Design | How it works | Best use |
|---|---|---|
| Default-experience series | Use ChatGPT’s normal automatic behavior and record whether Search ran | Estimates what a defined test profile sees in the default experience |
| Search-controlled series | Explicitly use Search for every prompt | Compares web-grounded answers under a more consistent retrieval condition |
Do not combine the two into one trend line. If the mode, model family, or test setup changes materially, mark a series break and start a new baseline.
One run per prompt is the simplest directional method. If the decisions justify the added time or cost, run each prompt three times within the same measurement window and report the proportion of runs where the brand appeared. Save all three results; do not keep only the most favorable one.
Use this copyable tracking table
Create one row for every answer. Copy these columns into Google Sheets, Excel, Airtable, or your database.
| Run ID | Date | Prompt ID | Exact prompt version | ChatGPT mode | Search used | Language | Market | Brand mentioned | Owned-domain citation | Brand recommended | Recommendation reason | Competitors shown | Visible source URLs | Answer archive | Site changes since prior run |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-07-23-P01-R1 | 2026-07-23 | P01 | v1 | Default | Yes | English | Denmark | Yes | No | No | — | Northstar; Delta | example.com/source | Internal capture | Baseline |
| 2026-07-23-P02-R1 | 2026-07-23 | P02 | v1 | Default | No | English | Denmark | No | No | No | — | Delta | None shown | Internal capture | Baseline |
The example rows are fictional. Replace them with your own complete records.
Keep the raw answer, not just the classification. A reviewer should be able to see:
- The answer wording and order.
- Which brand entity was referenced.
- The recommendation or exclusion reason.
- Visible citation and source URLs.
- Competitors and directories included.
- Any factual errors.
- Whether the run failed, refused, or returned no useful answer.
Store captures in an approved internal location. Do not put confidential customer questions, personal data, or sensitive account context into a public spreadsheet or shared conversation link.
Version prompts instead of silently editing them. If P04 changes from “small software company” to “B2B SaaS with 50 employees,” label the new wording v2 and do not treat it as a direct continuation of v1.
Calculate metrics that preserve the denominator
With ten valid prompts and one run per prompt, the core calculations are simple:
- Mention rate: prompts that mention the brand ÷ valid prompts.
- Owned-domain citation rate: prompts with at least one visible citation to your domain ÷ valid prompts.
- Recommendation rate: prompts that recommend or shortlist the brand ÷ valid prompts.
- Competitor inclusion rate: prompts that include each tracked competitor ÷ valid prompts.
- Source coverage: prompts that cite any first-party or credible third-party source about the brand ÷ valid prompts.
- Persistence rate: scheduled measurement periods where a prompt produced the same status ÷ periods measured.
Suppose a ten-prompt baseline produces four mentions, two owned-domain citations, and one recommendation. Report:
- Mention rate: 4/10 (40%)
- Owned-domain citation rate: 2/10 (20%)
- Recommendation rate: 1/10 (10%)
Do not report “40% visibility” without naming the metric and denominator. It hides whether the brand was merely listed, supported by a source, or actually recommended.
If one prompt fails, say 4 mentions across 9 valid prompts, with 1 failed run. Do not quietly count the failed answer as an absence or remove it without reporting coverage.
For repeated runs, report both prompts and observations. For example: “The brand was mentioned in 5 of 10 prompts and 11 of 30 total answer runs.” That shows breadth and volatility.
Useful diagnostic cuts include:
- Status by buyer intent.
- Status by market or language.
- Status by ChatGPT mode.
- Competitors shown for each prompt.
- Owned versus third-party citation domains.
- Reasons given for inclusion or exclusion.
- Prompt-level movement since the previous comparable check.
Misleading metrics include:
- A global ChatGPT mention count inferred from your own tests.
- One “average rank” across different prompts, modes, or markets.
- A composite score with undisclosed weights.
- Combining mentions, citations, and recommendations.
- Comparing a branded prompt with an unbranded prompt.
- Rerunning until the brand appears and recording only that answer.
- Claiming a website change caused movement because it happened first.
An aggregate score can be a useful summary if its inputs and denominator are visible. It should never replace the underlying answers.
Compare dates, modes, markets, and competitors correctly
Every comparison needs a shared unit.
| Comparison | Keep constant | Report separately |
|---|---|---|
| Date over date | Prompt version, mode, language, market, classification | Run date, answer, sources, and status changes |
| ChatGPT modes or models | Prompt, date window, language, market | Each mode or visible model as its own series |
| Markets or languages | Buyer intent and classification rules | Each exact localized prompt and location context |
| Competitors | The same saved answer set | Mention, citation, recommendation, and reason for each brand |
| Different AI engines | Prompt intent, measurement window, and rules | Each provider’s answers and source behavior; never blend them into “ChatGPT” |
For competitor analysis, classify every tracked brand from the same answer. This avoids giving your brand a favorable test while competitors face different prompts.
Do not overvalue list order. Generated answers do not have a stable search-results position system, and the first brand named is not automatically the strongest recommendation. Record order as context if it matters, but prioritize explicit fit, recommendation language, sources, and consistency.
When a competitor appears repeatedly, inspect why:
- Is its category clearer?
- Does it have a page that directly answers the buyer question?
- Does the answer cite current pricing, proof, reviews, or comparison evidence?
- Is the competitor being recommended—or merely named?
- Is your brand absent, excluded, or described inaccurately?
Use those observations to find an evidence gap. Do not copy a competitor page or assume that mentioning the same phrases will reproduce the answer.
Choose manual or automated monitoring
Manual tracking is a valid starting point. It is often better than buying a dashboard before you know which prompts and classifications matter.
| Method | Strengths | Limitations | Best fit |
|---|---|---|---|
| Manual spreadsheet | Low setup effort; close reading; flexible notes; easy to learn | Time-consuming; inconsistent capture; difficult to scale across markets and engines | A first ten-prompt baseline or occasional focused review |
| Automated monitoring | Repeatable scheduling; saved history; broader prompt, market, domain, or engine coverage | Requires cost, QA, transparent classifications, and review of raw answers | Ongoing measurement after the prompt set and business questions are stable |
Before choosing an automated tool, verify that it exposes:
- Exact prompts and prompt versions.
- Engine or provider and run timestamp.
- Full answer text or an inspectable capture.
- Visible citations and source URLs.
- Mention, citation, recommendation, and competitor classifications.
- Market, language, and mode context.
- Failed-run coverage.
- Historical comparison and export.
- A clear explanation of any score.
A dashboard that shows only a score can make measurement easier to look at while making it harder to audit.
Use Surfaced’s AI visibility monitoring page to see the product workflow for recurring buyer-question checks, competitor context, and visibility history. The educational method in this guide remains useful whether you track manually or use software.
Set a cadence and expect volatility
Choose a schedule based on how quickly you can act—not how often you can press regenerate.
- Weekly: useful while you are actively publishing important fixes, launching new evidence, or watching a fast-changing category.
- Monthly: usually enough for a stable website or a team that needs time to implement and evaluate changes.
- Ad hoc: appropriate after a material product, pricing, positioning, crawlability, or source change.
- Daily: often creates more noise than insight unless you have a specific operational reason and enough repetitions to analyze volatility.
Run the full set in one measurement window. Avoid testing three prompts on Monday and seven after a major site change on Friday, then calling the combined result a baseline.
Expect individual answers to move. Retrieval, available sources, product behavior, market context, and generated wording can change. One new recommendation is a lead for investigation, not proof of durable improvement.
Use a simple interpretation rule:
- One changed answer: inspect and save it.
- Repeated prompt-level change: check whether the new status persists across runs or periods.
- Movement across several relevant prompts: examine shared sources, reasons, and recent changes.
- Durable movement plus business outcomes: report the association, while keeping causal claims modest.
Do not erase a baseline when ChatGPT changes modes or a prompt stops working. Preserve the old series, document the break, and start the next one.
Turn the tracking log into the next action
The tracker earns its keep when it changes what you do.
| Pattern | Likely question | Practical next step |
|---|---|---|
| Absent across category and use-case prompts | Is the offer discoverable and unambiguous? | Check crawl access, category language, audience, use case, and the best answer page |
| Mentioned but rarely recommended | Is there enough evidence of fit? | Strengthen proof, limitations, comparisons, pricing, and scenario-specific content |
| Recommended but never cited to the owned domain | Which third-party sources carry the answer? | Verify their accuracy and make first-party evidence easier to find |
| Owned pages cited with incorrect facts | Are public facts stale or inconsistent? | Correct the canonical source and conflicting profiles, then recheck later |
| Competitor repeatedly recommended | What reason does the answer give? | Address the material evidence gap if it reflects a real buyer need |
| Status changes on isolated reruns | Is this signal or normal variation? | Add repetitions or wait for the next scheduled window before acting |
| Improvement across several frozen prompts | Did relevant evidence and sources also improve? | Record the association, preserve the change log, and continue measuring |
If you do not yet know why the brand is missing, use Why ChatGPT Is Not Recommending Your Business to locate the likely discovery, understanding, mention, citation, or recommendation gap. If the gap is clear and you need an implementation sequence, follow How to Get ChatGPT to Recommend Your Brand.
Crawler access can be part of the diagnosis, but it is not a visibility guarantee. OpenAI says OAI-SearchBot is used to surface websites in ChatGPT Search and is controlled independently from GPTBot, which is associated with potential model training. See OpenAI’s crawler documentation.
After a change, rerun the same prompt version under comparable conditions. Record “no movement” when that is the result. Honest negative findings prevent teams from repeating low-value work.
Frequently asked questions
Can I see how often ChatGPT mentions my brand in every user conversation?
No. Your controlled tests can measure only the answers you run and save. Do not present them as access to private conversations or as a global ChatGPT usage total.
How many prompts should I track?
Start with ten high-value buyer questions. A smaller stable set is more useful than hundreds of weak variations. Expand by distinct buyer intent, market, language, or product line only when you can maintain the added series.
Is a brand mention the same as a citation?
No. A mention names the brand. A citation links a visible source. A recommendation connects the brand to the buyer’s stated need. Track all three separately.
Should I force ChatGPT Search on for every test?
Either default behavior or an explicit Search series can be valid. Choose the question you want to answer, record the setup, and do not combine different modes into one trend.
How often should I check?
Weekly can suit active implementation periods; monthly is often more useful for a stable site. Consistency matters more than frequency. Choose a cadence that gives your team time to act and keeps the test conditions comparable.
Does a higher mention rate prove my website change worked?
No. Timing alone does not prove causation. Use frozen prompts, complete answer records, source changes, repeated checks, and a change log to build stronger evidence. Describe the result as an observed association unless you have a design that supports a causal claim.
Does allowing OAI-SearchBot guarantee a mention?
No. It helps make public pages eligible to be surfaced in ChatGPT Search, but OpenAI explicitly says there is no way to guarantee top placement. Access is a prerequisite, not a recommendation promise.
The defensible workflow is simple: freeze a meaningful prompt set, preserve every answer, keep the denominator visible, separate mentions from citations and recommendations, and act only on patterns that survive comparable rechecks.
Related Surfaced guides and tools
Check your own evidence
Find the first visibility gap worth fixing.
The free audit checks public website evidence and shows a prioritized first fix. Results are diagnostic, not a promise of future AI recommendations.