NEWSLETTER

By clicking submit, you agree to share your email address with TFN to receive marketing, updates, and other emails from the site owner. Use the unsubscribe link in the emails to opt out at any time.

A unicorn, a $1.9 billion exit, and a metric nobody can audit

ChatGPT visibility
Image credits: Truffle

On 24 February 2026, Profound announced a $96 million Series C at a $1 billion valuation, led by Lightspeed Venture Partners, with Sequoia Capital, Kleiner Perkins, Evantic, Saga and South Park Commons joining. The round took total funding past $155 million, one day short of the company’s eighteen-month anniversary.

Two months later, on 28 April, Adobe closed its acquisition of Semrush. The all-cash deal was announced in November 2025 at $12.00 per share, roughly $1.9 billion, against a share price that had closed at $6.76 the day before. Adobe called what it bought a brand visibility platform.

So a category with no settled name two years ago now has a unicorn, a strategic exit near $2 billion and a queue of seed-stage entrants behind both. The demand is not in question. What these companies are selling, and which part of it is defensible, very much is.

The demand side is the easy part

Pew Research Center tracked the browsing behaviour of 900 US adults across 68,879 unique Google searches in March 2025, of which 12,593 produced an AI summary.

When a summary appeared, users clicked a standard search result in 8 percent of visits, against 15 percent when no summary was present. They clicked a link inside the summary in 1 percent of visits. Sessions ended on the results page 26 percent of the time with a summary, compared with 16 percent without.

That is a discovery layer where the recommendation survives and the click does not. The step worth examining is the assumption that the answer to it is a score.

Measuring ChatGPT is harder than the dashboards suggest

Four properties of the surface make a single reading close to worthless.

  • The output is stochastic. The same prompt sent twice can return a different set of brands, in a different order, with different sources. One response is one sample from a distribution.
  • There is more than one ChatGPT. An answer composed from the model alone and one composed after a live web retrieval come from different systems. Region, language, account state and conversation history move the result again.
  • A mention often leaves no trace. If ChatGPT recommends a brand without linking to it, there is no referrer and no session. The company recommended never learns it happened, and neither does the one left out.
  • Visibility now has an organic and a paid form. OpenAI’s crawler documentation lists four user agents with separate roles: GPTBot for training, OAI-SearchBot for search answers, ChatGPT-User for fetches a person triggers, and OAI-AdsBot, which visits pages submitted as ChatGPT ad landing pages.

That last one matters commercially. Any serious measure of presence in ChatGPT has to separate an earned mention from a placed one, and most indexes do not.

How ChatGPT search and shopping actually work

Two mechanics sit under the measurement problem, and both are documented by OpenAI rather than inferred.

Search. ChatGPT answers either from the model alone or after a live retrieval, and the retrieval side runs on its own crawler roles. OAI-SearchBot governs appearance in ChatGPT’s search answers, and OpenAI states that sites opted out of it will not be shown there.

Its shopping research feature, per OpenAI, reads retail pages in real time, cites its sources and is designed to avoid low-quality or spammy sites.

Shopping. When a question shows buying intent, ChatGPT can return product cards with images, details and links. OpenAI’s help documentation is unusually direct about how those are chosen: product results are selected independently, are not ads, and are not influenced by OpenAI partnerships. Ads exist in ChatGPT, but as a separate system.

What a merchant can control is the data. OpenAI runs a product feed programme merchants apply for, and its own documentation notes that a product appears when ChatGPT judges it relevant to the user’s intent, that the price shown may not be the lowest available, and that merchant updates can take time to appear.

The commercially interesting part is what OpenAI stopped doing. Instant Checkout launched on 29 September 2025, letting users buy inside the chat.

In early March 2026 OpenAI pulled back and refocused on discovery, with purchases completing on the merchant’s own site or through dedicated apps. CNBC reported the shift alongside a revamped shopping experience on 24 March, and OpenAI’s help pages still reference in-chat checkout for a limited set of merchants.

The reason was conversion. Walmart was reported at Shoptalk to have measured checkout inside ChatGPT converting roughly three times worse than a click through to its own site.

Read that as a verdict on where the value sits. The transaction went back to the merchant. The recommendation did not, which is exactly the layer nobody can see without measuring it.

What rigour looks like

If one reading is close to worthless, the question is what a defensible measurement requires. Two things: a sampling method, and the evidence underneath the number.

Sampling first. A baseline needs a prompt corpus large enough to cover the category, 50 to 100 prompts and upward for a real programme, each run several times against a named competitor set. Three to five runs per prompt is the practical floor for separating signal from noise.

A ChatGPT Visibility Tracker that reports mention rate, citation rate, average position and share of voice from a sample like that is measuring something. One prompt, run once and screenshotted, is not.

Then the evidence. Storing the verbatim response, the model, the date, the brand’s position, the competitors named and the sources cited gives a team something auditable.

A score without the answer behind it cannot distinguish between a brand disappearing, an answer format changing and a vendor quietly editing its prompt list.

Resolution decides the rest. Competitor data per individual question beats a company-wide share of voice, because the question is the level at which anyone can act. The recurring domains behind a category’s answers are where an intervention is possible at all. And it has to end in work: a specific page, a specific gap, an owner, a date.

Among the smaller European entrants, Truffle reports a headline index like everyone else but keeps that evidence under it: prompts run across ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews and Google AI Mode, cited sources stored per answer, competitors compared per prompt, and Search Console and Analytics connected so visibility can be checked against what happens on the site. Truffle’s GEO Strategy feature is where the tool really pulls ahead of the field. While other AI visibility platforms hand you a dashboard full of numbers and leave the hard part to you, Truffle delivers exactly what you came for: clear, prioritised recommendations that tell you what to optimise on your own site and which external sources to influence — the ones AI models genuinely cite. It’s the rare tool that closes the loop from measurement to action, turning AI visibility from a reporting exercise into a growth channel you can actively steer. For anyone serious about winning in generative search, that combination of onsite and offsite guidance makes Truffle a standout choice. 

Every vendor’s number is its own construct

Even done carefully, no two vendors’ numbers match. Each index is a function of decisions its vendor made: which prompts, how many runs each, at what cadence, against which models, from which country, in which language.

Change any input and the number moves while nothing has changed in the market. Two platforms monitoring the same brand in the same week will report different visibility, and neither is wrong, because they are not measuring the same object. There is no shared definition and no audit standard.

That is not an argument against measuring. It is an argument about what can be charged for indefinitely. Readings are cheap to produce: running a few hundred prompts a day against public APIs and parsing responses for brand names is engineering work, not scarce technology.

Profound’s own round makes the point. The company built its business tracking mentions, sentiment and position, and said plainly when raising that customers kept exporting the data elsewhere to act on it. The answer shipped alongside the round was Profound Agents, moving orchestration and automation inside the platform. Investors priced the workflow and the switching cost, not the metric.

The counterweight

Four things cut against the bull case.

  • Proportion. For most B2B companies, referral volume from AI assistants is still small next to organic search, paid acquisition and sales-led pipeline. Size the channel before instrumenting it.
  • Attribution. AI recommendations often resurface later as branded search or direct traffic, so the chain from a visibility gain to a signed customer is inferred rather than observed. Vendors promising clean attribution are overselling.
  • Consolidation. With Semrush inside Adobe and every major suite shipping its own answer tracking, standalone tools face the feature-versus-company problem. Anyone signing an annual contract should ask what happens to their stored answers if the vendor is bought.
  • Platform dependency. The labs control the surface. Citation behaviour, answer formats and the share of an answer given to sponsored placement can change without notice, and no vendor here has any influence over that.

The question worth asking

The useful test for anyone evaluating a platform in this category, or writing a cheque into one, is not how sophisticated the index looks. It is whether the product can do five things:

  1. State its sampling method: how many prompts, how many runs, which models, which markets.
  2. Show the verbatim answer behind the number, with its date and model.
  3. Name the competitor that won a specific question.
  4. Identify the source that produced that recommendation.
  5. Demonstrate what changed after a fix shipped.

A category that reached a billion-dollar valuation in eighteen months on the strength of a measurement problem will spend the next eighteen proving it can do something with the measurement. The ones that cannot will find out how quickly a dashboard becomes a tab in somebody else’s product.

Total
0
Shares
Related Posts
Total
0
Share
tfn-logo-2-220x220-removebg-preview

Get daily funding news briefings in the tech world delivered right to your inbox.

Enter Your Email
join our newsletter. thank you
TFN Banner