AI Visibility

How AI Search Really Works (The Black Box, Opened)

We modeled the publicly observable AI answer workflow across Claude, ChatGPT, Gemini, and Perplexity as one four-stage workflow. The real systems are proprietary and differ by product. Here is what likely happens when an AI decides whether to name your brand.

By , NYFTY Labs AI Content Engine

When you ask an AI assistant a question, it rarely runs one search and reads one page. Modern answer engines run a pipeline: they decide whether to search at all, break your question into many smaller queries, fetch and re-rank candidate pages, pack the best passages into the model's context, then write an answer and attach citations. Understanding each stage explains the single most important fact for any brand: being named in the answer and being cited as a source are two separate outcomes, and you can win one without the other. This is the documented machinery behind the "black box," with the inferred industry practice flagged as such.

How AI Search Really Works From a user query to a cited, named answer in seven steps 1. Query User asks 2. Rewrite Fan-out terms 3. Retrieve Search index 4. Rank Score sources 5. Read Extract facts 6. Synthesize Draft answer 7. Answer Output result Two separate gates NAMED brand mentioned CITED linked as source Being named is not the same as being cited, chase both gates to fully show up in AI answers.
A seven-step flow of AI search, from query through rewrite, retrieval, ranking, reading, synthesis, and answer, ending at two distinct visibility gates: NAMED (your brand mentioned) and CITED (linked as a source).

Stage one: the engine decides whether to search

Before any retrieval happens, the model makes a routing decision: answer from its training data, or go fetch live information. In practice, engines lean toward searching for time-sensitive, comparative, or factual queries (prices, news, reviews) and lean toward internal knowledge for stable, timeless questions. This decision is situational and varies by platform: Perplexity grounds nearly every response in live sources, while ChatGPT searches more selectively. The takeaway is that if the engine never searches, your content cannot be cited no matter how good it is.

Stage two: fan-out turns one question into many

Rather than searching your exact words, the engine decomposes the prompt into multiple related sub-queries and runs them in parallel, a technique Google publicly calls "query fan-out" in its AI Mode announcement, powered by a custom version of Gemini (Gemini 2.5 at the May 2025 AI Mode launch), a model line that keeps moving through versions. Different sub-queries can be routed to different sources: the open web, a knowledge graph, shopping or local data, and structured feeds. The practical consequence is that you are not competing for one keyword anymore. You are competing across a spread of sub-questions the user never typed.

AI VISIBILITY 1Your oneprompt2Sub-queries3Open web4Knowledgegraph5Shopping /localYou compete across questions the user never typed NYFTYLABS
One prompt fans out into parallel sub-queries routed to different sources.

Stage three: fetch, re-rank, and pack the context

Each sub-query returns a set of candidate pages. Most engines are believed to re-rank those candidates by relevance and quality and keep only a small subset, discarding most of what they retrieved. The surviving passages, not whole pages, are then packed into the model's context window as the evidence it will read. Industry analyses consistently find that the large majority of retrieved pages are never used, so simply being retrievable is necessary but far from sufficient.

AI VISIBILITY 1Candidatepages2Re-rank3Keep smallsubset4Pack passages5ContextwindowMost retrieved pages are never used NYFTYLABS
Retrieved candidates are re-ranked and trimmed to a few passages packed into context.

Stage four: generate the answer, then attach citations

The model writes the answer from the packed passages, and citation can happen one of two ways. Some systems generate the answer and its sources together so the text is tied to evidence as it is written; others generate the answer first and attach supporting links afterward, an approach documented in research on systems like RARR. The second method, sometimes called post-hoc attribution, can produce a citation that supports a claim without being the true origin of the wording. This is why a cited link does not always mean that page is where the assistant "learned" the fact.

The key insight: named is not the same as cited

Two distinct things can happen to your brand in an AI answer. You can be mentioned, where the model names you in the recommendation itself, or you can be cited, where your URL appears as a linked source. These are decided at different stages by different signals, so they do not move together. Industry analyses of AI visibility suggest relatively few brands consistently earn both, meaning your research can inform an answer that then recommends a competitor by name.

AI VISIBILITYMentionedNamed in the answerPart of recommendationDecided at generationCitedURL as linked sourceAppears as evidenceDecided at retrievalvsNYFTYLABS
Being named and being cited are separate outcomes set by different signals.
Key takeaways
  • AI answers come from a pipeline, not a single search: decide-to-search, fan-out into sub-queries, fetch, re-rank, pack, generate, then attribute. Optimize for the whole chain, not one keyword.
  • Fan-out means you are judged across many sub-questions the user never typed. Google documents this as "query fan-out" in AI Mode, run on a custom Gemini model.
  • Most retrieved pages are never used. Being findable gets you into the candidate pool; surviving the re-rank and making it into the packed context is the harder bar.
  • Citations are not always proof of origin. Some engines attach sources after writing the answer, so a cited link may support a claim without being where the wording came from.
  • Being NAMED in the answer and being CITED as a source are separate outcomes governed by different signals. Track both, because you can win one and lose the other, and the named brand usually captures the buyer.

← More guides

The memory graph

How AI handles your data.

One simplified pipeline view, four engines, each with its own dial settings. Click any node to see how that engine retrieves, ranks, combines, and finally names (or skips) a brand. The two final tracks, named and cited, are different gates you win separately.

FAQ

Questions, answered.

They are two separate gates. CITED means a page from your site survived retrieval and got attached as a source. NAMED means your brand token actually appears in the written answer. In some systems the answer is generated first and citations are mapped on afterward; in others, citations are generated alongside the answer. Either way, you can be cited but not named, or named but not cited. Winning AI visibility means chasing both.

Not the way they did. AI engines turn your content into meaning-vectors and match them, chunk by chunk, against the meaning of the query (some engines also generate a hypothetical ideal answer to compare against, a technique called HyDE), leaning far more on meaning than on the user's exact words (most engines still blend in keyword matching, so exact names and identifiers remain one signal). You win by covering the concepts a perfect answer would contain in short, self-contained, front-loaded chunks, not by repeating the query terms.

Cover the ideal answer's concepts so your chunk survives retrieval and reranking, make sure your key fact lands in the first result set, and seed the same canonical fact across multiple independent sources. When several independent sources agree, the model is more likely to converge on your brand's name.

Want this working for your brand?
Definition

What is The AI answer workflow?

The AI answer workflow is the sequence of steps an AI engine like Claude, ChatGPT, Gemini, or Perplexity runs to turn a question into a written, sometimes-cited answer. NYFTY Labs models the publicly observable behavior of these proprietary systems as one four-stage framework: Decide, Fan-Out, Fetch, re-rank, Pack, Vote, and Attribute.

How it works

The engine first decides whether to search or answer from memory, rewrites the question into sub-queries, fetches the top handful of results, re-ranks them by meaning, packs the survivors into one buffer, generates (votes on) the answer word by word, and only then stamps citations onto the finished sentences. Being cited (a page survived retrieval) and being named (your brand token appears in the prose) are two separate gates, and in some systems the citations are reverse-engineered onto an answer that was written first.

Who it’s for

For marketers, founders, and content teams trying to understand why an AI does or does not mention their brand. The benefit is better decisions: instead of guessing, you know which stage you are failing at (retrieval, reranking, or naming) so you can fix the right thing rather than chasing keyword tactics that no longer apply.

In practice

A software company keeps losing to competitors in ChatGPT answers. Mapping it to the seven stages reveals their page survives retrieval and gets cited as a source, but their brand name never lands in the written sentence, so they seed the same canonical fact across several independent sites to sharpen the odds the model names them, not just links them.