When ChatGPT cites a page, it did not simply “search Google”. It fetched that page through one of several retrieval pipelines, and it records which one in the conversation data. Sprout SEO Extension reads that label and shows you the split.
This matters more than it sounds. Which pipeline an answer leans on tells you whether being crawlable, being licensed, or being a strong SERP performer is the thing that would actually change your visibility.
Purpose
- See which retrieval path ChatGPT used to fetch each source, and what that means for your SEO.
Where to Find It
- Per conversation: Popup nav →
AI Insights→ the Sourcing Pipelines section, shown as a doughnut chart. Hover any segment for an explanation. - Across all captures: the Sourcing pipelines card in the Bulk Analysis dashboard.
- Individual sources also carry a pipeline badge.
Pipeline data comes from ChatGPT. Gemini and Google AI Overviews do not expose an equivalent.
The Four Pipelines
| Pipeline | Shown as | What it is |
|---|---|---|
bright |
Bright Data | A commercial web scraper. Dominant for shopping, finance, weather and local results. |
oxylabs |
Oxylabs | A commercial web scraper, Bright Data’s rival. Skews toward regional and local press. |
labrador |
Labrador · licensed | An allowlist of licensed and established publishers: Reuters, WSJ, Wikipedia, arXiv and similar. Near-full-article extracts. |
serp |
Open web · SERP | The open-web search baseline. Mostly seen on news results. |
How to Read the Split
Heavy on Bright Data or Oxylabs. The answer was assembled by scraping live pages. Your page needs to be fetchable and parseable, and its content needs to answer the question in the page body. This is the closest thing to classic technical SEO mattering.
Heavy on Labrador. The answer leaned on licensed publishers. You are not going to out-optimise your way into that set. The lever is being covered by those publishers, or being the source they cite. Digital PR, not on-page work.
Heavy on Open web / SERP. Search ranking is feeding the answer fairly directly, so conventional rankings for the fan-out queries are the lever.
Mixed. Most answers are. Read the proportions and weight your effort accordingly.
Why This Beats Guessing
The common failure in AI-search work is to assume every answer is a rankings problem, spend three months on content and links, and see nothing change because the model was pulling from licensed publishers the whole time. The pipeline split turns that from an assumption into something you can check in a minute.
Recommended Plays
- Diagnose before you plan. Run a prompt set for your category, read the pipeline mix, and pick the strategy that matches it.
- Compare by intent. Buying-intent prompts and research prompts often source very differently. Run both sets and compare.
- Watch it over time. The mix shifts as models change their retrieval. A saved prompt list re-run monthly will show it.
Tips & Troubleshooting
- No pipeline data? The conversation had no web search in it, or you are looking at a Gemini or Google AI Overview capture. Only ChatGPT exposes this.
- An unfamiliar pipeline name? Unknown providers are shown with their raw label rather than being hidden, so new ones surface as they appear.
- These are labels ChatGPT publishes about its own retrieval. They describe how a source was fetched, not how much weight the model gave it.