What is deep research poisoning?
Deep research poisoning is an attack in which someone edits a public, user-generated web page that AI research agents habitually retrieve, so the agent's finished report repeats the attacker's claim. Cornell Tech researchers named the technique WARP, short for Web Agent Retrieval Poisoning. No access to the AI system is required.
The important detail is what the attacker does not need. No model access. No API key. No knowledge of your exact wording. No ability to publish a new site or buy a domain.
They only need to append a short passage to a page that already ranks and that agents already fetch. A Reddit comment is enough.
That is a different failure mode from the one you have probably been trained to look for. A hallucinated citation points at a URL that does not exist. A poisoned citation points at a real URL, on a real high-authority forum, with a real thread full of real humans. Click it and everything looks fine.
How did 13 words steer an AI research agent?
The researchers ran reconnaissance to find which user-generated pages an agent consistently fetches for a topic, appended roughly 13 words of chosen text to one of those pages, then re-ran a cluster of related queries. The agents' reports drifted toward the planted claim at conditional mention rates of 38% to 51%.
Here are the load-bearing facts, so you can check them yourself rather than take this article's word for it.
The research, in extractable form
--- Paper: Deep-Research Agents Can Be Poisoned via User-Generated Content, arXiv 2605.24245, posted 22 May 2026
--- Authors: Tingwei Zhang, Harold Triedman and Vitaly Shmatikov, Cornell Tech
--- Attack name: WARP, Web Agent Retrieval Poisoning
--- Open agents tested: STORM, Co-STORM, OmniThink
--- Poison payload: roughly 13 words appended to one existing user-generated page
--- Reported effect: conditional mention rates of 38% to 51% across a cluster of related queries
--- Attacker requirements: ability to post a comment on a page agents already fetch, nothing more
Note the phrase cluster of related queries. The attacker does not need to guess your prompt. They poison the topic, and any reasonable phrasing of a question about that topic routes through the same handful of pages.
Independent coverage followed in June 2026 from Search Engine Land and 404 Media, both framing the practical risk as commercial rather than academic: steering buyers toward a product or away from a competitor.
Why is "just check the citations" no longer enough?
Checking citations catches fabricated sources. It does not catch a real source whose content was written to be retrieved. If you open the link, see a live Reddit thread and tick it off as verified, you have confirmed that the page exists, not that the claim is supported by anyone credible.
Most practitioners run one of two verification habits, and both fail here.
The first habit is spot-checking two or three citations out of twenty. Poisoning is a probabilistic attack. If the planted claim shows up in roughly half of related reports, the odds that your sample lands on the one paragraph that matters are poor.
The second habit is trusting the agent's own confidence language. Deep research outputs read like consultancy decks. Confident structure is a formatting choice, not evidence, and a poisoned claim inherits exactly the same authoritative tone as a well-sourced one.
There is a practical reason this matters more in 2026 than it did a year ago. Deep research modes are now the default answer surface for a lot of work, not a novelty tab. Reports get pasted into decks, quoted in client emails and used to justify spend, often without the person downstream ever seeing which sources were involved.
Once a claim has been reformatted into a slide, it stops looking like a retrieval result and starts looking like a fact your team asserted. That is the moment a planted sentence becomes expensive.
There is a broader reliability backdrop here too. An audit of 111 million references across roughly 2.5 million papers reported tens of thousands of hallucinated, non-existent citations in 2025 papers alone. Fabrication and poisoning are separate problems, and a verification routine has to catch both.
How do you verify a deep research report in 10 minutes?
Sort claims by consequence, then verify only the load-bearing ones against source type rather than source existence. Four steps: mark the decision-driving claims, classify each cited source as primary or user-generated, re-run the same question with user-generated content excluded, and compare the two outputs for claims that vanish.
Step 1: Mark the claims that change your decision
A twenty-page report usually contains three or four claims that actually change what you do next. A price, a limit, a named vendor recommendation, a regulatory deadline. Highlight those and ignore the rest for now.
Step 2: Classify each source behind those claims
Split them into two buckets. Primary sources are vendor documentation, official pricing pages, filings, changelogs, published papers. User-generated sources are Reddit, Quora, forums, comment sections, community wikis, aggregator listicles.
A decision-driving claim resting only on user-generated sources is not verified. It is a rumour with a footnote.
Step 3: Re-run the question with user-generated content excluded
This is the step almost nobody does, and it is the one that exposes poisoning. Ask for the same research again, restricted to primary sources. Any claim that survives both runs is probably real. Any claim that only appears in the unrestricted run is where you look harder.
Step 4: Check the vendor's own page for anything commercial
For prices, limits, trial terms and feature availability, the vendor's live page is the only source that counts. Thirty seconds of checking beats a paragraph of confident summary every time.
What prompt forces your AI to audit its own report?
Paste the report back into the model and ask it to classify its own evidence rather than defend its conclusions. The prompt below does that. It is written to produce a table you can scan in under a minute, and it deliberately forbids the model from re-arguing the original answer.
Try this prompt
--- You are auditing a research report for retrieval poisoning risk, not rewriting it. Do not restate or defend the conclusions.
--- Below is a research report I generated. For every claim that would change a business decision, produce one table row with these columns: CLAIM, SOURCE TYPE (primary documentation / published research / news / user-generated), SOURCE URL, and RISK (high / medium / low).
--- Mark RISK as high whenever a decision-driving claim is supported only by user-generated content such as forums, Reddit, comment threads, community wikis or aggregator listicles.
--- After the table, list separately: (1) every claim you cannot trace to a specific URL, and (2) every commercial figure such as price, usage limit or trial term that I should confirm on the vendor's own page.
--- Finally, in one sentence, name the single claim that would do the most damage if it turned out to be planted.
--- Here is the report: [PASTE REPORT]
The last instruction is the useful one. It forces a ranking rather than a wall of caveats, and it tells you exactly where to spend your remaining five minutes.
If you want the reasoning behind why constraints like these change output quality so reliably, that is the same mechanism covered in our piece on context engineering.
Where does this verification workflow break down?
It breaks in four predictable places, and knowing them is what separates a workflow you keep from one you abandon after a week.
The model cannot always see its own retrieval trail. Some tools show every fetched URL, others show a curated subset. If the source list looks suspiciously tidy for a twenty-page report, treat the audit as partial rather than complete.
Excluding user-generated content costs you real signal. Forums are genuinely the best source for how a tool behaves after three months of daily use. The point is not to ban them, it is to stop them from silently carrying a decision on their own.
Self-audits are not adversarial. A model grading its own output has no incentive to look hard. Running the audit in a different model than the one that wrote the report is a small change that meaningfully improves the catch rate.
The attack surface is not only Reddit. Any user-editable, well-ranked page qualifies. Community wikis, Q&A sites, review sections, GitHub discussions and comment threads on high-authority publications all sit in the same category.
What should you test in the next 20 minutes?
Pick the last deep research report you actually acted on, run the audit prompt above on it, and count how many of your decision-driving claims trace back only to user-generated pages. Most practitioners find at least one. That single number tells you how much your current workflow was relying on trust.
Then run the more uncomfortable version of the test. Ask any AI assistant to recommend a vendor in your own category, and read which sources it leans on. If a forum thread you have never seen is doing the deciding, that thread is now part of your competitive landscape.
None of this means deep research is not worth using. It compresses a week of desk research into twenty minutes, and that trade is still overwhelmingly worth making. It means the twenty minutes you saved should not all be spent celebrating.
We know AI's cold edges. We know your real challenges. 28 years with UD, turning technology into a partnership with warmth.
Reviewed by the UD AI team.
🔍 See What AI Actually Says About You
If AI answers can be steered by a stranger's forum comment, it is worth knowing what AI currently says about your own brand and content. UD's AEO Auditor checks how answer engines read, cite and recommend your pages. And when you are ready to build a research workflow your team can trust every time, we'll walk you through every step, from source policy to prompt templates to sign-off.