This website uses cookies

Read our Privacy policy and Terms of use for more information.

TL;DR: A generated research report is right only if every link in it holds, so accuracy falls away as the steps pile up. In one benchmark's public filings, the same models score 83% on answering what a document says and 35% on answering what follows from it. Verifying the recent and the obscure misses it; decisions rest on public material. The mode that most resembles diligence fabricates the most citations. Mark every claim retrieved or composed, and re-source only the composed ones before money moves.

Want to listen to this article?

Subscribers to the Orion Playbook newsletter can listen to the AI-Generated Audio version of this article for free. Already a subscriber? Log in here

I bought industry analyst reports more than once, for a few thousand dollars each. The sellers were small firms working off public filings and press releases, guesses filling the gaps. The general picture was directionally right. The parts about my own company were wrong, and several real players were missing. We were PE-backed, so no numbers went to the street. Ours had been reconstructed from whatever those releases implied. Some overshot. Some undershot. Both in the same report. A report wrong in two directions at once can only be corrected by someone who already knows the answer.

Run an LLM-based deep research report today and you get that artifact again, free and better written. The skepticism those paid reports earned did not survive the price falling to zero. That is the problem, and careful people are guarding against a different one.

The advice already circulating is sophisticated enough to sound sufficient. Verify the recent, the proprietary, and the niche: anything past the cutoff, anything never in the corpus, anything too obscure to have been seen twice. The rule is tractable. It feels rigorous, and most operators have adopted some version of it. It is the right instinct aimed at the wrong variable.

Vals AI runs a benchmark called Finance Agent v2 against public company filings. In its August 13 results, the same frontier models on the same harness scored 83% on general quantitative extraction and 35% on financial modeling. The extraction questions ask what a document says. The modeling questions ask what follows from it. Nothing in the second set is newer, more private, or more obscure than the first. The only difference is how many steps have to hold.

Correctness in a research task is conjunctive. The answer is right only if every link holds. A directional claim is one or two links. A decision-grade claim is six, and all six have to land. The older explanation still holds: the corpus is thinnest exactly at the particular, and it is true as far as it goes. Yet the failure that bites an operator needs no obscurity at all. The material on which a business decision rests is usually public, and it is almost always composed.

The principle reappears on the citation side. In a University of Pennsylvania study of commercial models and deep research agents, the agents fabricated 10.7% of their citation URLs, compared with 4.8% for plain search-augmented models, because multi-step retrieval compounds errors rather than filtering them out. The mode that most resembles diligence invents the most. As the study's authors put it, "users treat these citations as evidence." That failure is largely correctable, and the industry is closing it, which is what makes the rest worse. A fabricated link is the one error that announces itself, and removing it takes the last visible cue with it. The paragraph about your own company reads exactly as well as the paragraph about the industry.

The benchmark flags certain facts as required, and getting any one of them wrong scores the whole answer zero, however much of the rest is right. Under that rule, the best available model sits at roughly half. That is how a board consumes an investment memo. There is no partial credit at the point of decision, and almost nobody reads the generated research under the standard by which they will be judged. They read it on partial credit, then act on it as though it had passed.

So mark the report before anyone acts on it. One reading pass, two marks. Every claim is retrieved, lifted from a source, or composed, assembled across steps. Only the composed claims get independently re-sourced before money or headcount moves. The tells: a therefore, a ratio, an adjustment, a comparison drawn across two documents. The cost is one read, which is what keeps the team out of pre-AI research economics. The analyst sizing a market and the board underwriting a thesis run the same pass. Better still, ask for the markers up front. The prompt below does that.

Nobody has separated the two explanations. No study has isolated the links that needed a rare fact from the links where every fact was right and only the assembly failed. Composition may simply be what thin data looks like from a distance, which would make this piece half right. The study that settles it does not exist yet. What holds either way is the tension underneath: speed of orientation and confidence in specifics are both real goods, and a rule that kills the first to protect the second costs more than it saves.

Last year I researched a property purchase in Europe, and the report had a sound grasp of the general rules. I never treated it as the source of truth. I read it on the flight over, and it bought me an educated conversation with the agent, the bankers, and everyone else in the room. It told me what to check. It did not tell me what was true. That is what a map is for, and it is why nobody builds on one.

All the images were generated with AI (ChatGPT Images, Gemini Nano Banana, Claude Opus) by Gérard Métrailler.

Food for your AI

Paste this into the instructions of your next deep research run. It moves the marking upstream, from your reading to the model's writing.

Tag every figure [DISCLOSED] if it comes from the source material, [DERIVED] if you calculated it from that material, or [ESTIMATED] if you inferred it, and show your reasoning. Never blur the three.

A gap you name is worth more than a number you invented, so convert every gap into a question I should ask.

Cite primary sources: the company's own filings, releases, and pricing pages. Never rest a figure on an aggregator or a forum. Where only a secondary is reachable, say so and tag it [ESTIMATED].

Sources

Vals AI. Finance Agent v2, benchmark results and methodology, version 2, updated 2026-08-13. https://www.vals.ai/benchmarks/fabv2. Accessed 2026-08-15.

Rao, Delip, Eric Wong, and Chris Callison-Burch. Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents. arXiv preprint, April 2026. https://arxiv.org/abs/2604.03173. Accessed 2026-08-15.

Badhe, Sanket, Deep Shah, and Nehal Kathrotia. Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications. arXiv:2602.16201, submitted 2026-02-18. https://arxiv.org/abs/2602.16201. Accessed 2026-08-15.

Reply

Avatar

or to participate