Guide
Before I file anything AI-researched, I run the research twice
Two runs of one deep-research task, locked to the same sources, landed about 3% apart on a number headed for a sworn filing. The gap between the reports told me which claims were facts and which were the model's judgment.
Key findings
- Same task, same source restrictions, run twice: both runs chose the same top comparable sales and still landed about 3% apart on value.
- The diff showed why: the runs pulled contradictory readings of the same public listing feeds (floor level, view quality, parking).
- When AI research feeds a sworn filing, a second independent run is the cheapest instrument that separates sourced facts from model inference.
Before AI research goes into anything I sign, I run the identical task a second time and diff the reports. Where they agree is evidence. Where they differ is inference.
The case that taught me this was a decline-in-value property-tax review for a condo — asking the county to temporarily lower an assessed value to what the market actually says it's worth. The number that research produces goes on a form signed under penalty of perjury. A single deep-research report (the AI mode that browses sources on its own and returns a cited report) reads as authoritative. It is an opinion with citations attached, and you cannot see which sentences are sourced and which are the model's taste. Run the identical task twice and diff the reports (put them side by side, mark every claim that changed) and suddenly you can.
How far apart do two identical runs land?
About 3% apart, in my case. Same instructions, same source restrictions, same evidence on top.
Both runs were locked to three source classes: the county's own forms, the state tax board's guidance, and public closed-sale records. Nothing else. Both independently ranked the same two nearest-in-time, same-complex sales as the strongest comparables (comps: recent sales of similar units, the evidence a value opinion stands on). If you skimmed either report alone, you would sign off. They look equally rigorous. They cite the same records.
The two value opinions still came out about 3% apart. On a filing where the number is the whole argument, that gap is not rounding noise. It is a measurement of how much of the report was judgment.
Why did the runs disagree about the same buildings?
Because public listing feeds contradict themselves, and a model silently picks a side.
The diff made it concrete. One run described the strongest comparable by its standout view quality. The other run read the same unit's structured listing fields and flagged it as ground-floor. One run treated parking as a wash across the whole comp grid; the other pulled the subject's public record, read it as having no garage, and applied a downward offset. Same units. Same feeds. Opposite readings.
Neither report marked its own reading as contested. That's my sharpest criticism of deep-research tools as shipped: each report was fluent, cited, and quietly opinionated. Nothing in either one separated "the record says this" from "I resolved a conflict in the record without telling you."
One catch surfaced only once. A single run flagged an extreme low-priced outlier sale as non-representative because the same agent sat on both sides of the transaction. The other run never mentioned that sale at all. A flag only one run raises is still a flag. Details here are generalized to protect the people involved; the mechanics are exactly as recorded.
What does the diff actually separate?
Three piles, each with its own action.
Where the runs agree and cite the same class of source: treat it as a source-backed fact, then spot-check the citation. The shared top comparables lived here.
Where the runs conflict (floor level, view quality, parking): everything in this pile is model inference wearing a fact costume. Each conflicted attribute goes to hand-verification against primary records (the parcel record, the owner's own documents) before it can touch a filing.
Flags only one run raised: keep all of them. Warnings get combined, never filtered down to the overlap.
This is the same instrument as running two trackers against one truth: two independently produced views of the same thing, where disagreement is the alarm. And it's the cheapest tool in my whole verification setup. The second run costs one prompt and a wait, and it produces the one thing a single run cannot: a visible boundary between sourced and inferred.
When is a second run worth it?
Whenever the output feeds a signature, a payment, or a deadline. I don't do this for brainstorming. I do it for filings.
Two honest limits. Both runs came from the same model family, so they share blind spots; agreement only means the runs failed to disagree, which is not the same as being proven right. A correlated error sails through the diff untouched. And the diff catches inconsistent inference, not systematic bias — if the model always over-weights view quality, both runs drift together and the diff stays quiet. When the stakes justify heavier machinery, I reach for a much larger review swarm instead.
But for the price of one extra prompt, I got a map of exactly which attributes in a sworn filing rested on a model's guess. The disagreement itself was the finding.
What's the last AI-researched number you acted on — and if you had run the task twice, which half of the report would have survived the diff?
Method & data
Method: two independent AI deep-research runs on a real decline-in-value property-tax review for a condo, diffed claim by claim, plus the project's status fileData: patterns and methods only; no names, dollar figures, addresses, case identifiers, or confidential content · Last checked: 2026-08-14
How this was made
AI-drafted, adversarially checked, human-directed. My AI assistant wrote this from the system's own records — the two dated deep-research reports from the condo tax review and the project's status file. A separate AI session then tried to break every claim against those records, and automated privacy and readability gates ran before publish. I direct this pipeline, own every boundary in it, and audit published pages on a rolling basis — if you find an error, tell me and it goes in the corrections log, dated, never silent.
I'm Ali — I run real life-and-work admin on AI agents, then check their work in the open. More at /about.
Published under my standards. Found an error? Tell me — corrections go in the corrections log, dated, never silent.
Cite this
@online{ali2026runtwice,
author = {Ali},
title = {Before I file anything AI-researched, I run the research twice},
date = {2026-08-14},
url = {https://alidoes.ai/run-the-ai-research-twice/}
}Caught something I got wrong? Send it directly. Confirmed corrections go in the corrections log.