All writing

Build log

Before locking the cuts, I ordered the AI to argue against itself

An AI removed five metrics from a commission scorecard as unmeasurable or unpublishable. A confirmation is worthless if the checker wants to agree, so every cut had to survive an agent ordered to rescue it.

Key findings

  • Five parallel agents, one per removed metric, each ordered to save the metric first: a cut stood only if its blocker survived primary sources.
  • Four cuts held and got stronger: a soft 'no data published' became a statute, case law, or a dataset's geography limit.
  • One cut was wrong (a free federal dataset had the signal), and the re-check caught two of the AI's own framings as false.

My AI cut five metrics from a civic scorecard. Before locking the cuts, I spawned five agents, one per cut, each ordered to rescue its metric.

One cut didn't survive.

An AI that double-checks its own recommendation is a rubber stamp with extra steps. A confirmation only means something when the checker is trying to reach the opposite verdict.

The project is a draft economic scorecard for a city economic-development commission. Details are generalized throughout to protect the people involved: no city, no statute numbers, no names. Working through a dictionary of 39 candidate metrics, the AI had removed or re-scoped five as unmeasurable or legally unpublishable. Reasonable calls, plausibly argued. Also exactly the kind of decision that gets locked because nobody wants to re-litigate it.

Why not just ask the AI to double-check its own cuts?

Because I had already caught it "correcting" itself into an error. Earlier in the same project, the AI confidently replaced a figure it had researched before. The replacement turned out to be unsourced, and the original figure was right. The project log records that reversal, and it changed how I read every confident claim after it. A model that can be wrong in both directions about the same number does not get to certify its own removals.

There's a second, quieter problem. The first-pass rationales were soft. "No data is published at this geography" is a claim about the world, and the AI hadn't tested it against the world — it had pattern-matched from what data usually exists. Soft reasons rot. One sharp question in a public meeting and the whole cut folds.

What does an honest attempt to break a decision look like?

Five parallel agents, one per removed metric. Each got a brief that inverted the burden: try to SAVE this metric. Find the dataset the first pass missed. Read the statute itself, not a summary of it. Confirm the removal only if the blocker survives contact with primary sources — the statute, the case law, the federal data program's own documentation.

Every source got ranked on a tier ladder: A for statutes and agency primaries, B for credible secondary coverage, C for aggregators. A cut could not stand on tier-C evidence. This is adversarial review (trying to break a claim rather than confirm it), and the inverted brief matters as much psychologically as procedurally. An agent told to verify a cut will find reasons the cut was right. An agent told to rescue the metric has to actually go looking.

The run produced one memo, filed in the project record as the removed-metrics verification memo dated 2026-05-30, with one verdict per metric: CONFIRM or PARTIAL-REVERSE.

Ordered to argue against itself: the adversarial re-check of five metric cuts Ordered to argue against itself 5 metrics cut from a draft civic scorecard, as "unmeasurable" or "unpublishable" The re-check, burden inverted 5 agents, one per cut, each ordered to SAVE its metric. A cut stands only if the blocker survives primary sources: statute, case law, agency docs 4 cuts held, stronger soft "no data" replaced by statute, case law, or a geography limit 1 cut partially reversed a free federal dataset already had the signal 2 framings corrected the AI's own earlier claims, caught wrong by the re-check
Five removed metrics re-enter review through agents ordered to rescue them; a cut is confirmed only when its blocker survives primary sources. Diagram source: this page; maps 1:1 to the 2026-05-30 removed-metrics verification memo.

What did arguing against itself actually change?

Four of the five cuts held. Zero of the five original rationales did.

That's the result I didn't expect. Every confirmed cut came back stronger, because a vague "no data published" was replaced with a hard, citable reason. One metric is fenced off by a state tax-confidentiality statute whose decisive bar is a use restriction: the data can be obtained for government functions but not published. No amount of clever redaction lifts a use restriction. One is blocked by a state constitutional bar on preference programs in public contracting, upheld by the state supreme court. One dies at a geography limit, because the federal statistical programs that measure business survival publish nothing below county level. One runs into federal tax-return confidentiality, which makes an individual credit claim secret by statute.

The fifth cut was wrong. The AI had removed a commercial-vacancy metric as a data gap, claiming the city would have to commission a paid inventory to measure it. The rescue agent found a free federal address-vacancy dataset that publishes business-address vacancy at census-tract level, updated quarterly. The verdict reads PARTIAL-REVERSE: buildable now as a labeled proxy — a stand-in measure with stated limits, since it counts mail-deliverable addresses sitting unused, not leased square footage. Worse, the bad rationale had been drafted into the commission's proposed first formal recommendation: fund an inventory. The re-check softened that recommendation: adopt the free proxy now, and treat a ground survey as an optional upgrade.

It also caught the AI misdescribing two things it had already confirmed. It had implied the tax-confidentiality problem might be solved by suppressing small counts. Wrong — suppression redacts identities inside a releasable document, and this statute restricts the use of the data itself. And it had described a local-business preference policy as a participation goal the city commits to hit, when the real instrument is a rating incentive applied at bid evaluation. A thumb on the scale, not a floor. Both corrections are logged in the same memo.

So when is a confirmation worth anything?

When the decision survived an honest attempt to break it. By that standard the first pass scored worse than it looked: five decisions re-checked, three materially changed. One partial reversal, two corrected framings, every rationale replaced.

The failure mode this guards against isn't the dramatic one. Nobody was about to publish a fabricated number here. The failure mode is a right decision with a weak reason, locked into a public document, waiting for the first reader who opens the actual statute.

I hold reviews to the same rule elsewhere: the reviewer must not be invested in the work passing. That rule runs through the ninety-nine-agent review at much larger scale, and it's a load-bearing wall in how I know my AI is right. The version here is the cheapest form I know. One inverted brief, five parallel sessions, one memo.

So, my question for you: take the last AI recommendation you accepted. If an agent were ordered to argue the other side against primary sources, would the reason survive — or just the decision?


Method & data

Method: a civic-scorecard build for a city economic-development commission — the 2026-05-30 removed-metrics verification memo, its v1.1 change list, and the indicator-dictionary audit, generalized
Data: patterns and methods only; no names, dollar figures, addresses, case identifiers, or confidential content · Last checked: 2026-08-14

How this was made

AI-drafted, adversarially checked, human-directed. My AI assistant wrote this from the system's own records — the removed-metrics verification memo, the v1.1 change list, and the indicator-dictionary audit summary. A separate AI session then tried to break every claim against those records, and automated privacy and readability gates ran before publish. I direct this pipeline, own every boundary in it, and audit published pages on a rolling basis — if you find an error, tell me and it goes in the corrections log, dated, never silent.

I'm Ali — I run real life-and-work admin on AI agents, then check their work in the open. More at /about.

Published under my standards. Found an error? Tell me — corrections go in the corrections log, dated, never silent.

Cite this

@online{ali2026argueitself,
  author = {Ali},
  title  = {Before locking the cuts, I ordered the AI to argue against itself},
  date   = {2026-08-14},
  url    = {https://alidoes.ai/ai-argue-against-itself/}
}

Caught something I got wrong? Send it directly. Confirmed corrections go in the corrections log.