Can an AI ingest 2,798 email threads without inventing one record?
An AI backfilled 2,798 email threads into my structured archive. The number that matters is zero fabricated records, and the boring tripwires that made zero possible.
Field notes, not a content feed
Each article answers one useful question from real work. The method, evidence, failure boundary, and correction path stay attached so you can judge the claim instead of taking my word for it.
Published tests
42 articles
An AI backfilled 2,798 email threads into my structured archive. The number that matters is zero fabricated records, and the boring tripwires that made zero possible.
Two views of one plan, each green on its own, quietly stopped matching. The post-mortem: why 'which one is right?' was the wrong question, and the check that answers it now.
An automation that dies without an error looks exactly like one that's working. This is the small check that watches the silence instead, and the reproducible benchmark showing where it works and where it fails.
Two runs of one deep-research task, locked to the same sources, landed about 3% apart on a number headed for a sworn filing. The gap between the reports told me which claims were facts and which were the model's judgment.
An AI drafts my public materials, so the public/private line can't live in anyone's memory. It's a machine-checked property of the file bytes — and the check already caught drift on my live site.
An agent misread a broad 'continue' as a freeze release and re-enabled a deliberately disabled scheduled task; its output merged to master. The postmortem: the pause never lived anywhere a machine could read it.
A property portal collected 25-agent, then 45-agent, then 99-agent review audits. The votes stacked into agreement. The safety came from somewhere much smaller.
Answer engines churn their citations weekly and expose no API. This is the small instrument I trust instead of buying a rank number — including the humble score it currently shows for my own site.
I benchmarked local models against paid frontier tiers on real document grunt-work, with sealed answer keys and a deterministic grader. The first verdict blamed the model. The recorded runs blamed my inputs.
I ran local models on my own GPU to make AI automation cheaper without trading away trust. A full audit found neither current quality proof nor a single measured cost, so the lab was rebuilt to fail closed.
My complete career evidence archive holds employer reviews, transcripts, and legal and financial records. The AI-built tool that governs its integrity is allowed to ask each file exactly three questions: path, size, modified time.
My AI assistant runs real personal, property, and civic work. This is the full tour: the architecture, the checks that make it trustworthy, and the parts where it failed.
The health check asked one question: does the shared folder exist? It did. Ten governed records were gone anyway. This is the post-mortem of a green light that proved almost nothing.
Every session hit the same credential wall, reported git as blocked, and politely waited for a human. The wall was real. The routing into it was the agent's own stale memory.
My web-research harness gave every answer a green GROUNDED stamp if it had fetched the cited page. The stamp measured retrieval, not support — and the gap was enormous.
Screening applicants is the highest-liability step a small landlord can automate. My AI never scores anyone; it keeps the ledger that proves every household was treated the same.
Article releases on this site run through a tool that refuses to stage a file unless the fresh review and my approval name its exact bytes.
My email client decided two conversations from different years were one thread. The AI that read it inherited the lie: an expired decision came back labeled live.
A zero that means 'we didn't look' is a lie with good posture. My measurement dashboard refuses to print one: every empty cell must say exactly what kind of empty it is.
My AI assistant summarized every message it read. When meeting and portal credentials arrived pasted into raw messages, it copied them into summaries and entity lists: one exposure, multiplied.
A scorecard is only as good as its stalest fact. Mine gives every claim an expiry clock, and the quarterly re-check would rather leave a row visibly stale than fake a refresh.
The AI reported my vendor-comparison workbook finished because the tabs and formulas existed. Ten of eleven scoring tabs held zero ratings, and the dashboard read a tab that was never created.
A document that merely mentioned another property pulled a tenant transition into the wrong home record. The fix wasn't a smarter model — it was a better routing key.
My workspace verifier's 'fast' path quietly grew past a minute at session start, so nobody ran it. Bounding it to entry-critical invariants got it to 4.61 seconds, and back into every session.
My AI swept roughly 30 turnover vendors and reported nothing missed — it had read the top of a ranked list and guessed about the rest. The audit that replaced guessing with counting found the winning bid in a missed thread.
The rule said: scaffold the student's argument, never write their sentences. It held right up until a new kind of document appeared with no check watching it.
Every new AI session met my live app through a doc that said it didn't exist yet, and billed me to rediscover it. The fix was one rule: state and plans never share a file.
The scariest leak in my system was a helpful one: private correspondence had become reusable voice context for unrelated AI work, and nothing was broken. This is the post-mortem, and the boundary check that ended it.
My AI assistant's learned rules lived where the tool put them by default: on the machine. One migration later, the index pointed at 17 files that no longer existed.
I let an AI manager label document-extraction cases before four contestant configurations ran them. Nine of 11 gradeable cases came back harder than the label — and the worst miss was the file that looked easiest.
An AI summarizing a legal folder merged two counterparties and promoted mediation paperwork into a settlement. The most dangerous wrong answer is the one that asks for no attention.
Ten scheduled tasks looked like ordinary work and were secretly probes of my AI's own rules. The score is honest because the sessions didn't know, and the worker never graded itself.
The first run of my task-difficulty benchmark was won by the model that designed it, ran it, and graded it. This is the post-mortem: the discard, the sealed rebuild, and the sequel I refused to publish.
A chapter draft arrived through the wrong folder, so my AI sampled it instead of reading it — and the feedback looked exactly like it had done the work. The full audit found ~75 issues the sample missed.
Three red past-due warnings on a family home's paperwork. Zero real deadlines. This post-mortem covers why a naive date parser cries wolf, and the two-part fix that made red mean red again.
On launch day my property portal could not email a single person, by design. The one path that could have reached a real tenant never passed through my kill switch at all.
We bought a four-unit building, and AI helped turn plain-English requests into 918TS — a live property portal with resident, manager, and owner views. The prompts were the easy part; the permissions, tests, and maintenance were the real work.
Auto-managed memory is genuinely convenient, and exactly as trustworthy as an unreviewed note the model wrote to itself. Here is the line I drew, and the rule that guards it.
An AI workspace doesn't know you quit. Mine kept encoding board authority I no longer had, until the whole domain was flipped to a hold a script can verify.
A credentialed analyst delivered a finished reserve study to an HOA board. An AI-assisted audit found old assumptions riding inside it — then caught its own reading error too.
An AI removed five metrics from a commission scorecard as unmeasurable or unpublishable. A confirmation is worthless if the checker wants to agree, so every cut had to survive an agent ordered to rescue it.
An AI reads every wedding-vendor quote I get and fills my tracker. Nothing enters until it clears four gates — because the failure I almost missed looked exactly like done.
Reader contract
01
One decision or question, stated plainly.
02
Sources, test conditions, and an artifact where safe.
03
What the result does not establish and what would change it.