The Bounded Verdict Ledger for Auditing AI Research Claims
A passing research dossier is not proof. The bounded verdict ledger traces every claim to its source and states exactly where the checking stopped.
Yash Sharma5 min readEvaluation and assurance
Research area 49: Independent adversarial audit of the research dossiers
In brief
- A passing dossier is not proof until every consequential claim is traced, rebuilt and graded independently.
- The bounded verdict ledger traces every claim to its source, rebuilds its math and assigns a graded verdict.
- Skipping the audit means funding decisions on evidence that was never actually checked.
A research dossier that reads like proof usually is not proof, and that gap decides whether an investment survives its first hard question. The independent adversarial audit of the research dossiers is what happens when a second, skeptical reader stops trusting a report's own confidence, retraces every consequential claim to its exact source, rebuilds the arithmetic by hand, and states plainly where the checking stopped. Executives who skip this step outsource the funding decision to whichever author wrote the most persuasive paragraph.
What it is
An adversarial audit treats a finished looking report as individually testable claims, not a single verdict to accept or reject. Here, nine specialist dossiers covering inference performance, routing, deployment, cost, human oversight, product value, system architecture, scientific computing and reliability were checked by a same run second reviewer against their own claims and arithmetic, producing 38 traced claims plus a further five source random supplement drawn from outside the highest risk set.
Four ideas carry the method. Trace every consequential claim to its exact source, version and stated scope, rather than accepting a footnote as sufficient. Grade each claim with a bounded verdict, supported, supported with limits, unresolved, contradicted, ineligible or unverifiable, so a report cannot hide an overclaim inside a passing grade. Target the highest risk claims on purpose, distribution guarantees, equivalence claims, formal proofs, human outcome studies, and also draw an unbiased random sample from everything else. Publish the audit's own boundary as loudly as its findings: a same run review sharing its author's assumptions is not an independent reproduction.
Organizations now fund decisions, not just publish decks, off these dossiers, so a misread number sizes a real budget or autonomy grant. The audit caught an incident of 54 consecutive successful calls summarized as 54 separate incidents, corrected, and a routing benchmark whose prose and table totals disagreed by 19 prompts and 525 instances, marked unresolved rather than guessed. It also caught a subtler error: a reported gain from 64.0 percent to 74.6 percent is 10.6 percentage points, about 16.6 percent in relative terms, phrases swapped often enough to make a modest gain sound bigger.
The architecture
I call the mechanism the bounded verdict ledger, a running record where every consequential claim earns a bounded verdict carrying its exact source, its exact scope, and whatever still sits outside what was checked.

The ledger runs in four moves. A claim enters as written, with whatever confidence its author gave it. Two checks run before it goes anywhere: one traces the claim to its original source, version and locator, and the other rebuilds any arithmetic behind it rather than trusting the report's own math, a step that mattered here because rebuilding a retry model's naive independence assumption changed a three attempt success estimate from .993141 to .8991 once a persistent share of impossible tasks was accounted for. The claim then lands on a graded rung, most often supported with limits, occasionally unresolved when the source is inconsistent or inaccessible, rarely contradicted when the wording overreaches the source. Only a claim that has cleared a rung, with its scope intact, is allowed to cross into a funding or deployment boundary.
The most common failure mode is not a fabricated number but a true number stripped of the condition that made it true, a study about a curated subset of problems reported as a claim about problems in general, or a formal proof about a reference model reported as a guarantee about whatever a team deployed.
Ownership belongs to a function that did not write the dossier and does not benefit from its conclusion, the way a finance team does not let a trading desk mark its own book. A same run second reader can catch real extraction errors, as it did here, but is not a substitute for a fully independent reviewer working from separate search and access. Fund tracing every high risk claim to its source, reconstructing arithmetic independently, and a small disclosed random sample of lower risk claims. Measure the ledger's verdict distribution: every claim coming back supported with no limits signals an audit that is not adversarial enough, not strong research. Two decisions belong to an executive alone: what counts as sufficient evidence when a claim is graded unresolved, and whether to fund the follow up study a dossier names before scaling a decision that rests on that gap.
What it is worth
The value shows up earliest in decisions an organization avoids making on evidence it did not have. Three hundred clean trials feel like proof of near perfect reliability, but zero observed failures across 300 trials still leaves a one sided 95 percent upper bound on the true failure rate of about 0.99 percent, and pushing that bound to one hundredth of a percent takes roughly 29,956 trials. A program that funds the 300 and calls the reliability question closed is pricing a guarantee it never bought.
The reviewed evidence supports specific, bounded findings: a corrected incident count, a narrowed formal proof scope, a disclosed arithmetic error in a retry model, and an unresolved benchmark accounting gap. It does not support treating any dossier, or this audit of it, as an exhaustive systematic review, an independent certification or a completed empirical validation, and no return figure is claimed here because none was measured. What it is worth is a funding decision that survives a second, harder look.
Questions leaders ask
- What does it mean to run a research dossier through an adversarial audit?
- It means treating every consequential claim as individually testable rather than accepting the report's own confidence. A second reader traces each claim to its exact source and version, rebuilds the arithmetic independently, and assigns a bounded verdict such as supported with limits, unresolved or contradicted, instead of one pass or fail grade.
- How many clean test runs actually prove an AI system is reliable?
- Far more than most teams assume. Three hundred independent trials with zero failures still leaves a one sided 95 percent upper bound on the true failure rate of about 0.99 percent, not proof of near perfect reliability. Pushing that bound to one hundredth of a percent takes roughly 29,956 trials.
- Can a report's own author audit it and still be trusted?
- Only as a qualified reference, not independent certification. A same run second reader can catch real extraction errors, including a corrected incident count and a narrowed proof scope here, but it shares the author's search process and source access, so genuine independent replication has to come from outside that shared context.
- What should a leader do when a dossier marks a finding unresolved instead of proven?
- Treat it as legitimate evidence about the limits of what is known, not a failure to smooth over. Fund the specific follow up study the dossier names, or narrow the business claim to what the evidence supports, rather than averaging conflicting numbers into a comfortable figure or upgrading unresolved to supported because a decision is waiting.


