Disagreement reveals alternatives.
The most frequent value is correct only 47% of the time. Evidence can distinguish alternatives that voting misses.
A harness for agentic verification
Same model. Better evidence.
Check what agent outputs disagree on—and what they all get wrong. Deliver a better artifact with a record of why.
Independent checks. Environmental evidence.
1 University of Cambridge · 2 Google Cloud AI Research
Gemini 3.5 Flash
average gain over a single rollout
Claude Opus 4.8
average gain over a single rollout
model–benchmark settings
best primary selection score
Documents, spreadsheets, code
and multi-file workspace tasks
01 / The observation
Repeated attempts expose different answers, but counting votes cannot tell us which claims to trust. And when every rollout agrees, the same model may simply be repeating the same mistake.
The most frequent value is correct only 47% of the time. Evidence can distinguish alternatives that voting misses.
Choosing another rollout cannot fix a claim they all get wrong. The shared value, interpretation, or omission needs a check of its own.
Claim-level analysis of ten-rollout Claude Opus 4.8 pools on APEX-Agents (§2). These are statistics about individual claims, not the probability that a complete artifact is correct.
02 / Inside the harness
The task asks for FY2025 revenue from the final financial report. Inspect the two sources of evidence below, in either order, then bring the findings together.
1 Candidate rollouts Three attempts from the same model
USD 100m
Source: draft report
Agrees with B
USD 100m
Source: draft report
Agrees with A
USD 120m
Source: final report
A different revenue figure
2 Independent investigations Separate contexts, shared evidence tools
100m or 120m? Trace the competing figures to the version the task actually requests.
All three say USD. Does that agreement survive a check against the source metadata?
Same model · separate investigation contexts · each produces an evidence record
3 Adjudication → delivery
Gather both evidence records to select a base and make a supported revision.
An interactive rendering of the paper’s illustrative example. The two investigators run independently; adjudication reads both records in a fresh context. Recorded runs follow in the case studies.
03 / Recorded runs
Each case is a recorded run from the paper’s evaluation: the pool the generator produced, what the two investigations checked in the workspace, and what was delivered.
Recorded runs from the main evaluation (Table 2). Sentences in plain type are our summaries. Text in the record panels is quoted from the verification record and shortened only where marked […]. Scores come from each benchmark’s own grader on a 0–1 scale; the verifier sees no scores, reference answers, or rubrics.
04 / The evidence
Average native benchmark score: 47.2 → 53.4
Single rollout → VeriHarness with evidence-backed revision
Native benchmark score · per-benchmark axes, non-zero origins
Table 2 · Same frozen pool of 10 rollouts for every method, same generator and verifier model, three verifier seeds. Each panel uses its own labeled vertical range, following the paper; the default axes do not start at zero. The average weights all five benchmarks equally. LLM-as-a-Verifier is the most recent prior method in the comparison.
| Method | Avg | APEX | WSB | WorkBuddy | SB-2 | JobBench |
|---|
± denotes cross-seed standard deviation. CLI deployments are listed separately. The grader-informed oracle bounds selection from the fixed pool; it does not bound revision. WSB: Workspace-Bench Lite. SB-2: SpreadsheetBench 2.
VeriHarness has the highest primary selection score in all ten model–benchmark settings, and revision improves on its selected base in every setting. With revision, it also has the highest score among the compared methods on every benchmark with both models.
05 / Learning what to check
A verification skill describes a reusable failure mode and a way to test it. Failure feedback can turn these procedures into a growing library of checks, with no change to the model or the verification protocol.
A learned spreadsheet skill asks the verifier to read the period in a growth-rate label and confirm that the formula uses the matching start and end columns.
Held-out tasks measure progress. Their scores never choose the library.
Evolving from the human-authored library adds 6.8 points; evolving from empty adds 11.0 points.
Claude Opus 4.8 · Final library comparison (Fig. 4) · Native scores, bars start at zero. Held-out subsets differ from the full-benchmark evaluation above.
What the harness adds
Verification capability depends on more than the base model. A workspace, active evidence checks, independent investigations, and reusable skills help the same model deliver a better artifact.
The protocol applies across reports, workbooks, code, and multi-file outputs, and most of the selection gain transfers to existing agent CLIs.
Scope of the evidence
The reported setting uses ten rollouts per task and two frontier models. Verification adds tool calls and latency; the study targets professional deliverables where that trade-off is useful.
The verifier receives no reference answers or grading rubrics at verification time. Its output is an artifact and an evidence record, not a calibrated scalar score or a guarantee of correctness.
Build on this work
@article{veriharness2026,
title = {VeriHarness: Scaling Agentic Verification
for Long-Horizon Tasks},
author = {Zhang, Caiqi and Han, Rujun and Wang, Zifeng
and CuiZhu, Zoey and Collier, Nigel
and Pfister, Tomas and Lee, Chen-Yu},
journal = {arXiv preprint arXiv:2610.00972},
year = {2026},
url = {https://arxiv.org/abs/2610.00972}
}