A harness for agentic verification

VeriHarnessScaling Agentic Verification
for Long-Horizon Tasks

Same model. Better evidence.
Check what agent outputs disagree on—and what they all get wrong. Deliver a better artifact with a record of why.

THE VERIFICATION HARNESSONE MODEL, THROUGHOUT
Resolve disagreement. Challenge consensus. Adjudicate with evidence. Blue candidate rollouts feed two independent investigations: a yellow disagreement resolver and a red consensus challenger. Both evidence records feed green adjudication, producing an artifact and its verification record. Candidate rolloutsSame model · different attempts ≠ = ResolveChallenge Which alternative holds up?DISAGREEMENT → CHECK What did everyone miss?CONSENSUS → TEST Adjudicate & deliverRevised artifact + verification record

Independent checks. Environmental evidence.

Caiqi Zhang1*Rujun Han2Zifeng Wang2Zoey CuiZhu2Nigel Collier1Tomas Pfister2Chen-Yu Lee2

1 University of Cambridge   ·   2 Google Cloud AI Research

* This work was done while Caiqi interned at Google Cloud AI Research.

+6.2 pts

Gemini 3.5 Flash
average gain over a single rollout

+6.4 pts

Claude Opus 4.8
average gain over a single rollout

10 / 10

model–benchmark settings
best primary selection score

5 benchmarks

Documents, spreadsheets, code
and multi-file workspace tasks

01 / The observation

Agreement is not evidence.

Repeated attempts expose different answers, but counting votes cannot tell us which claims to trust. And when every rollout agrees, the same model may simply be repeating the same mistake.

Disagreement reveals alternatives.

74%of disputed claims contain a correct candidate

The most frequent value is correct only 47% of the time. Evidence can distinguish alternatives that voting misses.

Consensus can hide a shared error.

34%of unanimous claims are incorrect

Choosing another rollout cannot fix a claim they all get wrong. The shared value, interpretation, or omission needs a check of its own.

Claim-level analysis of ten-rollout Claude Opus 4.8 pools on APEX-Agents (§2). These are statistics about individual claims, not the probability that a complete artifact is correct.

02 / Inside the harness

Two investigations.
One better artifact.

Interactive walkthrough

The task asks for FY2025 revenue from the final financial report. Inspect the two sources of evidence below, in either order, then bring the findings together.

Task   Report FY2025 revenue from the final financial report.

PAPER’S WORKED EXAMPLE · FIG. 3

1 Candidate rollouts Three attempts from the same model

Rollout A01

USD 100m

Source: draft report

Agrees with B

Rollout B02

USD 100m

Source: draft report

Agrees with A

Rollout C03

USD 120m

Source: final report

A different revenue figure

2 Independent investigations Separate contexts, shared evidence tools

Disagreement resolver

100m or 120m? Trace the competing figures to the version the task actually requests.

A check that separates the alternatives: inspect the report’s version history.

Consensus challenger

All three say USD. Does that agreement survive a check against the source metadata?

A check that tests the shared assumption: inspect the report’s currency metadata.

Same model · separate investigation contexts · each produces an evidence record

3 Adjudication → delivery

What should be delivered?

Gather both evidence records to select a base and make a supported revision.

0 of 2 evidence checks complete

An interactive rendering of the paper’s illustrative example. The two investigators run independently; adjudication reads both records in a fresh context. Recorded runs follow in the case studies.

03 / Recorded runs

Case studies

Each case is a recorded run from the paper’s evaluation: the pool the generator produced, what the two investigations checked in the workspace, and what was delivered.

Recorded runs from the main evaluation (Table 2). Sentences in plain type are our summaries. Text in the record panels is quoted from the verification record and shortened only where marked […]. Scores come from each benchmark’s own grader on a 0–1 scale; the verifier sees no scores, reference answers, or rubrics.

04 / The evidence

Better outputs with the same model.

Average native benchmark score: 47.2 → 53.4
Single rollout → VeriHarness with evidence-backed revision

+5.4 pts
Single rolloutLLM-as-a-VerifierVeriHarness

Native benchmark score · per-benchmark axes, non-zero origins

Table 2 · Same frozen pool of 10 rollouts for every method, same generator and verifier model, three verifier seeds. Each panel uses its own labeled vertical range, following the paper; the default axes do not start at zero. The average weights all five benchmarks equally. LLM-as-a-Verifier is the most recent prior method in the comparison.

Explore all methods and exact scores
All methods, Gemini 3.5 Flash. Mean and cross-seed standard deviation.
MethodAvgAPEXWSBWorkBuddySB-2JobBench

± denotes cross-seed standard deviation. CLI deployments are listed separately. The grader-informed oracle bounds selection from the fixed pool; it does not bound revision. WSB: Workspace-Bench Lite. SB-2: SpreadsheetBench 2.

VeriHarness has the highest primary selection score in all ten model–benchmark settings, and revision improves on its selected base in every setting. With revision, it also has the highest score among the compared methods on every benchmark with both models.

05 / Learning what to check

The model stays fixed. The skills improve.

A verification skill describes a reusable failure mode and a way to test it. Failure feedback can turn these procedures into a growing library of checks, with no change to the model or the verification protocol.

Experience becomes a checking procedure.

A learned spreadsheet skill asks the verifier to read the period in a growth-rate label and confirm that the formula uses the matching start and end columns.

  1. Inspect development-task failures and feedback
  2. Propose reusable checking and revision skills
  3. Screen for task-specific answers and identifiers
  4. Retain a library using development scores

Held-out tasks measure progress. Their scores never choose the library.

Evolving from the human-authored library adds 6.8 points; evolving from empty adds 11.0 points.

Claude Opus 4.8 · Final library comparison (Fig. 4) · Native scores, bars start at zero. Held-out subsets differ from the full-benchmark evaluation above.

What the harness adds

A new axis for scaling verification.

Verification capability depends on more than the base model. A workspace, active evidence checks, independent investigations, and reusable skills help the same model deliver a better artifact.

The protocol applies across reports, workbooks, code, and multi-file outputs, and most of the selection gain transfers to existing agent CLIs.

Scope of the evidence

Quality comes with an investigation cost.

The reported setting uses ten rollouts per task and two frontier models. Verification adds tool calls and latency; the study targets professional deliverables where that trade-off is useful.

The verifier receives no reference answers or grading rubrics at verification time. Its output is an artifact and an evidence record, not a calibrated scalar score or a guarantee of correctness.

Build on this work

Cite VeriHarness.

@article{veriharness2026,
  title   = {VeriHarness: Scaling Agentic Verification
             for Long-Horizon Tasks},
  author  = {Zhang, Caiqi and Han, Rujun and Wang, Zifeng
             and CuiZhu, Zoey and Collier, Nigel
             and Pfister, Tomas and Lee, Chen-Yu},
  journal = {arXiv preprint arXiv:2610.00972},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.00972}
}