Skip to main content
AI Interview Question
INTERVIEW GUIDEAI News8 questions5 min readOct 5, 2026

VeriHarness: why majority voting falls short for AI agents

Google Cloud AI Research's VeriHarness checks agent outputs against evidence, not votes: same model, no rubric, 6.2 to 6.4 points over a single rollout.

VeriHarness: why majority voting falls short for AI agents

Researchers from Google Cloud AI Research and the University of Cambridge posted VeriHarness to arXiv on 1 October 2026. It is a harness that turns an agent's own model into a verifier for long-horizon work such as reports, spreadsheets, and code patches. The verifier gets no reference answers, no grading rubric, and no stronger judge. Across five workspace benchmarks, it beat majority voting, LLM judges, and pairwise tournaments at picking the best output, and with revision it added 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.

It targets a common test-time habit: sample the same task several times and keep the answer most runs agree on, as majority voting and self-consistency do. The paper shows why that breaks on long tasks. On ten-rollout pools from Claude Opus 4.8 on APEX-Agents, 34% of the claims every rollout agreed on were judged incorrect. On claims where rollouts disagreed, 74% had a correct candidate somewhere in the pool, but the most frequent value was right only 47% of the time. Agreement is not evidence: consensus can hide a shared error, and voting can throw away the right answer.

VeriHarness gives the verifier the same model and tools as the generator, plus a workspace. Each rollout's delivered artifact, final state, and action trace sit there as files next to the task and its source data. The verifier can trace a claim to its source, recompute a number, or run code. Two investigations then run in separate contexts. A disagreement resolver finds claims the rollouts dispute and picks checks that best tell the candidates apart, discarding the ones the evidence contradicts. A consensus challenger takes claims everyone agreed on, proposes how they could be wrong, and tests that: it recomputes values, checks labels against source metadata, tests a shared reading of a file or period, and looks for requirements every rollout missed.

A fresh context of the same model then adjudicates. It reads both evidence records, picks a base rollout, writes an evidence-backed revision plan, and lists claims that are still unsettled. The delivered output comes with a verification record linking each claim to its check, evidence, and verdict. The paper's worked example: two rollouts report FY2025 revenue of 100m from a draft report, one reports 120m from the final report, and all three say USD. The resolver checks version history and backs 120m. The challenger checks source metadata and finds the currency should be EUR, an error no vote could catch.

Reusable verification skills carry the domain knowledge. A skill is a short description of a failure mode and how to check for it, with an optional script, and never a task-specific answer. One example: check that a growth rate uses the start and end columns its label names. The authors also let the verifier grow its skill library from failures on development tasks. On held-out tasks with Claude Opus 4.8, a library evolved from empty beat the empty library by 11.0 points on APEX-Agents and 5.7 on SpreadsheetBench 2, and also beat the human-written library. Evolving from the human library added 6.8 and 3.7 points over it.

The numbers come from five benchmarks: APEX-Agents, Workspace-Bench Lite, WorkBuddy Bench, SpreadsheetBench 2, and JobBench. Every method got the same ten rollouts per task. With Gemini 3.5 Flash, the five-benchmark average went from 47.2 for a single rollout to 47.5 with majority voting, 51.6 with VeriHarness selection, and 53.4 with revision. With Claude Opus 4.8 it went from 49.6 to 50.4, 53.7, and 56.1. VeriHarness had the best selection score in all ten model and benchmark settings. An agentic verifier with the same environment access but without the protocol and skills recovered only about half the gain, so tool access alone is not enough.

The ablations show what each part does. With Opus, the resolver alone took the average from 49.6 to 54.3 and the challenger alone to 52.6; together they reached 56.1. Running both in one context cost 0.8 points. On APEX-Agents with Opus, selection scored only 0.6 points above the pool mean when all rollouts agreed, and 6.1 points above it when they disagreed. The challenger never overturned a claim that every rollout had right. When shared errors survived, about 70% of the time it had checked a correct intermediate step while the error sat further downstream.

Cost is reported too. Selection cost $1.44 per task with Flash, a tenth of LLM-as-a-Verifier's cost at twice the gain, and $3.92 per task with Opus. Revision costs about three times as much as selection. Because calls are incremental turns over one workspace, 86 to 91% of input tokens were served from cache. The authors list the limits: it needs a pool of rollouts (ten per task here), it gives no scalar score per rollout, it was only tested with the same model as generator and verifier, and it adds latency. The code and roughly 26,000 rollouts, produced at a cost of over $100,000, are public.

Interview angle: expect this in agent evaluation and system design rounds. If asked "why not just use majority voting or self-consistency for agent outputs?", say that voting only measures agreement. On long tasks, rollouts can share the same error, and the right answer can sit in the minority. Then describe the fix: check claims against the environment, treat disagreement and consensus as two different jobs, keep them in separate contexts, and adjudicate with a fresh context. Contrast it with LLM-as-a-judge, which scores by reading outputs rather than gathering evidence. Mention why code and math are easier (tests and proof checkers exist) than reports and spreadsheets. Close with the trade-offs: ten rollouts plus verification cost money and time, prompt caching keeps it affordable, and the verification record gives you an audit trail and possible training signal.

https://arxiv.org/abs/2610.00972 https://veriharness.com https://github.com/google-research/veriharness https://huggingface.co/datasets/caiqizh/veriharness

VeriHarnessGoogleAI agentsAgent evaluationLLM-as-a-judgeTest-time scaling

Questions in this guide

Deep explanations with architecture diagrams for every question below.