---
title: VeriHarness
type: vocabulary
url: "https://www.envisioning.com/vocab/veriharness"
summary: "An agentic verification method that turns a generator's own LLM into a verifier with a workspace and evidence tools, resolving disagreement among rollouts and challenging claims they all agree on."
year: 2026
generality: 0.40
---

# VeriHarness

An agentic verification method that turns a generator's own LLM into a verifier with a workspace and evidence tools, resolving disagreement among rollouts and challenging claims they all agree on.
VeriHarness is a method for verifying the outputs of long-horizon LLM agent tasks using the same base model that generated them, without reference answers or grading rubrics at test time. Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, and Chen-Yu Lee introduced it in "VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks," posted to arXiv (2610.00972) on 1 October 2026. It is a preprint and has not yet been peer reviewed. The name is also used, unrelated, by an earlier and much smaller July 2026 paper, "Structured Feedback Improves Repair in an LLM Agent Loop" (arXiv:2607.14167), for a code-controlled agent-repair loop. The two papers share only the name.

The method starts from an observation about repeated sampling. When an agent's rollouts agree, that consensus can still conceal a shared error. When rollouts disagree, the disagreement often points toward the correct answer rather than away from it. VeriHarness turns the generator's own underlying LLM into an agentic verifier by giving it a workspace, evidence tools, and a set of reusable verification skills, with two roles. A disagreement resolver checks competing claims from different rollouts against evidence available in the task environment. A consensus challenger instead interrogates claims every rollout agrees on, actively searching for requirements the rollouts collectively missed. Findings from both roles guide which rollout, or revision of one, gets selected as the final output.

Across five long-horizon workspace benchmarks and two frontier models, the authors report that VeriHarness achieves the highest selection scores among tested baselines. Evidence-backed revision adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The paper also reports that its verification skills can self-improve from failure feedback, and that the protocol transfers to existing agent command-line tools with most of the gain preserved. The authors release roughly 26,000 rollouts generated across the five benchmarks and both models, produced at an estimated cost of over $100,000, to support further research on agentic verification.

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/veriharness)
