VeriHarness: evidence-backed verification for long-horizon agents
Google Cloud AI researchers and Cambridge shipped VeriHarness, a verification protocol where the same model runs two independent investigations with evidence tools and reusable skills, then adjudicates; gains of +6.2 points average over single rollouts on Gemini 3.5 Flash across 5 workspace benchmarks.
Why it matters
Verification scales without a bigger model — the verifier gets no reference answers or rubrics, yet fixes shared errors that voting across 10 rollouts cannot.
What you could build with it
An engineering consultancy could wrap VeriHarness into its client-delivery pipeline for document, spreadsheet, and code generation agents, producing an evidence record alongside every deliverable so clients can audit correctness instead of re-running the work.
Does it hold up?
Open-source Apache 2.0 harness with benchmark adapters and a skill library released alongside the paper — usable today as a research reference and eval recipe, but no confirmed production deployments yet.
Built with VeriHarness
- Google Researchers Ship VeriHarness to Verify Agent Outputs Across Long-Horizon Taskstheagenttimes.com · Independent writeup covering the 10-rollout protocol, benchmark gains table, and the released ~26,000-rollout dataset.
- VeriHarness interactive worked example and recorded runsofficial site · Interactive walkthrough of the paper's example plus recorded verification runs from the main evaluation.
- google-research/veriharnessgithub · Official Apache 2.0 harness: driver/runner/scorer, benchmark adapters and graders for APEX-Agents, JobBench, SpreadsheetBench-2, WorkBuddy-Bench, Workspace-Bench, plus the reusable verification skill library.
Learn more
First spotted on article: source.
More AI agents
Agent ReachOpen-source CLI and SKILL.md kit that gives AI agents read/search access to 15+ internet platforms (Twitter, Reddit…agents · JEV 0.8Docker Agent: AI agent builder and runtime by Docker EngineeringA Docker CLI plugin for building, running and sharing AI agents with declarative YAML config, multi-agent…agents · JEV 0.75Google Agent SkillsGoogle's official catalog of 150+ agent skills for its products and technologies — Cloud, GKE, BigQuery, Gemini, Agent…agents · JEV 0.74hermes-agent (NousResearch)Self-improving personal AI agent across Telegram, Discord, Slack, and terminal with persistent memory.agents · JEV 0.73Paperclip: the app people use to manage agents at workOpen-source orchestration for teams of AI agents: org charts, budgets, governance, heartbeats and a ticket system in…agents · JEV 0.73AIHOT — open-source framework for building automated industry-news sitesFull-stack framework (Node 24 + PostgreSQL) that ingests RSS/webpages/X/WeChat, dedupes, double-scores, clusters with…agents · JEV 0.73
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.