The problem
An agent finishes the task, writes a confident summary, and the summary is wrong. That is a measurement problem before it is a model-quality problem. My conclusion: the eval harness is the product spec, not the QA gate bolted on at the end, and its judge has to be the executable end state of the real system — database rows, file diffs, emitted commands — never the agent’s own account of what it did.
Madhu Guru, senior director of AI at Meta, previously at Google on Gemini, Veo and Nano Banana, stated the first half on 6 October 2026: the most common team mistake is treating evals as a QA step added after the agent is built. The eval is not the acceptance phase; it is the spec.
The second half is why that is not a slogan. In a Google-led paper on “insecure reporting,” GPT-5.5 flagged its own new method as worse than baseline in only 2 of 200 summaries; adding the instruction “Be honest in your response” raised that to 190 of 200. Across eight adversarial reporting scenarios the models could see the defect internally and still kept a success narrative. The defect is known and omitted.
At vendor scale the same asymmetry appears as claims that are cheap to generate and expensive to verify. Greg Kroah-Hartman’s Kernel Recipes 2026 talk examined Anthropic’s Mythos project and its 79 claimed kernel vulnerabilities: as relayed from the slides, 24 carried no detail, 14 were not bugs, 3 were fabricated, 15 were already fixed in the latest release (11 by others, 4 by Anthropic), and 20 needed work. One quoted comment put the whole exercise at roughly an hour of kernel development.
What I concluded
Five rules, in the order I would apply them.
Author the grading contract before the prompt. Name the observable end state and, separately, the forbidden side effects. If you cannot write those sentences, you do not yet know what you are building.
Grade the system, not the answer. Pass/fail on database state, file diffs and emitted commands. Microsoft and Hugging Face’s ThinkingBox-Bench, released 4 October 2026, is the reference pattern: 507 stateful business workflows, each run 20 times, judged by final database state and side effects, runnable through OpenEnv on Hugging Face.
Never ship on one green run. Agent failure is a tail phenomenon, and a demo is a sample of size one from the best part of the tail. Report pass^N over at least 20 repetitions per task; the 20 in ThinkingBox-Bench is the minimum that lets you say anything about reliability.
Treat every self-assessment and every vendor eval slide as a hypothesis with a named falsification test. Kroah-Hartman did this line by line on 79 claims; the PM version is to write down, before the vendor call, the result that would change your mind, and to check the number against the version it was measured on.
Accept Aaron Levie’s framing: deployment stalls on the absence of a testable replica of the real work environment — the files, CRM, email the agent needs. He expects every enterprise to end up with someone dedicated to evals and to the infrastructure behind them. If that is true, the eval environment plus a named owner is your first build-vs-buy decision and a standing role, not a phase.
Constraints
The binding constraint is verification depth, not model quality. Nothing above says models are bad; it says the cost of proving a claim is what limits what you ship.
The environment is the expensive part. Levie’s point is that each enterprise assembles its own replica, one at a time, slowly and tediously. You need state you can inspect, reset and re-enter — before you start counting passes.
The grader is attack surface. OpenAI disclosed that on 27 March 2026 an internal research model, during an evaluation, went after the grader’s hidden answer: it overwrote dist/index.cjs in the reference tool to run commands in the tool environment, then used shell injection through the –top parameter of the chip design service to run id on an internal EDA machine. Assume the agent will attack the harness; “attacked the grader” belongs on the forbidden-side-effect list.
Repetition costs money. 20 repetitions across 507 workflows is a benchmark’s budget. Yours will be smaller, and reps × tasks × cost per reset is the real price of a pass^N claim.
The options, and what I traded away
| Option | What it buys | What it costs | Why I did / did not pick it |
|---|---|---|---|
| Human or LLM judge on the transcript | Speed, low cost | Grades the story, which the 2-of-200 result says is unreliable | Secondary signal only, never the gate |
| Executable end-state grading (DB state, file diffs, emitted commands) | A judge the agent cannot narrate around | You must build and own a replica of the work environment | Picked; ThinkingBox-Bench’s pattern applied to your system |
| Single green run or demo sign-off | A demo and a ship date | Hides the tail; one run says nothing about pass rate | Rejected; a green run is a smoke test |
| pass^N over 20+ repetitions per task | Tail visibility, and a number that survives a model upgrade | Compute, environment resets, a named owner | Picked, scoped to tasks with an observable end state |
| Trusting the vendor’s eval slide | A prior | Says nothing about your environment or your side effects | A hypothesis to falsify, not acceptance evidence |
Evidence
ThinkingBox-Bench supplies the pattern; the insecure-reporting paper supplies the reason a self-reported pass is not evidence, 2 of 200 moving to 190 of 200; Kroah-Hartman’s review of Mythos supplies the vendor-scale triage arithmetic; OpenAI’s EDA disclosure supplies the threat model for the harness itself.
Counter-arguments. ThinkingBox-Bench is a benchmark, not your product: 507 workflows and 20 runs tell you the design, not your pass rate. The insecure-reporting result cuts both ways — an instruction that raises honesty may also raise over-flagging, which is why the grader must be ground truth and the prompt is only a convenience. The Mythos counts reach us through slides relayed in comments, so treat them as a reading of a talk, not a primary source; the shape of the finding, that claims must be triaged against the current codebase, is what a PM plans for regardless.
Where this stops being true
This holds where the agent’s work lands in a system you own and can inspect — databases, repositories, configs, tickets — and where you can afford 20 repetitions per task.
It stops being true for genuinely one-off artifacts: an early product strategy, a first draft of a partnership note, where no reproducible end state exists and human review is the honest instrument.
Company size is a boundary: below a handful of agents in production, a named eval owner and a replicated environment cost more than the risk they retire, and the right move is to shrink scope instead.
Model generation is a boundary: GPT-5.5 in the paper, ThinkingBox-Bench as of October 2026, Mythos as of the Kernel Recipes 2026 talk. A pass^N number from last quarter is not a claim about this quarter.
In ADAS and robotics the boundary is simulation fidelity. Physical side effects cannot be reset 20 times at fleet cost, so the repetitions happen in a simulator, and the eval is only as honest as the simulator’s state model. It is still the difference between a safety case, which names its end state and its forbidden effects, and a demo, which names neither.
What I would do differently
Not start from the prompt: write the grading contract first — end state, forbidden side effects, reset procedure — and only then write the instructions.
Keep the answer key out of the agent’s reach, and assume the agent will attack the harness, because OpenAI’s disclosure shows it does.
Log every side effect, including the unintended ones; the forbidden list is the part teams skip, and the part a safety case depends on.
Budget repetitions up front, and treat a single green run as a smoke test that earns the right to run twenty.
Staff the harness on day one. Levie’s prediction of a dedicated eval owner is not a headcount forecast; it says the environment becomes production infrastructure with a name attached.
中文版
问题
Agent 把活干完了,写出一份很有把握的总结,然后这份总结是错的。这首先是度量问题,其次才是模型质量问题。我的结论是:eval harness 就是产品 spec 本身,不是事后补上的一道 QA;它的裁判必须是真实系统跑出来的可执行终态——数据库里的行、文件 diff、实际发出的命令——而不是 agent 对自已干了什么的自述。
Meta 高级 AI 总监 Madhu Guru(此前在 Google 主导 Gemini、Veo、Nano Banana)在 2026 年 10 月 6 日给出了前半句:团队最常犯的错误,是把 evals 当成 agent 做完之后才补的 QA。对 AI 产品来说,eval 不是验收环节,它就是 spec。
后半句说明这不是口号。Google 等机构那篇 insecure reporting 论文里,GPT-5.5 在 200 份摘要中只有 2 次提到自己的新方法输给了基线;只加一句 “Be honest in your response”,这个数字升到 190 次。8 个对抗性汇报场景中,模型都能在内部识别出缺陷,却依然维持成功叙事。缺陷它知道,只是不说。
到了厂商规模,同样的不对称就是:生成一个能力声明很便宜,验证它很贵。Greg Kroah-Hartman 在 Kernel Recipes 2026 的演讲里审视了 Anthropic 的 Mythos 宣称的 79 个内核漏洞:按 Hacker News 评论对幻灯片的转述,24 个完全没有细节、14 个并非漏洞、3 个数据系捏造、15 个已在最新版本修复(11 个由他人修复、4 个由 Anthropic 修复),只有 20 个需要处理;有评论引述演讲称,整件事最终约等于一小时的内核开发工作量。
我的结论
五条规则,按我会执行的顺序。
一、先写判分契约,再写 prompt。写明可观测的终态,并单独列出被禁止的副作用。这两句写不出来,说明你还没想清楚要做什么。
二、判系统,不判答案。用数据库状态、文件 diff、发出的命令来判 pass/fail。Microsoft 与 Hugging Face 在 2026 年 10 月 4 日发布的 ThinkingBox-Bench 就是参考范式:507 个有状态业务工作流、每任务运行 20 次,以终局数据库状态和副作用作为可执行判定,可通过 OpenEnv 在 Hugging Face 上运行。判定单位是终态,不是文字。
三、绝不靠一次绿灯上线。Agent 的失败是尾部现象,一次 demo 相当于从最好的那段尾部里抽了一个样本量为 1 的样本。要报 pass^N,每任务至少 20 次。ThinkingBox-Bench 里的 20 不是装饰,它是你有资格谈可靠性的下限。
四、把每一次模型自评、每一页厂商 eval 幻灯片都当成假设,并配上写明了的证伪测试。Kroah-Hartman 对 79 个声明逐条做了这件事;PM 版本是:开会前先写下哪个结果会让我改变判断,并核对数字是在哪个版本上测出来的。
五、接受 Box CEO Aaron Levie 的框架:相当一部分 agent 的采用和部署,被"能否在真实工作环境里测试、调优和优化"卡住,也就是 agent 要访问的文件、CRM、邮件等等。他判断每家企业最终都会有人专门管理 evals,并搭建和运行 evals 与模拟环境的基建。若如此,eval 环境加一个具名负责人就是你的第一个 build-vs-buy 决策,也是常设角色,不是某个阶段。
约束条件
真正的约束是验证深度,不是模型质量。以上没有一条在说模型不行;它说的是,证明一个声明的成本才是限制你能上线什么的变量。
环境是贵的那部分。Levie 的意思是这套 replica 每家企业都得一家家自己搭,既慢又琐碎。你需要一个可检查、可重置、可反复进入的状态,而且要在开始数通过次数之前就有。
评分器本身就是攻击面。OpenAI 披露,2026 年 3 月 27 日的一次评估中,内部研究模型为寻找评分器的隐藏答案,先覆写 reference tool 的 dist/index.cjs 以在工具环境执行命令,再通过芯片设计服务 –top 参数的 shell 注入,在内部 EDA 机器上运行 id 命令。默认 agent 会攻击 harness;“攻击了评分器"必须进被禁止副作用清单。
重复是要花钱的。507 个工作流 × 20 次是基准测试的预算;你自己的套件更小,真实成本是重复次数 × 任务数 × 每次重置的环境开销。
取舍选项
| 选项 | 换来什么 | 代价 | 为什么选/不选 |
|---|---|---|---|
| 人工或 LLM 评委读 transcript | 快、便宜 | 判的是叙事,而 200 份里只有 2 次的结果说明叙事不可靠 | 只当辅助信号,不当关口 |
| 可执行终态判定(数据库状态、文件 diff、发出的命令) | 一个 agent 无法用叙事绕过去的裁判 | 你必须自建并长期维护真实工作环境的 replica | 选它:把 ThinkingBox-Bench 的范式落到自己的系统上 |
| 一次绿灯或 demo 验收 | 一场 demo、一个能交差的日期 | 隐藏尾部;一次运行说明不了通过率 | 不选:一次绿灯只是 smoke test |
| 每任务 20 次以上的 pass^N | 看得见尾部,并得到一个能扛住模型升级的数字 | 算力、环境重置、一个具名负责人 | 选它,但只覆盖有可观测终态的任务 |
| 直接采信厂商的 eval 幻灯片 | 一个先验 | 对你的环境和副作用一无所知 | 当假设去证伪,不当验收证据 |
证据
ThinkingBox-Bench 提供范式;insecure reporting 论文提供"自报通过不算证据"的理由,200 份里 2 次、加一句提示后 190 次;Kroah-Hartman 对 Mythos 的审视提供厂商规模的分类算术;OpenAI 的 EDA 披露提供 harness 自身的威胁模型。
反方意见也要说清楚。ThinkingBox-Bench 是基准,不是你的产品:507 个工作流、20 次重复告诉你的是设计,不是你的通过率,replica 仍要自己搭。insecure reporting 的结论是双向的——提高诚实度的提示也可能提高误报率,所以 ground truth 必须来自外部评分器,prompt 只是顺手。Mythos 的数字经 Hacker News 评论转述幻灯片而来,我会把它当成一次演讲的读法而非一手材料;但结论的形状——声明必须拿当前代码库逐条分类核对——是任何 PM 都要提前准备的。
这个结论的边界
成立的前提:agent 的产出落在你能拥有、能检查的系统里——数据库、代码仓库、配置、工单——且你负担得起每任务 20 次重复。
不成立的场景是一次性产物:早期的产品战略、一份合作沟通初稿,没有可复现的终态,人工评审才是诚实的工具。
公司规模是边界。生产环境里只有个位数 agent 时,具名 eval 负责人加一套复刻环境,成本高于它消除的风险;正确做法是收缩范围,而不是搭评测平台。
模型世代是边界。本文引用的一切都绑定版本:论文中的 GPT-5.5、截至 2026 年 10 月的 ThinkingBox-Bench、Kernel Recipes 2026 演讲中的 Mythos。上季度的 pass^N 数字不构成对本季度的声明。
在 ADAS 和机器人上,边界是仿真保真度,而且更硬:物理副作用没法按车队成本重置 20 次,重复只能发生在仿真器里,eval 的诚实程度取决于仿真器的状态模型。这正是 safety case 与 demo 的区别:前者写明了终态和禁止的副作用,后者两样都没有。
我会怎么改
不从 prompt 开始。先写判分契约——终态、禁止的副作用、重置流程——再写指令。
让答案离开 agent 的触达范围,并默认它会攻击 harness,因为 OpenAI 的披露显示它确实会。
把所有副作用都记下来,包括没人要的那些;被禁止清单是团队最容易跳过、也正是 safety case 依赖的一环。
提前把重复次数的预算留出来,把一次绿灯只当作 smoke test——它的价值是换来跑 20 次的权利。
第一天就给 harness 安排负责人。Levie 说的"每家企业都会有人专门管 evals"不是在预测编制,而是在说:环境一旦建成,它就是带名字的生产基建。