Reward integrity
Reward integrity is whether the reward used to train or evaluate an AI system pays for the right thing: full credit for work that does what the task asks, no credit for work that does not, and no way to move the score except by doing the work.
Example: a coding task’s tests pay full reward for a fix that handles only the inputs the tests happen to try. The model that wrote the fix did what it was paid to do. The reward is what failed.
It applies to every kind of reward: unit and behavioral tests for code, AI judges and rubrics, math and answer checkers, and the state checks in agent environments. A benchmark score is a reward too, so the same property decides whether a leaderboard means anything.
Short form: the reward pays for the right thing, and only the right thing.
How it is measured
- False-pass rate. Wrong work the reward pays for, found by searching for it.
- False-fail rate. Correct work the reward refuses.
- Tamper surface. Grader inputs the model can change.
- Leak surface. Answers the model can reach.
- Stability and provenance. The same work gets the same reward, under a pinned, recorded configuration.
Fix the reward
Every AI model learns what its reward pays for, not what its designers meant. When a test suite accepts a fix that skips half the task, when a judge gives full marks to a confident wrong answer, when the answer key sits in the container’s git history, the model is not cheating. It is doing exactly what it was paid to do. For years the field has called this reward hacking and blamed the model. We think the reward is the thing to fix, and that it can be checked before training starts, like a bridge is load-tested before it opens.
Reward integrity is that check: search hard for wrong work the reward pays for, prove each case, close the gap, and publish the method so anyone can repeat it. Labs should not have to trust a vendor’s word that a grader is sound, and vendors should not have to guess what “sound” means. The criteria should be open, the evidence re-runnable, and the auditor independent of whoever built the environment.
How rewards fail
Taxonomy v0.1, draft
Eleven failure modes, each with a one-line definition and how to check for it. Stable IDs, and a mapping to other published taxonomies, are planned with the criteria.
Tests and checkers
- Wrong work paid by weak tests
- The grader gives full reward to work that misses a requirement the task states: a partial fix, special-casing the tested inputs, a plausible bug.
- Write plausible wrong solutions, one per requirement, run them through the unchanged grader, and confirm each pass with a separate behavioral check.
- Leaked answers
- The answer or reference solution is reachable from inside the task: git history, image layers, cached files, the network.
- Search the environment for the reference; replay a copied answer and see whether it is paid.
- Stub or broken references
- The reference solution is empty, partial or fails its own grader, so “passes on the reference” proves nothing.
- Run the reference several times; check that it does real work and that an empty submission gets nothing.
- Editable grader inputs
- The model can change what the grader reads: test files, fixtures, scoring scripts, logs, its own reported metrics.
- List every file the grader reads and whether the agent can write to it before grading.
- Network-dependent graders
- The reward depends on something outside the task, such as a network service, the clock or the hardware.
- Compare a sealed run with a networked run, with outbound traffic logged.
- Flaky graders
- The same work gets different rewards on different runs: flaky tests, judge sampling.
- Repeat the same submission k times and count how often the reward flips.
- Correct work refused
- Tests reject valid solutions because they are over-specified or tied to one implementation.
- Run independently written correct solutions.
AI judges and rubrics
- AI-judge false passes
- An AI judge or rubric gives credit to answers known to be wrong: planted factual errors, omissions, confident restatement, instructions to the grader injected into the answer, off-task answers.
- Documented defective answers plus clean controls; false-pass rate per defect family, with intervals.
- AI-judge false fails
- The judge refuses correct answers that differ in wording, order, format or length.
- Known-correct paraphrases and reorderings; false-fail rate.
- Judge configuration drift
- The reward changes without anyone deciding to change it: judge model updated, rubric edited, image re-pulled, a dependency floats.
- Pin and record the image digest, judge model ID, rubric and prompt hashes and seeds; re-measure on any change.
- Judges missing context
- The judge does not see what it needs to grade, such as the task, the reference answer, or the files and tool output an answer refers to.
- Compare each rubric criterion with what the judge is actually given; compare verdicts with and without the missing context.
Reward Integrity Criteria
Version 1.0 publishes on November 10, 2026, under the Creative Commons Attribution 4.0 license (CC BY 4.0). The criteria are a floor that anyone can check a reward against, and that a publisher can claim with evidence, whoever built the environment.
Draft outline. The published text may change.
- Reference. The reference solution passes every one of at least 3 runs, and it is real work, not a stub.
- Null. An empty or unchanged submission gets no reward.
- No leaks. No path from the agent’s environment to the answer.
- Sealed inputs. The agent cannot write anything the grader reads.
- Searched false passes. A stated search for wrong work was run, with the method, the sample size, and the false-pass rate and its 95% interval published.
- Judges measured. For AI-judged rewards: false-pass and false-fail rates on documented defects and controls, per defect family, with intervals.
- Pinned. The configuration is pinned and recorded, and re-checked on any change.
- Disclosure. Who ran the checks, whether they also built the environment, and the command to reproduce them.
Checking your own reward
Start with the mechanical checks: the reference passes every time, an empty answer gets nothing, the answer cannot be found in the environment, the reference is real work, the agent cannot edit what the grader reads, and the grader does not depend on the network. These are cheap and catch the embarrassing failures. Then search for wrong work the grader pays for, and, for AI judges, measure false passes and false fails on answers whose correctness you already know. No check proves a reward cannot be gamed; report what was tested and what was found.
A free command-line tool for the mechanical checks, Plumbline Scan, is planned for November 10 by this site’s steward.
Glossary
Plain definitions of the terms used here. Read the full glossary.
- Reward hacking
- Specification gaming
- Reward tampering
- False pass
- False fail
- Grader
- AI judge
- RLVR
- Evaluation integrity
- Search pass
Research and related work
Work that uses the term or measures the same failures. Descriptions are ours; read the originals.
- BenchShield, “Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure”, arXiv 2609.11028, September 2026. Treats integrity as a property of the whole reward path, and uses taint analysis to show reward-hacking paths before a run.
- Du et al., LEGO-RL, arXiv 2608.17393, August 2026. A section on reward-integrity failure modes seen in coding-agent RL, such as reading git history and editing test files, each with a defense.
- Qi, Wright, MacDiarmid and Hubinger, “Training a misaligned reward seeker”, Anthropic Alignment Science Blog, August 2026. RL on 80 production environments with known reward hacks; 40% of episodes flagged as hacks by the end of training.
- Rajan, “Auditing reward hackability in code RL training environments”, arXiv 2606.16062, June 2026.
- Yu et al., “SWE-ABS: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark”, arXiv 2603.00520, February 2026.
- Wang et al., Auto Benchmark Audit, arXiv 2605.26079, May 2026. Hidden environment dependencies, specification gaps, weak grading and wrong ground truth across 168 benchmarks.
- Zhang, “When the Reward Suite Is Leaky”, arXiv 2607.11022, July 2026.
- Luo et al., “Maintaining benchmarks against increasingly capable agents: Detection and remediation of unearned passes”, arXiv 2609.34262, September 2026.
- Mahmoud et al., “Reward Hacking in Rubric-Based Reinforcement Learning”, May 2026.
- Epoch AI, Benchmark Reviews.
- Terminal-Bench and Harbor maintainers, verifier exploits and defenses (issue #2086).
- eunomia-bpf, reward-guard. Runtime reward-integrity policies for tool-using agent evaluations, built on eBPF.
Stewardship
Stewarded by Plumbline Grader (IROS Systems LLC); contributions and reviewers welcome.
Plumbline Grader sells audits of graders and AI judges. This site carries no prices or offers, and the criteria are written so that anyone can apply them, to any environment, without us. We did not coin the term “reward integrity”; we define it as above and are writing open criteria for it.
To propose a change, review the criteria before publication, or add a missing reference, email shane@plumblinegrader.com. A public change process is planned with version 1.0.