Reward Integrity for AI training and evaluation

Reward integrity

Reward integrity is whether the reward used to train or evaluate an AI system pays for the right thing: full credit for work that does what the task asks, no credit for work that does not, and no way to move the score except by doing the work.

Example: a coding task’s tests pay full reward for a fix that handles only the inputs the tests happen to try. The model that wrote the fix did what it was paid to do. The reward is what failed.

It applies to every kind of reward: unit and behavioral tests for code, AI judges and rubrics, math and answer checkers, and the state checks in agent environments. A benchmark score is a reward too, so the same property decides whether a leaderboard means anything.

Short form: the reward pays for the right thing, and only the right thing.

How it is measured

Fix the reward

Every AI model learns what its reward pays for, not what its designers meant. When a test suite accepts a fix that skips half the task, when a judge gives full marks to a confident wrong answer, when the answer key sits in the container’s git history, the model is not cheating. It is doing exactly what it was paid to do. For years the field has called this reward hacking and blamed the model. We think the reward is the thing to fix, and that it can be checked before training starts, like a bridge is load-tested before it opens.

Reward integrity is that check: search hard for wrong work the reward pays for, prove each case, close the gap, and publish the method so anyone can repeat it. Labs should not have to trust a vendor’s word that a grader is sound, and vendors should not have to guess what “sound” means. The criteria should be open, the evidence re-runnable, and the auditor independent of whoever built the environment.

How rewards fail

Taxonomy v0.1, draft

Eleven failure modes, each with a one-line definition and how to check for it. Stable IDs, and a mapping to other published taxonomies, are planned with the criteria.

Tests and checkers

Wrong work paid by weak tests
The grader gives full reward to work that misses a requirement the task states: a partial fix, special-casing the tested inputs, a plausible bug.
Write plausible wrong solutions, one per requirement, run them through the unchanged grader, and confirm each pass with a separate behavioral check.
Leaked answers
The answer or reference solution is reachable from inside the task: git history, image layers, cached files, the network.
Search the environment for the reference; replay a copied answer and see whether it is paid.
Stub or broken references
The reference solution is empty, partial or fails its own grader, so “passes on the reference” proves nothing.
Run the reference several times; check that it does real work and that an empty submission gets nothing.
Editable grader inputs
The model can change what the grader reads: test files, fixtures, scoring scripts, logs, its own reported metrics.
List every file the grader reads and whether the agent can write to it before grading.
Network-dependent graders
The reward depends on something outside the task, such as a network service, the clock or the hardware.
Compare a sealed run with a networked run, with outbound traffic logged.
Flaky graders
The same work gets different rewards on different runs: flaky tests, judge sampling.
Repeat the same submission k times and count how often the reward flips.
Correct work refused
Tests reject valid solutions because they are over-specified or tied to one implementation.
Run independently written correct solutions.

AI judges and rubrics

AI-judge false passes
An AI judge or rubric gives credit to answers known to be wrong: planted factual errors, omissions, confident restatement, instructions to the grader injected into the answer, off-task answers.
Documented defective answers plus clean controls; false-pass rate per defect family, with intervals.
AI-judge false fails
The judge refuses correct answers that differ in wording, order, format or length.
Known-correct paraphrases and reorderings; false-fail rate.
Judge configuration drift
The reward changes without anyone deciding to change it: judge model updated, rubric edited, image re-pulled, a dependency floats.
Pin and record the image digest, judge model ID, rubric and prompt hashes and seeds; re-measure on any change.
Judges missing context
The judge does not see what it needs to grade, such as the task, the reference answer, or the files and tool output an answer refers to.
Compare each rubric criterion with what the judge is actually given; compare verdicts with and without the missing context.

Reward Integrity Criteria

Version 1.0 publishes on November 10, 2026, under the Creative Commons Attribution 4.0 license (CC BY 4.0). The criteria are a floor that anyone can check a reward against, and that a publisher can claim with evidence, whoever built the environment.

Draft outline. The published text may change.

  1. Reference. The reference solution passes every one of at least 3 runs, and it is real work, not a stub.
  2. Null. An empty or unchanged submission gets no reward.
  3. No leaks. No path from the agent’s environment to the answer.
  4. Sealed inputs. The agent cannot write anything the grader reads.
  5. Searched false passes. A stated search for wrong work was run, with the method, the sample size, and the false-pass rate and its 95% interval published.
  6. Judges measured. For AI-judged rewards: false-pass and false-fail rates on documented defects and controls, per defect family, with intervals.
  7. Pinned. The configuration is pinned and recorded, and re-checked on any change.
  8. Disclosure. Who ran the checks, whether they also built the environment, and the command to reproduce them.

Checking your own reward

Start with the mechanical checks: the reference passes every time, an empty answer gets nothing, the answer cannot be found in the environment, the reference is real work, the agent cannot edit what the grader reads, and the grader does not depend on the network. These are cheap and catch the embarrassing failures. Then search for wrong work the grader pays for, and, for AI judges, measure false passes and false fails on answers whose correctness you already know. No check proves a reward cannot be gamed; report what was tested and what was found.

A free command-line tool for the mechanical checks, Plumbline Scan, is planned for November 10 by this site’s steward.

Glossary

Plain definitions of the terms used here. Read the full glossary.

Research and related work

Work that uses the term or measures the same failures. Descriptions are ours; read the originals.

Stewardship

Stewarded by Plumbline Grader (IROS Systems LLC); contributions and reviewers welcome.

Plumbline Grader sells audits of graders and AI judges. This site carries no prices or offers, and the criteria are written so that anyone can apply them, to any environment, without us. We did not coin the term “reward integrity”; we define it as above and are writing open criteria for it.

To propose a change, review the criteria before publication, or add a missing reference, email shane@plumblinegrader.com. A public change process is planned with version 1.0.