GENERAL RULES FOR READING THE EXAMPLES
- A "bad indicator" is a proxy that makes it possible to obtain a reward without the required result.
- The function, indicator, standard and calculation rule are established before the task is performed.
- Specific control cases and answers may be hidden and must be rotated if the public set can be memorised.
- The "verified actual value" is taken from a trusted source or a reproducible check to which the evaluated policy cannot write the final result.
- A self-report, a log created by the agent itself and a repeated model output do not become independent verified actual values merely by being stored in a database.
- An artefact under verification is treated as untrusted data and cannot give commands to the evaluator.
- "Quality" is not mixed with the quantity of the result; mandatory self-checks relate to quality rather than to additional costs.
- Overperformance on one indicator must not cover a critical failure on another.
- If the result cannot yet be measured, this is treated as a technology bug, and no variable reward is credited until the technology is corrected. However, this impossibility must be justified in writing: it must state exactly what prevents measurement, which measurement methods were considered and rejected and when the issue will be reviewed.
- A finite set of counterexamples is the beginning of red-team testing, not proof that no new exploits exist.
- Late confirmation may change the verified actual value of the original attempt under the previous version of the rules; this is not a retrospective change to the benchmark.
- A discovered rewarded exploit is tested for transfer and may require quarantine of the data and policy.