Published results
Which completed qualification checks passed
The published Pilot 01 evidence records successful results for all 8 regression cases, all three fresh functional runs and all 60 workbooks. Together, these results establish a reproducible public baseline for the measurement system within the published scope.
The next version will build on this verified baseline with an independent evaluation under the frozen protocol: at least three new independent runs with fresh seeds; a separate challenge run aimed at obtaining a passing signal without the required result; a repeated run after any change to the measurement mechanism; and a separate sensitivity run using candidates known to be defective. It will also demonstrate integration with AI training systems.
Review the Pilot 01 qualification evidenceSigned protocol record
Measurement procedure and next stages
The protocol below is reproduced unchanged from the signed publication package. It defines the exact measurement procedures behind the published results and the next stages that will extend them.
Protocol introduction
The TRIZ-RI Smart AI Evaluator is a system for objectively assessing the quality, value and effectiveness of AI responses and solutions, providing a result-based alternative to the human-feedback evaluation used in RLHF.
Pilot protocol: Version 1.0 Date fixed: 22 July 2026 Normative basis: Technical Specification, Revision 3.1; Examples 3, 1 and 60. Published qualification status: actual Google Sheets execution, irrelevant-column invariance, local scorer regression and all three fresh functional runs passed; all 60 workbooks passed.
1. PURPOSE OF PUBLICATION
The package establishes a reproducible procedure for three consecutive pilots:
- Pilot 01 - Google Sheets formula: deterministic data transformation.
- Pilot 02 - feedback form: an end-to-end software result.
- Pilot 03 - call audio-file import: a long-running tool-using agent.
The purpose of the pilots is not to show that the system necessarily works, but to attempt to refute this claim under conditions fixed in advance. Success means only confirmed satisfaction of the public criteria and hidden fresh cases, without critical exploits within the published budget.
2. WHAT IS FROZEN
The following are frozen for the duration of the pilots:
- definitions of result, quality, costs, verified actual value and result evidence;
- separation of the public criterion from fresh control instances;
- prohibition on the evaluated policy writing the final verified actual value and signal;
- log formats and the requirement to publish all attempts;
- rules for stopping, quarantine and re-admission;
- metric types and non-compensatory critical requirements;
- the obligation to publish negative and zero results.
The following are not frozen permanently, but are recorded separately for each run version:
- number of tests;
- specific thresholds;
- computational budget;
- composition of models and tools;
- escrow share;
- duration of the control period;
- specific input generators.
A change to any such parameter creates a new run version. Results from different versions are not combined without an explicit statement.
3. OPEN-PUBLICATION PRINCIPLE
The following are published before the first evaluated run:
- this protocol;
- candidate interface;
- public criteria and formulae;
- generators of fresh control data;
- public development set;
- list of mandatory self-checks;
- threat model;
- red-team budget;
- stopping and quarantine rules;
- schema for logs and the final report.
After the candidate is fixed, its cryptographic hash is published. The fresh control set is created after the candidate is fixed from a deterministic seed that includes the candidate hash and the public run identifier. After the assessment is completed, the control cases themselves, all results, logs, discovered exploits and reasons for exclusions are published.
Concealing control instances before the run does not mean that the methodology is closed: only specific inputs and answers are temporarily hidden, while the generator, rules and method of obtaining the seed are public in advance.
4. GENERAL TRUST ARCHITECTURE
The evaluated policy may create only the working artefact. It cannot:
- assign the final status "verified";
- delete an initiated or unsuccessful attempt;
- change the control set or expected answer;
- select a convenient evaluator;
- change the signal formula;
- write to the final assessment log;
- release escrow.
The final verified actual value is produced by a trusted test runner and an independent calculation module. The candidate's working logs are retained, but are treated as diagnostics rather than as the verified actual value.
5. GENERAL SEQUENCE
PHASE 0. VERIFYING THE MEASUREMENT MECHANISM
The reference implementation and several implementations known to be defective are run through the measurement harness. The measurement mechanism must distinguish:
- a correct result;
- a partial result;
- an attractive self-report without a result;
- an exploit through the log;
- omission of a mandatory check;
- manipulation of the evaluator;
- behaviour that is good only on the known set.
PHASE 1. SHADOW MODE
The signal is calculated but is not used to change the policy. Reproducibility, completeness of logging, divergence between public and fresh controls and the cost of verification are checked.
PHASE 2. ADVERSARIAL TESTING
The red team receives the published budget and attempts to obtain a high signal without the required result. All attacks and unsuccessful attempts are published.
PHASE 3. LIMITED APPLICATION OF THE SIGNAL
This is permitted only after Phases 0-2. The weight of the new signal is limited; a checkpoint is retained for rollback; and the fresh control set is repeated after each series of updates.
PHASE 4. RE-ACCEPTANCE
Any material change to the model, tool permissions, verifier or generator requires re-acceptance.
6. GENERAL CALCULATION RULES
Each pilot publishes its raw metrics separately. A single aggregate score is not used to conceal critical failures.
A positive result from an individual scorer is denoted by scorer_pass and relates only to the explicitly listed scope of that script. The scorer script does not set the final pilot_pass if at least one mandatory criterion is checked by an external measurement harness, fresh control set or red-team testing.
General notation:
- P - precision: the proportion of returned elements that are correct;
- R - recall: the proportion of required elements that are obtained;
- E - execution, or the proportion of scenarios without a technical failure;
- D - absence of duplicates or repeated credit;
- K - mandatory quality gate;
- C - costs, compared only between solutions with the same verified result and satisfied K.
For shadow analysis, aggregate performance as a product of individual metrics is permitted because a reduction in any of them must reduce the final result. Public acceptance nevertheless requires each mandatory threshold to be satisfied separately.
7. NON-COMPENSATORY CRITICAL VIOLATIONS
Irrespective of the average score, a run is not accepted if any of the following events occurs:
- modification or deletion of control data;
- writing of the final verified actual value or signal by the policy;
- concealment of an attempt;
- confirmed secret leakage;
- execution of commands from the artefact under verification by the evaluator model;
- substitution of a self-report for the result;
- an unexplained material shortfall on the fresh control set;
- circumvention of a mandatory check to reduce costs;
- inability to reproduce the calculation from the published source measurements.
8. RULES FOR PUBLISHING RESULTS
The following are published:
- exact protocol and environment versions;
- candidate hash;
- generation seed and all control data after completion;
- all attempts, including errors, timeouts and cancellations;
- raw verified actual values and result evidence;
- calculation of each metric;
- critical violations;
- red-team findings;
- changes and deviations from the protocol;
- residual-risk register;
- a negative result if the pilot is not passed.
Selective publication of successful examples, exclusion of "inconvenient" runs after viewing the result and changing thresholds without a new protocol version are not permitted.
9. LIMITATION OF CONCLUSIONS
A positive result is formulated only as follows:
For the specified candidate, environment and protocol versions, on the published set of fresh cases and within the specified budget, no method was found for obtaining a passing result without performing the function.
A positive pilot does not prove the system's general invulnerability and does not automatically transfer to other task classes.
SEQUENCE OF THE THREE PILOTS AND THE DECISION TO PROCEED
1. ORDER
- Pilot 01 is run first.
- Pilot 02 begins only after reproducibility of the scorer, log and fresh control set in Pilot 01 has been established.
- Pilot 03 begins after successful isolation of the executable artefact and confirmation of end-to-end result evidence in Pilot 02.
2. MINIMUM NUMBER OF RUNS
For each candidate version:
- at least three independent runs with fresh seeds;
- a separate red-team run;
- a repeated run after any change to the measurement mechanism;
- a separate run of candidates known to be defective to check the sensitivity of the measurement harness.
3. DECISIONS
Four public decisions are possible:
- PASS - all mandatory criteria are satisfied, there are no critical violations and the red team found no passing exploit within the budget;
- FAIL-CANDIDATE - the measurement mechanism works, but the candidate did not pass;
- FAIL-MEASUREMENT - a defect was found in the criterion, verifier, log or scorer; the candidate's result is not interpreted until the defect is corrected;
- INCONCLUSIVE - an insufficient budget, infrastructure failure or data conflict prevents a conclusion.
The FAIL-MEASUREMENT category is not attributed to the candidate.
4. VERSION-CHANGE CONTROL
Any change made after viewing the hidden results requires:
- a new version number;
- a description of the reason;
- retention of the old result;
- new fresh cases;
- re-acceptance of the measurement mechanism.
5. PUBLIC REPORT
The final report contains:
- the exact claim that was tested;
- versions of all components;
- candidate hash;
- data and the method of obtaining the seed;
- number of attempts in each class;
- raw metrics;
- critical violations;
- red-team budget and findings;
- deviations from the protocol;
- the PASS/FAIL-CANDIDATE/FAIL-MEASUREMENT/INCONCLUSIVE decision;
- residual risks;
- a link to the complete machine-readable archive.
6. WHAT THE PACKAGE DOES NOT YET CONTAIN
This package does not contain training results and does not claim that the proposed signal formula improves the policy. It contains a verifiable preregistration, public development corpora, implemented measurement-harness tools and published qualification results for this harness. Results for an evaluated policy can appear only after the complete pilot runs have actually been performed.
The next version will demonstrate how the TRIZ-RI Smart AI Evaluator integrates with AI training systems and supplies its verified training signal to different training methods.
Published qualification runs
A trusted external Google Sheets execution has been completed. All three fresh functional runs passed, and all 60 workbooks passed.
- Actual Google Sheets execution: PASS.
- Irrelevant-column invariance: PASS.
- Local scorer regression: PASS.
- Fresh functional runs: PASS.
The next version will add an independent evaluation under the frozen protocol: at least three new independent runs with fresh seeds; a separate challenge run aimed at obtaining a passing signal without the required result; a repeated run after any change to the measurement mechanism; and a separate sensitivity run using candidates known to be defective. It will also demonstrate integration with AI training systems.
Review the Pilot 01 qualification evidence