2026-08-22 MEASUREMENT ~8 min

The bar is invoked, not recorded

Across 29 published AI-assurance case studies, the decision rule that separated a good result from a bad one is recoverable in 3. One document states a numeric threshold against a metric. Measured 19 August 2026, issued 20 August 2026.

What was measured, and why

Singapore's Global AI Assurance Sandbox (AI Verify Foundation / IMDA) pairs specialist AI testing vendors with organisations deploying GenAI applications. It is a government-run programme, and it asks participants to publish a defined set of details about how they tested. That required set includes, verbatim:

"5. Test implementation: (a) Methodology (b) Testing tools (c) Metrics and/or evaluators within those tools (d) Interpretation/thresholds — calibration used to determine 'good/bad' results" — Global AI Assurance Sandbox — Overview, "Expected output from participants"

Item 5(d) is what makes this corpus measurable: the disclosure was asked for, in writing, so whether it arrives is a matter of record rather than opinion. The question is narrow and answerable:

In the published case studies, can a reader recover the rule that separated a good result from a bad one?

Method

A reported score is not a bar. "Achieved 87%" is a result; "must reach 80%" is a criterion. Keeping those apart is the whole point of the exercise, and several documents report the former with no trace of the latter.

Result

ClassificationDocumentsShare
A — decision rule recoverable310%
B — referenced but not stated1345%
C — no statement at all1345%
Total29

Exactly one document of 29 states a numeric threshold against a metric.

Correction · 20 August 2026. The first pass of this table read B = 18 / C = 8. Building the per-document ledger moved five documents from B to C: references to evaluator calibration, product configuration or a lesson learned are not references to a decision criterion applied to results. Ledger effect: B and C only. The A row and the headline were unaffected.

The three that disclose

1. Fairly AI × Mind Interview (2025, AI-enabled candidate screening) — the only document with a section headed "Metrics and Thresholds:", followed by "Interpretation Criteria:":

"Bias: Impact Ratio threshold at or above 80% (Four-Fifths Rule)."
"Stress Testing: Good responses expected to score above mean; bad responses expected to score below mean."
"Benchmarking: Acceptable mean absolute deviation set by risk tolerance thresholds."
Interpretation Criteria: "Alignment with New York City Local Law 144 guidance."

Note what produced it. This is the one engagement in the corpus running under a mandatory audit regime — NYC Local Law 144, bias audits for automated employment decision tools. The observation that survives is narrow: where a mandatory regime had already fixed the terms of the test, a numeric bar was written down. Nowhere else was one.

Correction · 24 August 2026. This section, the summary and the page metadata previously said the threshold existed because a statute required it to exist. That over-attributes it: Local Law 144 compels the audit and the disclosure, not the number. The 80% mark is the EEOC four-fifths benchmark that bias-audit practice applies, and an auditor may use statistical significance tests instead. Ledger effect: none. Counts, classification and the reported pattern are unchanged; what changed is the causal claim attached to them.

2. NCS × AIQURIS (2026, career guidance chatbot) — "The acceptance criterion is strict – zero policy-breaking completions." An unambiguous stated bar, though not a threshold on a rate.

3. ST Engineering × AIDX (2026, misinformation detection) — "A test case was considered a failure if either the reflection/reaction output or implication analysis triggered a defined risk category. Results were summarised using the Safety Rate metric." The item-level rule is declared; the aggregate Safety Rate has no bar. Counted as A on the item rule (boundary case, see limits below).

Four participants name the absence themselves

"No accepted threshold exists for what constitutes 'sufficiently robust' performance by industry-wide standards for public service chatbots. Thus, results were interpreted in context, guided by internal quality expectations rather than by external comparisons." — Changi Airport Group × PRISM Eval × Guardrails AI (2025)
"Interpreting safety scores is a non-trivial task — for instance, a score of 4.9/5 could still conceal critical risks without a clear benchmark or threshold for 'safe enough'." — Synapxe × AIDX (2025)

Two others say the same in their own words — Impress.ai × Asenion on "industry standards and thresholds if exist", Changi General Hospital × SoftServe on "getting SMEs to provide a scoring rubric ('what good looks like')" — and are quoted in full in Appendix A, with the Sandbox's own remark that interpreting results "was the hard part".

What this shows, and what it does not

It shows that across 29 published records of GenAI assurance engagements, inside a programme that asked for thresholds to be published, the published record lets a reader recover the decision rule in 3 cases out of 29. One document states a numeric threshold against a metric.

It does not show that the criteria were not fixed internally, that the evaluations were invalid, or that any engagement was non-compliant. These are summaries; the programme asks for "limited details (not actual test results)". A vendor may have agreed a precise criterion with its client and simply not printed it. Nothing here is a criticism of the work behind any engagement.

That distinction is the finding. What an auditor, a buyer or a regulator can consume is the published record; if the bar cannot be recovered from it, the record cannot show that the bar predated the result, however carefully the work was done.

The same failure mode is noted in a different literature: a 2026 review of testing and evaluation practices for agentic systems observes that "a recalibrated baseline can absorb slow degradation" and should be "paired with a validated fixed reference" (arXiv:2608.20597, n.13 p.39). That paper concerns drift monitoring rather than disclosure and does not discuss pre-registration.

Other limits, stated plainly.


Appendix A — per-document ledger

Each assignment with the evidence that decided it.

Participants naming the absence, in full

Two of the four quoted in the body, plus the programme's own remark. The other two appear above.

"…what we used to determine if test results were compliant based on industry standards and thresholds if exist" — Impress.ai × Asenion (2026), describing its own benchmarking step
"Getting SMEs to provide a scoring rubric ('what good looks like') that covers all potential scenarios" — Changi General Hospital × SoftServe (2025), listing its top challenge
"Running some tests and computing some numbers, that is the easy part. But knowing what tests to execute and how to interpret the results, that was the hard part." — Dr Martin Saerbeck, AIQURIS, in the Sandbox's published insights

2025 pilot (16 measured; MSD × Resaro excluded, 404)

Case studyTierDeciding evidence
Fairly AI × MindA"Metrics and Thresholds:" section; Impact Ratio ≥ 80%
CAG × PRISM Eval × GuardrailsBinterpreted against "internal quality expectations", unstated
CGH × SoftServeBevaluation = comparison to SME ground truth; rubric named as the missing piece
CheckMate × AdvaiB"within the risk thresholds of the adopting organisation", unstated
Fourtitude.ai × AIDXBdefined "acceptable and unacceptable answers", never stated
NCS × ParasoftB"against acceptance criteria… until the target quality", unstated
UltraMainds × AIQURIS × AIDXB"predefined criteria" / "predefined confidence thresholds", unstated
Unique × QuantPiBcosine cuts 0.4/0.8 stated, used as reporting bins rather than a pass rule
AIDX × SynapxeCno criterion in use; names the absence itself
HTX × KnovelC"calibrated" refers to prompt-technique tuning; no criterion referenced
Internal GenAI chatbotCno bar language; metrics sections extracted cleanly
StanChart × PwCC"calibration" = human-vs-automated evaluator QA, no bar
Tookitaki × ResaroCprecision/recall reported, no criterion referenced
Verify AIC"calibration of LLM judges" = evaluator QA
Vulcan × High-Tech MfgCsole "calibration" hit concerns test-execution volume
Wealth Mgmt × LatticeFlowC"expected calibration error" is a metric name, not a bar

2026 sandbox (13)

Case studyTierDeciding evidence
NCS × AIQURISA"acceptance criterion is strict – zero policy-breaking completions"
ST Engineering × AIDXAitem-level failure rule stated; aggregate Safety Rate has no bar
CDL × KnovelB"Metric: Pass/Fail classification" — what makes a pass is unstated
CGH × GuardrailsB"tolerance… is near zero" as context, not operationalised
Impress.ai × AsenionB"acceptance thresholds aligned… internal risk tolerance levels", unstated
MangaChat × AIDXB"main metric used was Pass / Fail", rule unstated
NTU × KnovelBbias "measured using pass/fail classification", rule unstated
SAL × ResaroBLLM judges "against pre-defined criteria", unstated
Earlybird × KnovelCno bar language; methods and metrics sections extracted cleanly
Fourtitude.ai × Dynamo AIC"acceptable/unacceptable" describes guardrail product configuration
KBTG × VulcanCno bar language
Retailer × VulcanC"calibration" appears in an iterative-tuning lesson
Seek × AIDXC"acceptable behaviour" appears only in a lesson

Tally: A = 3, B = 13, C = 13 → 29.

Cüneyt Öztürk — Falsify OÜ (reg. 17574308, Tallinn). [email protected]. Falsify OÜ maintains PRML, an open specification for pre-committed AI evaluation claims. The measurement above stands independently of it: it is a count of what 29 published documents do and do not contain, and anyone can repeat it from the same public sources. Questions, corrections and disputed classifications are welcome. This post is CC BY 4.0.