Public submission · 9 September 2026

Input on NIST AI 300-1 ipd: one proposed subfield, 6.5-(*).4

NIST's AI Standards Zero Drafts pilot invited input on Guidance and Templates for Public-Facing AI Documentation (NIST AI 300-1, initial public draft, July 2026) until 16 September 2026. This is what we sent, verbatim.

Context

The default AI model profile in Annex A records what was measured (6.5-(*)), on which intended use (.1), under distributional shift (.2) and how stable the results are (.3). No field records whether the criteria used to judge fitness were fixed before the reported results were produced. Our four comments propose one optional subfield and three consequential edits. The proposed text is vendor-neutral; PRML is named once, in the rationale of comment 3, as one example of an existing format.

Sent by email to [email protected] on 9 September 2026 with the comment table attached as CSV. Submissions become part of NIST's public record. Disclosure, as the draft requests: an AI assistant was used to help locate references and draft portions of the table; the author reviewed the cited material and takes responsibility for the substance.

The four comments

#TypeSectionCommentProposed change
1TechnicalAnnex A, Table A.3.1 (Default AI model profile), field 6.5-(*) "Evaluation Metrics and Results" and subfields 6.5-(*).1–.3The profile records what was measured (6.5-(*)), on which intended use (.1), under distributional shift or adversarial conditions (.2), and how stable or reproducible the results are (.3). It has no field indicating whether the criteria used to judge fitness were fixed before the reported results were produced. Stability and confidence intervals do not establish that ordering: a threshold selected after observing results can still yield stable results. For readers using public documentation to assess model suitability, it is useful to distinguish a pre-specified acceptance criterion from one selected after evaluation. The draft already provides version identifiers and cryptographic manifests for datasets (Table A.2.1, 1.5.1–1.5.2) and version identifiers and signatures for models (Table A.3.1, 1.6.1–1.6.2); an optional criteria-commitment reference could bind the evaluation criteria to those identified artifacts.Add an optional subfield to Table A.3.1: 6.5-(*).4 Pre-specified Evaluation Criteria. Description: "Indication of whether the criteria used to judge the model's fitness for the intended use — including metrics and, where applicable, comparators, thresholds, evaluation-dataset identity and randomisation seeds — were fixed before the reported results were produced, and, where available, a reference to a commitment record supporting that ordering (for example, a cryptographic digest of the criteria together with an independent timestamp or transparency-log entry)." Designation: Optional. Additional guidance: "Where populated, a commitment record should identify the evaluation dataset version (see Table A.2.1, 1.5.1–1.5.2) and the model version (see 1.6.1) to which the criteria applied. Where timing is asserted without an independently verifiable commitment record, the artifact should indicate that the timing is provider-attested rather than independently verifiable."
2TechnicalClause 5.3, Model documentation template, field 6 "Evaluation" (description beginning "may include descriptions of the evaluation protocols and datasets…")The template-level description lists protocols, datasets, results and qualitative analysis, but not the acceptance criteria against which results were judged, nor the version of those criteria. Without that, an artifact can be "fresh" in the sense of 4.2.5 while silently changing the bar between versions.Amend the description of field 6 to read: "…may include descriptions of the evaluation protocols and datasets, the acceptance criteria (metrics, thresholds and comparison rules) against which results were judged and the version of those criteria, documented analyses of risks and potential impacts, quantitative evaluation results, and qualitative analysis."
3TechnicalClause 4.2.6 Artifact interoperability (lines 430–432) and Clause 6 ProfilesIf subfield 6.5-(*).4 is adopted, its value is naturally machine-readable: a URI plus a digest and a timestamp. This makes the performance claims in 6.5 checkable by tools rather than by trust, consistent with the interoperability goal in 4.2.6. One example of an existing format is PRML, an open specification with the IANA-registered media type application/vnd.prml+yaml; the field should not prescribe PRML or any other particular format.Add to the guidance of 6.5-(*).4: "Where a commitment record is machine-readable, the artifact should include its identifier or URI and the digest, so that interoperable tooling can verify the ordering."
4TechnicalTable A.3.1, subfield 6.5-(*).3 "Evaluation Reliability Metrics and Results", guidance sentence "This information enables interested parties to determine whether performance claims are stable enough to support deployment decisions."The sentence implies stability is the reader's main uncertainty about a performance claim. For deployment decisions a second uncertainty is whether the bar was set before the result; the guidance could point to the proposed 6.5-(*).4 so the two are not conflated.Append: "Stability does not by itself establish that the criteria for a performance claim were fixed before the results were known; see 6.5-(*).4."

Cover email

Dear NIST AI Standards team,

Please find attached four comments from Falsify OÜ (Tallinn, Estonia) on the initial public draft of NIST AI 300-1, Guidance and Templates for Public-Facing AI Documentation. They concern a single gap in the default AI model profile: the documentation records what was measured and how stable the results are, but has no field for whether the acceptance criteria were fixed before the results were produced, or for a reference that would let a reader verify that ordering. Comment 1 proposes an optional subfield 6.5-(*).4 with replacement text; comments 2 to 4 are consequential edits.

The proposed text is vendor-neutral. Falsify maintains an open specification (PRML) for such commitment records and names it once, in the rationale of comment 3, as one example of an existing format; the proposal does not depend on it.

Disclosure, as requested in the Note to Reviewers: an AI assistant was used to help locate references and draft portions of the comment table. I reviewed the cited material and take responsibility for the substance and proposed text.

Falsify OÜ consents to this input becoming part of the public record.

Kind regards,

Sources