Evidence Check · free · no sign-up

Does this report show when the acceptance criteria were fixed?

Paste an evaluation report, model card, validation summary or audit extract. Five fixed checks, each answered only with text quoted from your document. No score. The result reads like the follow-up questions an examiner would ask.

This tool reads the document you provide. It does not determine what happened outside that document.

or try a sample: model card (typical) · validation summary (stronger)

We do not retain your document. The text is sent to an AI model provider for analysis. We retain only anonymous usage metadata (document length and the five labels), not the document text. Do not submit confidential or personal information unless you are comfortable with that processing.

Reading

Quotes are verbatim from your text. Where a check has no supporting text, the tool says so; it does not conclude that the evidence does not exist.

Need the timing itself to be independently checkable? PRML is an open format for committing evaluation criteria before results are known. Learn about PRML →

Have a document this reading does not fit, or a question about what an examiner would accept?

Ask a question →

How this works

Five questions are asked of every document, in the same order. Each answer carries a label, the verbatim text that supports it, a one- or two-sentence reading limited to that text, what an outside reader still cannot establish, and what kind of record would make the evidence stronger.

1. Are the acceptance criteria written down?

Metric, comparison rule, threshold, evaluation dataset. A number alone is a result, not a criterion.

2. Does the document claim the criteria were fixed before the results?

A claim ("defined prior to deployment") is recorded as a claim. The next check asks what stands behind it.

3. What kind of evidence supports the timing?

Self-asserted: a date, signature, memo or approval whose clock the provider controls. Internally traceable: a version-control reference, ticket or QMS record an insider could trace but an outsider cannot verify from the document. Independently time-evidenced: an RFC 3161 timestamp, transparency-log entry or other third-party time anchor the document identifies. The label follows what the document says the evidence is, not the name of a tool.

4. Are the criteria bound to the model version and dataset version?

Without the binding, a criteria record can be reattached to a different run later.

5. If the criteria were revised, does the document say when and why?

A revised bar with a visible history is stronger evidence than a bar that was never written down.

An AI model (Claude) reads the text under a fixed instruction; the server then discards any quote that is not a verbatim substring of your document and reports a check with no surviving quote as "no supporting text found". There is no score and no overall verdict, by design. Related reading: NAIC Exhibit C, field 11 · NIST AI 300-1, field 6.5 · Reproducible is not pre-registered. For a quick heuristic read of a single benchmark claim, the older Claim Check still works offline in your browser.