Notes · falsify.dev

Short notes on ML eval rigor — in print, not in tweets.

Essays, postmortems, and field reports on pre-registration, the PRML specification, and the wider question of what it would take for ML evaluation claims to be falsifiable in practice. New posts on no schedule. RSS.

Three notes so far — the list loads from /notes/posts.json: the Lock #2 post-mortem, model cards vs pre-registration, and the v0.2 RFC overview. Essays land when there's something useful to write rather than something topical to react to.