TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 30 September 2026, 01:00 UTC

Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study

What
Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
Who
arxiv.org
When
29 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.33843
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. It comes from a paper posted to arXiv on 29 September 2026. The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.

Why it counts

A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.