TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 14 September 2026, 07:00 UTC

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

What
When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline
Who
arxiv.org
When
14 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.12017
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. It comes from a paper posted to arXiv on 14 September 2026. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes and 27 failures despite zero gating failures, because known gaps are explicitly exempted from the gate. An identifier-free export represents 8,843 tool-attempt rows: 8,395 recorded durations and 448 missing values. Of the durations, 121 equal the signed 32-bit maximum and carry abandoned-client labels; inspected database code clamps elapsed lifecycle age. The pooled recorded 99th percentile is 2,147,483,647 ms, versus 38,118.31 ms among server-observed completed calls. This is a stratum contrast, not a treatment effect. In a documented single-trajectory compaction pilot, the reported follow-up input reduction is 94.39%, but the reduction across the trigger and follow-up calls together is 46.54%. We reproduce the descriptive calculations, verify 91 timing statistics through a separate weighted rational-arithmetic implementation, and execute 13 scoring-function and 12 analysis-verifier tests. Finite-completion bounds show how missing durations limit all-row timing statements without imputing values. The contribution is a source-linked case study and reusable verification package for separating gate policy, lifecycle timing, and request-level accounting from broader agent-performance claims. Historical provider runs and the full current pipeline were not independently reproduced; general capability superiority and population-level statistical significance are not established.

Why it counts

We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.