TSR · The desk
Science · 10 October 2026
Articles in this category in 10 October 2026.
TSR Desk · science · 10 October 2026, 01:00 UTC
Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent SystemsTSR Desk · science · 10 October 2026, 01:00 UTC
A 3D Characterization Framework for Intelligent Sequential Decision MakingTSR Desk · science · 10 October 2026, 01:00 UTC
Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive RiggingTSR Desk · science · 10 October 2026, 01:00 UTC
Humanize: Judgement Engineering for Agentic CodingTSR Desk · science · 10 October 2026, 01:00 UTC
MemTrial: Learning When to Trust Memory in LLM Portfolio AgentsTSR Desk · science · 10 October 2026, 01:00 UTC
RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact ReasoningTSR Desk · science · 10 October 2026, 01:00 UTC
Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot andTSR Desk · science · 10 October 2026, 01:00 UTC
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug FixesTSR Desk · science · 10 October 2026, 01:00 UTC
MultiWorldBench: Do Independently Controlled Views Describe One Shared World?TSR Desk · science · 10 October 2026, 01:00 UTC
ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable DocumentTSR Desk · science · 10 October 2026, 01:00 UTC
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running GameTSR Desk · science · 10 October 2026, 01:00 UTC
Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, BenchmarkTSR Desk · science · 10 October 2026, 01:00 UTC
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksTSR Desk · science · 10 October 2026, 01:00 UTC
\$OneMillion-Bench: How Far are Language Agents from Human Experts?TSR Desk · science · 10 October 2026, 01:00 UTC
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?TSR Desk · science · 10 October 2026, 01:00 UTC
An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and DataTSR Desk · science · 10 October 2026, 01:00 UTC
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs