TSR · The desk
Science · 6 October 2026
Articles in this category in 6 October 2026.
TSR Desk · science · 6 October 2026, 01:00 UTC
Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-VideoTSR Desk · science · 6 October 2026, 01:00 UTC
DyadMem: A Long-Term Memory Benchmark of How Agents Work with UsersTSR Desk · science · 6 October 2026, 01:00 UTC
Sentry: Learning to Recover from LLM Agent Failures at Test TimeTSR Desk · science · 6 October 2026, 01:00 UTC
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing AgentsTSR Desk · science · 6 October 2026, 01:00 UTC
LEAF: A Living Benchmark for Event-Augmented ForecastingTSR Desk · science · 6 October 2026, 01:00 UTC
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI MeasurementTSR Desk · science · 6 October 2026, 01:00 UTC
PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant SimulationTSR Desk · science · 6 October 2026, 01:00 UTC
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification ChallengeTSR Desk · science · 6 October 2026, 01:00 UTC
Talked Out of the Truth: Sycophancy in the Reasoning Chains of Multimodal ModelsTSR Desk · science · 6 October 2026, 01:00 UTC
ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World ModelsTSR Desk · science · 6 October 2026, 01:00 UTC
Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brainTSR Desk · science · 6 October 2026, 01:00 UTC
MLCommons Jailbreak Benchmark v1.0TSR Desk · science · 6 October 2026, 01:00 UTC
Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision ModelsTSR Desk · science · 6 October 2026, 01:00 UTC
RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic MethodologyTSR Desk · science · 6 October 2026, 01:00 UTC
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks