TSR · The desk
Science · 7 October 2026
Articles in this category in 7 October 2026.
TSR Desk · science · 7 October 2026, 01:00 UTC
Odyssey: A Closed-Loop Benchmark for Long-Horizon Real-World Driving with Explicit NavigationTSR Desk · science · 7 October 2026, 01:00 UTC
FREA: A Multi-Source Expert Benchmark for Reaction Feasibility VerificationTSR Desk · science · 7 October 2026, 01:00 UTC
EvalResearchBench: Can AI Agents Design Their Own Evaluations?TSR Desk · science · 7 October 2026, 01:00 UTC
Behavioral History Outperforms Descriptions of the Person for LLM Synthetic PersonasTSR Desk · science · 7 October 2026, 01:00 UTC
Scaling Participation in Modular AI SystemsTSR Desk · science · 7 October 2026, 01:00 UTC
TAPDreamer: Transferable Adversarial Patches for World Action ModelsTSR Desk · science · 7 October 2026, 01:00 UTC
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language ModelsTSR Desk · science · 7 October 2026, 01:00 UTC
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User RequestsTSR Desk · science · 7 October 2026, 01:00 UTC
Benchmarking Psychological Dynamics in Generative AgentsTSR Desk · science · 7 October 2026, 01:00 UTC
AECG: Asymmetric Experience Consolidation and Governance In Multi-Agent SystemsTSR Desk · science · 7 October 2026, 01:00 UTC
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-RewardTSR Desk · science · 7 October 2026, 01:00 UTC
When Agent Context Goes Stale: Incoherence in Volatile Agent ContextTSR Desk · science · 7 October 2026, 01:00 UTC
TRACE: Training-time Report-guided and Clinically Ordered Concept EditingTSR Desk · science · 7 October 2026, 01:00 UTC
MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual MultipleTSR Desk · science · 7 October 2026, 01:00 UTC
Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs