TSR · The desk
Science · 5 October 2026
Articles in this category in 5 October 2026.
TSR Desk · science · 5 October 2026, 07:00 UTC
SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large LanguageTSR Desk · science · 5 October 2026, 07:00 UTC
Deep learning-based prediction of time-resolved adhesive forces in viscoelastic HertzianTSR Desk · science · 5 October 2026, 07:00 UTC
ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code GenerationTSR Desk · science · 5 October 2026, 07:00 UTC
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest VerificationTSR Desk · science · 5 October 2026, 07:00 UTC
Rethinking the Evaluation of Harness Evolution for AgentsTSR Desk · science · 5 October 2026, 07:00 UTC
A Multi-Timescale Recursive Self-Improvement Engine for Open-Ended Persona GrowthTSR Desk · science · 5 October 2026, 07:00 UTC
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining UnderTSR Desk · science · 5 October 2026, 07:00 UTC
Evaluating the Retrieval Robustness of Large Language ModelsTSR Desk · science · 5 October 2026, 07:00 UTC
$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-TrainingTSR Desk · science · 5 October 2026, 07:00 UTC
Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPGTSR Desk · science · 5 October 2026, 07:00 UTC
A Language Model from 1913: Pretraining on Historical TextTSR Desk · science · 5 October 2026, 07:00 UTC
WAMpy: Efficient Synthesis of Prolog Programs in PythonTSR Desk · science · 5 October 2026, 07:00 UTC
LoGo: Local-Global Rewards for Consistent Long-Horizon Video GenerationTSR Desk · science · 5 October 2026, 07:00 UTC
GUI Agents for Continual Game Generation