TSR · The desk
Science · 19 September 2026
Articles in this category in 19 September 2026.
TSR Desk · science · 19 September 2026, 01:00 UTC
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary RevisionTSR Desk · science · 19 September 2026, 01:00 UTC
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language ModelsTSR Desk · science · 19 September 2026, 01:00 UTC
Evaluating Large Language Models for Symbolic Security Protocol AnalysisTSR Desk · science · 19 September 2026, 01:00 UTC
TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative VisualTSR Desk · science · 19 September 2026, 01:00 UTC
DeltaSelect: Affordable A/B Testing for Coding AgentsTSR Desk · science · 19 September 2026, 01:00 UTC
AutoResearch: Insight In, Hallucination OutTSR Desk · science · 19 September 2026, 01:00 UTC
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding AgentsTSR Desk · science · 19 September 2026, 01:00 UTC
Chronicle: Cut-Point Replay for Regression Testing of LLM AgentsTSR Desk · science · 19 September 2026, 01:00 UTC
CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health PatientTSR Desk · science · 19 September 2026, 01:00 UTC
Prediction-Powered Smoothing and Validation for Disaggregated AI EvaluationTSR Desk · science · 19 September 2026, 01:00 UTC
Attributing Preprocessing Invariance in Spectral Foundation ModelsTSR Desk · science · 19 September 2026, 01:00 UTC
A Neuropsychologically Grounded Evaluation of LLM Cognitive AbilitiesTSR Desk · science · 19 September 2026, 01:00 UTC
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation inTSR Desk · science · 19 September 2026, 01:00 UTC
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, andTSR Desk · science · 19 September 2026, 01:00 UTC
Not All Nodes Are Created Equal: Homophily-Aware Stratification for Stable GNN Evaluation