TSR · The desk
Science · 8 October 2026
Articles in this category in 8 October 2026.
TSR Desk · science · 8 October 2026, 01:00 UTC
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMsTSR Desk · science · 8 October 2026, 01:00 UTC
CASE: Cost-Aware Stopping for Efficient Long-Video AgentsTSR Desk · science · 8 October 2026, 01:00 UTC
CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health ResearchTSR Desk · science · 8 October 2026, 01:00 UTC
COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol LayersTSR Desk · science · 8 October 2026, 01:00 UTC
BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI ClinicalTSR Desk · science · 8 October 2026, 01:00 UTC
CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI AssessmentTSR Desk · science · 8 October 2026, 01:00 UTC
Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied VideosTSR Desk · science · 8 October 2026, 01:00 UTC
ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual MemoryTSR Desk · science · 8 October 2026, 01:00 UTC
Jev in Medicine: A Benchmark EvaluationTSR Desk · science · 8 October 2026, 01:00 UTC
Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries inTSR Desk · science · 8 October 2026, 01:00 UTC
Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence inTSR Desk · science · 8 October 2026, 01:00 UTC
Best-of-$N$ Guidance for Test-time Diffusion AlignmentTSR Desk · science · 8 October 2026, 01:00 UTC
Bi-objective chance-constrained evolutionary optimization for large-scale open-pit mineTSR Desk · science · 8 October 2026, 01:00 UTC
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code MigrationTSR Desk · science · 8 October 2026, 01:00 UTC
Authority-Bound Governance of Heterogeneous AI Security Decisions in Telecom and IoT NetworksTSR Desk · science · 8 October 2026, 01:00 UTC
SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs