TSR · The desk
Articles
TSR Desk · compute · 11 September 2026, 01:00 UTC
Compute-Bounded Security Assurance - Coverage, Verification, and Response under ResourceCompute-Bounded Security Assurance - Coverage, Verification, and Response under Resource Constraints
TSR Desk · science · 11 September 2026, 01:00 UTC
SloMoDeblur: A Large-Scale Smartphone Image Deblurring DatasetSloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset
TSR Desk · science · 11 September 2026, 01:00 UTC
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation SystemsRESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
TSR Desk · science · 11 September 2026, 01:00 UTC
RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion ModeRevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
TSR Desk · physics · 11 September 2026, 01:00 UTC
Reinforcement learning for Quantum Tiq-Taq-ToeReinforcement learning for Quantum Tiq-Taq-Toe
TSR Desk · science · 11 September 2026, 01:00 UTC
Improving 5G AI-RAN MCS Selection by Predicting RetransmissionsImproving 5G AI-RAN MCS Selection by Predicting Retransmissions
TSR Desk · compute · 11 September 2026, 01:00 UTC
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled RewardProof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
TSR Desk · science · 11 September 2026, 01:00 UTC
Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, LosslessScaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution
TSR Desk · science · 11 September 2026, 01:00 UTC
CityPlanner: A Sandbox Agent for Executable Urban PlanningCityPlanner: A Sandbox Agent for Executable Urban Planning
TSR Desk · science · 11 September 2026, 01:00 UTC
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models inFrom Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
TSR Desk · science · 11 September 2026, 01:00 UTC
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionJarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
TSR Desk · science · 11 September 2026, 01:00 UTC
GANDR: Claim Auditing for Verifiable Legal Answer GenerationGANDR: Claim Auditing for Verifiable Legal Answer Generation
TSR Desk · science · 11 September 2026, 01:00 UTC
Albedo Estimation via Latent Bridge MatchingAlbedo Estimation via Latent Bridge Matching
TSR Desk · science · 11 September 2026, 01:00 UTC
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI FrameworksAn Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks
TSR Desk · science · 11 September 2026, 01:00 UTC
Multi-Agent Agentic Graph Learning via Structural SignaturesMulti-Agent Agentic Graph Learning via Structural Signatures
TSR Desk · science · 11 September 2026, 01:00 UTC
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?
TSR Desk · science · 11 September 2026, 01:00 UTC
SpecBench: Measuring Reward Hacking in Long-Horizon Coding AgentsSpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
TSR Desk · science · 11 September 2026, 01:00 UTC
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMsFortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
TSR Desk · science · 11 September 2026, 01:00 UTC
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis AgentsSafe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
TSR Desk · science · 11 September 2026, 01:00 UTC
Can Artificial Intelligence Support Healthcare and Mental Health Through Early CyberbullyingCan Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
TSR Desk · science · 10 September 2026, 01:00 UTC
NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global CulturesNormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
TSR Desk · science · 10 September 2026, 01:00 UTC
A visual large language foundational model for medical image recognition usingA visual large language foundational model for medical image recognition using clinician-oriented social media
TSR Desk · science · 10 September 2026, 01:00 UTC
OdysSim: Building Foundation Models for Human Behavior SimulationOdysSim: Building Foundation Models for Human Behavior Simulation
TSR Desk · science · 10 September 2026, 01:00 UTC
Do Web Agents Investigate Before They Decide?Do Web Agents Investigate Before They Decide?
TSR Desk · science · 10 September 2026, 01:00 UTC
Aes3D: Aesthetic Assessment in 3D Gaussian SplattingAes3D: Aesthetic Assessment in 3D Gaussian Splatting
TSR Desk · economy · 10 September 2026, 01:00 UTC
Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on FunctionExperimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
TSR Desk · science · 10 September 2026, 01:00 UTC
XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
TSR Desk · science · 10 September 2026, 01:00 UTC
PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided SearchPAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search
TSR Desk · compute · 10 September 2026, 01:00 UTC
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous InferenceJetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
TSR Desk · science · 10 September 2026, 01:00 UTC
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMsTimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
TSR Desk · science · 10 September 2026, 01:00 UTC
EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI AssistantsEvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
TSR Desk · science · 10 September 2026, 01:00 UTC
Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted SkillWho Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
TSR Desk · science · 10 September 2026, 01:00 UTC
ARNAI: Artifact Removal Network based on Autoencoding and Inpainting for Robust Spinal ImageARNAI: Artifact Removal Network based on Autoencoding and Inpainting for Robust Spinal Image Segmentation and Measurement
TSR Desk · science · 10 September 2026, 01:00 UTC
Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-TrainingUnfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
TSR Desk · science · 10 September 2026, 01:00 UTC
Neural Symbollic Regression Using Deep Learning and Sparse ModellingNeural Symbollic Regression Using Deep Learning and Sparse Modelling
TSR Desk · science · 10 September 2026, 01:00 UTC
CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence RecordsCIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records
TSR Desk · science · 10 September 2026, 01:00 UTC
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise AgentsDI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
TSR Desk · science · 10 September 2026, 01:00 UTC
SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D GenerationSymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation
TSR Desk · science · 10 September 2026, 01:00 UTC
In-Place Instruction Following in Diffusion Language ModelsIn-Place Instruction Following in Diffusion Language Models
TSR Desk · physics · 10 September 2026, 01:00 UTC
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine ActivitiesSAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
TSR Desk · science · 9 September 2026, 07:00 UTC
When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to StatisticalWhen Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability
TSR Desk · science · 9 September 2026, 07:00 UTC
GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat RelianceGeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims
TSR Desk · science · 9 September 2026, 07:00 UTC
Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM BenchmarksBenchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
TSR Desk · science · 9 September 2026, 07:00 UTC
SWE-Test: Benchmarking LLM Vulnerability Discovery via Input PredictionSWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
TSR Desk · science · 9 September 2026, 07:00 UTC
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent EraAgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
TSR Desk · physics · 9 September 2026, 07:00 UTC
Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event CamerasEmo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras
TSR Desk · science · 9 September 2026, 07:00 UTC
ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL DevelopmentProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language
TSR Desk · science · 9 September 2026, 07:00 UTC
Explainable Deep Learning for Price-Trade Dynamics: From Black-Box Forecasts to EffectiveExplainable Deep Learning for Price-Trade Dynamics: From Black-Box Forecasts to Effective Parametric Models
TSR Desk · compute · 9 September 2026, 07:00 UTC
Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box EditingBait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks
TSR Desk · energy · 9 September 2026, 07:00 UTC
Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation ModelAssessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model
TSR Desk · physics · 9 September 2026, 07:00 UTC
Do Quantum AIs Dream in Paths? Path-Integral Slow Thinking through Grover InterferenceDo Quantum AIs Dream in Paths? Path-Integral Slow Thinking through Grover Interference
TSR Desk · science · 9 September 2026, 07:00 UTC
When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AIWhen Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems
TSR Desk · science · 9 September 2026, 07:00 UTC
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language ModelsLogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
TSR Desk · science · 9 September 2026, 07:00 UTC
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses toHow AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
TSR Desk · science · 9 September 2026, 07:00 UTC
Models That Know How Evaluations Are Designed Score SaferModels That Know How Evaluations Are Designed Score Safer
TSR Desk · science · 9 September 2026, 07:00 UTC
PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base InvariantsPiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base Invariants
TSR Desk · physics · 9 September 2026, 07:00 UTC
DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep LearningDUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning
TSR Desk · science · 9 September 2026, 07:00 UTC
Robust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A ReviewRobust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A Review
TSR Desk · compute · 9 September 2026, 07:00 UTC
AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page DocumentsAtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents
TSR Desk · science · 9 September 2026, 07:00 UTC
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM AgentsBeyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
TSR Desk · science · 8 September 2026, 01:00 UTC
EXAONE Forecast for FinanceEXAONE Forecast for Finance
TSR Desk · science · 8 September 2026, 01:00 UTC
Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language ModelsMolecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
TSR Desk · science · 8 September 2026, 01:00 UTC
A Deep Generative Model for Synthesizing Labeled Wireless SignalsA Deep Generative Model for Synthesizing Labeled Wireless Signals
TSR Desk · mathematics · 8 September 2026, 01:00 UTC
MABPD: Multi-Agent Bias Probing & Detection via Structured Argument DebateMABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate
TSR Desk · science · 8 September 2026, 01:00 UTC
PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video GenerationPRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation
TSR Desk · science · 8 September 2026, 01:00 UTC
LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
TSR Desk · science · 8 September 2026, 01:00 UTC
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following inMM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
TSR Desk · science · 8 September 2026, 01:00 UTC
Cross-modal triage network: a multimodal deep learning framework for severity-based triage andCross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
TSR Desk · science · 8 September 2026, 01:00 UTC
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool AgentsSiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
TSR Desk · science · 8 September 2026, 01:00 UTC
A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic GapA Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap
TSR Desk · science · 8 September 2026, 01:00 UTC
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive OptimizationAsk Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
TSR Desk · mathematics · 8 September 2026, 01:00 UTC
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual ContextsGSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
TSR Desk · mathematics · 8 September 2026, 01:00 UTC
AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-DimensionalAxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics
TSR Desk · science · 8 September 2026, 01:00 UTC
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language ModelsKnowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
TSR Desk · science · 8 September 2026, 01:00 UTC
Reinforcement Learning for improving Large Language Models' Catalan text simplificationReinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
TSR Desk · science · 8 September 2026, 01:00 UTC
AI Revealed PreferencesAI Revealed Preferences
TSR Desk · compute · 8 September 2026, 01:00 UTC
Substrate-Aware AI Agents: Execution Context as a First-Class InputSubstrate-Aware AI Agents: Execution Context as a First-Class Input
TSR Desk · science · 8 September 2026, 01:00 UTC
PetQA: Benchmarking Veterinary Knowledge and Clinical ReasoningPetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
TSR Desk · energy · 8 September 2026, 01:00 UTC
Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNetsPhase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets
TSR Desk · economy · 8 September 2026, 01:00 UTC
CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency MetricCPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric
TSR Desk · science · 7 September 2026, 07:00 UTC
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive MarketERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
TSR Desk · science · 7 September 2026, 07:00 UTC
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
TSR Desk · science · 7 September 2026, 07:00 UTC
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
TSR Desk · science · 7 September 2026, 07:00 UTC
A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in MultilingualA Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
TSR Desk · science · 7 September 2026, 07:00 UTC
Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-BallGlobal to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball
TSR Desk · science · 7 September 2026, 07:00 UTC
TeleTables: A Benchmark for Large Language Models in Telecom Table InterpretationTeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
TSR Desk · science · 7 September 2026, 07:00 UTC
Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer LanguageTechnical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
TSR Desk · science · 7 September 2026, 07:00 UTC
ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk ClassificationReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification
TSR Desk · science · 7 September 2026, 07:00 UTC
Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring ofAgentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
TSR Desk · science · 7 September 2026, 07:00 UTC
You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and PlayfulYou Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
TSR Desk · physics · 7 September 2026, 07:00 UTC
Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D DetectionPost Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection
TSR Desk · science · 7 September 2026, 07:00 UTC
When Does an Interpretation Count as Established? The Formation, Evaluation, and ResponsibilityWhen Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI
TSR Desk · science · 7 September 2026, 07:00 UTC
HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative RewardHLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
TSR Desk · science · 7 September 2026, 07:00 UTC
CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict MediationCC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
TSR Desk · science · 7 September 2026, 07:00 UTC
MaxKernel: Agentic Kernel Generation for TPUsMaxKernel: Agentic Kernel Generation for TPUs
TSR Desk · physics · 7 September 2026, 07:00 UTC
BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital BiomarkerBioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
TSR Desk · science · 7 September 2026, 07:00 UTC
A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a NewA Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
TSR Desk · science · 7 September 2026, 07:00 UTC
AlcaTRAz - Anchored Tree-Rule Defense Against JailbreaksAlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
TSR Desk · science · 7 September 2026, 07:00 UTC
Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis inAletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
TSR Desk · compute · 7 September 2026, 07:00 UTC
Iris: Climbing to the Search FrontierIris: Climbing to the Search Frontier