TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 2 October 2026, 01:00 UTC

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

What
Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving
Who
arxiv.org
When
1 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.38862
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints. It comes from a paper posted to arXiv on 1 October 2026. Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computational overhead and suffer from objective misalignment. Additionally, preference-based methods rely on strict pairwise annotations, limiting data utilization. To overcome these limitations, we propose EMPlan, an efficient multi-modal trajectory planning method powered by reward-guided fine-tuning. We design a hybrid architecture that combines sparse anchors with an offset refinement module for efficient multi-modal trajectory prediction. Sparse anchors provide coarse trajectory candidates with low latency, which are subsequently refined by the offset module for higher prediction accuracy. To enhance safety without incurring additional inference costs, we adopt a two-stage training paradigm consisting of pretraining and reward-guided fine-tuning. During fine-tuning, we leverage rule-based reward signals and unpaired preference supervision to refine the pretrained policy toward safer trajectory selection.

Why it counts

We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.