DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
AuthorsMerav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
AffiliationsOct 2026 Azimuth° Doppler Azimuth° Doppler Lateral Sensor Shift High Res Sensor Range [m] Range [m] Azimuth° Doppler Azimuth° Doppler Figure 1: Radar re-simulation with DyRAD. From recorded radar · Technion Cornell Tech NVIDIA Ground Truth Repositioned Object Range [m] Range [m]
Resources
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
Key results
DyRAD’s correlation, compared with 0.068 for the strongest baseline
Reference-detected objects recovered by DyRAD, versus 26.9% for the strongest baseline
Lateral-shift distance in the real-data round-trip evaluation, measured in meters
Direct rendering F1, compared with 0.101 for linear upsampling
Milliseconds to render a full RAD tensor on an NVIDIA RTX 4090
What the paper found
DyRAD addresses a central weakness in radar novel-view synthesis: prior methods either reconstruct range–azimuth measurements without motion cues or render Doppler only for static scenes. It represents a driving environment as static background reflectors and motion-tracked dynamic point reflectors, derives each reflector’s radial velocity from object tracks and sensor motion, and differentiably renders complete range–azimuth–Doppler tensors. A fixed, analytic point-spread function derived from the radar signal-processing chain prevents measurement blur from being mistaken for scene geometry, enabling rendering from displaced viewpoints and transfer to new sensor configurations without refitting. Evaluated on RADIal, Boreas, and a synthetic benchmark, DyRAD raises full-RAD correlation on RADIal from 0.068 for the strongest baseline to 0.272 and recovers 90.7% of reference-detected objects versus 26.9%. In real-data off-path testing after a 2 m lateral-shift round trip, it reaches a 91.7% foreground hit rate. Coarse-to-fine sensor transfer improves detection F1 from 0.101 with linear upsampling to 0.195 through direct PSF-aware rendering. The CUDA implementation runs on an NVIDIA RTX 4090, taking 7.2 minutes for 10k optimization steps and 5.65 ms to render a full RAD tensor, although the method still relies on object annotations and a ground-plane representation.
Original abstract
Reconstructing dynamic driving scenes from recorded sensor data supports closed-loop evaluation of autonomous driving systems by synthesizing observations beyond the original trajectory. Unlike cameras and LiDAR, radar measures radial velocity directly through Doppler. Yet existing radar novel-view synthesis fails to exploit this capability: methods addressing dynamic scenes reconstruct only range-azimuth tensors, while methods that render Doppler assume static scenes. Moreover, because radar processing spreads each reflection across multiple bins, existing representations absorb this spread into scene geometry, causing it to render incorrectly when the viewpoint moves. We present DyRAD, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors. Reflector velocities are derived from object tracks and projected onto the line of sight, making Doppler both a rendered output and supervision for those tracks. Crucially, we render reflectors through a fixed analytic point-spread function (PSF) derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation. Beyond improving scene reconstruction, this separation also enables zero-shot sensor-configuration transfer, allowing the same reconstructed scene to be rendered under different radar specifications without refitting. We evaluate DyRAD on RADIal, Boreas, and a synthetic benchmark across both on-path poses and displaced viewpoints untested by prior work. On RADIal, DyRAD recovers radar detections in 90.7% of reference-detected objects, compared with 26.9% for the strongest baseline.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.
A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
The study shows that video models may retain the right physical understanding even after generating the wrong motion—and that targeted internal writes can bring that knowledge back.