DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.
Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu
The study shows that video models may retain the right physical understanding even after generating the wrong motion—and that targeted internal writes can bring that knowledge back.
Bi-FlowGS lets generated videos and 3D Gaussian scenes iteratively correct each other so sparse-view reconstructions become both more visually complete and geometrically faithful.
Peiyu Liu, Dingxi Zhang, Federico Tombari, Marc Pollefeys, Christina Tsalicoglou, Daniel Barath
SplashSplat reconstructs fleeting real-world liquid splashes in 3D by combining multi-camera masks, coarse fluid motion, and dynamic Gaussian primitives.
AdaptVPR uses controlled generative scene changes and verification to create harder same-place images, making visual place recognition more robust to weather, lighting, and occlusions.
Ref-GeNVS helps generative vision systems create realistic new viewpoints of mirror-containing scenes by explicitly using reflections as additional virtual camera views.
Scal3R keeps online 3D reconstruction stable over long videos by querying poses against multiple past keyframes instead of relying on a single distant reference.
The study shows that claims about which vision-learning rules best match the brain can change dramatically depending on the image resolution used for evaluation.
Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun, Lihan Zhang, Bin Liang, Kam-Fai Wong, Haibin Huang, Chi Zhang, Xuelong Li
RefVideo-6M is a massive, quality-controlled dataset that teaches video-editing models to follow both instructions and visual references more reliably.
Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
GeoNeXt repurposes video generation models to learn depth and surface geometry jointly, achieving strong zero-shot results with far less training data.
4DAnyone turns an ordinary monocular human video into a consistent, renderable 4D reconstruction by coordinating diffusion-generated views across time and camera angles.
Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun Rhee
Mover360 enables users to move, insert, or remove objects in 360° panoramas with simple point, box, or mask controls while preserving realistic geometry and scene consistency.
This study shows that tomatoes, potatoes, and onions can help train face anti-spoofing systems by teaching detectors to recognize presentation artifacts rather than facial identity.
A physics-aware data pipeline and fast video diffusion model remove distracting glass reflections while introducing a benchmark for measuring progress.
Self-Geometry improves 3D vision foundation models at test time by using camera and multi-view geometry as self-supervision without requiring ground-truth labels.
Leonid Kuturin, Ilya Sotnikov, Mark Khusnutdinov, Mikhail Potemkin, Pavel Baranas, Aleksandra Korepanova, Alexander Kalashnikov
It shows that AI-image detectors may learn file-compression fingerprints instead of synthesis—and that attractive heatmaps are not automatically faithful explanations.
Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
ID-V2V edits a video's style and scene while preserving the people, expressions, gaze, and lip movements that make the original performance recognizable.
VidMap turns difficult, uncalibrated videos into reliable 3D camera trajectories by combining the best ideas from SLAM, SfM, temporal matching, and monocular depth.
Mykola Lavreniuk, Nataliia Kussul, Andrii Shelestov, Yevhenii Salii, Volodymyr Kuzin, Charlotte Julia Li-Xing Wang, Zoltan Szantoi
A globally scalable computer vision foundation model maps agricultural field boundaries across countries with major accuracy gains and practical nationwide deployment speed.
Yufei Cai, Xuesong Niu, Hao Lu, Kun Gai, Kai Wu, Guosheng Lin
MetaView uses diffusion, implicit geometry, and metric depth to generate spatially consistent novel views from a single image—even under large camera movements.
Zanyi Wang, Xin Lin, Haodong Li, Dengyang Jiang, Yijiang Li
This work shows that text-to-image models can do dense vision better by reading out depth, masks, and other per-pixel signals directly from their internal patch grid instead of forcing everything back into RGB.
Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
This paper argues that video generation models can double as powerful general vision learners, enabling a single pretrained model to tackle diverse perception tasks with strong data efficiency.
PixWorld unifies 3D scene reconstruction and generation in pixel space, adding geometry-aware supervision to improve both visual quality and 3D fidelity.
SCAIL-2 is a new end-to-end character animation system that learns motion transfer directly from video, using synthetic training data and preference tuning to produce more faithful animations.
Amir Mann, Gal Michael Harari, Merav Keidar, Or Litany
VideoMDM teaches a 3D human motion generator using only 2D keypoints from videos, getting surprisingly close to fully 3D-supervised models without needing 3D ground truth.
Goku introduces a million-scale dataset and benchmark for instruction-based video editing, plus a model that better handles both appearance changes and structural motion edits.
Sanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim
This paper turns dense cell segmentation from hundreds of clicks into one click per cell type by chaining prompts through SAM's built-in feature structure.
Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee, Eunbyung Park, Seon Joo Kim
This paper turns 3D scenes into object-level tokens instead of raw point clouds or Gaussians, enabling reconstruction, segmentation, editing, and retrieval from unposed images.
LAS2 shows that stereo matching can stay fast enough for real devices while still improving zero-shot accuracy, thanks to a smarter architecture and a carefully staged training recipe.
This work shows that video diffusion models can do more than generate clips—they can also reconstruct detailed two-hand motion from egocentric video surprisingly well.
Orest Kupyn, Goutam Bhat, Philipp Henzler, Fabian Manhardt, Christian Rupprecht, Federico Tombari
FLAT turns a single image into an explorable 3D scene by directly decoding surface triangles from diffusion latents, aiming for better geometry than Gaussian-based methods.
Koichi Namekata, Yash Kant, Zhizheng Liu, Ryan D Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu
This paper lets creators control video generation by tracking points across reference images and frames, enabling more precise compositing and camera motion in a single model.
Hongkai Zheng, Ta-Ying Cheng, Benjamin Klein, Yisong Yue, Zhuoning Yuan
Vera edits videos by generating only the changed layer and blending it back with the original, aiming to preserve everything that should stay the same while still allowing strong visual edits.
Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong
Memento improves long video generation by teaching the model to reconstruct recurring subjects from memory, helping characters stay consistent across shots and scenes.
This work proposes a new way to restore and sharpen AI-generated images using the original reference image, recovering lost fine details while fixing generation artifacts in one step.
Meixi Song, Dizhe Zhang, Hao Ren, Ruiyang Zhang, Bo Du, Ming-Hsuan Yang, Lu Qi
UniSHARP lets a single monocular view-synthesis system render sharp scenes across perspective, fisheye, and panoramic cameras by aligning them in a shared omnidirectional latent space.
Feng Qiao, Zhaochong An, Zhexiao Xiong, Serge Belongie, Nathan Jacobs
Track2View lets a video diffusion model re-render scenes from new camera paths by using paired 3D point tracks to keep motion and appearance consistent across time.
Alexander Veicht, Sunghwan Hong, Dániel Baráth, Marc Pollefeys
ZipSplat makes fast 3D scene reconstruction more flexible by clustering visual tokens into fewer, smarter Gaussians, cutting representation cost while improving quality.
Qixin Hu, Shuai Yang, Wei Huang, Song Han, Yukang Chen
LongLive-RAG turns previously generated video latents into a searchable memory, helping autoregressive models keep long videos more consistent instead of drifting over time.
Ulrich Prestel, Stefan Andreas Baumann, Nick Stracke, Björn Ommer
RayDer turns self-supervised novel view synthesis from brittle multi-module training into a single scalable transformer that learns from real-world video and performs strongly zero-shot on many benchmarks.
Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, Muhammet Esat Kalfaoglu
YOLO26 is a faster, cleaner evolution of the YOLO family that aims to deliver real-time detection and related vision tasks with better accuracy-latency tradeoffs and simpler deployment.
Leonhard Sommer, Artur Jesslen, Basavaraj Sunagad, Adam Kortylewski
This paper makes 3D object understanding more fine-grained by teaching a model to infer consistent parts and shapes across object categories from a single image, while also releasing a new large benchmark to measure that skill.
CubePart lets users generate 3D objects from text while specifying the parts they want, making AI-generated assets more useful for games, animation, and simulation.
This paper introduces a large-scale diffusion model that can generate and edit editable layered images, making layered visual composition much faster and more practical.
This paper shows how to train image and video editors without paired before-and-after data by bootstrapping guidance from the model itself and using flow matching to preserve structure.
Xiangtao Kong, Jixin Zhao, Lingchen Sun, Rongyuan Wu, Lei Zhang
This paper uses multimodal foundation models to fabricate realistic ground-truth images from degraded photos, creating a 100K-pair dataset that helps restoration models work better on real-world images.
Qi Xu, Yue Tan, Shihao Chen, Jiahao Meng, Anna Wang, Shunping Ji, Hao Fei, Jason Li
This paper tackles a harder version of video grounding where one text query can match multiple separate video segments, and it builds a benchmark plus training method to make multimodal models much better at finding all of them.
HorizonStream is a new streaming Transformer that helps 3D reconstruction stay stable over extremely long video sequences by separating short-range matching from long-range geometric memory.
This paper turns video editing into a fast, few-step generation problem, using streaming diffusion-style models and clever attention mechanisms to make edits both quicker and more controllable.