NTH
Research collection

Computer Vision research

Research on interpreting images and video, including recognition, geometry, and visual understanding. Compare methods through their reported evidence.

58 papers · Latest edition October 2, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Computer Vision papers

Newest editions first.

02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis
08Cv

Reflection-aware Generative Novel View Synthesis

GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

Ref-GeNVS helps generative vision systems create realistic new viewpoints of mirror-containing scenes by explicitly using reflections as additional virtual camera views.

Read analysis
12Cv

Video Generative Models as Geometry Learner

Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng

GeoNeXt repurposes video generation models to learn depth and surface geometry jointly, achieving strong zero-shot results with far less training data.

Read analysis
13Cv

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

Yudong Jin, Tao Xie, Qihang Zhang, Zehong Shen, Zhen Xu, Yujun Shen, Hujun Bao, Xiaowei Zhou, Yinghao Xu

4DAnyone turns an ordinary monocular human video into a consistent, renderable 4D reconstruction by coordinating diffusion-generated views across time and camera angles.

Read analysis
20Cv

Explaining AI-Image Detection: What the Heatmap Actually Shows

Leonid Kuturin, Ilya Sotnikov, Mark Khusnutdinov, Mikhail Potemkin, Pavel Baranas, Aleksandra Korepanova, Alexander Kalashnikov

It shows that AI-image detectors may learn file-compression fingerprints instead of synthesis—and that attractive heatmaps are not automatically faithful explanations.

Read analysis
22Cv

ID-V2V: Identity-Preserving Video Restylization

Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu

ID-V2V edits a video's style and scene while preserving the people, expressions, gaze, and lip movements that make the original performance recognizable.

Read analysis
25Cv

Delineate Anything v2: A Global Foundation Model for Field Delineation

Mykola Lavreniuk, Nataliia Kussul, Andrii Shelestov, Yevhenii Salii, Volodymyr Kuzin, Charlotte Julia Li-Xing Wang, Zoltan Szantoi

A globally scalable computer vision foundation model maps agricultural field boundaries across countries with major accuracy gains and practical nationwide deployment speed.

Read analysis
27Cv

Unified Video Dense Prediction from Disjoint Data

Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee

UniD uses diffusion-informed distillation to teach one video model eight scene-understanding tasks from separate datasets.

Read analysis
30Cv

Video Generation Models are General-Purpose Vision Learners

Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu

This paper argues that video generation models can double as powerful general vision learners, enabling a single pretrained model to tackle diverse perception tasks with strong data efficiency.

Read analysis
37Cv

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

Junpeng Jing, Ronglai Zuo, Zhelun Shen, Shangchen Zhou, Rolandos Alexandros Potamias, Stefanos Zafeiriou, Krystian Mikolajczyk, Jiankang Deng

LAS2 shows that stereo matching can stay fast enough for real devices while still improving zero-shot accuracy, thanks to a smarter architecture and a carefully staged training recipe.

Read analysis
40Cv

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

Koichi Namekata, Yash Kant, Zhizheng Liu, Ryan D Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu

This paper lets creators control video generation by tracking points across reference images and frames, enabling more precise compositing and camera motion in a single model.

Read analysis
41Cv

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Hongkai Zheng, Ta-Ying Cheng, Benjamin Klein, Yisong Yue, Zhuoning Yuan

Vera edits videos by generating only the changed layer and blending it back with the original, aiming to preserve everything that should stay the same while still allowing strong visual edits.

Read analysis
42Cv

Memento: Reconstruct to Remember for Consistent Long Video Generation

Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong

Memento improves long video generation by teaching the model to reconstruct recurring subjects from memory, helping characters stay consistent across shots and scenes.

Read analysis
44Cv

UniSHARP: Universal Sharp Monocular View Synthesis

Meixi Song, Dizhe Zhang, Hao Ren, Ruiyang Zhang, Bo Du, Ming-Hsuan Yang, Lu Qi

UniSHARP lets a single monocular view-synthesis system render sharp scenes across perspective, fisheye, and panoramic cameras by aligning them in a shared omnidirectional latent space.

Read analysis
46Cv

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang, Yi Hua, Ben Mildenhall, Christoph Lassner, Narendra Ahuja, Gengshan Yang

World Tracing is a new way to turn images into 3D by keeping pixels aligned while also hallucinating the hidden parts of objects and scenes.

Read analysis
47Cv

ZipSplat: Fewer Gaussians, Better Splats

Alexander Veicht, Sunghwan Hong, Dániel Baráth, Marc Pollefeys

ZipSplat makes fast 3D scene reconstruction more flexible by clustering visual tokens into fewer, smarter Gaussians, cutting representation cost while improving quality.

Read analysis
50Cv

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, Muhammet Esat Kalfaoglu

YOLO26 is a faster, cleaner evolution of the YOLO family that aims to deliver real-time detection and related vision tasks with better accuracy-latency tradeoffs and simpler deployment.

Read analysis
51Cv

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

Leonhard Sommer, Artur Jesslen, Basavaraj Sunagad, Adam Kortylewski

This paper makes 3D object understanding more fine-grained by teaching a model to infer consistent parts and shapes across object categories from a single image, while also releasing a new large benchmark to measure that skill.

Read analysis
52Cv

CubePart: An Open-Vocabulary Part-Controllable 3D Generator

Yiheng Zhu, Kangle Deng, Jean-Philippe Fauconnier, Inaki Navarro, Daiqing Li, Ava Pun, Yinan Zhang, Peiye Zhuang, Xiaoxia Sun, Maneesh Agrawala, Kiran Bhat, Tinghui Zhou

CubePart lets users generate 3D objects from text while specifying the parts they want, making AI-generated assets more useful for games, animation, and simulation.

Read analysis
56Cv

Towards One-to-Many Temporal Grounding

Qi Xu, Yue Tan, Shihao Chen, Jiahao Meng, Anna Wang, Shunping Ji, Hao Fei, Jason Li

This paper tackles a harder version of video grounding where one text query can match multiple separate video segments, and it builds a benchmark plus training method to make multimodal models much better at finding all of them.

Read analysis
57Cv

HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction

Chong Cheng, Peilin Tao, Nanjie Yao, Guanzhi Ding, Xianda Chen, Yuansen Du, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Zhengqing Chen, Hao Wang

HorizonStream is a new streaming Transformer that helps 3D reconstruction stay stable over extremely long video sequences by separating short-range matching from long-range geometric memory.

Read analysis