NTH

Mover360: Controllable Object Manipulation in 360° Panoramic Images

AuthorsHaoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun Rhee

August 26, 2026 2 min read
Watch on YouTube
The one-line take

Mover360 enables users to move, insert, or remove objects in 360° panoramas with simple point, box, or mask controls while preserving realistic geometry and scene consistency.

Key results

4B
Backbone model size

Mover360 adapts Flux2-Klein-4B-Base.

12,150
Training sequences

UE5-generated paired camera–object sequences.

60,750
Training pairs per epoch

Dynamically sampled Translation, Remove, and Insert pairs.

210
Synthetic benchmark tuples

Held-out UE5 test tuples with ground truth.

50
Real benchmark tuples

Real captured test tuples with ground truth.

95.5
UE5 Translation FID

Mover360 bbox result versus 108.9 for Insert-Anything.

What the paper found

Mover360 addresses object-level editing in 360-degree equirectangular panoramas, where wrap-around boundaries, latitude-dependent distortion, and global lighting make perspective editors unreliable. Unlike OpenAI’s GPT-Image-2, which can misjudge ERP geometry or lose context after perspective projection, Mover360 edits the full panorama natively for Translation, Remove, and reference-guided Insert. It adapts the Flux2-Klein-4B-Base rectified-flow diffusion transformer with LoRA, Qwen3 task prompts, DA2 depth conditioning, ERP-aligned circular padding, and a three-channel instruction map that supports a single target point, bounding box, or mask. A UE5 pipeline generates 12,150 paired camera–object sequences and 60,750 training pairs per epoch, while evaluation uses 210 synthetic and 50 real panorama tuples with ground truth for all three tasks. On the UE5 translation benchmark, Mover360’s bbox variant reaches an FID of 95.5, compared with 108.9 for Insert-Anything; for removal, its mask variant reaches 77.4 FID. Ablations show that depth conditioning improves geometric plausibility, especially object scale and support, while denser mask guidance improves reconstruction metrics. The framework generates a 1024 × 512 panorama in approximately 20 seconds, but remains vulnerable to large-area light sources, building-scale relocation, complex reflections, and open-ended language constraints.

Original abstract

We present Mover360, a controllable object manipulation framework for 360° images. Unlike perspective images, 360° images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis