Mover360: Controllable Object Manipulation in 360° Panoramic Images
AuthorsHaoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun Rhee
Resources
Mover360 enables users to move, insert, or remove objects in 360° panoramas with simple point, box, or mask controls while preserving realistic geometry and scene consistency.
Key results
Mover360 adapts Flux2-Klein-4B-Base.
UE5-generated paired camera–object sequences.
Dynamically sampled Translation, Remove, and Insert pairs.
Held-out UE5 test tuples with ground truth.
Real captured test tuples with ground truth.
Mover360 bbox result versus 108.9 for Insert-Anything.
What the paper found
Mover360 addresses object-level editing in 360-degree equirectangular panoramas, where wrap-around boundaries, latitude-dependent distortion, and global lighting make perspective editors unreliable. Unlike OpenAI’s GPT-Image-2, which can misjudge ERP geometry or lose context after perspective projection, Mover360 edits the full panorama natively for Translation, Remove, and reference-guided Insert. It adapts the Flux2-Klein-4B-Base rectified-flow diffusion transformer with LoRA, Qwen3 task prompts, DA2 depth conditioning, ERP-aligned circular padding, and a three-channel instruction map that supports a single target point, bounding box, or mask. A UE5 pipeline generates 12,150 paired camera–object sequences and 60,750 training pairs per epoch, while evaluation uses 210 synthetic and 50 real panorama tuples with ground truth for all three tasks. On the UE5 translation benchmark, Mover360’s bbox variant reaches an FID of 95.5, compared with 108.9 for Insert-Anything; for removal, its mask variant reaches 77.4 FID. Ablations show that depth conditioning improves geometric plausibility, especially object scale and support, while denser mask guidance improves reconstruction metrics. The framework generates a 1024 × 512 panorama in approximately 20 seconds, but remains vulnerable to large-area light sources, building-scale relocation, complex reflections, and open-ended language constraints.
Original abstract
We present Mover360, a controllable object manipulation framework for 360° images. Unlike perspective images, 360° images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.