K² · Artificial intelligence
AI already knows how objects move across camera angles — it just wasn't listening
The paper introduces MVTrack4Gen, a motion-aware training framework that uses multi-view point tracking as additional geometric and motion supervision for camera-conditioning-only novel-view video diffusion models.
JL
JJ
JoungBin Lee, Jaewoo Jung, Jongmin Lee et al.
11 authors · cs.CV
The 30-second scan
Point a camera at someone walking, then ask an AI to re-film that same moment from a different angle it never saw.
- The task addressed is synthesizing a novel-view video from a monocular reference video along a target camera trajectory, which requires both geometric consistency and motion fidelity.
- The authors state that existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos.
- The authors state that camera-conditioning-only methods can achieve high visual quality but often struggle to preserve geometric and motion consistency.