K² · Artificial intelligence
Show the AI one photo, and it animates the exact grip you wanted
The authors built IMAGIN-4D, a diffusion-based human-object interaction (HOI) generator that uses a reference image as a visual specification of the desired interaction snapshot.
SD
FB
Sai Kumar Dwivedi, Federica Bogo, Buğra Tekin et al.
9 authors · cs.CV
The 30-second scan
Tell an animation program "pick up the box and carry it across the room," and it can't know how.
- IMAGIN-4D addresses the ambiguity in existing HOI methods, where the same text prompt and trajectory can produce different grasps, approach directions, body poses, object poses, contacts, and body-object layouts.
- The method decomposes image conditioning spatio-temporally, extracting supervised interaction-state tokens for spatial conditioning and computing frame-aware tokens for temporal conditioning.
- IMAGIN-4D uses role-aware conditioning where text, waypoints, and interaction-state tokens use separate AdaLN streams, while frame-aware visual tokens cross-attend with motion tokens.