K2

K² · Artificial intelligence

Show the AI one photo, and it animates the exact grip you wanted

The authors built IMAGIN-4D, a diffusion-based human-object interaction (HOI) generator that uses a reference image as a visual specification of the desired interaction snapshot.

SD
FB

Sai Kumar Dwivedi, Federica Bogo, Buğra Tekin et al.

9 authors · cs.CV

arXiv preprintArtificial intelligenceJun 2026 · ~75s read

The 30-second scan

Tell an animation program "pick up the box and carry it across the room," and it can't know how.

  1. IMAGIN-4D addresses the ambiguity in existing HOI methods, where the same text prompt and trajectory can produce different grasps, approach directions, body poses, object poses, contacts, and body-object layouts.
  2. The method decomposes image conditioning spatio-temporally, extracting supervised interaction-state tokens for spatial conditioning and computing frame-aware tokens for temporal conditioning.
  3. IMAGIN-4D uses role-aware conditioning where text, waypoints, and interaction-state tokens use separate AdaLN streams, while frame-aware visual tokens cross-attend with motion tokens.