Liangyang Ouyang1,2, Ruicong Liu2, Xuangeng Chu2, Kaipeng Zhang2, Yoichi Sato1
1The University of Tokyo 2Alaya Lab
HelloWorld is a video world model that enables social interaction with in-world characters. With a single button press (F), users can prompt the on-screen character to respond toward the camera, e.g., turning to the viewer, waving, nodding, or speaking a short greeting, while maintaining high-quality scene and camera-trajectory reconstruction.
- Self-distillation training: the base video generation model is finetuned on data synthesized by itself, containing both social interactions and camera motion, so it learns camera-pose conditioning without degrading interaction quality.
- Training-free temporal control: at inference, a temporal cross-attention mask localizes the character's response to the
F-press window. - HelloWorldBench: a 400-sample benchmark with three social interaction metrics (ActAcc, TimeAcc, GazeDev) alongside three conventional metrics.
Coming soon.
For questions, please contact oyly@iis.u-tokyo.ac.jp or liangyang.ouyang@shanda.com.
