Paper Details
Summary
A general framework for cognitive robotics using a Perceive → Plan → Reason → Act pipeline. Perception and reasoning use VLMs for structured, constraint-aware plans; action generation uses a diffusion-based foundation world model to synthesize future trajectory videos; a lightweight embodiment-specific policy maps video actions to motor commands. Evaluated on mobile navigation and long-horizon manipulation, achieving up to 83% success on human-familiar tasks. Open-sourced, embodiment-agnostic design bridges foundation-model reasoning and physical execution.
Suggested Section
Embodied World Models or Reasoning & Planning
Checklist
Paper Details
Summary
A general framework for cognitive robotics using a Perceive → Plan → Reason → Act pipeline. Perception and reasoning use VLMs for structured, constraint-aware plans; action generation uses a diffusion-based foundation world model to synthesize future trajectory videos; a lightweight embodiment-specific policy maps video actions to motor commands. Evaluated on mobile navigation and long-horizon manipulation, achieving up to 83% success on human-familiar tasks. Open-sourced, embodiment-agnostic design bridges foundation-model reasoning and physical execution.
Suggested Section
Embodied World Models or Reasoning & Planning
Checklist