Skip to content

Add PhysicalAgent #6

Description

@keon

Paper Details

Summary

A general framework for cognitive robotics using a Perceive → Plan → Reason → Act pipeline. Perception and reasoning use VLMs for structured, constraint-aware plans; action generation uses a diffusion-based foundation world model to synthesize future trajectory videos; a lightweight embodiment-specific policy maps video actions to motor commands. Evaluated on mobile navigation and long-horizon manipulation, achieving up to 83% success on human-familiar tasks. Open-sourced, embodiment-agnostic design bridges foundation-model reasoning and physical execution.

Suggested Section

Embodied World Models or Reasoning & Planning

Checklist

  • Add paper entry to README
  • Verify links are working

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions