Iterative exploration
We introduce the task of interactive scene exploration, where a robot autonomously explores an environment and produces an Action-Conditioned Scene Graph (ACSG) that captures its underlying structure. Unlike a static reconstruction, the ACSG records both low-level geometry and semantics and high-level relationships that only become visible through physical interaction.
RoboEXP combines a Large Multimodal Model with an explicit memory design. It reasons about what to explore, how to interact, and when new evidence should update its model of the scene.