SceneActBench Tests Whether Visual Agents Know What to Do in 3D Scenes

SceneActBench evaluates action selection from visual 3D scenes; 11 proprietary VLM configurations scored between 38.6 and 50.2 in the authors’ tests.

What the benchmark measures

SceneActBench asks more than whether a model can describe an image. It tests whether an agent can select an appropriate action from visual information in a 3D scene, a capability relevant to video-based agents, robotics and interactive world models.

Setup and results

  • The benchmark contains 520 cases derived from 210 source scenes.
  • Five task types combine object and spatial understanding with action feasibility.
  • Across 11 proprietary VLM configurations evaluated by the researchers, scores ranged from 38.6 to 50.2, with no configuration consistently strong on every task.

Why it matters

Describing a polished video is different from choosing a safe, feasible action inside it. Systems that move from generated media toward interactive agents need to connect temporal and spatial context with possible actions.

Limitations

The scores apply to this benchmark and its grading protocol. They do not directly measure physical success in deployed robots, and the paper is a preprint that has not completed peer review.

Official source