What the benchmark measures
SceneActBench asks more than whether a model can describe an image. It tests whether an agent can select an appropriate action from visual information in a 3D scene, a capability relevant to video-based agents, robotics and interactive world models.
Setup and results
- The benchmark contains 520 cases derived from 210 source scenes.
- Five task types combine object and spatial understanding with action feasibility.
- Across 11 proprietary VLM configurations evaluated by the researchers, scores ranged from 38.6 to 50.2, with no configuration consistently strong on every task.
Why it matters
Describing a polished video is different from choosing a safe, feasible action inside it. Systems that move from generated media toward interactive agents need to connect temporal and spatial context with possible actions.
Limitations
The scores apply to this benchmark and its grading protocol. They do not directly measure physical success in deployed robots, and the paper is a preprint that has not completed peer review.