2026-06-30 · Video-MME-Logical
Video-MME-Logical: reasoning in moving worlds
Why controlled video tasks reveal failures that broad video benchmarks can hide.
The question
A model can appear capable on a broad video benchmark while relying on recognition, language priors, or static cues. Video-MME-Logical asks a narrower question: can the model preserve and manipulate state as a scene changes over time?
The design
The benchmark uses controllable scenes and task families such as counting, occlusion, tracking, rotation, and spatial transformation. Difficulty can be varied while the underlying rule remains inspectable.
What is public
The public repository links the paper, dataset, evaluation code, compact leaderboard, representative figures, and reproduction guidance. This note summarizes those materials; it does not introduce claims beyond them.