Benchmark · Video understanding · 2026
Video-MME-Logical
A benchmark for evaluating logical and temporal reasoning in video-capable multimodal models.
Video-MME-Logical tests whether multimodal models can understand what changes over time in a video, using controlled tasks such as object counting, tracking, occlusion, rotation, and transformation. It helps researchers diagnose temporal-reasoning failures and evaluate models for applications such as video assistants, embodied agents, robotics, and long-video understanding.
BenchmarkVideo reasoningMultimodal LLMTemporal logic
Project media
Read the project note ↗
Interactive challenge
Play the video and track the cup hiding the ball. The question appears when the shuffle ends.