Skip to content
Back to work

Benchmark · Video understanding · 2026

Video-MME-Logical

A benchmark for evaluating logical and temporal reasoning in video-capable multimodal models.

Video-MME-Logical tests whether multimodal models can understand what changes over time in a video, using controlled tasks such as object counting, tracking, occlusion, rotation, and transformation. It helps researchers diagnose temporal-reasoning failures and evaluate models for applications such as video assistants, embodied agents, robotics, and long-video understanding.

BenchmarkVideo reasoningMultimodal LLMTemporal logic
Benchmark construction pipeline
Benchmark construction pipeline
Magic transformation task
3D maze reasoning task
Interactive challenge

Play the video and track the cup hiding the ball. The question appears when the shuffle ends.

Object tracking with cups