Explore
Finding your world…
Which world was more convincing?
Rate a capability (optional)
Previous vote · models and replay
Rankings
Intervals reflect variation among voters in this fixed corpus. Provisional results need more evidence. Different continuation protocols and releases are evaluated separately.
How it works
Watch, choose a path, then compare.
Meet your worlds
We assign a scenario and two models. Read the challenge beneath the player, then press Play to watch both worlds together.
Choose each next step
After each clip, choose an action for both worlds. Playback waits for you at every step. Some scripted challenges have one continuation. The path length appears above the player; return to an earlier step in Path to try another branch.
Make your call
At the end, choose World A, World B, Tie, or Both poor. You can optionally rate a capability, or reveal the models without voting.
Keep voting
After each accepted vote, the models appear briefly and the next world starts. Previous vote keeps the last model cards and replay link. Guest voting stops after 25 votes; connected accounts have no total cap. A session ends when every available comparison has been rated. Shared replays are unranked.
Keyboard: Tab moves between controls; Enter activates a button. Space plays or pauses when you are not focused on a control. During loading, wait for the spinner to clear; use Retry if a media error appears.
About
World Model Arena asks how well a generative system preserves actions, remembered state, physical plausibility, spatial relationships, and continuity as its rollout grows longer.
One beginning, two futures
Models receive fixed conditioning images and identical action prompts. Each continuation inherits that model’s own preceding output. Branches are generated offline, then played back here. You choose a path through that saved history.
Five ways to look closer
- Action responsiveness
- Does the specified action have the expected visible consequence?
- Memory
- Do objects and places remain recognizable when you return?
- Physical plausibility
- Do gravity, contact, movement, and object permanence look believable?
- Spatial relationships
- Does the model follow directions such as left of, behind, inside, or larger than?
- Continuity
- Does the scene remain coherent as the sequence continues?
Your judgment matters
Rankings come only from votes. A regularized Bradley–Terry fit places pairwise judgments on an Elo scale; ties and both-poor judgments each count as an equal outcome. Matchups favor underrepresented models and avoid comparisons you have already rated. Twenty-five guest votes are available per browser identity; HF sign-in removes the total allowance. Repeated votes on the same model pair and path do not count again. Guest identifiers and keyed account hashes stay private. We do not store names or email addresses in voting records.
Inspect or contribute
Inputs, generated outputs, exact prompts, settings and provenance are openly inspectable. Model identities are hidden in the player until you vote or reveal them; determined visitors can identify outputs from public artifacts.
Submitters provide a versioned manifest and precomputed outputs through an HF Dataset pull request. No model-specific adapter is required. Image conditioning, complete lineage, fixed inputs, and publication rights are required. HF reviews the package and manually enables models.
These results measure a finite, initially small corpus. Visual plausibility is not proof of physical accuracy or native long-term memory.
Inspect the initial SANA camera runs. Ranked comparisons are being prepared with matching test conditions.
Citations
If you use the arena, cite the benchmark dataset and include the release and model versions you evaluated.
World Model Arena (2026). Benchmark releases, inputs, and generated rollouts. Hugging Face.
BibTeX
Ranking method: Ralph Allan Bradley and Milton E. Terry (1952). Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3–4), 324–345.
Model-specific citations and licenses are available on the model cards revealed after a match.
Arena operations
Private review, voting diagnostics, and release controls.