PLAYER-GUIDED EVALUATION FOR INTERACTIVE WORLD MODELS

PlayWorld

Benchmarking World Models with Agent Players over Long-Horizon Objectives

Kaixin Ding · Xi Chen · Minghong Cai · Zhiyuan Xu · Yiyang Wang · Yuxiang Lu · Junyi Li · Shuyang Chen · Yuan Gao · Xin Tao · Pengfei Wan · Hengshuang Zhao

Kuaishou Technology · The University of Hong Kong

171Cases
1,417Videos
10–60sRollouts
9+Models
01 / CURATED MODEL ROLLOUTS

Model demos.

Geometry Consistency60.1s
02 / FULL PLAYER PROCESS

Full Player process.

Complete recordings from world creation to adaptive action execution and final result preservation.

MODEL

Genie3

Google DeepMind

LS010Full process · 2:14
OE014Full process · 2:31
MODEL

HappyOyster

Alibaba

GC008Full process · 3:53
GC010Full process · 3:09
MODEL

HY-World2

Tencent

GC022Full process · 1:00
GC033Full process · 0:59
04 / LEADERBOARD

Model ranking.

Select one capability to view its ranking.

RankModelScore
1Genie32.12
2HappyOyster1.92
3LingBot-World21.82
4LingBot-World1.78
5HY-World21.61
6SANA-WM1.48
7Hunyuan-GameCraft-21.42
8HY-WorldPlay1.21
9Matrix-Game-3.01.14

Complete nine-model ranking from the current paper results. Equal scores share a rank.

VQA evaluation across four dimensions in PlayWorld

Scores range from 1 to 5. Overall is the unweighted mean across the four evaluation dimensions.

ModelGeometry consistency ↑Interaction fidelity ↑Insight evolution ↑Out-of-sight evolution ↑Overall ↑
Closed-source Models
Genie32.742.401.511.812.12
HappyOyster2.542.151.471.541.92
Open-source Models
LingBot-World2.112.231.331.431.78
LingBot-World22.042.131.951.161.82
HY-World22.142.061.131.091.61
SANA-WM1.721.891.131.161.48
Hunyuan-GameCraft-21.621.521.211.311.42
HY-WorldPlay1.121.631.011.081.21
Matrix-Game-3.01.301.251.001.001.14