Skip to content

Commit a7061ea

Browse files
xjxu21claude
andcommitted
Restructure page: drop Abstract, bullet-point Why WorldMark
- Remove Abstract section; its content now drives Why WorldMark as two problem bullets (✗) and four solution bullets (✓) - Hero stat cards expanded into detailed lists (models, case breakdown, metric families) - Action Suite: trajectory and VLM-selection figures side by side - Capability radar shown at full width - Remove Qualitative Examples section Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1 parent 83ee7f7 commit a7061ea

3 files changed

Lines changed: 84 additions & 62 deletions

File tree

docs/index.html

Lines changed: 84 additions & 62 deletions
Original file line numberDiff line numberDiff line change
@@ -99,6 +99,10 @@
9999
figcaption { font-size: 13.5px; color: var(--ink-faint); margin-top: 10px; text-align: center; line-height: 1.55; max-width: 720px; margin-left: auto; margin-right: auto; }
100100
.figframe { border: 1px solid var(--line); border-radius: var(--radius); padding: 14px; background: #fff; max-width: 50%; margin: 0 auto; }
101101
@media (max-width: 860px) { .figframe { max-width: 100%; } }
102+
.fig-row { display: flex; gap: 18px; align-items: flex-start; margin: 26px 0 8px; }
103+
.fig-row figure { flex: 1; margin: 0; min-width: 0; }
104+
.fig-row .figframe { max-width: 100%; }
105+
@media (max-width: 860px) { .fig-row { flex-direction: column; } }
102106

103107
/* ── Hero ────────────────────────────────── */
104108
.hero { text-align: center; padding: 64px 0 30px; }
@@ -126,13 +130,34 @@
126130
.btn.arena { background: linear-gradient(120deg, #5b2ee5, #9333ea, #d946ef); }
127131

128132
/* ── Stats strip ─────────────────────────── */
129-
.stats { display: flex; justify-content: center; gap: 8px; flex-wrap: wrap; margin: 44px auto 0; max-width: 900px; }
133+
.stats { display: flex; justify-content: center; gap: 14px; flex-wrap: wrap; margin: 44px auto 0; max-width: 1020px; }
130134
.stat {
131-
flex: 1 1 150px; padding: 18px 10px; text-align: center;
135+
flex: 1 1 260px; padding: 20px 22px; text-align: center;
132136
background: var(--bg-alt); border-radius: var(--radius);
133137
}
134-
.stat b { display: block; font-size: 30px; font-weight: 800; letter-spacing: -0.02em; color: var(--accent); }
135-
.stat span { font-size: 13px; font-weight: 600; color: var(--ink-soft); }
138+
.stat b { display: block; font-size: 34px; font-weight: 800; letter-spacing: -0.02em; color: var(--accent); }
139+
.stat span { font-size: 14px; font-weight: 700; color: var(--ink); }
140+
.stat ul { list-style: none; margin-top: 12px; text-align: left; }
141+
.stat li {
142+
font-size: 12.5px; line-height: 1.5; color: var(--ink-soft);
143+
padding: 5px 0 5px 16px; border-top: 1px dashed var(--line); position: relative;
144+
}
145+
.stat li::before { content: "·"; position: absolute; left: 4px; font-weight: 800; color: var(--accent); }
146+
147+
/* ── Bullet lists ────────────────────────── */
148+
.blist-head { font-size: 15px; font-weight: 800; color: var(--ink); margin-top: 26px; }
149+
.blist { list-style: none; margin-top: 10px; }
150+
.blist li {
151+
position: relative; padding: 9px 0 9px 30px; font-size: 15px; color: var(--ink-soft);
152+
border-top: 1px solid var(--line);
153+
}
154+
.blist li:first-child { border-top: none; }
155+
.blist li b { color: var(--ink); }
156+
.blist li::before {
157+
position: absolute; left: 0; top: 9px; font-weight: 800; font-size: 14px;
158+
}
159+
.blist.bad li::before { content: "✗"; color: #d43c3c; }
160+
.blist.good li::before { content: "✓"; color: #0a8a5f; }
136161

137162
/* ── Comparison table ────────────────────── */
138163
.cmp-scroll { overflow-x: auto; border: 1px solid var(--line); border-radius: var(--radius); max-width: 760px; margin-left: auto; margin-right: auto; }
@@ -270,7 +295,6 @@
270295
<div class="topbar-inner">
271296
<a class="brand" href="#top"><img src="static/images/favicon.png" alt="">WorldMark</a>
272297
<div class="topnav">
273-
<a href="#abstract">Abstract</a>
274298
<a href="#comparison">Why WorldMark</a>
275299
<a href="#suite">Benchmark</a>
276300
<a href="#metrics">Metrics</a>
@@ -328,9 +352,33 @@ <h1><span class="wm">WorldMark</span>: A Unified Benchmark Suite for Interactive
328352
</div>
329353

330354
<div class="stats">
331-
<div class="stat"><b>10</b><span>Models Evaluated</span></div>
332-
<div class="stat"><b>500</b><span>Standardized Cases</span></div>
333-
<div class="stat"><b>9</b><span>Deterministic Metrics</span></div>
355+
<div class="stat">
356+
<b>10</b><span>Models Evaluated</span>
357+
<ul>
358+
<li>Yume 1.5 &middot; HY-World 1.5 &middot; HY-GameCraft 1.0</li>
359+
<li>Matrix-Game 2.0 / 3.0 &middot; LingBot-World</li>
360+
<li>SANA-WM &middot; DreamX-World &middot; AlayaWorld &middot; Lyra 2.0</li>
361+
<li>Caption, pose, keyboard &amp; trajectory control formats&mdash;one adapter each</li>
362+
</ul>
363+
</div>
364+
<div class="stat">
365+
<b>500</b><span>Standardized Cases</span>
366+
<ul>
367+
<li>50 scenes: 25 photorealistic + 25 stylized</li>
368+
<li>Paired first-/third-person views (100 images)</li>
369+
<li>5 VLM-selected trajectories per image</li>
370+
<li>3 difficulty tiers, 20&ndash;60s</li>
371+
</ul>
372+
</div>
373+
<div class="stat">
374+
<b>9</b><span>Deterministic Metrics</span>
375+
<ul>
376+
<li>Action Dynamics: accuracy, purity, latency, stability (per axis)</li>
377+
<li>World Memory: local, global, revisit</li>
378+
<li>Visual Quality: perceptual, aesthetic</li>
379+
<li>Feed-forward only&mdash;bit-identical across runs</li>
380+
</ul>
381+
</div>
334382
</div>
335383
</header>
336384

@@ -345,34 +393,26 @@ <h1><span class="wm">WorldMark</span>: A Unified Benchmark Suite for Interactive
345393
</figure>
346394
</section>
347395

348-
<!-- ── Abstract ── -->
349-
<section class="wrap" id="abstract">
350-
<p class="kicker">Abstract</p>
351-
<h2>The user acts, and the world responds</h2>
352-
<p>
353-
Unlike text- or image-driven video generation, an interactive world model is driven by <em>actions</em>. Two obstacles stand in the way of fair and comprehensive evaluation.
354-
First, models take actions in incompatible formats&mdash;captions, camera trajectories, action functions&mdash;so no shared protocol has been established.
355-
Second, while existing benchmarks have advanced world memory and visual quality, action following is reduced to trajectory or direction error, which collapses a whole path into one number:
356-
not how quickly the world reacts to a command switch, nor how cleanly it moves along the commanded axis.
357-
</p>
358-
<p>
359-
<span class="wm">WorldMark</span> removes both obstacles. Per-model adapters translate a shared WASD-style vocabulary into each model's native control format,
360-
so <strong>ten heterogeneous models receive semantically identical instructions</strong> across 500 standardized cases spanning styles, viewpoints, and difficulty tiers; a new model costs one adapter.
361-
On this common ground we characterize action dynamics through a <strong>control-systems lens</strong>&mdash;direction accuracy, direction purity, response latency, and motion stability, each resolved per axis&mdash;alongside suites for world memory and visual quality.
362-
Together they expose differences existing protocols cannot see: the fastest responders are often the least stable; per-axis resolution reveals models that follow translation almost perfectly while barely responding to rotation;
363-
and the model with the best perceptual quality ranks last in translational direction accuracy and latency. We release all data, evaluation code, and model outputs.
364-
</p>
365-
</section>
366-
367396
<!-- ── Comparison ── -->
368397
<section class="wrap" id="comparison">
369398
<p class="kicker">Why WorldMark</p>
370-
<h2>What existing benchmarks cannot measure</h2>
371-
<p class="muted">
372-
Visual quality and memory are well served; action following is not. Prior protocols report a single direction or trajectory score,
373-
which cannot in principle capture <em>response dynamics</em>&mdash;latency, persistence, stability&mdash;and none is fully reproducible run-to-run.
374-
</p>
375-
<div class="cmp-scroll" style="margin-top: 20px;">
399+
<h2>The user acts, and the world responds</h2>
400+
401+
<p class="blist-head">Two obstacles stand in the way of fair evaluation</p>
402+
<ul class="blist bad">
403+
<li><b>Actions arrive in incompatible formats.</b> Captions, camera trajectories, pose strings, action functions&mdash;no shared protocol exists, so models cannot be driven with identical instructions on identical scenes.</li>
404+
<li><b>Action following is collapsed into one number.</b> Existing benchmarks advanced world memory and visual quality, but reduce control to trajectory or direction error&mdash;which shows neither how quickly the world reacts to a command switch, nor how cleanly it moves along the commanded axis.</li>
405+
</ul>
406+
407+
<p class="blist-head">How <span class="wm">WorldMark</span> removes them</p>
408+
<ul class="blist good">
409+
<li><b>One vocabulary, ten native interfaces.</b> Per-model adapters translate a shared WASD-style vocabulary into each model's own control format, so ten heterogeneous models receive <strong>semantically identical instructions</strong> across 500 standardized cases spanning styles, viewpoints, and difficulty tiers. A new model costs one adapter.</li>
410+
<li><b>Action dynamics through a control-systems lens.</b> An action command is a <em>step input</em>: we measure direction accuracy, direction purity, response latency, and motion stability&mdash;each resolved separately for translation and rotation&mdash;alongside suites for world memory and visual quality.</li>
411+
<li><b>Deterministic by construction.</b> Every estimator is feed-forward, so scores are bit-identical across runs: no SLAM bundle-adjustment variance, no sampled VLM judgments.</li>
412+
<li><b>Differences existing protocols cannot see.</b> The fastest responders are often the least stable; per-axis resolution reveals models that follow translation almost perfectly while barely responding to rotation; and the model with the best perceptual quality ranks last in translational accuracy and latency.</li>
413+
</ul>
414+
415+
<div class="cmp-scroll" style="margin-top: 30px;">
376416
<table class="cmp">
377417
<thead>
378418
<tr>
@@ -444,19 +484,20 @@ <h3>Action Suite &mdash; 15 trajectories, 3 functional tiers</h3>
444484
<div class="tier med"><b>Medium &middot; 40s, two-segment</b>Adds exactly one switch, without which there is no transient; provides the first round trips.</div>
445485
<div class="tier hard"><b>Hard &middot; 60s, three-segment</b>Patrol routes and 360&deg; rotations run past most models' conditioning window, where revisit and global memory fail.</div>
446486
</div>
447-
<figure>
448-
<div class="figframe"><img src="static/images/trajectories_3d.png" alt="The 15 standardized action sequences" loading="lazy"></div>
449-
<figcaption>The 15 standardized action sequences, from elementary translations and rotations to combined and cyclic trajectories.</figcaption>
450-
</figure>
451-
452487
<p>
453488
For each image, a VLM selects the five most plausible sequences from the library, identifying physical constraints&mdash;lateral obstacles in a corridor make prolonged strafing implausible&mdash;so
454489
models are not penalized for refusing to walk through a wall. Across 100 images this yields <strong>500 evaluation cases</strong>.
455490
</p>
456-
<figure>
457-
<div class="figframe"><img src="static/images/vlm_traj.jpg" alt="Context-aware action selection via VLM reasoning" loading="lazy"></div>
458-
<figcaption>Context-aware action selection. A VLM analyzes the initial image to identify physical constraints and selects plausible action sequences from the predefined library.</figcaption>
459-
</figure>
491+
<div class="fig-row">
492+
<figure>
493+
<div class="figframe"><img src="static/images/trajectories_3d.png" alt="The 15 standardized action sequences" loading="lazy"></div>
494+
<figcaption>The 15 standardized action sequences, from elementary translations and rotations to combined and cyclic trajectories.</figcaption>
495+
</figure>
496+
<figure>
497+
<div class="figframe"><img src="static/images/vlm_traj.jpg" alt="Context-aware action selection via VLM reasoning" loading="lazy"></div>
498+
<figcaption>Context-aware action selection. A VLM analyzes the initial image to identify physical constraints and selects plausible action sequences from the predefined library.</figcaption>
499+
</figure>
500+
</div>
460501
</section>
461502

462503
<!-- ── Metrics ── -->
@@ -654,30 +695,13 @@ <h2>Ten models, four splits</h2>
654695

655696
<div class="wrap" style="padding: 0;">
656697
<figure>
657-
<div class="figframe"><img src="static/images/capability_profiles.png" alt="Radar capability profile of each model over thirteen columns, Real vs Stylized overlaid" loading="lazy"></div>
698+
<div class="figframe" style="max-width: 100%;"><img src="static/images/capability_profiles.png" alt="Radar capability profile of each model over thirteen columns, Real vs Stylized overlaid" loading="lazy"></div>
658699
<figcaption>Capability profile of each model over the thirteen reported columns, First-Person Real (solid) against Stylized (dashed). All axes run 0&ndash;100 in the same order in every panel; the gap between the two outlines is that model's stylization penalty. Lyra&nbsp;2.0 is the only near-convex profile&mdash;one all-round model and no second&mdash;and the stylization penalty falls hardest on the models that lead on Real.</figcaption>
659700
</figure>
660701
</div>
661702
</section>
662703

663704
<!-- ── Qualitative examples ── -->
664-
<section class="wrap" id="examples">
665-
<p class="kicker">Qualitative Examples</p>
666-
<h2>What each failure looks like</h2>
667-
<p class="muted">
668-
One high- and one low-scoring case per metric, holding scene, command, and sampling fixed within each pair&mdash;the only difference is the model.
669-
The Action Dynamics rows carry a signal panel as well as frames, because their failures are in the <em>motion</em> rather than the pixels: six stills cannot show that a model kept turning after the command changed.
670-
</p>
671-
<figure>
672-
<div class="figframe"><img src="static/images/metric_examples_action.jpg" alt="High- and low-scoring examples for each Action Dynamics metric with flow signal panels" loading="lazy"></div>
673-
<figcaption>Action Dynamics examples. The panel at left is the signal the metric reads; shading marks what the flow on the commanded axis <em>should</em> do, and the dashed rule in Response Latency is the command switch. Green marks the high-scoring model, red the low-scoring one.</figcaption>
674-
</figure>
675-
<figure>
676-
<div class="figframe"><img src="static/images/metric_examples_rest.jpg" alt="High- and low-scoring examples for memory and quality metrics" loading="lazy"></div>
677-
<figcaption>Memory and quality examples, drawn from the span each metric actually reads: adjacent frames for Local Memory, outbound and return halves for Revisit Memory, the full trajectory otherwise. These failures are visible in the frames themselves.</figcaption>
678-
</figure>
679-
</section>
680-
681705
<!-- ── Findings ── -->
682706
<section class="wrap" id="findings">
683707
<p class="kicker">Key Findings</p>
@@ -737,12 +761,10 @@ <h2>BibTeX</h2>
737761

738762
<!-- ── Right side nav ── -->
739763
<ul class="side-nav">
740-
<li><a href="#abstract">Abstract</a></li>
741764
<li><a href="#comparison">Why WorldMark</a></li>
742765
<li><a href="#suite">Benchmark Suite</a></li>
743766
<li><a href="#metrics">Metrics</a></li>
744767
<li><a href="#results">Results</a></li>
745-
<li><a href="#examples">Qualitative Examples</a></li>
746768
<li><a href="#findings">Key Findings</a></li>
747769
<li><a href="#arena">Arena</a></li>
748770
<li><a href="#bibtex">BibTeX</a></li>
-617 KB
Binary file not shown.
-818 KB
Binary file not shown.

0 commit comments

Comments
 (0)