INDENG 242B final project, Spring 2026

A strong AI.
Against whom?

The top-scoring volleyball policy lost more often than it won against the runner-up. A broken ranking — or a different question? Compare 18 trained policies using recorded evidence. Test a fixed matchup with paired seeds. Then run the models live in your browser.

18
trained policies
4,860
benchmark matches
10
common ranking seeds
System design — how evaluation, inference and reproducibility fit together

These are two different execution paths. Recorded evidence explains the frozen Python benchmark; browser play and experiments run new JS simulations. Neither path is presented as the other.

Run models, keep the experiment separate

  1. to select models. Both load before a new pair is committed.
  2. A 30Hz physics clock is separate from rendering and asynchronous ONNX inference. Live play may reuse the last completed action.
  3. to freeze paired seeds and both side assignments before execution.
  4. Each match owns a Worker. Results, failures and unrun slots are explicit; exported protocols can be re-run inside the same versioned browser engine.

Why not one universal strength score?

An opponent-balanced average and a direct matchup answer different questions. We retain both, disclose ten common seed blocks, and avoid converting low variation into high strength. The simultaneous Arena difference band does not apply to every new metric.

Why not reproduce the official match live?

The browser has different random sequences and action timing. Python recordings preserve what happened in an official match; local Workers support repeatable browser experiments, not cross-engine equivalence.

How can a static site reject bad evidence?

Source and trace hashes check exact bytes; schemas and semantic checks reject inconsistent identities or outcomes. Native SHA-256 has a tested fallback for local HTTP. Same-origin hashes check consistency, not independent authenticity.

What is tested, and what is not claimed?

Generators recompute derived artifacts. Tests cover source mutations, stale requests, fixed schedules, cancellation, and Chrome/WebKit interactions. Fixed canaries and the ten selected re-recordings do not mean every historical match is rerun on every release. Training-seed variance and causal ablations require additional experiments.

Loading the featured policy…