INDENG 242B final project, Spring 2026

Slime volleyball,
played by gradient descent.

We trained eighteen deep-RL policies on the same tiny volleyball court, then froze a 4,860-match benchmark to settle which one is actually best. Everything runs in your browser — watch the policies play, check the numbers, or pick up the keyboard and challenge them yourself.

18
trained policies
4,860
benchmark matches
0
servers involved
Preparing live inference…
— fps
Loading…
5–5 LIVE · ON-DEVICE
Loading…
Your browser does not support canvas.
Loading match…
Exhibition series Reference 0 · Trained 0 · Draws 0 Opening match
Policy Inspector per-action outputs · rolling 200 frames
Model controls & Policy Inspector
Inspect side
Output view

PPO exposes policy logits. DQN and Rainbow expose Q-values; “Centered” removes each frame’s mean. The RNN reference has neither and is not charted. Keyboard when a side is human — left: A/D move, W or Space jump · right: / move, jump · R restart.

Loading the featured policy…