INDENG 242B final project, Spring 2026
Slime volleyball,
played by gradient descent.
We trained eighteen deep-RL policies on the same tiny volleyball court, then froze a 4,860-match benchmark to settle which one is actually best. Everything runs in your browser — watch the policies play, check the numbers, or pick up the keyboard and challenge them yourself.
- 18
- trained policies
- 4,860
- benchmark matches
- 0
- servers involved
Model controls & Policy Inspector
PPO exposes policy logits. DQN and Rainbow expose Q-values; “Centered” removes each frame’s mean. The RNN reference has neither and is not charted. Keyboard when a side is human — left: A/D move, W or Space jump · right: ←/→ move, ↑ jump · R restart.