Compose
Which one sounds better?
Two architectures compose the same brief — identical length, key, temperature and random seed — and you hear them blind. Validation loss says which model fits the corpus. Only a listener says which one is worth hearing. Votes update an Elo rating for each model.
The arena needs two trained models. 0 is trained so far.
Train another from the Train movement — the transformer is the interesting opponent, and at ~3.9M parameters it trains far faster than the 179M-parameter LSTM + attention.
Leaderboard
| Model | Elo | Games | W / L / D | Win rate |
|---|
Every model starts at 1500. A win moves the rating by up to 24 points, scaled by how surprising the result was — beating a higher-rated model is worth more. Ratings are only meaningful after a few dozen votes.