🍌 BananaMind Arena
Blind, side-by-side battles between very small language models. Vote, and the winner's Elo goes up.
Instruct Arena
Both models answer your message through their own chat template.
Models in this arena: SmolLM2-135M-Instruct (134.5M) · Supra2-100M-Instruct (100.7M) · Supra2-Medium-Instruct (25.4M) · BananaMind-2-Nano-Chat (10.0M) · BananaMind-2-Mini-Chat (25.2M) · BananaMind-2-Medium-Chat (49.6M) · BananaMind-2-Pro-Preview-Chat (139.0M)
Instruct Leaderboard
7 instruct / chat models · 22 rated battles. Ratings are a Bradley-Terry maximum-likelihood fit over the whole battle record, re-solved after every vote, so beating a strong model is worth more than beating a weak one. Unplayed models sit at 1000. Both bad scores as a draw.
Base Arena
Both models continue your text — these are base models, so write a prompt they can complete rather than an instruction.
Models in this arena: BananaMind-2-Micro (3.2M) · BananaMind-2-Nano (12.1M) · BananaMind-2-Medium (55.9M) · BananaMind-2-Pro (159.9M) · Supra1.5-50M-Base-exp (51.8M) · Supra2-Medium-Base (25.4M) · Supra2-100M-Base (100.7M) · SmolLM2-135M (134.5M) · GPT-X2.5-135M (135.0M) · Syn-2.6M (2.6M) · Boris-1.3-75M (77.4M) · Boris-1.7-D60M-n30M (90.0M) · Zero-v0.1-150M (151.6M)
Base Leaderboard
13 base (text-completion) models · 6 rated battles. Ratings are a Bradley-Terry maximum-likelihood fit over the whole battle record, re-solved after every vote, so beating a strong model is worth more than beating a weak one. Unplayed models sit at 1000. Both bad scores as a draw.
Every prompt, both responses and every vote are logged to the public bucket Banaxi-Tech/bananamind-arena-logs — assume anything you type here is publicly accessible. Don't submit personal or sensitive information.