Arena

Evaluation platform born at Berkeley where the public blind-compares two AI models' answers and votes, feeding a ranking based on human preference.