Judge Arena: Benchmarking LLMs as Evaluators

Running Agents 112 Judge Arena 💻 112 View and compare open‑source AI model rankings with ELO scores

Source: Hugging Face — Published — Category: Models

🔗 Read full article on Hugging Face →