Judge Arena: Benchmarking LLMs as Evaluators
Running Agents 112 Judge Arena 💻 112 View and compare open‑source AI model rankings with ELO scores
🔗 Read full article on Hugging Face →
Source: Hugging Face — Published — Category: Models