LLM Evaluation & Benchmarks
Leaderboards, eval harnesses, and benchmark platforms.
AgentsProof
AgentsProof is an evaluation SDK that wraps your LLM and tool calls with a decorator, grades each run against rules you define in plain…
Agent Island
Built by the Stanford Digital Economy Lab and described in arXiv paper 2605.04312, Agent Island puts language models into a shared…