ClientCoded
ClientCoded runs adversarial multi-turn conversations against your agent, scores every exchange across 10 published dimensions, and fires a…
Leaderboards, eval harnesses, and benchmark platforms.
ClientCoded runs adversarial multi-turn conversations against your agent, scores every exchange across 10 published dimensions, and fires a…
Oqoqo runs agent evaluation experiments on managed cloud infrastructure, spinning up sandboxed machines per task so runs don't bleed into…
The tool lets you rank frontier models on a 0–100 AGI Score, explore per-domain breakdowns across Thinking, Doing, and Communicating, and…
AgentsProof is an evaluation SDK that wraps your LLM and tool calls with a decorator, grades each run against rules you define in plain…
Langdrift runs your agent prompts across multiple locales and compares behavior — checking whether tool calls, response structure, and…
The published detection module runs three frontier LLMs across four analyst personas against 1,205 labeled NetFlow units, scoring each…
The core loop is simple: you submit a prompt, two models respond anonymously, you vote for the better answer, and Arena logs the result…
The platform combines trained human evaluators with automated metrics across three surfaces: multi-step agent workflows, SaaS AI features…
Proctor wraps each agent execution in a Linux sandbox that cuts off access to hidden tests, fix history, and network egress, so the agent…
Built by the Stanford Digital Economy Lab and described in arXiv paper 2605.04312, Agent Island puts language models into a shared…