Skip to main content
AIDiveForge AIDiveForge
Save tools:Log inSign up
Visit AgiRanker

Get This Tool

License: License: unverified

Share This Tool

Compare This Tool
📋 Embed this tool on your site

Copy this code to embed a compact tool card:

AgiRanker

FreeOpen Source

Pricing

Model
Free

Summary

Composite AI benchmark scores are only as trustworthy as the methodology behind them — and most leaderboards bury that methodology three clicks deep, if they publish it at all. AgiRanker is a public web leaderboard built around a single, openly documented AGI Score that aggregates ten major benchmarks into one number, with every weight and correction logged publicly.

The tool lets you rank frontier models on a 0–100 AGI Score, explore per-domain breakdowns across Thinking, Doing, and Communicating, and reweight the formula yourself to stress-test whether the ranking changes when you care more about coding than knowledge. A value-for-money view plots capability against published API list prices, so you can see which model gives the most capability per dollar without running your own evals. The Reasoning category flags itself: only two benchmarks cover it, and most models have just one data point — the site surfaces this rather than hiding it. There is no API, no self-hosting path, and no programmatic data export described anywhere on the site.

Bottom line: Pick AgiRanker when you need a defensible, transparent starting point for comparing frontier models without building your own benchmark aggregator — but plan a separate pipeline the moment you need programmatic access, custom benchmark ingestion, or scores outside the ten sources it tracks.

Community Performance Report Card

No community ratings yet. Be the first to rate this tool!

Best For: Researchers tracking verifiable AI capability trends, Analysts needing transparent composite scores, Users evaluating model performance per dollar
  • Publicly logged corrections and a versioned methodology page, which means you can audit exactly why a score changed between versions — something you cannot do with lab-self-reported rankings.
  • Weight customization in the Explorer adjusts all composite scores on the fly, so you can validate whether a model's top ranking holds under your team's actual domain priorities rather than the default weighting.
  • Unverified scores are flagged rather than estimated, which means contested or unconfirmed numbers do not silently inflate a model's position the way they do on aggregators that fill gaps with proxies.
  • A value-for-money scatter plot pairs capability scores against published list prices, so you can identify which model delivers the most capability per dollar without running cost experiments yourself.
  • CC BY 4.0 data licensing, so benchmark data can be reused in downstream analysis without negotiating permissions — unlike proprietary leaderboards that restrict redistribution.
  • There is no API and no described data export path, so any team that needs to pull scores into a dashboard, trigger alerts on ranking changes, or feed benchmark data into an internal tool has to scrape the site manually — and maintain that scraper every time the site updates.
  • The Reasoning category covers only two benchmarks with sparse model coverage, and the site itself warns that single-cell rankings in this domain mislead. Teams evaluating models specifically on reasoning-heavy tasks — math, formal logic, novel problem-solving — get a number that the vendor explicitly says should not be trusted at face value.
  • Coverage is limited to the ten benchmarks AgiRanker has activated. Teams tracking a model that dominates on a benchmark outside that set — or whose internal eval suite differs from the included sources — get a score that misrepresents relative performance for their workload. At that point, teams either build a parallel aggregation layer or move to a platform that accepts custom benchmark ingestion, at which point AgiRanker becomes a reference check rather than a primary tool.

About

Platforms
Web
API Available
No
Self-Hosted
No
Last Updated
2026-08-14T03:22:28.766Z

Best For

Who it's for

  • Analysts needing transparent composite scores
  • Users evaluating model performance per dollar

What it does well

  • Compare frontier AI models on a unified AGI-progress metric
  • Explore benchmark coverage and source attribution
  • Analyze value for money across capability and API pricing
Help improve this page

Add notes, reviews, and benchmarks so the next visitor gets a clearer picture.

Sign in to contribute

Spotted incorrect or missing data? Join our community of contributors.

Sign Up to Contribute

Frequently Asked Questions

Is AgiRanker free?
Yes — AgiRanker is fully free to use. There is no paid tier.
Is AgiRanker open source?
Yes. AgiRanker is open source.
When was AgiRanker released?
AgiRanker was first released in 2026.
What platforms does AgiRanker support?
AgiRanker is available on: Web.

Most leaderboards bury methodology three clicks deep

Composite AI benchmark scores are only as trustworthy as the methodology behind them. AgiRanker aggregates ten benchmarks into a single 0-100 AGI Score with every weight and correction logged publicly, and it performs no imputation of missing data.

Explorer and value view

The tool ranks frontier models on the AGI Score and shows per-domain breakdowns for Thinking, Doing, and Communicating. Users can reweight the formula on the fly to test whether rankings shift when priorities favor coding over knowledge. A separate value-for-money plot compares capability against published API list prices.

Transparency on gaps

The Reasoning category flags itself because only two benchmarks cover it and most models have just one data point. Unverified scores are surfaced rather than estimated or hidden.

Limitations

There is no API and no data export, so teams that need to pull scores into dashboards must scrape the site and maintain the scraper on updates. The site itself warns that single-cell rankings in Reasoning should not be treated as reliable.

Who it is for / who should skip it

Best for researchers tracking verifiable AI capability trends, analysts needing transparent composite scores, and users evaluating model performance per dollar. Skip it if you require an API or automated data feeds.