Methodology
ModelCap is a continuously refreshed public-data index for AI models. The primary language-model rank is a capability estimate that does not require ModelCap to run or host a model. Price, popularity, downloads, likes and provider breadth never decide capability. Every input comes from a named public feed, every identity and configuration stays traceable, and modeled values are labeled separately from measurements.
ModelCap Rank
Current rank is a fixed, contiguous position—1, 2, 3 and so on—among current, canonical, exact language-model identities. All-versions archive rank is a separate contiguous ordering across both current and superseded canonical identities; it is not a model’s frozen historical rank at retirement. Open-weight repositories can enter without a ModelCap-hosted evaluation, but only when their artifact identity is pinned. Image and audio catalogues stay separate because their task units are not capability-comparable. The index follows one deterministic order:
exact admitted evidence: conservative confidence-bound capability scoreverified child or quantization: explicit penalty from the exact parentno verified lineage: calibrated architecture + release-era fallbackunsupported, moving, or ambiguous identity: remain unranked
Measured rows use the existing Bayesian Evidence Score normalization. Its multi-family likelihood uses free public evaluations — 80% general preference · 12% coding · 5% agents/tools · 3% reasoning. When Overall is present it supplies 80% of the likelihood and the available specialist families share 20%. Overall-led models are not demoted merely because a specialist board has not tested them. The breadth adjustment applies only to specialist-only evidence. A narrow source receives a wider interval and less evidence support than broad multi-family evidence. The active v4.1 board sorts a conservative lower confidence bound and publishes separate contiguous Current and All versions positions; it does not expose v5 rank probabilities. Evidence age lowers confidence; observations below 1% combined confidence are excluded from scoring. Modeled rows show a modeled label, confidence, and interval. Sparse static metadata produces a deliberately broad estimate; unsupported or ambiguous identity produces no rank.
The four families use Arena Overall; controlled SWE-bench and Arena Coding; Arena Agent and BFCL; and verified ARC-AGI-2 / ARC-AGI-3. Each source becomes:
source score = 80% competitive placement + 20% achievement. Placement is a midpoint-smoothed source percentile (rank 1 of 100 is 99.5, not fake perfection). Achievement keeps each board’s published scale: pass rates as percentages; Arena ratings via the inverse of Arena’s Bradley–Terry display transform; Agent causal lift with 50 as neutral.
When no exact public evaluation matches, the active v4 Index first accepts an exact, verified, acyclic lineage edge and applies its explicit configuration penalty. Without verified lineage, the architecture-and-release-era fallback must pass its independent-anchor, support, proximity, calibration, and uncertainty gates; otherwise ModelCap publishes no score or rank. Popularity, providers, downloads, price, and unverified model-card evaluations remain separate context and cannot raise capability.
Arena Overall prefers an exact Default result; only when Default is absent can a named configuration supply that family. Specialist sources select their highest published score. Every configuration remains separate on the detail page. Identity confidence, sample size, interval width and actual source publication age determine evidence coverage.
In the active v4 Index, exact identity means the named catalogue product, not an assertion that every public board ran one identical endpoint configuration. The product score aggregates the admitted observations while disclosing each board’s configuration. ModelCap claims an exact endpoint configuration only after every contributing observation has a reviewed mapping; the current v4 artifact does not make that claim. A Hugging Face commit is retained separately as artifact metadata and is never relabeled as the external evaluation revision.
Inside the measured-evidence subsystem, a BES result is labeled Confirmed with Arena Overall, at least 2families and evidence mass ≥ 50%; otherwise it is Provisional. Those labels describe the measured input, not a second public ordering. ModelCap keeps one conservative score ordering with separate contiguous Current and All versions positions.
Filtering or sorting the table does not renumber a model. Rank stays stable within the selected scope. Under v4.1 and v5, switching between Current and All versions can display a different position because those are two independently contiguous boards; only the predecessor v4 snapshot has a single all-version scope.
The design follows published work on partial rankings with missing benchmark scores and robust cross-benchmark rank aggregation. ModelCap’s published layer is the source mapping, family weights, source-scale transforms, configuration selection and evidence adjustment—not a hidden model making subjective guesses.
Arena rating
The homepage Arena rating column shows one official external Arena result per model. When Arena publishes multiple configurations, ModelCap uses the one with the highest published Arena rating and displays that rating—not Arena’s global source rank. It does not average configurations or alter the winning result with price, context size or feature checklists. The source rank, configuration name, vote count, confidence interval and snapshot time remain on the detail page.
ModelCap joins the complete current Arena text board, then falls back to Arena’s reproducible archive if the live board is unavailable. Matches require the same publisher and full model identity. A reviewed one-way registry handles official dated names, full checkpoint names, provider renames and named configurations; fuzzy and family-name matching are never used.
A newly available model without any admitted result is labelled Awaiting evaluation. It can receive a ModelCap Index position from exact verified lineage or from the active v4 architecture-and-release-era fallback only when its support and uncertainty gates pass. Its Arena rating remains blank, and the index is explicitly labeled Inherited or Estimated rather than Measured. A named configuration such as High or Max counts as published quality evidence only for that exact configuration. When a possible cross-source identity is not reliable enough, it is labelled Identity match under review rather than attaching another configuration’s result.
Evaluation evidence is configuration-specific. Preview versions, reasoning levels, quantizations, and dated releases are not silently merged into a benchmark result. The catalogue folds an endpoint only when a reviewed registry declares it an identical alias of a canonical route; unreviewed lookalikes remain separate. Confidence intervals describe uncertainty in the community rating; identity-match confidence separately describes how certain the catalogue-to-evaluation name match is.
Related evaluations such as High, X-High, Max, thinking modes, or explicit token budgets appear as separate configurations on the canonical model’s detail page. The highest Arena community rating supplies the separately labelled homepage Arena value. The general-capability family follows the Default-first rule above and uses a named configuration only when Default is absent. The original source result is never relabeled as a default score.
Board metrics also stay separate. For example, LMArena Agent publishes an agent score with observations and sessions—not the Arena preference rating and vote count. ModelCap labels each board and normalizes its rank before combining evidence; it never averages raw, incompatible scales.
Three evidence layers
ModelCap no longer asks one number to mean three different things. Every model page separates:
- Measured ModelCap Index
- Admitted public observations bound to one exact catalogue product, with every contributing board configuration, source field size and publication provenance disclosed separately. ModelCap did not run the model, and v4 does not claim one exact endpoint configuration across those observations.
- Inherited or Estimated ModelCap Index
- The active v4 method applies an explicit configuration penalty to exact verified lineage, or uses its gated architecture-and-release-era fallback. Both remain visibly distinct from exact measured evidence, and v4 abstains when fallback support is insufficient.
- Deployment readiness
- Artifact, licence, provider and evaluation-provenance metadata. It measures how inspectable and deployable an artifact appears, never capability, safety certification or production readiness.
Hugging Face provenance
For an exact Hugging Face repository id, the sync records the immutable commit SHA, architecture, model type, total and explicitly reported active parameters, dtypes, base lineage, licence/access state, library, task, declared training datasets and repository dates. Inference-provider pricing, context, tool support, time-to-first-token and throughput stay in a separate operational feed.
Structured model-card evaluation rows are useful discovery leads, but they are not automatically trustworthy benchmark observations. Every repository row preserves the observed artifact SHA, dataset, dataset revision, configuration, task, metric, source file and verification status when published. The artifact SHA is not treated as the revision evaluated by an external board unless that source explicitly binds it. It remains quarantined and excluded from scoring until an independent admission review confirms exact identity, comparable harness and immutable dataset provenance. An unpinned dataset revision cannot pass that gate.
Cold-start posterior and validation
The cold-start estimator uses only models with admitted observed capability as anchors. Quarantined Hugging Face result values are not features. The launch proxy uses a narrow allowlist: architecture, model type, log total and active parameters, modalities, release-time proximity, exact lineage, and explicit quantization. Its calibration excludes anchors from the same publisher and lineage family so the system cannot validate a model by learning from another spelling of the same lineage.
Each result publishes its evidence state, support, score interval, expected rank, rank interval, and top-5/top-10 probabilities. Out-of-domain distance, missing static fields, weak lineage, and historical residual error widen uncertainty rather than granting an optimistic score. Unsupported or ambiguous identities still fail closed. An Estimated result remains visibly different from Measured evidence everywhere it is displayed.
Release validation is temporal: the evidence ledger freezes the first modeled posterior and compares it only with a strictly later admitted measurement. The cutover protocol additionally requires complete publisher, lineage, and architecture-family holdouts. Launch gates cover top-10 recall, frontier misses, pairwise ordering, interval calibration, group error, identity safety, and time from a complete manifest to an estimate.
This follows the practical lesson from tinyBenchmarks: a carefully calibrated subset can be informative. It also preserves the caution from research on predicting frontier model performance: extrapolation beyond observed model families is fragile and must be visible.
Deployment readiness metadata
The readiness index is a deterministic, componentized summary of artifact reproducibility, access/licence clarity, deployability and evaluation provenance. Missing metadata lowers published coverage and stays listed on the model page. A high value does not mean the model is capable, safe, compliant, cheap or production-certified; a low value may simply mean the publisher has not exposed enough machine-readable detail.
ModelCap Core pilot
The repository now contains a versioned ModelCap Core pilot protocol and an active-sampling planner. Given a direct-run budget, the planner chooses two roles: observed residual anchors spread across the capability range, and unobserved targets prioritised by uncertainty, out-of-domain risk, board impact and lineage diversity. Unselected models receive an estimate with an interval or an abstention—they are not silently treated as tested.
The 100-item pilot allocates general, coding, agent and reasoning tasks, pins generation settings and configuration identity, repeats requests, captures request/response hashes and provider ids, and requires a sealed holdout plus inter-rater audit for subjective items. Item text remains private to reduce contamination; the public artifact is the protocol, manifest hash and aggregate evidence. Planning is local and makes no paid API calls: npm run benchmark:plan -- --budget 12.
Running the selected models is a separate measured-evidence gate because it requires provider credentials, a sealed item bank and explicit spend authority. Pilot results still cannot affect rank until the same identity, completeness, holdout and provenance checks used for external sources pass. The operator commands and private-artifact boundary are documented indocs/modelcap-core-operations.md.
Market Gravity
A 0–100 composite of four broadly available signals. It measures market presence—not intelligence, and not a fictional global market cap. Market Gravity remains on model pages as transparent secondary context; it no longer determines the main ranking. See what this does not measure.
Usage
55%OpenRouter weekly popularity across the full model catalogue.
Liquidity
25%Independent providers versus the model's open or closed cohort, weighted by uptime.
Open reach
15%Hugging Face 30-day downloads for open models; neutral for closed models.
Freshness
5%Time since first public availability, on a six-month half-life.
Usage uses OpenRouter’s public weekly popularity order, which covers the full live catalogue. Because OpenRouter exposes only an ordinal without an API key, ModelCap scores that ordinal with explicit logarithmic decay and never invents a token count. Vercel’s spend, token and request exports cover only a small top-ten slice, so they are excluded from Market Gravity rather than treating missing models as zero.
Liquidity compares independent provider breadth within open and closed cohorts and modulates it with measured uptime. Open reach uses the Hugging Face 30-day download percentile for open weights; closed models receive the neutral midpoint because no comparable Hub download channel exists. Freshness has only 5% weight, enough to surface a real launch but not enough to crown one without activity.
Best value
Cheap is not the same as good. A model must first be in the top Arena preference quartile and have at least 1,000 community votes. Models without quality evidence cannot qualify.
Cost uses a published 3:1 input-to-output workload: 750,000 input tokens and 250,000 output tokens per one million combined tokens. The blended price is therefore (0.75 × input price) + (0.25 × output price).
From the eligible set, ModelCap publishes the price-performance frontier. A model is removed if another model has both an equal-or-higher quality rating and an equal-or-lower blended price, with at least one strict advantage. This prevents a poor but extremely cheap model from winning and avoids pretending that one arbitrary quality-per-dollar ratio fits every budget.
Release radar
The release radar is separate from the rankings and never changes a model's ModelCap score, position, or availability date. It scans the complete active feed under Polymarket's AI Releases tag and discovers model-release families from the market questions themselves. The card covers the active workweek through the following Friday, rolls forward every Sunday, and refreshes live odds every minute on Sunday and weekdays. Model names are not manually maintained.
To qualify, an item needs an active Yes/No contract with a dated “released by” question, a matching market deadline, public-release resolution rules, and a quote updated within four hours. A high-quality signal also needs a two-sided spread within five cents, at least $5,000 of liquidity, and at least $500 in trailing 24-hour volume. A current thinner contract can appear when its spread is within ten cents and liquidity is at least $100. If an otherwise current public-release contract has a one-sided order book, the radar can show its live indicative Yes price when liquidity is at least $100; it does not present that as a midpoint or use it for a date estimate. Displayed figures are market signals, not ModelCap forecasts or promised launch dates.
A market-derived date is stricter still: at least three compatible, liquid “released by” cutoffs must bracket the 25th, 50th, and 75th percentiles at useful cadence. The card then labels the interpolated 50% crossing as a market estimate and shows its 25–75% window. A lone deadline contract—or a lower-liquidity market deadline—remains a dated market projection rather than a manufactured estimate.
The homepage publishes up to five families whose qualified market deadline falls inside that two-workweek window, ordered by the implied probability at the latest qualifying cutoff, then quote quality, activity, and liquidity. Related contracts such as Claude Opus, Sonnet, and Haiku are collapsed to one “Next Claude” family without mixing their probability ladders. If fewer than five families qualify, the card keeps its layout but explicitly leaves the remaining capacity empty instead of pulling in later dates or inventing rows. A last-good snapshot may bridge a short upstream interruption for no more than 30 minutes and is never carried across a Sunday window change.
We do not turn anonymous posts, social-media speculation, or unverified leaks into public model-release claims. “Recently available” remains a separate source-backed catalogue observation.
Sources
Public, unauthenticated feeds supply catalogue, pricing, community results, venue activity and open-weight reach. Each sync writes a dated snapshot; last successful run 7 Aug 2026, 20:30 UTC.
The homepage's trending-model card is a separate, same-origin live feed from Hugging Face. It shows the provider's current text-model trend order, not an invented universal activity score; the card shows the source check time and labels any retained result as stale. Trend position and downloads never affect a model's capability rank.
- https://openrouter.ai/api/v1/modelsOpenRouterokcatalogue, pricing, context, modality, release dateFetched 7 Aug 2026, 20:30 UTC400 rows
- https://openrouter.ai/api/v1/models?sort=most-popular&output_modalities=textOpenRouter Popularityokweekly text-model popularity order (ordinal only)Fetched 7 Aug 2026, 20:30 UTC298 rows
- https://vercel.com/api/ai/leaderboard-export?dataset=models&modality=textVercel AI Gatewayokdaily top-10 request, token and spend share; latest reading and 7-day changeFetched 7 Aug 2026, 20:30 UTCPublished 7 Aug 2026, 00:00 UTC1538 rows
- https://openrouter.ai/api/v1/models/:id/endpointsOpenRouter Endpointsokper-provider pricing, measured uptime, and bounded retained provenanceFetched 7 Aug 2026, 20:30 UTC
- https://huggingface.co/api/models/:repoHugging Facepartialrevision, access and license, architecture, parameters, lineage, card metadata, reported evaluation candidates, downloads, likesFetched 7 Aug 2026, 20:30 UTC
- https://router.huggingface.co/v1/modelsHugging Face Inference Providersokexact model id, provider status, context, price, tool and structured-output support, first-token latency, throughputFetched 7 Aug 2026, 20:30 UTC130 rows
- https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetLMArenaokstyle-controlled overall community rating, confidence interval, votes, rankFetched 7 Aug 2026, 20:30 UTCPublished 6 Aug 2026, 00:00 UTCRevision d485cd3f6a51385 rows
- https://arena.ai/leaderboard/textArena Liveokcurrent style-controlled overall community rating, confidence interval, votes, rankFetched 7 Aug 2026, 20:30 UTCPublished 6 Aug 2026, 13:00 UTC385 rows
- https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetLMArena Agentstale-cacheagent score, confidence interval, observations, sessions, rank, configurationFetched 7 Aug 2026, 20:30 UTC0 rows
- https://arena.ai/leaderboard/text/codingArena Codingokcurrent style-controlled category rating, confidence interval, votes, rankFetched 7 Aug 2026, 20:30 UTCPublished 6 Aug 2026, 13:00 UTC380 rows
- https://www.swebench.com/SWE-benchokcontrolled bash-only coding score, source rank, model configurationFetched 7 Aug 2026, 20:30 UTCPublished 26 Feb 2026, 00:00 UTC47 rows
- https://gorilla.cs.berkeley.edu/leaderboard.htmlBFCLokfunction and tool-use overall accuracy, rank, configurationFetched 7 Aug 2026, 20:30 UTCPublished 12 Apr 2026, 00:00 UTC109 rows
- https://arcprize.org/leaderboardARC Prize · ARC-AGI-2okverified semi-private abstract-reasoning score, source rank, reasoning effortFetched 7 Aug 2026, 20:30 UTCPublished 7 Aug 2026, 20:10 UTC190 rows
- https://arcprize.org/leaderboardARC Prize · ARC-AGI-3okverified interactive-reasoning score, source rank, reasoning effortFetched 7 Aug 2026, 20:30 UTCPublished 7 Aug 2026, 20:10 UTC26 rows
Why these benchmarks
More benchmark rows do not automatically make a better score. A source enters ModelCap only when it is public and free to read, machine-readable, dated or versioned, explicit about model configuration and evaluation harness, broad enough to compare current frontier models, and precise enough for an exact identity match. A new source also needs a declared 0–100 achievement transform before the sync will accept it.
SWE-bench Verified measures resolved real-world GitHub issues, but the result evaluates a model together with its agent scaffold. ModelCap therefore uses the official controlled bash-only board rather than mixing unrelated scaffolds. BFCL contributes structured tool-use accuracy; Arena Agent contributes independently measured behavior from real workflows; and ARC contributes only independently verified results.
ModelCap monitors additional high-quality work such as LiveBench, Humanity’s Last Exam, Terminal-Bench and SWE-bench Pro. They are not silently folded in when a public table mixes tool settings, scaffolds, self-reported figures or incomplete closed-model coverage. They can be admitted in a future methodology version once one comparable, structured feed clears the same publication gate.
Identity and catalogue policy
OpenRouter currently contributes 400 routing entries, while exact Hugging Face Router discovery can add independently identified repositories. After normalization ModelCap tracks 378 model identities; 104 currently hold a public category position. These policies explain how identities are grouped, retained, ranked, or separated:
- Alias pointers
- Entries like ~anthropic/claude-opus-latest resolve to a concrete model already listed. Counting both would double every frontier release.
- Meta-routers
- openrouter/auto picks a real model per request. It is infrastructure, not a model.
- Free mirrors
- A :free entry is the same weights at a promotional price. It folds into its parent as a free-tier flag.
- Superseded versions (109)
- A newer family member removes an older release from the contiguous Current board while preserving that exact supported identity in the separately contiguous All versions archive. Each ranked identity still needs its own evidence, verified lineage, or gated architecture-and-release-era fallback.
- Identical variants
- Only a reviewed one-way registry can declare a speed route, preset, or provider alias equivalent to one canonical product. Similar fields alone never collapse a new model.
- Adult fine-tunes
- A small denylist of roleplay-oriented publishers, so the board stays safe to open at work. Not a quality judgement.
- Non-language models
- Image, audio and video generators are ordered on separate catalogue boards from published deployment signals. Those positions are not capability ranks, and token prices do not compare across media.
- Models with no live endpoint
- The catalogue keeps listing models after every provider has dropped them. They can retain a universal Index position when exact capability support remains, while their page clearly reports that no current serving endpoint is available.
How change is measured
Every sync writes a dated snapshot of each model’s ModelCap rank, Market Gravity score and provider count, plus language-model output token price. The next compatible sync diffs those readings within the same category. Rank movement appears directly beneath the far-left rank. A methodology change starts a new movement series rather than comparing unlike scores.
A model with no prior reading is marked New and reports no change, which is the honest answer. Change pills are hidden rather than shown as zero, because a grey 0.00% would claim a measurement that was never taken.
What this does not measure
One universal definition of quality. Community preference is a useful overall signal, not proof that one model is best at coding, factual reasoning, tool use, latency-sensitive work, or a particular language. ModelCap publishes its component scores and source rows so the overall summary never hides where two models differ.
No composite proves one model is best for every task. ModelCap is an overall summary, not a claim that the top model wins every coding language, latency target, factual domain or workflow. Source-level evidence remains visible on each model page.
Throughput and latency. These are provider- and deployment-specific, not inherent model constants. ModelCap shows Hugging Face Router telemetry only for an exact repository/provider record and leaves it blank elsewhere. OpenRouter still returns those fields as null to unauthenticated clients, so the two feeds are never merged into a fictional global speed.
Release dates are catalogue dates. “First seen” is when a model appeared on a public endpoint, which is the earliest anyone could call it — usually a little after the lab’s own announcement.
Venue coverage is not the whole market. OpenRouter represents one routing venue, not direct contracts, every cloud, or consumer chat apps. Hugging Face reach and provider availability are also partial views. Market Gravity is a transparent market-presence index, not a global usage or revenue estimate.
Found something wrong? Every number on this site traces to a published source above, and the sync and scoring code is in scripts/. If a figure looks off, the arithmetic is checkable rather than a matter of trust — which is the whole point of publishing this page.