A reporting layer
over evaluation
infrastructure.

Evaluation Cards is a collection of reported model–benchmark results, organized under a five-level rollout hierarchy and four interpretive signals computed over the joined record.

Corpus snapshot · August 25, 2026
6,488
Models
Tracked across reporting sources
256,190
Reported results
(model, benchmark, metric) triples
43
Reporting organizations
Distinct evaluator initiatives in this corpus
901
Model developers
Distinct model-publishing organizations
80
Benchmark families
Top of the rollout hierarchy
936
Single benchmarks
855 slices · 917 metrics

Interpretive signals

Four signals computed over each (model, benchmark, metric-path) record and aggregated to the corpus level. Per-record instances appear on every model and benchmark page.

Benchmark families

All 80
Five-level rollout hierarchy

Every score resolves to an explicit path through this hierarchy, so aggregate claims drill down to the evidence supporting them.