Evaluator Bench
Evidence current to 15 Sep 2026

Independence, scored.

Evaluators, government institutes, vendors and benchmarks — plus a watchlist of expected entrants — scored on independence from the labs they test. Every score is built from public signals and ledger rows you can inspect.

Money is traced by hops: hop 0 is a lab; hop 1 is a lab investor, observer, board member, employee, or contractor; higher hops run through principals or funders. "Lab-tied" means hop 0 or hop 1. Scores are gated ordinal projections, not procurement truth.

Set the weights for your situation; the ranking updates as you move them.

What matters to you

Domain
Role
Type
EvaluatorDomainsIndependenceScore

Watchlist (not ranked)

Expected entrants scored on the same rubric but excluded from the ranking and from every average: a hypothetical composite and an initiative announced with no evaluations yet.

The rubric

Eight dimensions, each scored 0 to 4 from public evidence, then weighted. The dimensions follow the AI Evaluator Forum's AEF-1 operating conditions and the financial-audit independence rules that Illinois SB 315 imports for frontier AI, with two additions the field tends to skip: who owns the evaluator, and whether it sells fixes to the companies it grades.

Signals

What raises and lowers an independence score. The list is deliberately concrete: each item is something you can verify from a filing, a contract term, a system card, or a published policy.

Raises the score

  • A written policy refusing money from frontier developers, including donations directed by their staff.
  • A published conflict-of-interest policy: no outcome-contingent fees, no side grants or investments from labs being evaluated, mandatory recusal for anyone with a financial interest.
  • No single funder above a stated share of budget; funders disclosed by name.
  • Publication rights fixed in the contract before work starts, with redaction limited to security-sensitive or privileged material and a public statement of what was redacted.
  • A record of publishing findings the lab did not like: scheming, reward hacking, shutdown resistance, safeguard failures.
  • Open evaluation code and task suites so results can be re-run by others.
  • On-site, weights-level, or training-time access rather than a few weeks on an API.
  • Access that does not depend on the lab's goodwill: statute, regulation, or a court-enforceable agreement.
  • Membership in a standards body (AEF-1, AVERI pilots) and disclosure of operating conditions for each evaluation.
  • Cooling-off periods before staff move to labs; equity in labs disclosed or prohibited.
  • A client base spread across many developers, so no lab is a dominant revenue source.
  • Double-blind or secure-enclave protocols that let the evaluator work without the lab seeing the test items.

Lowers the score

  • The lab pays for the evaluation of its own model. Common, and the single most under-discussed conflict in the field.
  • Investors shared with the labs: a venture firm that backs both the evaluator and OpenAI, or a chip vendor buying the evaluator while anchoring a lab's IPO.
  • Majority ownership by a frontier developer, or status as a business unit of one.
  • Leadership holding a board seat, safety-committee chair, or advisory role at a lab the organization evaluates, even with recusal.
  • Selling defenses, guardrails, or monitoring products to the same labs it evaluates: the audit-plus-consulting problem that Sarbanes-Oxley separated in accounting.
  • A contract that forbids disclosing who funded a benchmark or evaluation.
  • The lab sets the scope, the time window, and which questions are out of bounds.
  • The lab reviews drafts for tone and emphasis, not only for security redactions.
  • Only aggregated or lab-summarized results reach the public.
  • Access that the lab can decline to renew with no consequence.
  • Political or budgetary dependence that can redirect a government institute's mandate within a year.
  • Co-authoring research with a lab while also serving as its external evaluator.
  • Free tokens and compute from the lab, when they are a material share of operating capacity.
  • Heavy talent flow in both directions between the evaluator and the labs.
  • No track record: a pledge to evaluate is not an evaluation.

How good is the evidence

The scores measure what the public record shows, and the public record is largely what the evaluators say about themselves. Among the 112 unique sources cited by 295 signals for the ranked population, 60 (54%) are tier-3 self-published and 10 (9%) are tier-1 sources (regulatory filings or public indexes). 41 of 295 signals (14%) carry an exact quoted span. The evidence mix varies: AVERI has 5 self-published sources out of 7, with 0 tier-1; SaferAI 2 of 5, with 1 tier-1; and METR 7 of 21, with 5 tier-1. An evidence-based score rewards silence, and self-published sources can be replaced only where independent reporting exists — which, for most evaluators, it does not. That scarcity is a finding, not a data gap to be patched: the field's independence cannot be verified from outside the field. Scores built mostly on self-report are flagged by their evidence tier on every scorecard.

Cases that set the bar

Six engagements from the last two years, read for what they reveal about each dimension. Longer treatments and sources are in the repo under paper/.

Where we are

Every mature high-stakes industry grew a third-party assurance layer, and most of them started the way AI has: voluntary, paid by the assessed party, with the assessed party choosing scope. The ladder marks the year each regime first reached each stage. Tap a cell to see the milestone behind it and its source; tap a regime name for its full history.

under 10 years into the regime 10 to 50 over 50 not reached rollback assessed party absorbed the assessment Download the static figure

A caveat the ladder cannot show: it records first arrivals, so a one-clause rule in one state fills a cell the same way Sarbanes-Oxley does, and it hides the order things happened in. The three views below fix both.

Read financial audit left to right and you get 158 years from the first mandate to the first rule separating auditors from consultants. Read the AI row and you get four. Across regimes, rules have followed public failures within a few years in the coded record (triggers were selected with hindsight, so this is a description, not a causal finding); the stretch before the first failure is where a regime's shape is decided. The oversight and independence-rule columns are the ones AI has not filled. Data and method: data/industries/, bench/timeline.py, paper/.

Paths, responses, and mechanisms

Three views that keep the nuance the ladder drops. Hover or tap any glyph, point, or cell to read what it stands for and where it comes from.

The order, not the calendar

Each regime as a sequence of moves. Reversals show up in red, the assessed party absorbing the assessment function in amber. The right column ranks regimes by how closely their opening moves match frontier AI's so far, using edit distance over the sequences.

Paths: the order each regime did things, ignoring the calendar Consecutive repeats collapsed. Right column: similarity of each regime's opening to frontier AI's path so far (1 = identical, 0 = nothing shared). voluntary trigger mandate standards oversight independence access/publication delegation payer shift rollback Financial statement audit M1844 R1856 V1880 M1900 A1934 T1938 S1939 D1990 T2001 O2002 I2002 similarity 0.36, opening MRVMATSDTOI Ship classification V1760 O1968 T1999 I2009 similarity 0.22, opening VOTI Boilers and pressure vessels V1866 T1905 M1908 S1914 O1919 similarity 0.22, opening VTMSO Electrical and consumer product safety V1894 S1903 M1972 O1988 T2007 M2008 similarity 0.33, opening VSMOTM Pharmaceutical safety and efficacy M1906 T1937 M1938 T1961 D1962 M1962 S1962 T1976 O1978 P1992 I1998 A2007 similarity 0.22, opening MTMTDMS Aircraft certification M1926 T1956 M1958 S1965 D2005 T2018 I2020 similarity 0.22, opening MTMSDTI Credit rating agencies V1909 D1970 T1970 O1975 T2008 I2010 similarity 0.44, opening VDTOTI Nuclear safety and safeguards M1957 A1970 I1974 A1978 V1979 T1979 A1997 similarity 0.22, opening MAIAVTA Automobile crash safety V1959 T1965 M1966 S1967 A1979 similarity 0.33, opening VTMSA Food safety M1906 T1993 M1996 V2000 T2008 O2011 similarity 0.22, opening MTMVTO Cybersecurity assurance S1985 O1999 M2004 O2011 T2013 similarity 0.11, opening SOMOT Sustainability (ESG) reporting assuran S1997 V2003 T2015 M2022 S2024 similarity 0.33, opening SVTMS Dietary supplements (US) R1994 V2001 T2003 A2006 S2007 similarity 0.33, opening RVTAS Crypto exchange proof of reserves V2014 T2014 R2022 T2022 M2023 similarity 0.44, opening VTRTM Platform and hiring-algorithm audits T2018 M2021 A2022 I2023 T2025 similarity 0.22, opening TMAIT Frontier AI evaluation V2022 D2023 V2024 T2025 S2025 A2025 T2026 M2026 A2026 reference path: VDVTSATMA Source: data/industries/. Generated by bench.figures.

Fast is not the same as strong

Every trigger incident paired with the next rule that followed. Most responses arrive within three years, and most are mandates or standards rather than structural changes; the two structural responses in the set are the PCAOB after Enron and lab inspection under Good Laboratory Practice after the IBT fraud. Frontier AI's two responses so far score 2 and 1.

After a failure: how fast the next rule came, and how strong it was Each point is one trigger incident. Hollow points: no rule followed. Frontier AI in teal. 1 2 3 4 0 5 10 15 20 years from trigger to next rule instrument strength (1 disclosure to 4 structural) no rule

Who pays, selects, sees, publishes, and watches

The configuration the ladder cannot show. Read the AI row against the others: the only other regime with the same five-cell pattern is dietary supplements, the regime that removed pre-market review by statute. The match is on a coarse categorical coding, so it is suggestive rather than conclusive.

Mechanisms today: the incentive configuration of each regime Red marks the configuration that concentrates control in the assessed party. Hover a cell for the state before reform. Who pays Who selects What they see What the public reads Who watches assessors Financial statement audit assessed party assessed party deep public regulator * Ship classification assessed party assessed party deep summary or mark regulator * Boilers and pressure vessels opposing exposure opposing exposure deep summary or mark regulator * Electrical and consumer product safe assessed party assessed party deep summary or mark regulator * Pharmaceutical safety and efficacy assessed party regulator or state * deep public * regulator * Aircraft certification assessed party regulator or state * deep summary or mark * regulator Credit rating agencies assessed party * assessed party * deep * public regulator * Nuclear safety and safeguards assessed party regulator or state embedded * public regulator Automobile crash safety opposing exposure * opposing exposure * deep public * regulator * Food safety assessed party assessed party deep summary or mark * regulator * Cybersecurity assurance assessed party assessed party deep summary or mark regulator * Sustainability (ESG) reporting assur assessed party assessed party shallow public regulator * Dietary supplements (US) (same as AI) assessed party assessed party shallow summary or mark none Crypto exchange proof of reserves assessed party assessed party shallow public regulator * Platform and hiring-algorithm audits assessed party assessed party deep * summary or mark * regulator * Frontier AI evaluation assessed party assessed party shallow summary or mark none * changed since the regime's origin. Source: data/industries/ mechanisms. Generated by bench.figures.

Money and ties, every evaluator

The ledger view: each evaluator's traced inflows by distance from a frontier lab, ties within two steps, and how much of the record has been re-derived from primary sources rather than imported. Read the confirmed column before the others; most rows are still leads. Rows and sources: data/ledger/. Rebuild with python -m bench exposure --json.

Two patterns worth reading off the matrix. First, the two organizations with the most complete self-disclosure, Transluce and AVERI, score lower on funding than several that disclose less; an evidence-based score rewards silence unless the empty cells are read as gaps, which is why the confirmed and second-hop columns sit beside the rows. Second, the same three or four funders (Coefficient Giving, the Survival and Flourishing Fund, the Audacious Project, the EU AI Office) sit behind most of the nonprofit evaluators, so a question about one evaluator's independence is often a question about one funder's.

Traced money by distance from a frontier lab, per evaluator Cell: number of inflow rows at that distance; the amount below is summed within the largest measure at that distance. Undisclosed amounts count as rows only. Right: board, advisor, donor, investor, or office ties within two steps of a lab, and the share of rows re-derived from sources (confirmed) rather than imported. Lab (0) Lab-tied (1) Two steps (2) Three steps (3) Public budget Donor not public TiesConfirmedNegatives2nd hop METR 1in_kind $0.4M 3recommendation $0.8M 3commitment $17.0M 6 5/7 6 5/6 RAND Corporation 1recommendation $1.0M 1commitment $38.0M 0 2/2 0 2/2 FAR.AI 1* 2recommendation $0.9M 1 2*grant $0.5M 0 4/6 0 5/6 Redwood Research 1daf_grant $1.3M 1 0 1/2 0 2/2 Irregular (formerly Pattern Labs) 1* 1investment $80.0M 1 2/2 1 2/2 Gray Swan 1* 1investment $40.0M 0 2/2 0 1/2 SecureBio 1* 1grant $17.2M 1recommendation $0.8M 0 3/3 0 3/3 Apollo Research 1* 1* 1 2/2 0 1/2 Epoch AI 4* 1daf_grant $0.6M 3grant $24.5M 0 4/8 0 7/7 Palisade Research 2grant $2.1M 0 1/2 0 2/2 Transluce 2* 1* 1 3/3 2 3/3 Andon Labs 1* 0 1/1 0 1/1 UK AI Security Institute 2*grant $7.5M 1* 0 2/3 1 2/2 US CAISI (NIST) 2grant $25.0M 0 1/2 2 1/1 EU AI Office 1 0 1/1 1 1/1 AVERI 1* 2* 1* 0 4/4 2 3/4 SaferAI 1* 1recommendation $0.3M 1 0 2/3 1 3/3 Scale AI (SEAL) 1investment $14.3B 1 1/1 1 1/1 MLCommons 1* 0 1/1 1 1/1 Hugging Face Open Alignment Initiati 1investment $12.9B 3 1/1 1 1/1 Center for AI Safety 2recommendation $1.4M 0 2/2 0 1/1 Microsoft AI Red Team 1* 0 1/1 0 1/1 Dreadnode no ledger rows yet; funding and ties not traced Humane Intelligence no ledger rows yet; funding and ties not traced Holistic Agent Leaderboard (Princeto no ledger rows yet; funding and ties not traced Assurance firms (Big Four and peers) 0 0/0 1 0/0 EquiStamp 1 0 1/1 1 1/1 Nemesys Insights 1* 1* 0 2/2 1 2/2 * includes undisclosed amounts. Source: data/ledger/. Generated by bench.figures. Imported rows are leads until confirmed.

The graph itself

Every entity in the ledger placed by its distance from a lab, with money in teal and roles in amber. Dashed lines are imported rows not yet re-derived. Hover a name to isolate its neighbourhood; click it to open the entity page and walk the graph hop by hop.

The funding graph, laid out by distance from a lab Teal lines are money (transfers); amber lines are roles. Solid: confirmed; dashed: imported. Hover a name to isolate it and its neighbours; hover a line for the row. Every element is a row in data/ledger/. Labs Direct lab ties Two steps Three or more Public or unattributed Evaluators Amazon Anthropic G42 Google / Google DeepMind Meta Microsoft OpenAI Thinking Machines Lab xAI AI Safety Fund (Frontier Model Alexandr Wang Andreessen Horowitz Anthropic employees (personal Ben Mann D.E. Shaw Ventures David Farhi Dustin Moskovitz Frontier-lab employees and alu Holden Karnofsky Jaan Tallinn Jeffrey Ladish Macroscopic Ventures (formerly Miles Brundage Neil Chowdhury Nvidia OpenAI employees (personal hol OpenAI Foundation Peter Mattson Sequoia Capital Zico Kolter Adam Gleave Artificial Intelligence Underw Clement Delangue Coefficient Giving (formerly O Good Ventures Foundation Jason Droege Longview Philanthropy Marco Mascorro Paul Christiano Rajiv Dattani Survival and Flourishing Fund Conrad Stosz Dan Hendrycks Mike McCormick Alec Radford Alignment Research Center Constellation Research Center ELMA Philanthropies Emerson Collective EU budget (Digital Europe Prog Fifty Years (50Y) Founders Pledge Founders Pledge (frontier AI f Gates Foundation Halcyon Futures Hillspire (Schmidt family offi Hudson River Trading Jacob Hilton MacKenzie Scott Magarac Venture Partners Obvious Ventures Redpoint Ventures Renaissance Philanthropy Salesforce Ventures Samsung Next Schmidt Sciences Silicon Valley Community Found Skoll Foundation Snowflake Ventures Swish Ventures Sympatico Ventures The Audacious Project (TED) UK government (DSIT) US government (NIST appropriat Valhalla Foundation Vanguard Charitable Wing Venture Capital Y Combinator Andon Labs Apollo Research AVERI Center for AI Safety Epoch AI EquiStamp EU AI Office FAR.AI Gray Swan Hugging Face Open Alignment In Irregular (formerly Pattern La METR Microsoft AI Red Team MLCommons Nemesys Insights Palisade Research RAND Corporation Redwood Research SaferAI Scale AI (SEAL) SecureBio Transluce UK AI Security Institute US CAISI (NIST)

Browse by type

Every item has its own page. Dockets: contestable-claim drafts for future Epistemedia review (none submitted; drafts are not evidence). Status: coverage, confirmed shares, open questions, data downloads. Evaluators: the scorecard table and a page per organization with its dimensions, signals, sources, ledger rows and a focused funding graph. Entities: every funder, lab, investor and person in the ledger, with money in and out and roles held. Regimes: the sixteen assurance regimes with milestones, mechanisms and path signatures. Sources: each cited source with its tier, audit status, and everything that cites it.

How the scores are made

Each evaluator gets a 0 to 4 on eight dimensions using only public evidence: filings, funding announcements, system cards, published policies, contracts described in reports, and press. The weighted total is scaled to 100. Scores are gated by default: any 0 on a dimension caps the total at 40, any 1 caps it at 60 — independence has floors, not just averages. Uncheck the gate for the raw compensatory score (sensitivity analysis). The lab-procurement preset is the pre-registered confirmatory view; the other presets are sensitivity checks, not alternative truths. Confidence tags flag where evidence is thin. An entry is not an endorsement, and a low score is not an accusation; it means the public record does not yet show the safeguards that would earn a higher one.

Independence is one axis. Competence, domain coverage, staffing, and turnaround are others, and a highly independent evaluator with no cyber team is the wrong pick for a cyber evaluation. Use the domain filters alongside the score.

Scores move when evidence moves. Send a contract term, a policy, or a correction and the entry updates with the source attached. This dataset was built by a single curator with AI assistance (Claude, Anthropic — itself a frontier developer in the ecosystem scored here); curator disclosure, including the fields still awaiting confirmation, is in DISCLOSURE.md. Corrections received through 29 Sep 2026 are filed as signals with the date received, and evaluators whose scores changed after publication receive a right-of-reply packet before the next batch.

How to submit evidence Status and data Dockets