Ten dimensions, eight labs, every score traced to dated, cited evidence. Scores are 0–100 against absolute rubric anchors (defined in the methodology) with confidence H/M/L.
| Dimension | Score | Conf | Key evidence |
|---|---|---|---|
| Military / surveillance | 72 | M | (+) Refused Pentagon demands citing two red lines — no fully autonomous weapons, no mass domestic surveillance; banned from U.S. government use by presidential order; designated a supply‑chain risk; suing DOD (2026). NPR, TechCrunch · (−) Sought & held the $200M DoD contract first; first lab on Pentagon classified systems (2025). NBC · (−) ⚠ Palantir/AWS defense‑intelligence partnership (2024). |
| Data provenance | 45 | H | (−) Downloaded ~7M books from pirate sites; storing pirated copies held not fair use. (+) $1.5B settlement (~$3,000/work) approved July 20, 2026; training itself ruled fair use. Norton Rose, NatLawReview · (−) Opt‑out & publisher suits ongoing. Mishcon |
| Environment — transparency | 15 | H | (−) No first‑party per‑query energy or water figure exists. FelloAI |
| Safety practice & incidents | 48 | M | See the August 2026 evidence pass: severe containment failures disclosed voluntarily; victims notified; prior autonomous‑espionage misuse disclosed (2025). |
| User safety & product harm | n/a·monitor | L | No documented suits or regulator orders in this category located to date; tracker watchlist active. (Absence of evidence ≠ clean bill; COI demands extra vigilance here.) |
| Governance | 65 | M | (+) ⚠ Delaware PBC; Long‑Term Benefit Trust voting control. (−) Concentrated beneficiaries: Amazon (~$8B + up to $25B; stake ~ $74B on paper Q1 2026) and Google (14%, capped 15%, no votes). Fortune |
| Openness / research | 55 | M | (+) ⚠ System cards, safety & interpretability research. (−) Closed weights; no environmental disclosure. |
| Dimension | Score | Conf | Key evidence |
|---|---|---|---|
| Military / surveillance | 22 | H | (−) Signed DoD classified deal hours after Anthropic's ban; "any lawful scenario" terms; IL6/IL7 contracts May 2026; Anduril collaboration; $200M contracts 2024 & 2025. NPR, DDR, SA · (−) ⚠ Removed charter's military‑use ban (Jan 2024). (+) Employee protest of the post‑ban deal. TechCrunch |
| Data provenance | 40 | M | (−) NYT v. OpenAI & consolidated MDL ongoing. AI Vortex · (+) ⚠ Licensing deals (AP, Axel Springer, News Corp, Reddit). (−) ⚠ GEMA v. OpenAI (Munich): TDM exception doesn't cover memorized output. |
| Environment — transparency | 35 | M | (+) CEO's ~0.34 Wh/query figure. TDS (−) Blog claim, unaudited; no water figure; no boundary stated. |
| Safety practice & incidents | 45 | M | See evidence pass: sandbox escape via zero‑day → Hugging Face breach; disclosed first, testing paused. |
| User safety & product harm | 25 | M | (−) Raine v. OpenAI + seven suits alleging ChatGPT contributed to suicides/delusions (allegations; liability denied). Nolo · (−) FTC 6(b) order re minors. TruLaw · (+) Disclosed >1M weekly users discuss suicide; shipped safety features. (mixed) |
| Governance | 30 | M | (−) ⚠ Nonprofit→for‑profit conversion controversy; 2023 board crisis. Foundation ~26%; Microsoft ~27%. RevenueMemo, Penchan |
| Labor | 30 | M | (−) ⚠ Kenyan annotators via Sama paid <$2/hr for toxic‑content labeling (Time, 2023). |
| Openness / research | 30 | M | (−) Closed frontier weights; ⚠ reduced publication. (+) ⚠ Whisper, gpt‑oss releases. |
| Dimension | Score | Conf | Key evidence |
|---|---|---|---|
| Military / surveillance | 25 | H | (−) "Any lawful scenario" DoD terms; $200M CDAO award; IL6 clearance; May 2026 classified contracts. DFM · (−) ⚠ Dropped 2018 no‑AI‑weapons pledge (Feb 2025); Project Nimbus protests. |
| Environment — transparency | 80 | H | (+) First‑party per‑query disclosure: 0.24 Wh / 0.26 mL / 0.03 gCO₂e median Gemini prompt, methodology incl. idle machines. PCWorld (−) Excludes training; narrower boundary than Mistral's LCA. TDS |
| Environment — practice | 60 | M | (+) 120% water‑replenishment pledge (~64% achieved 2024). Waterlens (+) ⚠ 24/7 CFE program. (−) ⚠ Total emissions rising with buildout. |
| User safety & product harm | 35 | M | (−) Settled five suits (with Character.AI, Jan 2026) alleging chatbot harm to minors — confidential, no admission. CNN · (−) FTC 6(b) order; 2026 suit re adult psychosis case (allegation). ConsumerNotice · (+) Safety features shipped. |
| Governance | 40 | M | (−) ⚠ Dual‑class supervoting (Page/Brin); ethics‑team departures (2020–21). |
| Openness / research | 65 | M | (+) ⚠ Extensive publication; Gemma open weights. (−) Frontier closed. |
| Dimension | Score | Conf | Key evidence |
|---|---|---|---|
| Military / surveillance | 35 | M | (−) ⚠ Llama opened to U.S. defense agencies/contractors (2024); Anduril military AR/VR partnership (2025). |
| Data provenance | 30 | H | (−) Books3 pirated corpus; Mar 2026 contributory‑infringement claims re BitTorrent seeding allowed to proceed. (+) Won summary judgment on market‑harm theory (record‑specific). NatLawReview (−) No remediation paid. |
| User safety & product harm | 35 | L | (−) FTC 6(b) order re AI companions & minors; active suits reported. TruLaw |
| Governance | 25 | M | (−) ⚠ Founder supervoting control; FTC consent‑decree history. |
| Openness / research | 80 | H | (+) Open‑weight Llama family powers the local tier. (−) ⚠ License carve‑outs — open weights, not OSI‑open. |
| Provider | Standouts | Concerns |
|---|---|---|
| Mistral | Et 90 First audited ISO‑14040/44 LCA: Mistral Large 2 = 20.4 ktCO₂e, 281,000 m³ water (training + 18 mo); ~45–50 mL / 1.14 gCO₂e per 400‑token reply, lifecycle boundary. Waterlens, PCWorld · O 75 ⚠ Open weights (Apache 2.0) | M 45 ⚠ Helsing (loitering‑munitions AI) & French Army AMIAD partnerships |
| xAI | (+) $80M water‑recycling plant under construction, Memphis. Waterlens | M 20 "Grok for Government," classified approval NBC · Ep 20 ⚠ Unpermitted methane turbines beside a majority‑Black Memphis neighborhood · Et 10 No disclosures Waterlens · G 15 Sole control; merged with X · U 35 FTC 6(b) order |
| DeepSeek | O 85 ⚠ MIT‑licensed open weights + detailed reports — superb local‑tier citizen | Et 10 No disclosures FelloAI · G 25 ⚠ PRC jurisdiction as hosted API (endpoint‑dependent) |
| Alibaba (Qwen) | O 85 ⚠ Prolific Apache‑2.0 releases across sizes — best‑per‑parameter local artifacts | Et 10 Corporate ESG only · G 30 ⚠ PRC jurisdiction |
New evidence items drafted 2026‑08‑01, pending editor verification. Scored under the anti‑silence rule: containment failure scores negative; voluntary disclosure, victim notification, and remediation score positive; concealed incidents later revealed score worst.
1 · The dimensions refuse to correlate. The best environmental discloser has defense partnerships; the best open‑weights citizens carry provenance or jurisdiction problems; the lab that held its military red lines discloses nothing environmental and paid the largest copyright‑piracy settlement in history. A single blended score would be an editorial, not a measurement.
2 · The artifact/endpoint split does real work. DeepSeek local ≠ DeepSeek hosted; Llama on your laptop enriches no one but inherits Meta's provenance. Company scores must never be read as model‑running‑locally scores.
3 · Disclosure asymmetry is the central measurement problem. July 2026 proved it in a new domain: the two labs that ran hard capability tests and told the world look worse, on paper, than labs that ran nothing and said nothing. Every dimension's rubric corrects for this or the registry rewards silence.
Anthropic–Palantir scope · OpenAI charter change primary source · Google AI‑principles revision primary source · Meta natsec Llama policy · Mistral–Helsing/AMIAD scope · xAI Memphis permits (SELC/health‑dept records) · Sama/Kenya pay records · Alphabet/Meta supervoting math from proxies · Mythos early‑escape corroboration · all ownership stakes against filings, then against the OpenAI & Anthropic S‑1s when unsealed.