← Integrity.ai · Working Paper
Integrity Registry · Document 03

Methodology & Schema.

How scores are made, contested, and corrected. The registry's credibility rests on this document more than on any score it produces.

Schema v0.2 · 2026‑08‑01 · license: CC‑BY‑SA (proposed)

Six design principles

PRINCIPLE 1Scores are computed, never asserted. Every score is a function of dated, cited evidence items. Change the evidence, the score recomputes in CI, and the changelog explains why. No editor can hand‑set a number.
PRINCIPLE 2Three layers, separately scored. The artifact (model: provenance, license, openness), the endpoint (who serves it, where, on what grid — energy and water live here), and the entities (who receives money and data, traced through an ownership graph). A routing decision composes all three. Llama on your laptop ≠ Meta's hosted API ≠ Meta the corporation.
PRINCIPLE 3 · THE ANTI‑SILENCE RULEMissing disclosure is penalized; honest disclosure never scores worse than silence. Non‑disclosers get pessimistic defaults on impact estimates and low transparency scores. In the safety dimension: containment failures score negative, voluntary disclosure scores positive, and a concealed incident later revealed by outsiders scores worst of all. A registry that punishes honesty teaches the industry to stop telling us.
PRINCIPLE 4Every environmental number carries its system boundary. Google's 0.24 Wh (datacenter operation, incl. idle) and Mistral's ~45–50 mL water (full lifecycle incl. upstream) describe different boundaries and are not directly comparable. Figures without a stated boundary are treated as estimates with widened uncertainty. Ranges always; false precision never.
PRINCIPLE 5Evidence decays; scores are versioned. A 2023 scandal fades on a half‑life; an active contract doesn't decay while active. Scores carry timestamps and confidence, and registry versions are immutable so any citation can be checked against what was known when.
PRINCIPLE 6Behaviors, not characterizations. The registry records "removed military‑use ban," "settled for $1.5B," "FTC issued a 6(b) order" — never "unethical," never "evil," never "cover‑up" in its own voice. This is both an epistemic standard and the defamation shield.

Evidence status & language rules

StatusMeaningEffect on scores
verifiedPrimary source in hand (filing, court record, first‑party disclosure)Full weight
reportedCredible press; primary source pendingReduced weight; queue for verification
disputedContested or single‑stream; parties disagreeHeld from scoring; displayed as disputed
recalled‑unverified ⚠Entered from research memoryZero weight until verified

Litigation adds its own ladder: a filed lawsuit is an allegation; a settlement is a resolution (admission status noted — most include none); only rulings, admissions, and regulator findings are established. In the user‑safety dimension: victims' names appear only as they appear in public case captions, no graphic detail ever, and every entry requires two editors' sign‑off. The registry documents accountability, not tragedy.

The ten dimensions

DimensionWhat it measures (anchors abridged; full rubric in schema)
Military & surveillance90–100: contractual public red lines (no autonomous weapons, no mass surveillance) that survive pressure. 30–59: defense contracts on broad "lawful use" terms. 0–29: kill‑chain participation, surveillance products, or bans removed then weapons‑adjacent partnerships signed.
Data provenance90–100: licensed/consented corpus, creators compensated. 60–89: mixed corpus with licensing programs and remediation paid. 0–29: pirated corpora without remediation; contributory infringement.
Environmental practiceWhat they do: 24/7 clean‑energy matching, verified water positivity, siting — vs. unpermitted generation and water‑stressed siting without mitigation.
Environmental transparencyWhat they disclose: audited per‑query lifecycle figures with stated boundary (90+) down to no first‑party figures at all (0–29). Scored separately from practice so silence can't impersonate cleanliness.
LaborAnnotation/moderation supply‑chain pay and conditions, pay‑equity disclosure, union posture.
GovernanceAccountability structures: trusts, PBC status, board independence — vs. supervoting shares, sole control, nonprofit‑conversion conduct, tax structures.
Openness & researchOpen weights, system cards, safety research publication, incident disclosure.
Jurisdiction & sovereigntyLegal regime of data access, retention, training‑on‑user‑data defaults, censorship layers. Endpoint‑dependent: collapses for open weights run locally.
Safety practice & incidentsContainment failures, autonomous‑capability incidents, third‑party misuse — scored on failure severity AND disclosure conduct (anti‑silence rule, Principle 3).
User safety & product harmDocumented user harm, regulator actions, design‑choice evidence (engagement vs. safety), and remediation — under the strictest evidence rules in the registry.

How a routing decision uses this

endpoint_utility =
    w_quality · predicted_quality(task, artifact)
  + w_cost    · cost_score(endpoint)
  + w_energy  · energy_score(endpoint, live_grid)   # per SUCCESSFUL task,
  + w_water   · water_score(endpoint, region)        #   incl. escalation risk
  + w_ethics  · Σ_d  u_d · entity_score(beneficiaries(endpoint), d)

subject to: hard vetoes ("never route to military_score < 30"),
            budgets, data‑sensitivity ("this never leaves the device"),
            agentic‑containment weighting when a task grants tools/network

beneficiaries(endpoint): weighted traversal of the ownership graph —
  spend → operator (1.0) → parents/investors (× stake × passthrough),
  cycles damped; control structures weighted over passive capital

Conflict‑of‑interest policy

Drafting assistance for v0 came from Claude (Anthropic) — a scored entity whose model family also appears in the July 2026 safety evidence. Accordingly: Anthropic‑related entries carry a standing COI flag and receive adversarial review by editors with no Anthropic relationship; no AI‑drafted claim can reach verified status without a human checking the primary source; and the long‑term goal is an editorial board with zero lab affiliations and a triage pipeline running on open models. The COI is disclosed everywhere the affected content appears, including this sentence.

Corrections

Any person or company named may submit corrections with sources to braxton@integrity.ai. Corrections get their own changelog entries, published alongside what changed and why. Being scored low is not grounds for removal; being scored wrong is grounds for immediate correction.

Full machine‑readable schema (v0.2 YAML)
# INTEGRITY REGISTRY — Schema v0.2 (abridged for display; canonical file in repo)
dimensions: [military_surveillance, data_provenance, environment_practice,
  environment_transparency, labor, governance, transparency_research,
  jurisdiction_sovereignty, policy_lobbying, safety_practice, user_safety]

entity:   {id, name, type: lab|parent|cloud_operator|investor|individual|nonprofit,
           jurisdiction, ownership: [{holder_entity_id, stake_pct, control_notes, source}],
           conflict_flags: []}

evidence: {id, entity_id, dimension, polarity: positive|negative|mixed,
           weight: 1-5, date, decay_halflife_months,   # null = no decay while active
           summary,            # documented behavior only, neutral language
           sources: [url],     # primary preferred
           status: verified|reported|disputed|recalled_unverified}

artifact: {id, developer_entity_id, license, open_weights,
           provenance_evidence_ids, distillation_lineage}

endpoint: {id, artifact_id, operator_entity_id, region, grid_zone,
           energy_per_mtoken_wh: {low, high, method, system_boundary, source},
           water_per_mtoken_ml:  {low, high, method, system_boundary, source},
           data_retention, trains_on_user_data_default}
           # method: first_party_disclosed | third_party_measured |
           #         parameter_estimated | pessimistic_default
           # system_boundary: gpu_only | datacenter_incl_pue | lifecycle_incl_embodied

score:    {entity_id, dimension, score: 0-100, confidence: low|med|high,
           evidence_ids, computed_at, registry_version, peer_percentile}

user_safety.evidence_classes:
  [filed_suit, regulator_order, settlement, ruling,
   admission_or_disclosure, remediation]