← Integrity.ai · Working Paper
Integrity Registry · Document 04

Data Pipeline &
News Tracker.

How the registry stays true: primary sources over databases, licensed access over scraping, and a human review queue that nothing skips. A static scorecard is misinformation with a delay — this document is the cure.

v0.2 · 2026‑08‑01

Source policy: the legitimacy map

Rule zero: no paywall circumvention, no scraping in violation of terms of service — not only for legal reasons, but because a credibility product cannot be built on misappropriated data. Everything load‑bearing must trace to a primary source anyway.

Tier A — free, primary, citable

SourceYieldsAccess
SEC EDGAR (full‑text search + RSS)Form D private placements (issuer, amounts, dates); S‑1 cap tables; 13F institutional holdingsFree API
DEF 14A proxy statementsExact supervoting math for Alphabet, Meta — replaces every ⚠ on founder controlFree
CourtListener / RECAPFederal dockets and filings, with alerts by party name — auto‑feeds every litigation dimensionFree API
Company newsroom / press‑release RSSFunding rounds with named investors — the raw material the paid databases themselves parseFree
FTC, DOJ, state‑AG feedsOrders, investigations, consent decreesFree
Companies House (UK) · INPI/RCS (FR)Mistral and EU‑entity filingsFree / nominal
Lab blogs & safety reportsIncident disclosures, environmental disclosures, system cardsFree
GDELT / news RSSDiscovery layer for the trackerFree API
Electricity Maps / WattTimeLive grid carbon intensity — feeds the router itselfFree tier / paid

Tier B — paid, licensed, discovery only

A Crunchbase API license (used through the API, per its terms) is worthwhile as a discovery index: it enumerates rounds and investors that we then verify in Tier A before anything becomes evidence. PitchBook likewise if budget allows. Press coverage (TechCrunch et al.) is cited and linked as reported evidence like any journalism. Registry rule: no evidence item may rest solely on a Tier B database row. Paid indexes tell us where to look; primary sources are what we cite.

The tracker

FEEDS ─────────────────► TRIAGE ────────────► REVIEW QUEUE ─────► REGISTRY
CourtListener alerts      local open model      human editor        versioned
EDGAR filing alerts       drafts candidate      verifies sources,   merge →
FTC / state‑AG feeds      evidence items:       sets status,        scores
press‑release RSS         {entity, dimension,   approves/rejects    recompute →
GDELT entity queries       polarity, weight,    (2 editors for      changelog
watchlist co‑mentions      summary, sources,    user‑safety)        entry
                           proposed_status}

Publishing architecture

braxtonjarratt.com/…/index.html   ← the essay (stable)
      │ links to
registry-scorecard.html      ← Document 01 · scores + evidence
registry-individuals.html    ← Document 02 · follow the money
registry-methodology.html    ← Document 03 · rubric, rules, schema
registry-pipeline.html       ← Document 04 · this file

next step: github.com/integrity-ai/registry             ← the living system
  entities/  evidence/  artifacts/  endpoints/  scores/  CHANGELOG.md
  (scores recomputed in CI from evidence; every change is a public diff;
   these HTML pages become rendered views of the repo)

The five HTML files above are self‑contained and relatively linked: drop them in one directory anywhere under your domain and every cross‑link works. When the GitHub repo goes live, these pages become its rendered front‑end and the version numbers start meaning something cryptographic rather than editorial.

Build order

  1. CourtListener + EDGAR + press‑release ingestion (a Claude Code project): watchlist of ~40 entities, candidate‑evidence drafting, review queue as PRs. This converts the registry from documents into a system.
  2. S‑1 readiness: parsers and templates staged so the ownership analysis re‑issues within hours of unsealing.
  3. Estimator module: EcoLogits‑derived energy/water ranges keyed to endpoint region + live grid feeds — shared between the registry's endpoint files and the router prototype.
  4. Editorial board: two reviewers minimum, no lab affiliations, named publicly, before anything drops the DRAFT banner.