NODE & NORM · RESEARCH INFRASTRUCTURE
What can the surviving evidence tell us about an AI control?
Research brief · Explore a dossier · Methods & documentation · Node & Norm projects
Development release: v0.3.0-dev · Release history
An investigation says a human could intervene. A trace records an override command. Neither statement, on its own, establishes that the command changed what the system did.
The Control Evidence Corpus (CEC) makes that evidentiary boundary inspectable. It connects each control proposition to a dated source, a precise evidence location, and a preserved judgment. Its research question is whether independent reviewers can use those records to reconstruct control operation reliably.
Important
Development infrastructure · 0.3.0-dev. Three synthetic demonstrations and one unrated source-reconnaissance note are available. A separate TAE/HIT sandbox adds sixteen scripted development conditions with preserved run records. The CEC also includes an eight-packet missingness-boundary design and a reproducible 32-response mock rehearsal; it contains no live results for that design. There are zero independently coded empirical events and no approved corpus release. Software checks do not establish scientific validity.
These are invented demonstrations of the data contract. They are not study results.
| Surviving record | Proposition | Recorded value | Inspect |
|---|---|---|---|
| An override command; no downstream acknowledgement | Did the action propagate? | not_reported |
Unresolved override |
| A command and a downstream stop acknowledgement | Did the action propagate? | yes |
Acknowledged stop |
| A policy assigning an approval role; no runtime record | Was intervention attempted? | not_reported |
Declared authority |
A missing record stays distinguishable from an affirmative finding that a control did not operate. Even a documented stop leaves mitigation unresolved without evidence about consequences and alternatives.
For a real source: the Tempe reconnaissance note shows why a documented takeover route, a timed action, and a safety outcome need separate questions. It is an AI-assisted source map awaiting independent review.
The candidate gap is reproducible reconstruction of individual control opportunities from incomplete, dependent public evidence. Existing work already addresses incident uncertainty, audit evidence, human oversight, safety arguments, and AI forensics. The novelty review records substantial overlap and unresolved comparisons.
CEC earns a research contribution only if it improves what reviewers can establish, or yields a defensible finding about the limits of reconstruction.
| Test | Comparison or evidence | What would count against CEC |
|---|---|---|
| Added value | Same packets presented as a narrative, a simple checklist, and CEC | No useful gain in supported reconstruction, or excessive preparation and review burden |
| Reproducibility | Independent control enumeration and initial human ratings | Unstable control boundaries or persistent category confusion |
| Validity | External construct review; later controlled cases with known states | Reliable labels that fail to distinguish the intended control states |
| Reuse | Exact-version records with retained uncertainty and corrections | Downstream interpretation loses provenance or turns missingness into certainty |
Read the comparative pilot design and research decision gates. Thresholds and study roles remain unfilled; no favorable result is presumed.
flowchart TB
S["Sources, located evidence, and dependence"] --> P["Frozen event packet"]
P --> A["Independent initial ratings"]
A --> J["Separate adjudication and disagreement"]
J --> R["Reviewed release for downstream reuse"]
Each analytical observation binds one event, one control objective, one opportunity, and one relevant period. A source claim remains separate from a researcher judgment. Review and release steps in this diagram are research requirements; the executable demonstration uses synthetic records.
| If you want to… | Start here | Then inspect |
|---|---|---|
| Assess whether this deserves a study | Research brief | Pressure test |
| Understand a record | Example gallery | Codebook |
| Reproduce intervention and recovery tests | TAE/HIT sandbox | Run evidence |
| Help run the pilot | Pilot workbench | Sampling · Reliability |
| Inspect or extend the implementation | Reproducibility guide | Schemas |
| Use future findings | Integration contract | Data and rights |
| Review the research authority | Charter · Agenda | Governance |
Results figure and formulas · Disagreement review · Successor paired-case draft
These descriptive counts come from public synthetic packets and an AI-authored reference. The successor cases are unrun development proposals.
The four-question live result is now available separately from the earlier three-question figure. Different questions and cases prevent a direct improvement comparison.
The instruction comparison found no net primary paired agreement advantage for the shorter package, with incomplete coverage.
Jev request preparation projects the three synthetic demonstrations into bounded questions about action, propagation and mitigation. The inspectable requests preserve missingness categories and exclude existing ratings. The first live synthetic run records agreement with an AI-authored reference, invalid responses and unresolved disagreements; it supplies no independent accuracy finding. See the preparation decision for the proposed evaluation and its limits.
Python 3.11 or later; one pinned dependency. No model credentials or external inference service required.
python3 -m pip install -r requirements.txt
python3 scripts/validate.py examples/synthetic.json
python3 scripts/render_dossier.py examples/synthetic.json --view review --output build/dossier.md
python3 -m unittest discover -s tests -vThe renderer produces a readable, explicitly synthetic dossier. The full workflow reproduces every example and checks documentation links. Empirical publication remains unavailable in this development implementation.
Repository map
research/ Contribution, literature, comparative pilot, source reconnaissance
examples/ Runnable synthetic bundles, generated dossiers, blind packet views
docs/
methods/ Constructs, codebook, sampling, reliability, validation
policies/ Sources, rights, ethics, data, AI use
reference/ Reproduction, downstream contract, continuation provenance
schemas/ Explicit object contracts
scripts/ Validation, rendering, example generation, synthetic export
tests/ Research-integrity and presentation safeguards
data/ Reserved for reviewed empirical records; currently empty
The documentation index covers every governing and supporting document. The original Charter and Agenda remain at the root, with section navigation added.
Next milestone: a reviewable, staffed development pilot. Track concrete prerequisites in the roadmap. Contribute through documented review; cite the exact development commit using CITATION.cff.
