Skip to main content
← BUILD 001COMPLETE RECONSTRUCTION · 2026-10-02
CONSENSUS LAUNCH PILOT / REVISE

Human Review Design Framework

A deterministic browser-local diagnostic that separates “a human was present” from the conditions required for meaningful review: competence, timing, evidence, authority, escalation, records, capacity, and evaluation.

RELEASE DECISIONREVISE

Keep and launch as a flagship after observed comprehension testing. The working capability and reasoning are strong; the unvalidated score must remain visibly heuristic.

01 / EXACT QUESTION

What is this prototype trying to learn?

Can a team describe an AI-assisted workflow precisely enough to determine whether its human reviewer can understand the case, intervene before consequence, change or stop the outcome, escalate uncertainty, and leave an inspectable record?

What exists now

  • A working A-route tool and a separate interactive B-route exist.
  • The score is deterministic and its implementation is inspectable.
  • A local JSON record preserves the assessment and engine version.
  • Primary-source provenance and exact claim boundaries are published.
  • Adversarial, privacy, keyboard, responsive, metadata, and motion fixtures exist.
02 / PEOPLE + AUTHORITY

Who does what—and who is allowed to decide?

Workflow owner

Describes the real AI-assisted decision and is accountable for implementing any redesign.

Reviewer

Needs domain competence, usable evidence, time, independence, and actual intervention authority.

Affected person

May bear the consequence. The prototype does not ask them to supply personal data or claim to represent their experience.

Escalation owner

Receives cases the reviewer cannot resolve and has authority to pause, investigate, or redirect the workflow.

Assurance owner

Checks disagreements, overrides, errors, complaints, and outcomes over time.

Prototype

Asks structured questions, applies disclosed arithmetic, prioritizes gaps, and exports a local record. It never approves the underlying AI system.

Authority boundary

  • The tool may identify missing control elements and suggest design actions.
  • Only the organization can assign a qualified reviewer, intervention rights, stop authority, workload, evidence access, and escalation ownership.
  • A “Strong” result cannot authorize deployment, certify compliance, or replace legal, safety, domain, worker, or affected-person review.
03 / TECHNICAL SYSTEM

From input to output, without magic.

01

Intake

A native HTML form collects a use description, consequence band, and eight control answers. Optional detailed fields capture reviewer, evidence, standard, trigger, impact, and failure path.

02

Normalization

Each answer is mapped deterministically to 0, 1, or 2. Missing answers map to 0.

03

Assessment

Eight dimensions are summed to 16. Product thresholds are Strong ≥13, Partial ≥8, Weak <8.

04

Priority pass

Every sub-2 dimension becomes a repair candidate. High-consequence cases move evidence, authority, and timing weaknesses first.

05

Renderer

The browser displays the grade, reasons, prioritized fixes, and a human-readable review protocol.

06

Local record

An explicit action downloads JSON. No account or server submission is required.

Exact decision procedure

  1. Read competence, timing, evidence, authority, escalation, record, capacity, and testing.
  2. Map the strongest declared condition to 2, a partial condition to 1, and absent, late, powerless, unsupported, or unanswered to 0.
  3. Add the values. There is no model inference, hidden weighting, probability, or external API call.
  4. Assign 13–16 Strong, 8–12 Partial, or 0–7 Weak. These cutoffs are RN heuristics.
  5. Create one repair statement for every dimension below 2.
  6. For High consequence, place evidence, authority, and timing gaps ahead of other repairs.
  7. Generate a protocol using only original declarations; preserve unspecified fields as unspecified.
  8. Keep the result in client state and export only on explicit request.

State model

UNANSWEREDThe form is available; no assessment exists.
WEAK / 0–7Multiple control conditions are missing or weak.
PARTIAL / 8–12Some conditions exist, but at least two points remain unresolved.
STRONG / 13–16Most declared conditions are strong. This is not a safety, legality, effectiveness, or deployment decision.
EXPORTEDThe visitor explicitly downloaded the local JSON record.
04 / TRUTH BOUNDARIES

What is real, synthetic, missing, or prohibited?

REAL

The form, scoring code, priority logic, protocol, keyboard path, local export, and automated fixtures are implemented.

SOURCE-DERIVED

The questions synthesize recurring oversight concepts from NIST, EU, and ICO materials.

RN HEURISTIC

Equal weighting, 0/1/2 values, total, thresholds, wording, and consequence ordering were designed by RN.

MISSING

No first-time-user study, deployed workflow, outcome data, reviewer-performance data, or longitudinal results.

PROHIBITED

Do not call the score compliance, certification, risk clearance, proof of effective oversight, or permission to deploy.

PRIVATE BY DESIGN

Assessment state is browser-local; an export is not proof the control operated in practice.

Known failure modes

Rubber-stamp reviewA named human can still defer to automation. Only observed operations can reveal this.
Self-report inflationOwners may choose idealized answers. Require documents and sample-case review.
High total hides a critical zeroAdd non-compensable stop gates for consequence-specific requirements.
Generic standardDetailed mode captures a standard, but the quick score does not independently gate it.
Late interventionReview after irreversible consequence cannot function as prevention.
No affected-person pathwayThe tool does not create notice, appeal, remedy, or participation.
05 / PRIMARY EVIDENCE + COUNTEREVIDENCE

What each source supports—and where it stops.

NIST AI RMF 1.0 — Core and Appendix C ↗

Exact locator: GOVERN 2.1 and 3.2; MAP 2.2, 3.4–3.5; MEASURE 1.2–1.3; Appendix C

Used for: Roles, oversight planning, decision information, proficiency, and control evaluation.

Does not establish: NIST does not prescribe this score, weighting, thresholds, or certify this implementation.

Regulation (EU) 2024/1689 ↗

Exact locator: Articles 13 and 14

Used for: Information, limitations, automation bias, interpretation, override, reversal, and interruption in applicable high-risk contexts.

Does not establish: Applicability varies. The prototype is not a legal assessment.

UK ICO — AI audit toolkit: Human review ↗

Exact locator: Human review control measures; ways to meet expectations; options to consider

Used for: Meaningful intervention, competence, time, independence, authority, logs, fallback, and evaluation.

Does not establish: Completing this prototype does not meet ICO expectations by itself.

UK ICO — Individual rights in AI systems ↗

Exact locator: Solely automated decisions; meaningful human input; automation bias and rubber-stamping

Used for: Counterevidence to the assumption that a nominal human makes intervention meaningful.

Does not establish: UK data-protection context; not legal validation of this scoring method.

06 / VERIFICATION

What has been tested, and what has not.

Automated and inspectable checks

  • A late, powerless reviewer cannot receive Strong.
  • All strongest responses produce 16/16 while preserving the non-certification boundary.
  • Threshold fixtures prove 13 is Strong and 12 is Partial.
  • High-consequence fixtures promote evidence, authority, and timing.
  • Empty submission visibly produces 0/16.
  • The core path makes no external request and export is user-initiated.
  • Keyboard, 320px, metadata, reduced-motion, and inactive-scene focus paths are tested.

Human research ledger

COMPLETECode behavior, source provenance, adversarial fixtures, privacy boundary, and automated interaction paths.
NOT CONDUCTEDObserved first-time use with non-specialists.
NOT CONDUCTEDSessions with reviewers in consequential workflows.
NOT CONDUCTEDQuick Check versus Detailed Design comprehension.
NOT CONDUCTEDOutcome study of whether repairs improve review quality.
NO INVENTED RESEARCH

An automated test can prove deterministic behavior. It cannot prove that a first-time visitor understands the language, that a reviewer resists automation bias, or that a repair works for disabled people. Those gates remain open until observed sessions exist.

07 / RECONSTRUCTION GUIDE

Rebuild the capability step by step.

  1. Define eight stable dimensions with three observable states each.
  2. Write states in workflow language, not abstract governance language.
  3. Implement a pure 0/1/2 mapper separate from rendering.
  4. Sum and label values while marking thresholds as heuristics.
  5. Generate repairs from exact weak dimensions.
  6. Add consequence-aware ordering without changing the score.
  7. Build the protocol from declarations; never invent missing details.
  8. Use native labels, selects, and fieldsets.
  9. Keep state local; version the explicit export.
  10. Test missing input, false reassurance, thresholds, ordering, privacy, keyboard, and 320px.
  11. Publish direct support, synthesis, and RN invention separately.
  12. Observe first-time users before claiming broad comprehension.
08 / PRODUCTION + STORY

What must happen before this becomes a launch story.

Production requirements

  1. Recruit at least five first-time users across nontechnical, operational, and governance backgrounds.
  2. Measure whether each can explain the result, reject certification framing, and name the first repair.
  3. Interview at least three real reviewers; test whether a decisive control is missing.
  4. Decide and document non-compensable gates for high-consequence contexts.
  5. Add an affected-person pathway model or mark it outside scope.
  6. Freeze a versioned scoring specification and export it.
  7. Repeat accessibility, privacy, and adversarial tests in production.
  8. Only then select creative direction and write the launch story from findings.

Three future creative territories—not designs

These are briefs to test only after the product and human-research gates above are complete. No carousel direction is selected here.

TERRITORY 1

THE EMPTY CHAIR — the gap between “human present” and actual authority becomes the visual system.

TERRITORY 2

THE CONTROL ROOM — evidence, timing, intervention, escalation, and record become physical switches that must connect.

TERRITORY 3

THE RUBBER STAMP — apparent oversight repeats until the hidden control architecture breaks it.

Open release gates

  • Observed comprehension and task-success evidence.
  • Reviewer-domain validation.
  • Decision on non-compensable stop gates.
  • Affected-person contest and remedy scope.
  • Creative territory comparison and selection.