Skip to main content
WHAT IT IS

The working A prototype: use its controls to test the build question.

WHY IT MATTERS

Define what “good AI” means for a specific task and consequence—not one headline score.

RN'S ROLE

RN conceived the question, structured the evidence boundaries, designed the experience, implemented the prototype and documented its limits.

BOUNDARY

This independent prototype demonstrates an approach; it is not a deployed client system, professional advice or proof of real-world outcomes.

HOW TO USE THIS PAGE

Work through the controls with a situation of your own rather than the sample values. The tool responds to what you put in, so the useful output comes from real input, and everything is processed in your browser as you go. Any sample content you find already loaded is there to show the shape of a filled-in state, and you can clear it and start again at any point.

WHERE TO GO NEXT

The build overview sets out the question, the purpose, the role and the evidence status in one place, and the public record behind it states what changed, what another person can reuse and what supports the system. If you want to judge how far this build should be trusted, the record is the page to read, not this one.

BUILD 038-A

Domain AI Evaluation Studio

Define what “good AI” means for a specific task and consequence—not one headline score.

ALL REQUIRED SLICES / contract 038.2.0

SYNTHETIC COMPARISON COMPLETE

1 of 2 meet every declared metric, conservative interval, and slice gate

Frozen metric contract

accuracy50%higher is better · minimum 80 percent · synthetic classification task
abstention25%higher is better · minimum 75 percent · synthetic review queue
reviewability25%higher is better · minimum 80 percent · synthetic reviewer rubric

Deterministic worst-slice ranking

1MODEL-BMEETS DECLARED GATESweighted 86.75 · worst rare-x-language (84.5)
2MODEL-ADOES NOT MEET GATESweighted 80.58 · worst rare-x-language (76.5)

Boundary: Scores, intervals, evidence and representativeness flags are synthetic and browser-local. Ranking is stable with model-ID tie-breaks; no result establishes statistical power, population coverage, generalization, safety, certification, or approval.