The working A prototype: use its controls to test the build question.
Define what “good AI” means for a specific task and consequence—not one headline score.
RN conceived the question, structured the evidence boundaries, designed the experience, implemented the prototype and documented its limits.
This independent prototype demonstrates an approach; it is not a deployed client system, professional advice or proof of real-world outcomes.
Work through the controls with a situation of your own rather than the sample values. The tool responds to what you put in, so the useful output comes from real input, and everything is processed in your browser as you go. Any sample content you find already loaded is there to show the shape of a filled-in state, and you can clear it and start again at any point.
The build overview sets out the question, the purpose, the role and the evidence status in one place, and the public record behind it states what changed, what another person can reuse and what supports the system. If you want to judge how far this build should be trusted, the record is the page to read, not this one.
Domain AI Evaluation Studio
Define what “good AI” means for a specific task and consequence—not one headline score.
SYNTHETIC COMPARISON COMPLETE
1 of 2 meet every declared metric, conservative interval, and slice gate
Frozen metric contract
Deterministic worst-slice ranking
Boundary: Scores, intervals, evidence and representativeness flags are synthetic and browser-local. Ranking is stable with model-ID tie-breaks; no result establishes statistical power, population coverage, generalization, safety, certification, or approval.