Intended outcome
Compare baseline and candidate evaluation pass rates, show the percentage-point change and test it against a customer-entered review threshold.
The working area above saves numeric inputs in this browser and exposes every formula. It does not decide fairness, safety, compliance or release approval.
What must be ready before release?
- A named owner has defined the task and the organization’s intended use.
- The customer has authority to use every supplied record, source or connection.
- Use the browser workspace only for internal preparation while its release checks remain open.
Inputs
The browser calculator accepts only the aggregate numeric inputs listed below.
- Baseline cases (required). Enter the number of evaluated cases in the approved baseline run.
- Baseline cases passed (required). Enter cases that met the baseline run’s stated acceptance criteria.
- Candidate cases (required). Enter the number of evaluated cases in the candidate run.
- Candidate cases passed (required). Enter cases that met the candidate run’s same acceptance criteria.
- Maximum allowed regression in percentage points (required). Enter the organization’s review threshold. This is not supplied as a legal or technical standard.
Outputs
Every result must identify its supplied facts, assumptions and unresolved items.
- Baseline and candidate pass rates.
- Percentage-point change.
- Comparison against the entered regression threshold.
- Commercial scope if accepted. One bounded check of supplied data, with results and an exportable report.
How will it work?
- 1Define the scope
Choose the exact evaluation run comparator task, responsible owner and intended use.
- 2Add supported inputs
Provide only the information listed for the record workspace. Unsupported material must remain unresolved.
- 3Inspect the result
Review the result, assumptions, source basis and exceptions before relying on any output.
- 4Approve or export
A responsible person decides whether the record is complete enough for the organization’s next step.
Sources and review basis
These official sources frame the working tool. They do not determine which requirements apply to a customer or decide the result.
- NIST AI Resource CenterNational Institute of Standards and Technology. Operational resources for AI evaluation and documentation.Testing, evaluation, verification and validation resources. Reviewed 2026-09-24.
- Artificial Intelligence Risk Management Framework 1.0National Institute of Standards and Technology. A voluntary US baseline for AI risk records, responsibilities, evaluation and lifecycle decisions.Govern, Map, Measure and Manage. Published 2023-01-26. Reviewed 2026-09-24.
- Generative Artificial Intelligence ProfileNational Institute of Standards and Technology. A cross-sector profile for identifying and managing risks specific to generative AI.Generative AI risks and suggested actions. Published 2024-07-26. Reviewed 2026-09-24.
Questions about the tool
What does the evaluation run comparator require?
The browser tool lists every supported input before work begins. Start with baseline cases, baseline cases passed, candidate cases, candidate cases passed, and maximum allowed regression in percentage points.
What does the evaluation run comparator produce?
It produces a complete on-screen calculation with visible formulas and a downloadable JSON calculation record. Paid document exports remain closed during release review.
Will the evaluation run comparator make the final decision?
No. Confirm inputs and rules. Review exceptions and approve consequential actions.
Can I buy the evaluation run comparator export?
No. The browser workspace works locally, but checkout remains closed until the release checks and commercial infrastructure pass.
What will still need review?
Confirm inputs and rules. Review exceptions and approve consequential actions.
This browser workspace does not provide legal advice, certify compliance, make a filing or authenticate a customer decision.
Unsupported cases must stop or remain visibly unresolved.
When will it be available?
The browser tool is available for local preparation during release review. No paid export or checkout is available.
Planned price $9.99 per report. Checkout closed. This planned price is not an active offer and does not create a right to purchase.
Commercial release remains closed until the database, payment and end-to-end release checks pass.
Common searches
These are recognized names and task phrases for this tool. Search uses them locally in this browser.
- Evaluation run comparator
- Evaluation run comparator
- artificial intelligence
- AI governance
- machine learning
Show 4 more search terms
- LLM
- model governance
- NIST AI RMF
- AI evaluation and testing
Ask about this tool
Email the product team without sending customer records, documents or case facts.
Email hello@businesscompliancetools.com