AI Core B · Official guidance Platform: AI

AI Evaluation

AI evaluation is a repeatable way to test an AI system. It uses clear tasks, test data, scoring rules, and safety limits.

See how it works
You might call it AI evalmodel evaluation

See how it works

Original worked exampleAI Evaluation

Another example

A team tests candidate models with 300 saved support cases and answers approved by people. It scores quality, safety, delay, and when a model refuses to answer. An impressive demo alone is not an evaluation.

Main parts

  1. 01versioned dataset
  2. 02scoring rubric
  3. 03release threshold

Use it when

Use this term when clear rules are used to compare models, prompts, search changes, or releases.

Do not use it when

Do not shrink an evaluation to one total score or a few prompts with no saved version.

Name used in code

dataset × rubric × evaluator × threshold

Before you ship

Check the model and prompt versions, source data, versioned evaluation set, and measures. Verify links to evidence, tool permissions, privacy, common failures, decline-to-answer and fallback behavior, human review, monitoring, cost, speed, and rollback.

Request you can copy

Outcome: Use or evaluate AI Evaluation to make AI behavior measurable and tied to evidence. User context: A team tests candidate models with 300 saved support cases and answers approved by people. It scores quality, safety, delay, and when a model refuses to answer. An impressive demo alone is not an evaluation. AI method or concept: AI Evaluation. Why it fits: Use this term when clear rules are used to compare models, prompts, search changes, or releases. Do not use it when: Do not shrink an evaluation to one total score or a few prompts with no saved version. AI requirements: Define the input and source evidence. Set model and tool permissions. Use versioned evaluation data and measures. Define failure, decline-to-answer, privacy, speed, and cost limits. Operational safeguards: Keep a clear trace and hide sensitive log data. Show users a safe fallback. Mark steps that need human review. Define how to roll back the model or prompt. Acceptance criteria: Record baseline and target measures on a versioned evaluation set. Test edge cases and hostile inputs. Verify fallback, monitoring, permissions, and rollback. Evidence and limits (evidence boundary): It is official for that source. Other platforms or teams may use the term in another way. Unknowns to confirm: Target task, model and version, evaluation owner, source data, risk limit, tool permissions, and production fallback.

Check this request

Review the current use of AI Evaluation. Definition: AI evaluation is a repeatable way to test an AI system. It uses clear tasks, test data, scoring rules, and safety limits. Release checks: Check the model and prompt versions, source data, versioned evaluation set, and measures. Verify links to evidence, tool permissions, privacy, common failures, decline-to-answer and fallback behavior, human review, monitoring, cost, speed, and rollback. Before changing code, report the evidence you found, gaps, severity, and the smallest safe fix.

B
How official is this term?

Official guidance

The named platform, framework, project, or expert group has an official guide for this term.

It is official for that source. Other platforms or teams may use the term in another way.

Scope
NIST AI risk and evaluation guidance
Document status
stable
Checked on
2026-07-30

Evidence sources & scope

Authority source + Term reference · NIST · stable Artificial Intelligence Risk Management Framework (AI RMF 1.0) Scope: AI evaluation, measurement, oversight, and risk management Role here: Direct term reference Source covers: canonical name, definition, usage guidance, avoidance guidance

Copy it yourself

The browser could not copy this. Select the request below and copy it yourself.