NIST TEVV-Athlon: what the new AI evaluation framework changes for projects
NIST is seeking comments on AI 200-2, a draft TEVV-Athlon framework for designing AI evaluations around organisational goals and risks.
Short answer
The NIST AI 200-2 draft introduces TEVV-Athlon, a way to build an assessment for a specific AI system rather than another universal leaderboard. It starts with an organisational objective, selects test events and tools, produces measurement blocks, and connects evidence to a decision. For an enterprise, the important change is documenting why a measurement exists, in which context it was collected and which decision it can support.
NIST announced the draft on 7 August 2026 and opened comments until 6 October 2026. AI 200-2 is therefore a draft for public input, not a mandatory standard or a compliance certificate. The proposal covers statistical models, large language models, multimodal models and agentic systems (NIST TEVV-Athlon announcement, AI 200-2 draft).
Nexxom diagram: an assessment links an objective, events and tools, measurement blocks, and a decision. It does not represent a test result.
What changes beyond a public leaderboard
A public leaderboard usually answers a narrow question: how did a model perform on a defined task set, with a specified method and period? That information can help filter candidates, but it does not automatically describe your business language, permissions, data, response targets, interface or cost of failure.
TEVV means testing, evaluation, verification and validation. NIST describes this family of practices as a way to produce evidence that a system meets individual or organisational goals while limiting negative impacts. TEVV-Athlon does not impose one score. It provides an architecture for composing an assessment around an objective and context (NIST AI Risk Management Framework).
The NIST Generative AI Profile and the AI RMF Playbook already provide risk-management guidance. TEVV-Athlon sits where a team has to turn that guidance into observable evidence.
The draft's four ideas
The NIST page describes a four-stage method for developing a customised assessment. The full draft should be read for definitions and examples, but its operational principles are already useful.
| Stage | Question to document | Expected output |
|---|---|---|
| Objective | What outcome does the organisation need? | Measurable objective and scope |
| Events and tools | Which situations and instruments create observations? | Scenarios, datasets, scripts or human guides |
| Measurement blocks | Which dimensions are measured and how? | Metrics, criteria, traces and calculation rules |
| Decision | What action follows the result? | Accept, retry, change or stop |
The key is traceability between all four levels. A precision metric can be correct and still be useless if it supports no decision. A qualitative review can be valuable if it captures a severe error that an average score hides.
Events, tools and measurement blocks
NIST describes a TEVV-Athlon as an assessment in which systems are tested through Events and Tools that produce data about Blocks related to measurement concepts. These terms belong to the draft and should not be presented as a certification.
An event could be a user request, a difficult document, a tool failure, a model change or a long interaction. A tool could be a held-out dataset, a script, a human-rating rubric, a simulator or a telemetry instrument. A block collects observations for a dimension such as faithfulness, latency, security, coverage or robustness.
| Object | Example for a document assistant | Risk if poorly defined |
|---|---|---|
| Event | Question against an updated internal policy | Test set is too easy or obsolete |
| Tool | Versioned corpus and rating protocol | Result cannot be reproduced |
| Measurement block | Correct citation and no invented fact | Aggregate score hides a critical error |
| Decision | Publish, request review or block | Measurement has no operational consequence |
Teams should retain the model version, prompt, corpus, evaluation code and system context. A result can only be compared over time if those variables are identified.
An enterprise reading
TEVV-Athlon does not ask organisations to replace every existing test. It helps connect them. A regression test, human review and production alert can become components of one assessment if their objective and method are documented.
For model procurement, the framework encourages teams to write criteria before looking at a public score. For an agent, it encourages measurement of task success, tool-call quality, refusals, escalations and side effects. For a multimodal system, it helps separate perception, reasoning and action errors.

Static Plotly matrix produced by Nexxom. Values from 1 to 3 are an editorial heuristic about context, traceability and adaptability, not a NIST measurement.
What teams can prepare now
An organisation can start before the final draft:
- Name the decision: what result allows deployment, retry or stopping?
- Define context: users, languages, data, channels, tools, volumes and response constraints.
- Separate errors: response quality, source faithfulness, security, fairness, availability, cost and unauthorised action.
- Attach evidence to each error: example, trace, review, automated test or production measure.
- Version the protocol: model, prompt, corpus, configuration, code and date.
- Define thresholds and follow-up: accept, correct, escalate, disable or reassess.
This discipline complements Nexxom’s article on evaluating an LLM beyond public leaderboards and the piece on agent traces in production. It does not automatically turn an internal assessment into regulatory evidence.
Questions worth asking before you adopt the framework
The draft is most useful when an organisation turns its abstract terms into a working brief. Ask which users are represented by the test population, which failures matter enough to stop a release, and which observations are independent from the data used to tune the system. Ask whether the same assessment can be repeated after a model update, a retrieval change or a new tool permission.
Also document the boundary of the system under test. A model-only score does not cover the prompt, retrieval layer, guardrails, tools, user interface and human escalation path. An agent may pass a language test and still fail when an API returns a delayed, empty or ambiguous result. A multimodal system may recognise an object accurately while making the wrong operational decision. The assessment should state which component produced each observation and which component owns the fix.
Finally, define how evidence will be shared. Procurement teams need comparable criteria, operators need actionable alerts, and governance teams need a record of assumptions and limitations. One report can serve all three audiences if it keeps the objective, method, data, result and decision separate. This is the practical value of a flexible framework: it makes an evaluation reusable without pretending that every use case has the same risk.
TEVV-Athlon and AITE are different initiatives
NIST also presents AITE, an evaluation programme with a sequestered test environment and blind data. The AITE page says its August 2026 evaluation period begins with image-analysis tasks in quantum science, genomics and public safety (NIST AITE).
AITE is a test programme and environment for announced tasks. TEVV-Athlon is a framework for designing assessments around different systems and objectives. An enterprise can learn from both, but should not attribute AI 200-2 concepts to AITE or describe TEVV-Athlon as an already certified testbed.
| Initiative | Nature | Enterprise use |
|---|---|---|
| TEVV-Athlon | Draft methodological framework | Design a contextual and traceable assessment |
| AITE | NIST programme and testbed | Observe an announced evaluation programme |
| Public leaderboard | Comparative result from a defined protocol | Filter options before an internal test |
| Production test | Measurement in the live system | Monitor drift, incidents and business effects |
Limits and cautions
The draft is still open for comments. Its terms, scope and examples may change. NIST specifically invites feedback on the definitions of testing, evaluation, verification and validation, the flexibility of the framework, gaps in covered activities and usefulness for emerging systems (NIST consultation questions).
Do not claim that an assessment “aligned with TEVV-Athlon” proves compliance. Do not publish a score without its corpus, versions, method and limitations. Do not replace human review or security validation with one aggregate number.
Teams can compare their protocol with the NIST AI standards programme and document gaps. The NIST framework is voluntary; applicable obligations depend on sector, country, contract and use case.
Conclusion
TEVV-Athlon offers a practical idea: a useful assessment starts with a decision, not a leaderboard. Enterprises can already map objectives, events, tools, measurement blocks and operational follow-up, then retain the evidence in a versioned register. AI 200-2 deserves monitoring and, for relevant organisations, a careful reading before the comment period closes on 6 October 2026.
For a deeper comparison, connect the method to your LLM reasoning tests and your structured-output controls. The goal is not to add a label to testing, but to make every AI decision defensible through understandable evidence.
Primary sources verified on 24 August 2026
- NIST, TEVV-Athlon announcement and AI 200-2 consultation
- NIST AI 200-2 draft
- NIST, AI Risk Management Framework
- NIST, Generative AI Profile
- NIST, AI RMF Playbook
- NIST, AI Technology Evaluation
- NIST, AI Standards
- NIST, ITL AI Program
- NIST, AI Research

