CHAI released the third version of its Risk Categorization Tool in June 2026, developed by its Risk Work Group, replacing the single risk score of earlier drafts with nineteen individually rated modifiers across two domains. The tool now walks a health system through nineteen named risk modifiers: nine under Life & Patient Safety and ten under Technology & Data.
And each one gets its own low, medium or high read rather than being folded into a single composite number. There is no dashboard score at the end of this. What comes out is a map of where a use case needs a closer look and CHAI pushes hard toward a deeper follow-up review the moment even one modifier lands on the high end. For any AI tool sitting inside a live clinical conversation, that modifier-by-modifier structure matters more than a single risk score ever could.
1. Four Risk Phases, Only the First Is Covered Here
CHAI frames AI risk management as four connected phases, and this tool only handles the first one.
1.1 The Four Phases in Order
Risk Categorization classifies a use case as low, medium or high during pre-deployment and sets how much rigor everything downstream needs.
Risk Assessment follows for anything that lands high, a closer look sized to the organization's own appetite for risk.
Risk Mitigation is where controls actually get built and written down.
Risk Monitoring is the ongoing watch over performance and safety once a tool is live.
1.2 Where the Framework Stops Today
The v3 guide is candid about a real gap. CHAI has not yet published a formal Risk Assessment methodology for phase two and tells health systems to lean on their own process in the meantime. That is a maturity signal worth noting, a national coalition shipping a tiering tool before it has finished the harder layer underneath it, rather than sitting on the whole framework until it is complete.
2. Nine Modifiers Score How Close the AI Sits to the Patient
The Life & Patient Safety domain asks a reviewer to work through modifiers like:
- Distance from Patient,
- Decision Autonomy,
- Consequences of Failure,
- Use Context and Complexity,
- React Time,
- Breadth of Potential Harm,
- Cross-System Propagation Risk
- and Population Sensitivity or Disparity Risk.
2.1 Distance from Patient and Decision Autonomy
A single medical interpretation platform can score differently on the same modifier depending on the conversation it is carrying. Confirming a scheduling appointment in Spanish sits close to the low end of Distance from Patient, with a human still deciding what happens next. Walking a patient through informed consent for surgery does not sit there and Decision Autonomy shifts too, from a human clearly deciding to a human supervising a conversation happening in real time.
2.2 React Time and Breadth of Potential Harm
On Care Culture Talks, Dr. Sam Frenkel described how competing for scarce interpreter resources in the ER shapes triage decisions in ways nobody teaches in medical school. That is exactly the kind of pressure React Time and Breadth of Potential Harm are built to surface before go-live, not after a bad outcome, since an ER visit gives a clinician far less time to catch an interpretation error than a scheduled outpatient consult does.
3. Ten Modifiers Score What's Underneath the Model
The Technology & Data domain covers:
- Use of Sensitive Data,
- Accuracy of Data,
- Completeness of Data,
- Veracity of Data,
- Data Transparency,
- Sufficiency and Representativeness of Data Used for AI Model Training and Operation,
- AI Model Security Vulnerabilities,
- AI Model Lifecycle Management and Updates,
- AI Monitoring and Incident Detection,
- plus AI Detection and Traceability. Two of these land squarely on interpretation platforms.
3.1 Sufficiency and Representativeness
This modifier asks whether training data cover the populations a tool will actually encounter, which is the same question as asking whether a platform's 295+ languages and dialects include enough Haitian Creole, Tagalog or regional Spanish variation to avoid gaps nobody notices until a patient does.
3.2 AI Detection and Traceability
This modifier asks whether a patient can tell when AI is influencing what they are hearing. That question sits close to what the 2026 NEJM Catalyst study of Spanish-speaking surgical patients at Brigham and Women's Hospital found: patients wanted AI for speed and privacy and a human interpreter for emotionally complex consults, which only works if the patient actually knows which one they are getting.
4. What CHAI's Own Worked Example Shows
CHAI's v3 guide includes a completed example rather than just a blank template, a scheduling chatbot used in outpatient, primary care settings.
4.1 The Scheduling Chatbot Scores Low
CHAI's own worked example walks an AI-assisted patient scheduling chatbot through the tool. Most modifiers come back low across the board. The reason is structural, not incidental. In this example, a human confirms the appointment once the AI finishes scheduling it. Decision Autonomy stays low because nothing happens without that final human step. Distance from Patient stays low because the tool never touches the clinical conversation itself.
4.2 Swap in Informed Consent and the Score Flips
Replace that same tool with an interpretation platform handling informed consent instead of scheduling and several modifiers move immediately. Distance from Patient moves toward high because the AI is now directly inside patient care, not adjacent to it.
Population Sensitivity or Disparity Risk becomes central rather than an afterthought, since the patients affected are, by definition, the ones with limited English proficiency. That contrast is the entire point of scoring modifiers individually instead of asking one blanket question like "is this AI tool risky."
Bottom Line
CHAI's Risk Categorization Tool (v3) provides a framework that supports a health system's own judgment. Its nineteen modifiers exist to score risk at the level of how a tool actually gets used, not as a single number. The real test is not whether a vendor has read this framework. It is whether they can walk through all nineteen modifiers for their own product without backing off, before a governance committee ever asks them to.