NAIC's Draft AI Exam Lets Insurers Set Their Own Materiality Bar

Version 5.0 of the NAIC's AI Risk Evaluation Supplement starts state exams with a model inventory scoped by a materiality threshold the examiner sets or asks the insurer to supply and explain. Comment letters are in, and the working group meets October 8.

By Rajesh Beri·October 5, 2026·12 min read
Share:
A state insurance examiner's conference table with a thick printed binder open to a spreadsheet-style model inventory, a blank form line labelled with a ruler beside it, and a stack of comment letters held by a binder cl

Illustration generated using AI

At a US insurer, the next state exam of your AI will probably start with a model inventory, scoped by a materiality threshold the examiner sets or asks you to pick and justify. That is the core of version 5.0 of the NAIC's AI Risk Evaluation Supplement, the questionnaire state examiners can use to review insurers' AI. Adoption would not oblige any state to use it, but it would give examiners a standard set of questions to reach for. An insurer that arrives with a written threshold and an inventory built to it gets to argue the scope of the exam from its own document; one that improvises both under a deadline will be arguing from the examiner's.

The draft is not final. Comments closed on September 29, and NAIC posted 26 comment letters on October 2. The Big Data and AI (H) Working Group meets on October 8 and November 2 to keep working on it. The industry's asks are on the record now, and so is the consumer side's pushback.

What Version 5.0 Actually Asks For

The supplement is four exhibits that regulators can use in sequence, from a headcount of AI use down to the data behind a single model. NAIC staff's summary of changes and the draft text reproduced in the comment package lay out the structure:

  • Exhibit A quantifies your use of AI. Part One is a grid of operation areas (marketing, underwriting, ratemaking, claims, utilization review, fraud, reserves, catastrophe triage, reinsurance and more), with counts of AI systems in use, systems implemented in "the Past X Months" (the regulator fills in X), and counts of AI models with direct consumer impact or material financial impact, split into GLMs, generative/agentic AI and other AI/ML.
  • Exhibit A Part Two is a request list. It asks for an explanation of your materiality calculation, documentation of your risk assessment process, and a Model Inventory listing each model's name, how it is used, its use case, the program area it sits in, its inherent risk level, and its consumer and financial impact.
  • Exhibit B covers your AIS Program, the governance program the NAIC's December 2023 model bulletin told insurers to build. It now carries a new question (3d) on explainability and transparency, a new question (3o) on materiality, and a new question (3p) on how you oversee third-party models.
  • Exhibit C goes model by model through your high-risk models, and Exhibit D asks about the data those models use.

Staff say the Model Inventory request was implied in version 4.0 and is now explicit. Version 5.0 also adds definitions for agentic AI, AI Model, Direct Consumer Impact, Material Financial Impact and materiality. A model inventory in this context is a list of every AI model in the scope the examiner sets, with its use, owner area, risk rating and impact, kept current enough to hand over on request.

The Materiality Threshold Is Yours to Defend

Version 5.0 lets the examiner either set a materiality threshold or ask you to supply one, and if you supply it, you have to explain it. The draft's Exhibit A regulator instructions say regulators "should either specify a materiality threshold to guide company responses to the Supplement or request that the company provide a threshold that it used in completing this Exhibit," and the form carries a blank line for it. The general guidance says regulators "may opt to rely on a company's internal materiality thresholds provided the information supporting a company's materiality, including an explanation for why that threshold was considered appropriate."

The definition is borrowed from the Financial Condition Examiners Handbook. NAIC staff describe the result as "a dynamic where a company may be asked to complete Exhibit A responses based on a materiality threshold the company specifies but also that it discloses to regulators."

That borrowing is the weak point, and the industry said so. AHIP's redline notes that the handbook's materiality concept "is in the context of financials," which leaves a gap for consumer-impact questions that are "not financial in nature." NAMIC went further and asked that any threshold "be determined by the company rather than the insurance department," because materiality "is not a defined standard in existing law." ACLI and APCIA both asked for a presumption that an insurer's own documented materiality framework stands unless the regulator gives a reason to depart from it.

The case for that position is a fair one. A threshold built inside your model risk program is tied to your book, your products and your controls, and a single number imposed from outside would sweep in every low-stakes chatbot. The consumer representatives take the opposite view. Their letter calls it "an inherent conflict of interest" that AI model risk criteria "are set by the insurance company," and argues for a uniform definition of high-risk models.

Both readings point to the same Monday task. If the regulator defers to your threshold, it will be the one you wrote down before the exam, with a rationale attached. If the regulator sets its own, you will want a documented threshold to argue from.


Who Wants What Changed on Thursday

The trade groups agree on most of their asks, and they line up against the consumer representatives on scope. Reading across the letters in the comment package:

  • On sequencing, ACLI asks the supplement to say "expressly" that Exhibits A through D operate "sequentially and as progressively escalating inquiries," with C and D used only when earlier answers identify a specific material risk. APCIA says it generally supports sequential review with Exhibit A as the threshold assessment. The draft itself only says regulators "may wish to first use Exhibit A."
  • On the reporting window, ACLI wants one definition of "currently in use" and says "the preceding twelve months provides a reasonable and administrable timeframe," instead of letting each regulator pick X.
  • NAMIC and ACLI want generalized linear models (GLMs) excluded, and APCIA asks the working group to consider whether conventional GLMs should be distinguished from AI systems. NAMIC's argument is that GLMs produce "static fixed scoring formulas" that have been regulated through rate filings "since the 1990s," and it cites EIOPA's recommendation that GLMs and GAMs be excluded from the EU AI Act's high-risk classification. Version 5.0 went the other way: its new guidance says GLMs "are not without risk of causing unfair discrimination or other adverse consumer outcomes" and gives them their own columns in Exhibit A.
  • ACLI asks that Exhibit A Part Two be reassessed because "a broad model inventory at this stage may expand the scope and burden." AHIP wants "low-risk customer service tools, such as chatbots that summarize or provide already approved information" left outside scope.
  • NAMIC says grouping generative and agentic AI "may result in a quantification that raises unnecessary alarm" and proposes three categories: predictive, generative and agentic.
  • From the other side, the consumer representatives want the human-in-the-loop checklist to ask how long reviewers actually spend, citing reports that review of model outcomes "may be less than 10 seconds."

NAMIC also flags a procedural problem: the substantial changes to Exhibit A in version 5.0 "have not been tested or piloted in this form." The pilot ran in 12 states, from California and Colorado to Virginia and Wisconsin, and ACLI notes that "many pilot participants tested only Exhibit A."

Do not plan around version 5.0 being the text that gets adopted. The working group's timeline, as reported by AI Guardian, runs through versions 6.0 and 7.0, with 7.0 the one likely to be considered for adoption at the Fall National Meeting on November 14 to 17 in Grapevine, Texas. What survives that process is uncertain. The inventory request looks like the safest bet: ACLI wants it narrowed to models with direct consumer or material financial impact, and APCIA wants insurers allowed to submit inventories they already keep, but neither letter asks regulators to stop asking for one.

Third-Party and Generative Models Are Where Inventories Break

Question 3p asks how you oversee third-party models, and that is the question most insurers will find hardest to answer with evidence. Your pricing vendor's credit-based score, your claims-intake classifier, the LLM behind your adjuster assistant and an agent built on a frontier model API are all models you did not build. Under version 5.0, they still go in the inventory if they meet the threshold, and they still sit under an AIS Program you have to describe.

The industry knows this is the soft spot. ACLI's letter says an insurer should only have to provide information "reasonably available through its contractual rights and ordinary vendor-oversight processes," and that the supplement should not imply you must produce "source code, training data, model weights, proprietary validation materials" you have no right to obtain. APCIA asks for guidance on "the minimum information insurers are expected to obtain from third-party vendors."

Both letters argue over how much an insurer must hand over and accept that the question will be asked. In practice, what you can tell an examiner about a vendor's model is limited to what your contract lets you request. A related NAIC group, the Third-Party Data and Models (H) Working Group, is separately drafting a regulatory framework for third-party data and model vendors in P&C pricing and underwriting, and version 5.0 borrows its definition of third parties from the Third Party Registration Framework.

Explainability (question 3d) lands hardest on the same vendor models. A GLM's coefficients can be printed. A generative model's answer to "why did the adjuster assistant recommend this reserve" cannot, and APCIA's letter asks that explainability expectations be "commensurate with both the technology involved and the information reasonably available to the insurer." If the honest answer is "we rely on human review," expect the follow-up the consumer representatives are pushing: how long does that review take, and how often does the reviewer overrule the model? Research on how reviewers trust vague AI explanations and on approval gates that rubber-stamp suggests that answer will not always flatter you.

Claims and intake vendors also change hands. When Adlib bought Paperbox, insurers using its claims-intake models got a new counterparty. Your inventory should record who owns each third-party model today, and your contract should give you notice when that changes.


What to Do Before the November Meeting

This Week:

  1. Read the version 5.0 clean draft with your chief compliance officer and model risk lead, and mark every Exhibit A row your company would put a non-zero number in.
  2. Dial in to the October 8 working group call (11:00 AM ET, per the working group page) or assign someone to take notes on how regulators respond to the GLM, sequencing and 12-month asks.
  3. Ask whether your company sits in one of the 12 pilot states, and if so, pull whatever you already sent examiners during the pilot. The draft says regulators may accept prior submissions "if the prior response is still current and applicable."

This Month:

  1. Write down a materiality threshold and the rationale for it, tied to consumer impact as well as financial impact, and get it approved by whoever owns model risk. Do this before an examiner asks, so the threshold reads as policy and not as a number picked for the exam.
  2. Build the inventory to the seven fields in Exhibit A Part Two: model name, how it is used, use case, program area, inherent risk level, consumer impact, financial impact. If you already have an inventory in an MLOps or governance tool, map its fields to these. Our guides to regulated-industry MLOps and AI governance platforms cover tools such as Credo AI and ServiceNow AI Control Tower that keep one.
  3. Record inherent risk (the risk before controls) for each model, since the draft asks for inherent risk specifically.

Before the Fall National Meeting:

  1. List every third-party model in the inventory and check each contract for what you can actually request: model documentation, validation results, change notices, data sources. Where the right is missing, raise it at renewal using a vendor security review that asks for it by name.
  2. Draft your answer to question 3d on explainability for one generative or agentic use case, end to end, and see whether it holds up without vendor help.
  3. Decide whether your trade association's GLM position matches your own. If your pricing GLMs already go through rate filing review, document that path so you can point to it whatever the final text says.

The Bigger Picture

The supplement is how examiners will check compliance with guidance that already exists. The NAIC adopted its AI model bulletin in December 2023, and 24 states and Washington, D.C. have adopted it since. That bulletin asked for a written AIS Program and oversight of third-party vendors, data and models. What it did not supply was a standard set of questions an examiner could ask to check. Version 5.0 is that question set, and Exhibit A is the first page of it.

Banks went through the same sequence with model risk: guidance first, then examination procedures that turned it into document requests. Insurers that treated the bulletin as a policy-writing exercise will find the supplement asks for evidence, starting with a list. For a banking parallel on vendor models, see our piece on model validation in AML alert disposition after Socure bought Fravity.

The threshold debate will be argued on October 8 and November 2. Build the inventory and write the threshold down now, so that whichever text is adopted, you are filling in a form you already have the answers for.

Continue Reading

Share:

Frequently Asked Questions

What is the NAIC AI Risk Evaluation Supplement?

It is a set of four exhibits state insurance examiners can use to review an insurer's AI during market conduct, financial analysis and financial exams. Exhibit A counts AI use and requests a model inventory, Exhibit B covers the AIS Program, Exhibit C covers high-risk models and Exhibit D covers model data. Version 5.0 was exposed for comment until September 29, 2026.

Who sets the materiality threshold in the NAIC AI exam supplement?

Under version 5.0, the regulator either specifies a materiality threshold for Exhibit A or asks the company to provide the threshold it used. If the company supplies it, the company must disclose it and explain why it was appropriate. The definition borrows from the Financial Condition Examiners Handbook.

What must an insurer's AI model inventory include under version 5.0?

Exhibit A Part Two asks for each model's name, how it is used, the broader use case, the operation or program area, its inherent risk level based on the company's risk assessment, and its consumer impact and financial impact. Regulators specify which program areas are in scope.

Are GLMs covered by the NAIC AI Risk Evaluation Supplement?

In version 5.0, yes. The draft gives GLMs their own columns in Exhibit A and says they are not without risk of unfair discrimination. NAMIC and ACLI asked the working group to exclude them, and APCIA asked it to consider distinguishing them from AI systems. The trade groups cited decades of rate filing review and EIOPA's recommendation to exclude GLMs and GAMs from the EU AI Act's high-risk classification.

When will the NAIC adopt the AI Risk Evaluation Supplement?

The Big Data and AI (H) Working Group meets October 8 and November 2, 2026 to continue work on it. Its timeline runs through versions 6.0 and 7.0, with 7.0 the version likely to be considered for adoption at the NAIC Fall National Meeting, November 14 to 17, 2026.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe