NewInteractive Agents Live

Guide

How to Run AI in Regulated Industries

A strategic framework for 2026 and beyond: how to put production AI on regulated systems you can govern deterministically, prove on demand, and operate under your own control, before the obligations arrive.

InteractiveAI16 min read

The reprieve that isn't

The deadlines moved. The obligations did not.

If you read only the headlines, 2026 looks like a year of breathing room. The European Union's Digital Omnibus, proposed in November 2025 and provisionally agreed by EU lawmakers in May 2026, is set to push the AI Act's high-risk obligations back: standalone high-risk systems would have until 2 December 2027 to comply, and AI embedded in regulated products until 2 August 2028. During negotiations the Commission had proposed tying the start to standards readiness; lawmakers instead settled on fixed calendar dates, to give the market clarity. Formal adoption is expected before August 2026. The deadline is shifting.

The exposure is not. Every regulated interaction your AI handles between now and 2027 is still defensible today under the rules you already operate under, and the systems that will be provable in 2027 are the ones being architected for proof now. A later deadline is not less exposure.

And the pressure is asymmetric. In some places the supervisor is stepping back precisely where the technology is moving fastest. In April 2026 the US banking agencies, the Federal Reserve, the FDIC, and the OCC, rescinded their long-standing model-risk guidance (SR 11-7) and replaced it with a principles-based framework that places generative and agentic AI outside its scope, with AI-specific guidance flagged to follow. Read that carefully. When a regulator declines to tell you how to govern an agent, the obligation to demonstrate control does not disappear. It lands on you, without a template.

And the standards a compliance team can actually lean on are arriving, not retreating. ISO/IEC 42001, the first AI management system standard, is now certifiable by accredited bodies, though as of 2026 it is not a harmonised standard under the EU AI Act and confers no presumption of conformity, so it supports your readiness rather than discharging your obligations. NIST's AI Risk Management Framework and its Generative AI Profile give you a vocabulary of risks and controls. The agent-specific attack surface is now mapped: the OWASP Top 10 for Agentic Applications landed in December 2025, naming risks like agent goal hijack, tool misuse, and identity and privilege abuse, and NIST opened an AI Agent Standards initiative in February 2026. Sector rules kept moving too: DORA came into force across EU financial services in January 2025, and the US proposed the first major HIPAA Security Rule update since 2013.

So the forward-sensing read for 2026 is not "we have until 2027." It is: the rules are multiplying, the agent-specific ones are still forming, and in the gap the burden of proving control sits with the operator. Provability is a now problem.

The category error, regulated edition

“Is the model accurate enough” is a question about a component, not about a deployment.

Ask most teams in a regulated industry how they are evaluating AI and they are answering "is the model accurate enough" or "is it safe enough." Those are questions about a component. They produce a component answer: a strong benchmark, a good demo, a model card. They do not tell you whether a governed agent did real work on a real system of record, enforced the rule that your license depends on, and left a record you could hand an auditor.

So here is the reframe the method turns on, in its regulated form:

The decision is not which model is most capable. It is whether the operating layer around the model can prove, govern, and be controlled to the standard your regulator holds you to.

This is where new language earns its keep. The bar is not explainability. It is provability.

Explainability is retrospective and probabilistic: after the fact, you reconstruct a likely reason a model produced an output. Provability is structural and verifiable: for the rules you can express as deterministic checks, the rule was enforced at the moment it mattered, by a layer around the model rather than by the model's own judgment, and the trace shows it. Explainability is a story you tell after the fact. Provability is the record that the rule fired, and the means to produce it.

A model you can explain is not a rule you can prove.

The Trust test: five questions a regulated deployment has to pass

Five questions inside the Trust lens decide whether you can deploy at all.

The full framework has twelve dimensions across three lenses. A board can stop at the lenses; an operator drills into the dimensions. In a regulated operation, five questions inside the Trust lens decide whether you can deploy at all. Each one names an ache anyone who has tried to put AI on a regulated system will recognize, the strategic question behind it, and the 2026 standard it answers to.

Can you make the rule fire, every time it must

The ache: the rule that your license depends on lives in a prompt, and you are trusting a probabilistic model to remember it under load. Most of the time it does. "Most of the time" is not a control.

The question: are your high-criticality rules, KYC, anti-money-laundering, responsible gaming, suitability, model risk, enforced by a deterministic layer around the model, at every control point, the input, the retrieved context, each tool call before it executes, and the output before it ships? Of these, the tool-call gate matters most: it is the last checkpoint before the agent takes an irreversible action on a real system. Determinism here means the rule fires regardless of what the model generates, for rules expressible as deterministic checks. Rules that cannot be reduced to a check, a suitability judgment, a responsible-gaming concern, are not forced into a gate; they are routed to human oversight, which is the subject of a later question. This is the property the EU AI Act's Article 12 logging requirement and an ISO 42001 management system are built to evidence.

A guardrail you hope holds is not governance. Governance is the rule that fires whether the model cooperates or not.

Can you prove it, on demand

The ache: an incident happens, the regulator asks what occurred, and the honest answer is a multi-week reconstruction across logs that were never designed to answer the question.

The question: for any interaction, can you produce the full path from the decision back to the policy that produced it, on demand? The EU AI Act's Article 12 requires high-risk systems to keep automatic event logs across their lifetime, with a retention floor of at least six months, for exactly this reason. The test is not raw speed, it is whether the answer comes from a purpose-built record or a reconstruction across logs that were never designed for the question. The regulator's question is simple: what did the AI do, on whose authority, and to which account? A system built to answer it produces the record directly, rather than rebuilding it after the fact.

Can it be attacked, and is the agent itself governed

The ache: you secured your users and your network, but the agent is a new kind of actor, with credentials, tool access, and the ability to be talked into things, and most stacks have not caught up.

The question: is tool access least-privilege and granted at runtime, is the agent governed as its own identity rather than riding a human's, and is the runtime built against the catalogued agent attack surface, agent goal hijack (which subsumes prompt injection), tool misuse, memory and context poisoning, identity and privilege abuse, and rogue or over-empowered agents? This is the ground the OWASP Top 10 for Agentic Applications, published in December 2025, and the NIST agent-standards work now define. An agent you cannot constrain at the tool call is an agent you cannot deploy.

Can you set how much freedom it has

The ache: the choice is presented as all or nothing, a rigid script that cannot handle the long tail, or a free-reasoning agent you cannot trust on the regulated core.

The question: can you set, per policy, where the agent must be strictly deterministic and where it may reason, and can you raise its autonomy progressively, under human oversight, domain by domain, rather than all at once? The EU AI Act's Article 14 requires high-risk systems to be designed for effective human oversight. The right architecture lets you start in a validated, human-in-the-loop mode and ramp to live as the evidence accrues, with the regulated core held deterministic the whole way.

Does your data answer to your law

The ache: the contract says your data sits in an EU region, and you assumed that settled sovereignty. It did not.

The question: do you know whose law can compel access to your data, not just which region it physically sits in? Residency is geography. Sovereignty is jurisdiction. Under the US CLOUD Act of 2018, a US-controlled provider can be compelled to produce data in its possession, custody, or control regardless of where that data is stored, which means an EU data center run by a US provider gives you residency, not sovereignty. The structural answers are deployment you control, cloud, private cloud, or on-premise and air-gapped, regional residency, no provider-side retention of your prompts and outputs for training (distinct from the audit logging you keep on your own side), and encryption keys you hold outside the provider's control rather than in the provider's key vault, so that for data at rest the provider can produce only ciphertext. Keys held inside the provider's own key-management service do not achieve this, and metadata and any plaintext present while the agent is processing remain in scope, which is why genuine sovereignty for an operating AI system points toward deployment you control, up to on-premise and air-gapped.

Residency tells you where your data sleeps. Sovereignty tells you whose law can wake it.

Ownership and Leverage still decide the long game. Who can change the agent's behavior and how fast, whether the spend compounds across use cases, whether you stay free of lock-in: these are the difference between AI that keeps pace with your obligations and AI that ages behind a queue. And the 2026 problem is rarely a single agent. It is a fleet: dozens of agents across several vendors and runtimes, with no common inventory, no consistent policy, and no single answer to "which agents touched this customer, and under whose authority." Governing that fleet, one inventory, agent identity at scale, consistent enforcement across runtimes, and a cost-per-outcome you can hold to a budget, is the Leverage problem that decides whether regulated AI stays auditable and affordable as it spreads. If your compliance team needs an engineering ticket to change a rule, your AI cannot keep pace with your obligations. But in a regulated operation, Trust is the gate you pass first, before the rest is even worth scoring.

InputRetrieved contextTool callthe highest-value gate: the rule firesbefore the action executesOutputpolicy enforced before this step proceedsaudit ledger: every decision traced to the policy that produced itthe regulator's question,answered in seconds: what didthe AI do, on whose authority,to which account
A single regulated interaction passes through a deterministic gate at every control point, the input, the retrieved context, each tool call before it executes, and the output. Each gate writes to an audit ledger beneath. The rule is enforced at every gate, not hoped for at the end.

How this stays honest

The rules of this comparison are stated in the open, and they cut against us too.

Comparison content in this market is mostly worthless, because the author chose the criteria to win. So the rules here are stated in the open, and they cut against easy claims.

Concede freely. Hyperscalers lead on security and compliance certifications, hold the broadest portfolios, and now run generative models on-premise and air-gapped, which is real sovereign deployment. Systems integrators bring deep regulated-industry delivery and industry accelerators, though the reusable assets are the firm's, re-sold to the next client, not owned by you. The foundation labs ship zero-retention agreements, regional residency, and healthcare-grade data terms, though that coverage is typically per endpoint and not enabled by default. No category is dismissed, and several reach the top mark on their home dimension.

The criterion cuts at the operating layer itself, too. Any platform you bring in to run regulated AI is a third party, and under DORA that is a governed concentration risk you remain accountable for. The honest answer is control rather than a promise: your data stays yours and portable, traces and datasets export in open formats and can be routed to your own store, so the record lives under your retention and access controls; the team accountable operates the system and changes its behavior directly; and the deployment is located to meet your data policy. That is a different exposure than a black box you can neither inspect nor relocate, and it is the exposure the readiness check below is built to expose.

Score only what is proven. ISO 42001 is certifiable but not yet harmonised under the EU AI Act, so it is readiness, not a presumption of conformity, and this guide says so. Where a capability is designed but not yet evidenced in production, it is named as direction, never scored as fact. A method that wants to survive a hostile read from a skeptical CISO or a regulator has to visibly refuse to overclaim.

The claim that survives all of that is narrow and strong. Each of these properties, deterministic guardrails, audit logging, agent identity, sovereign deployment, now has credible point suppliers, so the claim is not that any one of them is rare. It is that the hard part is enforcing them together, with deterministic precedence, on one operating layer your team controls, and evidencing it in a live regulated operation. That combination, not any single capability, is what one approach reaches and the others do not.

Sector by sector

The lenses are universal. What “provable” has to mean is not.

The lenses are universal. The binding rule, and what "provable" has to mean, changes by sector.

Financial services.
Statistical and machine-learning models in US banking have long fallen under SR 11-7's broad definition of a model, with its demands for inventory, independent validation, and board-level governance. In April 2026 the agencies rescinded SR 11-7 and replaced it with a principles-based framework that places generative and agentic AI outside its scope. That does not lighten the load, it moves it: the burden of demonstrating control over an agent now sits with you, without a prescribed template. Across the EU, DORA has applied since January 2025, pulling third-party ICT providers, an AI platform among them, into scope for operational resilience and oversight. Provable here means a model and agent inventory, deterministic enforcement of suitability and disclosure rules, and a resilience and audit posture you can show on demand.
iGaming.
Identity verification, anti-money-laundering, and responsible gaming are enforced on every interaction or the license is at risk. This is one of the most demanding live tests of regulated AI, and it is where the proof in this guide comes from. France's ANJ licenses online sports betting, horse racing, and poker, and holds operators to per-interaction KYC, AML, and responsible-gaming obligations, including active detection of at-risk players. The question is not whether the AI is clever, it is whether those policies fire on every triggered interaction with a trace behind each one.
Healthcare and pharma.
Protected health information is governed by HIPAA, and the first major Security Rule update since 2013, proposed in January 2025 and still pending as of mid-2026, would make encryption and multi-factor authentication explicit and require written annual verification of safeguards from business associates, which includes any AI vendor handling PHI. Provable means every access and action on PHI is logged, gated, and attributable, and your AI vendor can pass that verification.
Insurance.
Under the EU AI Act, AI used for risk assessment and pricing in life and health insurance for natural persons is classed high-risk (Annex III), which brings the logging, transparency, and human-oversight obligations directly onto underwriting and claims. Provable means an adverse pricing or coverage decision can be traced, on demand, to the rating factor and the policy that drove it, for the high-risk classes the Act names.
Public sector.
Citizen-facing and essential-service uses carry both high-risk classification and acute sovereignty pressure. The combination makes deterministic governance and sovereign, auditable deployment non-negotiable rather than nice to have.
Telecom.
High-volume customer operations under data-protection and consumer-protection rules. The challenge is less any single exotic regulation than governing correctly at enormous scale, where a small per-interaction failure rate becomes a large compliance problem.

The proof: a regulated operation, in production

A national licence, strict KYC and AML, and Interactive Agents in production.

Betsson France runs one of the more demanding regulated support operations in European iGaming, under a national license with strict KYC, AML, and responsible-gaming requirements. It deployed Interactive Agents into production across customer support, KYC, and compliance, integrated end to end with its player-management, identity-verification, and ticketing systems.

The Trust properties above are not theoretical there. Compliance policy is enforced on every triggered interaction, with the high-criticality rules evaluated in a deterministic engine and precedence resolved before context reaches the model. Every policy match, tool call, and response is logged as a hierarchical trace, so the question "what did the AI do, on whose authority, and to which account" is answerable in seconds rather than reconstructed over weeks. On Betsson France's reported numbers, around 80% of customer issues are resolved by AI end to end, and KYC checks that once took days resolve in minutes.

And it is operated by the business, not waited on. When a rule changes, a new responsible-gaming check, a tightened bonus-eligibility condition, a tone change on a sensitive topic, the accountable team writes the policy, tests it, and deploys it, with no engineering ticket. Engineering owns the platform and integrations underneath; the team that owns the obligation owns the behavior. That is the regulated case for operate, not wait, made in production.

The Regulated AI Readiness Check

Ten questions to run against every vendor, and against building it yourself.

A strategy is only worth something if you can act on it. Run these ten questions against any vendor on your shortlist, and against your own build-it-yourself option, scored as honestly as the rest.

  1. Determinism. Are your high-criticality rules enforced by a layer around the model, or are you trusting the model to remember them?
  2. Control points. Is enforcement applied wherever your risk actually lives across the interaction, including before each consequential action executes, or only at the end?
  3. Provability. For any past interaction, can you produce the path from the decision to the policy that produced it, on demand?
  4. Audit latency. When a regulator asks what the AI did, on whose authority, and to which account, does the answer come from a purpose-built record or a weeks-long reconstruction?
  5. Operating control. When a rule changes, can compliance, risk, or operations change it directly through a controlled process, or does it wait in an engineering queue?
  6. Agent security. Is tool access least-privilege and policy-gated, and is the agent governed as its own identity?
  7. Human oversight. Can you set where the agent must be deterministic and where it may reason, and raise autonomy progressively under human control?
  8. Sovereignty. Do you know whose law can compel access to your data, not just which region it sits in, and can you hold the keys?
  9. Versioning. Can you roll back a behavior change and prove what the agent did under the prior version?
  10. Independence. If you left this vendor, do your policies, data, and logic leave with you in a usable form?

Read the result like this. Mostly no: you are waiting on someone else to carry your burden of proof, and you will feel it the day an incident or an audit arrives. Mostly "yes, inside one suite": you will pass until the question crosses a system boundary the suite does not own. Mostly yes across determinism, provability, and operating control: you are built to run regulated AI, not hope it holds. A well-architected in-house build can score well here too. That is the honest result, and the same questions show you what reaching it would take.

What is actually due now

The AI you can run is the AI you can govern, prove, and operate yourself.

Strip away the framework and the strategy is one sentence. In a regulated industry, the AI you can actually run is the AI you can govern deterministically, prove on demand, and operate under your own control.

The obligations are not getting lighter. They are multiplying, the agent-specific ones are still being written, and in the meantime the burden of proving control sits with you. The deadline moved. The burden of proof did not.

If you have run the readiness check and the answer points to an operating layer your own team controls, the next move is not another year of evaluation. Take the one regulated domain that matters most, scope a pilot against these questions, and treat the readiness check as the acceptance test: a governed agent, on your real systems, enforcing the rules your license depends on, with a trace behind every decision and your team holding the controls. It is the lowest-risk way to test a decision this consequential, with your own data, under your own regulator, and to find out whether the answer holds before you scale it.

Sources

Every regulatory claim above, with its primary source.

Regulatory and standards claims, verified June 2026. Re-verify quarterly; this landscape moves fast.

Bottom line

Capability is not your constraint. Provability is.

In a regulated industry, the AI you can actually run is the AI you can govern deterministically, prove on demand, and operate under your own control. The deadline moved. The burden of proof did not.

Three next steps, in order

Test the decision on your own systems.

Primary

Take the one regulated domain that matters most and scope a pilot against the readiness check, with your own data, under your own regulator.

Secondary

See deterministic enforcement, provable traces, and operating control on a governed agent, walked through live.

Tertiary · Reader-forked

Go deeper