NewInteractive Agents Live

Insight

How to build an effective AI strategy

An AI strategy is a decision about how your company will operate AI, not which AI to buy. Six decisions, one discipline, and the questions each person in your boardroom will ask you.

InteractiveAI35 min read

For the CEO, the COO, the CFO, the CTO/CIO, the CISO and DPO, and the function heads who will run it.

01

The use case

Where does AI create value first?

COO + the function head

02

The context

What must the AI know, where does that knowledge live today, and how does it become operating context?

The function head + CTO

03

The team

Who builds it, and who owns it after go-live?

COO + CHRO

04

Governance

How do we bound what it can do, see everything it did, and catch what it got wrong?

COO + Legal

05

Security and compliance

What will our CISO, DPO and regulator demand?

CISO + DPO

06

Infrastructure

Which operating layer do we run on, and who is accountable in production?

CTO/CIO + CFO

The one discipline: every use case, at every stage, answers three numbers. What it gains you. What it costs. How fast you see it. A proposal that cannot answer all three is not a strategy. It is a wish.

The one test to carry into every vendor room

Every serious contender will now tell you that your business experts can edit the AI without engineering, and most can demonstrate something like it. So stop asking who can edit, and ask what survives. Almost nobody passes all five.

  • Authored, not scraped Rules written deliberately by accountable people, not inferred from a pile of documents.
  • Editable by the people who own the operation Not by ticket. Directly.
  • Versioned and testable Like code: see what changed, test before live, roll back.
  • Loaded per situation Not dumped into one giant prompt.
  • Portable If you terminate the vendor tomorrow, your rules, tests and decision records leave with you in a form someone else can run.

Decision 2 explains each one. Take the list as it stands, and use it on us too.

Four questions before you read on

Answer these about your own company. Most leadership teams fail two.

  1. In your top use case, can you say what is already automated today, and what is not?
  2. Six months after go-live, can a named operations person change a business rule without an engineering ticket?
  3. Can you take one AI decision from last week and reconstruct why it happened?
  4. For your number-one use case, do value gained, cost to run and time to value exist on paper?

A no is not a failure. It is the chapter you should read first.

Who wrote this, and what we sell

Before the method, the disclosure, because "neutral" guides that hide their author insult your intelligence. This guide is published by InteractiveAI. We build the platform where domain experts and developers build and run AI agents together, governed the way the six decisions describe, and we run our own operation on it. The Betsson results in decision 3 run on our platform.

So we are not neutral: we built the company on the conviction that the operate-don't-wait model wins, and this guide is that conviction written down. Everything in it is still method, and the method does not care who you hire. Use the five-question test above on us too.

The frame

Most AI strategies fail before they start, and almost never for a technical reason.

Start with a question you can answer about your own company in ten seconds. Six months after your last AI pilot went live, one of the business rules it runs on changed. Who changed the AI, and how long did it take?

If that answer is uncomfortable, you already know why most of this fails. Two statistics say the same thing. MIT's 2025 study of enterprise AI found 95 percent of generative AI pilots produced no measurable impact on the P&L. Research by IDC for Lenovo, counting differently, found that for every 33 proofs of concept launched, four reached production. Both have been argued with, fairly: small samples, short windows, contested definitions of failure. Halve either number and the conclusion holds.

This is also not a wave you can sit out. Gartner expects 40 percent of enterprise applications to have task-specific AI agents embedded by the end of 2026, up from under 5 percent. So the question stops being whether your company runs AI agents. It becomes whether you operate them or merely receive them.

The same analysts expect over 40 percent of agentic AI projects to be cancelled by the end of 2027, and the reasons they give are worth reading twice: escalating costs, unclear business value, inadequate risk controls. Not model quality. Every one of those is an operating failure. McKinsey's global AI survey points the same way from the opposite direction: the strongest differentiator it finds between companies getting real value from AI and companies stuck in pilots is whether they redesigned the workflow end to end, and only about one company in five has.

The pattern behind the failures is a shopping decision. Someone picked a model, a tool or a vendor, ran an impressive demo, called it a strategy. Then it met real customer data, real edge cases, real compliance questions, and a real operations team that was never in the room. The pilot that wowed the executive committee in week two was quietly shelved in month six. Nobody could say why it answered the way it did. The one person who knew the rule it kept getting wrong could not change it. And what it did on day 90 was exactly what it did on day one, because nothing it encountered ever made it better.

An AI strategy is a decision about how your company will operate AI. Not which AI to buy.

One thing raises the stakes. The last generation of AI answered questions. This generation does work: it reads your systems, decides, and acts. It updates the account, issues the refund, files the report. That carries real upside and real risk in a way an answer machine never did, and the playbook you used for analytics tools does not cover it.

Which layer do you refuse to rent?

There is AI your team operates and adapts: your people write the rules, see every decision, improve it weekly. And there is AI you wait on: a vendor's queue, an engineering backlog, a black box that ages while your business moves.

One clarification, because a sharp reader will push back here. You will wait on some things, and you should. Waiting on your model provider is fine: models are interchangeable and they improve without you. Waiting on a cloud provider is fine. What is not fine is waiting on someone else to change your own business rules, because your rules are the part that is only yours. So the line is not build versus buy. It is: which layer do you refuse to rent?

The discipline

Economics is not a chapter at the end. Three numbers run through every decision, and each has a trap in it.

Take the six decisions in order. Infrastructure comes last on purpose: it has to satisfy everything the first five produce, and choosing it first is how companies end up with a specification written by a vendor.

  • What it gains you Split it in two, always. Cashable: cost you stop paying, revenue you can trace, a loss you can show did not happen. Capacity: hours returned. Capacity becomes cash only when someone commits to absorbing it, fewer contractors, no next hire, more volume on the same team. Get that commitment named alongside the projection, or report it as capacity and never as savings.
  • What it costs Not the license, and not one number. One-off: integration build, context authoring in internal expert days, security and impact assessments, training. Run-rate, per year: usage, the rule-owner time your improvement cadence consumes, maintenance, audit evidence. Count internal days at a real loaded rate, because they are usually the largest line and appear on no vendor's quote. And state which lines scale with volume, because in usage-based pricing your cost grows with your success.
  • How fast you see it Two dates. First measured result, in weeks. Then the payback month, when cumulative value passes cumulative all-in cost including the one-offs. A case with no payback month is a case nobody finished.

You will use these to pick the first use case, judge the operating layer, and run the operation afterwards. If a proposal cannot answer all three, send it back.

The room

Most AI strategies do not fail on technology. They fail on an empty chair.

The compliance officer who was never consulted arrives in week ten with a veto. The CFO discovers the usage pricing after the pilot succeeds. The function head whose team the AI runs alongside hears about it from a slide. Every one of those is preventable in the first meeting.

CEO

The ambition and the operate-vs-wait choice

Without them: AI stays a side project, funded but unowned

COO

The use case, the ownership model, what stays human

Without them: An AI nobody in operations wants or trusts

The function head

The rules the AI runs on, and their team's new role

Without them: Quiet resistance; the people who know the work disown it

CFO

The three numbers, and the cost model at scale

Without them: Sticker shock at rollout; success killed by its own invoice

CTO / CIO

The operating layer, integrations, model strategy

Without them: A stack nobody can extend, or a build that never ships

CISO / DPO

Access and identity, lawful basis, the impact assessment, retention

Without them: The week-ten veto

Legal / Compliance

What must stay deterministic, what evidence exists on demand

Without them: A regulator question with no answer

CHRO

The role changes and the consultation plan

Without them: The floor hears it as a rumour

Procurement

The supplier path and the real timeline

Without them: A nine-month onboarding you discover in week six

Works council or employee representation

Informed and consulted at the planning stage

Without them: A deployment blocked after it is built

Take this table into your next AI meeting and check the chairs. It is the cheapest risk reduction available to you.

Decision 1 · The use case

Pick the use case you can prove, fast, in production, and be willing to say no to the rest in writing.

Every company has an AI use-case list: forty ideas from workshops, ranked by enthusiasm. The list is not the problem. Picking without a method is, because then you pick the flashiest idea instead of the one that pays.

Start from the outcome, not the capability. Never "what can AI do?" Always "what outcome are we buying?" Fewer escalations. Faster onboarding. A compliance process that stops eating senior time. If you cannot name the outcome, the use case does not exist yet.

Then check what is already automated

This is the fork most strategies skip. Almost every operation has an easy layer: repetitive, well-documented work that basic automation or a decent knowledge base can already handle. If that layer is still manual, that is your quick win: fast, cheap to prove, and it builds the muscle for everything after. If it is already automated, do not pick "the use case" at department level. Go one level deeper, because inside any function there are five to fifteen distinct workflows. Pick the workflow, not the department.

Take customer support. In many support operations, 20 to 40 percent of contacts are questions a knowledge base can already answer. If yours still handles those by hand, start there. If that layer is done, "automate support" is no longer a use case. The real candidates sit one level down: identity document checks, promotion disputes, refund handling, login recovery. Each has its own volume, rules and risk profile, and treating them as one blob is how pilots get scoped to fail.

Score every candidate on four axes

  • Impact Volume times value times the share the AI will actually handle end to end. That third factor is where optimistic cases die, and 100 percent is never the answer.
  • Speed to deployment Weeks, not quarters. Fewer systems touched, faster proof. Check it against your own calendar, because every operation has weeks when nothing ships: peak season, year-end, a migration, an audit.
  • Provability A crisp baseline beats a fuzzy one every time, even at equal impact.
  • Adjacency Winning here should make the next two candidates cheaper. A win that leads nowhere is a demo with a budget.

Then force the three numbers onto the top three. Survivors are your real shortlist, normally two or three items.

Two filters before you commit. If the rule is exact and the input structured, write code, not a prompt. Eligibility maths, threshold checks, anything that must be right to the cent belongs in deterministic logic the AI calls; use the model where language and judgment are the actual work. And check what you already own, because your incumbent platforms are embedding agents. If one will cover a workflow inside your payback horizon, that workflow is not your use case. What is left, the work that runs on your rules, is.

Decision 1 produces a ranked shortlist, not a lock: candidates get eliminated in decisions 2 through 5 on missing context, missing owners, or a data constraint your DPO names. That is the method working, and it is why you scored three.

Start small, and write down the noes. One workflow, one segment or brand, in production, with real users. A narrow win in production is worth ten broad decks. And a use case that is not worth doing gets named as such, in writing, with a reason. That is what stops the forty-idea list resurfacing every quarter.

Ask your CFO and COO

  • Which three candidates did we score, and why did the winner win?
  • What are its three numbers, and who stands behind them?
  • What is already automated today, and are we sure we are not rebuilding it?
  • What did we say no to, and where is that written down?

Decision 2 · The context

The AI has read the internet. It has not read you, and your particulars are the whole game.

This generation of AI is astonishingly capable and completely ignorant of your business. It does not know what "VIP" means in your book, when your refund policy bends, or which regulator watches which product line.

And your particulars are mostly unwritten. Take your best operator, the one you would trust with your hardest customer at the worst moment, and ask where what they know is actually written down. Some sits in a wiki nobody updated since the last reorg, some in training decks that contradict it. The most valuable part is written nowhere: it lives in their head, and it walks out of the building every evening.

What context actually means

Not a data lake, not "we uploaded our documents." Walk through what your best operator carries:

  • Your vocabulary A "VIP", a "chargeback", a "priority case" mean something specific here. The AI must use your definitions, not the internet's.
  • Your rules If this, then that, and who owns the rule.
  • Your procedures The multi-step processes, exceptions included.
  • Your approved language The phrasings legal signed off, inserted verbatim rather than paraphrased.
  • Your live situation Who this customer is, what their history says, which flags are raised now.
  • Your systems and documents What the AI may look up, what it may act on, and what it may cite. Three different permissions.

From knowledge to operating context

Knowing what the AI must know is half the decision. The other half is translation. A rule written as a paragraph in a wiki is knowledge. A rule written as a condition and an action, with a named owner and a test, is operating context. The first informs. The second operates.

  • Extract Out of heads and scattered documents, workflow by workflow, starting with the use case from decision 1. Not the whole company. One workflow, completely. Check two things while you are in there: the state of the data the AI will read, because it inherits every gap and duplicate in your systems of record, and whether you are allowed to reuse what you are mining. Historic tickets were collected to serve customers, not to train an AI, and in Europe that makes reuse a purpose question your DPO answers before you start.
  • Author Write it as the pieces above, each owned by a named person. This is where knowledge stops being tribal and becomes an asset.
  • Operationalize Make it live: versioned, testable, reaching the AI at the right moment. A rulebook the AI cannot run on is documentation, and you are not building documentation.

The exercise pays before the AI does: half the time, writing the rules down exposes that two teams have been running two different rulebooks. Done right it is the most durable thing your AI programme produces. Models come and go. Your authored operation does not.

Then the part most companies get wrong: who ends up holding it? If extracting your context means handing it to a vendor or burying it in an engineering codebase, you have not kept your knowledge. You copied it and gave the working copy to someone else, and from that day every correction goes through someone else's queue. Who owns it, and what your team's job becomes, is the next decision.

One note on the fourth criterion from the test at the top of this page, because it is the one vendors talk around. An AI given a thousand instructions at once reliably follows only some, and the ones it drops are not the ones you would choose. This is measured, not theoretical: adherence to large instruction sets degrades in every frontier model tested, and a bigger context window does not repeal it. Ask any vendor to show you exactly how the right rules reach it at the right moment.

Ask your CTO and function head

  • Where does our operational knowledge live today, honestly?
  • If our three most experienced people left next month, what fraction walks out with them?
  • Who would author the rules, and who checks them?
  • When a rule is wrong in production, what is the path from we noticed to it is fixed, and how many people does it pass through?

Decision 3 · The team

Six months after go-live, the refund threshold moves. Who changes the AI, and how long does it take?

If the answer is "our ops lead edits the rule, tests it, ships it, same day," you own an operation. If it is "we file a ticket," you own a dependency, and a dependency cannot keep pace with a business that changes weekly.

That is not a criticism of your engineers. The queue exists because they are rationing scarce capacity across the whole company, and the point of this model is to take recurring rule changes off their plate so they can do the integrations only they can do. Bring them that framing, and name the integration capacity you need, in weeks, in writing, in week one. It is usually the scarcest input in the plan.

The three default answers, and what each costs

  • Outsource it Fast start, real expertise. But the knowledge and the edit button live outside your company, and three years in you have paid for an asset that compounds on their balance sheet.
  • Build it in engineering You own the code, but the people who know the rules cannot touch it, and the backlog becomes the speed limit of your AI.
  • Give everyone AI tools Your experts are hands-on, which is right, but there is no shared system of record: forty prompts in forty drawers, no audit trail, no answer to "why did it say that?"

Two keep control and lock out the experts. One keeps the experts and gives up control. The premise underneath all three, that you must choose, is false.

The model that works: the people who know the business and the people who know the systems build and run the AI together, in one governed environment the company owns. Your domain experts author the behaviour, the policies, procedures and vocabulary from decision 2, in language they own. Your developers wire the systems: integrations, data, guardrails. Both work on the same artifacts with the same audit trail. When the operation changes, the ops person changes the rule. When a system changes, the developer changes the wiring. Neither waits on the other.

It has a price, and you should know it first. It costs your best operators' time, permanently, as part of the job rather than a project, and it only works where someone genuinely owns the rules. If you cannot name that person and free up their week, pick one of the other three honestly instead of doing this one badly.

This is not theory. Betsson, one of Europe's largest gaming operators, runs its French brands' customer operation this way: around 75 percent of eligible customer issues resolved end to end by AI, with payment-sensitive actions deliberately excluded and routed to a person, and more than 30 percent of would-be tickets eliminated at root cause, measured as contact volume on the same reasons before and after the fix. Read the number precisely, because a qualified number is the only kind worth quoting. But the numbers are downstream of this sentence:

I write the policy, test it, deploy it. No engineering ticket, no waiting cycle.

Mehdi Aounallah, Betsson France

The person who knows the rule owns the rule.

Be straight about the people, because they will be straighter than you

A handful of your best operators get a genuinely better job: they stop handling every case and start deciding how every case is handled. It is a real promotion and it is not a big number. A 200-seat floor produces perhaps eight to fifteen rule-owners, not two hundred. Most of the floor sees something else: the easy work goes, the hard work stays, and the headcount growth they expected does not arrive. In most companies that lands as slower hiring and unbackfilled attrition rather than redundancies, but say which one it is in your company, and say it early.

Two points people skip: rule-authoring is a distinct skill, closer to analysis than to service, so select for it rather than assuming it; and a rule-owner is a quarter to a half of a person off the floor for good, so put that in your cost to run before your CFO finds it.

In Europe this is also a duty. The AI Act requires measures supporting AI literacy among the people who operate these systems, enforceable nationally from 3 August 2026, and your rule-owner training is that measure: document it and date it. And in Germany, the Netherlands and France, a system that changes how work is organised or measured triggers formal employee consultation, with co-determination or consent rights in the first two. Start that in week one. It is on the critical path, not a formality.

Ask your COO and CHRO

  • After go-live, who exactly can change the AI's behaviour, and what is their job title?
  • Run the refund-threshold test: rule change to live fix, how long, how many hands?
  • Who approves a rule change, and does that depend on the rule's risk tier?
  • What happens to the team this touches, and have we started the consultation the law requires?
  • If we part ways with every external partner tomorrow, what do we still hold, and in what format?

Decision 4 · Governance

An AI that answers questions can embarrass you. An AI that acts on your systems can hurt you.

Once it is issuing refunds and updating accounts, "it's usually right" stops being a governance model. Start from the belief that anchors this chapter: if you cannot trace an AI's decision back to the rule that produced it, you cannot govern it. And if you cannot govern it, you cannot put it in front of your customers.

Enforcement, not a PDF

A policy document next to the system governs nothing. Ask where rules are enforced. The answer must be: at every point where something enters or leaves the AI. What comes in. What it retrieves. Every action on another system, checked before it executes rather than reviewed after. What goes out. "The AI usually remembers the compliance rule" is not a sentence you want to say to a regulator.

Two kinds of rule live here, and the difference is the whole chapter. A rule that guards an action, no refund over 500 euros without a human, is enforced in code at the moment of execution: the system cannot skip it, whatever the model produces. A rule that guides judgment, how to weigh a goodwill case, is an instruction to a probabilistic system, and you make that one reliable by testing it rather than asserting it. Ask any vendor which of your rules is which. Anyone who says all of them are guaranteed is selling you something.

Traceability

For any decision the AI made last quarter you should be able to reconstruct it: what it was given, which rules were in force, which enforcement points fired, what it did on which system, what a human approved. That is a record. Attributing a model's exact wording to one rule is not available from anyone, and a vendor who claims otherwise is telling you about their honesty rather than their architecture.

Every decision is traceable, testable, explainable, and auditable.

Mehdi Aounallah, Betsson France

That is an operations manager in a regulated industry describing his own AI, and it is the standard to hold anyone to, including us, for the rules you can express as deterministic checks. The rest is routed to a human, by design.

Two more things while you are here. Set the retention period deliberately, because the AI Act sets a floor where it reaches you while data protection law says no longer than necessary, and your sector rules usually decide the real number. And require that traces export into your own logging stack: an audit trail you can only read inside a vendor's console is not one you own.

Autonomy is a dial, set per task

Not one global setting. The same AI should handle a password reset end to end, draft-but-not-send a goodwill credit, and hand a withdrawal straight to a human. Some things stay human because you decided they stay human, and no accuracy statistic overrides that.

Decide those carve-outs now, in writing. And know that the list is not only an operating preference: under European data protection law, a decision taken solely by machine that significantly affects someone, a refused refund, a blocked withdrawal, a closed account, needs a lawful basis and a real person the customer can escalate to. Write it with your DPO, not after them. Then design the human's side properly, because the law asks whether the person overseeing has the competence, authority and time to reach a different answer. A reviewer with forty seconds a case and no mandate to overrule is not oversight, and a regulator will say so.

Then design the failure, because you will have one

"It goes wrong the way you designed it to" is a comforting sentence and a false one. It goes wrong at 22:40 on a Saturday, to a customer who screenshots it. What you control is the ceiling on the damage and the speed of recovery, so decide five things before go-live, in writing: what triggers an alarm and who receives it out of hours; a kill switch that has been tested rather than designed; blast-radius caps, actions per hour and value per action, so a bad rule is expensive in the hundreds and not the millions; rollback to the previous behaviour version, timed in a drill, plus what the AI does when a system is down or it is simply not confident, including what the human it hands to receives; and who tells the customer, who tells the regulator. Those clocks are real: 72 hours to the supervisory authority for a personal data breach, 15 days for serious AI incidents and 2 for the worst. Your vendor contract has to get you the news in time to make your own clock.

Governance built this way does not slow you down. Teams with turn-by-turn control ship faster, because they can put AI on real operations without betting the brand on model behaviour.

Ask your COO and Legal counsel

  • Show me one real decision from last week, reconstructed: what it was given, the rules in force, the controls that fired.
  • Which rules are enforced in code, and which are instructions we trust the model to follow?
  • What is our current human error rate on this workflow, honestly measured, and what rate would we accept from an AI that is traceable when our people are not?
  • What are our always-human actions, and has our DPO signed that list?
  • When it goes wrong badly, who outside this company do we tell, in how many hours, and does our vendor contract get us the news in time?

Decision 5 · Security and compliance

Meet these questions in month one and you have a design conversation. Meet them in month ten and you have a veto.

The decision, in one line: collect the security and compliance requirements before choosing anything, and treat them as selection criteria rather than obstacles.

Your CISO's questions

An AI that acts is a new actor inside your perimeter, and it deserves the rigour you would apply to a new starter with system access and no notice period. What can it access, and is that access scoped per task or blanket? What identity does it act under: its own auditable identity, or a shared service account that could hide anything? Can we revoke it in minutes, without a change ticket? And what happens when the manipulation arrives inside the work, hidden in a ticket body, an attached document, a supplier email the AI reads? That is the live attack class, precisely because the AI reads your systems, it is not solved by better prompts, and the only durable answer is limiting what a hijacked turn is authorised to do. Ask for the red-team results and the blast radius, not a reassurance.

Your DPO's questions

On what lawful basis are we processing this, and is it the purpose we collected it for? Is the impact assessment done, and dated before go-live? Where does the data live, where is it processed, and whose law reaches it? What is our retention period for decision records? Who are the processors and sub-processors, including whichever model provider sits underneath? And when a customer asks why the AI did that, what do we hand them, and who is the human they escalate to?

Residency is where your data sleeps. Sovereignty is whose law can wake it.

Both matter and they are different controls. Residency is answered by choosing a region. Sovereignty is answered by who controls the provider, who holds the keys, and where the system runs, because a provider under US jurisdiction can be compelled to produce data it holds wherever that data sits. Get both answers in writing before you sign.

What is already due, whatever your roadmap says

The market spent this summer talking about the EU AI Act being delayed. Part of it was. From 2 August 2026 you must tell people when they are dealing with an AI rather than a person, and that duty applies to ordinary customer-facing AI including yours. It is not the high-risk regime, which the Digital Omnibus, now Regulation (EU) 2026/1744, moved to 2 December 2027 for standalone systems and 2 August 2028 for AI embedded in regulated products. The AI literacy duty is nationally enforceable from 3 August 2026. And data protection law has applied in full since 2018, which is where most of your real obligations already live.

So the deferral changed your deadline, not your exposure. Build for four demands, not one. One customer asking why, entitled to a human. One assessment finished before go-live: where processing is likely to be high risk the impact assessment is a legal precondition, and in several member states the employee consultation is too. One supervisor asking for the record, which decision 4 gives you as a byproduct. And one clock, in hours, for when it fails. The audit everyone talks about arrives last. The other three are what stop projects.

Ask your CISO and DPO

  • What systems can the AI touch, and under what identity?
  • Where is our data stored and processed, and whose law reaches it?
  • Does anything we send leave our control or train a third party's model?
  • If a regulator asks why did it do that, what do we hand them, and how fast?

Decision 6 · The infrastructure

Everything you decided so far is a specification. Now you choose what has to meet it.

The capstone decision, last on purpose. And here is where most companies make the category error: they frame it as "which model?" or "which tool?" The models are converging. Stanford's AI Index has the leading models sitting within a few points of each other on most standard evaluations, close enough that for general business work model choice is not your edge, and whichever you pick will be leapfrogged within a year. Real gaps do persist on the hardest reasoning and on long-running agentic work, which is exactly why you test on your own cases rather than on a leaderboard. The tools are the demo layer. The real question is one level down.

Which operating layer will own your AI outcomes across the whole lifecycle, and who is accountable when it runs in production?

That layer is where everything in this guide lives: where context is authored and versioned, where your experts and developers work, where rules are enforced turn by turn, where the traces live, where integrations act on your systems. Whoever runs it runs your AI.

What to demand of it

  • Model independence Swap models without rebuilding: rules, context, integrations and tests stay, the engine changes. Expect to re-tune and re-validate against your own test cases, a week of work rather than a rewrite. You want this for price, capacity, deprecation and residency, not because one model is smarter. If switching means starting over, you bought a lock-in, not a layer.
  • Integrations that move numbers "We connect to your CRM" is a brochure sentence. The test: can it act end to end, look up the account, apply the rule, execute the change, in production, under the governance from decision 4?
  • A place your experts can work If only engineers can touch the layer, you rebuilt the ticket queue with extra steps.
  • Governance built in, not bolted on Retrofit governance fails exactly when you need it.
  • A clear split of regulatory roles, in the contract If your people author and brand the behaviour, you may become the provider of that AI system under the EU AI Act rather than merely its deployer, with heavier obligations. Often the right trade for the control, but allocate it before you sign rather than discovering it in an audit. Same for the processor terms, the sub-processor list including whichever model provider sits underneath, incident notification fast enough for your own clocks, and an exit that returns your rules and records in usable form.

Check what you already bought

Your incumbent platforms are shipping agents and the account team will offer theirs at no marginal cost, already security-reviewed. Score it in the same grid, honestly, on the thing that matters: your operation is not shaped like one product. A single refund case touches ticketing, the payment provider, the CRM and a risk flag. Four suite agents means four rulebooks, four audit trails and nobody who owns the outcome. Ask each vendor to run one real end-to-end case crossing two of your other systems, under one trace, and watch what happens.

Then build, buy, or partner, against the team you actually have. Building your own is real if you have a platform team ready to own a product for years. Buying a closed stack gets speed and costs the expert ownership you just committed to. Partnering on a layer you run keeps ownership and shortens the road. No universal answer, only the honest inventory of your team against the specification you wrote.

Now the three numbers bite hardest

Two cost curves matter more than any license fee. The flat line: AI you wait on does not stand still, the vendor improves it, but it cannot track you, and the gap between what the product does and what your operation requires widens every quarter, because closing it is a feature request in someone else's backlog. The rising line: AI your team operates compounds, because every correction and every added workflow accrues to an asset you hold. Note what does not fall: total run cost rises with volume and with each workflow. The number that must fall is cost per outcome, and that is the line for the board report and for a vendor forecast in writing.

The distance between those lines three years out dwarfs the line items on the quote. And make proof a term rather than a sentiment: name the metric, the baseline, who instruments it, the date. Then price your own side, because whatever fee is at risk, your internal cost is larger and no clause refunds it.

For formal scoring, a published instrument exists. The Right AI Infrastructure carries twelve evidence-gated dimensions across Ownership, Trust and Leverage, scored on what a contender can prove today. Use it, or steal from it. Either way you score, with criteria you own, and nobody scores for you.

Ask your CTO and CFO

  • If we switch models in a year, what survives and what do we rebuild?
  • Can it execute a real end-to-end case on our systems today, on a hard case rather than a happy path? Show it on our shortlist's second-best reference, not their best.
  • What is the all-in cost per outcome at 2x volume, at 10x, and at the volume where this stops working?
  • Three years out, what do we own that we did not have under each option? And what does it cost to stop?

After launch

The AI that ships on day one is the weakest it will ever be. Or it should be.

Whether it improves is not a property of the model but of the operation around it, and it is decided by whether anyone owns the loop.

  • Watch the right things System health tells you it is up, not that it is good. Instrument the business result, plus drift, the slow slide a month of trend lines reveals. Watch the handoff hardest, because containment tells you what the AI kept and nothing about what it handed over. Track repeat contact rate, satisfaction on handed-off cases as a separate line, and whether your person inherits full context or starts from zero. Containment rising while escalation satisfaction falls is not progress. It is a queue you made worse.
  • Expect your people's work to get harder on day one Take the easy half of the cases away and what remains is the hard half. Handling time rises, and so does emotional load, because nobody gets an easy one between two difficult ones. Re-baseline handling time and occupancy for the new mix, re-cut the forecast, tell team leads their targets are changing and why, and check any bonus scheme tied to volume or handling time before it becomes a grievance. If your first sign of success is falling satisfaction and rising attrition on the human queue, you measured the easy half.
  • Correct at the right layer Three things could have failed: the rule, the data it read, or the integration it acted through. Only the first is fixed by authoring, and a loop that only edits rules will grind rules while the real problem is a stale record. Once you know it is the rule, the fix is an edit by the person who knows the business: written, tested, shipped, often same day. That property, a correction is an edit and not a project, separates AI that compounds from AI that decays.
  • Know it works before it ships Editing rules at operator speed is only safe if every edit is tested automatically. Build a set of real past cases with the answers you would have wanted, a few hundred is enough, and run every change against it before it goes live, against a pass bar agreed in advance. High-risk changes go to a slice of traffic first, and the rule's risk tier decides who approves it. That same test set tells you a new model is safe to switch to and catches drift early. Without it, same-day edits are same-day risk.
  • Keep a cadence with named owners A weekly thirty-minute review run by the rule-owners. Put the AI's KPI into those owners' objectives, because a cadence in nobody's targets is the first meeting cancelled in a busy month. You probably already have the team: your quality function samples, scores and coaches, so point it at the AI with a scorecard rewritten for AI-handled work. It stops coaching one agent at a time and starts coaching the rule that governs thousands of cases. And budget for the rulebook itself, because authored rules accumulate and go stale.

One thread, three altitudes

Reporting is where most AI programmes go vague, so make yours concrete. The same reality at three zoom levels: decision level, daily, every action reconstructable, for operations and compliance, the raw material of trust; operation level, weekly, KPIs, drift and corrections shipped, for the function head and rule-owners, the raw material of improvement; board level, monthly, the three numbers as trends, for you and your CFO, the raw material of the scale decision.

If the number in the board pack cannot be drilled down, through the operational view, to individual decisions, you are not reporting results. You are reporting hope.

And close the loop on the money. The three numbers you used to pick the use case become the monthly report, measured in production, on a page, to the same CFO who signed the projections. Three rules keep it honest: baseline before launch, so value is distance from a measured start; count only what moved, hours actually redeployed rather than capacity theoretically freed; and carry the all-in run cost, so the ROI survives due diligence. Then add the two derived numbers that make it a decision tool: cost per outcome this month against last, and payback actual against projection.

One honesty note: none of this is the AI "improving itself." Reliable self-improving AI is not something anyone can responsibly sell you today. What compounds is your people improving the AI, with a system that makes their corrections fast, safe and cumulative. Less magical, far more valuable.

The calendar

Strategy becomes real on a calendar, and it runs on two clocks rather than one.

Day one, start both. The decision clock is yours and moves fast. The clearance clock is not yours and is longer: vendor security review, the impact assessment, third-party risk, the data processing agreement, procurement onboarding, and employee consultation where your markets require it. In most European companies of this size that runs six to twelve weeks and it, not the build, sets your go-live date. Start it in week one with a scope sketch, not in week six with a signed contract.

  • Weeks 1 to 3, decide the six In order, with the right people in the room. Output: a one-page strategy, the shortlist with three numbers each, the written noes, the always-human list, the security and data requirements collected before any vendor conversation, the impact assessment started, the consultation opened, and the budget line named with its owner. An approved use case with no funding route waits for January.
  • Weeks 4 to 6, pressure-test the case Make the three numbers survive scrutiny: baseline measured before anything launches, cashable value separated from capacity, one-off and run-rate split, a payback month, and the number at which you stop, written down before you start, with a named owner and a date for that call.
  • Weeks 7 to 13, prove it in production, one rung at a time Move the autonomy dial in one direction only, on evidence. It answers into a log and nobody sends it. It drafts and a human sends. It acts on a narrow slice with every case reviewed. It acts with sampling, and the sample shrinks only when the numbers earn it. Shadow mode is not a delay, it is your cheapest week: you find the wrong rules with no customer in the room. The impact assessment is signed before the first real customer. And if you run nights and weekends, decide who can restrict or stop the AI on a Sunday without waking an executive.

The goal is not a perfect agent. It is a proven loop: value in the KPI, corrections flowing from your own experts, the three numbers re-measured against projection. That, running, is a strategy the board can see, and expanding it becomes evidence instead of faith.

One sizing note, because you will have to sell this internally. Make the first step fit a budget you already control. A fixed-scope assessment and one workflow is an operating decision. A platform programme is a capital committee and a nine-month wait.

The prize

The first use case is the most expensive one you will ever build. Everything after gets cheaper.

Not because the technology improves, but because of what the first one leaves behind. By the time it runs in production you hold five assets that never need building twice: the context, your vocabulary, rules and procedures, a meaningful share of which applies to the next workflow as-is; the team, experts who know how to author and developers who know the wiring; the governance, because CISO, DPO and legal have already reviewed the identity model, the access model and the always-human list, so the next approval is a delta review; the integrations, systems of record already connected under controls already trusted; and the method, so the next selection takes days rather than months.

You are not buying AI projects. You are standing up an AI capability.

Projects have budgets and endings. A capability has a rhythm: each quarter more workflows live, more of the operation authored, cost per outcome falling while coverage grows.

One condition decides your roadmap. Cheaper holds while you stay inside the same systems, the same data class and the same approvals. Cross any of those and you are close to paying first-build again. The second market is cheaper, not free: rules and method travel, approved wording does not, since each language needs local legal to sign it. So sequence by what reuses your assets, not by what is loudest. Betsson's pattern is exactly that: prove it on one operation, then expand across brands on the assets the first win created.

And this is the competitive math. Two competitors start together. One runs AI projects: each starts near zero, flat-lines, gets renegotiated yearly. The other compounds, its operation broader and its cost per outcome lower every quarter. For a year they look similar from outside. By year three the gap is wide, and here is the honest version of what it is worth. Can a competitor buy their way in? Partly. They can buy a platform, hire your rule-owners, even acquire a company. What they cannot buy is three years of their own people's corrections against their own edge cases. Not an unbridgeable moat. A head start that costs real money to close, which is what a moat is worth in practice.

Ask yourself, every quarter

  • What did the latest use case cost against the first, and why?
  • How many workflows and named rule-owners does our operation have now, versus last quarter?
  • What is the trend in cost per outcome, and in the share of the operation that is authored?
  • Which team joins next, and what do they inherit for free?

The check

Fourteen questions. Every no is a hole, and the chapter numbers tell you where to dig.

Before you take any of this to the board, run your current plan against these.

The use case

  • Have we scored at least three candidates on impact, speed to deployment and provability?
  • Do we know what is already automated today, so we are not rebuilding it?

The context

  • Can we name where our operational knowledge lives today, and its three biggest gaps?
  • Do we have a path from what our people know to written, testable rules with named owners?

The team

  • After go-live, can a named operations person change a rule without an engineering ticket?
  • Does every rule have an owner with a name, not a department?

Governance

  • Can we reconstruct any single AI decision: what it was given, the rules in force, the controls that fired, the action taken?
  • Is the always-human list written, signed by our DPO, and enforced by design?

Security and compliance

  • Has our CISO seen the access and identity model, before the vendor decision?
  • Is our impact assessment signed and dated before the first real customer interaction?

The infrastructure

  • Have we scored our shortlist, including our own build option and whatever our incumbent platforms already ship, against criteria we wrote?
  • Do we have a set of real test cases and a pass bar that every rule change and every model change must clear?

The three numbers

  • Does the case name a payback month, a cashable share, and the number at which we stop?
  • Is there a date on which those numbers get re-measured in production, and a name attached?

Now count the noes

  • 0 to 2 You have a strategy. Go pressure-test the number-one case and set the stop rule.
  • 3 to 6 The normal score for a company that has run a pilot. The plan is sound and the gaps are all in the same place: nobody has produced the three numbers on a real baseline. That is the next piece of work, and it is the one that decides funding.
  • 7 or more Do not start with a vendor conversation. Start with decisions 1 and 2, on one workflow, and get an owner named. A shortlist built on this many unknowns will be scored on the wrong criteria.

Fourteen yeses does not mean the strategy is finished. It means it has no silent holes.

The honest part

A guide that only says where a model works is marketing. Here is where this one does not.

It does not pay below a certain volume: a workflow running a few hundred times a month will not return the authoring effort. It does not work where no rules exist and nobody will own writing them, because there is nothing to author. It does not work mid-restructuring, since you would author a process you are about to delete. And it does not work where the domain expert has no free hours, however willing they are.

If you recognise your company in that list, the honest next step is not a pilot. It is fixing one of those four things first.

The next piece

Almost everything here you can do yourself, and you should. One piece resists.

If your score landed in the normal band, your gap is the three numbers on your number-one use case, and that is the artifact your CFO will attack hardest. Producing it means measuring a real baseline before anything launches, separating cashable value from capacity and getting someone to commit to absorbing the capacity, pricing your own people's days at a loaded rate, and scoping at workflow level rather than department level.

Every one of those requires the people who are currently running the operation, which is why in most companies this phase takes a quarter of part-time argument and arrives at a number nobody fully believes. That is the moment where an outside pass earns its fee: not for the strategy, which you now have, but for the measurement, done by people who have produced this number for operations like yours and who can tell you, credibly, that the honest answer is do not proceed.

The companies that win with AI are not the ones that bought the best model, but the ones that built the operation around it, owned by their own people, provable turn by turn, better every week.

The only way to have a three-year-old compounding operation is to have started three years ago. The second-best time is now, with one workflow, on purpose.

Sources

Every number on this page, with its origin and its caveat.

  • MIT NANDA, The GenAI Divide: State of AI in Business, 2025 The 95 percent figure on pilots without measurable P&L impact. Preliminary research, contested on sample size and on how failure is defined.
  • IDC for Lenovo, CIO Playbook 2025, March 2025 For every 33 AI proofs of concept launched, four reached production.
  • Gartner, August 2025 40 percent of enterprise applications expected to feature task-specific AI agents by the end of 2026, up from under 5 percent in 2025.
  • Gartner, June 2025 Over 40 percent of agentic AI projects expected to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Based on a poll of more than 3,400 organisations investing in the technology.
  • McKinsey, global State of AI survey End-to-end workflow redesign is the strongest differentiator between organisations capturing value from AI and those stuck in pilots. About one company in five has done it.
  • Stanford HAI, AI Index Convergence among leading models on standard evaluations.
  • Regulation (EU) 2026/1744, the Digital Omnibus on AI Official Journal 24 July 2026, in force 27 July 2026, amending Regulation (EU) 2024/1689. High-risk obligations for standalone Annex III systems deferred to 2 December 2027, and for AI embedded in regulated products to 2 August 2028. Article 50 transparency obligations apply from 2 August 2026 and were not deferred.
  • Regulation (EU) 2024/1689 and Regulation (EU) 2016/679 The EU AI Act, Articles 4, 25, 26 and 50. The GDPR, Articles 5, 22, 28, 33 and 35.
  • Betsson France Figures and quotations published with the client's approval.

Verified July 2026. Regulatory positions move; check the current text before relying on any date here.

The next step

Run the method on your own operation.

A short, fixed-scope assessment takes your number-one use case and comes back with the three numbers, the KPIs a pilot would commit to, and, if the honest answer is do not proceed, that answer in writing.