NewInteractive Agents Live

Domain Experts view

Make your AI measurably better over time.

You decide what a good answer looks like, score the runs that matter, and correct the ones that miss. Every correction becomes a test case you can check the next version against.

This week's reviewrefund assistant
Scoring
Right answer94%up from 88%
Right tone91%up from 88%
Made something up1%down from 4%

Your corrections seed the test cases the next version is checked against. 2 runs flagged for your review.

You set the bar · Corrections become test cases · On the model you already have · Fully reversible

An agent that can't get better is a liability with a launch date.

Every run tells you something. The loop turns what you learn into the next, better version of the agent, and the people who know the work own every step of it. The control and enforcement that make that safe come from one environment you share with your engineers.

Evaluate
Annotate
Refine
Ship
01 · Evaluate

Catch the regression before your customers do.

You set the baseline for a good answer once, in plain language. An evaluator turns that baseline into a score on every run; evaluation runs your evaluators across the cases that matter. So you see the miss, and the reason for it, before a customer ever does.

Set the baseline

Define what a good answer looks like once, in plain language.

Evaluators produce scores

An evaluator checks each run against your baseline and produces a score, with the reason, not just a number.

Evaluation, before it spreads

Run your evaluators across live runs to catch a drop in quality, not next quarter's complaints.

See these scores land on live runs in Governance
Live runrefund assistant
scored

“Refund a $90 order from 34 days ago?”

Your baselineRefund within 30 days
Result✗ Approved · misses the bar
Whyignored the 30-day window

Caught on a live run, before a customer ever saw it.

02 · Annotate

The people who run the agent decide what good looks like.

When a run is wrong, the people who know the domain correct it, and that correction becomes the ground truth the next version is measured against. No labeling vendor, no engineer translating your standards into code.

Correct the miss

Fix the run that's wrong, in your own words.

It becomes ground truth

Your correction is what the agent is measured against from now on.

Experts, not vendors

The people who know the work, not an outsourced labeling team.

Needs your reviewrefund assistant

Agent said

Refund approved.

Your correction

Decline: past 30 days, and not a faulty item.

Mark as ground truth

Now a test case the next version has to pass.

03 · Refine

Fix the behavior in plain language.

A miss is usually a missing rule or the wrong context, not a worse agent. You adjust it yourself, in plain language in the Co-Pilot, and it takes effect on the model you already have, no retraining.

You author the fix

Change the rule or the knowledge yourself, in the Co-Pilot.

It's a context change

The agent's behavior is its context, governed and reversible.

Live on the next run

On the model you already have, no retraining, no rebuild.

How experts author context in the Co-Pilot
Refund policyplain language

Refunds are allowed within 30 days of purchase.

+ Exception: faulty items can be refunded any time.

Saved by Maria · Support Lead✓ live on the next run

On the model you already have. No code, no retraining.

04 · Ship

Ship the better version without an engineering ticket.

Prove the change against the runs that mattered, then promote it: no ticket, no release train, no waiting weeks. Because the agent already runs here, the better version is one decision away, and you decide who owns what.

Yesterday's traces, tomorrow's test cases

Promote any run you cared about into the set the next version must pass.

Run it before you ship

Check the candidate against that set, then a person makes the call to promote.

No ticket, no wait

It ships on the infrastructure the agent already runs on.

The infrastructure it ships on
Before you shiprefund assistant · v2
Checking

Checked against 142 saved cases

140 passed▲ +0.06 vs current version
Partial refund after 30 daysNeeds your review
Disputed chargebackNeeds your review
A person makes the call to promote.Promote

Then the next run starts the loop again.

Improvement

Make AI that compounds, not decays.

Bring a live workflow. We'll map how your experts set the bar, score the runs, and turn corrections into a loop that makes the agent measurably better.