AI Operations · Glossary
EvalOps
EvalOps is architectural requirement 3 of an Interactive Agent: the operating practice of structured evaluation at scale, run against authored Policy criticality and instrumented through LLM-as-a-Judge, evaluation suites, and per-turn evaluation.
EvalOps is a practice, not a vendor product. The practice runs evaluations continuously against production decisioning and against pre-deployment changes to the policy library. The instruments of the practice are concrete. LLM-as-a-Judge scores agent outputs against authored criteria. Evaluation suites bundle the criteria that matter for an agent into a structured artefact run on a schedule and on every change. Per-turn evaluation runs at the granularity of the conversational turn, not at the session boundary. Domain experts who own the affected Policies and Routines author the evaluation criteria; the AI Operations Lead governs the Suite composition. EvalOps is what gives the AI Operations Lead a quantitative grip on agent behaviour. It is the evidence layer beneath the autonomous improvement loop and the human governance gate.