Skip to content
AI Solutions

The reason our AI survives production.

AgentOps is the operational layer around AI in production. iLeaf's Governor Engine enforces what an agent may do, how much it may spend and how long it may take, escalates to a human instead of guessing, and records why every action was taken — so a decision can be reconstructed months later during an audit.

Ask Lia about agentops & mlops

What this includes

Governor Engine
Runtime policy: permitted tools, spend ceilings, latency budgets and blocked actions, enforced per agent.
Cost control
Per-agent and per-tenant budgets with model routing, so the token bill is a decision rather than a surprise.
Evaluation gates
Prompt and model changes tested against your own task suite in CI before they can reach production.
Audit trail
Goal, tools called, reasoning and outcome recorded for every run and retained for review.
Escalation paths
Defined routes to a human on low confidence, limit breach or repeated failure. Nothing is silently retried.
Model lifecycle
Versioning and rollback for models, prompts and retrieval indexes, independent of application deploys.
Mapped to the frameworks you are audited against
The controls above line up with the EU AI Act's requirements for high-risk systems and the NIST AI Risk Management Framework — logging, human oversight, traceability — so a procurement questionnaire is answered from what already runs, not written afterwards.

How we run it

Every phase ships something usable on its own, so you are never holding a half-finished system waiting on the next milestone.

  1. Write the policy first

    What the agent may never do gets agreed before it is built, in plain language, with your risk owners in the room.

  2. Instrument from the first run

    Cost, latency, tool errors and escalations are emitted from day one, not added once something goes wrong.

  3. Gate the pipeline

    Evaluations run in CI, so a regression is caught by a failing build rather than by a customer.

  4. Rehearse the failure

    We test escalation and rollback deliberately, because an untested failure path is not a failure path.

What we build it with

  • OpenTelemetry
  • Grafana
  • Prometheus
  • Temporal
  • Kubernetes
  • Python
  • TypeScript
  • PostgreSQL
  • Redis
  • GitHub Actions
  • Terraform
  • On-premise GPU
  • Governor Engine

Questions we get asked

Is this a product we license?

It is delivered as part of the engagement, running in your own environment on open components — not a black box you rent. You keep the configuration, the logs and the ability to run it without us, which matters because governance you cannot inspect is not governance.

Which numbers do you actually run this against?

Task completion accuracy, agent handoff rate, time-to-first-token, semantic faithfulness of answers to their sources, and cost per workflow. They are instrumented from the first run rather than added after something goes wrong, because an agent that is fast, cheap and confidently wrong looks healthy on every metric except the one that matters.

Does this help with the EU AI Act or NIST AI RMF?

Yes, though we are engineers rather than your compliance function. Both regimes ask for the same underlying things — records of what an automated system did and why, meaningful human oversight, and traceability from a decision back to its inputs. The Governor Engine and audit trail produce that as a by-product of running, so evidence is exported rather than reconstructed. Your legal team still owns the classification and the filing.

What does the audit trail actually capture?

For each run: the goal, the tools called with their arguments, the model and prompt version, the reasoning, the cost, the latency and the outcome — including whether a limit was hit or a human was involved. Enough to reconstruct a decision months later for a regulator or an incident review.

Can you govern agents we built ourselves?

Often yes. If your agents call tools through an interface we can sit in front of, the Governor can enforce limits and record actions without rewriting them. We start with an assessment of how your agents invoke tools, since that determines what is enforceable.

Let’s talk about agentops & mlops.

Tell us what you are running and what it needs to do next. We will tell you honestly whether we are the right team for it.

Talk to a solutions lead