AI evals | 3 October 2026 | LLMOps

Your AI eval suite can become technical debt.

Maddipalli Gopalakrishna · AI / ML Engineer

60 seconds · music only

In yesterday's post I argued that production AI failures should become regression opportunities. A fair objection came back: if every failure turns into a permanent test, doesn't the suite eventually become the bottleneck?

It does, if the rule is “failure → eval”.

  1. 10
  2. 100
  3. 1,000
  4. 10,000
  5. 100,000 evals

At that scale you get duplicated tests, slow CI, expensive LLM-as-judge runs, stale and flaky evals, and a release gate nobody trusts.

Failure to eval cannot be automatic

The step between a failure and an eval has to be triage, not a copy:

  1. Failure
  2. Fingerprint
  3. Known / novel
  4. Risk
  5. Human review if needed
  6. Promote
  7. Tier

Take an illustrative batch of 10,000 production failures:

GroupWhat happens
9,700 match known clustersUpdate the existing eval's metadata: frequency, severity, last seen. No new test.
250 are variationsExtend an existing eval only if coverage is actually missing.
49 are low-impact noiseMonitor; revisit if they recur or escalate.
1 has no cluster matchNovelty queue. It is never dropped for having a count of one.

Ten thousand failures, and possibly one new test.

Deduplication has a blind spot

A failure that appears once has no cluster to join, and a frequency-based filter quietly drops it. That one case might be a new attack pattern, an authorization gap, an unusual RAG failure or a tool behaviour nobody has seen before. That isn't always true, since many singletons are noise. But it's the reason frequency alone shouldn't decide promotion.

Deduplication controls volume. Novelty detection protects coverage.

Each failure gets a handful of signals: severity, business impact, recurrence, novelty, safety and security impact, customer impact, reproducibility and triage confidence. Combine them however your team decides:

Risk = f(severity, impact, recurrence, novelty, safety, confidence)

This is a conceptual framework, not a standard formula. The only rule that matters is that recurrence is one input among several.

Who triages

Both. Automation handles fingerprints, embeddings, clustering, near-duplicate detection, recurrence counts, metadata extraction, a first severity guess, novelty scoring and routing. People handle security, safety and policy-sensitive cases, novel behaviour, ambiguous cluster assignments, low-confidence calls and the promotion of critical behavioural boundaries.

  1. Failure
  2. Automated triage
  3. Confidence / risk check

Known, low-risk cases are handled automatically. Novel or high-risk cases go to a person.

The triage needs evals too

Once automation decides what gets promoted, the triage step becomes a system with its own failure modes:

Triage failureWhat it looks like
False mergeTwo genuinely different failures end up in one cluster.
False splitOne underlying failure becomes several clusters.
Missed noveltyNew behaviour is classified as known.
False noveltyKnown behaviour is escalated as new.
Bad prioritizationA critical failure is ranked low.
Missed promotionAn important boundary never reaches permanent coverage.

Audit a labelled sample of triage decisions on a schedule, and track incidents that an existing cluster should have caught.

Tier the suite

Not every eval needs to block every deployment.

TierContents and cadence
T1 Release gateCritical deterministic checks, security boundaries, high-impact regressions, core agent behaviour, mandatory policy. Small, fast, blocking. Runs on every release.
T2 CI regressionKnown historical failures, tool use, RAG quality, routing, structured output, workflow scenarios. Runs on every change.
T3 Deep evalsLong trajectories, adversarial and red-team cases, expensive LLM-as-judge, multi-turn, large retrieval benchmarks. Scheduled or before major releases, usually non-blocking.
T4 Production discoveryTraces, drift, unusual trajectories, novel tool use, rare failures. Continuous; its output is the novelty queue.

A novel, high-risk singleton can earn a place in the release gate. Most promoted cases belong in CI regression.

Evals need a lifecycle

  1. Candidate
  2. Active
  3. Critical
  4. Monitored
  5. Redundant / stale
  6. Archived

Review an eval when it turns redundant, stale, flaky, superseded or too expensive for its tier. Version the suite (eval-suite-v1.0, v1.1, v2.0) and keep a record per eval: the behavioural boundary, originating incident, reason added, owner, severity, tier, last failure, last review and the model and system versions it ran against. Archiving keeps that lineage; it doesn't delete history. These fields are a suggestion, not a standard schema.

Build a platform, not a pile

Four-layer eval governance platform: production observability, failure intelligence, eval governance and evaluation execution, with security, cost, observability, governance and auditability rails.
Conceptual reference architecture · not a product or standard

Evals start to look like any other test platform: lifecycle, versioning, ownership, observability, cost controls and release policy.

The challenge isn't collecting more evals. It's preserving the right behavioral boundaries without allowing the evaluation system itself to become the bottleneck.

How are you deciding which production failures deserve permanent regression coverage?