CoreWeave Forge | 2 October 2026 | Agent engineering

Your agent's most valuable dataset is its production failures.

Maddipalli Gopalakrishna · AI / ML Engineer

65 seconds · music only

A failure that ends in a monitoring dashboard teaches the next version nothing. Most agent teams still run the lifecycle they inherited from classic ML:

  1. Build
  2. Test
  3. Deploy
  4. Monitor

That shape assumes one output per request and a model that changes rarely. Agents break both assumptions. They take dozens of actions per task, and the interesting failures sit in the middle of those actions, not at the end. Production behavior has to flow back into the next version on purpose:

  1. Run
  2. Observe
  3. Curate
  4. Improve
  5. Evaluate
  6. Release
  7. Repeat

What CoreWeave shipped

On 30 September 2026 CoreWeave launched Forge, which it describes as a development layer that runs that loop in one environment. Three details matter for this argument, quoted from CoreWeave's own pages:

Forge “runs the entire AI loop – run, observe, curate, improve, evaluate and repeat – in one connected environment.”
Agent Lens: “Each detected failure becomes a test case in a growing evaluation set. Run a candidate fix against the entire set … and catch regressions before promoting the change to production.”
Domain experts tune the LLM judge and review disagreements with human scores “before the judge automatically evaluates every new production trace.”

Forge brings together Weights & Biases Models, OpenPipe's post-training work and the marimo notebook project. CoreWeave also publishes its own performance figures; I've left them out, because they are vendor benchmarks I can't check. The rest of this piece is my reading of the engineering lesson, and none of it depends on using Forge.

Every failure type maps to an eval

The useful move is to name the failure precisely enough that it becomes testable.

FailureBecomes
Hallucinationan unsupported claimGroundedness eval
Bad retrievalirrelevant evidenceRetrieval-quality eval
Wrong toolincorrect selectionTool-routing eval
Bad tool argumentsright tool, wrong parametersStructured tool-call eval
Trajectory failureright answer, unsafe or wasteful pathTrajectory eval
Policy violationa forbidden actionPolicy eval

The last two rows are the reason final-answer scoring falls short for agents. An agent can land on the right answer after calling a tool it had no business calling. A grader that only reads the answer marks that run as a pass.

Evaluate the layers, not just the answer

A classic eval is input → model → output → score. An agent run has a plan, tool calls, observations and state in between, and each layer can fail independently.

L6
System
latency, cost, reliability, task completion, human escalation
L5
Safety / policy
authorization, sensitive data, forbidden actions
L4
Trajectory
a reasonable path, recovery from failure, no needless repetition, workflow constraints
L3
Tools
correct tool, correct arguments, sequencing, error handling
L2
Retrieval
recall, precision, evidence relevance, citation correctness
L1
Output
correctness, relevance, groundedness, completeness

Software already solved the shape of this

Engineers trust one rule: a bug that reached production gets a test, so it can't ship twice.

  1. Bug
  2. Reproduce
  3. Write test
  4. Fix
  5. Test passes

The AI version adds a trace and a gate:

  1. Failure
  2. Capture trace
  3. Reproduce
  4. Create eval
  5. Fix
  6. Regression eval
  7. Release gate

Every important production failure should be able to become a regression test. Not every failure deserves one. The ones that recur, cost money or touch policy do.

Keep the loop controlled

A feedback loop is not permission for a system to rewrite itself in production. Every improvement still goes through dataset curation, engineering review, evaluation, regression testing, security checks and a release gate. CoreWeave's own pages describe curation as human-in-the-loop, which is the right default. The goal is an agent that gets better from production under control, release by release.

Which failure type is hardest for your team to turn into a repeatable eval?