CoreWeave Forge | 2 October 2026 | Agent engineering
Your agent's most valuable dataset is its production failures.
Maddipalli Gopalakrishna · AI / ML Engineer
65 seconds · music only
A failure that ends in a monitoring dashboard teaches the next version nothing. Most agent teams still run the lifecycle they inherited from classic ML:
- Build
- Test
- Deploy
- Monitor
That shape assumes one output per request and a model that changes rarely. Agents break both assumptions. They take dozens of actions per task, and the interesting failures sit in the middle of those actions, not at the end. Production behavior has to flow back into the next version on purpose:
- Run
- Observe
- Curate
- Improve
- Evaluate
- Release
- Repeat
What CoreWeave shipped
On 30 September 2026 CoreWeave launched Forge, which it describes as a development layer that runs that loop in one environment. Three details matter for this argument, quoted from CoreWeave's own pages:
Forge “runs the entire AI loop – run, observe, curate, improve, evaluate and repeat – in one connected environment.”
Agent Lens: “Each detected failure becomes a test case in a growing evaluation set. Run a candidate fix against the entire set … and catch regressions before promoting the change to production.”
Domain experts tune the LLM judge and review disagreements with human scores “before the judge automatically evaluates every new production trace.”
Forge brings together Weights & Biases Models, OpenPipe's post-training work and the marimo notebook project. CoreWeave also publishes its own performance figures; I've left them out, because they are vendor benchmarks I can't check. The rest of this piece is my reading of the engineering lesson, and none of it depends on using Forge.
Every failure type maps to an eval
The useful move is to name the failure precisely enough that it becomes testable.
| Failure | Becomes |
|---|---|
| Hallucinationan unsupported claim | Groundedness eval |
| Bad retrievalirrelevant evidence | Retrieval-quality eval |
| Wrong toolincorrect selection | Tool-routing eval |
| Bad tool argumentsright tool, wrong parameters | Structured tool-call eval |
| Trajectory failureright answer, unsafe or wasteful path | Trajectory eval |
| Policy violationa forbidden action | Policy eval |
The last two rows are the reason final-answer scoring falls short for agents. An agent can land on the right answer after calling a tool it had no business calling. A grader that only reads the answer marks that run as a pass.
Evaluate the layers, not just the answer
A classic eval is input → model → output → score. An agent run has a plan, tool calls, observations and state in between, and each layer can fail independently.
- L6
- System
- latency, cost, reliability, task completion, human escalation
- L5
- Safety / policy
- authorization, sensitive data, forbidden actions
- L4
- Trajectory
- a reasonable path, recovery from failure, no needless repetition, workflow constraints
- L3
- Tools
- correct tool, correct arguments, sequencing, error handling
- L2
- Retrieval
- recall, precision, evidence relevance, citation correctness
- L1
- Output
- correctness, relevance, groundedness, completeness
Software already solved the shape of this
Engineers trust one rule: a bug that reached production gets a test, so it can't ship twice.
- Bug
- Reproduce
- Write test
- Fix
- Test passes
The AI version adds a trace and a gate:
- Failure
- Capture trace
- Reproduce
- Create eval
- Fix
- Regression eval
- Release gate
Every important production failure should be able to become a regression test. Not every failure deserves one. The ones that recur, cost money or touch policy do.
Keep the loop controlled
A feedback loop is not permission for a system to rewrite itself in production. Every improvement still goes through dataset curation, engineering review, evaluation, regression testing, security checks and a release gate. CoreWeave's own pages describe curation as human-in-the-loop, which is the right default. The goal is an agent that gets better from production under control, release by release.
Which failure type is hardest for your team to turn into a repeatable eval?