Agentic system design · 30 September 2026
Typed decisions, not paragraphs.
What Jev gets right about agentic systems.
Maddipalli Gopalakrishna · AI / ML Engineer
Most agent failures I've debugged weren't reasoning failures. They were parsing failures.
The model wrote a good paragraph. Then the software downstream had to guess what it meant, whether it was sure, and what to do next.
TypeSafe AI's new model, Jev, starts from the opposite end. Instead of writing text you then have to parse, it returns a typed decision with a probability attached. It's a small idea with large consequences for how we design agent systems.
What Jev actually is
TypeSafe calls Jev the first of its “System One” models. The name borrows from fast, intuitive thinking: quick, structured calls rather than long deliberation.
The contract is simple:
- Input: unstructured state, meaning documents, records and context.
- Output: a value from a set of options defined in advance, plus a probability for that value.
- Training: a method TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD). It's aimed at probabilities that are honest, rather than at answers people prefer.
TypeSafe describes the product as a frontier-intelligence function call.
TypeSafe reports large speed and cost gains over general-purpose LLMs on these tasks. It quotes response times of 70 to 500 milliseconds. These are the vendor's own benchmarks from an early-access release, and its announcement discusses their limits openly. I'd treat them as promising, not settled.
Why decision tasks need structure
Much of what we call “agentic AI” is actually a string of small decisions:
- Is this record complete?
- Do these two sources agree?
- Which of four categories does this project fall into?
- Is the evidence strong enough to go on, or should a human take over?
General-purpose LLMs answer these in prose. Then we write fragile code to pull a decision out of the prose. Then we bolt on a confidence score the model was never trained to produce.
That's the gap Jev targets. If the set of possible outputs is fixed and the probability is calibrated, the step where prose gets translated into a decision disappears.
Where this meets work I've already shipped
I've been building toward the same idea from the application side, without a model like Jev.
In NTO Operations Copilot, a research coach for Notice to Owner researchers, every answer follows one fixed format: Answer, Why and Next step. When a customer's claim conflicts with a recorded document, the system doesn't pick a winner. It keeps both, labels where each came from, and sends the case down an approved escalation path.
In my work-order orchestration project, a deterministic Python layer decides what happens next. The AI agents propose; the code decides.
Both are the same instinct: don't let prose drive the workflow. Jev moves that instinct into the model itself.
A blueprint for putting a decision model into an agent pipeline
This is how I'd approach adding a typed decision model to an agent pipeline. It's a design plan, not a report of a finished integration.
Ingestion: state the decision before choosing a model
- Write out the output type first. Every allowed value, including NEEDS_HUMAN_REVIEW.
- Label where each input came from: customer claim, official record, or a system-generated value.
- Keep the model read-only and restricted to approved tools, like the read-only MCP boundary in NTO Copilot.
Optimization: send each task to the right tool
- Fast, typed decisions go to a System One-style model.
- Explanations and open-ended questions stay with a general LLM.
- Hard rules stay in plain code. No model should decide a legal deadline.
Validation: decide what “sure enough” means
- Set a confidence threshold for each decision type, based on how costly a mistake is.
- Below the threshold, hand off to a human. Don't retry until the output looks better.
- Check the probabilities against real outcomes. A calibration claim is worth only as much as your own measurements.
A sketch of the decision boundary
This is an illustrative sketch of the pattern, not Jev's API. TypeSafe's early-access interface may look different.
from dataclasses import dataclass
from enum import Enum
class GCResolution(Enum):
MATCHES_RECORD = "matches_record"
CONFLICT_PRESERVE = "conflict_preserve"
INSUFFICIENT_EVIDENCE = "insufficient_evidence"
@dataclass(frozen=True)
class Decision:
value: GCResolution
probability: float # calibrated confidence from the decision model
sources: tuple[str, ...]
# Stricter threshold where a wrong call is expensive.
THRESHOLDS = {
GCResolution.MATCHES_RECORD: 0.95,
GCResolution.CONFLICT_PRESERVE: 0.80,
GCResolution.INSUFFICIENT_EVIDENCE: 0.0, # always safe to stop
}
def route(decision: Decision) -> str:
"""The policy lives in code; the model only proposes a decision."""
if decision.probability < THRESHOLDS[decision.value]:
return "escalate_to_human"
if decision.value is GCResolution.CONFLICT_PRESERVE:
return "cc_first_confirmation_path"
if decision.value is GCResolution.INSUFFICIENT_EVIDENCE:
return "request_missing_evidence"
return "continue_research"The model returns a typed value with a probability. Plain code turns that into the next step. A human handles everything the code doesn't clear.
Where this fits on this site
My portfolio follows the same rule as my systems: clear structure, stated uncertainty, no inflated claims.
Every case study separates what's implemented from what's a target. The goal is the same one Jev aims for at the model level: output a reader, or a machine, can rely on without guessing.
Let's talk
If you're designing agent systems where a wrong decision costs real money, where autonomy should stop is the most important design question.
I'd like to compare notes. Connect with me on LinkedIn and tell me where your agents still depend on prose to make decisions.