01 Trigger Event
Lenny's Newsletter's latest "How I AI" episode featured Claire Vo (founder of ChatPRD) demoing a product called Jev, made by TypeSafe AI. Jev describes itself as a "decision model, not a language model" — input priced at $0.04/M tokens, with output tokens completely free. Claire demoed a few numbers: 9 cents to analyze 1,700 PRs, $4 to process 200,000 classifications, and total Jev spend for an entire week was under $10.
She repeatedly referenced a paired pattern: Jev handles classification/clustering/routing, while frontier models handle reasoning. The original article mentioned the name "GPT-6 Astra" — I couldn't verify this against any public roadmap. It might be a podcast transcription error or a gap in my knowledge, so I'll hedge on that.
02 What This Really Means
The real significance here isn't "yet another AI startup launching a new product" — it's a structural shift in token economics: inference is being decoupled into two layers of abstraction.
Over the past 18 months, the language model has been a single abstraction — you give it a prompt, it spits out text. For classification, you ask the model to output a field in JSON, then pay for it out of output tokens. This abstraction assumes output is a product of "generation," and "generation" is expensive.
What Jev-style products abrogate is this: most production workloads don't need generation at all — they need constrained decisions. Give me a piece of text; I need a category, score, probability, or routing label — a 1-of-K categorical output. This isn't generation; it's retrieval-on-steroids.
Price comparison (Claude Haiku 4.5 ~$1/$5, Gemini Flash Lite ~$0.10/$0.40, GPT-4.1 mini ~$0.40/$1.60, all per M input/output):
A typical scenario: 100K input + 50 token output for classification
- Claude Haiku 4.5: $0.10 + $0.25 = $0.35
- Gemini Flash Lite: $0.01 + $0.02 = $0.03
- Jev: $0.004 + $0 = $0.004
The gap is over 80x. This isn't optimization — it's an order-of-magnitude reconstruction. And Jev's output format is a fixed schema (category/score/probability), with no risk of "the model wanting to say more," so retry rates should be far lower than traditional LLMs.
This structural shift, I think, goes further than model routing (routers distributing requests across different models) — routing is request-level splitting, while decision models tear apart the inference shape itself.
03 Historical Analogy
The 2014 Lambda vs EC2 split was an early version of the same pattern: same AWS, two abstractions, two different price elasticities. Lambda was mocked at the time for "cold start latency / limitations," but a decade later it's one of AWS's fastest-growing services. The point isn't that Lambda is "better" than EC2 — it's that it unlocked a class of workflows EC2 fundamentally couldn't cover (event-driven, short-duration, bursty).
The 2017-2019 TPU vs GPU split is another parallel: the same underlying class of compute, different shape. TPUs vastly outperform GPUs on inference, but NVIDIA's CUDA moat kept TPUs running only within the Google ecosystem.
Where do Jev-style decision models sit? I'm inclined to say it's the Lambda position, not the TPU position. Reason: the foundation is likely still a transformer (or some kind of distilled transformer); it doesn't require specialized hardware and can run at the software level. The moat isn't in infrastructure — it's in productization. Whoever first turns this workflow (classification → reasoning) into developers' default mental model wins. This is the same move AWS made when they packaged serverless as "you don't have to think about servers anymore."
Counter-risk exists too: if "decision model" is essentially just Llama 3.1 8B or Phi-3-mini wrapped in a constrained decoding layer, then the moat is near zero, and OpenAI/Anthropic can roll out a "structured output mode" at any time and eat this category directly. I haven't run internal benchmarks myself, so I can't conclude.
04 What This Means for AI Builders
Several decisions worth adjusting this month/quarter:
First, recalculate unit economics. Run your production workload once and mark out the parts that "actually only need classification/scoring/routing." If that share exceeds 30%, you're burning money for nothing. A 100K-token classification task, run 10,000 times a day, migrated from Haiku to a decision-class model, can save ~$1M annually — this isn't an exaggeration, it's the 80x gap above multiplied out directly.
Second, the routing layer is shifting from "optional" to "standard." The standard pattern for production systems: a cheap classifier (decision model) up front for screening, a frontier reasoner (Claude/GPT/Gemini) in the back handling only the high-difficulty requests that survived the screen. This two-tier pattern is isomorphic to 2023's RAG retrieval-then-generation — use cheap tools to narrow the problem space, then use expensive tools for the core.
Third, monitor your own workflow migration. Claire herself discovered that in January, 100% of her AI usage was coding-related; by September it had dropped to 40%, with the rest filled by agents and media categories. This isn't a personal phenomenon — this is the entire builder community's workflow shifting from "I directly use frontier models" to "I orchestrate multiple models." Application-layer founders' product roadmaps should account for this trend.
Fourth, watch out for lock-in. Free output sounds enticing, but the input-side price is the platform's only revenue lever. If 90% of your workload runs on a decision model, your pricing power is zero — the platform can raise the input price at any time. This is reverse lock-in, structurally identical to AWS egress fees. In contract negotiations, input price should be tied to volume — don't sign fixed unit prices.
05 Counter-arguments / Risks
I may be misjudging in several places; I'm writing them out to remind myself.
First, "decision model" is likely not a new architecture — it's a wrapper. TypeSafe hasn't published model parameters, isn't open source, and has no benchmarks. I haven't run Jev internally; the 80x gap above is based on official pricing and doesn't account for quality differences. A 9-cent PR analysis demo: real-world workload complexity is usually 5-10x higher than the demo, plus retries and fallback, so actual unit cost will be significantly higher than the displayed numbers. Don't calculate ROI from marketing demos.
Second, frontier model vendors' response will be underestimated. Anthropic and OpenAI can (and already are) rolling out "structured output modes" + drastically cutting output prices. Claude's prompt caching has already done similar tiered pricing. If frontier vendors cut output token prices to $0.10/M instead of $0, decision models' price advantage evaporates. My personal bet is that frontier vendors will do this, within 6-12 months.
Third, the name "GPT-6 Astra" — I haven't seen it in any public information. If it was a slip in the podcast or a transcription error, then I haven't verified the frontier model tier's reference point, and some of the numerical comparisons above may need correction. If GPT-6 actually exists and Astra is a tier, that's another story, not covered here.
Fourth, decision models' quality ceiling may be low. They excel at structured output, but they're not good at reasoning. Once the task gets slightly complex (e.g., multi-hop classification, judgment requiring world knowledge), it falls through to frontier models. In other words, decision models don't replace frontier models — they supplement them. This weakens their TAM claim. The real killer use cases are still on frontier models; decision models are just a filter. Long term, this might be a router-layer play (e.g., Martian / Not Diamond / OpenRouter's approach), not an independent category.
After reading it back, I think what I most need to watch for is points one and two — pretty demo numbers + frontier vendor counterattack. That combination will most likely compress the window for independent decision model vendors like TypeSafe. Those who make it will either be acquired and integrated by frontier vendors, or develop a composite layer of routing + decision (becoming an orchestration platform like Not Diamond / Martian). For pure single-point decision models, I'm neutral-to-skeptical.