Back to list
Sep 19, 2026
2.7K
0
0
Other LLM

TypeSafe AI Launches Jev, a New System One Model

Ex-OpenAI researcher Diogo Almeida's TypeSafe AI released Jev in early access, a structured-decision model priced at $0.042 per million input tokens.

#TypeSafe AI#Jev#System One Model#Structured Output#AI Agents
TypeSafe AI Launches Jev, a New System One Model
AI Summary

Ex-OpenAI researcher Diogo Almeida's TypeSafe AI released Jev in early access, a structured-decision model priced at $0.042 per million input tokens.

Introduction

TypeSafe AI, a startup founded by former OpenAI researcher Diogo Almeida, introduced its first "System One Model" on September 15, 2026, in a blog post titled "Introducing System One Models and Jev." The post is explicit about what Jev is not: "Jev is neither small nor an LLM." TypeSafe built it instead for fast, structured decisions that software can act on directly, rather than open-ended text generation, after two years in stealth. Almeida writes that at OpenAI he helped build the methods that made language models useful at following instructions, work that became the research behind ChatGPT. TechCrunch reported on September 18 that he left OpenAI two years ago, and that demand was high enough that the company briefly lost the ability to serve users from its API.

Feature Overview

A new stack built around structured output. TypeSafe says Jev combines a new model architecture, a parallel sampler, and a training method it calls Reinforcement Learning for Calibrated Decisions (RLCD). Rather than producing strings token by token, Jev generates all of its possible outputs in parallel within a single query, and every answer must match a predefined schema of type-safe structured values. TypeSafe's stated goal is intelligence comparable to existing LLMs on these "System One" tasks, achieved at what the company describes as two orders of magnitude greater speed and efficiency. The naming reflects that framing: "System One" comes from Daniel Kahneman's fast-thinking concept in Thinking, Fast and Slow, and Jev is named after economist William Stanley Jevons, whose Jevons paradox describes efficiency gains increasing overall resource use.

Calibrated confidence on every answer. Each response ships with a calibrated probability score, which TypeSafe positions as a way for developers to programmatically decide when to trust an automated decision. The system supports up to 255 possible choices per decision; beyond that, TypeSafe uses a two-stage score-then-choose process.

Pricing and latency. Input tokens cost $0.042 per million ($42 per billion), and TypeSafe does not charge for output tokens. The company reports 70 to 500 millisecond end-to-end response times, and says this translates to 40x to 200x faster responses than general-purpose LLMs on System One-shaped queries.

Headline benchmark claims, with caveats attached by TypeSafe itself. The company's own workflow evaluations report Jev running "193.6x faster, 444.6x cheaper" than the models it was compared against. TypeSafe discloses that these workflows were built by its own model-capabilities team, used the average of GPT-6 Astra and Fable 5.1 predictions as the reference answer, and represent "the higher end of real world gains" rather than a typical result; the company also states "some bias could exist" given that its own staff designed the comparison. TypeSafe separately chose not to publish Jev's results on public benchmarks.

Synthetic training data. TypeSafe says all of Jev's training data is synthetic and produced in-house, and that it does not train on customer data.

Usability Analysis

Jev is available only in early access to developers pulled off a waitlist, and the two public demos TypeSafe has shown -- a bot playing Doom at roughly 10 queries per second for about $7 an hour, and a Wikipedia-navigating "Wikiracing" agent whose steps can involve hundreds to thousands of links, more than the 255-choice limit, so TypeSafe falls back to its two-stage process -- both operate on structured text state rather than images. Early third-party tests reported by TechCrunch suggest real gains on narrow, classification-style tasks: Vercel engineer Pranit Sharma said swapping an OpenAI-based classifier for Jev produced results five to 18 times faster with greater accuracy, and Bryo AI's Nikhil Mudholkar found Gemini slightly more accurate but 10 to 20 times more expensive than Jev in an email-classification test. Earendil's Armin Ronacher pointed out a tradeoff built into the design: because Jev hands back a probability rather than a definitive answer, developers still have to decide what a 50% versus a 95% confidence score should mean for their workflow, and he suggested model routing -- using Jev's speed and cost to decide when to escalate to a larger model -- as a practical use case.

Pros and Cons

The type-safe output format removes an entire category of parsing errors, since Jev's possible answers are defined in advance rather than generated as free-form text that downstream code must interpret. A calibrated confidence score accompanies every answer, giving developers a probability they can use directly to decide when to trust an automated decision versus escalate it. Early developer reports cited by TechCrunch describe meaningful speed and cost advantages over general-purpose LLMs for narrow, classification-style tasks, including the Vercel test showing responses five to 18 times faster with greater accuracy. Parallel generation of all possible outputs in a single query, rather than token-by-token decoding, is what TypeSafe credits for its claimed 40x to 200x speed advantage on System One-shaped queries.

The headline efficiency figures come from TypeSafe's own workflow evaluations, built by its own model-capabilities team and measured against the average of two other companies' models, and the company itself flags them as an upper bound rather than a typical result. An unverifiable pricing claim underlies the cost comparison, since TypeSafe itself acknowledges it cannot prove its pricing is not subsidized. TypeSafe declined to publish Jev's performance on public benchmarks and has not disclosed its model architecture, leaving outside observers cited by TechCrunch to suspect it may build on an existing open-weight LLM. Factual accuracy is not what the zero-hallucination claim covers, since it addresses only schema-level type matching, and TypeSafe describes the figure itself as not empirical.

Outlook

TypeSafe says it plans to release more versions of System One Models in new modalities beyond the structured text state its current demos use, which would extend the approach toward tasks involving images or other inputs. TechCrunch notes that outside observers suspect Jev's undisclosed architecture is built on top of an open-weight LLM, while TypeSafe's FAQ says only that Jev is "neither small nor an LLM" without describing the design; how the company responds to that scrutiny, and whether it eventually publishes results on public benchmarks or third-party evaluations beyond its own workflow tests, will shape whether Jev's efficiency claims hold up outside the vendor's own comparisons. Broader adoption will likely depend on how well the 255-choice cardinality limit and text-only demos generalize to the messier, higher-cardinality decisions many production systems actually need.

Conclusion

Jev is TypeSafe's attempt to carve out a model category distinct from LLMs: fast, calibrated, structured-decision outputs rather than open-ended generation, backed by early third-party reports of meaningful speed and cost advantages on narrow tasks. Its most significant caveats -- benchmark figures built and evaluated internally, an undisclosed architecture, and an early-access-only, waitlist-gated release -- mean the strongest claims remain largely vendor-reported for now. It is most relevant to developers building high-volume classification, routing, or decision workflows who can tolerate an early-access product and are willing to validate Jev's calibrated probabilities against their own data before relying on them in production.

Editor's Verdict

TypeSafe AI Launches Jev, a New System One Model is a workable proposition that fills a clear gap, even if it doesn't fundamentally change the landscape.

The strongest case for paying attention: the type-safe output format removes an entire category of parsing errors, since Jev's possible answers are defined in advance rather than generated as free-form text that downstream code must interpret. That alone raises the bar for what readers should expect in this space. Reinforcing that, a calibrated confidence score accompanies every answer, giving developers a probability they can use directly to decide when to trust an automated decision versus escalate it — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the new model class, called System One Models, is built to make fast structured decisions that software can call directly, trading open-ended text generation for predefined, type-safe output values. On the other side of the ledger, one constraint is real rather than a marketing footnote: the headline efficiency figures come from TypeSafe's own workflow evaluations, built by its own model-capabilities team and measured against the average of two other companies' models, and the company itself flags them as an upper bound rather than a typical result. It should factor into any serious decision. Layered on top of that, an unverifiable pricing claim underlies the cost comparison, since TypeSafe itself acknowledges it cannot prove its pricing is not subsidized — which narrows the set of teams for whom this is an obvious yes.

For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, the smart move is to track its trajectory and revisit once the rough edges are filed down. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • The type-safe output format removes an entire category of parsing errors, since Jev's possible answers are defined in advance rather than generated as free-form text that downstream code must interpret.
  • A calibrated confidence score accompanies every answer, giving developers a probability they can use directly to decide when to trust an automated decision versus escalate it.
  • Early developer reports cited by TechCrunch describe meaningful speed and cost advantages over general-purpose LLMs for narrow, classification-style tasks, including a Vercel test showing responses five to 18 times faster with greater accuracy.
  • Parallel generation of all possible outputs in a single query, rather than token-by-token decoding, is what TypeSafe credits for its claimed 40x to 200x speed advantage on System One-shaped queries.

Cons

  • The headline efficiency figures come from TypeSafe's own workflow evaluations, built by its own model-capabilities team and measured against the average of two other companies' models, and the company itself flags them as an upper bound rather than a typical result.
  • An unverifiable pricing claim underlies the cost comparison, since TypeSafe itself acknowledges it cannot prove its pricing is not subsidized.
  • TypeSafe declined to publish Jev's performance on public benchmarks and has not disclosed its model architecture, leaving outside observers cited by TechCrunch to suspect it may build on an existing open-weight LLM.
  • Factual accuracy is not what the zero-hallucination claim covers, since it addresses only schema-level type matching, and TypeSafe describes the figure itself as not empirical.
Advertisement

Comments0

Key Features

1. Structured, type-safe outputs instead of free-form text, matched to a predefined schema. 2. Calibrated confidence score attached to every answer for up to 255 possible choices. 3. Parallel generation of all outputs in a single query rather than token-by-token decoding. 4. Pricing of $0.042 per million input tokens with free output tokens, and 70-500ms response times. 5. Training exclusively on in-house synthetic data, with no training on customer data.

Key Insights

  • The new model class, called System One Models, is built to make fast structured decisions that software can call directly, trading open-ended text generation for predefined, type-safe output values.
  • TypeSafe's founder Diogo Almeida named the category after Daniel Kahneman's 'System One' fast-thinking concept and named Jev itself after economist William Stanley Jevons, whose Jevons paradox describes efficiency gains increasing overall resource use.
  • Jev pairs a new model architecture with a parallel sampler and a training method TypeSafe calls Reinforcement Learning for Calibrated Decisions, generating every possible answer in a single query instead of token by token.
  • TypeSafe prices Jev at $0.042 per million input tokens with free output tokens, and reports 70 to 500 millisecond end-to-end response times.
  • TypeSafe's own workflow evaluations report Jev running 193.6 times faster and 444.6 times cheaper than the compared models, but the company itself calls these figures 'on the higher end of real world gains' and built the reference workflows internally.
  • TypeSafe says Jev cannot produce a malformed output because its schema matching is guaranteed, though the company describes its 0% hallucination figure as not empirical and deliberately declined to publish results on public benchmarks.
  • TechCrunch reports that Almeida is tight-lipped about Jev's architecture and that outside observers suspect it is built on top of an open-weight LLM; TypeSafe's own FAQ says only that Jev is "neither small nor an LLM."
  • TechCrunch cites early developer tests: a Vercel engineer reported Jev replacing an OpenAI-based classifier ran five to 18 times faster with greater accuracy, while a Bryo AI executive found Gemini slightly more accurate but 10 to 20 times more expensive in an email-classification comparison.

Was this review helpful?

Share

Twitter/X
Advertisement