Jev: TypeSafe AI's New Model That Judges Instead of Generates

Jev is a new AI model that returns structured judgments instead of generating text, designed for automated decision-making inside software.

Jev: TypeSafe AI's New Model That Judges Instead of Generates

If you follow trends in the AI world, chances are you have already come across Jev, a new AI model by TypeSafe AI. It is trending on X, and once you understand the reason behind it, you will want to try it out for yourself.

TypeSafe AI came out of two years in stealth on September 15, 2026, backed by $40 million in seed funding. Founder Diogo Almeida spent years at OpenAI, where he helped build the research behind ChatGPT. In the company’s own words, models have been superhuman at chat for years, so why has all the automation not caught up? That question is what led Almeida to build TypeSafe: AI infrastructure meant not for conversation, but for decisions inside software.

In this article, we will explain what Jev is, why it matters, and how it is different from ChatGPT, Claude, and other LLMs you already use.

What is Jev?

Jev is not a chatbot. It cannot write emails, poems, or code. In fact, it cannot generate text at all. According to TypeSafe, that is the point.

Models like ChatGPT, Claude, and Gemini generate responses one token at a time. This makes them flexible and conversational, but also slower and inherently open-ended. Jev takes a different approach. It analyzes a situation once and returns a fixed, predefined answer with a confidence score. Instead of generating a response, it makes a judgment.

Because Jev does not generate text sequentially, it can process its output in parallel. This makes it faster, more predictable, and potentially cheaper to run.

You can find more details on the official TypeSafe AI documentation.

How it works?

A regular LLM writes one token at a time, and each new token depends on the one before it. This chain is what makes these models slower when used inside software, and it is unavoidable as long as the output is a sentence.

Jev skips this entirely. It reads the situation once, then answers every question you asked in a single pass, all at the same time.

Let’s take an example of a customer support ticket that needs to be routed and prioritized. A traditional system might rely on keyword rules or manual review. With Jev, a developer can evaluate multiple aspects of the ticket at once, such as which team should handle it and whether it is urgent:

from typesafe_sdk import Choice, Noul, TypeSafeClient

response = TypeSafeClient().system_one(
    state=ticket,
    questions={
        "department": Choice(
            instructions="Which team should handle this ticket?",
            criteria={
                "billing": "Payment issues",
                "technical": "Bugs or integrations",
            },
        ),
        "is_urgent": Noul(
            instructions="Is this ticket time-sensitive?"
        ),
    },
)

response.answers["department"].choice   # "technical"
response.answers["is_urgent"].noul      # 1.0

Jev analyzes the ticket once and returns both results together. There is no need to process each question separately or wait for one response before evaluating the next.

This one call also shows the three types of questions Jev can handle:

  • Choice: Selects one option from a predefined set, such as the appropriate support department
  • Score: Assigns a value to something based on a defined scale
  • Noul: Returns a probability between 0 and 1 for a yes/no question, such as whether a ticket is urgent

These question types can be combined freely in a single request, with results returned simultaneously rather than generated one after another.

One consequence of this is worth pausing on. Adding more questions to a single call barely changes how long it takes to get an answer. Asking Jev one question or thirteen questions takes roughly the same amount of time. TypeSafe’s own cookbook reports that batching thirteen questions into one request came out 12.2 times cheaper and 10 times faster than asking them one at a time, with no change in the answers.

Jev in Action: Examples

Jev has launched with limited access, and since we do not have hands-on access yet, we are showing examples from X to see how it performs across different tasks.

Example 1: Scoring Leads and Outreach at Scale

One builder, Roman, tested Jev on 700 high-intent leads paired with personalized outreach messages. In 40 seconds and for about $0.09 total, Jev predicted how each message would perform, attached a confidence score, and flagged mismatches between leads and messages. The same builder noted Jev can also score leads, read buying signals, match prospects to the best-fit message, and identify which campaigns are likely to perform based on the data.

Example 2: Reading 464,720 Research Papers

One researcher, DevaiahShrithan, ran every arXiv AI abstract from 1993 to 2026 through Jev, asking five questions of each: does it claim state of the art, did it release code, is it written in LLM style, what type of paper is it, and how hyped is the language. The output is a chart of how AI research writing has changed over three decades, built from 2.3 million individual judgments.

Example 3: Triaging 3 Million Session Replay Events

Another builder, Tarasshyn, pointed Jev at 3 million session-replay events. In 40 seconds and for $2.17, it reviewed 3,247 sessions, caught 132 rage clicks, 116 dead clicks, and 95 JavaScript errors, then opened 213 draft fix pull requests. The interesting part is the last step: the judgments were good enough to act on automatically, not just to display on a dashboard.

Jev Evals

TypeSafe backed the launch with some striking numbers. Input costs $0.042 per million tokens, and output is free since there is nothing generated to bill for. At that price, a bot playing Doom and making ten decisions a second runs about $7 an hour.

TypeSafe was upfront that every one of these figures comes from its own evals, using its own workflows and reference answers, with no independent verification yet. That transparency did not stop critics from calling Jev little more than a well-packaged classifier, a category of tool that has existed for years.

What that critique misses is that classification was never the real claim. Calibration was. A calibrated model’s confidence scores mean something specific: the answers it scores at 0.8 are right about 80% of the time. That is what lets code act automatically on high-confidence answers and escalate the rest, which is the difference between a demo and something production-ready.

The most convincing result, though, was not one of the flashy demos. TypeSafe’s re-ranking cookbook takes 40 legal queries, builds a 30-passage shortlist for each using ordinary keyword search, then asks Jev a single question per query-passage pair. That one extra step more than tripled top-1 accuracy.

Jev 2

Like everything else here, the number is TypeSafe’s own. But the underlying pattern still holds: at this price and speed, calling a model ten times a second stops being reckless.

Hallucinations, RLCD, and Where Jev Breaks

TypeSafe’s “no hallucinations” claim needs a caveat. Jev cannot return a value outside the set you defined, so it can never hand your code something unexpected or off-format. However, it can still choose the wrong option from within that set, which means errors are bounded in shape but not eliminated in substance. Developers should treat high-confidence thresholds as a filter rather than a guarantee, and build escalation paths for answers that fall below their required confidence level.

Conclusion

Jev represents a meaningful shift in how AI can be embedded in software pipelines. Rather than generating open-ended text, it returns structured, calibrated judgments at speeds and costs that make real-time automated decision-making practical. Whether it lives up to its benchmarks under independent scrutiny remains to be seen, but the architectural approach — parallel output, bounded answer sets, and confidence scoring — addresses real limitations of current LLMs in production environments. For developers building classification, routing, or triage workflows, Jev is worth watching closely as access opens up.