![]()
For the past few years, most people have equated AI with chatbots: you ask a question and get a paragraph back. Then, on 15 September 2026, a US start-up called TypeSafe AI released Jev, a model that deliberately cannot chat. Its launch video drew roughly 40 million views on X, and the company announced a US$40 million seed round at the same time. Jev's pitch is simple: on classification, routing and scoring work, it beats typical large language models on speed and cost by one to two orders of magnitude. This article explains what Jev is, how it fundamentally differs from LLMs, which business scenarios suit it, how to evaluate it, and what to watch out for when adopting it.
What is Jev?
Jev is built by TypeSafe AI. Its founder and CEO, Diogo Almeida, spent about four years at OpenAI working on RLHF, InstructGPT, ChatGPT and GPT-4. After leaving, the question he kept asking was this: models have been superhuman at chat for years, so why is business automation still so slow to arrive? His answer was that chat models are simply not designed to be used directly by software. The team spent two years redesigning the model architecture, sampling method and training method, and launched a new category they call "System One models." Jev is the first public model, currently available in early access.
Two ideas behind the name
"System One" comes from psychologist Daniel Kahneman's book Thinking, Fast and Slow: the brain has a fast, intuitive "System 1" and a slow, deliberate "System 2." Most everyday judgments are made by System 1, such as spotting at a glance that an email is spam. Jev is built specifically for this kind of fast judgment. The name "Jev" comes from economist William Stanley Jevons and the "Jevons paradox": when steam engines made coal use far more efficient, total demand for coal went up rather than down. TypeSafe expects AI to follow the same path, with every order-of-magnitude drop in the cost of intelligence unlocking an order of magnitude more use cases.
The fundamental difference from LLMs
A typical LLM outputs text, generated one token at a time, with each token waiting for the one before it. Text is flexible: it can be an answer, code or nonsense. For software to use it, the output has to be parsed and validated first, and there is always a chance it goes off the rails. Jev works completely differently. It does not generate text; instead it returns structured values within an answer space defined in advance, such as "yes or no," "one of A, B or C," or "a score from 0 to 100," and it produces all of its answers in parallel in a single pass. Because the answer format is locked in beforehand, Jev cannot produce type errors, nor can it "hallucinate" a response outside the defined options. TypeSafe describes it as a function call with frontier-level intelligence: messy data in, decisions your code can use directly out.
Confidence scores: the key to real automation
Another important feature is that every Jev answer comes with a calibrated probability and confidence score. It sounds like a technical detail, but it decides whether a task can actually be automated. Suppose a model is right 95% of the time but never tells you which answers fall in the wrong 5%; you would not dare let it run unattended. If instead the model can honestly say "I'm not sure about this one," a company can set a clear rule: handle high-confidence cases automatically and send low-confidence cases to a person. TypeSafe developed a training method it calls Reinforcement Learning for Calibrated Decisions (RLCD) for exactly this purpose, so that higher confidence means higher accuracy.
Speed and cost: how big are the numbers?
According to TypeSafe, Jev's end-to-end response time is about 70 to 500 milliseconds, while frontier LLMs usually take several seconds or longer. On pricing, input costs USD 0.042 per million tokens and output is free; by comparison, typical LLM input ranges from USD 0.20 to USD 10 per million tokens, with output often several times more expensive. On the company's own workflow evaluations, Jev was up to about 190 times faster and roughly 1/440 of the cost of LLMs. Note that these are the company's own figures, and TypeSafe itself says they sit at the higher end of real-world gains. Businesses should treat them as an indication of potential, not a guarantee.
What can a business use Jev for?
Jev is best suited to work with a clear answer space and a high volume of repeated judgments. First, classification and routing, such as sending customer-service queries to the right department automatically or filing documents into the right category. Second, scoring and ranking, such as scoring sales leads or assessing the risk level of applications. Third, content moderation, such as detecting spam, inappropriate content or leaks of sensitive data. Fourth, acting as a gatekeeper for AI agents: when an AI assistant is about to run a command such as resetting a database, Jev can instantly judge whether the action is irreversible or off-task compared with what the user actually asked, then decide whether to allow or block it. Fifth, real-time applications: because responses take around a tenth of a second, it can be used in games and simulations where speed is critical, and TypeSafe has even demonstrated Jev playing the classic game Doom in real time.
What it is not suited for
Understanding Jev's limits matters as much as understanding its strengths. Jev does not write articles or emails, cannot hold open-ended conversations, and is not suited to work that needs long reasoning or creative generation; those remain the strengths of LLMs. Each Jev choice also supports up to about 255 options, so larger option sets need to be handled in stages. It currently works only with text-based data and does not yet support images. Most importantly, Jev is still in early access, so its long-term stability and whether its pricing is sustainable will take time to prove.
How should a business owner evaluate and adopt it?
A practical approach has four steps. First, audit your existing AI and manual workflows and identify which are really "multiple-choice questions," such as classifying, judging or scoring; these are candidates for a decision model. Second, pick one high-volume workflow with relatively clear rules for a small pilot, such as routing customer-service queries, and compare accuracy, speed and cost against your current approach. Third, design a confidence threshold: above a certain level, process automatically; below it, hand over to a person, and spot-check results regularly. Fourth, adopt a "small model plus large model" architecture, where a decision model like Jev handles the high volume of fast judgments and an LLM only handles the parts that genuinely need writing and reasoning. This can cut costs sharply and make the whole system more stable.
Conclusion: AI doesn't have to chat
Jev is a reminder that the value of AI lies not in how much it can say, but in whether it can fit reliably into daily operations. Most business automation needs are large numbers of small, repeated judgments, and handling them with an expensive chat model is like delivering takeaway in a sports car. Choosing the right tool for each job is what a real AI strategy looks like. To learn how to plan the right AI architecture for your company and deploy AI staff into real business workflows, visit ai.ud.hk to explore UD's AI Staff solutions.
懂AI,更懂你|UD相伴,AI不冷