In this section I review one AI-powered application and demonstrate how it can be used to create new value.
Most of us have got used to thinking about AI in two ways: chat, where you ask and it answers, and agents, where you hand over a longer task. This month a different kind of model got a lot of attention. Jev, from a new company called TypeSafe AI, does not write at all. You give it some text and a set of questions with their possible answers, and it gives back a decision and how confident it is - in less than half a second, and at a very small fraction of the cost of the models most of us use.
How Jev works
The questions can be "which of these options fits best?", "where does this fall on a scale?" or "is this statement true?", and you can ask many of them about the same text at once. TypeSafe calls this a "System One" model, after the quick, gut-level thinking that the psychologist Daniel Kahneman called System 1. Simon Willison, a developer who writes widely about AI models, summarized it as "unstructured state in, typed probabilistic decisions out".
The answers come back as data that software can use directly, with a probability for each option, so there is no text to read or parse. In its launch post, TypeSafe priced Jev at $0.042 per million input tokens, with output free.
One developer, Ryan Vogel, used it to sort 1,700 of his own emails by category, priority, spam and whether they need a reply, for 18 cents in total. TypeSafe was founded by Diogo Almeida, who worked on InstructGPT and ChatGPT at OpenAI, and it is reportedly already in talks to raise money at a valuation above $10 billion, only days after its first round.
Where small judgments become cheap
A model like this makes a small judgment very cheap. Many decisions in an organization were never worth a person's time, or a call to a large model, so they were made on a sample or not made at all. I think the useful question for leaders is where in your organization someone repeatedly reads something, makes a small judgment, and then takes a predictable next step. A few examples:
- Routing incoming support tickets and leads to the right team.
- Checking every document, or every output of your AI agents, against your rules. You could even turn each of your team's tenets into a yes/no question and check the work against it.
- Analyzing a whole archive, such as a year of customer emails or meeting notes, instead of a sample.
- Answering right away in services that are built on a slow queue, like "we'll get back to you with a quote by the end of the day".
Checking work against rules is a good illustration of the trade-off. Every (the publication) planted seven mistakes in twelve passages of text. Jev caught six of them in 0.35 seconds. Claude Fable 5.1 caught all seven, in 8.83 seconds and at about 580 times the cost. When checking is that cheap, you can check every sentence and not only a sample, and send the uncertain cases to a person or a larger model.
A task seems to be a good fit when you can write down the possible answers in advance, when there is a lot of it, when a wrong answer is cheap or easy to catch, and when the evidence fits in a reasonably short text.
My experience so far
I tried Jev myself this week, in a side project: a remake of the classic space game Elite for my Mac. Jev makes the quick decisions in the game - understanding the spoken orders I give the ship's computer, choosing what the other pilots do in a fight, and judging how a negotiation is going - while Claude writes the dialogue. That split, a fast model for the many small decisions and a language model for the words, is also the pattern TypeSafe recommends.
And Jev did not stay alone for long. At its DevDay on September 29, OpenAI announced a Decisions API, a version of its low-cost GPT-6 Luna model that picks one of your predefined answers in about 150 milliseconds. I think that is a sign this kind of fast decision is becoming table stakes.
Limitations
You should be aware of its limitations, though. Jev gives you a number without any reasoning, it reads your question very literally, it is weak with numbers and dates, and early testers have already found bias in its answers. Simon Willison, for example, asked it which cities were good places, and found that it ranked Cupertino at the top and East Palo Alto at the bottom. Hence, it might not be the best choice for hiring or money decisions or recommendations. A good general piece of advice, if you want to try it out and see whether you can trust it, is to label around 50 examples yourself and compare its answers with yours.
Go deeper
If you want to dive deeper, watch the AI Daily Brief episode. Nathaniel Whittemore groups the real use cases into six categories, from analyzing archives and triaging inboxes to checking work against rules. He also covers the limits TypeSafe itself points out, and gives practical tips for writing questions Jev can answer well.
For the builder's view, including the 1,700-email example, Greg Isenberg's episode with Ryan Vogel walks through several demos built on Jev in its first days.
Your action step
Pick one of those repeated small judgments in your organization, like the examples above. Write down the possible answers in advance, then label around 50 real examples yourself. Run the same examples through Jev (or OpenAI's Decisions API, once you have access) and compare its answers with yours. If they match often enough, and a wrong answer is cheap to catch, you have found a judgment you can now make on every item instead of a sample.
If you want to map where small, repeated judgments could create value in your organization, that is the kind of question I take on in AI strategy advisory engagements and in sessions as an AI keynote speaker and workshop facilitator.