TypeSafe Jev is a System One classification model that dramatically improves LLM integration by outputting structured probability scores instead of generating text tokens. This eliminates hallucinations and JSON parsing errors, allowing backend applications to reliably evaluate complex data and route workflows with predictable, type-safe decision logic.
Last week, a revolutionary new model became available for beta testing in the world of AI & Machine Learning: Jev by TypeSafe AI, created by Diogo Almeida (one of the inventors of ChatGPT and RLHF).
The problem it solves comes down to this:
Large Language Models (LLMs) are fundamentally chatbots. They are fantastic at having natural conversations, but when you try to integrate them into application logic, things get messy fast.
Using LLMs to guide a software workflow or make decisions can be a real headache. You practically have to beg the model to give you exactly what your code needs, in a strict format that the application can read and process.
Sometimes it just completely ignores the details of your request. It might add conversational fluff around the data or confidently make up an answer just to give you a response, leaving your code to deal with the fallout. It's often a hit or miss unless you pay a fortune for the most advanced models out there (and even then, there are no guarantees).
System One models, like Jev, solve this by not responding with text at all. Instead, Jev expects a highly specific data format containing the data to analyze and a list of questions about it with a list of possible answers for each question. In return, it gives you a structured response that consists of probabilities for each predefined answer to your questions. Each of these probabilities represents the likelihood that the question can be answered with yes/no or another predefined value. This way, your code can decide for itself how reliable the response actually is.
This is the JSON object you would send to Jev to decide which department should handle a support ticket:
{
"model": "jev-latest",
"state": "My screen is cracked, how much to replace it?",
"questions": {
"department": {
"type": "choice",
"instructions": "Which department should handle this ticket?",
"criteria": {
"hardware_repair": "Physical damage to devices",
"software_support": "App crashes and OS issues",
"billing": "Invoice and payment inquiries"
}
}
}
}
The response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "hardware_repair",
"confidence": 1.0,
"probabilities": {
"software_support": 0.0,
"billing": 0.0,
"hardware_repair": 1.0
}
}
},
"usage": {
"input_tokens": 343,
"output_tokens": 43
}
}
As a bonus, using Jev is much faster and cheaper than using generative LLMs for the same job. This is what makes Jev a real game changer!
How much does TypeSafe Jev cost?
With Jev, you only pay for the input tokens (the data you send it to analyze). The output tokens (the answers) are completely free. On top of that, the current price for those input tokens is an incredible 238 times cheaper than with Claude Fable 5.1!
According to TypeSafe AI, this isn't a bait-and-switch tactic to get users on board - the current pricing is already profitable for them.
Speed
Because Jev is essentially a classifier model rather than a generative one, it uses parallel processing to process your data much faster than traditional LLMs. And since it returns significantly less data without generating long text responses, the total processing time is much shorter.
Is TypeSafe Jev reliable for production applications?
Because the exact structure of the response is determined by you, and every result has a reliable, calibrated probability attached to it, you don't have to deal with hallucinations. You decide the questions and the possible answers, so Jev can't hallucinate any other options. As long as you ask the right questions and let your code take action based on those probability scores, your application will be incredibly stable and reliable.
Availability
For now, there is a waiting list to sign up and create an account.
Luckily, RedhotCoding got early access and tried it out for you.
Putting it to the test: application for checking rental contracts
I wanted to give it a try at once, and came up with the idea of using it to find issues with rental contracts.
The basic idea is as follows: let's say you're building an app to help people check their rental contracts for any potentially problematic clauses added by their landlord or agency.
- First, you let the user upload an image or a digital version of the rental contract.
- In the backend, the contract is converted to raw text and cleaned up.
- The application adds whatever information is needed to pass it to an AI system for processing.
- The AI system returns the needed data to be processed by the application code
The "traditional" approach would be to ask a generative LLM to go through the document and point out any issues. This would probably work okay, but:
- it would be slow and expensive
- you'd have to ask the LLM to return structured data, so it can be processed by your application code. It would not necessarily adhere to that.
- there would very likely be hallucinations from the LLM, so it would not always be reliable
You could solve some of those issues by using distillation:
- first use a frontier LLM to create a large dataset of contracts where the issues are being pointed out
- programmatically verify the structured data output of the LLM
- optionally let humans verify the data for accuracy
- use the dataset to fine-tune a smaller, cheaper and faster model to do the same job (If you go this route, consider using a tool like ConvoManager to easily manage your AI datasets)
This would still not be foolproof, and would require a significant financial and time investment.
This is where the promise of Jev sounds incredibly appealing: backend processing for such an application could theoretically be handled by Jev very easily and reliably.
Let's walk through the workflow:
- First, you let the user upload an image or a digital version of the rental contract.
- In the backend, the contract is converted to raw text and cleaned up.
-
The application asks Jev questions about the contract, for example:
"Does the contract contain any hidden administrative fees or excessive penalties for late payments?"
"Does the contract allow the landlord to withhold the security deposit at their sole discretion without providing an itemized list of deductions?"
It also provides the options and criteria for the possible answers. - Jev provides probabilities for all answers to the questions and a confidence score for the most likely answer.
It cannot hallucinate in the sense that it cannot come up with answers that you haven't provided options for. It will always select one of the options you've given it to choose from.
It also can't ignore any of them because it must give a probability for each question and answer. - The application lists all the potential issues (for example, the ones with a confidence score above 70%) and shows the confidence score as a percentage. These issues can then be color-coded for a better user experience.
Sounds great! But how does Jev actually perform? Can it really live up to the hype?
Let's try it out.
For the first test, I asked my favorite chatbot to create a problematic rental contract:
1. RENT AND TERM:
The monthly rent shall be $2,200, payable on the 1st of each month.
The lease term begins on October 1, 2026, and ends on September 30, 2027.
Upon expiration, this agreement shall automatically renew on a month-to-month basis unless either party provides written notice 60 days prior.
2. FEES AND PENALTIES:
A mandatory non-refundable administration fee of $300 is required upon signing.
If rent is late by more than 3 days, a penalty of $100 per day shall apply.
3. MAINTENANCE AND REPAIRS:
The tenant agrees to keep the premises in good order.
Tenant is strictly responsible for all maintenance and repairs up to $500 per incident, including HVAC system maintenance, plumbing clearing, and appliance repairs.
4. ENTRY BY LANDLORD:
The Landlord reserves the right to enter the premises at any time without prior notice for inspections, maintenance, or showing the property to prospective buyers or tenants.
5. SECURITY DEPOSIT:
A security deposit of $4,400 is required upon signing.
Landlord retains the right to withhold any portion of the deposit for any reason deemed necessary at the sole discretion of the Landlord, with no itemized accounting required upon move-out.
Then I sent this json payload to Jev:
{
"model": "jev-latest",
"state": {
"document_type": "Residential Lease Agreement",
"contract_text": "RESIDENTIAL LEASE AGREEMENT\n\n1. RENT AND TERM:\nThe monthly rent shall be $2,200, payable on the 1st of each month.\nThe lease term begins on October 1, 2026, and ends on September 30, 2027.\nUpon expiration, this agreement shall automatically renew on a month-to-month basis unless either party provides written notice 60 days prior.\n\n2. FEES AND PENALTIES:\nA mandatory non-refundable administration fee of $300 is required upon signing.\nIf rent is late by more than 3 days, a penalty of $100 per day shall apply.\n\n3. MAINTENANCE AND REPAIRS:\nThe tenant agrees to keep the premises in good order.\nTenant is strictly responsible for all maintenance and repairs up to $500 per incident, including HVAC system maintenance, plumbing clearing, and appliance repairs.\n\n4. ENTRY BY LANDLORD:\nThe Landlord reserves the right to enter the premises at any time without prior notice for inspections, maintenance, or showing the property to prospective buyers or tenants.\n\n5. SECURITY DEPOSIT:\nA security deposit of $4,400 is required upon signing.\nLandlord retains the right to withhold any portion of the deposit for any reason deemed necessary at the sole discretion of the Landlord, with no itemized accounting required upon move-out."
},
"questions": {
"auto_renewal": {
"type": "noul",
"instructions": "Does the lease agreement contain a clause stating that the contract will automatically renew upon expiration?"
},
"non_standard_fees": {
"type": "noul",
"instructions": "Does the agreement mandate any non-standard fees?"
},
"late_penalty": {
"type": "noul",
"instructions": "Does the agreement impose a financial penalty if the rent payment is late?"
},
"condition_penalty": {
"type": "noul",
"instructions": "Does the agreement impose a financial penalty for not keeping the house in good condition?"
},
"repair_liability": {
"type": "noul",
"instructions": "Does the contract place financial liability for minor routine repairs, maintenance, or structural/appliance fixes onto the tenant?"
},
"entry_without_notice": {
"type": "noul",
"instructions": "Does the landlord reserve the right to enter the property without providing standard reasonable advance notice?"
},
"deposit_withholding": {
"type": "noul",
"instructions": "Does the contract allow the landlord to withhold the security deposit at their sole discretion or without providing an itemized list of deductions?"
},
"guest_restrictions": {
"type": "noul",
"instructions": "Does the contract impose limits on how long a guest may stay, or require landlord approval for any overnight visitors?"
},
"negligence_waiver": {
"type": "noul",
"instructions": "Does the agreement contain a clause where the tenant is forced to waive their right to sue the landlord for negligence, personal injury, or property damage caused by the landlord's failure to maintain the property?"
},
"arbitrary_rent_increase": {
"type": "noul",
"instructions": "Does the agreement contain a clause allowing the landlord to unilaterally increase the monthly rent amount before the fixed lease term has expired?"
}
}
}
This is what I got back:
{
"model": "jev-1.13.0",
"answers": {
"auto_renewal": {
"type": "noul",
"noul": 0.99
},
"non_standard_fees": {
"type": "noul",
"noul": 0.92
},
"late_penalty": {
"type": "noul",
"noul": 0.99
},
"condition_penalty": {
"type": "noul",
"noul": 0.4
},
"repair_liability": {
"type": "noul",
"noul": 0.97
},
"entry_without_notice": {
"type": "noul",
"noul": 0.98
},
"deposit_withholding": {
"type": "noul",
"noul": 0.99
},
"guest_restrictions": {
"type": "noul",
"noul": 0.03
},
"negligence_waiver": {
"type": "noul",
"noul": 0.05
},
"arbitrary_rent_increase": {
"type": "noul",
"noul": 0.03
}
},
"usage": {
"input_tokens": 862,
"output_tokens": 199
}
}
Let's break down Jev's evaluation and see how accurate these probabilities actually are, question by question:
- auto_renewal (0.99): Spot on. The contract explicitly states it will "automatically renew on a month-to-month basis". Jev is 99% certain.
- non_standard_fees (0.92): The contract mandates a "$300 administration fee". It perfectly fits the definition of a non-standard junk fee. Jev rightly says 'yes' with 92% confidence.
- late_penalty (0.99): Completely accurate. The "$100 per day" late penalty is unambiguous.
-
condition_penalty (0.4): This is a tricky one. The contract says the tenant must "keep the premises in good order", but there is no explicit financial penalty attached just to the condition of the house. Jev correctly leans to 'no', but the 40% score indicates some doubt. I don't like that. We'll have to find a solution for this.
- repair_liability (0.97): Highly accurate. Jev correctly identifies that forcing the tenant to pay "up to $500 per incident" covers minor routine repairs and appliance fixes.
- entry_without_notice (0.98): Perfect match. The contract explicitly says the landlord can enter "without prior notice".
- deposit_withholding (0.99): Perfect match again. The contract states the landlord can withhold the deposit "for any reason" with "no itemized accounting required".
- guest_restrictions (0.03), negligence_waiver (0.06), and arbitrary_rent_increase (0.03): Extremely confident "no" answers. Because the contract text completely omitted these issues, Jev correctly identified that they are absent, giving them all a 6% probability or less.
As you can see, Jev didn't guess or hallucinate. It parsed the legal text and mapped it directly to the exact questions we asked, providing very accurate probabilities for them!
But it's not perfect yet. It needs tweaking.
When we change the question slightly to:
"Does the agreement explictly state that a financial penalty will be imposed for not keeping the house in good condition?"
The Jev API now returns this for the condition_penalty:
"condition_penalty": {
"type": "noul",
"noul": 0.08
}
Great! That's much better!
Removing any ambiguity in the questions clearly helps.
Conclusion:
Thanks to the power of Jev, in combination with AI-assisted coding, this is now a simple application that could be built in no time, whereas it would normally take substantial time and investment.
This is just a simplified example, of course.
In reality, the app would need to take location into account for example, because different places have different standards and laws.
The application would probably need to make a first request to figure out the legal jurisdiction, and then a second request with questions specifically tailored to those local laws and standards.
What is the catch with using Jev?
There are a few disadvantages to using System One models:
1. You must know all options beforehand
Because Jev generates probabilities instead of text, it cannot discover unexpected issues or extract information from the input data. You need to know all the possible issues you want to check in advance and list them explicitly. Jev can then give you a probability for each option. But it cannot tell you anything you didn't anticipate.
In our rental contract validator, this means if the contract contains a completely wild, predatory clause that we haven't specifically defined a question for, Jev is completely blind to it. Our application will simply not be able to indicate it to the user.
This can be solved by combining Jev with a standard generative LLM for open-ended discovery.
You would need multiple API calls to achieve this safely:
- One call to a standard AI model to freely read the contract and extract or summarize any "other unusual clauses" it finds.
- Another call to Jev, dynamically generating questions based on the LLM's summary, to double-check those specific findings and tell you exactly how likely they are to be true.
- If Jev confirms the LLM's findings with high probability, you can confidently show them to the user. If not, you discard the hallucination.
I can't help but feel like this is a band-aid for a fundamental limitation of the model: since the core design of Jev is to evaluate specific hypotheses instead of generating text, it will never handle open-ended discovery on its own.
I can imagine many use cases where you need both high reliability AND open-ended analysis or information extraction.
The proposed architecture of using a generative LLM as a "proposer" and Jev as a "verifier" is likely an area that will see a lot of traction going forward though. The solutions that TypeSafe proposes are not perfect, but they can improve the results substantially.
2. It's not easy to know WHAT DATA is responsible for the answers.
Going back to our rental contract example: Jev is great at flagging problematic clauses, but the application cannot easily tell you where in the contract it found that information.
You can give the user a list of issues and a confidence score, telling them how likely it is that this problem exists in the contract. But you cannot easily make it so that a click on the issue shows the user exactly where the issue can be found in the contract.
This is a side-effect of the first issue - because the model doesn't output text, you can't get exact quotes as references, which can be limiting for certain use cases.
TypeSafe does offer some workarounds for this in their documentation, which can be helpful, but they don't seem like perfect solutions either.
Conclusion
Even though Jev is only TypeSafe's first System One model, and I assume it's still being improved, it is already incredibly promising. It got me really excited about eventually using it in production applications. Of course, a lot more testing is required to be confident about how reliable the model is for specific use cases.
Solutions to the issues I pointed out exist, and even though they're not perfect, they can be used.
Over time, more procedures and architectural patterns as well as code snippets and perhaps even libraries will surely be created that will make it much easier to use these models in production.
After all, the simplicity, speed and low cost of this approach make it incredibly attractive as a potential solution for building reliable and cost-effective AI applications.
I believe this could become an important part of AI for application development and automation (or at least one of the missing pieces of the puzzle).
I will be running more tests and I'll post about the results here soon.

