A good AI workflow can have several kinds of intelligence in it. Ordinary code can calculate a quantity. A small model can check whether an inquiry includes a required detail. A more capable model can help with a difficult interpretation. A person can make the decision that requires professional judgment.
The engineering work is choosing what belongs where—and testing whether each choice earns its cost.
Give Jev a small, clear job
Jev is TypeSafe AI’s decision model. Its interface accepts context and bounded questions, then returns typed decisions or scores. That makes it a candidate for routing an inquiry, checking a missing fact, or assessing the relevance of a document. Writing a customer email is a separate job. See TypeSafe’s introduction.
Consider a flooring inquiry with room measurements, a product preference, and a few site notes. A useful question might be: “Does this request state that a moisture assessment has been completed?” The answer can help an application identify what is missing. It does not establish that a floor is suitable for installation.
This is the kind of bounded task we compare in the Built Correct AI scorecard. The detailed lab keeps the original input, questions, model version, evidence, timing, and usage with each attempt. Our initial Jev integration used jev-1.13.0 on September 23, 2026.
Start with the cheapest approach that does the job
A required date already stored in a form does not need a model to discover it. A calculation should use code. A rigid rule may be enough to handle a predictable phrase. A model becomes useful when people express the same meaning in different ways and a rule misses too much.
The right comparison is the same task, on the same cases, with the same expected answers. Compare a rules baseline, Jev, and a general model at the decision stage. Keep the rest of the workflow fixed. Then inspect the mistakes as well as the bill.
A lower API price is a reason to test a model. It is not proof that the completed work is cheaper. A model that asks unnecessary questions can create more work for the estimator or customer.
Where retrieval helps
Retrieval brings relevant records or documents into a workflow. For flooring, that could mean looking up guidance for the selected product. Access rules, product identity, and source selection come first. A model can then assess a narrow question against that evidence.
Do not add retrieval when the application already has the required fact. Do not assume retrieving a relevant page makes the final decision correct. Test whether the evidence actually changes the outcome in a useful way.
Our public flooring examples include source links and explicit product checks. Code owns quantities, product constraints, and the next-step policy. Jev is one candidate for the language decision in between.
A confident answer still needs an evaluation
TypeSafe’s confidence documentation explains how its confidence signal is derived. A high score is not an independent verification of correctness.
Independent studies show why workflow-specific testing matters. ASSAY-001 found different calibration behavior across two intent datasets. A separate 495-claim RAG verification comparison found that its error types differed between the models tested. Those results describe their experiments; they do not predict performance on a flooring company's inquiries.
For our own use, the useful questions are concrete: Did we miss a required detail? Did we send an unnecessary question? Did a bad decision get through? How much review did a person need to do?
Count the complete outcome
TypeSafe’s September 15 launch announcement listed Jev at $0.042 per million input tokens, with no output-token charge. That is a dated provider rate, not a measured cost for a complete Built Correct workflow.
The complete picture includes the calls actually made, retries, other processing, and human review. We keep usage-based API estimates separate from recorded review effort so a small model bill does not hide a larger operational cost.
Our public cases are engineering examples, not a customer benchmark. They make the comparison inspectable. A business rollout needs representative cases and independent review by people who know the work.
See the AI scorecard · Read how we evaluate cost and quality
Built with perspective.
The Built Correct team.

