

Jev Use Cases: Five Decisions You Shouldn't Be Paying a Frontier Model to Make
This is post 2 of 3 in the Building with Maia & Jev series. Post 1 covers building the shared Jev connector and pipelines used in the examples below.
There's a particular kind of waste that has crept into data pipelines over the last two years, and most teams haven't costed it yet.
Somewhere in your stack, a large language model is being asked a question with four possible answers. Does this ticket go to billing or technical. Is this review positive or negative. Does this document mention an encryption standard. The model generates a paragraph of reasoning, emits a JSON object, and you parse one field out of it and throw the rest away. You paid for every token of that reasoning, you waited for it, and because the output was generated text rather than a constrained value, you also wrote a validator in case it came back malformed.
That is three problems wearing one coat. The first is cost, and it compounds with volume in a way that pilots never reveal. The second is latency, because generating reasoning you discard is time your pipeline spends waiting. The third is the one that actually causes incidents: a free-text response to a closed-choice question can always come back as something that isn't on the list.
In the previous post I built a shared connector and pipelines for Jev, TypeSafe's typed decision model, inside Maia Foundation. This post is about what I put through it — five worked examples, three question types, and what it cost to run.
The Three Question Types
Jev is not a chat model and it doesn't generate prose. You give it some state — text or JSON — and a set of predefined questions, and it returns typed answers. There are three shapes, and knowing which one a problem is saves a lot of redesign later.
Choice is a categorical pick from a list you define, up to 255 options. It returns the selected option, a probability for each option, and a confidence value. Routing, triage, and classification are all choice problems.
Score is a position on an ordered rubric between 2 and 10 levels. It returns an expectation across the scale rather than a single label, which is more useful than it sounds when the middle of your scale is where the interesting cases sit.
Noul — short for Bernoulli, and easy to mishear as "null" — is a single probability between 0 and 1. It's the shape for "does this document contain X", where you want a degree of belief rather than a hard yes or no.
The reason the distinction matters: the response payload differs between them, which is why the unpack step in the shared pipeline is its own component.

Example 1: Support Ticket Triage
The first pipeline creates sample tickets in Snowflake, loops them through the shared child pipeline with a table iterator, and writes Jev's answers back.
The configuration asks two questions of every ticket. Which team should handle it, with rules I could state in a sentence — payment, subscription, and refund go to billing; bug, error, integration, and data issues go to technical; pricing goes to sales. And what priority it should carry.
A ticket reading Charged twice for subscription, requesting a refund comes back routed to billing with a priority, and I can sample the output in the canvas without leaving the pipeline.

So far this is a classifier, and classifiers are not new. The part worth your attention is what happened next.
Example 2: Changing the Logic Mid-Flight
I decided, while the pipeline was running, that I also wanted a Docs department. Not a new pipeline — a change to the decision logic in the one I had.
So I told Maia: identify where a ticket relates to the documentation team, and update this pipeline to accommodate that.
Maia already had the context. It could see which orchestration pipeline I had open, it could read the YAML, and it understood the routing logic inside it. It went into plan mode on its own, came back with a proposal — add a docs criterion to the department question, update the canvas notes, and here are the components affected — and I accepted it.
Then I asked it to run the pipeline again and show me the results. A new ticket, API docs outdated, references removed, landed in the Docs department. Billing, technical, and sales were untouched.

That is the difference between a model you call and a model your agent can reconfigure. The questions sent to Jev are configuration, and the configuration is something Maia can read, reason about, and change on request — with a plan I approve first. The classifier isn't the interesting part. The fact that adjusting it took one sentence and one approval is.
Example 3: Scoring Content Quality
The second question type, applied to three research articles.
Two questions: how clear and well-written is this content, and how technically detailed and accurate is it. Jev returns a clarity score and a technical depth score, each with its own confidence.
This was a small corpus — a couple of paragraphs each — so treat it as a shape rather than a result. But the shape generalizes to any content pipeline where you need to rank a large volume of unstructured text on a consistent rubric, and where having a confidence value per score tells you which items need a human.

Example 4: Binary Compliance Checks
The third type, applied to policy documents. Does this document specify an encryption standard for data at rest. Does it define a deletion process, a retention process, access controls.
What comes back is not a bare yes or no but a numeric certainty against each question, which is the right output for a compliance workflow. A document at 0.97 on encryption and 0.31 on retention tells you exactly which section a human should read. A hard boolean would have told you nothing about where to look.
The data here is synthetic — I generated it for the walkthrough.

Example 5: Putting All Three Together
Product reviews, using every question type in one pass.
Sentiment as a choice across positive, neutral, negative, and mixed. Satisfaction as a score. Then three nouls: is a defect mentioned, is a refund requested, would they recommend.
R-001 came back positive, strong satisfaction, no defect, no refund, would recommend. R-002 medium-positive, defects not mentioned. R-003 strongly negative, low score, high defect, refund requested.
None of which is surprising, and that's the point — on clear cases you want boring correctness, and you want it cheap. The ambiguous middle is where you'd spend your evaluation effort.
The fifth example ran churn signals over energy customers, with status, payment type, fuel type, and balance as the state, returning low, medium, medium-high, and very high churn indicators. Also synthetic, though it's the one that reads most like a live utility extract, so I'll say so plainly.

What It Actually Costs
Here's where the argument stops being architectural.
Across the examples above, my TypeSafe console recorded 40 requests, 30,821 tokens, and $0.0011 of spend. That's $0.0000275 per decision, or about 36,000 decisions per dollar. Jev is priced at $0.042 per million input tokens, and output tokens are unmetered.
Forty requests proves nothing about production, so I ran the arithmetic forward using my own measured token profile — 655 input and 116 output tokens per decision — against published list prices for three models you might otherwise reach for.
At 10 million decisions a month, the gap between Jev and Haiku 4.5 is about $145,000 a year. Against GPT-4o mini — already a cheap model — it's 6x.
Two honest caveats. Those are list prices at the time of writing and they move. And the comparison assumes the chat models emit a similar number of output tokens, which is generous to them: asked for reasoning, they emit considerably more, and output is where Jev's free tier bites hardest.
The more useful way to read the table is not "Jev versus a frontier model" but "Jev versus the cheapest model that already clears your accuracy bar." That's the comparison worth running on your own data, and it's the one I'd want to see before moving a production workload.
Where Jev Is the Wrong Tool
Jev's schema guarantee is structural — it cannot return a value outside the set you defined, and TypeSafe describes type errors as mathematically impossible. That guarantee covers the shape of the answer and says nothing about whether the answer is right. Worth being precise about, because the two get conflated.
TypeSafe publishes a candid list of weak spots and I'd read it before building anything load-bearing. The ones that matter most in a data pipeline:
It is not a calculator. Counting, arithmetic, and numeric comparison are unreliable. Implement mathematical logic in code, not in a question.
It reads dates as text, not as ordered quantities. Which of two dates comes first, how far apart they are, whether one falls in a window — all unreliable. This one will catch people, because date logic looks like a decision.
It treats state as data, not as hostile. Content written to steer the model can move the answer. If your input is user-supplied, that's a threat model you own.
It always answers. Independent testing has found out-of-scope inputs classified confidently into whichever category was on offer. If your taxonomy doesn't include an explicit "none of these" option, you won't find out when something falls outside it.
And the broader caution: TypeSafe's headline speed and cost multipliers are vendor-run against vendor-authored workloads, and independent reproductions have landed well below them. The price is verifiable and I've verified it. The intelligence comparison is not settled, and anyone telling you it is hasn't looked.
What's Next
Five examples on synthetic data prove the integration works. They don't prove it survives real documents.
The next post takes a corpus of actual inbound claim documents — emails, call transcripts, portal payloads, and scanned PDFs, all arriving in the same S3 bucket in different formats — and has Maia build the full pipeline: ingestion and cleaning into bronze, Jev decisions in silver, and governed data products out the other side. Metadata-driven, so the logic scales to claim types and source formats it hasn't seen yet.
That one has messy inputs, a confidence threshold that actually matters, and an answer key to score against. It's the interesting one.
This was post 2 of 3 in the Building with Maia & Jev series. Next up is post 3: The Prompt Was Three Sentences. The Build Was a Governed Claims Data Product. If you missed the start, post 1 is Building a TypeSafe Jev Connector and Shared Pipelines in Maia.
See how Maia automates decisions like this.

Related Resources
Data management



