

The Prompt Was Three Sentences. The Build Was a Governed Claims Data Product.
This is post 3 of 3 in the Building with Maia & Jev series, and the final one. Posts 1 and 2 cover the shared connector and the five worked examples this build relies on.
The demo that convinces nobody is the one where the input is clean.
Every agentic data engineering walkthrough you've seen, including the two I published before this one, starts from a table someone created for the purpose. Sample tickets. Three research articles. Synthetic reviews with tidy columns. The logic works, the output is correct, and a practitioner watching it thinks the same thing I'd think: yes, but my data doesn't arrive like that.
Mine arrives as an S3 bucket with four file formats in it. Emails with headers and quoted replies. Call transcripts that machine transcription has mangled. Portal submissions where the entire claim sits in one free-text field. Scanned intake sheets where OCR has turned a policy number's 1 into a capital I.
So this time I gave Maia the bucket, a short prompt, and nothing else — no guidance on how to handle any individual format — and had it build the thing properly: ingestion and cleaning into bronze, Jev decisions in silver, a star schema in gold, with tests, documentation, and lineage. It took about 90 minutes and processed 440 claims.

What I Actually Asked For
The prompt was deliberately thin, because the interesting question is how much design the agent does rather than how well it follows instructions.
Two phrases in there did most of the work. "Metadata driven and not just a single feature build" is what stopped Maia writing four parallel ingestion pipelines. And "consult with me as you need" is what got me a plan to argue with rather than a finished build to unpick.
I set balanced permissions rather than full autonomy — I wanted to be asked about anything consequential — and put it in plan mode to buy thinking time.
The Plan, and the Two Corrections It Needed
Maia came back with a bronze, silver, and gold architecture, plus something I didn't expect and now look for: a list of risks and assumptions, with questions attached. I answered those, and then largely left it to work.
I intervened twice.
The first was modelling. Maia had centered the gold layer on the claimant. That's a defensible read of claims data and it's wrong for this purpose — the grain should be the claim, because a claimant can file many and the decisions we're making are per-claim. One sentence of feedback, and it restructured.
The second was testing. I asked for the pipelines to be tested properly rather than built and declared done.
Both corrections are the kind a data engineer makes in a design review, which is roughly what the exchange felt like. Neither required me to touch a pipeline.

Bronze: A File Registry Before Any Ingest
This is the design decision I'd have hoped for and wouldn't have bet on.
Rather than loading files directly, Maia built a metadata registry first — it catalogs every file in the bucket into a Snowflake table, recording format type and file pattern, and then passes those down into a single ingestion engine.
That inversion is what makes the thing extensible. A new format doesn't need a new pipeline; it needs a row in the registry and a branch in the engine. When I said "metadata driven, not a single feature build", this is what I meant, and I didn't have to explain it twice.
Inside the ingestion engine, each format takes its own path. PDFs go through the AWS Textract components, which turn them into semi-structured output stored as a VARIANT column in Snowflake. JSON goes through the S3 load component into a JSON staging layer. Text and email files use the S3 load components too, with different handling for each.
I gave no guidance on any of that. Maia selected those components from the Maia Foundation palette itself, which is the part worth sitting with — the component library is the reason the agent can make reasonable choices without being told. Textract is one option among several; the platform also supports Azure and Snowflake document services, a range of vector databases, and model endpoints including Vertex and OpenAI where a pipeline needs them.

Silver: Questions as Data, Not as Code
The silver layer is where Jev does its work, and the structural choice here is the one I'd most want another practitioner to copy.
The questions put to Jev are not written into a pipeline. They're seeded into a Snowflake table, and the pipeline loops that table through the shared Jev component I built in the first post — the one where the model, the state text, and the questions are all parameters.
So the unstructured content comes in from bronze, gets married up with whichever questions apply, and the same component answers all of them regardless of whether a given question is a choice, a score, or a noul.
The consequence: changing what you ask about a claim is a data change, not a code change. Add a row, and the next run evaluates it. That's the difference between a pipeline that embeds a decision and a pipeline that embeds a decisioning capability, and it's why the shared component was worth building separately before any of this existed.


Gold: A Star Schema With Confidence Attached
Maia joined Jev's answers back to the raw data and built a facts and dimensions model with the claim as the fact grain.
Sampling the output, you can follow a claim left to right: its ID, the IDs you'd use to match it against a central claims system, and then the evaluated fields — claim category, liability party, routing decision, fraud risk, severity, and complexity. For one claim, the category came back cargo damage.
Every one of those carries a score and a confidence value. That matters more than the classification itself. A routing decision at high confidence can be actioned; the same decision at 0.4 belongs in a review queue. Because confidence is a column, that policy is a WHERE clause rather than an architecture.
The gold layer then derives a prioritization score from fraud risk, severity, and approval risk — and I can override it, which is the right default for anything feeding a human workflow.
One column I didn't ask for and will keep: total input and output tokens consumed per claim. Cost observability in the fact table. Most teams discover their per-unit inference cost from a billing surprise; here it's queryable alongside the thing it priced.
Maia also generated architecture diagrams, which I'd have skipped and shouldn't have — they're how the next person understands this without reading nine pipelines.

What the Decisions Cost
Six questions across 440 claims is 2,640 decisions.
Measured on my own TypeSafe account earlier in this series, a Jev decision on this workload costs $0.0000275 — Jev is priced at $0.042 per million input tokens with output unmetered. Carrying that rate forward, the evaluation for the entire run comes to roughly $0.07. About 16 hundredths of a cent per claim.
That projection is derived from my earlier measurement rather than read off the console for this run, so treat it as an estimate — and note that you don't have to, because the per-claim token counts are sitting in the fact table. Query them and you have the real number.
For scale, the same 2,640 decisions on list prices elsewhere: about $0.44 on GPT-4o mini, $1.28 on Gemini 3.5 Flash-Lite, $3.26 on Claude Haiku 4.5. Small sums at 440 claims. At a realistic daily claims volume they stop being small, which is the whole argument for putting a typed decision model in the silver layer rather than a general one.
The Part I Didn't Ask For Twice
I asked for tests once, in the original prompt, and then had to ask again.
What I got in the end was a test pipeline for each stage — including one that asserts the bronze layer registers file metadata correctly, which is exactly where this design would fail silently if it failed at all. A registry that quietly misses a format doesn't error; it just processes fewer files than you think.
Maia wrote those tests, ran them, and documented every pipeline. I'd still rather it had treated testing as part of "finished" without the second prompt, and that's a fair criticism to make of my own prompt as much as of the agent: "test all of the pipelines" sat at the end of a long instruction, which is the worst place to put a requirement you care about.
Promotion, Lineage, and What Happens When It Breaks
A build that lives in one branch isn't a data product. The rest of the lifecycle is the part demos usually skip.
From my local branch I committed the changes, pushed to the remote repository, published to a QA environment with a named artifact version, and scheduled it. Hourly works; sub-five-minute works if you want new files picked up close to arrival.
Then the bit I'd actually use day to day. Searching the catalog for claims brings up column-level lineage — I can trace the liability confidence column back through the transformation pipeline that produced it to the ingest that fed it. Alongside that is pipeline observability: what ran, when, what succeeded.
When I opened a failed pipeline, it showed the downstream dataset affected, and offered root cause analysis as an agentic task. That's the honest end state of this kind of build. Pipelines fail. The question is whether the platform tells you what broke, what it broke downstream, and proposes the fix.
Where I'd Push Back on My Own Build
Three things I'd want fixed before this carried real claims.
There was no central claims database to match against. The gold layer has the matching IDs but nothing to match them to, so the reconciliation step is unproven. Worth noting that fuzzy matching is itself a decision problem — if you receive email and text dumps with no reliable key, Jev is a reasonable tool for resolving them to a claim, and that's a build I haven't done.
I set no confidence threshold. Every answer was written, whatever its confidence. The platform gives me the column and I didn't use it. Independent testing of Jev has found out-of-scope inputs classified confidently into whatever categories were on offer, so a claim type the taxonomy doesn't cover would come back looking decisive. A "none of these" option plus a review route is the fix, and it isn't optional in production.
And validity is the claim I'd scrutinise hardest. The prompt asked whether a claim is valid, with rules for how to reason about it. Jev will answer that question every time. Whether its answer is good enough to act on is a question about your data and your labels, not about the model's architecture, and the only way to know is to score a run against a labelled set. I have the answer key for this corpus. That evaluation is the next thing I'd publish, and I'd be cautious of anyone presenting a pipeline like this without it.
What This Series Was Actually About
Three posts, and the pipeline was never the point.
The first built a connector and shared components. The second put five decisions through them. This one handed over a messy bucket and a short prompt and got back a tested, documented, scheduled data product with lineage.
What changed between post one and post three isn't Maia's capability. It's that the reusable pieces existed — a shared Jev component, a skill describing how to use it, questions held as data. The 90 minutes only looks impressive because of the afternoon that came before it.
Which is the ordinary truth about agentic data engineering, and the part that doesn't demo well: the leverage comes from the components your team already shares. Build those, and the next data product is a conversation.
This was post 3 of 3 in the Building with Maia & Jev series, and the final one. If you missed the start, post 1 is Building a TypeSafe Jev Connector and Shared Pipelines in Maia.
See how Maia builds data products like this.

Related Resources
Data management



