The box that should stay empty

September 2026

Rent one of the strongest models, ask it to fill a dental insurance breakdown, and it fills boxes that should stay empty. GPT-6 Luna filled in about half of them in our test. So we trained a small model of our own. On the same 10 test sheets, it left 96.9% of them empty.

It’s called Altadena 1.5, and it’s in training. It isn’t filling in any office’s sheets yet. Here’s the test, and where it still falls short.

The sheet and the pile.

If you run a dental office, you know the breakdown sheet. Before a visit, somebody works out what the patient’s plan covers: coverage percentages, how often, age limits, deductibles and maximums. At a dental group we work with, that sheet has about 290 boxes.

Our overnight check already fills it in. Every night it logs into each insurer’s portal and saves everything it sees. Then rules we wrote by hand for each insurer turn that pile into the sheet. The model is the next step. The check will still grab the data, but the model will read it and fill in the sheet, with no rules per insurer.

Reading is the hard part. For one patient, the portal hands back machine answers, web pages and plan PDFs, and a typical file is a novel’s worth of text. Only about 1 line in 20 holds an answer. And about half the answers aren’t written down directly. The portal says “covered age 14 and over,” and the box wants “>=14.”

The answer key.

How do you teach a model this, and grade it? You need an answer key. We built 127 breakdown sheets, and our AI auditors proved every filled box against the portal data and the plan’s published documents. A box they couldn’t prove stayed blank. That’s 36,813 boxes in all. The model trained on 117 verified sheets, and we held back the other 10 for the test.

The key also sets a floor. Guess the usual answer for each box without reading anything, and you’ll get 59% right. Guess from the same insurer’s other patients, and you’ll get 68%. A model only counts as smart if it clearly beats that.

Three tries.

We tried Jev first, TypeSafe’s fast model that picks answers from a list. It reads about 32,000 word-pieces, or tokens, at a time, and it got lost in long, messy files. It scored about 60% untuned and 77% tuned. Those figures come from smaller test groups, not the 10 sheets below.

Next came OpenAI’s GPT-6 Luna, a big model that thinks before it answers and reads the whole file in one go. We gave it a one-page rulebook of what each box means, and it did well on boxes with an answer. But it also filled in about half the boxes that should stay empty. And it can’t tell you how sure it is.

Then we built our own. It’s a 4-billion-parameter language model, small enough to run on one graphics card. We gave it a small add-on that learns from the answer key, and 290 answer buttons, one per box. Each button can only pick from the sheet’s own list of answers, like “80%,” “2x/ 1yr,” “<=18” or blank. It can’t make one up. Every button also says how sure it is. It reads the whole file once and presses all 290 buttons in one go.

Training ran overnight on one rented graphics card: five rounds of about two hours each, about $26 in total. One round fixed a quiet problem. Some files are huge, and cutting them short during practice had been teaching the model to guess.

The same 10 sheets.

Here are both models on the same 10 held-out sheets, which neither had seen.

Same 10 held-out sheets
Altadena 1.5 GPT-6 Luna
All boxes right, including the ones that should stay empty 93.0% 70.9%
Boxes that should stay empty, left empty 96.9% 48.3%
Boxes with an answer, right 90.4% 86.6%
Boxes with an answer, filled 97.4% 92.1%
Wrong among what it filled 7.3% 5.9%
Says how sure it is Yes No
Time per patient ~6 s ~72 s
Cost per 1,000 patients ~$9 ~$13
Same 10 held-out sheets Altadena 1.5 GPT-6 Luna All boxes right Altadena 1.5, all boxes right: 93.0% 93.0% GPT-6 Luna, all boxes right: 70.9% 70.9% Should stay empty, left empty Altadena 1.5, boxes that should stay empty, left empty: 96.9% 96.9% GPT-6 Luna, boxes that should stay empty, left empty: 48.3% 48.3% Boxes with an answer, right Altadena 1.5, boxes with an answer, right: 90.4% 90.4% GPT-6 Luna, boxes with an answer, right: 86.6% 86.6% 0% 50% 100%
The first three rows of the table, drawn. Longer is better on all three.

Across every box, ours got 93.0% right and GPT-6 Luna 70.9%. But the gap is widest on the boxes that should stay empty. On boxes with an answer, the two are closer: 90.4% against 86.6%. Ours is faster, too, at about 6 seconds a patient against about 72, and it costs about $9 per 1,000 patients against about $13. It beat GPT-6 Luna on 9 of the 10 insurers in the test, and it lost the tenth by about 1 box in 100.

One line runs the other way. Of the boxes each model filled, ours got 7.3% wrong and GPT-6 Luna 5.9%. That’s what the confidence score is for.

You can tell our model to fill a box only when it’s at least 95% sure, and to mark the rest verify. At that bar it fills about 3 boxes in 4, and fewer than 1 in 100 of those are wrong. A stricter bar means fewer mistakes and more boxes left for your team. Our goal is 80% filled with fewer than 1 in 100 wrong. We aren’t there yet. GPT-6 Luna can’t work this way, because it can’t say how sure it is.

What’s still hard.

Four things are still hard. We’ve only tested it on insurers it practiced on, and new portals come next. Copay plans are hard for every model we tried, and each scored about 70% on them.

Some files are longer than it could practice on, because the chip ran out of memory. And the answer key disagrees with itself in a few places, like “12 months” in one place and “1 year” in another. No model can score 100% until we fix that.

What comes next.

First, we’ll map the sheet with that group: every box, every allowed answer and every rule, written down once. Then we’ll save every portal page and plan document, so the data always holds the answer.

Next, we’ll build more than 1,000 verified sheets. One model drafts each sheet and shows the line it read every answer from, and a second fills it in on its own. Where the two disagree, our AI auditors settle it, and people spot-check a random sample. Then we’ll train it for more rounds on a bigger chip with room for whole files, and test it on insurers it has never seen.

Last, we’ll turn it on with verify flags. Until then, the overnight check keeps running on its hand-written rules. Once it’s on, the model will fill in what it’s sure of and flag the rest for the front desk, and every box the desk fixes becomes a new lesson. A box marked verify beats a box filled in wrong.

If your front desk still builds breakdowns by hand, write to us. We’ll go through a blank copy of your sheet with you, box by box.