AI Agents are CPUs. Automation needs a GPU.

October 1, 2026

Documents pass through a decision model; uncertain decisions go to a person.

AI Agents work like a processor: one instruction at a time, with you as the interrupt. The work left to automate in a business needs something else: one pass over all the data, a confidence on every answer, and questions only where the model is unsure.

An agent and a decision model work through 120 tasks. The diagrams are illustrative; counts and thresholds are examples. Open animation.

Ask an agent to change how payroll rounds overtime and you can watch it work. It reads a folder, greps for “overtime,” opens a document, edits it, then stops to ask whether the state rule or the federal rule applies. You answer twenty minutes later because you were in a meeting. It reads two more files, edits, runs the tests, reads the failures and asks you something else.

That loop is impressive, and I use it all day. But it isn't how most work in a business will get automated. The work that's left is in the gray area: is this procedure covered, does this pay run comply, which bucket does this claim go in. It isn't rules and it isn't creative. It needs a fast answer, a confidence you can act on and the line the answer came from. A loop that fetches one file at a time and stops for every doubt is the wrong machine for it.

An agent is a CPU

You can think of this agent loop as two queues. The retrieval queue decides what to read next: a file, a search, a page of docs. The action queue decides what to write from what it read: an edit, a command, a test run. Each step depends on the result of the one before, so the steps run one after another.

A classic CPU works the same way. Fetch an instruction, decode it, execute it, then the next. It waits on memory because loads are slow. It can't run far ahead because a branch depends on an earlier result. When it needs something from outside, an interrupt stops the pipeline.

In the agent, you are the interrupt.

The agent reads, writes and waits on you, one step at a time. Open animation.

Look at where the time goes. The reading and writing take minutes. The two questions take most of the half hour, because the person who knew the answers was doing something else. That isn't a speed problem you fix with faster tokens. It's the shape of the work.

We're already faking the GPU

Two habits that spread among people who use agents every day give away what they want instead.

The first is grill-me, a skill that makes the agent interview you about a plan before it writes anything. It works because it moves every question to the front. Answer them once and the agent can run without stopping. That's batching: gather the inputs before the pass instead of handing them over one interrupt at a time.

The second is subagents. A workflow fans the task out to five agents, each with its own slice. Those are threads. You get five cores, but each one still runs the same read, think, write loop and still blocks when it needs something another thread has. The run finishes when the slowest one does.

Grill-me asks everything up front, then subagents split the work. Open animation.

Both help, and I use both. They're workarounds. Neither changes the machine.

Graphics had this exact problem

The first OpenGL drew geometry in immediate mode. Between glBegin and glEnd you called glVertex once per vertex, every frame, from the CPU. A mesh with a million vertices was a million function calls before anything reached the screen.

Vertex buffers (OpenGL 1.5, 2003) and programmable shaders (OpenGL 2.0, 2004) changed the contract. You upload the mesh into GPU memory once. You bind a small program that says what to do with each vertex and each pixel. Then you issue one draw call, and the GPU runs that program over everything in parallel. Immediate mode was deprecated in OpenGL 3.0 in 2008.

Comparison between immediate mode & buffers/shaders. Open animation.

What changed wasn't only speed. The CPU stopped describing the work step by step. It handed over the data and a small, specialized program, and asked for everything at once.

The GPU model of work

Now do the same thing with a decision at work:

  • Teach it one job. Take a small model and train it on the decisions your team already makes, the way a new hire learns by watching the person who's done it for years. It doesn't need to know everything. It needs to know this.
  • Give it everything at once. Hand over the whole file up front: the claim, the policy, the history. Then it doesn't have to go back for one more page every few minutes.
  • Make it pick from a list. Don't ask for an essay. Ask it to choose one answer, like covered, not covered or needs approval, and to say how sure it is.
  • Only ask when it's unsure. If it's sure, the answer goes through. If it isn't, a person gets one clear question to settle it.
Every claim goes in at once and every answer comes out with a confidence level. Open animation.

That last part is grill-me built into the model. The model doesn't ask because its loop needs input to continue. It asks because, for that record, it is genuinely unsure, and it can say so with a number.

Decision models already work this way. TypeSafe’s Jev turns text and application state into typed judgments with probabilities. OpenAI announced a Decisions API at DevDay on September 29: you supply context and questions with finite, predefined answers for classification, routing or choosing an agent’s next action. It launched in limited preview. The shape is converging: a question, a set of answers, one call.

Every answer makes it better

Every question the model asks comes back with an answer from someone who knows. We add those answers to the training set and retrain, so the next version already knows what the last one had to ask. The goal is for fewer records to turn into questions as the model improves.

Answered questions become training examples for the next version. Open animation.

This is how the model remembers. An agent keeps its memory in notes it has to read back before every task, and anything that didn't make it into the notes is gone. Here the answers go into the weights themselves, the skeleton the whole model is built on. The model doesn't look anything up. It just knows.

A spicy thought about code

Writing software is where I currently work with agents. Watch one write a function and what you see is typing: read a file, write a line, run a test, ask.

Now picture describing the function and getting it back whole, in one pass, with the two decisions you didn't specify raised as questions and a confidence beside each.

Comparison between a traditional coding agent and decision model. Open animation.

That isn't how code gets written today. But I think the line-by-line loop is the immediate mode of software, and it won't be the last way we write programs.

Confidence and provenance are the product

For a business, the answer alone isn't worth much. A payroll lead can't approve a number she can't trace. An insurance coordinator can't tell a patient their crown is covered because a model sounded sure.

Two things make an answer usable. The confidence has to be calibrated: when the model says 0.95, it should be right about 95 times in 100, so the threshold where a person steps in means something. And every value needs its source, the page and the line it was read from, so the person who approves it can check it in seconds.

Every value points back to the line it was read from. Open animation.

A general model in an agent loop gives you neither by default. It gives you text, and it sounds just as sure when it's guessing. A small model trained on one job, scored against answers your operators agreed on, can give you both.

As sure as the people who do the work

Ask an operator how they decide and they'll tell you the rules. What they can't tell you is everything they actually do: the payer they always double-check, the code they know gets denied, the exception they make without thinking about it. Those decisions never get written down, but they're in the data. Every past case records what was decided.

The model learns from thousands of those cases, so it picks up the implicit decisions along with the written ones.

Implicit decisions cover part of the work. Open animation.

That's the bar: the model is held to the standard of the people doing the work today. Where it can't meet that standard, it doesn't guess. It asks one of them and learns.

How we build it

This is how we automate work for our clients. We fine-tune small open-source models, one per workflow: payroll, compliance, insurance verification. For a dental practice, that's reading a patient's plan from the payer's portal and filling the benefits breakdown (maximums, deductibles, frequencies, coverage percentages) with every value pinned to where it was read.

Every one of these models runs the same way. Load the record and its documents, run one pass, keep the confident answers with their sources, and send a person the questions that are left.

The model is the fast part

Training isn't what takes time. Two things before it do.

The first is access. The data lives in payer portals, shared inboxes, scanned PDFs, spreadsheets and an ERP somebody configured years ago. Getting it out in a usable form, with permission, for every record, is most of the integration work.

The second is the golden set. To know whether a model is good, you need answers you trust to score it against, and the only people who can give you those are the operators who do the work today. We sit with them, go through real cases and agree on the right answer for each, including the ones where they disagree with each other at first. It's slow, manual work, and it's the part that decides whether the model is any good.

Data access and operator knowledge are what makes or breaks a model. Open animation.

Once both exist, training is the short step, and the model runs over every record in one pass.

What's next

We're about to announce the first set of models built this way, with the evaluation behind each one: the golden sets, the scores and where they fall short. We'll publish the numbers before making claims about them.

If you have a workflow in that gray area, the one a person checks by hand because no rule covers it, bring it. Mulholland Technologies will tell you what data it needs and who would have to grade it.

By Marijn Brussel. Originally published on LinkedIn on October 1, 2026.