There's a humanities guy sitting next to you. Hasn't written a line of code in his life — "repository" means something about a library to him. And his AI produces clean, working output. Yours — twenty years of engineering experience, the same model — generates garbage ten times in a row, and you sit there correcting it like a misbehaving intern.
First thought: he got lucky. His tasks are simple. What can he possibly extract from a model when he doesn't understand what's happening under the hood? Comfortable thought. Wrong one.
You hired a junior and you're surprised it acts like a junior
Look at how you work with the model. You explain the task in terms of code, architecture, and libraries — because you can, because that's your native language. The model proposes broadly sound ideas and botches the details. You catch the mistake because it's in your zone of competence, and you correct it. It messes up again. You correct again.
This creates an impression: the AI is a diligent but dim partner you need to supervise. A junior you hired and sat next to. The dialogue runs direct: you ↔ model. Every mistake passes through exactly one checkpoint — your eyes. Everything it does wrong, you see. Everything you see irritates you. That's where the feeling that the model is stupid comes from: you're sitting in the same room with its rough drafts and corrections, watching the entire raw stream of consciousness from first attempt to final.
This isn't a dumb approach. It's the approach of someone who can speak the model's language — and therefore didn't build anything on top. Why bother, when you can explain directly? The luxury of qualification.
The humanities guy doesn't have that luxury — and it saves him
He can't hold a dialogue with the model at the level of architecture and libraries. It wouldn't even occur to him — he doesn't have the vocabulary to correct the model on substance. He is physically incapable of sitting next to a junior and walking through where the interface went wrong.
And here's the turn. The absence of that capability doesn't make him helpless — it pushes him toward a different architecture. Since he can't validate the model himself, he builds a system where one model validates another. Not human ↔ model, but model ↔ model, directly.
He assembles a loop or a graph of agents. One model proposes a solution. Another validates it. A third checks the architecture. A fourth writes tests. Roles are encapsulated — from analyst to tester, each with a narrow mandate. And over all of it he hangs things that don't argue and don't get tired: deterministic linters and automated tests that catch mistakes objectively and return them to the model for correction. Not "this looks off to me," but a red run that won't go green until the code is fixed.
Notice what he does not do. He doesn't sit next to the junior. He removed himself from the hot loop entirely. He doesn't watch the model produce its first broken draft — another model and a linter watch that for him.
The same garbage, behind a closed door
Now the main point. It looks like the humanities guy's model is "smarter" — it works, after all. Here's the truth: the model produces exactly as much garbage as yours. The same crooked first drafts, the same stupid mistakes, the same "proposed → failed → rewrote" iterations.
The difference isn't model quality. It's the same model. The difference is whether the garbage is visible from the outside.
Yours: the entire raw stream passes through your eyes — you are the observation point, sitting inside the loop, watching every draft. His: the internal model-to-model dialogue is hidden in logs. One model produced nonsense, another rejected it, a linter caught the error, a third rewrote it — all of it simmering behind a closed door. What reaches the human is only what has already passed internal quality control. A polished result. He has no sense the model is struggling — it struggles, of course it does, mightily — but he doesn't see it. He sees the output, not the kitchen.
You optimized your dialogue with a junior. He pulled the dialogue out from under his own observation and kept only acceptance testing. The same unreliable component: yours is exposed with its guts open, his is wrapped in scaffolding that absorbs its failures before they reach a human.
Who actually built an engineering system
Lay out how each of you treats your executor.
You treat the agent as a junior colleague. One executor, one communication channel, one controller — you. Relationship model: "competent intern under supervision." Supervision is manual, so it scales to exactly one head and one focus of attention.
He treats the agent as a contractor team. The kind of team where all roles are encapsulated from analyst to tester, with a quality-assurance loop running inside. Relationship model: "I set the task, I do acceptance on the output, and how they negotiate among themselves is their business and their logs." Control is built into the structure, not hanging on one person's attention.
Here's the fork, and you should swallow it whole. This isn't "smart engineer vs. dumb humanities guy." It's a fork between two relationship models with the same tool. The one that produces stable results is not the one that demands more qualification. It demands less: you don't need to argue with the model on substance, you need to build a loop that argues with itself.
The amateur built the mature scaffolding; you didn't
The final sting, and I won't soften it. The person who "doesn't understand our work" accidentally built more mature engineering scaffolding around an unreliable component than you did. Role separation, automated quality control, feedback loops for correction, acceptance at the boundary — this is exactly what we engineers have spent decades building around any unreliable part of a system. He arrived at it not from brilliance but from necessity: he couldn't do it any other way. You didn't arrive at it because you could — and you stopped at "junior under supervision."
He understands exactly zero about engineering. He just treats the model as what it is: an unreliable team that needs to be surrounded with control, not a colleague who needs to be trained. And you, with all your qualification, still talk to it like a person — so you see every mistake with your own eyes and call it a stupid model.
The model is the same. What's different is where you put your quality control: on your own shoulders or into the loop.