AI Companies Are Now Building Little Robot Departments, and Somehow That's the Sane Option

Share
AI Companies Are Now Building Little Robot Departments, and Somehow That's the Sane Option

There's a new trend in AI, and it is exactly what it sounds like if you've ever worked at a company with too many meetings. Instead of asking one AI model to do everything, researchers are now spinning up several of them at once, each with a job title, and making them talk to each other until the work gets done. A researcher agent goes and digs things up. A coder agent writes the actual code. A critic agent sits there and picks holes in what the coder just wrote. An executor agent runs the thing and reports back on whether it caught fire. Put them in a shared workspace, let them pass notes back and forth, and you get something that behaves less like a chatbot answering a question and more like a very caffeinated project team that never sleeps, never takes a lunch break, and occasionally forgets what it was doing halfway through.

Why bother splitting one brain into five

The pitch is pretty simple once you say it out loud. A single AI model, no matter how good, has to hold everything in its head at once: the original question, the research it dug up, the code it's writing, and whether any of it still makes sense forty steps later. That's a lot to juggle, and juggling too much tends to be where things quietly go wrong. Split the job across several agents, each with its own clean workspace and a narrow lane to stay in, and each one gets to focus on being good at one thing instead of okay at everything.

Anthropic, which has actually built and shipped this kind of system, lays out three situations where handing work off to a team of agents beats using just one. The first is keeping unrelated information from cluttering an agent's head, so a customer support agent trying to fix your login problem isn't also wading through your entire order history. The second is parallelization, meaning several agents can go explore different angles of a research problem at the same time instead of one agent plodding through them one after another. The third is specialization, since stuffing one agent with fifteen or twenty different tools tends to make it worse at using any of them, the AI equivalent of a Swiss Army knife so overloaded it can't cut anything cleanly anymore.

The part where it gets expensive and a little unhinged

According to Anthropic's own writeup, running the same task through a multi-agent setup instead of a single agent burns through three to ten times more tokens, which is a polite way of saying three to ten times the cost. Where does all that extra money go? Mostly into agents talking to each other, duplicating context, and summarizing results for each other, which is basically the AI version of a status meeting that could have been an email.

It gets weirder. Anthropic found that if you split up the work by phase, planning first, then building, then testing, agents lose context every time the baton gets passed, and the coordination overhead can eat more tokens than the actual work does. In one internal experiment, subagents reportedly spent more effort coordinating with each other than doing the job they were coordinating about. There's also a delightful failure mode nicknamed "early victory," where a verification agent runs a couple of quick tests, decides everything looks fine, and declares success without bothering to check the parts that would have actually broken. Which is, again, extremely reminiscent of a coworker who says "yep, all good" without opening the file.

So no, this is not some flawless robot dream team humming along in perfect harmony. It is closer to a genuinely capable but slightly chaotic small team, complete with communication overhead, someone cutting corners, and a budget that quietly balloons if nobody's watching it.

When it actually works, it really works

And when it's tuned right, it apparently does work. Anthropic's 2026 agentic coding report claims one of its coding agents implemented a complex set of changes across a 12.5 million line codebase in seven hours with 99.9 percent accuracy, the kind of task that would eat a human engineering team's entire sprint. A recruiting company called Fountain reportedly used a hierarchical setup of multiple coordinated agents and saw screening speed up by 50 percent, onboarding get 40 percent faster, and candidate conversions double. The same report notes that developers currently lean on AI for something like 60 percent of their work, but can fully hand off only zero to 20 percent of tasks. Translation: the machines are doing a lot of the typing, but somebody still has to stick around and make sure they didn't quietly wander off a cliff.

The bigger shift the report describes is that engineers are being nudged out of typing code and into something more like managing it, reviewing architecture, setting direction, and stepping in when the AI team needs a decision only a human is allowed to make. It also sounds exactly like the plot of every story where the humans hand over the keys "just for the routine stuff" and then discover, a few years later, that the routine stuff was actually all of it.

The honest takeaway

Multi-agent AI systems are not magic and they are not, contrary to how they get marketed, simply "more AI equals more better." They're a real engineering tradeoff. Done carefully, with clear boundaries between agents and someone watching the token bill, they can genuinely tackle problems too sprawling or too long for one model to hold in its head. Done carelessly, you get an expensive committee of AIs in-fighting over context, congratulating itself on a job half-finished, and burning ten times the electricity to arrive at roughly the same answer a single, well-prompted model could have given you for a tenth of the price.

So sure, keep building the robot departments. Give them job titles, let them talk to each other, let them run projects too big for one model alone. Just maybe don't leave the room while they're in there deciding among themselves what counts as "done."

Sleep well.

Sources