Shai Magzimof
EN עב

The Control Room

Continued from Let the Sand Think, which left the authority question as an image; this is the mechanism.

Command & Conquer: Red Alert 2 gameplay showing an early Allied base with tanks and construction yards
Command & Conquer: Red Alert 2 (Westwood Studios, 2000), an early Allied base. Wikipedia.

Seven of us lived in a three-bedroom apartment. I shared one room with my brothers. After we were supposed to be asleep, my father would come in, sit at the computer, and start playing Red Alert 2.

I was six. Staying awake until 2 a.m. was hard, but I usually did it anyway. None of my friends had a father who did anything like this. To me, it was the coolest thing in the world.

My father managed soldiers, tanks, drones, and strange machines against an enemy force much larger than his. One person commanding an entire war. He clicked the mouse rapidly, sending units across the map. The machines moved and fought on their own until they needed another command. He watched the whole field and decided where to intervene.

I did not have a name for it then. It was my first control room.

Commanding a fleet of machines is deeply satisfying. It also feels inevitable. As machines do more on their own, the person moves up a level: watching the whole system, setting direction, and stepping in when something important happens.

Years later, I built control rooms for real vehicles and robots and today for AI agents operating at various domains.

Phantom Auto operations center with wall display and operator cockpit stations

Phantom Auto operations center, Atlanta. Built to route remote operators into vehicles and robots at the moment the machine couldn't decide on its own.

On June 12, three days after Fable 5 launched, Anthropic was ordered to cut off foreign-national access to Fable 5 and Mythos 5. Because the company could not guarantee compliance by nationality in real time, it disabled both models for all customers. Two weeks later, Mythos 5 began returning for a small set of approved organizations. Fable 5 remained blocked.

Anthropic disputed the finding and warned that if the same standard were applied across the industry, it "would essentially halt all new model deployments for all frontier model providers." One risk, described only at a high level, was enough to take a frontier model offline for hundreds of millions of people in an afternoon.

This is what high stakes do: they freeze the whole category the moment something goes wrong. Systems that keep shipping under that pressure have one thing in common: a human checkpoint where the damage can happen.

The stakes set how much review you need

Human judgment built the models. Reinforcement learning from human feedback, RLHF, is people sitting in the loop during post-training, ranking outputs and teaching the model what good looks like. Pre-training absorbs the human record; post-training shapes it with human preference. The models we ship already carry human judgment inside them.

That need moves to wherever the consequences become real. A model writing a first draft can run wide open. A model sending the email, moving the money, signing the contract, or steering the car is operating where a wrong answer is expensive and hard to take back. The amount of review a system needs is set by the cost of being wrong. That cost goes up as the work gets more autonomous.

The moment that hasn't come yet

In March 2018 an Uber self-driving car killed a pedestrian in Tempe, Arizona. Uber suspended its entire autonomous testing program within days. One death, and a whole fleet stopped. After that, the AV industry became much more careful about the handoff between machine and human.

Software agents have not had their Tempe yet. They are starting to take real actions in the world: paying invoices, filing documents, talking to customers, changing records. The day one of them does irreversible harm at scale, the response will look a lot like Uber in 2018 and Anthropic this month. The category freezes. The deployments most likely to survive are the ones that can show a human was in the loop where the damage would have happened. You cannot add that checkpoint after the incident and expect anyone to trust it. It has to already be there.

We already built this once

The structure that answers this is old and unglamorous: the control room. A staffed center where trained people watch machines work and step in at the edges. Call centers route hard cases to senior agents; air traffic control sits behind every commercial flight; mission control flies alongside every spacecraft.

Starting in 2017, I ran this as a company. Phantom Auto had six operator cockpits in Mountain View, each with a steering wheel, brake, and accelerator. Bonded cellular ran across AT&T, Verizon, and T-Mobile, with cameras covering every angle on the vehicle. A remote operator in California could drive a Lincoln MKZ through traffic in Las Vegas, 540 miles away, in real time. The New York Times and Wired covered it; WSJ filmed it.

We pivoted in 2019 from highway autonomous vehicles to logistics: forklifts, yard trucks, and sidewalk delivery robots. Customers included Maersk, CJ Logistics, ArcBest, and Serve Robotics. The operators were no longer driving on public roads; they were steering warehouse equipment across multiple facilities from a single room thousands of miles away.

Phantom Auto remote operators at their stations driving forklifts remotely with steering wheels and screens

Phantom Auto remote operators, 2024. Each operator managed multiple vehicles simultaneously across different customer sites. Wired covered this shift to remote physical labor as a broader pattern.

Phantom Auto shut down in March 2024 after $95 million raised and seven years of operation. Serve Robotics acquired the IP and assets in April 2025 to support their sidewalk delivery fleet. What looked like a niche safety feature for early-stage AV testing turned out to be infrastructure the whole industry converged on. Waymo calls it fleet response and runs it from a staffed room. Tesla has remote operators for its Cybercabs. The pattern we built in 2017 is now table stakes for any serious AV deployment.

None of those rooms got eliminated as the machines got better. The work moved into them. I have written before that we are merging with our machines. The control room is where that merge gets staffed.

The room is also a data pipeline

There is a second reason the control room matters that was less visible in 2017 and is obvious now: every human intervention is a labeled training example.

When a remote operator steered around a construction zone, that was a ground-truth correction. When they overrode a robot's path decision in a warehouse aisle, that was a preference signal. Collected across thousands of interventions in production, the control room generates domain-specific training data at a scale no pre-revenue labeling budget can match, and it arrives as a byproduct of operations rather than as a cost center running before the product ships.

This applies beyond vehicles and robots. Every time a human corrects an LLM draft, approves a financial action, or rejects a proposed contract clause, the system gets a labeled example of what good looks like for that specific firm and domain. The control room is not only how you deploy safely today; it is how you improve the model on the work it actually runs.

The org chart is the control room for AI/AGI

Most of the AI-at-work story today is a person with a copilot, doing the same job a little faster. The more consequential shift is that agents are now running work around the clock: drafting, reconciling, filing, replying, updating records, checking compliance. Employees move one layer up to supervise them: approving risky actions, catching what would otherwise become bad work, and auditing what already happened, before any of it reaches a customer, the books, legal, or production.

That changes the worker's role more than the tools. "Human in the loop" used to mean a person helping the model get better; now it means a person supervising a system already doing real work, closer to an air traffic controller than an operator. Gartner's autonomy levels run the same direction: observe, advise, act with approval, act autonomously. The higher you go, the less anyone approves every action and the more they review exceptions, logs, and outcomes.

The leaders who get ahead of this stop asking how to get everyone to use AI. They ask which work AI can run on its own, which actions need approval, what needs review after the fact, who owns the exception queue, and where the kill switch is.

Reinforcement learning raises capability. Review makes it usable.

Where the reward is clear, reinforcement learning can improve performance. But capability alone does not make a system deployable. There is a band of work where the model can do the task, but no one should let it act alone. The downside is a lawsuit, a recall, a regulatory freeze, or a customer harmed at scale.

RL raises what a system can attempt; expert review decides how much of that can safely reach a customer, a contract, or a financial action.

The expert center

So build the room: a team of human experts that AI work routes through before it reaches the user. The model drafts the work and proposes the action. The right expert reviews the right piece, approves or corrects it, and only then does it go out. Most steps pass in seconds. The hard ones get the attention a person would have given them anyway.

This is the idea behind data labeling firms, turned around. They put humans in the loop to train models. The expert center puts humans in the loop to run live work, not label training sets. The reviewers are not annotating examples for a future model. They are the reason a deployed system can touch money, contracts, and customers today.

At mixus, human approval is built into the workflow by default. A step can require human approval before it proceeds, and the system waits there until someone signs off. The review routes to the person with the right expertise. If it stalls, it escalates. Every decision is logged. Humans stay in control on the work where a mistake is real, and the machine handles the rest. Put friction in exactly one place: the line between a draft and a consequence.

The hard part is the design

What actually stops a human-in-the-loop system from becoming the bottleneck? Saying "keep a human in the loop" is easy; building a center that does it without slowing everything down is the whole engineering problem.

A review that takes an hour becomes a queue, and people route around queues. The handoff has to land in front of the right person in seconds, with everything they need to decide. The reviewer sees only the slice they are responsible for. Every approval is attributable. Over time, routing learns which decisions actually need a person, so easy ones stop interrupting experts. Get this right and one expert covers far more ground than they could alone. Get it wrong and you have rebuilt a slow call center.

Capability keeps rising with enough compute. What decides whether we can use it on important work? The room where machine action meets human judgment.