On August 13, DeepSeek open-sourced its agent harness. Twenty-two thousand GitHub stars in an hour and a half; past sixty-five thousand within a day. The README says:
Everything is a plugin. Every run is traceable.
This is fascinating — I honestly didn’t expect DeepSeek to go this far. And the first thing it made me think about wasn’t code. It was robot commissioning — the line item that eats nearly forty percent of an industrial automation project’s cost — and how it ought to be done. Because that sentence says, in the language of software, something the robot-commissioning trade has been doing for thirty years without ever giving it a name.
First, What a Harness Is
The core loop of an agent is simple: the model reads context, decides which tool to call, the tool runs, the result goes back into context, look again. A few dozen lines gets you a working loop. The hard part is the ring around the loop: who allows it to call this tool? What gets checked before the call? Who confirms the result afterward? When something goes wrong, how do you replay what it saw at the time?
That ring is the harness. The word means, literally, the tack you put on a horse — the gear that lets a fast animal be controlled, stopped, and traced. The horse is the model. The harness is the discipline around it.
DeepSeek’s version is radical: there is no privileged core. The model adapter, the tool registry, the session log, the agent loop itself — all plugins. Adding a capability doesn’t mean editing the core; it means mounting a plugin beside the others. But “everything is a plugin” as a slogan isn’t what I want to talk about. What matters to me are the two mechanisms underneath.
First: checks can be layered outside a tool without touching the tool. The Cordis framework underneath has something called waterfall — essentially around-middleware. Each listener receives (…args, next). It can inspect the arguments and then call next() to let the call through; it can decline to call next() and short-circuit; or it can take the downstream result, alter it, and pass it back up. The lifecycle document describes tool execution as “ordered pre, concurrent execute” — pre-checks run in order, and only what passes runs concurrently.
Which means: adding a new safety check to a tool does not require opening that tool’s source. Checks are external, and they stack. Hold onto this — on the robot side it becomes a matter of life and death, and I’ll come back to it.
Second: every run can be replayed. Everything the model sees — system prompts, reasoning, tool calls and results, subagent scheduling, every context injection — goes into an append-only session log. Resume, fork, search, and replay all operate on the same event stream. When something breaks, you can see exactly what it saw and why it decided what it did.
Put plainly: the parts that build capability (model, tools) and the parts that constrain it (checks, log) are structurally separated, and the second can grow independently of the first.
Three Rounds of Collision Checking, and None of Them Included the Part
A while ago I took a real, large die-cast structural part and, using an AI agent, ran the entire chain offline — from CAD to a two-arm robot deburring path. Parting-line pickup, curve sampling, tool-point generation, dual-arm layout optimization, arm assignment, visualization — all the way to watching the robots move frame by frame on screen.
It ran. But running wasn’t the real takeaway. The real takeaway was watching how the errors got found.
The process surfaced a dozen or so substantive defects. Most of them I caught by glancing at the screen: the part was lying upside down; the burr direction pointed inward; tool points flew a meter outside the part; the two arms kept trading off back and forth along one segment, which is unworkable as a process. A handful were caught by the system itself, because something didn’t match real data.
The one that bothered me most was collision checking. The agent did three rounds. All three passed. All three left the workpiece out of the collision model — it only checked whether the robot would hit itself. Deburring means reaching into the part’s cavities; the most typical interference is hitting the part. Three rounds of self-checking missed the most basic collision in the trade.
A check designed by the agent that wrote the code cannot catch the failure modes that agent didn’t think of.
This isn’t about the model not being smart enough. Swap in a stronger model and it writes prettier collision checks — that still only cover the cases it thought of. The builder’s blind spots get copied, intact, into the verification the builder writes. That’s a structural problem, not a capability problem.
The old hands have always known this — they just don’t say it this way. Anyone who has commissioned a robot knows the ritual: teach, dry run, single-step, low speed, full speed. Every step is a check independent of the person who wrote the program. The dry run checks the trajectory; single-step checks each point; low speed checks the dynamics. What the ritual is really saying is: the brain that wrote the program is not to be trusted, so wrap it in layer after layer of checks that don’t depend on it.
That’s a harness. The shop floor was using one for thirty years before software gave it a name.
Why a Robot’s Harness Has to Be Harder
Up to here the two worlds look like they’ve merely arrived at the same place. But the robot side has a boundary the software world doesn’t.
Every error in my exploration — the upside-down part, the tool point a meter out, the missed collision — what did it cost? Rerun. A few minutes. In the world of pure software agents, failure is cheap, reversible, and infinitely retryable, so you can build boldly, fail fast, fix later.
The moment that program goes down to a real robot, the tool point a meter out is no longer a bug. It’s a crash. A bent fixture, a wrecked end effector, a stopped line — and, on a bad day, someone hurt. The fault tolerance of the offline stage does not exist on the production line.
That boundary changes the design requirements for the harness. At least three of them.
One: verification must be independent of generation. In software, letting the agent write its own tests is an acceptable shortcut — if the tests miss something, you run one more round. In robotics that shortcut is the direct cause of accidents. Real verification has to come from outside the agent: in-line inspection, comparison against real-machine measurement, or a second, fully independent implementation for cross-checking. Verification coverage must not fall below automation coverage. Every ring you add to automation enlarges the region where the system can quietly produce wrong results; verification has to keep pace, or you’re only widening the blind spot.
Two: the human’s position is structural, not transitional. Most of my dozen-odd defects were “glance at the screen and see it.” That kind of perception — is the part oriented right, is the direction flipped, does the path look sane — no automation replaces today, and I doubt it should any time soon. The right division of labor: AI owns the speed of building and iterating; humans own perception and anchoring to reality. The first can be accelerated without limit. The second cannot be dropped. Some semantics can only be supplied by a person; the job is to make that fast (one click instead of an hour by hand), not to eliminate it.
Three: every time you add a piece of automation, add a check that can independently catch it being wrong. I now write this sentence into every project charter. It sounds like common sense, but common sense without structural support doesn’t survive the second project — under deadline you add the feature first and leave the check for “next time.” The reason the harness’s waterfall mechanism caught my attention is exactly that it moves this discipline from relying on conscientiousness to relying on structure: checks are external, stackable, and addable without modifying the tool. Get the structure right, and the discipline holds.
The Asset Isn’t the Software. It’s the Harness.
Before that exploration, I thought the output would be “a piece of software that automatically generates robot paths.” Afterward I understood the software itself was worth little. What was actually valuable were the few lines inside each tool that read: “you only know to write it this way after you’ve hit the wall.”
A certain coordinate frame is off by a fixed transform — miss it and you’re off by a few degrees the whole way, with symptoms that point misleadingly elsewhere. A certain geometry can’t use the standard simplification under this kind of operating condition, or it false-alarms end to end. A certain degree of freedom is redundant, and locking it throws away a dimension for nothing.
These things share one trait: they can’t be derived from documentation or experience. They can only be hit. And once hit — and encoded into a check — the next part never pays that tuition again. The compounding isn’t in the features. It’s in the constraints that stop charging you tuition.
So the asset isn’t the software. It’s the harness. And a harness can only be fattened inside real projects. R&D detached from projects produces things that are self-consistent and wrong — because nobody knows what to check.
DeepSeek’s line, translated for the shop floor, comes out like this: capabilities are pluggable; discipline must leave a trace. However fast the builder runs, the verifier cannot be itself.
Software just wrote that law into code. Robotics has been learning it, one wrecked fixture at a time, for thirty years.





