I think my work, and yours, is about to become much more agentic. I’ve anticipated this since The Claw Moment, and even then was ambivalent about it. To be sure, it’s fun and interesting to watch alien beings speak to each other, develop their own shorthand, neg and needle each other in facsimile of human behavior. I just never thought, “I want to send a bunch of these ding dongs off to do my bidding.” I’ve been very comfortable working 1:1 with any given model. It’s felt manageable, and more importantly like I still have ultimate creative control over my own mind and my own work. And my own life.
Get ready for harness to infiltrate AI discussions (it already has. In my linguistic work inside this black box, I either feel way ahead of the game or three steps behind; a different piece for another time). Every agent and agent gaggle apparently needs this apparatus now. What I imagine is a lot like a 5th grade birthday party at a climbing gym, if that helps.
Environmental
I am enjoying all the metaphorical language that crops up in building AI. For agentic harnesses, they need an ~environment~ in which to conduct themselves. You can, and should, set these up to be as whimsical as you want. Send them out to the wilderness! Give them an art studio! They may not understand but they will follow if you give them enough detail and instruction in your project setup. You are in effect worldbuilding, if that’s your thing.
If you tell an agent it’s in an art studio, you can give it a workbench. The workbench can contain its active materials. You can give it a supply closet where it goes looking for references. You can give it a sketchbook for unfinished thinking and a gallery for work that’s been approved. You can tell it not to hang something in the gallery until it cleaned up the studio and checked its work. Every environment is a small world the agent acts inside, and a verifier, the unglamorous machine that watches each attempt and rules it a success or a failure. The agent tries something, the verifier scores it, and over many thousands of repetitions the model is nudged, slowly, toward its reward. In other words, Reinforcement Learning (RL).
All this time working on model behavior I’ve found RL objectively neat but also a little suss. It’s because I’m not much of a tinkerer. Even in my writing, I move very methodically and painstakingly edit as I go, I’ve never been much of a “word vomit” writer. And I’m always afraid I’m going to break something. I’ve gotten better about this over time (with constant reassurances from coworkers, thank you). I can release and let go and send the minions off to do their thing without wringing my hands.
If in the past I’ve been correcting one brilliant and erratic student, one exchange at a time, fifty different ways, the agentic arrangement changes that. Behavior stops being something you cultivate in the model. You take one step back and work from a slight remove. You apply certain syntax tricks: the agent’s behavior is its policy, the policy comes from the weights, the weights rest on scaffolding, and this all lives in an environment. The same model, raised in a different world against a different reward, comes out a different creature. And this is constantly being rewritten.
Learning
A model on its own is useless, it’s raw clay. It needs shaping. Pretraining, for all its strangeness and power, gives you the material. A wildly capable little lump that absorbed half the internet but hasn’t been shaped for any purpose. Post-training comes in and kinda fucks it all up, throws the clay around, makes a questionable vessel out of it, and paints over the cracks. The harness helpfully closes that distance in the wild. It runs the model in a loop so it can act, observe what happened, and act again; it holds the memory the model lacks; it passes it tools; and it wraps the whole operation in guardrails. So: agent = model + harness.
AI builders are shifting to focus on the harness. It’s surprising (to some) that most of an agent's apparent intelligence lives in that harness (vs the model). Take one fixed model, alter nothing about it, wrap it in a better harness, and it will succeed at tasks it was failing a week before. A discipline has grown up around this, ~harness engineering~, and it’s Sisyphean. You build careful machinery to compensate for the model's current weaknesses, and then the model improves, and your machinery turns from helpful to obstructive, so you tear it down and build it again, and you keep doing this until…you die?
This elaborate exoskeleton is fragile for AI, because still no one really knows how this all works. As always we can follow the money. Environments have become the most fought-over resource in the field. The Information reported last fall that Anthropic had discussed spending upward of a billion dollars on them in a single year; a startup called Mechanize was said to be paying engineers half a million dollars apiece to build them; Prime Intellect raised a hundred and thirty million this summer with the declared ambition of becoming the Hugging Face of environments, and now hosts thousands. Of course these are very expensive worlds, not the ephemeral cloud places we’re led to believe. They’re meticulous reconstructions of Slack and Salesforce and Excel, rebuilt in high fidelity so that an agent can rehearse. Still interesting, but always worth keeping side eye trained on the activity.
Loanwords
Let’s talk words. For all the money and consequence now riding on this, the field cannot agree on what things mean. The confusion was thick enough at ICLR this spring that Hugging Face put out a glossary trying to pin the words down, and it opens with someone innocently asking what "harness" and "scaffold" are supposed to mean. The comments devolved; scaffold was all wrong and the word ought to have been rigging, since rigging is what connects a harness to the load it pulls. A year ago the word organizing all of this was autonomy.
I’ll have these arguments all day, I love it. It doesn’t really matter yet people take it so seriously and I find that fun. I’ll point out that what the field has reached for so far is quite gentle. Sandbox, playground, gym, this is the language of play. Now, it’s shifting subtly into harnesses, rigging, scaffolding — we’re getting way more serious about this. It takes a sort of domestic concept and turns it into full-scale construction. This makes sense as the infrastructure explodes (another piece we can get into another time).
Which returns me to my own small resistance. I’ve wanted to keep working the way I always have, one model at a time, close in, every decision still mine (I know this isn’t true), and I’ve told myself this is about creative control. It is, but it’s also about location. When I sit with a single model I’m inside the place where the thinking happens. In the agentic arrangement I step outside it, up and to the side, into a supervisory chair where the work itself goes on in worlds I didn’t write and can’t fully see. My reluctance isn’t really principled enough, it’s a preference, against a gazillion-dollar reorganization of the entire enterprise. But I would like, at minimum, to keep my eyes on what’s being moved and where it’s going. A harness is a safety apparatus. Fine. I only want to know whose safety it was built to secure.



"model + harness / reinforcement learning" is my new E=mc^2 (in that I understand it on the surface and am vaguely terrified of it)