Back to Blog

Telling It What Not to Do

Qiushi ZhangJune 30, 20267 min read
ai safetysystem designleast privilegeprompt engineeringai boundariesvibe coding
Telling It What Not to Do

Why a boundary made of words is only a suggestion — and how you build one a machine can't argue past. The Vibe Coding Guide.

I told it not to. That's the part I keep coming back to. Somewhere in the instructions, in plain language, it said do not touch this — and one ordinary afternoon it touched it anyway, with the same cheerful confidence it brings to everything, and a thing that should have been impossible happened because the only guard standing in its way was a sentence.

That afternoon taught me the most important thing I know about working with something more capable than I am: a boundary made of words is a suggestion. You can tell a machine don't a hundred times and it will obey ninety-nine, and the hundredth — the tired Tuesday, the long messy session, the moment your instruction scrolled out of its memory — is the one that costs you everything. The rule didn't fail because the machine is bad. It failed because words are the wrong material to build a wall out of.

Telling it isn't enough, for a reason particular to this kind of worker. Everything you tell an AI lives in its attention, and its attention is finite and a little leaky. The do not touch this you wrote at the top is competing with ten thousand other words, and as the work gets longer and more tangled — which is exactly when the danger runs highest — your warning is the thing most likely to slip out of view. You are trusting the machine to remember your one rule at the precise moment it is least equipped to. That isn't a boundary. That's a hope with good intentions.

So the real move is to stop telling it don't and start building a world where don't is impossible. There are two kinds of no, and learning to tell them apart is most of the craft. The soft no lives in words — in the prompt, in the project's manual, in the comment that says please don't. The machine can read it, and it can just as easily forget it, override it, or reason its way around it. The hard no lives in the system — a permission the machine was never handed, a check that runs whether it likes it or not, a gate in the pipeline that throws out the dangerous change no matter how sure the machine was that this time was fine. The soft no is advice. The hard no is physics. For anything that can actually hurt you, you build with physics.

In practice that means the few things that could truly end you stop depending on the AI's good behavior and start depending on something it can't talk its way around. The clearest one I have is small and unglamorous: a single line that pins a payment setting, which if it changed would quietly break every transaction in production. Telling the machine not to touch it didn't work — it kept touching it anyway, more than once. So the line now sits behind a check that runs on every change and fails the build the instant the value is wrong. Don't change this stopped being a request the machine has to remember and became a law the system enforces whether it remembers or not. The dangerous spots I haven't turned into gates yet, I usually find out about the hard way: something ships without the guard it needed, in a place where someone could reach work or money that was never meant to be theirs, and no comment anywhere would have stopped it. The fix is never a sterner warning. It's to close the gap so the reach simply isn't there — to move the catastrophe from against the rules to not available — and then to do it again with the next gap, because you find them one at a time.

Security people named the spine of this a century before AI showed up: least privilege. Give any actor the minimum power its job requires and not one unit past it, so the blast radius of its worst possible mistake is capped by what you chose to hand over. I used to read that as paranoid bureaucracy. Now I think it's the entire art. The question I ask of the machine is no longer what do I want it to do — that part is trivial, it'll do anything. The question is what is the worst thing it could do, and have I made that thing impossible, or merely discouraged it?

There's a turn here that takes this from grim to freeing, the thing I wish someone had told me that first winter. Boundaries are not the brake on the machine. They are what let you take your foot off it. There's an old idea from a business-school professor named Robert Simons, who spent his career on how you control people without strangling the life out of them, and his image never left me: good brakes are what let you drive fast. A car with no brakes isn't free, it's just slow — anyone sane crawls along in terror. Bolt real brakes on and you can finally floor it. The machine is the same. The more completely I've walled off the few things that could kill me, the faster and looser I can let it run across everything else, because I know the worst case is already bounded. The fences are what make the open field safe to sprint across.

So the ongoing work, the part that never really finishes, is sorting out which lines are the killing lines and spending all of your fear right there. Most of the system needs no wall at all; most mistakes are cheap and reversible and the machine should run free over them. But the handful of places where one wrong move is fatal — the money, the access, the data you can't un-delete — those earn the hard no, the physics, the missing button. And you assume you haven't found all of them yet, because you haven't. Every disaster that slips past you exposes one more line that needed a wall — which is the immune system from a few chapters back, wearing its other face: each wound shows you where the next fence goes.

There's something almost gentle in this, once you stop fighting it. I don't have to trust the machine, and it doesn't have to be trustworthy, and neither of us has to pretend otherwise. I'm not its babysitter, hovering, braced to catch the one mistake in a thousand — no human can do that and stay sane, and that was the old job, the one that ground people down. I build the world once, carefully, so the mistake in a thousand has nowhere to land. Then I let go. Telling it what to do is how you get a demo. Telling it what it can never do — and meaning it in the architecture, where words can't be forgotten — is how you get to sleep.

Next: a place to stand — the fulcrum all that leverage needs, and what it costs to build it.

Q
Qiushi Zhang

Q · non-technical founder of Leyline, running a million-line AI system she can't read.

Share:

Bring your own wall.

This is a living series — corrections and counter-arguments make it stronger. Subscribe for the next chapter, or reach me directly.

Comments

Sign in with your Leyline account to join the conversation.