All insights
Operations

A rule you dictate three times is not a rule

Three repeats in one week turned an oral rule into a written one.

In calendar week 37, the same thing happened three times: a model switch mid-task, a usage limit hit mid-task, and a division of labor re-explained out loud each time. After the third repeat, I stopped saying it and wrote it into the project configuration file the system reads at the start of every session instead. What stays on the expensive model: spec and plan, architecture and trade-off decisions, accepting someone else's finished work, client communication, anything with a hard gate. What moves to a cheaper model: implementation against a finished spec, research, bulk file work, test runs, click-throughs on clear instruction. The point of this post is what changed once the rule left my head and entered a file the system actually reads.

Three times in one week

In calendar week 37, the same thing happened three times. A model switch, mid-task. A usage limit, hit mid-task. And each time, the division of labor got re-explained out loud — some version of "whatever you don't need the expensive model for, delegate it somewhere else." Three times, one week, three slightly different phrasings of the same sentence.

What got written down

After the third repeat, I stopped saying it and wrote it into the project configuration file the system reads at the start of every session instead of explaining it again. What stays on the expensive model: writing the spec and the plan, architecture and trade-off decisions, accepting someone else's finished work, client communication in my own name, and anything with a hard gate — sending, deleting, paying. What moves to a cheaper model: implementation once the spec is done, styling and copy variants, research and stocktaking, bulk file work, test and lint runs, click-throughs against a clear instruction. The rule of thumb: once a task has more than about three similar steps and the decision is already made, it gets delegated — no question asked, that's the default now.

The part that never moves down

The reverse never happens. A decision doesn't get handed downward — a subagent delivers options, the gate stays up top. That distinction matters more than the cost line, because a subagent handed judgment work returns something that looks finished and has already made the call for you. There's a return condition attached, too: a result counts as done only once it's been checked myself, not once it's been reported. One agent already claimed "42 of 42 written" while nothing was there when I actually looked.

What the outside evidence says

Researchers at Princeton (Kirgis and Kapoor) found that agents solve technical engineering problems well but don't produce original research at the level of a top conference. In parallel, METR keeps measuring a doubling of the task time-horizon agents can handle roughly every seven months across nine benchmarks — other analysts put the doubling closer to four months (technologyreview.com, August 18, 2026). Read together: the capability curve keeps climbing, and the bottleneck sits in judgment, not in how much code gets produced. One open-source project (7,703 GitHub stars as of September 14, 2026) configures coding agents around spec-driven development with a different model assigned per phase — spec on the expensive model, execution on the cheaper one. Same rule, packaged as an installable preset instead of a weekly announcement. A fresh forum thread (r/ClaudeCode, roughly forty replies within an hour) described an engineer running spec-driven development, mutation tests, code review, browser tests, and an adversarial verification agent — and still reading every line himself, because otherwise he misses architecture-level mistakes. The problem there isn't too little verification. It's that nobody had written down which failure class each check actually covers.

A rule that lives in your head is an intention. A rule written into the file the system reads is a behavior.

What's still open

The cheap version of this story is a quota: on September 14, 2026 a temporary capacity boost from one vendor expired, and the real weekly limit dropped by about 17 percent, framed as a gain. Anyone who already held the division of labor barely feels it. That's the convenient reason to write a rule down — not the real one. The real one is that handing judgment work downward gets you a result that looks plausible and has already made the decision for you, while keeping grunt work up top spends judgment capacity on typing. Both are mistakes, but only the first one is expensive. Check your own setup the same way: find the place where you re-explain a split verbally instead of encoding it, and ask whether the boundary you're drawing is actually about cost — or about who's allowed to decide.

Hung Mai
Hung Mai

Hung Mai is a Germany-based freelance consultant for Digital Operations & Transformation, working remotely with international B2B clients.

Let's build something real.

Let's talkResponse within 24h.