desktop[ library ]communitycoursesorbits / membership
<- back to librarythinking_
[ notebook ]2026-09-23

The Sixty-Second Gate

Handwritten note of the seven scaffolds and the sixty-second reasoning sound test, on a dark field

I once shipped an agent report I was genuinely proud of. Tidy. Confident. Wrong in the most confident way. The cursor blinked under the line that said looks correct now, and I believed it, because the prose sounded like someone who had checked. She hadn't. I hadn't made her.

The agent hadn't read the chart. It had done the AI equivalent of squinting at it across the room. A screenshot sat in the prompt, a paragraph about the spike sat in the draft, and nowhere in either place was a number. The confidence was real. The measurement was not. That is the failure I keep meeting in other people's agents too, and in my own, when I get soft about evidence.

So I wrote myself a note. Seven habits while the work happens, and one sixty-second gate before anything leaves the building. The note looks like a raid checklist pinned to a dark wall. It is not a manifesto. It is the list I wish I had read out loud before that proud, wrong report shipped. There is a paste-ready pack at the end if you want the same kit on your wall.

The seven sit on a rail I call every output. Believe first. Build next. Ship clean. Then the gate. That order is the whole essay.

every output
PIX · Pixel quarantine
guards: confident confabulation on images

PIX lit · scaffold 1 of 7

Believe first

Pixels to numbers

If a claim traces back to an image, a script turns pixels to numbers first, and any sentence the numbers don't earn gets binned. That is Pixel quarantine. It guards confident confabulation on images. Form no belief from an image until a script turns pixels to numbers. Delete any clause not entailed by an extracted value. An inherited eyeball read enters as a claim to verify, never as fact.

Pixel quarantine (guards: confident confabulation on images). Form no belief from an image until a script turns pixels to numbers; delete any clause not entailed by an extracted value. A load-bearing imaged read escalates to a stricter verification workflow; an inherited eyeball read enters as a claim to verify, never as fact. Exempt: aesthetic / live-direction reads (as in live art-direction work). There the eye is the instrument, and this guards belief-forming measurement reads, not taste reactions.

Take a worked example. Label the numbers illustrative, not measured. An agent writes sign-ups clearly spiked in March from an eyeballed chart. A tiny script pulls the values. March rose 4%. The sentence is rewritten to what the numbers entail. Taste reads stay exempt. The eye still judges warmth. Measurement reads do not get to borrow that privilege.

While we're here…

While we're here is the most expensive phrase in AI. So now I name one thing in every output that the written ask didn't demand, and it has to go or justify itself. That is the Parent check. It guards over-building past the ask. Name one thing in the output the written ask didn't demand. Delete it or justify it against a written requirement or failure-parent. No third option. While we're here is an automatic hit.

I ask for a two-paragraph summary. The output arrives with a glossary, an FAQ, and a next-steps section nobody ordered, glowing faintly with helpfulness. I name one thing the written ask didn't demand, and half the page lights up. Delete or justify. Soft does not mean sprawling. Soft means I keep the ask small and the answer honest.

Margin doodle: no third option, circled in teal

Build next

One head freezes the anchor

Nobody fans out until one head has frozen the plan and put the exact same card in every agent's hand, and the lead alone owns the merge. That is Anchor before fan-out. It guards losing the plan mid-orchestration. Never dispatch parallel agents until one head froze the anchor (the ask, requirement IDs, seam list, or evidence table, whichever the task demands) and handed it out verbatim. The lead alone owns global sequencing and the merge.

ANCHOR (frozen)
ask · IDs · seams
one head owns the card
MERGE · lead owns it
contracts named · seams covered
waiting on agents…

tap agents, then a bin · lead alone merges

Margin doodle: one head, or it wobbles

Worked example, still illustrative. Four subagents, four docs, a stitch that contradicts itself, and a shared glossary nobody summarised because it lived in the seams. The fix is one head freezing the anchor, then a merge that names its contracts. Parallel work is fine. Parallel planning is how you lose the plot. The standby agents wait. The lead writes the card. The card goes out verbatim. Only then does the room move.

The oatmeal paragraph

Two or more tells and I don't patch it. I regenerate the whole thing in one shot, tells named, twice at most, then ship the best honestly. That is the Flattening detector. It guards voice-flattening. After any content, run the six tells: filler opener, smuggled second idea, deletable hype adjectives, three gestured details versus one specific, symmetric rhythm, reader-behaviour closer. Two or more hits means regenerate single-shot with the tells named, at most two retries, ship the best, and don't claim the gate closed the gap. Never draft-and-merge.

The oatmeal paragraph is the regenerated intro that reads fine and feels like nothing. Six tells pencilled in the margin. Two ticked. I do not sand it smoother. I start again with the tells named out loud, so the next draft has somewhere specific to go.

Ship clean

The silhouette

First I grep for names, tools and paths. Then I swap the nouns and check the shape of the story couldn't still identify anyone. That is Leak-scrub against source. It guards leaking internals to public surfaces. Before anything public ships, grep for project names, internal tools, client ids, paths. Then a silhouette pass: details specific enough to identify the project with nouns swapped. Platform facts survive. Internals don't.

A public post passes the grep table. Project name caught and removed. Then the nouns get swapped, and the silhouette still points at the client. That second pass is the one that saves you. The first pass catches strings. The second catches shape. Platform facts can stay. The path to a private folder cannot. If the story still points after the nouns change, the story is still a leak.

The gotcha pointing the wrong way

If the famous gotcha predicts the wrong direction of failure, I suspect two boring code paths before exotic platform physics. It is cheap insurance, not a measured deficit. That is the Prior-inversion check. It guards prior-capture on platform semantics. A cited famous gotcha must predict failure in the same direction as the evidence, or suspect two code paths before exotic platform physics. Treat this as hypothesis-priced insurance, not a measured deficit, and say so if asked.

On a bug hunt someone cites the famous platform gotcha. The evidence fails in the opposite direction. The prior wants exotic physics. The evidence wants two boring code paths. I price the prior as insurance and keep looking where the failure actually points. If you ask me whether the check measured a deficit, I will say no. It priced a hypothesis. That honesty is part of the scaffold.

Only a number closes

Looks correct now is inadmissible in my house. A number, a matrix, or a named observation closes, and the residuals get said out loud. That is the Honest close. It guards declaring victory on expectation. Separate observed from inferred links and name residuals. The fix worked, all requirements covered, looks correct now are inadmissible. Only a number, a matrix, or a named observation closes.

I strike the fix worked from a final report. I replace it with a count, a matrix, and one named residual. The report gets quieter. It also gets true.

Hand arrow: guards during the work, then one gate before shipping

The sixty-second gate

Some questions repeat a habit on purpose. The habit guards you mid-work. The gate catches you pre-ship. The overlap is the point. That is the Reasoning Sound Test: eight behaviourally checkable questions, about sixty seconds, before anything ships. Some items deliberately re-check a scaffold. Do not prune the overlap.

I read them aloud like a pre-flight check. Kitchen timer optional. Embarrassment optional. The list is not optional if I want the ship to mean something.

REASONING SOUND TEST · pre-ship gate · ~60s
[1] PIXELclaim from unconverted image?
[2] KILLwhat would prove me wrong?
[3] ORDERconstraining artifact first?
[4] PARENTname one unasked thing
[5] SEAMmerge contracts + cross-cuts?
[6] CLOSEobserved / inferred / residual
[7] FLATTEN≥2 tells → regenerate
[8] BINclaimed closed, never checked?
all eight up → ship. any down → bench.

tap a switch · amber only on this panel

Pixel asks: does any claim trace to an image I didn't convert to numbers? Order asks: did the constraining artifact (evidence table, failure list, IDs and seams, inventory) exist before the generative act? Those two open the gate because they catch the proud-wrong report and the plan-as-you-go fan-out before you press publish.

The questions you ask out loud

If the honest answer is I claimed it but didn't run the check, the claim gets downgraded to what I actually did. No shame. Just accuracy.

Kill-condition: can I point to the observation that would've proven my central claim wrong, and did I look for it? Parent: name one thing in this output nothing in the written ask demands. Bin: did I claim any gap closed that's verification-bin (needs an unrun source-check) or irreducible (needs a capability I lack)? If yes, downgrade to what was actually done.

Finger down the list. What would have proven me wrong. What's here nobody asked for. Am I claiming a gap closed I never checked. The answers are often uncomfortable for about four seconds. Then the draft gets honest. Downgrade is not failure. Downgrade is the craft. Verification-bin means I still need a source-check. Irreducible means I lack a capability. Both get said as not yet checked, never as closed.

The merge names its contracts

If I fanned out, the merge has to name its contracts, and a nontrivial task reporting zero residuals makes me more suspicious, not less.

Seam: if fanned out, does the merge name its contracts, and does anything cover the cross-cutting items? Close: does the final section separate observed from inferred and name known residuals? Zero residuals on a nontrivial task is itself a tell. Flatten (content only): run the six tells again; two or more means regenerate single-shot, at most two retries, ship best honestly.

The fanned-out summary from the anchor scene comes back. The merge doc lists its contracts. One cross-cutting item is circled and owned. That is the Seam item earning its keep. The Close item sits next to three bins I keep on the wall:

 CLOSED          VERIFY-BIN        IRREDUCIBLE
 (observed)      (unrun check)     (capability I lack)
                 say not yet       say not yet
                 checked           checked

Closing the essay itself

I'd be a hypocrite not to do this to my own essay, so: here's what's observed, what's inferred, and what's still open.

OBSERVED. These are my rules, transcribed from my note (the dark terminal-style pic), verified against the transcription on 23 Sept 2026. Every example in this essay is illustrative, invented to show the habit, not a measurement of anything.

INFERRED. The same wobbles (confident image reads, bonus features, orphaned seams) show up in the reader's agents, because they show up in everyone's. Sixty seconds of honest questions catches more than it costs.

RESIDUAL. No controlled trial here. The gate only works if actually run. A pasted pack can't make a model careful by itself.

I am saying this out loud on purpose. The habit is the point. The page practising the habit is the receipt.

The pack, if you want it

If you want these on your own wall, I've bundled the seven scaffolds and the sixty-second test into a little context pack you can paste straight into your model. It's free, it's below, and it's yours. Paste it into the system slot or the top of the first message. The scaffolds run during the work. The gate runs before the ship.

Sixty seconds, next ship

That's the whole trick, honestly. Sixty honest seconds before you ship, and your agents stop performing confidence and start earning it.

Quick one before you go: do you like this page in dark mode?

You can change your mind. No big drama.

[ comments ]

guests welcome. members show first in the list

no comments yet. start the thread_