The Swarm That Refused to Choose
The reviewer said keep. The final decision still selected nothing.
That happened in a test of Brainstorming Swarm V2, the skill now available on GitHub. Two discovery lenses examined an invented fairground wayfinding problem. Two approaches went through separate critiques. One survived a revision. The campaign was structurally complete.
But the proposed route needed shade. The supplied notes established a staffed location and a step-free route. They did not establish shade there. A good critique of the proposal could not turn that missing fact into evidence. The selected approach stayed empty, and the next step became a bounded check.
That is the behaviour we wanted from a swarm: more minds on a design, with a way to stop them from mistaking agreement for proof.
A favourable critique assesses an approach. It does not supply a missing fact.
Many lenses, one decision
A solo brainstorm can uncover a strong idea. A swarm can expose the idea to different pressures before anyone builds it. Failure, access, cost, maintenance, and lived experience do not ask the same questions. Neither should the agents.
Parallelism alone does not give you independent thought. If every agent sees the others' answers too early, the room converges on the first plausible story. If every leg sends a full transcript back, the person making the decision has to excavate it.
The skill gives each leg a bounded job and asks for a digest. Source facts get their own anchors. Inferences cite the facts they rely on. A reviewer states what a cited fact means for the proposal separately from the fact itself.
- 01FrameOutcome, locks, sources, stop pointOne owner
- 02RouteCheck the host before task material movesPreflight
- 03DiscoverDistinct lenses return bounded digestsParallel if availableAccessFailureCost
- 04ProposeCompeting approaches expose assumptionsParallel if available
- 05ChallengeA different reviewer keeps, revises, or killsCross-review
- 06DecideAudit each necessary fact before selectionOne synthesis
- 07Human gateApprove the design or leave it openStop
The run has seven stages:
- Frame: Write the outcome, constraints, allowed sources, privacy boundary, and stop point.
- Route: Check what this host can actually run. Record the model and effort controls that matter before sending task material.
- Discover: Give two or three distinct lenses the same bounded question. Validate each digest and its factual clauses before accepting it.
- Propose: Develop competing approaches, with assumptions, necessary conditions, risks, and effort visible.
- Challenge: Give each approach to a different reviewer. Keep, revise, or kill it. A material revision needs a new review.
- Decide: Use accepted work only. Check every fact the preferred approach needs. If a necessary condition is unknown, leave the selection open.
- Human gate: Present the design and stop. Design approval may lead to a PRD. The written PRD gets its own review before task planning.
The method requires a validation barrier between submitted work and the next dependent stage. The leg stays incomplete while its facts are checked. A validation event enters the journal first. Only then is acceptance recorded and dependent work allowed. The checker verifies the recorded order; it cannot prove the host obeyed that order at runtime. A neat digest is not accepted merely because it exists.
Try it on a problem you can check
Follow the First Run guide. Copy the whole skill directory into your agent's skill location. Keep SKILL.md, references/, and templates/ together. If your host has no skill loader, open SKILL.md with the agent and provide the linked files. The method does not grant an agent permissions it lacks.
For a first run, give it a small invented problem with supplied facts. The repository uses a civic library:
Use Brainstorming Swarm V2 to compare wayfinding options for an
invented civic library. It has two floors and a narrow entrance wall.
Consider visitor accessibility, staff maintenance, and cost. Use only
these supplied facts, label assumptions, and stop for my design decision.
Use independent agents if available; otherwise label the result solo_analysis.
Create a campaign folder with SCOPE.md, LEDGER.md, a campaign.json role map, an empty events.ndjson, and a digests/ directory. The scope says what the agents may read and do. The role map records who performed each leg. The ledger points to the current and earlier revisions. The event journal records validation and acceptance in order. The digests are the work product, not a pasted chat transcript.
From the copied skill directory, you can run the included structural checker:
python3 scripts/check_campaign.py path/to/campaign
Before accepting each leg, run python3 scripts/check_campaign.py path/to/campaign --leg g1 with that leg's actual ID and check each cited source yourself. Record validation while the leg is incomplete, then acceptance. Run python3 scripts/check_campaign.py path/to/campaign --events before dependent work. The full command above checks the completed campaign's links, files, shapes, and recorded event order. It cannot tell you whether a citation is true or two agents actually thought independently. You still read the evidence and make the call. If the host cannot spawn independent agents, the same method can run in solo_analysis, clearly labelled as one agent's analysis.
The mistake that changed the method
An earlier internal synthetic run made an attractive jump. A location was staffed, so the final choice treated it as shaded. An independent review vetoed that choice. The source did not say what the decision needed it to say.
We changed the method to name decision dependencies. A preferred route now has to list the conditions it needs, and each condition has to be marked supported, unverified, or contradicted against a direct source anchor. A favourable reviewer verdict cannot fill a missing source fact.
In the public corrected run, the first discovery digest was accepted, then revoked before any draft began. A later check found that it inferred shade from a place name. Its revision made shade explicitly unverified. A design also changed materially. Its earlier critique no longer counted, so a fresh critique followed. The final choice remained open because shade and full pilot cost were still unverified. An independent reviewer passed the corrected artifact trail and source reasoning. The public trial and review are there to inspect.
Discovery
Design and critique
The ledger points to current revisions. The journal keeps the order of validation and acceptance.
That trail is why the journal exists. If a critique can be invalidated by a revision, the old verdict must remain visible. If an accepted fact is corrected, anything built on it has to be reconsidered. The files let another reader see what changed and why.
How we made the public skill
We started with a product brief and a task plan, then built the portable contract in dependency waves. Separate agents worked on the written method, the structural checker, invented examples, and host guides. Reviewers tested the seams between those pieces: Can a missing judge look complete? Can an unsupported source claim survive a valid JSON shape? Does a revision leave its old verdict looking current? Failed cases changed the method and its checks.
We then ran the skill against invented source packets. The first fairground decision exposed the shade error. The corrected campaign preserved its revisions and stopped without selecting a design. That gave us evidence about the method's behaviour, while leaving host-specific claims open.
For publication, we assembled a fresh repository from reviewed files instead of carrying private development history into the public Git objects. Literal scans checked paths and credential patterns. An infrastructure lens and an identity lens checked for recognizable fingerprints that a token scan could miss. Their reports remain private. The public scrub summary describes the boundary and the checks.
What public means here
The skill is MIT licensed. It is a portable written method with templates, host guides, examples, and an optional checker. It is not a scheduler or a paid service.
Its host guides name Claude, Codex, Grok, and Muse as targets. The support matrix marks every exact host and model row candidate. The fairground exercise passed an artifact audit; it did not prove native child identity, actual concurrent computation, cancellation behaviour, or support on two host families.
You can use the method today with the capabilities your own host actually provides. Record when a run is serial. Record when it is solo. Stop when a required route cannot be verified. The useful output is a design decision with its uncertainty still attached.
The swarm did its job when it refused to choose.
[ comments ]
guests welcome. members show first in the list
no comments yet. start the thread_