In our last post we made the argument for why multi-agent design works: a language model left alone gives you the modal answer, personas are a cheap way to sample from different regions of its distribution, and the value shows up when diversity is followed by adjudication. That was the theory. This is the practice manual — the actual mechanics and team topologies we run at Lone Star Observatories — followed by a live, clean-room experiment: two simple, universal design briefs, each given to a single directive prompt and to our iterative multi-team process. Every agent was a raw API call with no system prompt and no context of any kind, and the full unedited transcripts ship alongside this post.
Two results, up front, and one frame that holds them together: design is search over a solution space, and only one of these processes searches. A directive prompt buys you a point — one design, sampled from the center of the distribution: fluent, polished, and by construction the design everyone else asking the same question already has. Fluency is not design; the mode is where a search starts, not where it ends. The team buys you a map — the same space with its explored-and-rejected regions marked, its walls located, and its tradeoffs chosen on purpose. First result: the basin has real gravity — on the open brief, six independent samples landed in recognizably the same territory, down to the directive agent and the team converging on the same product name, and our own red-team agent said it to the team's face: "four people, one insight, dressed five ways. That's convergence, not creativity." Without structure, even conditioned samples fall back toward the average; a lone prompt never leaves it at all, because leaving requires the argument it doesn't contain. Second result: everything that makes the team's deliverable a design happened after the mode — five rejected regions with reasons marked on them, a wall nobody had noticed (weather), two ideas revealed to occupy incompatible coordinates so the synthesis had to choose, and an honest "this part of the map is unsurveyed." One process measures the prior. The other explores the space the prior sits in.
The frame: design thinking, mechanized
None of what follows is freestyle multi-agent enthusiasm — it is the classic design thinking loop, Empathise → Define → Ideate → Prototype → Test, run as a machine. Our squads execute it as Define → Ideate → Prototype → Test → Empathise → iterate-or-ship under a Facilitator who owns pacing and the iteration decision, with one deliberate twist: empathy comes around twice — as persona conditioning before anyone sketches, and as a formal radical-empathy phase after testing whose written output is explicitly allowed to contradict the original problem definition. The loop closes only when the Facilitator judges the definition has stopped moving.
The strategies underneath everything
Five principles drive how we do design work with Claude. Everything else is implementation.
Design is search, not generation
A brief has a solution space, not an answer. One agent samples it once, near the center. We place multiple samples in different regions, then select. Divergence and convergence are separate mechanical steps — never one prompt.
Condition the samplers, don't heat them
Diversity comes from identity, not temperature. Every member gets a name, age, culture, and MBTI type. Identity reliably moves the reading of a brief; as the experiment shows, escaping the idea-basin takes the next principle too.
Make disagreement structural
"Consider other perspectives" is a sticky note. A roster where perspectives are assigned — one agent allowed only to attack — is a meeting where they show up and argue.
The filesystem is the memory
Every phase writes files the next phase reads. That's what makes iteration real: pass N pushes off pass N−1 instead of restarting from the same floor.
The human sits at the gate
No mid-session approval gates. The team runs autonomously; the Board sees the finished package — and the dissent that survived. Human attention is spent on verdicts, not supervision.
The five topologies
"Multi-agent" is not one architecture. We run five distinct shapes, chosen by the kind of question being asked. Filled sapphire nodes are facilitators/orchestrators; open nodes are persona agents; brick nodes are adversaries; the dashed node is the human.
1 · Facilitated round-table
The Design Thinking Squad: 4–8 personas around a Facilitator running Define → Ideate → Prototype → Test → Empathise → iterate. Every phase ends in a rendered team discussion, in voice, by name. Best for: "what should this even be?"
2 · Pipeline with debate gates
The BA Team: Intake → Story Mapper → Analyst → Tester → Sober Engineer → Design Lead, with the whole team critiquing each deliverable before the next role begins. Best for: turning a settled direction into stories, acceptance criteria, and briefs.
3 · Panel-and-judge
When the answer is an artifact, don't debate — compete. Five separately-conditioned teams shipped five complete Vision Pro concepts (Copernicus, Galileo, Halley, Hypatia, Tycho); selection deferred to the Board. Best for: directions where anchoring on the first concept costs the most.
4 · Adversarial pair
One agent asserts; a second — with standing to say "not proven" — attacks. Used small (a red-team pass inside a session, as below) and large (verification before anything ships). Best for: claims that matter.
5 · Hierarchy
A Programme Office of standing personas commissions design-thinking runs, turns output into business cases, and spawns BA/UX/delivery teams with briefs. The org chart is genuinely stateful, not a fresh cosplay per session. Best for: programmes, not projects.
Choosing between them
Open question? Round-table.
Settled direction? Pipeline.
High-stakes artifact? Panel-and-judge.
Claim that matters? Adversarial pair.
Many workstreams? Hierarchy.
And underneath every one of them: the filesystem, carrying the argument between phases.
The mechanisms that make the magic
Topologies are the org chart. The actual search moves — the things that reliably pull a session somewhere a lone prompt doesn't go — are three mechanisms we use everywhere, and each is a way of sampling a different region of the space on purpose.
1 · Weird-persona injection. A professional design team samples professional-designer space, so we deliberately seat people who have no business being in the room. When five of our teams took an astrophotography app's desktop pages through a design charrette, every team's core (facilitator, UX lead, domain analyst) was flanked by a professional astronomer, a professional photographer, a novice, and a space-obsessed teenager — and every pass was critiqued by a design-snob review board: a fine-art painter, an infographic designer, and an Apple-school web designer, none of whom owed the design anything. The credits are traceable in the reports: the painter is why every page's final form treats the image as the center of gravity ("everything else is a mat around it" was, verbatim, "every fine-art-painter critique, every page"); the infographic designer contributed the honesty spine — every number carries its reference frame; every threshold is published and editable. The novice and the teenager are boundary probes: they ask for the thing professionals forgot to want and refuse the thing professionals forgot is confusing. A weird persona is a cheap coordinate transform — the same model, teleported to a far region of the space, reporting back what the design looks like from there.
2 · Identity swapping. The cheapest usability lab in existence: at the Test phase, the design team becomes the test users. The instruction, from a real session's facilitator, is one line: "Step out of yourselves. Be the tester. Be honest, including when it hurts." Agents who spent three phases advocating for the design are re-conditioned into the target personas from the test plan and made to meet their own prototypes cold. It works for the same distributional reason everything here works: an identity is a position in the space, and swapping identities moves the observer instead of the artifact — the designer-persona knows what the button is for; the user-persona it becomes genuinely doesn't, and says so, in character, with feelings. It's also what keeps persona conditioning from hardening into persona lock-in: nobody keeps one seat for a whole session. The Empathise phase then formalizes what the swap surfaced into the refined problem definition.
3 · The subtractive charrette (remove-a-component). Generative models accrete — ask for a refinement and you will overwhelmingly get additions. Deletion is a search operator the model almost never applies to its own output, so we force it structurally: in the charrette above, five teams passed each page around a Latin square — every team touched every page — and each pass ended with a hard rule: delete one fresh major component and hand the blank to the next team, who must decide what best deserves the reclaimed space. Thirty agents, thirty files, zero re-used deletions. The same class of component kept getting cut and never missed — activity logs, KPI walls, storage bars, standing settings forms — and the through-line the deletions revealed became design law: most "panels" are an intermittent need masquerading as a standing one; only things that change while you watch earn permanent pixels. The final report's summary is the mechanism's whole defense: "the page got quieter and more honest at every step, and never lost a capability — each deletion was a relocation, not an amputation."
Map these back to the framework and the shape of the whole system becomes visible: weird personas widen Ideate, identity swapping powers Test and feeds the Empathise return leg, and the subtractive charrette is what iteration looks like when you refuse to let it mean accretion.
The outputs we demand
A topology without deliverable discipline produces vibes. Every phase has a named artifact with a schema, and the artifacts are designed to be arguable with — they preserve disagreement and provenance rather than laundering everything into consensus prose. This is what one real session directory looks like:
The experiment
Claims like these deserve a test you can inspect, so we ran one under clean-room conditions: two simple, universal briefs; every agent a raw API call to the same model with no system prompt — the prompt is the entire context; persona prompts specify identity, role, and schema only; downstream agents receive prior outputs verbatim, not summarized; one pass, no cherry-picking. Full transcripts with per-call token counts: every agent's unedited output.
Brief 1 (open-ended): the first night
"Design an experience around a person's first night out with their first telescope."
The directive agent produced "First Light" — kit plus app, one promise: "you will see something breathtaking tonight, within 30 minutes of stepping outside." It is fluent and confident: daylight assembly with a finderscope-alignment checkpoint, one "hero object" instead of a menu, a deliberate 60 seconds of interface silence at the eyepiece, a second-session metric, even a quotable thesis ("the product's most important design decision is knowing when to disappear"). Hold that inventory in mind — every one of those moves is about to reappear, unprompted, in the blind samples below. Nothing here was found. It was recalled.
The four persona ideators, running blind, read the same brief four different ways:
"The first night is not an experience problem — it's a failure problem… Saturn's rings sell themselves. Frustration is what needs engineering."
Ingrid · first instinct
"Don't design the first night with a telescope. Design the first night with the sky, where the telescope earns its entrance."
Dayo · first instinct
"Underneath all of it is this fragile, embarrassing hope: I wanted to feel wonder tonight, and instead I feel stupid… The person's relationship with their own hope is the product."
Mercedes · first instinct
"Let's be honest about what this brief actually is: a churn problem wearing a romance costume."
Ken · first instinct
Notice what happened: the readings genuinely diverged — commissioning, expectation, shame, retention — but the ideas underneath kept converging: everyone wants a daylight rehearsal, a single guaranteed target, a designed second night. The directive agent, with no persona at all, had found most of the same moves. Five samples, one basin. The red team is the agent that noticed — and Fig. 6 shows what she did about it.
Ingrid
Daylight rehearsal — align the finder before dark
Date-keyed target card with a sketch of what you'll actually see
Comfort checklist · laminated fix sheets on the mount
"Design the second night, not the first"
Dayo
Court-the-sky-first sequencing; "blob appreciation" reframe
The Locked Cap
"clear skies are perishable; burning them on principle is design vanity"
Twin First Light (paired strangers)
four dependencies, zero in our control
Disappointment Postcard
Mercedes invented it too — "convergence, not creativity"
Mercedes
Protect hope, not features; whisper companion (audio, no screen)
One Guaranteed Win; re-entry ritual ("tell someone what you felt")
Postcard to yourself
duplicate; deferred payoff doesn't fix immediate churn
Ken
One-target-no-menu guarantee; funnel framing
48-hour hook; second-session KPI
Shareable proof-of-win badge
"Ken designed the funnel and forgot the dark"
Priya · red team
"Weather. Nobody said the word." — the cloudy first night
The Moon isn't always up — the guarantee evaporates
Scope of control unstated (app? packaging? hardware?)
Groupthink: "four people, one insight… all anecdote, zero research"
Synthesis → "First Light: a two-night, screen-dark protocol"
- Scope stated plainly: printed in-box kit + optional audio companion — what a telescope maker actually controls
- The Ken/Mercedes contradiction resolved: the phone exists before sunset and after pack-up only — audio and red-light print at the eyepiece
- Clouds get a designed path: a cloudy first night is a rehearsal night — "waiting is the hobby's first skill"
- The single target becomes a date-keyed ranked ladder (Moon → planet → double star → cluster) — the promise never rests on one object
- Daylight rehearsal as persuasion, not a padlock: "skip this and tonight will fail"
- Ten minutes of naked-eye orientation first — Dayo's instinct, minus his hostage mechanism
- 48-hour hook and second-session-within-7-days as the north-star metric
- And an admission: the expectation-gap thesis is anecdote — diary studies with 20 real first-nighters before scaling
Brief 2 (basic): a page about a star
"Design a web page that shows information about a single star — take Betelgeuse as the example."
The three drafters split on what the page is for. Ruth (ISFJ): the real visitor arrives from a search — "is Betelgeuse going to explode?" — anxious about looking foolish; answer their question first, in plain words. "Clarity is kindness." Beto (ENFP): this is a deathwatch — the page should leave you "small, lucky, and a little haunted." Anja (INTJ): comparators and honest uncertainty are the page — distance shown as a ±90-light-year band, mass as a range ("show the range, not a fake midpoint"), a live light curve because a static magnitude misleads, and a "How we know" section: "provenance is what separates reference from decoration." The facilitator named the real disagreement — what the first ten seconds promise: reassurance, awe, or rigor — fused all three into one hero, and killed Beto's "could explode before you finish reading this page" hook with the sharpest sentence of the run: manipulating the anxious reader we exist to reassure is "a trust debt no engagement metric repays."
Condition A · directive, one agent
"The design becomes the data" — a variability-pulsing portrait.
Betelgeuse
A dying giant, 548 light-years away.
The Great Dimming — scrollytelling; the page itself dims
Life & Death — supergiant stage, impending supernova
Sky Finder — find it in Orion tonight
Condition B · team synthesis
The question answered in the hero; uncertainty as first-class content.
This star will explode — sometime in the next 100,000 years.
Nothing dangerous. Everything wonderful. 🔊 BEET-l-jooz
The light in your eye left it around the time the Ming dynasty ruled. The band is the honest answer — parallax is hard for a star this bloated.
↻ flip: how we measured thisThe return leg, run clean-room: empathy, randomness, subtraction, and the sun
The two briefs above exercised only the Define → Ideate slice, and we said so. So we ran the rest of the loop under the same clean-room conditions. The two first-night concepts from brief 1, verbatim, met four identity-swapped user personas, each living a different real night: Marcus (34, two kids with twenty minutes of patience, clear with a half Moon), Eleanor (61, app-averse, completely clouded out), Priya (24, anxious perfectionist, moonless sky), Tyler (27, phone-native, under a parking-lot light dome). The directive concept was frozen — one pass is its definition — and tested once. The team concept entered the house loop: each iteration ran the empathy round (four blind persona nights), injected one outside voice from a rotating weird deck, and ended in a facilitator revision with a mandatory deletion and a refined definition allowed to contradict the original. And the iteration count was tied to the sun: a fixed 420-second night clock. Nobody chose the number of iterations. The clock did.
Eleanor's clouded-out night is the whole argument in one persona. Against the directive design she wrestled a QR code at three distances and received: "Cloud cover 100%. No targets visible tonight." Full stop. Her verdict — yes to a second night, "but because I've waited since I was ten and the clouds will break, not because this design earned it; it planned one perfect night and forgot that weather, knees, and grandmothers exist." Against the team design she found the Cloudy Night path on paper, ran the daylight rehearsal anyway, and had her first light in the afternoon — a pigeon on a chimney pot: "The finder agreed with the eyepiece because my hands made them agree… The clouds took the sky, not that."
An honest caveat before the score: all twenty simulated nights ended in SECOND NIGHT: yes — synthetic users are agreeable, and we won't pretend otherwise. The signal lives in the clause after the yes. The directive design's yeses arrive despite it; the team design's arrive because of named mechanisms: "only because the fix sheet made failure survivable" — "the design let me fail privately, fix it myself, and still win" — "when it broke, the paper let me fix it in front of my kids instead of failing in front of them."
And the loop ate its own darlings, on evidence. v3's deletion was the Whisper Companion — the audio guide the original synthesis was proudest of — with the ruling that "its best feature — the silence — becomes the default state of the whole design." v4 then deleted the Star-Hop panel v3 itself had just created: the loop correcting its own previous iteration, something a frozen artifact cannot do by definition. The weird deck earned its seats line by line — the lighthouse keeper's assemble it blind once, the kindergarten teacher's fire drill ("introduce heroes before the crisis"), the magician's sealed envelope ("people return to collect promises"), the controller's pre-committed aborts ("a go-around never feels like quitting, because we rehearsed it as success"). The shipped v5 opens: "Paper-only, any-sky, two-night protocol… Every session ends on a printed card, never on a shrug." Cost of the whole return leg: 28 calls, 105,085 tokens, one 420-second night — every addition tagged with the persona or voice that earned it.
Scoring it honestly
Total tokens, measured (relative to directive)
Wall-clock time, measured (seconds)
Directive (1 agent)Team (4–6 agents)
| Directive (1 agent) | Team (4–6 agents) | |
|---|---|---|
| Brief 1: core ingredients | Daylight setup · one target · silent moment · 2nd-night metric | Same ingredients, reached independently |
| Brief 1: weather | Assumes a clear night — the promise breaks on clouds | Named as blind spot; cloudy-night path designed |
| Brief 1: screens in the dark | App at the eyepiece (red mode + photo prompt) | Contradiction named, resolved: phone banned after sunset |
| Brief 1: assumptions | Shipped unexamined | "All anecdote, zero research" — validation step added |
| Brief 2: signature move | Variability-pulsing page | Same pulse, invented independently — the modal move |
| Brief 2: trust layer | None | Answered question · uncertainty bands · "How we know" · hook killed for cause |
| Killed ideas w/ reusable reasons | 0 | 5 (+4 mandatory deletions in the return leg) |
| Eleanor's clouded-out night | "It planned one perfect night" — dead on arrival | Cloudy path + daylight first light — "the clouds took the sky, not that" |
| Second-night yeses (all 20 said yes) | 4 — despite the design | 16 — because of named mechanisms |
| Iterations | 1, by definition | 4 — the count chosen by the sun |
| Self-correction | None possible | v4 deleted what v3 created |
| Provenance | None | Every element attributed, incl. the weird voices |
A directive prompt buys you a point. The team buys you the map.
What we are not claiming
Same disclaimers as last time, because they still hold. This is orchestration, not new capability — every call in every condition hit the same model. Two briefs plus one return-leg run is an illustration, not a study, and the team condition consumed an order of magnitude more inference. The simulated users all voted yes — synthetic testers are agreeable, which is why we scored the reasons, not the verdicts, and why real diary studies remain in the shipped concept's own plan. Prototyping is the one phase of the loop not demonstrated here. And the directive baseline is not a strawman — which is precisely the hazard worth naming: its outputs are fluent, plausible, and heavily overlapped with the team's instincts, and nothing in the artifact discloses that no search happened. The unexplored average arrives wearing the confidence of a finished design. That is why "it looks done" is the least trustworthy property a design can have.
One more disclosure, in the spirit of this series: our first attempt at this experiment ran as subagents inside our own workbench, and a diagnostic probe showed those agents were inheriting our organization's full project context — they were not the blank agents the comparison required. We threw that run away and rebuilt everything as raw, context-free API calls. The failure mode is worth naming because it is the multi-agent version of the modal trap: your agents are only as independent as the context you didn't notice they share.
We keep it dark out there by resisting the easy, probable, beige answer. Sometimes that takes six agents and an argument. Sometimes it takes one agent and the discipline to leave it alone. Knowing which brief is which is the actual skill.
Full unedited transcripts, per-call token counts, and run metadata: transcripts.md. Every call in both conditions ran on claude-fable-5, 2026-08-11, single pass, no system prompt. Diagram palette validated for color-vision-deficiency separation and contrast in both light and dark themes.