# When Human Brainstorming Beats LLM-Assisted Ideation

Charlotte Higgins · August 24, 2026

> When Human Brainstorming Beats LLM-Assisted Ideation. Four engineers thinking aloud together produce roughly half the unique ideas th...

| Takeaway | Detail |
| --- | --- |
| Live group brainstorming loses to solitary work on both quantity and quality | Four engineers thinking aloud together yield roughly half the unique ideas those same four produce silently apart (Diehl and Stroebe, 1987); the mechanisms are turn-taking queues, embarrassment about voicing half-formed ideas, and unequal contribution levels |
| Osborn's volume doctrine set the benchmark the group format then failed to meet | Osborn's BBDO team ran 7 sessions within one month totaling 225 ideas — roughly 32 per session — on the bet that sheer idea volume raises the odds of valuable ones (Osborn, 1953) |
| LLM assistance flips the trade-off: higher individual novelty, lower collective diversity | GPT-4 assistance lifted individual novelty by up to about 10 percent while making everyone's output measurably more alike (Doshi and Hauser, Science Advances, 2024) — example-induced fixation reinstalled at scale |
| Exposure order, not tool choice, decides the outcome: hints early, intervention at fixation, structure at convergence | Co-design across five in-person workshops with 28 participants (FAccT '26) yielded guidance that AI offer hints, not solutions, during early ideation, initiate interaction only when participants face fixation or saturation, and absorb tedious process tasks rather than the ideation teams value doing themselves |

Four engineers thinking aloud together produce roughly half the unique ideas those same four generate silently apart. Diehl and Stroebe's 1987 finding refuses to die because it indicts the session's mechanics, not its people: turn-taking makes ideas queue, embarrassment keeps half-formed notions unsaid, and unequal participation lets loud voices crowd out quiet ones. Even brainstorming's founding arithmetic assumed volume was the point — Alex Osborn's BBDO team logged 225 ideas across seven sessions in one month, roughly 32 per session.

Then came the 2024 inversion. Doshi and Hauser reported in Science Advances that GPT-4 assistance lifted individual idea novelty by up to about 10 percent — while making everyone's output measurably more alike. Read from a design-for-manufacturing desk, the model removes brainstorming's oldest bottleneck, turn-taking, and quietly reinstalls its oldest failure mode: example-induced fixation, where early machine suggestions anchor every subsequent pass.

So exposure order, not tool choice, decides whether a team banks the throughput gain or pays the homogenization penalty. A Shah-metric scorecard names the exact conditions under which the live human hour still earns its place, and a 2026 FAccT study — five workshops, 28 participants — supplies the sequencing: hints before solutions, intervention at fixation or saturation, structure at convergence, machines on tedious process tasks while humans keep the ideation they value.

![Winding dirt footpath climbing through misty pine forest](https://static.mm-ais.com/article-images-ai/when-human-brainstorming-beats-llm-assis-ai-bf8358e1.jpg)
Winding dirt footpath climbing through misty pine forest

## Turn-Taking Is the Tax

A live brainstorm runs on a queue: one mouth, several listeners, and everyone else generating into a mental buffer they never fully unload. That queue is the tax, and it is priced in the exact currency this guide scores. Fix the accounting first. According to Shah, Vargas-Hernandez and Smith's 2003 framework in Design Studies, an idea set is scored on four axes at once: novelty weights the leaf nodes of a concept genealogy tree by depth, so a deeper departure from known solutions scores higher; variety counts distinct occupied branches at each hierarchy level, from physical principle down to embodiment; quality is a feasibility-times-value product; quantity is a raw count. Ideas-per-hour is one axis of four — maximizing it while collapsing the other three wins nothing.

The biggest line item is production blocking. Because only one person speaks at a time, waiting suppresses generation: according to Diehl and Stroebe's 1987 experiments, four-person interacting groups yielded roughly half the unique ideas of the same four people working silently apart, and their follow-up attribution experiments pinned the largest share of that loss on blocking itself. The same research program found solitary sessions produced superior-quality ideas, with embarrassment about voicing half-formed thoughts and unequal member contribution covering most of the remaining gap (per Charles Leon's summary of the Diehl-Stroebe findings).

Blocking is not the only charge. According to Jansson and Smith's 1991 experiments, showing designers an existing example solution before ideation reduces the novelty of everything produced afterward — design fixation. An LLM suggestion is exactly such an example, delivered with more authority than a foam mockup: one engineer pastes the model's concept into the shared document, and the whole team now orbits the model's training-distribution center. Current design-tool guidance reaches the same conclusion from the human side — AI should offer hints rather than solutions during early ideation, and should intervene only once participants are already fixated or saturated (arXiv:2604.27997v1).

What the model buys back is the queue. Parallel asynchronous querying lets each engineer generate at personal fluency — roughly one idea every one to two minutes in short bursts — with zero turn-taking. Throughput climbs without any change in individual creativity; the gain is arithmetic, not cognitive.

The model's own fee lands on variety. Autoregressive sampling concentrates probability mass on common solution archetypes, so many independent queries converge on the same handful of concepts unless deliberately diversified through varied prompts, temperature settings, and best-of-N selection. Diversification is a demonstrable lever: according to Straub, Khan, Jay, Cabral and Linde's December 2025 multi-agent study, persona choice alone shapes which idea domains agents explore, and curated personas produce both depth and cross-domain coverage (arXiv:2512.04488v2).

Synthesis, and the claim the rest of this guide tests: the net effect of adding an LLM equals blocking removal — positive on quantity — minus fixation and homogenization — negative on variety. Strict dominance over the whiteboard would require winning all four Shah axes simultaneously; the published record shows a trade, not a win. Which side of the ledger dominates turns on exposure order — whether anyone sees model output before the manual pass — and that is precisely the variable the decision framework ahead manipulates.

| Mechanism | Shah axis hit | Evidence and magnitude |
| --- | --- | --- |
| Production blocking (live room) | Quantity down | Diehl and Stroebe (1987): four-person interacting group yields roughly half the unique ideas of the same four people apart |
| Apprehension and unequal contribution | Quality down | Diehl-Stroebe program: solitary sessions yield superior-quality ideas; embarrassment and uneven participation explain the residue |
| Design fixation (any shown example) | Novelty down | Jansson and Smith (1991): pre-ideation examples reduce subsequent novelty; an LLM paste acts as the example |
| Turn-taking removal (solo plus LLM) | Quantity up | Solo fluency of roughly one idea every one to two minutes in short bursts, zero queue |
| Autoregressive homogenization | Variety down | Queries converge on common archetypes; countered by varied prompts, temperature, best-of-N, persona curation (arXiv:2512.04488v2) |
| Net effect | Sign set by exposure order | Humans diverge first, model expands second; keep the live room only when feasibility judgment or buy-in is the deliverable |

![Turn-Taking Is the Tax — When Human Brainstorming Beats LLM-Assisted Ideation](https://static.mm-ais.com/article-images-pixabay/when-human-brainstorming-beats-llm-assis-7ea20d89.jpg)

## The Published Ledger

Read the ideation literature as a two-column ledger and one entry lands in the debit column every single time: collective variety. The credit column shifts from study to study — sometimes fluency, sometimes individual novelty — but the debit never clears. As of 2026, the load-bearing entries still fit in one table, and none of them shows strict dominance.

The oldest entry predates large language models entirely. According to Girotra, Terwiesch and Ulrich, writing in Management Science in 2010, virtual asynchronous groups produced roughly 50 percent more ideas per participant than equivalent face-to-face groups, with higher average quality and more top-rated ideas. For anyone trained on Shah's metric battery, this is the control condition that keeps getting skipped: removing synchrony alone — no AI anywhere — already buys the throughput teams now credit to chatbots. An evaluation of LLM assistance that omits this baseline is measuring the medium, not the model.

The AI-era entries then split cleanly along metric lines. According to Doshi and Hauser in Science Advances (2024), across 293 professional writers, GPT-4 best-of-five assistance raised individual novelty ratings by up to about 10 percent, with the largest gains among the least-creative third — while pairwise similarity across all stories rose enough to shrink collective diversity by double digits. Assistance lifted the floor of the distribution while pulling the corpus toward its center: the canonical throughput-up, variety-down result.

According to Si, Yang and Hashimoto at Stanford (2024), the gains survive the harshest test available: more than 100 NLP researchers scored research-proposal ideas in blinded expert review, where LLM-drafted ideas came out significantly more novel (p below 0.05) but slightly less feasible, and reviewers distinguished AI-drafted from human-drafted ideas barely better than chance. Two readings matter for a design team. Blinding removed any anti-AI penalty, so the novelty signal is real. And feasibility — the axis on which concept screens live or die — did not improve.

Two further entries close the escape hatches. According to Anderson, Shah and Kreminski (ACM Creativity and Cognition, 2024), participants co-developing story premises with GPT-4 produced idea sets measurably less novel and less diverse than unassisted controls on embedding-distance measures — counter-evidence arrived at with a cheaper operationalization than rater panels, yet pointing the same direction as the human-scored studies. And according to Vinchon and colleagues in the Journal of Creative Behavior (2023), ChatGPT matched or exceeded human brainstormers on fluency but trailed on flexibility, the spread of ideas across categories, in a non-engineering, think-aloud setting — so the variety deficit is not an artifact of design-school samples.

The myth this ledger retires is the comfortable one: that LLM-assisted ideation strictly dominates the whiteboard session — more ideas per hour at equal novelty, variety, and quality. Five studies, three task domains, fourteen years apart, and not one row shows dominance. Every row prices a trade, and the currency is variety.

| Entry | Setting | Credit column | Debit column | Ledger reading |
| --- | --- | --- | --- | --- |
| Girotra, Terwiesch & Ulrich, 2010, Management Science | Virtual asynchronous vs. face-to-face groups | Roughly 50% more ideas per participant; higher average quality; more top-rated ideas | Synchrony itself, removed by design | Throughput gain predates AI |
| Doshi & Hauser, 2024, Science Advances | 293 professional writers; GPT-4 best-of-five | Individual novelty up to ~10%; largest lift in least-creative third | Collective diversity down double digits as pairwise similarity rises | The canonical trade |
| Si, Yang & Hashimoto, 2024, Stanford | 100+ NLP researchers; blinded expert review of proposals | Novelty significant at p < 0.05 | Feasibility slightly lower; AI detection near chance | Novelty survives scrutiny; feasibility does not |
| Anderson, Shah & Kreminski, 2024, ACM Creativity and Cognition | Story premises co-developed with GPT-4 | Assisted output volume | Novelty and diversity on embedding-distance measures | Counter-evidence, convergent operationalization |
| Vinchon et al., 2023, Journal of Creative Behavior | Think-aloud brainstorming vs. ChatGPT | Fluency matched or exceeded | Flexibility across idea categories | Deficit generalizes beyond engineering |

Read the table column-wise before choosing a mode. If the deliverable is raw idea count, individual-plus-model wins — and the 2010 row shows individual-plus-asynchrony captured most of that gain before models existed. If the deliverable is variety or feasibility judgment, four of five rows flag a loss, which is precisely the case for keeping humans in the loop ahead of the model. How to act on that split is the next section's job; this ledger's job is to make sure nobody cites these five studies as proof the live room is dead.

![The Published Ledger — When Human Brainstorming Beats LLM-Assisted Ideation](https://static.mm-ais.com/article-images-pixabay/when-human-brainstorming-beats-llm-assis-71e98be6.jpg)

## Five Modes, One Winner

Run all five credible ideation modes through one scorecard and the ranking is not close: the humans-first hybrid takes the composite, LLM-solo takes exactly one column, and the mode most teams instinctively try first — pasting model output onto the projector before anyone thinks independently — finishes dead last. The card below applies the standard 1-to-5 rubric anchored to the effect sizes in the ledger above; on every column, 5 is favorable, so the fixation-risk cell is scored inversely (5 means the mode best resists anchoring).

| Mode | Ideas per person-hour | Shah novelty | Shah variety | Quality–feasibility | Fixation risk | Team buy-in |
| --- | --- | --- | --- | --- | --- | --- |
| Live group brainstorm | 1 | 3 | 2 | 4 | 3 | 5 |
| Nominal group (silent individual generation, then pooling) | 3 | 3 | 4* | 3 | 4 | 3 |
| LLM-solo ideation | 5 | 3 | 2 | 2 | 2 | 1 |
| LLM-seeded live group (model output shown first) | 2 | 2 | 1 | 3 | 1 | 2 |
| Humans-first hybrid (silent divergence → LLM expansion → human-only convergence) | 4 | 4 | 4 | 5 | 3 | 4 |

*Best among human-only modes; the hybrid matches it only because its divergence phase is also silent. Read the winners off the columns: LLM-solo owns throughput, the nominal group owns variety and fixation resistance among human-only options — a lineage Lucidchart formalizes in its catalog of twelve structured techniques, where round-robin and brain-netting both enforce silent, independent generation — and the hybrid owns novelty, quality-feasibility, and the composite. The seeded room loses outright on arithmetic, not taste: it pairs the worst fixation score on the card with no throughput edge over the hybrid, because a room reading aloud from a shared prompt queue pays the turn-taking tax quantified above while inheriting the model's clustering. Equal-weighting all six columns, the composites run: hybrid 24, nominal 20, live 18, LLM-solo 15, seeded 11.

The headline verdict, stated numerically: for a mechanical design team whose deliverable is a feasible concept portfolio, the humans-first hybrid wins the composite score; LLM-solo wins only when the deliverable is raw candidate volume on a well-characterized subsystem under a deadline shorter than 24 hours. Two thresholds flip the choice. When feasibility uncertainty is high — new physics, regulatory exposure, unproven manufacturing processes such as an unqualified powder-bed cycle — weight the quality and variety columns heavily and keep live human convergence, because no prompt sweep can certify a manufacturable cross-section. When the concept space is mature incremental redesign — a cast bracket in a qualified alloy, a gate-review refresh — the variety penalty is cheap and LLM-heavy modes dominate.

| Mode class | Marginal cost | Elapsed latency | Binding constraint |
| --- | --- | --- | --- |
| LLM-solo sweep | Single-digit dollars of API spend for a 500-prompt run | Minutes | None worth naming |
| Humans-first hybrid | Same single-digit API spend, plus one convergence hour | Roughly a day | Engineer attention at convergence |
| Any live-room mode (live, nominal, seeded) | Zero API spend | Up to a week of calendar lead time to align five cross-functional schedules | Scheduling latency and facilitator quality |

That last row reframes the debate: money is irrelevant at these magnitudes, so latency and attention are the binding constraints on mode choice. Facilitation load compounds it — according to Substack commentary on running these sessions, "you also need to have a good facilitator within the group," a scarce skill you rent by the hour in every live format.

To reproduce the card yourself: rate each mode one to five on each Shah axis anchored to the published effect sizes, multiply into a weighted composite where variety carries double weight for exploratory programs, and declare the winner from the arithmetic rather than preference. The weighting is not arbitrary — according to the canonical definition on Wikipedia, brainstorming's core outputs are "the volume and variety of ideas," so any scorecard that prices volume but not variety misstates the trade. The myth this kills: the whiteboard is not dead; it is demoted to a convergence instrument, and the genuinely obsolete artifact is the model-seeded room.

Next action: before booking any room, compute the composite for your program's weights. Book the hour only if quality-feasibility and buy-in outweigh throughput in your arithmetic.

![Five Modes, One Winner — When Human Brainstorming Beats LLM-Assisted Ideation](https://static.mm-ais.com/article-images-pixabay/when-human-brainstorming-beats-llm-assis-974366a9.jpg)

## What the Data Doesn't Tell You

Before you paste the hybrid sequence into a mechanical-design brief, audit what the published record actually measured. The strongest LLM-versus-human results come from creative-writing prompts and research-proposal tasks — domains where a concept is a paragraph, not a part with tolerance stacks. No large-sample peer-reviewed study yet scores CAD-ready engineering concepts on all four Shah metrics. The two closest engineering-flavored candidates a literature search surfaces — a biological-inspiration-versus-brainstorming novelty comparison, and a theory-driven experiment scoring novelty, feasibility, and value — returned only ResearchGate security pages on retrieval, so verification stops at the abstract. Importing these numbers to brackets and gearboxes is extrapolation, not citation.

The judge mismatch compounds the domain mismatch. Novelty and quality ratings in these experiments came mostly from crowdworkers or mixed panels. Engineering quality turns on manufacturability judgments — will this wall thickness survive injection molding, does that undercut require a side-action — that raters without design-for-manufacturing experience cannot make. A crowdworker scoring "quality" on a gearbox sketch is rating plausibility, not producibility. For our purposes, published quality scores may sit closer to noise than signal.

Model drift dates everything else. Every cited experiment ran on 2023-to-2024-era checkpoints; by 2026 the frontier differs, and both the throughput gain and the homogenization penalty may have shifted in either direction. Treat every published effect size the way you treat a material datasheet tied to a specific lot — a dated measurement, not a constant.

There is also a metrology problem inside the scorecard itself. Shah novelty and variety scores depend on where the analyst draws genealogy-tree boundaries and how deep the decomposition goes. Two analysts scoring the same hundred-concept set can diverge substantially, and most papers do not report inter-rater reliability for the tree itself. Part of the variety deficit the literature keeps finding could be an artifact of tree-drawing convention rather than a property of the mode.

And the scorecard omits what live rooms actually manufacture: shared mental models, organizational buy-in, and early surfacing of the resident skeptic's objection. According to a preliminary arXiv report (arXiv:1308.4978), walking brainstorming even helps a team sustain mental energy across a session — a benefit no Shah metric prices. A mode can lose on ideas per hour and still be the right meeting to hold when alignment, not idea count, is the deliverable. That is exactly the carve-out the sequencing rule preserves; nothing here rescues the dead-whiteboard reading either, because the ledger above records a trade, not a win.

Finally, one risk no single-session laboratory study can detect: feedback loops. If the whole industry adopts LLM ideation, future training corpora fill with today's model output, potentially compounding homogenization across model generations. Today's homogenization penalty reflects models trained mostly on human text; tomorrow's models train partly on model text.

| Evidence gap | What the literature measures | What a bracket-and-gearbox program needs | How to treat the number |
| --- | --- | --- | --- |
| Task domain | Creative writing, research proposals | CAD-ready mechanical concepts | Directional only |
| Rater pool | Crowdworkers, mixed panels | Design-for-manufacturing judgment | Possible noise |
| Model vintage | 2023–2024 checkpoints | The checkpoint you deploy in 2026 | Dated measurement |
| Scoring protocol | Analyst-drawn genealogy trees | Inter-rater reliability on the tree | Fragile comparison |
| Deliverable priced | Ideas per hour, four Shah metrics | Buy-in, shared models, surfaced objections | Outside the metric |
| Training corpus | Predominantly human-written text | Corpora increasingly model-written | Unmeasured compounding |

Before citing any ideation paper in a design review, run a three-point audit: confirm the task involved physical artifacts, confirm raters held design-for-manufacturing experience, and confirm the paper reports reliability for the genealogy tree itself. If any answer is no, label the figure directional and size your own pilot before committing the schedule to it.

## Worked Case

Put the ledger on a real part before trusting it. Take a five-person team — two mechanical engineers, one manufacturing engineer, one electrical engineer, one industrial designer — tasked with generating mounting concepts for an EV battery-tray bracket inside a single 90-minute window in a spring 2026 program review. Arm A runs the classic shape: 60 minutes of live brainstorming plus 30 minutes of cleanup. Arm B runs the humans-first sequence: 30 minutes of silent individual sketching, then 30 minutes of LLM expansion seeded only with the team's own sketches, then 30 minutes of human-only convergence. Same people, same prompt, same clock.

Arm A computes the way the blocking literature predicts. Five engineers working alone at typical individual fluency would produce roughly 55 nominal ideas; turn-taking in the room collapses that to about 25 to 30 unique concepts once duplicates merge. For scale, Osborn's own BBDO team averaged about 32 ideas per session across seven sessions in one month, according to *Applied Imagination* as recounted in Beagatchalian's essay "Cooking Up A Brainstorm" (Medium, September 15, 2022) — the live room has sat in this band since the 1950s. Assign the arm illustrative Shah scores: novelty mean 2.4 out of 5 on leaf-weighted scoring, variety spanning 2 of 6 physical-principle categories, and 3 concepts clearing a feasibility-times-value screen of at least 12 out of 25.

Arm B changes the arithmetic at the seams. The silent half-hour yields about 50 unique human ideas — nearly double the live room, with no queue to unload into. The LLM expansion pass adds roughly 80 candidates, of which about 25 percent land inside existing clusters and are discarded on embedding or genealogy dedupe, leaving about 35 genuinely new leaf nodes. Human-only convergence screens the pooled set to 6 strong concepts. Totals: near 85 unique candidates and about 110 raw ideas in the same 90 minutes.

Read the composite honestly: Arm B takes all four axes in this projection — quantity about 110 versus 75 raw, novelty 3.1 versus 2.4, variety 4 versus 2 principle categories, quality 6 versus 3 strong concepts. The mechanism matters more than the sweep: the top survivors originate in the human silent phase and are merely stress-tested by the model, which preserves feasibility instead of trading it for fluency. Had the deliverable been cross-functional feasibility judgment or stakeholder buy-in rather than idea count, the canonical rule keeps the live room; a bracket sprint needs concepts, so the hybrid takes it.

Now the falsifying arm. Run the identical session as Arm C with LLM output shown first and predict the flip: the team anchors on the model's archetypes — for a bracket, flanged bolt tabs, adhesive bonds, welded bosses — so raw counts stay high because generation is cheap, but variety collapses back toward 2 categories, matching Arm A's worst axis. Tool presence is constant between Arms B and C; sequencing is the only variable. The model amplifies whatever diversity the humans bring it and flattens whatever diversity it preempts.

Treat every arm-level number above as an illustrative projection anchored to the published effect sizes, not a measurement. The validation step is cheap: timestamp and origin-tag every concept as human, LLM, or hybrid from minute zero, compute your own Shah scores on the tagged pool, and re-run the A/B against each new model generation — the anchoring failure in Arm C will not stay calibrated to old checkpoints. Skip the tags and you cannot tell a sequencing win from a tool win.

| Metric | Arm A (live) | Arm B (humans-first) | Verdict |
| --- | --- | --- | --- |
| Raw ideas | ~75 | ~110 | B, +~35 |
| Unique candidates | 25–30 | ~85 | B, ~3x |
| Novelty mean (leaf-weighted, /5) | 2.4 | 3.1 | B |
| Variety (of 6 physical principles) | 2 | 4 | B |
| Strong concepts (feasibility x value >= 12/25) | 3 | 6 | B |

## How to Choose Well

"The whiteboard is dead" is the wrong moral to draw from the throughput numbers. The published record shows a trade, not a win, so the operative question is never which mode wins but what the session owes the program. Answer that, and five rules collapse into one decision tree you can run before booking any room.

Rule 1 — humans before the model, always. Silent individual generation comes first; nobody sees model output until the team's own round has closed. Example-exposure anchoring is the failure mode: once engineers see the model's candidates, their next proposals regress toward those archetypes, converting the model from an expander of the space into an attractor inside it. According to the collaboration-mode study, whose Result 2 finds that collaboration mode shifts the diversity of idea generation, ordering is the variable that carries the diversity effect — break the sequence and you forfeit the very variety premium that justified manual ideation.

Rule 2 — match the mode to the deliverable. If the deliverable is ideas per hour on a well-characterized subsystem — fastener patterns, thermal stacks, cable routing — cancel the live session and run asynchronous LLM-assisted individual ideation; the async hybrid wins that assignment outright. If the deliverable is feasibility judgment or cross-functional buy-in, the live manual brainstorm wins despite its lower count, because manufacturing, electrical, and industrial-design constraints surface as arguments, not list items. The taxonomies agree these are distinct trades: according to EdrawMind, individual and collective sessions call for different technique sets, and according to Break Out of the Box, methods split into individual, group, structured, and asynchronous families. Format inside the manual round stays flexible — according to the software-development brainstorming study, a preliminary finding holds that walking can carry an effective session.

Rule 3 — dedupe before you count. Pool human and LLM candidates together and cluster them on a genealogy tree or an embedding-similarity pass; count only distinct leaf nodes as novel. Six phrasings of one magnetic-latch concept are one leaf, not six. Raw lists that double-count archetype variants inflate apparent LLM productivity, because near-neighbor density is precisely where generative sampling clusters — score the pruned tree, not the paste buffer.

Rule 4 — cap and diversify the model's seeds. Hold the LLM to about three candidates per physical-principle branch and force prompt and temperature variation across branches. Uncapped sampling piles mass onto high-density modes; the cap converts that mode-seeking tendency into deliberate breadth. Treat it as a designed experiment: branches are factors, seeds are replicates, and you are purchasing coverage, not volume.

Rule 5 — re-baseline yearly. Every twelve months, re-run a fixed 90-minute A/B — manual versus hybrid, same brief, same scorers — against whatever the current frontier model is, logging your own Shah scores. Calibration decays faster than teams expect: according to the FAccT '26 program, the study anchoring this debate was accepted in April 2026 for presentation June 25–28, 2026 in Montreal, so the freshest peer-reviewed evidence is already aging against newer checkpoints. Published batteries also disagree — the DCC'14 line scores novelty, feasibility, and value — which is why internal consistency beats borrowed effect sizes.

| Gate | Trigger condition | Action | Why it holds |
| --- | --- | --- | --- |
| 1 · Sequencing | Anyone has already seen LLM output | Discard the exposure; rerun the silent individual round first | Anchoring flips the model from expander to attractor; Result 2 ties mode to diversity |
| 2 · Mode | Deliverable is ideas per hour on a well-characterized subsystem | Cancel the live session; run asynchronous LLM-assisted individual ideation | EdrawMind: individual and collective sessions require different technique sets |
| 2 · Mode, alternate | Deliverable is feasibility judgment or cross-functional buy-in | Keep the live manual brainstorm despite its lower count | The DCC'14 battery scores feasibility as its own axis, not a byproduct of count |
| 3 · Counting | Combined human-plus-LLM list is ready to score | Cluster on a genealogy tree or embedding pass; count distinct leaf nodes only | Duplicated archetype variants inflate apparent LLM yield |
| 4 · Sampling | Prompting the model for candidates | About 3 per physical-principle branch; vary prompt and temperature across branches | Converts mode-seeking into deliberate breadth |
| 5 · Currency | Twelve months elapsed or the frontier model changed | Rerun the fixed 90-minute A/B, manual versus hybrid; log your own Shah scores | Every effect size binds to a dated model generation |

Run the gates top to bottom; the first matching row assigns your mode, your format, and your scorecard.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | Run the silent solo round first: before anyone opens a chat window, have every participant generate ideas alone in writing — no speaking, no shared screen. | Diehl and Stroebe showed thinking aloud together yields roughly half the unique ideas of working apart, because turn-taking queues ideas and embarrassment suppresses half-formed ones. Solo-first banks the unanchored baseline. |
| 2 | Pool and dedupe the solo lists before any model sees them — merge duplicates, note who contributed what, and lock the document. | This creates your human-only reference set, so you can later tell whether LLM assistance genuinely raised novelty or merely made everyone's output more alike, the Doshi and Hauser homogenization effect. |
| 3 | Only after the pool is locked, bring the model in with hints, not solutions: ask for categories, provocations, and missing angles rather than finished idea text. | The FAccT '26 workshop guidance is explicit — AI should offer hints during early ideation. Full solutions handed out early reinstall example-induced fixation, where the first suggestion anchors every pass after it. |
| 4 | Set a fixation trigger: instruct whoever operates the model to intervene only when the team stalls — repeated near-duplicates, circular threads, visible saturation — not continuously. | Co-design across the five workshops found interaction initiated at fixation or saturation helps, while unprompted, constant suggestion erodes collective diversity. Timing of exposure, not tool choice, decides the outcome. |
| 5 | Assign the model the tedious process work — clustering, tagging, deduplicating, formatting the backlog — while humans keep generating and judging. | Workshop participants consistently wanted to keep ideation themselves and hand off process chores. Offloading structure preserves throughput without surrendering the creative core to the machine. |
| 6 | Spend the live cross-functional hour on feasibility filtering, not generation: score the pooled-and-expanded list against a Shah-metric scorecard (novelty, variety, quality), then convene engineering, design, and manufacturing to kill or advance candidates. | A live room loses to solitary work on raw idea count — Osborn's own volume doctrine was never met by the group format. Its defensible deliverable is cross-functional feasibility judgment, so reserve the session for exactly that. |

## Frequently Asked Questions

**How many fewer ideas do live brainstorming groups actually produce compared to the same people working alone?**

Diehl and Stroebe's 1987 experiments found four-person interacting groups yielded roughly half the unique ideas of the same four people working silently apart, with production blocking pinned as the largest share of that loss.

**What was the actual idea output behind Osborn's volume doctrine at BBDO?**

Osborn's BBDO team logged 225 ideas across seven sessions within one month — roughly 32 per session — on the bet that sheer idea volume raises the odds of valuable ones.

**How much does GPT-4 assistance lift individual idea novelty, and who benefits most?**

Across 293 professional writers, GPT-4 best-of-five assistance raised individual novelty ratings by up to about 10 percent, with the largest gains among the least-creative third, though pairwise similarity rose enough to shrink collective diversity by double digits.

**Did asynchronous virtual groups already beat face-to-face brainstorming before any AI tools existed?**

Girotra, Terwiesch and Ulrich reported in Management Science in 2010 that virtual asynchronous groups produced roughly 50 percent more ideas per participant than equivalent face-to-face groups, with higher average quality and more top-rated ideas.

**At what point in a session should an AI tool step in instead of offering solutions upfront?**

Co-design guidance from five workshops with 28 participants holds that AI should offer hints rather than solutions during early ideation and initiate interaction only once participants face fixation or saturation.

**How can teams keep everyone's LLM-assisted ideas from converging on the same handful of concepts?**

Varied prompts, temperature settings, best-of-N selection, and curated personas counter autoregressive convergence on common archetypes, with Straub, Khan, Jay, Cabral and Linde's December 2025 multi-agent study showing persona choice alone shapes which idea domains agents explore.

## Quick answers

| How does live group brainstorming compare to solitary work on idea output? | Four engineers thinking aloud together yield roughly half the unique ideas those same four produce silently apart, due to turn-taking queues, embarrassment about voicing half-formed ideas, and unequal contribution levels (Diehl and Stroebe, 1987). |
| --- | --- |
| What benchmark did Alex Osborn's original brainstorming sessions set? | Osborn's BBDO team ran 7 sessions within one month totaling 225 ideas — roughly 32 per session — on the bet that sheer idea volume raises the odds of valuable ones. |
| What trade-off did Doshi and Hauser find with GPT-4 assistance in 2024? | GPT-4 assistance lifted individual novelty by up to about 10 percent while making everyone's output measurably more alike, reinstalling example-induced fixation at scale. |
| According to the 2026 FAccT co-design study with 28 participants across five workshops, how should AI behave during ideation? | AI should offer hints rather than solutions during early ideation, initiate interaction only when participants face fixation or saturation, and absorb tedious process tasks rather than the ideation teams value doing themselves. |
| What variable decides whether a team banks the LLM throughput gain or pays the homogenization penalty? | Exposure order, not tool choice, decides the outcome — specifically whether anyone sees model output before the manual pass. |

Also worth reading: **Train your team on AI concept tools that actually ship**: [Train your team on AI](https://graftconcepts.com/blog/train_your_team_on_ai_concept_tools_that_actually_ship.php) · **How AI concept generation sharpens product-market fit in 2026**: [How AI concept generation sharpens](https://graftconcepts.com/blog/how_ai_concept_generation_sharpens_product_market_fit_in_2026.php) · **AI concept generation for startups on a tight budget**: [AI concept generation for startups](https://graftconcepts.com/blog/ai_concept_generation_for_startups_on_a_tight_budget.php)

### Related reading

- [Structured Brainstorming: Techniques to Boost Team Creativity](https://graftconcepts.com/blog/structured_brainstorming_techniques_to_boost_team_creativity.php)
- [Break Free from Solo Brainstorming: AI-Powered Concept Generation for Real-World Impact](https://graftconcepts.com/blog/break_free_from_solo_brainstorming_ai_powered_concept_generation_for_real_world_impact.php)
- [AI-Augmented Ideation Is Reshaping Product Development](https://graftconcepts.com/blog/ai_augmented_ideation_is_reshaping_product_development.php)
- [Generative Design: 35% Speedup Only Under Right Conditions](https://graftconcepts.com/blog/generative-design-35-speedup-only-under-right-conditions.php)
- [Prompt engineering for AI product concept generators](https://graftconcepts.com/blog/prompt_engineering_for_ai_product_concept_generators.php)
- [GenSolver v4 Cuts Bracket Mass 16%, Swapping to Ti-6Al-4V](https://graftconcepts.com/blog/gensolver-v4-cuts-bracket-mass-16-swapping-to-ti-6al-4v.php)

### Latest

- [Generative Design: 35% Speedup Only Under Right Conditions](https://graftconcepts.com/blog/generative-design-35-speedup-only-under-right-conditions.php)
- [Structured Brainstorming: Techniques to Boost Team Creativity](https://graftconcepts.com/blog/structured_brainstorming_techniques_to_boost_team_creativity.php)
- [AI-Augmented Ideation Is Reshaping Product Development](https://graftconcepts.com/blog/ai_augmented_ideation_is_reshaping_product_development.php)
- [Prompt engineering for AI product concept generators](https://graftconcepts.com/blog/prompt_engineering_for_ai_product_concept_generators.php)

Canonical: https://graftconcepts.com/blog/when-human-brainstorming-beats-llm-assisted-ideation.php
Markdown: https://graftconcepts.com/blog/when-human-brainstorming-beats-llm-assisted-ideation.php/index.md
