Thank you, Azeem and Rohit. This piece, and the Karpathy council it builds on, directly shaped a tool I built, so thank you. The finding that councils show groupthink turned out to be the most useful thing I have read about combining multiple models.
What I took from your approach was the diagnosis. Decomposing answers into atomic "idea cards," clustering them, and blind judging their quality made something measurable: peer review followed by blending keeps the consensus while quietly dropping minority ideas. It also showed that the naive "best answer" picker preserves more uniqueness than the blend does. That reframed the whole problem for me.
Here is where I tried to resolve some of the shortcomings.
First, an explicit preservation protocol, built in. Before any synthesis, I build an idea ledger with provenance and rank each idea on its own merit, never on how many models raised it, so that a brilliant idea from one model can outrank a mediocre consensus one. Minority ideas are flagged as essential to include.
Second, a red team that rescues rather than only critiques. One pass hunts for valuable ideas raised by only one model that a consensus summary would otherwise drop, and then argues to keep them. This is a direct counter to the hidden profile problem.
Third, the survival metric is the product itself, not an afterthought. Every run reports unique versus consensus idea survival for both the baseline, which follows Karpathy's design, and my preserving variant, using the same inputs. Therefore the groupthink gap is visible on every run rather than merely assumed.
Fourth, genuinely decorrelated models. I ran a mixed Western and Chinese panel, namely Opus 4.8, Gemma, Kimi, and DeepSeek via OpenRouter and Ollama (for free local models). On a prompt asking for unconventional ideas, the baseline blend deleted exactly the unconventional ones, such as blockchain incentives, contribution art avatars, and curiosity gap designs, and most of these had come from the models that were not Western. The preserving pipeline kept all of them. On that run, the gap fell from 55% to 0%.
One boundary is deliberate. Preserved does not mean endorsed. Several of the rescued ideas are precisely the ones that most need human judgment, given the coercion and undue influence risk in my regulated domain. Therefore the tool surfaces them flagged with guardrails rather than smoothing them into the recommendation. The council widens the option space; however, a human still adjudicates.
I am also mindful that this is still early work, and the underlying evidence is one informal experiment. It is directionally convincing and consistent with human group dynamics; however, it is a signal to design against rather than settled science. I am grateful for the prompt to do exactly that.
I am now going to build an advisory board council with personas that Gianni suggested. It was only list to do already ….
> This piece, and the Karpathy council it builds on, directly shaped a tool I built, so thank you. The finding that councils show groupthink turned out to be the most useful thing I have read about combining multiple models.
Thank you!! This makes me quite happy! And very excited to see you build on this. Do put this out too when you're ready, I'd be quite interested to see it.
Very interesting follow-on to Rohit's experiments. Thanks for that!
With that said, let me quibble with your characterization of your red team's rescuing unique &/or valuable ideas as "a direct counter to the hidden profile problem." By this I mean to quibble with your characterization, not with the value added by the red teams' actions you've assigned. Those actions are a separate issue, and seem a good idea.
As explained in Rohit's exposition, I take it that the "hidden profile problem" is more analogous to the zero-day exploits found by Mythos. I.E. There are a number of factors, ideas in this case, existing independently of a coherent whole. In the zero-day case this is a whole exploit cobbled together from discrete weaknesses not exploitable in isolation; in this idea case the coherent whole is some composite idea or approach to a challenge or problem to which a solution is desired.
My point here is that in the hidden profile situation, the desired "idea" or approach is not necessarily contained in any one of the approaches suggested by a single LLM, and thus cannot simply be "rescued." Rather, it consists only of particles or components of some approach that might be greater than the sum of its' parts. It is highly desirable that such "compositions" emerge rather than lay fallow, but insofar as these compositions are inherently novel, similarly their emergence is novel and for all that, unlikely. Nor would it seem a rescue function is conceptualized in such a way as to reveal them. Am I mistaken to think of this as a second-order problem? Whether first- or second-order, whether fanciful or actual, this possibility would seem to merit investigation.
Summing up, what I'm suggesting here is that this hidden profile circumstance is a separate problem worthy of separate consideration since the dimension in which it exists is perhaps orthogonal to that of coherent "complete" single-LLM-sourced ideas.
Thank you, this is a sharp and fair correction, and I think you are right.
You have caught a genuine conflation in my wording. What the red team rescues are complete, single-source ideas: things one model stated whole that a consensus blend would drop. That is a real failure mode, and it is essentially what Rohit's experiment operationalised and measured (idea-card survival). However, you are right that the classic hidden profile construct is about something deeper. The desired answer is distributed across members as fragments, held whole by no one, and it only emerges if those fragments are pooled and composed. Rescuing intact ideas does not reach that, so my phrase "a direct counter to the hidden profile problem" overstated it. More precisely, what we counter is the loss of complete minority ideas, which is adjacent to, but not the same as, the composition problem you describe.
What makes your point even sharper for my setup is that the anti-confabulation guardrails actively work against emergent composition. I added a faithfulness check that flags any claim in the synthesis that is "not traceable to any single member." A genuine hidden-profile composite is, by definition, traceable to no single member. So, my current design would flag exactly the kind of emergent, greater-than-the-parts idea that the hidden profile problem wants surfaced and treat it as a possible fabrication.
There is a real tension between "do not invent" and "synthesise something no one said alone," and the composition problem lives precisely in that gap. So, I agree it is a separate, second-order problem, roughly orthogonal to the dimension we have been working in. Addressing it would need a different mechanism: a deliberate composition or cross-pollination pass that tries to assemble fragments across models into candidate composites, plus some way to judge whether a composite is genuinely valuable rather than merely novel. That is harder, partly because it pulls against the same convergence-avoidance and fidelity controls that keep the rest honest, and partly because there is no ground-truth ledger for an idea that did not previously exist. I had not separated these cleanly. It deserves its own investigation, and I am grateful for the nudge.
I am going toa dd this as a separate artifact and make it exempt from the “faithfulness” check since by design they aren't traceable to one model), so the synthesis stays grounded while I still see "what might emerge from combining across models”
Excellent. I look forward to hearing about your struggle with this. It seems to be "a worthy problem!" I confess I wouldn't know how to begin! The goal I can see, the path... not so much. And of course another problem with my anticipated "hearing of" your efforts and hoped for solution is this: a key characteristic of Substack is the on-rushing FLOOD of content! What I lack, & so far as I know does not exist, is a cogent way to manage these conversations! While this isn't the topic here, insofar as this actual topic is one of the most interesting to have come up recently, any mechanisms, tips, tricks, etc U know of that might help, I'm all ears!
Neat experiment and the callout on asymmetric information is really interesting. Did you try a version where the job of the picker is to facilitate? That is, find those places of asymmetry and have them defend their view to the group before then getting the group view?
Thanks !! And separately yes, but asymmetry by itself isn’t enough either. It’s some artful combination that matters, which changes from topic to topic.
Two things that I would try, of which one is possibly a special case of the other:
1. Ensure that the choices are driven by a rubric, and the rubric is explicit and built up front by the council so that it is clear what is being judged.
2. A special case that is using personas and asking the individual council members to behave like that personabWhen making the choice . Personas have an implicit rubric or choices, and although they are a little bit of a black box, they restrict the ambit of the choice in a way that is more controllable.
In real life, this is also how often design thinking processes do things. I tend to like using human correlates when designing collectively intelligent decision systems. I haven't tried this one, but it is worth giving it a shot. Happy to chat more if needed or appropriate.
For validation I asked Claudine to check with some of my advisory boards that I have held recently and the point is well made and I have got a convergence machine. Claudine has rewritten the skill - with my help to capture distinct ideas. Can we model emergence/ innovative ideas?
We can, with effort, get the models to help us think through more permutations of situations and therefore emergent ideas, but for now there needs to be a fair bit of scaffolding if we're to go further into the unknown.
Thank you, Azeem and Rohit. This piece, and the Karpathy council it builds on, directly shaped a tool I built, so thank you. The finding that councils show groupthink turned out to be the most useful thing I have read about combining multiple models.
What I took from your approach was the diagnosis. Decomposing answers into atomic "idea cards," clustering them, and blind judging their quality made something measurable: peer review followed by blending keeps the consensus while quietly dropping minority ideas. It also showed that the naive "best answer" picker preserves more uniqueness than the blend does. That reframed the whole problem for me.
Here is where I tried to resolve some of the shortcomings.
First, an explicit preservation protocol, built in. Before any synthesis, I build an idea ledger with provenance and rank each idea on its own merit, never on how many models raised it, so that a brilliant idea from one model can outrank a mediocre consensus one. Minority ideas are flagged as essential to include.
Second, a red team that rescues rather than only critiques. One pass hunts for valuable ideas raised by only one model that a consensus summary would otherwise drop, and then argues to keep them. This is a direct counter to the hidden profile problem.
Third, the survival metric is the product itself, not an afterthought. Every run reports unique versus consensus idea survival for both the baseline, which follows Karpathy's design, and my preserving variant, using the same inputs. Therefore the groupthink gap is visible on every run rather than merely assumed.
Fourth, genuinely decorrelated models. I ran a mixed Western and Chinese panel, namely Opus 4.8, Gemma, Kimi, and DeepSeek via OpenRouter and Ollama (for free local models). On a prompt asking for unconventional ideas, the baseline blend deleted exactly the unconventional ones, such as blockchain incentives, contribution art avatars, and curiosity gap designs, and most of these had come from the models that were not Western. The preserving pipeline kept all of them. On that run, the gap fell from 55% to 0%.
One boundary is deliberate. Preserved does not mean endorsed. Several of the rescued ideas are precisely the ones that most need human judgment, given the coercion and undue influence risk in my regulated domain. Therefore the tool surfaces them flagged with guardrails rather than smoothing them into the recommendation. The council widens the option space; however, a human still adjudicates.
I am also mindful that this is still early work, and the underlying evidence is one informal experiment. It is directionally convincing and consistent with human group dynamics; however, it is a signal to design against rather than settled science. I am grateful for the prompt to do exactly that.
I am now going to build an advisory board council with personas that Gianni suggested. It was only list to do already ….
> This piece, and the Karpathy council it builds on, directly shaped a tool I built, so thank you. The finding that councils show groupthink turned out to be the most useful thing I have read about combining multiple models.
Thank you!! This makes me quite happy! And very excited to see you build on this. Do put this out too when you're ready, I'd be quite interested to see it.
One important point that I forgot to add- I built this council on top Andrew Ng's Aisuite - https://github.com/andrewyng/aisuite.
Very interesting follow-on to Rohit's experiments. Thanks for that!
With that said, let me quibble with your characterization of your red team's rescuing unique &/or valuable ideas as "a direct counter to the hidden profile problem." By this I mean to quibble with your characterization, not with the value added by the red teams' actions you've assigned. Those actions are a separate issue, and seem a good idea.
As explained in Rohit's exposition, I take it that the "hidden profile problem" is more analogous to the zero-day exploits found by Mythos. I.E. There are a number of factors, ideas in this case, existing independently of a coherent whole. In the zero-day case this is a whole exploit cobbled together from discrete weaknesses not exploitable in isolation; in this idea case the coherent whole is some composite idea or approach to a challenge or problem to which a solution is desired.
My point here is that in the hidden profile situation, the desired "idea" or approach is not necessarily contained in any one of the approaches suggested by a single LLM, and thus cannot simply be "rescued." Rather, it consists only of particles or components of some approach that might be greater than the sum of its' parts. It is highly desirable that such "compositions" emerge rather than lay fallow, but insofar as these compositions are inherently novel, similarly their emergence is novel and for all that, unlikely. Nor would it seem a rescue function is conceptualized in such a way as to reveal them. Am I mistaken to think of this as a second-order problem? Whether first- or second-order, whether fanciful or actual, this possibility would seem to merit investigation.
Summing up, what I'm suggesting here is that this hidden profile circumstance is a separate problem worthy of separate consideration since the dimension in which it exists is perhaps orthogonal to that of coherent "complete" single-LLM-sourced ideas.
Thank you, this is a sharp and fair correction, and I think you are right.
You have caught a genuine conflation in my wording. What the red team rescues are complete, single-source ideas: things one model stated whole that a consensus blend would drop. That is a real failure mode, and it is essentially what Rohit's experiment operationalised and measured (idea-card survival). However, you are right that the classic hidden profile construct is about something deeper. The desired answer is distributed across members as fragments, held whole by no one, and it only emerges if those fragments are pooled and composed. Rescuing intact ideas does not reach that, so my phrase "a direct counter to the hidden profile problem" overstated it. More precisely, what we counter is the loss of complete minority ideas, which is adjacent to, but not the same as, the composition problem you describe.
What makes your point even sharper for my setup is that the anti-confabulation guardrails actively work against emergent composition. I added a faithfulness check that flags any claim in the synthesis that is "not traceable to any single member." A genuine hidden-profile composite is, by definition, traceable to no single member. So, my current design would flag exactly the kind of emergent, greater-than-the-parts idea that the hidden profile problem wants surfaced and treat it as a possible fabrication.
There is a real tension between "do not invent" and "synthesise something no one said alone," and the composition problem lives precisely in that gap. So, I agree it is a separate, second-order problem, roughly orthogonal to the dimension we have been working in. Addressing it would need a different mechanism: a deliberate composition or cross-pollination pass that tries to assemble fragments across models into candidate composites, plus some way to judge whether a composite is genuinely valuable rather than merely novel. That is harder, partly because it pulls against the same convergence-avoidance and fidelity controls that keep the rest honest, and partly because there is no ground-truth ledger for an idea that did not previously exist. I had not separated these cleanly. It deserves its own investigation, and I am grateful for the nudge.
I am going toa dd this as a separate artifact and make it exempt from the “faithfulness” check since by design they aren't traceable to one model), so the synthesis stays grounded while I still see "what might emerge from combining across models”
Excellent. I look forward to hearing about your struggle with this. It seems to be "a worthy problem!" I confess I wouldn't know how to begin! The goal I can see, the path... not so much. And of course another problem with my anticipated "hearing of" your efforts and hoped for solution is this: a key characteristic of Substack is the on-rushing FLOOD of content! What I lack, & so far as I know does not exist, is a cogent way to manage these conversations! While this isn't the topic here, insofar as this actual topic is one of the most interesting to have come up recently, any mechanisms, tips, tricks, etc U know of that might help, I'm all ears!
Neat experiment and the callout on asymmetric information is really interesting. Did you try a version where the job of the picker is to facilitate? That is, find those places of asymmetry and have them defend their view to the group before then getting the group view?
Thanks !! And separately yes, but asymmetry by itself isn’t enough either. It’s some artful combination that matters, which changes from topic to topic.
Two things that I would try, of which one is possibly a special case of the other:
1. Ensure that the choices are driven by a rubric, and the rubric is explicit and built up front by the council so that it is clear what is being judged.
2. A special case that is using personas and asking the individual council members to behave like that personabWhen making the choice . Personas have an implicit rubric or choices, and although they are a little bit of a black box, they restrict the ambit of the choice in a way that is more controllable.
In real life, this is also how often design thinking processes do things. I tend to like using human correlates when designing collectively intelligent decision systems. I haven't tried this one, but it is worth giving it a shot. Happy to chat more if needed or appropriate.
For validation I asked Claudine to check with some of my advisory boards that I have held recently and the point is well made and I have got a convergence machine. Claudine has rewritten the skill - with my help to capture distinct ideas. Can we model emergence/ innovative ideas?
We can, with effort, get the models to help us think through more permutations of situations and therefore emergent ideas, but for now there needs to be a fair bit of scaffolding if we're to go further into the unknown.