CreativityNeuro: Steering Language Model Weights to
Improve Divergent Thinking and Reduce Mode Collapse

Samuel Schapiro1
Univeristy of Illinois, Urbana-Champaign Core Francisco Park
Center for Brain Science, Harvard University
CBS-NTT Program in Physics of Intelligence, Harvard University
Prior Computers
Felix Sosa
Prior Computers
Lav R. Varshney
AI Innovation Institute, Stony Brook University


Abstract

Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect. Here, we introduce CreativityNeuro, a data-free method for enhancing divergent thinking in LLMs via contrastive weight steering. We evaluate our method across multiple creativity assessments and report several main findings. On the Divergent Association Task (DAT), a vocabulary-space creativity test, CreativityNeuro improves performance by up to 14 human percentile points. Next, in a large-scale human evaluation (N=720) on the Alternative Uses Test (AUT) and the Task Task, CreativityNeuro achieves significant improvements in originality, surprise, and creativity, transferring to longer-form and more open-ended tasks. Importantly, we find that across all three tasks, CreativityNeuro demonstrably reduces measures of mode collapse. Moreover, activation steering achieves comparable performance to CreativityNeuro on the DAT, but it does not transfer to the AUT and Task Task, demonstrating the effectiveness of weight-space steering in generalizing to unseen tasks. In conclusion, CreativityNeuro improves divergent thinking and reduces mode collapse without requiring behavioral data, re-training, or gradient-based fine-tuning, providing a straightforward way to enhance LLM performance in creative domains.

Figure 1: CreativityNeuro (CN) pipeline. Given a pair of contrastive creative prompts, CN computes parameter importance scores, selects a sparse subset of creativity-relevant parameters, and applies a scaled weight perturbation—without requiring behavioral datasets or gradient-based finetuning. CN improves divergent thinking across various tasks. Subplot (b) visualizes CN thinking outside of the “box”(i.e., the convex hull of baseline DAT responses), despite baseline responses falling within CN’s convex hull in subplot (a).

1 Introduction↩︎

Recent advances in large language models (LLMs) have renewed interest in a longstanding question: how can we understand and enhance creativity in intelligent systems? [1]. While this question has deep roots in cognitive science [2][11], it is now increasingly studied in the context of large-scale generative models [12][14]. Recent work has begun to assess the capacity for LLMs to engage in creative and open-ended tasks [15][20], where a recurring issue has surfaced: models tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect [21].

Within the creativity literature, a common distinction is made between divergent thinking, the capacity to generate multiple diverse solutions to a problem, and convergent thinking, the ability to find a single correct solution that unifies multiple diverse stimuli [5], [22]. Studying ways to enhance divergent thinking offers a promising pathway to encourage greater diversity and novelty in model responses, combating the homogenization issues that have emerged thus far. Here, we introduce a weight-space steering method that improves divergent thinking in LLMs. Our method outperforms prior approaches—including decoding, prompting, and activation steering—and generalizes better to unseen tasks, without requiring behavioral data or gradient-based fine-tuning. In detail, our main contributions are as follows:

  1. In 3, we introduce CreativityNeuro, a data-free method for steering creative behavior.

  2. In 4, we find that CreativityNeuro significantly improves divergent thinking on the Divergent Association Task (DAT), outperforming baselines such as prompting, activation steering, and decoding baselines.

  3. In 5, we conduct a large-scale human evaluation on the Alternative Uses Test (AUT) and Task Task (TT) and find that CreativityNeuro improves originality, surprise, and creativity on the AUT and TT, while activation steering transfers poorly to the AUT and TT.

  4. In 6, we find that CreativityNeuro reduces mode collapse across all three tasks.

  5. In 7, we find evidence that divergent thinking and factual reasoning are non-separable in weight space.

2 Related Work↩︎

We start by briefly reviewing related work before introducing our method in 3.

Evaluating the creativity of LLMs. Previous work has evaluated LLMs on divergent creativity assessments–including the DAT [23], AUT [5], the Task Task [24]–and in various real-world settings like scientific ideation [15], [16] and open-ended user queries [21]. On the DAT, LLMs can achieve scores well into the 90th percentile of humans [18], [19], whereas [25] studied GPT-3 on the AUT and concluded that humans exhibited greater creativity, with model responses showing weaker originality. Lastly, [24] found that model-generated goals on the Task Task achieved similar creativity ratings as human-generated goals, as assessed by a large panel of human raters.

Improving the creativity of LLMs. Most similar to this work, [26] proposed an activation steering method to amplify the creativity of LLMs, although improvements were only established for a single model, task, and human annotator. Our study is the first to demonstrate a steering method to improve creative behavior and validate its effectiveness in a large-scale human study. Apart from steering, various other approaches to improving LLM creativity include prompting frameworks [19], [27], [28], varying decoding parameters such as temperature [29], and reinforcement learning (RL) on preference data [30]. Unlike steering and RL-based approaches that typically require labeled behavioral data, our method operates entirely data-free.

Figure 2: CreativityNeuro
Table 1: Examples of contrastive prompt sets. Each set contains creative (\(\mathcal{P}^\text{cre}\)) and non-creative (\(\mathcal{P}^\text{non-cre}\)) prompts used for parameter importance scoring in [alg:creativityneuro].
Set Creative (\(\mathcal{P}^\text{cre}\)) Non-Creative (\(\mathcal{P}^\text{non-cre}\))
story Write the first line of a story that makes the reader question reality. Create a typical story beginning that establishes setting and character clearly.
problem What would an alien civilization’s approach to this problem look like? What best practices should be followed when solving this?
minimal Surprise me. Be precise.

3 Method↩︎

Recently, [31] has shown that parameter importance methods can be used to identify and amplify weights involved in mathematical reasoning, improving scores on the MATH benchmark by 4–17% [32]. Unlike mathematical reasoning, which can be elicited and evaluated on structured benchmarks such as MATH and GSM8K, creativity is a property of responses, not of questions. Open-ended prompts, such as those in [21], admit a wide range of potentially creative completions, but novelty and usefulness are measured on the outputs [1], [12], [13] rather than the inputs themselves. Therefore, our key methodological innovation is a framework for extending MathNeuro to a cognitive domain where the target behavior can be prompted but no structured dataset exists, making CreativityNeuro entirely data-free.

Contrastive Prompt Sets MathNeuro relies on questions drawn from MATH and GSM8K to obtain inputs for parameter importance scoring. Because no analogous dataset exists for creativity, we instead construct contrastive prompt sets: short instructions that direct the model toward creative (\(\mathcal{P}^\text{cre}\)) versus non-creative (\(\mathcal{P}^\text{non-cre}\)) behavior. We use six such sets spanning a range of styles—dat, storytelling, ideation, problem solving, open-ended, and minimal—where minimal contains only two- to five-word instructions (e.g., Surprise me vs Be precise). Representative examples are given in 1, and all six prompt sets are given in 2. As a result, CreativityNeuro does not require datasets, behavioral generations, scored responses, or labeled examples, making it entirely data-free, unlike [31].

Parameter Importance Scoring We use the same Wanda-style [33] parameter importance scoring as [31], restated here for completeness. This is done by taking the product of weight magnitude and activation norm, summed across a set of prompts \(b\) and their token positions \(t\): \(S_{\ell,ij} = \sum_{b,t} |W_{\ell,ij}| \cdot \|\mathbf{x}^{(b,t)}_{\ell,j}\|_2\), where \(\mathbf{x}^{(b,t)}_{\ell,j}\) is the \(j\)-th input activation at layer \(\ell\) for token \(t\) in prompt \(b\). We compute importance scores on creative \(\mathcal{P}^\text{cre}\) and non-creative prompts \(\mathcal{P}^\text{non-cre}\). Then, we isolate creativity-specific parameters by selecting the top \(\rho\) percent of weights ranked by creative importance that do not also appear in the top \(\rho\) percent for non-creative prompts. This set difference operation (\(C_\ell \setminus N_\ell\)) ensures we identify parameters uniquely associated with creative behavior. At inference time, we multiply weights in \(C_\ell \setminus N_\ell\) by a scaling factor \((1 + \alpha)\). The hyperparameters \(\rho\) and \(\alpha\) control the importance threshold and scaling strength, respectively. The full procedure is given in 2.

4 Experiments on the Divergent Association Task↩︎

Figure 3: CreativityNeuro (CN) improves divergent thinking across models and prompt sets. Given a human reference distribution [19] (N = 9{,}297, \mu = 78.26, \sigma = 6.73), we report: (a) DAT human percentile (\pmSEM) averaged across T \in \{0.9, 1.0, 1.2\} for CN, CAA, and the strongest sampling-based baselines; dashed lines show cross-model means for CN and CAA. (b) Heatmap showing human percentile improvement (\Delta%ile) for CreativityNeuro models across prompt sets, with statistical significance (p < 0.05) at each of the temperatures tested (0.9, 1.0, 1.2) denoted by an asterisk. (c) CDF showing DAT scores for CreativityNeuro models on the best performing prompt set.

We test instruct-tuned models across three open-weight model families (Phi, Llama, Qwen), totaling six models at 3B, 4B, 7B, 8B, and 14B sizes: LLaMA (3.2-3B-Instruct, 3.1-8B-Instruct) [34], Qwen-2.5 (7B-Instruct, 14B-Instruct) [35], and Phi (3.5-mini-Instruct (4B), 3-medium-4k-Instruct (14B)) from Microsoft [36]. For each model we use the best \((\rho, \alpha, \text{prompt set})\) configuration with \(\geq 120\) valid CN samples at all three temperatures \(T \in \{0.9, 1.0, 1.2\}\). Full hyperparameter sweep details are in 12.

Task The DAT asks participants to generate \(N\) words that are as semantically distant from each other as possible [23]. Given a set of \(N\) words \(W := \{w_1, w_2, \dots, w_N\}\) with corresponding GloVe embeddings \(V := \{\mathbf{v}_1, \mathbf{v}_2, \dots, \mathbf{v}_N\} \subseteq \mathbb{R}^{300}\), the DAT score is the average pairwise semantic distance among all distinct pairs of those \(N\) words: \[\textrm{DAT}(W) := \frac{100}{N(N-1)} \sum_{i \neq j}^N (1 - \cos(\mathbf{v}_i, \mathbf{v}_j))\] Following [23], we use the 840B-token GloVe embeddings [37] as our semantic space. A participant is asked to name \(N=10\) words, and the first \(7\) valid words are kept [23]. Full prompts are given in 13.

Baselines Existing studies have found that temperature-scaling and prompting can influence DAT scores [18], [19]. To ensure the CN intervention leads to a meaningful improvement over such techniques, we compare against a broad set of baselines, including prompting; varying decoding parameters such as top-p nucleus sampling, top-k sampling, temperature, and repetition penalty; as well as activation steering via contrastive activation addition (CAA; [38]), which injects a steering vector \(\mathbf{v}_\ell = \bar{\mathbf{h}}_\ell^{+} - \bar{\mathbf{h}}_\ell^{-}\) into the residual stream during decoding. For activation steering, contrast pairs are obtained from top- vs.bottom-quartile DAT responses by score, creating a “divergent thinking” direction in the residual stream.2 Full decoding strategy and activation steering hyperparameter settings are given in 10. All settings are evaluated across all six models at \(T \in \{0.9, 1.0, 1.2\}\) until \(N{=}120\) valid DAT responses are obtained.

4.1 Results↩︎

CreativityNeuro achieves robust performance improvements across prompt sets CN improves DAT performance across all six models and prompt sets (3), outperforming all sampling-based baselines. Panel (b) of 3 shows \(\Delta\) Percentile for the best \((\rho, \alpha)\) per model–prompt combination: while the dat prompt set produces the most consistent gains, non-dat prompt sets also yield statistically significant improvements, suggesting CN is able to identify weights controlling divergent thinking behavior, rather than localizing DAT-specific task knowledge.

CreativityNeuro outperforms activation steering without needing behavioral data CN (94.1 avg) slightly outperforms activation steering (93.9 avg) on DAT percentile (3a); however, activation steering requires scored DAT responses to construct its steering vector, while CN uses only creative and non-creative prompts with no generation or scoring. Prompt-only activation steering (omitted from the figure; see footnote in [par:dat95baselines]) averaged only 87.8, comparable to prompting (87.9), suggesting that the behavioral data is essential for activation steering to be competitive, whereas CN achieves stronger performance from prompts alone.

CreativityNeuro improves DAT scores across all six models and prompt sets, outperforming prompting, sampling-parameter, and activation steering baselines.

5 Experiments on the Alternative Uses Test and Task Task↩︎

Here, we evaluate CreativityNeuro on more complex divergent thinking tasks than the DAT. The Alternative Uses Test (AUT), a standard instrument in the psychometrics literature, asks participants to generate creative uses for a common object [5]. We administer the AUT using a standard set of objects in the creativity literature: brick, paperclip, and fork. Following [25], uses are scored on originality, surprise, and utility. We also evaluate CreativityNeuro on the Task Task (TT), which assesses the ability to generate novel challenges or goals that themselves expect novel solutions [24]. Participants design creative game show challenges that would be fun to attempt, entertaining to watch, and difficult enough to be interesting. Following [24], we evaluate responses on creativity and originality.3 Full prompts are in 13.

Figure 4: Cohen’s d (\pmSE) from intra-participant z-scored human ratings on the AUT and TT. (a–c) AUT Originality, Surprise, Utility. (d–e) TT Creativity, Originality. Black outlines indicate p < .05, and green shading marks the positive-effect region.

Models and Inference For each of the six models in 4, we select the CreativityNeuro configuration (\(\rho\), \(\alpha\), prompt set) that produced the largest statistically significant DAT improvement. Then, for each task, we sample 40 stimuli (20 baseline, 20 creative) at temperature \(T=1.0\), \(\text{top-}k = 0\), and \(\text{top-}p = 1.0\). We compare CreativityNeuro against activation steering, the strongest non-CN baseline from 4.

Human Experiment Design We evaluate AUT and TT stimuli via human ratings on Prolific. Human studies use a between-subjects design, where every participant rates 10 stimuli (5 baseline, 5 creative) in randomized order with condition labels hidden. We ensure balanced allocation via automated participant-to-slot assignment so that each stimulus receives exactly 10 independent ratings. Ratings on the AUT and TT are given on continuous 0–100 sliders, and to control for individual differences in scale usage, we compute intra-participant \(z\)-scores: for each participant \(i\) and dimension \(d\), \(z_{i,d,s} = (r_{i,d,s} - \bar{r}_{i,d})\,/\,\sigma_{i,d}\), where \(\bar{r}_{i,d}\) and \(\sigma_{i,d}\) are computed across all stimuli that participant rated on that dimension. We recruit 30 participants per (model, task, method) triple, reaching N=720 total human reviewers. Effect sizes and significance (\(t\)-tests4) are computed on \(z\)-scored ratings for baseline vs.creative, where each participant is treated as an independent sample.

Figure 5: Top-rated CreativityNeuro vs.baseline generations from the same model. Intra-participant z-scores averaged across raters (N{=}30 per cell). (a) CreativityNeuro responses on the AUT tend to score higher on originality and surprise. (b) CreativityNeuro challenges on the Task Task.

5.1 Results↩︎

CreativityNeuro improves originality, surprise, and creativity We report results in 4. On the AUT, CreativityNeuro achieves uniformly positive originality effects across all six models (avg.\(d = +.36\)), with four reaching significance. The effects are even stronger for surprise (avg.\(d = +.43\)), with five models reaching significance. On the Task Task, CreativityNeuro obtains strong originality gains (avg.\(d = +.40\)), with the strongest effects in Phi-14B (\(d = +.61\)) and Qwen-14B (\(d = +.56\)), and moderate improvements in overall creativity (avg.\(d = +.24\)). Although CreativityNeuro degrades AUT utility, this is a predictable consequence of the novelty–utility tradeoff already present in baseline responses. See Appendix 14 for detailed analysis of this tradeoff.

CreativityNeuro generalizes better to the AUT and TT than activation steering Additionally, while activation steering (93.9 avg) performs comparably to CreativityNeuro (94.1 avg) on the DAT (3a), activation steering fails to generalize to the AUT and Task Task. These results are consistent with recent work that has found weight-space steering generalizes further out-of-distribution than activation steering on sycophancy and value alignment tasks [39]. Moreover, activation steering effectiveness has been shown to vary significantly by behavior type, with more complex behaviors like embodying persona archetypes and public figures proving more difficult to steer [40]. Techniques such as context-dependent [41], [42] and learned activation steering [43] have been proposed to remediate such issues, and may be explored in future work.

CreativityNeuro generalizes to open-ended creative tasks judged by human raters, improving measures of originality and surprise, while activation steering does not exhibit reliable transfer.

Figure 6: Measures of mode collapse across tasks. Baseline / CAA / CN shown left-to-right (teal / green / orange). (a) DAT vocabulary entropy. (b) DAT top 10 word share. (c) AUT and TT embedding homogeneity. (d) Cross-family vocabulary overlap.

6 Evaluating Mode Collapse Across Tasks↩︎

Instruction-tuned LLMs are known to suffer from mode collapse–the tendency to concentrate open-ended outputs on a narrow set of semantic clusters [21], [44]. In this section, we test whether CreativityNeuro reduces word-level mode collapse on the DAT and response-level collapse on the AUT and TT.

Metrics On the DAT, we pool outputs across temperatures \(T \in \{0.9, 1.0, 1.2\}\) and compute each model’s vocabulary entropy \(H\) and probability mass assigned to each (model, condition \(\in \{\) baseline, CAA, CN \(\}\))’s own top 10 most-frequent words. On the AUT and TT, where responses span multiple sentences, we adopt the intra-model repetition (mean pairwise cosine similarity within a model’s responses) and inter-model homogeneity (mean pairwise cosine similarity between responses from different models on the same query) metrics from [21], also using openai/text-embedding-3-large for response embedding.

6.1 Results↩︎

Baseline models suffer significantly from mode collapse on the DAT. Baseline models concentrate \(25.5\%\) of generated tokens on average on each model’s top-10 most frequent words (6b, averaged across \(T \in \{0.9, 1.0, 1.2\}\)). Furthermore, \(19\) words appear in the top 30 vocabulary of at least three of the six tested models, as shown in 6d. Three of these words—galaxy, quasar, xylophone—appear in the baseline top 30 of all three model families (LLaMA, Phi, Qwen) simultaneously, and two more—glacier, nebula—appear in the baseline top-30 of \(\geq 4\) of the six models.

CreativityNeuro reduces word-level mode collapse on the DAT. As shown in 6b, CreativityNeuro reduces the top 10 share by \(10.2\) pp on average and increases vocabulary entropy by \(0.59\) nats (\(+10\%\), 6a). Activation steering (CAA) produces a smaller but qualitatively similar effect: top-10 share drops by \(6.6\) pp (\(0.255 \to 0.189\)), and vocabulary entropy increases by \(0.40\) nats (\(+7\%\)).

CreativityNeuro and activation steering both reduce mode collapse on the AUT and TT In 6c, we report relative reductions vs.baseline annotated above each CAA bar (green) and CN bar (orange). Both CN and CAA reduce homogeneity on each task. On the AUT (sentence-length responses), CAA reduces intra-model repetition by \(6.6\%\) and inter-model homogeneity by \(4.7\%\), larger than CN’s \(2.4\%\) and \(3.3\%\) respectively. On the TT (paragraph-length responses), CN reduces intra-model repetition by \(5.5\%\) and inter-model homogeneity by \(6.3\%\), larger than CAA’s \(3.0\%\) and \(2.5\%\).

CreativityNeuro reduces semantic mode collapse on both word-level (DAT) and embedding-level (AUT, TT) measures. Activation steering produces a similar but smaller effect on the DAT and TT, while matching CreativityNeuro on the AUT.

7 Are Divergent Thinking and Factual Reasoning Separable in Weights?↩︎

Previously, [31] demonstrated that MathNeuro could improve MATH benchmark scores without degrading factual reasoning on the Massive Multi-task Language Understanding (MMLU) benchmark [32]. Here, we study whether improved divergent thinking interferes with factual reasoning on MMLU.

Setup For each of the six models, we take the top performing \((\rho, \alpha\), prompt) configuration from 4 and measure 5-shot accuracy on 500 MMLU questions, across 5 seeds per model, under two mask construction techniques:

  1. Default: \(P^\text{cre} \setminus P^\text{non-cre}\) uses the default creative and non-creative prompts from 4.

  2. MMLU-protected: \(P^\text{cre} \setminus (P^\text{non-cre} \cup P^\text{MMLU})\) adds 20 randomly sampled MMLU prompts \(\mathcal{P}^\text{MMLU}\) to the negative contrast set in 2, further removing any weight whose importance is above the \(1-\rho\) percentile for MMLU.

In summary, the MMLU-protected masks attempts to explicitly separate creative weights from MMLU weights. If divergent thinking and factual reasoning are fully separable, the MMLU-protected mask should preserve MMLU scores without degrading DAT scores.

7.1 Results↩︎

Divergent thinking and factual reasoning are non-separable in weight space Default masks reduce MMLU accuracy by \(-3.13\) pp on average (7). Meanwhile, MMLU-protected masks gain an additional \(+1.33\) percentile \(\Delta\)DAT on top of default masks, but further reduce MMLU scores by \(-0.71\) pp (7). Surprisingly, adding MMLU prompts to the negative contrast set further degrades MMLU accuracy, despite reducing the size of the masks by \(\sim 2 \times\), providing evidence that divergent thinking and factual reasoning are functionally entangled and non-separable in weight space.

Figure 7: Parameter importance across masks at \rho{=}0.1. The default mask \mathcal{P}^{\mathrm{cre}} \setminus \mathcal{P}^{\mathrm{noncre}} and the MMLU-protected mask \mathcal{P}^{\mathrm{cre}} \setminus (\mathcal{P}^{\mathrm{noncre}} \cup \mathcal{P}^{\mathrm{MMLU}}) are the dotted regions on the left. Annotated \DeltaDAT (percentile) and \DeltaMMLU (pp) are cross-model means \pm SEM.

Our results are consistent with broader findings in the mechanistic interpretability literature. Namely, it has been shown that individual weights can be entangled in multiple distinct functions (a phenomenon known as polysemanticity), which supports the ability of neural models to represent more features than they have neurons (superposition) [45], [46]. Creativity research more broadly suggests a need for separation between generation and selection [13]. The success of multi-agent creative systems [47] and multi-stage prompting techniques that decouple creative exploration from constraint satisfaction [27] can be interpreted within the context of these results: if weights controlling divergent and convergent thinking are entangled, simultaneously eliciting strong divergent and convergent abilities may be challenging or even impossible. Therefore, separating these steps across agents or prompts can provide stronger overall performance.

We find evidence, consistent with the broader mechanistic interpretability literature, that divergent thinking and factual reasoning are functionally entangled in model weights.

8 Limitations and Future Work↩︎

Our evaluation is restricted to a finite set of divergent thinking benchmarks and metrics, which capture only certain aspects of creativity. As [48] notes, divergent thinking is not synonymous with creativity and should best be thought of as a measure of creative potential. Our comparison to activation steering focuses on a standard CAA-style method and does not exhaust the space of possible activation-space interventions, such as context-dependent or learned approaches. Lastly, we found evidence suggesting divergent thinking and factual reasoning are non-separable in model weights—understanding whether this entanglement reflects a fundamental architectural constraint or arises as an artifact of the Wanda-style importance technique remains an important open question. Designing architectures that learn unified factored representations [49] and disentangle divergent thinking from factual reasoning would be a valuable and timely direction for future work.

9 Full Sets of Contrastive Prompts↩︎

Each of the six prompt sets contains 10 creative and 10 non-creative exemplars. Below we show 3 representative examples from each set.

Table 2: Full contrastive prompt sets (3 examples each). Each prompt set contains 10 creative (\(\mathcal{P}^\text{cre}\)) and 10 non-creative (\(\mathcal{P}^\text{non-cre}\)) exemplars used for parameter importance scoring in [alg:creativityneuro].
Set Creative (\(\mathcal{P}^\text{cre}\)) Non-Creative (\(\mathcal{P}^\text{non-cre}\))
dat List 10 common English nouns that are as unrelated in meaning as possible. Avoid any shared topic or category. Output only the nouns, separated by commas. List 10 common English nouns that are as closely related in meaning as possible and clearly fit into a single narrow topic. Output only the nouns, separated by commas.
Give me 10 one-word English nouns that are extremely far apart in semantic meaning. They should not fit into a single theme. Give me 10 one-word English nouns that are very strongly associated with each other and belong to the same specific domain.
Produce 10 everyday English nouns that share no obvious connection with one another. Each noun should come from a very different domain. Produce 10 English nouns that are tightly connected around a single clear theme (for example, all parts of a computer or all items in a kitchen).
story Write an unusual opening sentence for a short story that subverts reader expectations. Write a standard opening sentence for a fairy tale.
Create a story beginning that combines two unrelated concepts in a surprising way. Create a typical story beginning that establishes setting and character clearly.
Write the first line of a story that makes the reader question reality. Write the first line of a conventional mystery novel.
ideation List 5 unusual uses for a brick that nobody has thought of before. List 5 common uses for a brick in construction.
Generate unconventional solutions to reduce traffic in cities. Generate standard solutions to reduce traffic in cities.
What are some surprising ways a library could be repurposed? What are the traditional functions of a library?
problem How might you solve this problem in a way that seems counterintuitive at first? What is the most efficient way to solve this problem?
What would an alien civilization’s approach to this problem look like? List the standard steps for addressing this type of issue.
If you had to solve this with resources from a different era, what would you do? What best practices should be followed when solving this?
open Generate a joke about electric vehicles. Explain nuclear fission like I am five years old.
Create the first verse of a wedding vow. What is Bukhara? Provide a paragraph-long explanation in layman’s language.
Write a song about a guy named Jacob working at a call center making jokes. In a few sentences explain what threats do scams pose to individuals?
minimal Invent something. State a fact.
Surprise me. Be accurate.
Be weird. Be precise.

10 Baseline Sweep Settings↩︎

In 4, we compare CreativityNeuro against a set of baseline techniques.

  1. Prompting: We present creative exemplars from each of six prompt sets (2) as in-context guides

  2. Decoding Parameter Sweeps

    1. top-\(p\) nucleus sampling, \(p \in \{0.8, 0.85, 0.9, 0.95, 1.0\}\)

    2. top-\(k\) sampling, \(k \in \{10, 25, 50, 100, \text{disabled}\}\)

    3. repetition penalty, \(\theta \in \{1.0, 1.1, 1.2, 1.5, 2.0, 3.0\}\)

  3. Activation Steering: We sweep the injection layer suffix from single-layer to the final 50% of layers and find that injecting into the final 30% works best. We sweep \(\alpha \in \{0.1, 0.2, 0.3, 0.4, 0.5, 1.0, 2.0, 4.0\}\), selecting the best \(\alpha\) where all three temperatures yield \(\geq 120\) valid samples.

11 Layerwise Ablation Studies↩︎

11.1 Suffix vs.Prefix↩︎

Figure 8: Layerwise ablation: suffix vs.prefix vs.single-layer. Each panel shows the % of full CN DAT effect recovered as a function of the number of layers k with CN weights applied. Solid lines: suffix (last k layers). Dashed lines: prefix (first k layers). Dotted lines: single-layer (one layer at a time, plotted by layer index). Diamond markers indicate the fewest layers from the back achieving \geq​95% of the full effect.

To localize the CN effect across the network, we compare three layerwise interventions (8): (i) suffix, applying CN weights to only the last \(k\) layers; (ii) prefix, applying CN weights to only the first \(k\) layers; and (iii) single-layer, applying CN weights to one layer at a time. For each condition, we generate \(N=120\) valid DAT samples at \(T=1.0\) and report the percentage of the full CN effect recovered in terms of DAT scores.

11.1.0.1 Suffix.

Applying CN weights to a suffix of layers (layers \(k\) through \(L{-}1\)) recovers the full effect with roughly half the network. On average, 51% of layers (from the back) are needed to reach  100% recovery, ranging from 29% for Qwen-7B (last 8/28) to 81% for Llama-8B (last 26/32). The remaining models fall in between: Llama-3B (last 14/28, 50%), Phi-14B (last 20/40, 50%), Qwen-14B (last 22/48, 46%), and Phi-4B (last 17/32, 53%). However, this analysis is only conducted on DAT scores, and it is unclear whether the last 51% of layers are sufficient to recover 100% of the scores on the AUT and Task Task.

11.1.0.2 Prefix.

Applying CN weights from the front (layers \(0\) through \(k\)) is far less efficient. On average, 78% of all layers must be included before the prefix condition reaches 95% recovery. For Qwen-14B (48 layers), the prefix condition never reaches 95% even when all layers are included (94% at \(k=48\)). This asymmetry provides evidence that the CN effect is concentrated in later layers of the residual stream.

11.1.0.3 Single-layer.

No single layer is sufficient to recover the full CN effect. The best individual layers recover 50–72% of the effect (e.g., Q7B layer 19: 72%, L3B layer 15: 64%, P14B layer 20: 60%), with top contributors generally appearing in middle-to-late layers. The gap between the best single layer and the suffix provides evidence that the CN effect requires cooperation across multiple late layers, consistent with findings that representations of human-interpretable concepts or behaviors may span multiple layers [46], [50].

12 Hyperparameter Sweeps↩︎

CreativityNeuro introduces two hyperparameters: the importance threshold \(\rho \in (0, 1]\), which controls the fraction of parameters selected by the importance mask, and the scaling factor \(\alpha > 0\), which controls the magnitude of the weight perturbation applied to masked parameters. To identify effective \((\rho, \alpha)\) configurations for each model, we conduct systematic grid sweeps evaluated on the DAT. For each model, we generate CN weight masks using six different prompt sets: dat, story, ideation, problem, openended, and minimal. Each prompt set produces a separate importance mask. We then evaluate every combination of keep ratio \(\rho \in \{0.01, 0.05, 0.1, 0.2\}\) and scaling factor \(\alpha \in \{0.1, 0.5, 1.0, 2.0\}\) at three sampling temperatures \(T \in \{0.9, 1.0, 1.2\}\), with \(\text{top-}k = 0\) and \(\text{top-}p = 1.0\) (i.e., untruncated sampling). For LLaMA-3.1-8B and Phi-3.5-mini, after initial experiments revealed that \(\alpha > 0.5\) were too high, we additionally tested a finer alpha grid \(\alpha \in \{0.01, 0.05, 0.075, 0.1, 0.5, 1.0\}\) to probe the conservative regime identified by [31]. Each configuration generates \(n = 120\) valid DAT samples (10-word lists with all words present in the GloVe vocabulary).

For each \((\rho, \alpha, T)\) triple, we compute the mean DAT scores for the CreativityNeuro (\(\overline{\text{DAT}}_{\text{CN}}\)) and baseline (\(\overline{\text{DAT}}_{\text{base}}\)) models, then convert each to a human percentile using the distribution from [19]. [fig:sensitivity_dat,fig:sensitivity_story,fig:sensitivity_ideation,fig:sensitivity_problem,fig:sensitivity_openended,fig:sensitivity_minimal] show \(\Delta\) Percentile \(= P(\overline{\text{DAT}}_{\text{CN}}) - P(\overline{\text{DAT}}_{\text{base}})\) (averaged across temperatures) as a function of \((\alpha, \rho)\) for each of the six prompt sets, across all six models. Positive values indicate that CN moves the model’s DAT performance upward in the human score distribution.

Figure 9: Sensitivity to (\alpha, \rho) — DAT prompt set. \Delta Percentile averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.
Figure 10: Sensitivity to (\alpha, \rho) — Story prompt set. Mean \DeltaDAT averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.
Figure 11: Sensitivity to (\alpha, \rho) — Ideation prompt set. Mean \DeltaDAT averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.
Figure 12: Sensitivity to (\alpha, \rho) — Problem prompt set. Mean \DeltaDAT averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.
Figure 13: Sensitivity to (\alpha, \rho) — Open-ended prompt set. Mean \DeltaDAT averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.
Figure 14: Sensitivity to (\alpha, \rho) — Minimal prompt set. Mean \DeltaDAT averaged across three temperatures (T \in \{0.9, 1.0, 1.2\}) for each model. Color scale is shared across panels and centered at zero.

13 Evaluation Prompts↩︎

In this section, we share the full prompts used to generate and/or post-process the stimuli for the DAT, AUT, and TT.

13.1 Divergent Association Task (DAT)↩︎

The DAT prompt instructs models to generate 10 maximally dissimilar nouns:

You will be asked to name exactly 10 English nouns.
Output exactly 10 words, separated by commas, and nothing else.
Return ONLY: word1, word2, ..., word10
DO NOT write explanations. DO NOT think step by step.

Be as original and unusual as possible. Avoid common or closely related words.
Now name 10 English nouns that are as different from each other as possible.

13.2 Alternative Uses Test (AUT)↩︎

The AUT uses a system prompt and a user prompt, with {object} replaced by one of 3 standard objects: brick, paperclipfork.

You are participating in a creativity test. Your task is to generate creative, unusual, and original uses for common objects. Be imaginative and think outside the box.

List 5 creative and unusual alternative uses for a {object}.
Be specific and creative. List each use on a new line, numbered 1 through 5.
Only list the uses, no explanations.

13.3 Task Task (TT)↩︎

The Task Task uses a system prompt and a user prompt with three in-context examples from [24].

You are a creative game show designer. Your task is to invent fun, original, and entertaining challenges that would be exciting for contestants to attempt and audiences to watch.

You’ve been recruited to help design challenge tasks for a new game show! Your job is to come up with a new creative, silly, and fun task for humans to solve.

Here are a few example tasks:

1.Your goal is to: Throw a teabag into a mug from the farthest distance. You can use: Anything you can reasonably expect to find in a house, garage, and garden shed.

2.Someone has squeezed all of the toothpaste out of the toothpaste tube. Your goal is to: Get as much of the original toothpaste back into the empty toothpaste tube as possible. You can use: Anything you can reasonably expect to find in a bathroom.

3.Your goal is to: Transfer water between two fishbowls using only the supplied items. You cannot move the fishbowls. You can use: a chocolate bar, a rubber glove, a baguette, a snorkel, a cardboard tube, and a plate of pasta.

Now it’s your turn! Create your own creative, silly, and fun task for future participants to solve. Specify the goal, scoring criteria, and any materials or constraints. Describe it in exactly 3–4 sentences as a short paragraph — do not use lists or bullet points. Do not comment on how entertaining, creative, or fun the task would be.

13.4 Task Task Post-Processing↩︎

Raw Task Task generations are post-processed using GPT-4o to normalize formatting before human evaluation.

You are an editor cleaning game show challenge descriptions for a human evaluation study.

Make these MINIMAL edits:

1. REMOVE any task title at the start (e.g., ‘In "Baking Bonanza,"’). After removing, capitalize the first word of the remaining text.

2. Convert ALL third-person references to SECOND PERSON (e.g., “contestants must” \(\rightarrow\) “you must”, “the player” \(\rightarrow\) “you”).

3. REMOVE any metacommentary sentences (e.g., “This creative juggling act combines laughs with a dash of physics comedy.”).

CRITICAL RULES:
- Do NOT summarize, condense, or shorten the task description
- Do NOT change the creative content or game mechanics
- Do NOT add any text — only edit or remove
- Output ONLY the cleaned description with no preamble or explanation

Clean this game show challenge description: {text}

14 Novelty–Utility Tradeoff Analysis↩︎

Among baseline AUT stimuli, originality and utility are negatively correlated (\(r = -.49\), \(p < 10^{-7}\)), as are surprise and utility (\(r = -.64\), \(p < 10^{-13}\)). As predicted by fundamental novelty–utility tradeoffs established in the creativity literature [13], increases in subjective ratings of novelty tend to be accompanied by decreases in perceived utility. Therefore, CN’s utility reduction is an expected byproduct of steering towards high-novelty responses, rather than an independent failure mode. To confirm this, we perform a hypervolume (HV) analysis [51] over all responses in the three-dimensional rating space (originality, surprise, utility) and find that the CN HV indicator is 15.2% larger than baseline (\(p = .18\)), suggesting CN obtains a moderate (but non-significant) gain in Pareto efficiency as well.

15 Disclosure of Large Language Model Usage↩︎

In this paper, large language models (LLMs) were used to assist in the code implementation, plotting and figure generation, reporting (but not analysis) of experiment results, and for comprehensive surveys of related work.

References↩︎

[1]
M. A. Boden, The Creative Mind: Myths and Mechanisms. Routledge, 2004.
[2]
L. A. Quetelet, “A treatise on man and the development of his faculties (a facsimile reproduction of the english translation of 1842 with an introduction by solomon diamond).” 1842.
[3]
F. Galton, Hereditary genius: An inquiry into its laws and consequences. D. Appleton & Company, 1870.
[4]
J. Hadamard, An Essay on the Psychology of Invention in the Mathematical Field. Courier Corporation, 1954.
[5]
J. P. Guilford, Psychological Bulletin THE STRUCTURE OF INTELLECT,” 4, 1956.
[6]
S. Mednick, The associative basis of the creative process. Psychological Review, vol. 69, no. 3, p. 220, 1962.
[7]
A. Koestler, The Act of Creation. Macmillan, 1964.
[8]
D. K. Simonton, Creativity in Science: Chance, Logic, Genius, and Zeitgeist. Cambridge University Press, 2004.
[9]
Dietrich, Arne, The Cognitive Neuroscience of Creativity,” Psychonomic Bulletin & Review, vol. 11 (6), pp. 1011–1026, 2004.
[10]
G. Fauconnier and M. Turner, The Way We Think: Conceptual Blending and the Mind’s Hidden Complexities. Basic Books, 2008.
[11]
A. Rothenberg, Flight from Wonder: An Investigation of Scientific Creativity. Oxford University Press, 2014.
[12]
M. L. Maher, Evaluating creativity in humans, computers, and collectively intelligent systems,” in Proceedings of the 1st DESIRE network conference on creativity and innovation in design, 2010, pp. 22–28.
[13]
L. R. Varshney, Mathematical limit theorems for computational creativity,” IBM Journal of Research and Development, vol. 63, no. 1, Jan. 2019, doi: 10.1147/JRD.2019.2893907.
[14]
S. Schapiro, J. Black, and L. R. Varshney, Transformational Creativity in Science: A Graphical Theory,” arXiv preprint arXiv:2504.18687, 2025.
[15]
C. Si, D. Yang, and T. Hashimoto, Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers,” Sep. 2024, [Online]. Available: http://arxiv.org/abs/2409.04109.
[16]
C. Si, T. Hashimoto, and D. Yang, The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas,” Jun. 2025, [Online]. Available: http://arxiv.org/abs/2506.20803.
[17]
A. Sanyal, S. Schapiro, S. Shashidhar, R. Moon, L. R. Varshney, and D. Hakkani-Tur, Spark: A system for scientifically creative idea generation,” arXiv preprint arXiv:2504.20090, 2025.
[18]
A. Bellemare-Pepin et al., Divergent Creativity in Humans and LLMs,” 2024.
[19]
H. Wang et al., A large-scale comparison of divergent creativity in humans and large language models,” Nature Human Behaviour, 2025, doi: 10.1038/s41562-025-02331-1.
[20]
Q. Wang, D. Downey, H. Ji, and T. Hope, SciMON: Scientific Inspiration Machines Optimized for Novelty,” Jun. 2024, doi: 10.18653/v1/2024.acl-long.18.
[21]
L. Jiang et al., “Artificial hivemind: The open-ended homogeneity of language models (and beyond),” arXiv preprint arXiv:2510.22954, 2025.
[22]
A. Dietrich, Types of creativity,” Psychonomic Bulletin and Review, vol. 26, no. 1, pp. 1–12, Feb. 2019, doi: 10.3758/s13423-018-1517-7.
[23]
J. A. Olson, J. Nahas, D. Chmoulevitch, S. J. Cropper, and M. E. Webb, Naming unrelated words predicts creativity,” Apr. 2021, doi: 10.1073/pnas.2022340118/-/DCSupplemental.y.
[24]
J. Chu, J. Hu, and T. D. Ullman, “The task task: Creative problem generation in humans and language models,” in Proceedings of the annual meeting of the cognitive science society, 2024, vol. 46.
[25]
C. Stevenson, I. Smal, M. Baas, R. Grasman, and H. Van Der Maas, Putting GPT-3’s Creativity to the (Alternative Uses) Test,” in International conference on computational creativity, 2022, [Online]. Available: http://osf.io/vmk3c/.
[26]
M. L. Olson, N. Ratzlaff, M. Hinck, S. Tseng, and V. Lal, Steering Large Language Models to Evaluate and Amplify Creativity,” Dec. 2024, [Online]. Available: http://arxiv.org/abs/2412.06060.
[27]
M. H. Nguyen and A. Singla, “Divergent-convergent thinking in large language models for creative problem generation,” arXiv preprint arXiv:2512.23601, 2025.
[28]
R. Morain and D. Ventura, Is Prompt Engineering the Creativity Knob for Large Language Models? in Proceedings of the 16th international conference on computational creativity (ICCC’25), 2025.
[29]
M. Peeperkorn, T. Kouwenhoven, D. Brown, and A. Jordanous, Is temperature the creativity parameter of large language models? arXiv:2405.00492, 2024.
[30]
X. Wei et al., “Igniting creative writing in small language models: Llm-as-a-judge versus multi-agent refined rewards,” in Proceedings of the 2025 conference on empirical methods in natural language processing, 2025, pp. 17171–17197.
[31]
B. R. Christ, Z. Gottesman, J. Kropko, and T. Hartvigsen, Math Neurosurgery: Isolating Language Models’ Math Reasoning Abilities Using Only Forward Passes,” Jun. 2025, [Online]. Available: http://arxiv.org/abs/2410.16930.
[32]
D. Hendrycks et al., Measuring Mathematical Problem Solving With the MATH Dataset,” Nov. 2021, [Online]. Available: http://arxiv.org/abs/2103.03874.
[33]
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,” arXiv preprint arXiv:2306.11695, 2023.
[34]
A. Grattafiori et al., “The llama 3 herd of models.” 2024, [Online]. Available: https://arxiv.org/abs/2407.21783.
[35]
A. Yang et al., “Qwen2.5 technical report.” 2025, [Online]. Available: https://arxiv.org/abs/2412.15115.
[36]
M. Abdin et al., “Phi-3 technical report: A highly capable language model locally on your phone.” 2024, [Online]. Available: https://arxiv.org/abs/2404.14219.
[37]
J. Pennington, R. Socher, and C. D. Manning, GloVe: Global Vectors for Word Representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1532–1543, doi: 10.3115/v1/D14-1162.
[38]
N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner, “Steering Llama 2 via contrastive activation addition,” arXiv preprint arXiv:2312.06681, 2024.
[39]
C. Fierro and F. Roger, “Steering language models with weight arithmetic,” arXiv preprint arXiv:2511.05408, 2025.
[40]
T. Bas and K. Novak, “What can we actually steer? A multi-behavior study of activation control,” arXiv preprint arXiv:2511.18284, 2025.
[41]
J. Li, Y. Li, and K.-H. Huang, “Steering vector fields for context-aware inference-time control in large language models,” arXiv preprint arXiv:2602.01654, 2026.
[42]
B. W. Lee et al., “Programming refusal with conditional activation steering,” arXiv preprint arXiv:2409.05907, 2024.
[43]
P. Rodriguez et al., “LinEAS: End-to-end learning of activation steering with a distributional loss,” NeurIPS, 2025.
[44]
J. M. Springer et al., “Annotations mitigate post-training mode collapse.” 2026, [Online]. Available: https://arxiv.org/abs/2605.09995.
[45]
N. Elhage et al., “Toy models of superposition,” Transformer Circuits Thread, 2022, [Online]. Available: https://transformer-circuits.pub/2022/toy_model/index.html.
[46]
L. Sharkey et al., “Open problems in mechanistic interpretability.” 2025, [Online]. Available: https://arxiv.org/abs/2501.16496.
[47]
Y.-C. Lin et al., “Creativity in llm-based multi-agent systems: A survey,” in Proceedings of the 2025 conference on empirical methods in natural language processing, 2025, pp. 27572–27595.
[48]
M. A. Runco, Commentary: Divergent Thinking Is Not Synonymous With Creativity,” Psychology of Aesthetics, Creativity, and the Arts, vol. 2, no. 2, pp. 93–96, May 2008, doi: 10.1037/1931-3896.2.2.93.
[49]
A. Kumar, J. Clune, J. Lehman, and K. O. Stanley, “Questioning representational optimism in deep learning: The fractured entangled representation hypothesis,” arXiv preprint arXiv:2505.11581, 2025.
[50]
J. Lindsey, A. Templeton, J. Marcus, T. Conerly, J. Batson, and C. Olah, “Sparse crosscoders for cross-layer features and model diffing,” Anthropic, 2024. [Online]. Available: https: //transformer-circuits.pub/2024/crosscoders/index.html.
[51]
Y. Cao, B. J. Smucker, and T. J. Robinson, “On using the hypervolume indicator to compare pareto fronts: Applications to multi-criteria optimal experimental design,” Journal of Statistical Planning and Inference, vol. 160, pp. 60–74, 2015, doi: https://doi.org/10.1016/j.jspi.2014.12.004.

  1. Corresponding author: schapironietzsche@gmail.com↩︎

  2. We also tested a prompt-only CAA variant using forward-pass activations from the same creative vs.non-creative prompt sets as CN, but it underperformed the behavioral-data variant on all models and is omitted for clarity.↩︎

  3. [24] studies additional dimensions, such as difficulty, how fun the task is to do, and how fun it is to watch, but here we restrict focus to creativity and originality, as these are most relevant to the goal of evaluating divergent thinking.↩︎

  4. We confirm responses follow a normal distribution before applying the \(t\)-tests.↩︎