When I ask a local model for thirty song ideas, I turn its thinking mode off. Reasoning models are better at most tasks, so this sounds backwards. Brainstorming, though, asks for spread, and reasoning is trained to reduce it.
What thinking mode changes
Models such as Qwen3, DeepSeek-R1 and several recent Gemma and gpt-oss releases can write a private chain of reasoning before the answer. Training rewards that reasoning for reaching correct, defensible answers. On a maths problem that's what you want. The model explores, checks itself, and converges.
A brainstorm has no single correct answer. The useful output is a wide set of candidates, including some odd ones, for a person or a scoring step to choose from. When the model reasons first, it tends to plan the list: pick categories, cover them evenly, and discard ideas that look weak. The planning produces tidy lists that read like the same list every time you ask.
There's also a sampling side. Vendors often recommend different settings for the two modes. Qwen's card for Qwen3-30B-A3B suggests temperature 0.6 and top-p 0.95 with thinking, and temperature 0.7 and top-p 0.8 without (model card). Whatever the mode, the diversity you get depends on those settings as much as on the reasoning, so compare modes at matched settings before drawing conclusions.
How to disable thinking in Qwen3 and Ollama
The switch depends on the model and the server:
- Qwen3 through Hugging Face templates: pass
enable_thinking=Falsetoapply_chat_template. With thinking enabled,/no_thinkin a message switches it off for that turn, but the card notes the two don't behave identically (Qwen3 card). - Ollama: send
"think": falsein the chat request, or use--think=falseon the command line (Ollama thinking docs). Ollama ignores the flag for models that don't support it, so check what came back. - Other servers: look for a chat-template flag or a reasoning-effort setting. Some models can only reduce reasoning, not remove it.
Whichever you use, confirm it worked. A model that "doesn't think" but returns a long preamble, or a server that puts all the text in a thinking field and leaves the answer empty, will quietly break a pipeline. My model-calling layer treats an empty answer with a full thinking field as an error for exactly this reason.
A split that works
I don't give up reasoning. I move it.
- Generate wide: thinking off, the model's recommended non-thinking sampling or a little warmer, many candidates.
- Remove duplicates and near-duplicates with embeddings.
- Judge narrow: thinking on, or a different model, scoring each idea against criteria written down in advance.
Reasoning is good at evaluating, and a divergent generator is good at supplying things worth evaluating. Doing both in one call gets you a careful model editing its own ideas before you see them.
Measure it on your own task
The effect depends on the model and the task, so measure it. This experiment takes an afternoon.
- Pick one brainstorming prompt from your real work, for example "30 names for a feature that lets teams approve AI drafts".
- Run it 10 times with thinking on and 10 times with thinking off, at matched sampling settings. Then run it 10 more times with thinking off at each model's recommended non-thinking settings.
- Pool each condition's ideas. Measure spread two ways: the share of distinct ideas after normalizing case and punctuation, and the average pairwise cosine distance of their embeddings (any local embedding model will do).
- Ask a person to mark ideas they'd actually consider, blind to the condition.
- Record time and tokens per run. Thinking costs both.
Report all three numbers together. A condition that wins on spread but loses on "ideas a person would consider" is producing noise. With song ideas, thinking off has given me more ideas worth reading per minute of generation. That's one person's taste on one kind of prompt, and it's why the experiment includes a blind human pass.
When to keep thinking on
Keep it for brainstorms with hard constraints ("names under 12 characters that aren't existing trademarks in our list"), where checking matters more than spread; for technical options that need to be feasible; and whenever the list will be used without a person or a judge in between.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real