Reasoning Models Explained: When Thinking Longer Helps
Fundamentals

Reasoning Models Explained: When Thinking Longer Helps

Reasoning models spend extra tokens working before they answer. Here is what that changes mechanically, which tasks it helps, and where it is just an expensive delay.

A reasoning model generates a working-out phase before it answers. Those intermediate tokens are produced by the same next-token machinery as everything else, but they are trained rather than prompted, and you pay for them.

Whether that trade is worth it depends almost entirely on the task, and the gap between where it helps and where it does not is larger than most teams assume.

Three ideas that combined

The field assembled reasoning models out of parts that already existed.

Chain of thought was the prompting observation: asking a model to work step by step improved accuracy on multi-step problems. Each intermediate token becomes part of the context for the next, so the model conditions on its own partial work instead of trying to jump to an answer.

Test-time compute was the reframing: instead of only scaling parameters and training data, spend more compute at inference. A model that generates more tokens before answering has more opportunity to explore, check and correct.

Reinforcement learning was the step that made it native. Rather than prompting the model to reason, train it to, rewarding reasoning traces that lead to verifiably correct answers. OpenAI o1-preview, released in September 2024, was the first widely available commercial model with this behaviour internalised, and the approach spread rapidly across labs afterwards.

The result is a model that produces its own scratchpad by default, with a length it decides based on apparent difficulty.

Why verifiability shapes the capability profile

The training signal for reasoning comes from problems where correctness can be checked automatically — mathematics, competitive programming, unit tests, formal proofs. You can generate many attempts, grade them mechanically, and reinforce the traces that worked.

That constraint explains the shape of the gains almost perfectly. Reasoning models are strongest on tasks that resemble the training signal: problems with a definite answer, reachable by decomposition, where an intermediate error is detectable.

They are much weaker on tasks with no such structure. Writing quality, judgement calls, summarisation and open-ended advice do not improve much from extended deliberation, because there is nothing for the model to check its work against. This is not a temporary limitation; it follows from how the capability was created.

What extra thinking actually buys

Reliably helps:

  • Multi-step mathematics and logic, where the answer is a composition of verifiable steps.
  • Algorithmic problems with a correctness criterion.
  • Debugging, where the model can hypothesise a cause, trace consequences, and rule it out.
  • Planning under stated constraints, where a candidate plan can be checked against each constraint.
  • Constraint satisfaction generally — scheduling, allocation, configuration.

Rarely helps enough to justify the cost:

  • Extraction and classification. The answer is in the input. Deliberation adds latency and occasionally talks the model out of a correct first instinct.
  • Format conversion and rewriting. Mechanical transformations do not have steps to reason about.
  • Recall. Thinking longer does not add a fact that was never in the weights. A long confident derivation from a false premise is harder for a reviewer to catch than a short one.
  • Simple tool selection. If the right tool is obvious, deliberating about it is pure overhead in an agent loop.

The costs are not only money

Tokens. Reasoning tokens are billed as output, which is the expensive direction, and they can dwarf the visible answer. A short reply may be preceded by thousands of tokens you never see. Budget on total output, not on response length.

Latency. Time to first visible token can stretch from under a second to tens of seconds. That is fine for a batch job and unacceptable for an interactive autocomplete. Some products bridge the gap by streaming a reasoning summary so the user sees progress.

Opacity. Several providers hide raw reasoning traces and expose only summaries, so you cannot always inspect the path to an answer. Where traces are visible, treat them as a rationalisation of the process rather than a faithful log of it — the trace is generated text too.

Context consumption. In multi-turn work, reasoning tokens compete for the same window as everything else. Long-context models help here: GLM 5.2 pairs a 1M-token context with up to 128K output tokens, which is the kind of headroom extended deliberation actually needs.

Effort levels and when to turn it down

Most reasoning models now expose some control over how much thinking to do — an effort or budget parameter, or a toggle. Use it as a routing decision rather than a global setting.

A workable policy in a mixed workload: default to low or no reasoning, and escalate on signals. Escalate when a first attempt fails validation, when the task involves more than a couple of dependent steps, when constraints must be satisfied simultaneously, or when the cost of a wrong answer is high. Do not escalate because the input is long — length is not difficulty.

One practical note on prompting: the old chain-of-thought instructions are unnecessary and sometimes counterproductive with these models. Telling a reasoning model to "think step by step" duplicates what it already does. Give it the problem, the constraints and the success criteria, and let it allocate its own budget.

The cost dynamics are also why flat-rate access is worth considering for exploratory or agent-heavy work — reasoning token counts are hard to predict per request, which makes per-token budgets awkward to plan against. That is a billing preference, not a capability argument.

How to decide, concretely

  1. Run your task on a fast non-reasoning model first and measure accuracy. This is your baseline and it is often good enough.
  2. Run the same task on a reasoning model. Compare accuracy, total output tokens and end-to-end latency.
  3. Compute cost per correct answer, not cost per call. A model that is three times the price and halves the error rate may still be cheaper once you count retries and human review.
  4. If the gain is small, route selectively: cheap model by default, reasoning model on failure or on flagged inputs.

The mistake worth avoiding is treating reasoning as a general upgrade. It is a specific capability with a specific shape, bought with tokens and latency. On the tasks it fits, the improvement is substantial. On the tasks it does not, you have bought a slower, more expensive version of the same answer.

Common questions

Should I always use a reasoning model?

No. They help on multi-step problems with checkable answers and add cost and latency everywhere else. Extraction, classification and format conversion generally get nothing from extended deliberation.

Do I pay for reasoning tokens?

Yes, typically at output rates, even when the trace is hidden from you. A short answer can be preceded by thousands of billed tokens, so budget against total output rather than visible response length.

Should I still prompt for step-by-step thinking?

Not with a reasoning model. It already produces its own working-out, and explicit chain-of-thought instructions are redundant and occasionally disruptive. Supply the problem, the constraints and the success criteria instead.

Similar articles

Chain of Thought: Why Thinking Out Loud Actually Helps
Fundamentals
Fundamentals·8 min read

Chain of Thought: Why Thinking Out Loud Actually Helps

Asking a model to reason step by step measurably improves accuracy on some tasks and wastes tokens on others. The mechanism, and when it is worth the cost.

Read
Test-Time Compute: Buying Accuracy With Tokens Instead of Size
Fundamentals
Fundamentals·9 min read

Test-Time Compute: Buying Accuracy With Tokens Instead of Size

Spending more computation at inference can substitute for a larger model. How the trade works, where it pays off, and what it does to your latency budget.

Read
Batch Size and Throughput: The Trade Behind Every Token Price
Fundamentals
Fundamentals·9 min read

Batch Size and Throughput: The Trade Behind Every Token Price

Batching is why per-token prices are low and why latency varies under load. How it works, where it stops helping, and what it means for your requests.

Read