All articles Research

Reasoning Models and Test-Time Compute

How the field shifted from scaling training to spending compute at inference — chain-of-thought as computation, RLVR, and the sampling, verification, and search strategies that turn extra thinking into better answers.

For most of the deep-learning era, the recipe for a better model was simple: more parameters, more data, more training FLOPs. Reasoning models break that framing. They show that once a model can generate a long, structured chain of intermediate steps, you can buy additional accuracy at inference time — by letting the model think longer, sample more candidates, and select among them — without touching the weights at all. This article explains the mechanics of that shift, the training regime (RLVR) that made it reliable, the concrete test-time strategies you can deploy, and the tradeoffs that decide when extra thinking is worth paying for.

Test-time compute turns inference from a fixed cost into a dial: spend more tokens or more samples, and — up to a point — buy more accuracy.

The shift: train-time vs test-time scaling

The scaling laws that dominated the last decade describe how loss falls as you increase parameters, data, and training compute together. That relationship is real, but it is also front-loaded: the cost is paid once, during pretraining, and every inference thereafter is cheap and fixed. Test-time scaling is a complementary axis. Instead of making a single forward pass smarter through bigger weights, you let a fixed model do more work per query — generating longer reasoning, drawing multiple samples, or searching over candidate solutions.

The two axes are not rivals; they compose. A model still has to be pretrained well enough to produce coherent reasoning steps. But once it can, test-time compute often gives a steeper accuracy-per-dollar curve on hard problems than an equivalent spend on more pretraining would — because the extra compute is targeted precisely at the queries that need it.

DimensionTrain-time scalingTest-time scaling
When compute is spentOnce, during trainingPer query, at inference
What changesModel weightsAmount of computation per prompt
Cost profileLarge fixed cost, cheap fixed inferenceCheap to enable, variable per-query cost
AdaptivitySame effort for every inputCan spend more on hard inputs, less on easy ones
Typical leverParameters, tokens, epochsReasoning length, samples, search width
Main riskDiminishing returns, data/compute limitsLatency and token cost; returns plateau

Chain-of-thought as computation

Chain-of-thought (CoT) prompting — asking a model to work through intermediate steps before answering — was originally framed as a prompting trick that improved accuracy on arithmetic and multi-step reasoning. The deeper interpretation is that each generated token is an opportunity to do computation. A transformer has a fixed depth per forward pass; it cannot perform an unbounded number of sequential operations to answer in a single step. By emitting intermediate tokens, the model uses its own output as a scratchpad, effectively unrolling a longer computation across many forward passes.

This reframing matters. If reasoning tokens are computation, then the length of the reasoning trace is a compute budget you can adjust. Short traces solve easy problems cheaply; long traces give the model room to decompose, backtrack, check intermediate results, and reconsider — behaviors that look like search carried out in natural language.

From prompted CoT to trained "thinking"

Early CoT was elicited by prompting ("let's think step by step"). Modern reasoning models internalize the behavior: they are trained to produce a long reasoning trace by default, often demarcated from the final answer so that the trace can be truncated, hidden, or budgeted independently. The important qualitative change is that these traces are not just longer — they contain self-correction: the model proposes an approach, notices an error, and revises. That behavior is what makes additional thinking tokens pay off rather than merely padding the output.

Reference points: o1 and DeepSeek-R1

Two systems anchor the public conversation about reasoning models.

  • OpenAI o1 was the first widely deployed model explicitly marketed around inference-time reasoning. Its central claim was that accuracy on hard math, coding, and science problems improves as the model is allowed to spend more compute thinking before answering. The detailed reasoning trace is kept internal, and users interact with a summarized result.
  • DeepSeek-R1 is notable because its training methodology was described openly. It demonstrated that long-reasoning behavior — including emergent self-verification and longer traces on harder problems — can be induced primarily through reinforcement learning against checkable rewards, rather than requiring an enormous corpus of hand-written reasoning demonstrations. Its companion exploration (a variant trained with RL from a base model with minimal supervised warm-up) showed that reasoning can emerge from RL alone, though a small amount of supervised data improves readability and stability.

The takeaway from both is the same: the capability to reason at length is trainable, and once present, it exposes an inference-time knob that trades compute for accuracy.

RLVR: reinforcement learning with verifiable rewards

The training ingredient that made reasoning reliable is reinforcement learning with verifiable rewards (RLVR). The idea is to restrict RL to domains where the correctness of a final answer can be checked automatically and cheaply:

  • Math — the final numeric or symbolic answer can be compared against a known solution, or checked with a computer-algebra system.
  • Code — a candidate program can be run against unit tests; passing or failing is an objective signal.
  • Formal tasks — anything with a machine-checkable acceptance condition (a proof checker, a constraint solver, a parser).

Because the reward comes from a verifier rather than a learned preference model, it is far less susceptible to reward hacking on the final answer, and it scales without human labeling. The model samples reasoning traces, the verifier scores the final answers, and a policy-gradient method (commonly a PPO-style or group-relative variant) pushes the policy toward traces that lead to correct answers.

Why verifiability is the crux

RLVR works precisely where the reward is unambiguous. That is also its main limitation: the technique is strongest in math, code, and formal reasoning, and weaker in open-ended domains — summarization, dialogue quality, subjective judgment — where "correct" is not machine-checkable. A subtle failure mode is that the outcome reward only checks the final answer, not the reasoning; a model can reach the right answer through flawed steps, so a correct answer does not guarantee a sound trace.

Test-time strategies

Given a trained reasoning model, several inference-time strategies convert extra compute into accuracy. They can be combined.

Longer reasoning (sequential compute)

The simplest lever is to let a single trace run longer — a larger thinking budget. This is sequential compute: one chain that decomposes and self-corrects more thoroughly. It benefits problems that genuinely require many dependent steps, but returns fall off once the model has "found" the solution.

Best-of-N sampling

Draw N independent samples (with temperature > 0) and select one. This is parallel compute. Selection requires a criterion — a verifier, a reward model, or agreement among samples. Best-of-N shines when the model can sometimes reach the right answer; more draws raise the chance that at least one is correct, provided you can identify it.

Self-consistency (majority voting)

A special case of best-of-N that needs no external verifier: sample many traces, extract each final answer, and take the majority vote. It exploits the observation that correct answers tend to be reached by many independent reasoning paths, while errors are scattered. It works well when answers are discrete and comparable (a number, a label, a short expression) and poorly for free-form outputs where "the same answer" is hard to define.

Verifier / reward-model reranking

Instead of voting, score each candidate with a separate model and keep the top-ranked one. Two flavors matter:

  • Outcome reward models (ORMs) score the final answer only.
  • Process reward models (PRMs) score intermediate steps, giving denser signal and enabling step-level selection or pruning.

Reranking can beat majority voting when a good verifier exists — especially in domains where the "majority" answer is systematically biased — but it adds the cost and failure modes of a second model.

Tree / search-style exploration

Rather than sampling whole traces independently, treat reasoning as a search tree: expand promising partial traces, score them (often with a PRM), prune weak branches, and optionally backtrack. Beam search, lookahead, and Monte-Carlo-Tree-Search-style variants fall here. Search allocates compute adaptively toward promising regions of the solution space and can outperform naive best-of-N at equal budget, at the cost of substantial orchestration complexity and many partial-trace evaluations.

A generic inference pipeline

Most of these strategies fit one template: generate reasoning, optionally produce multiple candidates, evaluate them, and select an answer. The diagram below shows the general flow; a plain long-CoT model is just the degenerate case where N = 1 and selection is trivial.

flowchart TD A[Prompt / problem] --> B[Generate reasoning trace] B --> C{Test-time budget?} C -->|Single trace| D[Take final answer] C -->|Multiple samples| E[Sample N traces in parallel] E --> F{Selection method} F -->|Majority vote| G[Self-consistency] F -->|Score candidates| H[Verifier / reward-model rerank] F -->|Expand and prune| I[Tree search over partial traces] G --> J[Selected answer] H --> J I --> J D --> K[Return answer] J --> K

Methods, cost, and when to use them

The strategies differ sharply in cost and in the conditions under which they help. A rough guide:

MethodCompute typeRelative costNeeds a verifier?When to use
Longer reasoning trace Sequential Low–moderate (more tokens, one call) No Problems with many genuinely dependent steps; default first lever
Best-of-N sampling Parallel ~N× generation, plus selection Yes (to pick the winner) Model sometimes succeeds and you can score outputs
Self-consistency (majority vote) Parallel ~N× generation, negligible selection No Discrete, comparable answers (math results, labels)
Verifier / reward-model rerank Parallel + scoring ~N× generation + N scoring passes Yes (ORM or PRM) A trustworthy verifier exists; majority vote is biased
Tree / search exploration Adaptive parallel Highest; many partial-trace evaluations Usually (PRM for step scores) Large solution spaces; budget and orchestration available

A practical rule: reach for sequential compute (longer thinking) first because it is the cheapest to enable, add self-consistency when answers are discrete and you lack a verifier, and escalate to reranking or search only when a good verifier justifies the extra parallel cost.

Distilling reasoning into smaller models

Running a large reasoning model with a big thinking budget is expensive per query. A powerful follow-on result is that the behavior can be compressed: generate long, high-quality reasoning traces with a strong teacher model (ideally filtered so you keep only traces whose final answers verify as correct), then fine-tune a much smaller student on those traces. The student learns to imitate the reasoning style — decomposition, self-checking, backtracking — and inherits a large share of the teacher's reasoning ability at a fraction of the inference cost.

The DeepSeek-R1 work popularized this: distilled dense models trained on traces from the larger RL-trained model showed strong reasoning on math and code benchmarks despite being far smaller. The practical implication for engineers is a deployment pattern — run expensive RL and long-reasoning generation once with a large model, then serve distilled students that are cheap enough for production. It is worth noting that distillation transfers the observable trace behavior; the student does not receive the reward signal directly, so its ceiling is bounded by the quality and coverage of the teacher's traces.

Practical tradeoffs and when thinking stops helping

Test-time compute is a dial, not free accuracy. The costs are concrete:

  • Latency. Longer traces mean more sequential decoding steps, which directly increases time-to-answer. Best-of-N and search can be parallelized across samples, trading money for wall-clock time, but longer single traces cannot.
  • Token cost. Reasoning tokens are billed like any other. A model that thinks for thousands of tokens before a one-line answer can cost far more per query than its output length suggests, and sampling N traces multiplies that.
  • Diminishing and negative returns. Accuracy versus compute typically rises then flattens. Beyond a problem-dependent point, extra thinking adds cost without accuracy — and can even hurt, as an over-long trace talks itself out of a correct early answer ("overthinking"). Self-consistency and best-of-N saturate once the correct answer already dominates the sample distribution.

Knowing when to stop

Because the plateau is problem-dependent, the mature move is to spend adaptively rather than using a fixed global budget:

  • Route by difficulty — cheap short-CoT (or a small model) for easy queries, large budgets only for hard ones.
  • Use an early-exit signal — stop sampling once majority vote stabilizes or a verifier is confident.
  • Cap the thinking budget per query and expose it as a tunable knob, so latency- and cost-sensitive paths can dial it down.

The engineering goal is not maximum thinking; it is spending the right amount of compute on each query.

Takeaways

  • Reasoning models make inference compute a first-class scaling axis, complementary to (not a replacement for) train-time scaling.
  • Chain-of-thought is best understood as computation: reasoning tokens unroll a longer, self-correcting computation across forward passes.
  • RLVR — RL against automatically checkable rewards in math and code — is what made long reasoning reliable and label-efficient; its limits are domains where correctness is not machine-checkable.
  • o1 and DeepSeek-R1 are the anchoring reference points; DeepSeek-R1 additionally showed the training recipe and distillation into smaller models openly.
  • Test-time strategies form a toolkit — longer traces, best-of-N, self-consistency, verifier reranking, tree search — with very different cost profiles and preconditions.
  • Extra thinking has diminishing and eventually negative returns; the discipline is adaptive, budgeted compute, not maximal compute.
← Back to all articles