← Back to Glossary

Model Reasoning

The Core Distinction: System 1 vs System 2

Borrowing from Daniel Kahneman’s Thinking, Fast and Slow, we categorize LLMs by how they allocate compute during inference:

Characteristic Non-Reasoning Models (System 1) Reasoning Models (System 2)
Paradigm Next Token Prediction Search / Planning / Verification
Latency Low (Single forward pass) High (Multiple forward passes / tokens)
Compute Scaling Training-time only Inference-time (Test-time) Scaling
Failure Mode Hallucination, snap judgments Overthinking, loops, excessive verbosity
Typical Use Chat, Creative Writing, Classification Math, Code, Logic, Multi-step Planning

1. Non-Reasoning Models (Standard LLMs)

These are the “classic” LLMs (GPT-3.5, LLaMA 2/3 Base, Mistral, Claude 2/3 Sonnet/Haiku). They perform one forward pass per token.

How they work

  1. Input Context -> Model -> Probability Distribution over Vocabulary.
  2. Sample Next Token.
  3. Repeat until Stop Token.

The “Fast Thinking” Limitation

When to use them


2. Reasoning Models (LRMs / System 2 Models)

These models (OpenAI o1 / o3, DeepSeek R1, Google Gemini 2.0 Flash Thinking, Qwen QwQ) are trained to generate intermediate reasoning tokens (often called a Chain of Thought - CoT) before producing the final answer.

The Architecture Shift: Inference-Time Scaling

The Scaling Law has moved:

Key Insight: You can take a smaller model and make it smarter at inference time by letting it “think longer” (generate more reasoning tokens).

How they work (High Level)

  1. Prompt -> Reasoning Phase (Hidden or Visible CoT) -> Answer Phase.
  2. The model generates tokens like: “Let me think… Step 1: … Step 2: … Wait, I need to verify… Okay, the answer is X.”
  3. These reasoning tokens attend to themselves, allowing the model to iterate, backtrack, and verify.

Training Recipe (RL on Verifiable Rewards)

  1. Base Model: Strong pre-trained LLM.
  2. Supervised Fine-Tuning (SFT): On high-quality CoT traces (distilled from larger models or human written).
  3. Reinforcement Learning (RL): The Secret Sauce.
    • Environment: Code Execution (Unit Tests), Math Verifiers (Symbolic Equivalence), Formal Proof Checkers (Lean/Coq).
    • Reward: Binary (Correct / Incorrect). No “partial credit” for pretty reasoning.
    • Algorithm: GRPO (Group Relative Policy Optimization) or PPO.
    • Emergent Behaviors: Self-correction (“Wait…”), Backtracking, Verification, Planning.

Types of Reasoning Output

Type Description Example Models
Hidden CoT Reasoning tokens generated but hidden from user (summarized). Used for safety/IP protection. OpenAI o1, o3
Visible CoT Full raw reasoning stream shown to user. Great for debugging/trust. DeepSeek R1, QwQ, Gemini Thinking
Structured/Tool Use Reasoning interleaved with tool calls (Code, Search). Agentic Systems, Claude 3.5 Sonnet (Computer Use)

3. The Trade-offs: Why not always use Reasoning Models?

Cost & Latency

“Overthinking” / Reasoning Collapse

Distillation Gap


4. Benchmarking: The New Frontier

Standard benchmarks (MMLU) saturate. Reasoning models are evaluated on:

Benchmark Domain Why it needs Reasoning
AIME / AMC Competition Math Multi-step derivation, verification
Codeforces / LiveCodeBench Competitive Programming Planning, debugging, edge cases
GPQA Diamond PhD Science Multi-hop knowledge retrieval + logic
ARC-AGI Abstract Reasoning Pattern discovery, generalization
SWE-Bench Software Engineering Repository navigation, test-driven dev

5. Practical Guide: Which to Choose?

flowchart TD
    A[Start] --> B{Is the task deterministic / verifiable?}
    B -- Yes (Math, Code, Logic) --> C{Do you need high throughput / low cost?}
    C -- Yes --> D[Non-Reasoning + Few-Shot CoT Prompting]
    C -- No (Quality Critical) --> E[Reasoning Model (o1, R1, QwQ)]
    B -- No (Creative, Chat, Subjective) --> F[Non-Reasoning Model (Sonnet, GPT-4o, LLaMA 3)]
    B -- Maybe (Complex RAG, Agent) --> G[Reasoning Model for Planning + Non-Reasoning for Execution]

Hybrid Architectures (The Future)


6. Summary Cheatsheet

Feature Non-Reasoning (GPT-4o, Sonnet, LLaMA 3) Reasoning (o1, R1, QwQ)
Mental Model Intuition / Pattern Match Search / Simulation / Proof
Prompting Few-shot, Explicit Instructions Minimal (“Solve this”), let it think
Context Window Standard (128k - 1M) Massive Output needed (Reasoning consumes context)
API Cost $ / 1M Tokens $ / 1M Tokens (but 10-50x tokens used)
Debugging Hard (Black Box) Easy (Read the CoT)
Safety Refusal based Reasoning-based alignment (Model argues with itself)

Key Papers / Resources