Model Reasoning
The Core Distinction: System 1 vs System 2
Borrowing from Daniel Kahneman’s Thinking, Fast and Slow, we categorize LLMs by how they allocate compute during inference:
| Characteristic | Non-Reasoning Models (System 1) | Reasoning Models (System 2) |
|---|---|---|
| Paradigm | Next Token Prediction | Search / Planning / Verification |
| Latency | Low (Single forward pass) | High (Multiple forward passes / tokens) |
| Compute Scaling | Training-time only | Inference-time (Test-time) Scaling |
| Failure Mode | Hallucination, snap judgments | Overthinking, loops, excessive verbosity |
| Typical Use | Chat, Creative Writing, Classification | Math, Code, Logic, Multi-step Planning |
1. Non-Reasoning Models (Standard LLMs)
These are the “classic” LLMs (GPT-3.5, LLaMA 2/3 Base, Mistral, Claude 2/3 Sonnet/Haiku). They perform one forward pass per token.
How they work
- Input Context -> Model -> Probability Distribution over Vocabulary.
- Sample Next Token.
- Repeat until Stop Token.
The “Fast Thinking” Limitation
- No Scratchpad: They cannot “work out” a problem internally. The answer must be encoded in the weights or derived in a single pass through the residual stream.
- Brittle on Composition: If a problem requires Step A -> Step B -> Step C, the model must know the answer to C immediately after seeing the prompt. It cannot “discover” A then B then C.
- Sycophancy: Tendency to agree with user premises even if wrong, because “disagreeing” requires a reasoning step the model doesn’t take.
When to use them
- Creative Tasks: Writing, brainstorming, style transfer (where there is no “right” answer).
- Low Latency / High Throughput: Chatbots, classification, extraction.
- Knowledge Retrieval: “What is the capital of France?” (Facts are compressed in weights).
2. Reasoning Models (LRMs / System 2 Models)
These models (OpenAI o1 / o3, DeepSeek R1, Google Gemini 2.0 Flash Thinking, Qwen QwQ) are trained to generate intermediate reasoning tokens (often called a Chain of Thought - CoT) before producing the final answer.
The Architecture Shift: Inference-Time Scaling
The Scaling Law has moved:
- Old: Performance ∝ Training Compute (Params × Tokens).
- New: Performance ∝ Training Compute + Inference Compute.
Key Insight: You can take a smaller model and make it smarter at inference time by letting it “think longer” (generate more reasoning tokens).
How they work (High Level)
- Prompt -> Reasoning Phase (Hidden or Visible CoT) -> Answer Phase.
- The model generates tokens like: “Let me think… Step 1: … Step 2: … Wait, I need to verify… Okay, the answer is X.”
- These reasoning tokens attend to themselves, allowing the model to iterate, backtrack, and verify.
Training Recipe (RL on Verifiable Rewards)
- Base Model: Strong pre-trained LLM.
- Supervised Fine-Tuning (SFT): On high-quality CoT traces (distilled from larger models or human written).
- Reinforcement Learning (RL): The Secret Sauce.
- Environment: Code Execution (Unit Tests), Math Verifiers (Symbolic Equivalence), Formal Proof Checkers (Lean/Coq).
- Reward: Binary (Correct / Incorrect). No “partial credit” for pretty reasoning.
- Algorithm: GRPO (Group Relative Policy Optimization) or PPO.
- Emergent Behaviors: Self-correction (“Wait…”), Backtracking, Verification, Planning.
Types of Reasoning Output
| Type | Description | Example Models |
|---|---|---|
| Hidden CoT | Reasoning tokens generated but hidden from user (summarized). Used for safety/IP protection. | OpenAI o1, o3 |
| Visible CoT | Full raw reasoning stream shown to user. Great for debugging/trust. | DeepSeek R1, QwQ, Gemini Thinking |
| Structured/Tool Use | Reasoning interleaved with tool calls (Code, Search). | Agentic Systems, Claude 3.5 Sonnet (Computer Use) |
3. The Trade-offs: Why not always use Reasoning Models?
Cost & Latency
- Token Explosion: A reasoning trace can be 10x-100x longer than the answer.
- Cost: Cost = (Input + Reasoning + Output) Tokens × Price.
- Latency: Sequential generation means you wait for all reasoning tokens before the first word of the answer.
“Overthinking” / Reasoning Collapse
- Simple prompts (“Hi”) trigger massive CoT (“The user said Hi. This is a greeting. I should respond politely…”).
- Fix: Routing classifiers or “Reasoning Budgets” (e.g., max_reasoning_tokens: 100).
Distillation Gap
- You cannot easily distill a Reasoning Model into a Non-Reasoning Model of the same size.
- The reasoning process (the search) is what creates the intelligence, not just the final weights.
- Distilling o1 -> 7B model yields a 7B model that outputs CoT style text but lacks the search capability (verification/backtracking).
4. Benchmarking: The New Frontier
Standard benchmarks (MMLU) saturate. Reasoning models are evaluated on:
| Benchmark | Domain | Why it needs Reasoning |
|---|---|---|
| AIME / AMC | Competition Math | Multi-step derivation, verification |
| Codeforces / LiveCodeBench | Competitive Programming | Planning, debugging, edge cases |
| GPQA Diamond | PhD Science | Multi-hop knowledge retrieval + logic |
| ARC-AGI | Abstract Reasoning | Pattern discovery, generalization |
| SWE-Bench | Software Engineering | Repository navigation, test-driven dev |
5. Practical Guide: Which to Choose?
flowchart TD
A[Start] --> B{Is the task deterministic / verifiable?}
B -- Yes (Math, Code, Logic) --> C{Do you need high throughput / low cost?}
C -- Yes --> D[Non-Reasoning + Few-Shot CoT Prompting]
C -- No (Quality Critical) --> E[Reasoning Model (o1, R1, QwQ)]
B -- No (Creative, Chat, Subjective) --> F[Non-Reasoning Model (Sonnet, GPT-4o, LLaMA 3)]
B -- Maybe (Complex RAG, Agent) --> G[Reasoning Model for Planning + Non-Reasoning for Execution]
Hybrid Architectures (The Future)
- Routing: Small classifier decides: “Route to Fast Model” vs “Route to Reasoning Model”.
- Budget Forcing: “Think for exactly 500 tokens, then answer.”
- Speculative Reasoning: Fast model drafts reasoning, Reasoning model verifies.
6. Summary Cheatsheet
| Feature | Non-Reasoning (GPT-4o, Sonnet, LLaMA 3) | Reasoning (o1, R1, QwQ) |
|---|---|---|
| Mental Model | Intuition / Pattern Match | Search / Simulation / Proof |
| Prompting | Few-shot, Explicit Instructions | Minimal (“Solve this”), let it think |
| Context Window | Standard (128k - 1M) | Massive Output needed (Reasoning consumes context) |
| API Cost | $ / 1M Tokens | $ / 1M Tokens (but 10-50x tokens used) |
| Debugging | Hard (Black Box) | Easy (Read the CoT) |
| Safety | Refusal based | Reasoning-based alignment (Model argues with itself) |
Key Papers / Resources
- “Learning to Reason with LLMs” (OpenAI o1 Blog, 2024)
- “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning” (2025)
- “Scaling Laws for Neural Language Models” (Kaplan et al., 2020) vs “Test-Time Scaling” (Snell et al., 2024)
- “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” (Wei et al., 2022)