Reasoning & Test-Time Compute · 2022
Chain-of-Thought Prompting Elicits Reasoning
Showed that prompting a large model to emit intermediate reasoning steps before its answer unlocks multi-step reasoning that direct-answer prompting fails at, without any fine-tuning.
Editorial record
Plain-language summary
By putting a few exemplars that spell out step-by-step worked solutions into the prompt, the model imitates that format and reasons through arithmetic, commonsense, and symbolic problems one step at a time. The benefit appears mainly at large model scale and substantially raised accuracy on benchmarks like GSM8K math word problems. It made intermediate-computation prompting a standard, training-free way to get harder reasoning out of existing models.
Knowledge graph
Relationships
Antecedents
ExtendsEvidence: Direct
Self-Consistency Improves Chain-of-Thought Reasoning
Self-consistency votes over many CoT samples
P-221
ExtendsEvidence: Direct
Structured Reasoning: Least-to-Most / PoT / Tree of Thoughts
ToT/L2M/PoT structure reasoning beyond a linear chain
P-222
Depends onEvidence: Strongly supported
DeepSeek-R1: Incentivizing Reasoning via RL (RLVR)
RLVR learns long chains of thought
P-224
ChallengesEvidence: Direct
Reasoning Correctives: Faithfulness / Overthinking / Contamination
CoT can be unfaithful; reasoning correctives
P-226
CombinesEvidence: Direct
ReAct: Synergizing Reasoning and Acting in LMs
ReAct combines reasoning with acting
P-230
ExtendsEvidence: Strongly supported
Solving Quantitative Reasoning Problems with Language Models (Minerva)
Minerva extends chain-of-thought to quantitative and mathematical reasoning
P-522
Conceptual ancestorEvidence: Strongly supported
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
s1: Simple Test-Time Scaling
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
LIMO: Less is More for Reasoning
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Builds on the reasoning lineage in the archive
freshness sweep 2026
Conceptual ancestorEvidence: Strongly supported
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Builds on the reasoning lineage in the archive
freshness sweep 2026
Descendants
EnablesEvidence: Direct
Language Models are Few-Shot Learners (GPT-3)
Chain-of-thought emerges only at scale
P-220
Source record
Provenance
- Record ID
- P-220
- Record created
- 2026-07-13
- Last reviewed
- 2026-07-14
- Record version
- 2
- https://arxiv.org/abs/2201.11903
- arXiv:2201.11903
Citation caveat: Citation metadata is approximate and marked unverified in the source dataset.