Prompting, reasoning and verification

Intermediate

Choose techniques by model and task, and check whether they improve measured outcomes.

Last updated: Sep 13, 2026

Beyond the Basics

Start with a direct task, the necessary source material and a way to check success. Reasoning models already spend internal computation on a problem; classic instruct models may benefit from worked examples. Compare approaches using the same inputs, model version and evaluation criteria.

Chain of Thought

Chain-of-thought prompting is a historically studied technique for eliciting intermediate reasoning, especially in classic instruct models. For reasoning models, follow the provider guidance: start with direct instructions rather than requesting a hidden thought transcript.

Example: “Solve the problem. Give the result, relevant assumptions and a short check that I can verify.”

Important: CoT is Not Universal

A useful answer explanation and a faithful account of internal reasoning are different things. Do not assume that an explicit chain of thought improves every model or task.

  • -Reasoning models generally do not need “think step by step”; check the model-specific guidance.
  • -A convincing explanation can still be wrong or omit the actual influence on a decision.
  • -Measure accuracy, latency and token use against a direct-prompt baseline.

Few-Shot Learning

Provide multiple examples to establish patterns the model should follow.

Start without examples. If a format or boundary is repeatedly misunderstood, add a few representative cases and retest held-out inputs.

Self-Consistency

Sample multiple candidate answers and aggregate comparable final answers. Shared mistakes can win a majority, so voting is not a substitute for external verification.

For an arithmetic task, compare candidate answers with an exact calculator rather than treating the majority as proof.

Task Decomposition

Break complex tasks into smaller, manageable sub-tasks.

Solve sub-tasks independently, then combine results.

Tree of Thoughts (ToT)

Tree of Thoughts (2023) is a search framework: generate candidate intermediate states, evaluate them and explore or prune branches. Its reported results apply to the paper’s tasks and models.

How It Works

Generate multiple reasoning branches at each step. Evaluate promising paths, prune dead ends, and backtrack to explore alternatives.

Best For

Planning problems, puzzles, creative tasks requiring exploration, and problems where the first approach may not be optimal.

Implementation sketch: generate candidates → evaluate constraints → choose or backtrack → continue until the stopping criterion. A single “consider three approaches” prompt does not implement this search.

Graph of Thoughts (GoT)

Graph of Thoughts (2023) is an orchestration framework that can combine and revisit intermediate outputs in a graph. It does not establish that a language model reasons like a human brain.

Key Feature

Unlike linear CoT or tree-structured ToT, GoT allows combining insights from different reasoning paths and revisiting earlier conclusions.

Best For

Complex problems with interdependencies, synthesis tasks, and problems where partial solutions need to be combined.

Implementation sketch: run independent branches → store their outputs → merge selected results → evaluate → refine. The graph and execution policy are part of the application.

Cost-Benefit Analysis

Advanced prompting techniques increase token usage, latency, and API costs. Understanding when these tradeoffs are worthwhile is crucial for production systems.

Worth the Extra Cost

  • +Complex reasoning: math, logic, multi-step analysis
  • +Tasks with a trustworthy external checker and a measured benefit from extra computation
  • +Problems where accuracy matters more than speed

Often Not Worth It

  • -Simple classification or extraction tasks
  • -High-volume, latency-sensitive applications
  • -Tasks where simpler prompts already achieve high accuracy

Tip: Start with simple prompts and add complexity only when needed. Measure the accuracy improvement against the cost increase to make informed decisions.

Additional Techniques

Role Assignment

Specify responsibility and audience when useful. Calling the model an expert does not establish domain competence.

Explicit Constraints

List what the model should NOT do to prevent common errors.

Self-Verification

Ask for a short check, then verify with source documents, executable tests or another independent reference. Self-review alone is not proof.

🧠

Verify a candidate answer

Check two constraints locally; no invented model comparison.

Check an answer with equations

A local arithmetic checker, not a model comparison. An explanation can sound convincing and still fail a constraint. Reasoning models usually do not need a “think step by step” instruction.

A bat and a ball cost 110 cents together. The bat costs 100 cents more than the ball. What does the ball cost?

Bat price derived from the difference
10 + 100 = 110
Total price
10 + 110 = 120 ≠ 110

The total constraint fails. Adjust the candidate.

Key Takeaways

  • 1Separate classic instruct prompting from reasoning-model guidance.
  • 2Examples must solve an observed problem and survive held-out tests.
  • 3ToT and GoT include external search or orchestration, not just a longer prompt.
  • 4Always consider cost vs. benefit—advanced techniques increase token usage
  • 5Start simple, add complexity only when accuracy requires it

Primary sources