Skip to main content
Workforce LibreTexts

3: AI Reasoning

  • Page ID
    67038
  • \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \( \newcommand{\dsum}{\displaystyle\sum\limits} \)

    \( \newcommand{\dint}{\displaystyle\int\limits} \)

    \( \newcommand{\dlim}{\displaystyle\lim\limits} \)

    \( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)

    ( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\id}{\mathrm{id}}\)

    \( \newcommand{\Span}{\mathrm{span}}\)

    \( \newcommand{\kernel}{\mathrm{null}\,}\)

    \( \newcommand{\range}{\mathrm{range}\,}\)

    \( \newcommand{\RealPart}{\mathrm{Re}}\)

    \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)

    \( \newcommand{\Argument}{\mathrm{Arg}}\)

    \( \newcommand{\norm}[1]{\| #1 \|}\)

    \( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)

    \( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)

    \( \newcommand{\vectorA}[1]{\vec{#1}}      % arrow\)

    \( \newcommand{\vectorAt}[1]{\vec{\text{#1}}}      % arrow\)

    \( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \( \newcommand{\vectorC}[1]{\textbf{#1}} \)

    \( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)

    \( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)

    \( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)

    \( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)

    \(\newcommand{\longvect}{\overrightarrow}\)

    \( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)

    \(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)

      Chapter 2 treated the single request: how to word it, what to supply with it, and why the same request can produce different answers. This chapter takes up the two things that most often separate a frustrating exchange from a productive one. The first is treating an interaction as a conversation rather than a transaction. The second is asking the model to work a problem through in visible steps instead of answering in one move. Both techniques share an origin in the mechanism described in Chapter 1. Because a model predicts each word from everything currently in front of it, anything it has already written becomes part of what it reads next. That single fact is what makes step-by-step reasoning work at all, and it explains both the power and the limits of the techniques below.

      Reasoning Vs. Learning

      It is worth restating a point from Chapter 1 that students frequently misremember. A model does not learn while you are talking to it. Training happened beforehand; the exchange you are having now does not update it. This raises an obvious question: how does such a system work through a problem it has never encountered? Part of the answer is that when a model is asked to reason step by step, it writes its reasoning out as text, and each step it writes becomes part of the context it reads before producing the next one. The analogy is exact and worth keeping. A person facing a difficult arithmetic problem does not become more intelligent by writing it down; they offload the burden of holding intermediate steps in memory, and the paper becomes part of the thinking process. A model asked to show its work is doing the same thing. A familiar puzzle illustrates the effect. A bat and a ball cost $1.10 in total; the bat costs $1.00 more than the ball; how much does the ball cost? The intuitive answer, ten cents, is wrong — working it through gives five cents. A model asked to answer immediately frequently makes the same intuitive error a person does. The same model, asked to work through it in steps, generally catches it. One clarification matters more than it may first appear. The model is not reasoning privately and then reporting a conclusion. There is no hidden deliberation occurring behind the visible text. The reasoning exists only because it is being written into the context, one word at a time, by exactly the same prediction process that produces every other word. Many current tools now perform this step-by-step work automatically without being asked, because it reliably improves accuracy — but the underlying mechanism is unchanged.

      From Answers to Processes

      Asking directly for a final answer is a reasonable default for simple requests and a poor one for complex tasks, because it gives the model no opportunity to catch itself. Specifying a process instead of requesting an outcome changes the result measurably on multi-step calculations, planning tasks, and any problem where an early error propagates. The simplest form is to add a phrase such as "work through this step by step" or "show your work" to the prompt, which forces the intermediate reasoning into the output where it can be inspected. A second and complementary technique is to treat the first output as a draft: rather than starting over when a response is unsatisfactory, ask the model to critique its own output and then revise it. This two-step sequence frequently produces a better result than attempting to specify everything correctly in one pass, and it mirrors how people revise their own work. Both techniques have a cost in time and length, and neither is universally appropriate. They earn their cost where errors are expensive or the problem genuinely has several stages. For a simple factual lookup, or for open-ended creative work where structure constrains rather than helps, imposing a visible process adds little.

      What Step-by-Step Prompting Does and Does Not Do

      The technique described above is known in the industry as chain-of-thought prompting, and it comes in two forms: adding a phrase requesting step-by-step reasoning with no examples, or supplying one or two worked examples that include the reasoning rather than only the answer. It helps because a difficult problem may not resemble anything in the model's training data while its individual sub-steps almost certainly do. Working through the steps allows the model to route through material it has genuinely encountered before, rather than attempting a leap it cannot make. A further refinement, useful when reliability matters more than cost, is to generate several independent step-by-step answers to the same question and take the most common result. Two limits deserve emphasis, and both are easy to forget precisely because the output looks so convincing. Step-by-step prompting improves how a model works through what it has; it cannot supply information the model never had, and it will not rescue a task the model fundamentally cannot do. On such tasks it produces confident, well-structured, incorrect reasoning. A visible chain of steps is therefore a signal worth checking, not evidence that the answer is right — and a longer, more elaborately reasoned response is not automatically a better one.

      Why the Same Technique Works Better on Different Models

      A practical complication follows. Give two different models the same complex problem and the same step-by-step instruction, and one may work through it cleanly while the other becomes tangled in its own logic and produces confident nonsense. The largest single factor is scale. A model trained on more data with more internal capacity holds a richer representation of language and reasoning patterns, and can therefore follow a chain of reasoning further before losing the thread. Smaller models are not incapable of the technique; they simply have less depth to draw on and lose their way sooner. The composition of the training material matters as well: a model trained on a substantial body of scientific writing, mathematics, and code tends to reason more reliably than one trained on a narrower mixture, and some developers deliberately reward multi-step reasoning during training to strengthen the capability. The practical conclusion is worth stating plainly, because it is easy to over-invest in prompting technique alone. Prompting is analogous to software and the underlying model to hardware: the best step-by-step prompt available still runs better on a more capable model. Where reasoning reliability genuinely matters, the choice of tool is at least as consequential as the wording of the request.

      A Different Kind of Tool: Calculators, Search, and Other Outside Help

      In the context of AI reasoning, the word "tool" refers to an outside resource a model reaches for when reasoning alone will not get the job done. Step-by-step prompting, for all its benefit, cannot fix every kind of error, and arithmetic is the clearest case. A language model is not a calculator; it is a system that predicts the most probable next word, and the most probable answer to a math problem is not always the correct one. This limitation runs deeper than sloppy prompting can explain: a model can write out a clean, convincing chain of steps and still land on the wrong number, because the internal process that actually produces its answer does not necessarily match the reasoning it describes on the page. The bat-and-ball problem earlier in this chapter is a preview of the same issue at a larger scale — arithmetic is exactly the kind of precise, rule-bound task a probabilistic word predictor was never built to guarantee.

      This diagram titled "Tool Use For Complex Mathematics" illustrates how an AI model handles complex math problems by failing to process partial differential equations directly, resulting in an error, and instead making an API call to an external calculator tool to execute the calculation and ingest the results.

      Illustration generated by Google's Nano Banana 2 image generation model.

      A tool, in this sense, is an external system the model can call on for information or action: a calculator, a search engine, a calendar, a database. Handing a model tools is a qualitative shift, not just an incremental improvement — it moves the system from one that can only produce answers to one that can take actions and retrieve current facts. Ask a model unaided about yesterday's news and it cannot know, because its knowledge stops at its training cutoff; give it a search tool and it can recognize the need, retrieve current information, and fold the result into its answer. Ask it to compute a loan payment unaided and it may confidently produce the wrong figure; give it a calculator tool and it recognizes the arithmetic, hands the numbers off, and reports the exact result the calculator returns rather than its own estimate. The clearest way to keep this straight is a brain-and-hammer distinction. The tool does not make the model smarter; it makes the system more capable, and the two roles stay separate. The model's job is the brain's work: understand what is being asked, recognize which part of the problem it cannot handle reliably on its own, decide which tool fits, phrase the request in a way the tool can use, and weave the tool's answer back into a natural, readable response. The tool's job is the hammer's work: do one narrow thing and do it exactly right every time, whether that is calculating, searching, or checking a calendar. A model that produces a perfect calculation is not a math genius; it is a system that knew it was bad at arithmetic and smart enough to delegate. The power of a reasoning model is demonstrated by delegation, not calculation.

      important icon   The power of a reasoning model is demonstrated by delegation, not calculation.

      The model-tool exchange itself follows the same principle that opened this chapter: a model can only act on what is currently in front of it. It cannot reach out and run a calculator or a search on its own; what it can do is write, in plain terms, what it needs and from which tool, and hand that request to the surrounding system actually connected to the calculator or the search engine. That system is typically called a harness, and it is designed to respond to these requests and returns the result as new text for the model to read — at which point the model does exactly what it has been doing throughout this chapter: it reads what is now in front of it, including the tool's answer, and continues from there, either giving a final response or recognizing that another tool is needed first. Nothing about the model changes between a plain conversation and one backed by tools; what changes is that it has somewhere to send a request and something reading the reply. This is also where a familiar caution returns. A tool call is only as reliable as the model's decision to make it and its reading of what comes back: choosing the wrong tool, mis-stating the request, or misreading a garbled result all produce confident, well-formatted answers that are simply wrong. Tools ground a model in facts and precise computation it cannot generate on its own, but they do not remove the need to check the output — which is exactly the subject of Chapter 4.

      Conclusion

      The through-line of this chapter is that what is termed "reasoning" in artificial intelligence comes from a models its ability to read what it has just written and compare it with stated requirements and provided examples, not from any hidden deliberation. Everything practical follows from that: asking for visible steps helps because the steps become context; asking for self-critique helps because the critique becomes context; and a longer conversation succeeds or fails largely on whether its accumulated context is useful or cluttered. Two cautions carry forward into Chapter 4. Step-by-step work improves how a model handles what it has, and cannot repair what it never had. And confidence, structure, and length are properties of the writing rather than evidence of correctness — which is exactly why the next chapter treats verification as a skill in its own right.


      3: AI Reasoning is shared under a not declared license and was authored, remixed, and/or curated by LibreTexts.

      • Was this article helpful?