An ordinary language model produces its answer more or less immediately, one token after another, with no separate stage for working anything out. A reasoning model such as o3 inserts a step before the answer. It generates an internal chain of working, often a long one, examines its own intermediate conclusions, and only then writes the response you see.
Two things follow from that design, and they explain almost everything about when o3 is worth using. It is markedly better at problems where the answer depends on several dependent steps. And it is slower and more expensive, because that hidden working is real computation that somebody pays for.
The trade is therefore simple to state: you are buying accuracy on multi-step problems with time and money. Whether that is a good deal depends entirely on the problem.
Across strategy, finance and planning tasks, the pattern of where a reasoning model pulls ahead is fairly consistent.
Anything where step three depends on getting step two right. Working through a pricing change and its knock-on effects on margin, capacity and cash. Reconciling figures that should agree and do not. Planning a sequence of work with real constraints on people and time. A standard model tends to produce something that reads correctly and falls apart on inspection; a reasoning model is more likely to notice its own contradiction.
Handing it a plan, a contract summary or a set of assumptions and asking what breaks is where it feels most valuable. It is genuinely good at surfacing the objection you had not considered, and this is low-risk because you are using it to generate questions rather than answers.
Weighing options against multiple criteria with trade-offs between them. Not because the conclusion is authoritative, but because the working is visible enough to argue with.
What it is not is an oracle. It reasons more carefully within what it knows; it does not acquire facts about your business it was never given, and a well-reasoned answer built on a wrong assumption is a more persuasive kind of wrong.
Most business writing does not benefit. Drafting an email, rewriting a service page, summarising a thread, replying to a review: a standard model does these as well, faster and for less. Paying reasoning-model prices for a routine draft is simply waste.
It also cannot rescue a vague question. "How do I grow my business" produces a long, careful, generic answer. The extra reasoning amplifies a good question and does nothing for a bad one.
And it is worth knowing that a more careful reasoning process does not eliminate confident errors. OpenAI's own documentation for this generation noted that on some evaluations the model asserted more incorrect statements than the previous one, which is a useful corrective to the assumption that "thinks harder" means "is right more often". Verification remains your job.
The hidden working is billed and it takes time. Instead of an answer in a couple of seconds you may wait considerably longer, and the token cost per question is a different order from a standard model.
That changes how you should use it. Reasoning models suit deliberate, occasional, high-value questions: the sort you would previously have set aside an hour to think through. They do not suit high-volume automated work, interactive chat, or anything a customer is waiting on. A sensible pattern for a small business is a cheap model for the daily flow and a reasoning model reserved for the handful of decisions each month that actually matter.
The habits that improve output from a standard model can actively hurt here.
Then read the reasoning it exposes and disagree with it. The value is in the argument you can inspect, not in the conclusion you accept.
o3 and the reasoning models that followed it are a real advance on a specific class of problem, and largely irrelevant to the rest. For a local business the practical use is narrow but genuine: a small number of consequential decisions each month where having a careful, argumentative second opinion is worth a few pounds and a few minutes of waiting. For everything else, the cheaper model was already sufficient.
It generates an internal chain of reasoning before answering rather than producing the response directly. That extra working makes it noticeably better on problems with several dependent steps, such as financial analysis or planning under constraints. It also makes it slower and more expensive, so it is poorly suited to routine drafting and everyday chat.
For a handful of consequential decisions each month, plausibly yes. For daily writing, summarising and customer replies, no, because a standard model handles those as well for a fraction of the price and the wait. The sensible arrangement is a cheap model for routine volume and a reasoning model reserved for genuine decisions.
On multi-step logical and quantitative problems, generally yes. It does not eliminate confident errors, and OpenAI's own documentation for this generation noted higher rates of incorrect assertions on some evaluations than the previous one. A carefully reasoned answer built on a wrong assumption is more persuasive, not less wrong, so verify anything consequential.
Give it context and constraints rather than instructions about technique. Telling it to think step by step is redundant and can hurt, since the reasoning already happens. Supply real figures, real limits and what a good answer looks like, then ask what would have to be true for its conclusion to be wrong.
Want this handled for you?
Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.
Get my free audit → or book a 15-min callWe help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.
Get Your Free AI Marketing Audit →