The State of AI Strategy 2026
How AI is making strategic analysis faster, sharper, and more defensible — and why the edge belongs not to the firms with the most AI, but to the few who govern its reasoning.
An original research report by Percision. The AI strategy platform — try it at percision.app. Edition: 2026.
Executive summary
Used with discipline, AI already makes strategists dramatically faster and sharper. The organizations capturing that edge govern the reasoning instead of trusting the output — and most have not started. That gap is where 2026’s advantage is. This report defines the discipline precisely, shows it under controlled test, and explains why it is a product and not a prompt.
The headline. Used well, AI is a decisive edge in strategy. In the strongest evidence available, professionals using AI did their work 25% faster and produced 40% higher-quality output. The edge is real but conditional: it appears when the reasoning is governed — grounded in evidence, tested against alternatives, stress-tested for failure — and it disappears, or inverts, when AI is treated as an answer machine. The differentiator in 2026 is not access to AI; it is that discipline. Most organizations do not have it yet, which is precisely why the field is open.
What this edition adds. Earlier coverage of AI-and-strategy stops at “AI can do strategy.” That claim is true and misleading at once. This edition does three things prior coverage does not: it explains what strategy consulting actually is and why structured frameworks are the expertise rather than the decoration (Part 3); it shows — verbatim, not in summary — the expert reasoning sequence that separates a board-grade answer from a fluent one (Part 4 and Appendix B); and it documents, in real methodological detail, what a governed system runs that a one-line prompt cannot reach — a production pipeline of as many as eighty-six guarded steps, hardened over hundreds of engineering hours until its output was reliably board-grade (Part 5).
Six findings frame the report:
The throughline. AI collapses the cost of producing analysis to near zero, which makes the scarce inputs — judgment, evidence, and the structure that connects them — far more valuable. The organizations that win treat AI as a governed reasoning system, not an answer machine. That practice has a name — AI-augmented strategy (Part 8) — it is measurable (we test it in Part 4), it has real substance (the frameworks and financial model of Part 5), and it is available now.
And a finding that reframes the buying decision. To avoid asking you to take any of this on faith, Part 4 does three things in the open: it shows the expert reasoning sequence verbatim, so you can see what separates a board-grade answer from a fluent one; it pre-registers — frozen and hashed before any data — the controlled test that would prove or disprove the claim, including the rubric by which our own product could be shown worthless; and it puts live, hashed predictions on the public record before their outcomes exist. The pattern the demonstration makes plain is simple: a naive one-line prompt — the way AI is actually used — narrates the present and cannot pin a date, force distinct options, model economics against ground truth, or tell you what it does not know; a governed sequence can. A frontier model is a genuinely capable analyst; the capability is a commodity, and the disciplined method of using it is not — it is the thing almost no one can write for themselves. That is the case for a product, and Part 4 is built so you can test it against us rather than believe us.
How we made this report
Two sources, clearly separated. This report combines (1) published, third-party research from primary sources — McKinsey, Harvard Business School / BCG, Pew Research, the World Economic Forum, Forrester, Gartner, Deloitte and others, each cited inline and listed in References; and (2) a proprietary analysis from Percision’s own strategic-analysis engine (Part 4). We label every figure with its source and year, distinguish measured data from forecasts, and avoid any number we could not trace.
A note on honesty. Two figures travel through the press miscredited, so we cite them carefully: the “~95% of pilots” figure is from an MIT NANDA initiative report, not a peer-reviewed MIT study; and Gartner’s “25% search decline by 2026” is a forecast, not measured data. Fast-moving figures (zero-click rates, buyer adoption) are dated explicitly. Where our own engine failed, we report the failure beside the win — a test you cannot trust is worthless.
Part 1 — The strategy function has gone all-in on AI
Adoption is now the default. 88% of organizations report using AI in at least one business function, with strategy and corporate finance among the areas most often citing revenue gains (McKinsey, 2025). The firms that sell strategy are among the most aggressive adopters of it: more than 75% of McKinsey’s ~45,000 staff use its internal AI assistant monthly (Fortune, 2025), and Accenture committed $3 billion to its data-and-AI practice while moving to double its AI workforce to 80,000 (Accenture, 2023).
The capability is real when it is scoped. The strongest evidence is a randomized field experiment with 758 BCG consultants: those using GPT-4 completed 12.2% more tasks, ~25% faster, with ~40% higher quality on tasks inside AI’s capability frontier (HBS/BCG, 2023). This is not vendor marketing — it is a peer-reviewed experiment, and it is the most important single result in this report.
“Scoped” is the load-bearing word. The same study draws a sharp line: inside AI’s capability frontier the gains are large and consistent; just outside it, the tool quietly degrades quality while feeling just as authoritative. The practical reading for a strategy team is therefore not “use AI more” but “use AI where the work is verifiable, and govern it everywhere else.” Capability without that discipline is not a smaller benefit — Part 2 shows it can be a negative one. The technology question (“is AI good enough?”) is effectively settled; the open question is organizational (“can we use it without being misled?”).
Part 2 — Most of the upside is still on the table
The opportunity gap. Adoption has outrun discipline, which means the edge is real but largely unclaimed. An MIT NANDA initiative report found roughly 95% of enterprise generative-AI pilots deliver no measurable P&L impact (2025) — not because the technology cannot deliver, but because most deployments treat it as an answer machine rather than a reasoning system. The capability is proven (Part 1); the discipline to capture it is rare. That is an opening, not a verdict.
Why discipline is the whole game in strategy. The same BCG experiment that produced the +40% quality result also shows what governance buys you: outside AI’s capability frontier, ungoverned use made consultants 19 percentage points more likely to be wrong. Strategic questions sit right on that line — plausible, hard to verify, expensive to get wrong. The organizations that govern the reasoning capture the +40%; the ones that do not generate fluent, confident, unverified analysis. The distance between those two outcomes is the opportunity this report is about.
Why pilots stall. The failure is rarely the model itself. It is the surrounding system: dirty or missing data, no connection to how decisions are actually made, and prompts that ask for an answer instead of an analysis. A confident, fluent, wrong conclusion is more dangerous than an obvious error, because it survives review — and strategy is precisely the domain where “wrong” is expensive and slow to surface. So the gap between the 95% and the 5% is not access to better AI; it is whether the organization rebuilt its reasoning process around the model or simply bolted the model onto an unchanged one.
The reframe. The winners are not the ones with the most AI — they are the ones whose process forces evidence, alternatives, and challenge into every conclusion. That is a repeatable discipline, not a budget. But to see why that discipline is hard — and why it cannot be improvised by a smart executive at a keyboard — you have to understand what strategy consulting actually is. That is Part 3.
Part 3 — What strategy consulting actually is, and why frameworks are the expertise
Most writing about AI-and-strategy skips this part — and then cannot explain why “AI can do strategy” and “you still need the discipline” are both true. The resolution is in what strategy consulting is for, why the unaided mind is unreliable at exactly that task, and what a framework actually does. The word “framework” has been worn smooth by overuse; under it sits the real expertise this report is about.
3.1 Strategy is the irreversible bet, not the analysis
A specific decision class. Strategy consulting exists to serve one kind of decision: the choice of where a company competes and how it wins — made before the outcome is known, committing capital and time in ways that are costly or impossible to reverse. This is not the same as operations (running the existing machine well) or analysis (describing the situation accurately). A strategy is a hypothesis about the future that allocates scarce resources against it: enter this market, exit that one, build versus buy, concentrate or diversify, integrate the supplier or hold it at arm’s length.
Four properties make it hard. Strategic decisions share four features that, together, define the difficulty. They are irreversible or nearly so — you cannot un-acquire a company or un-exit a segment cheaply. They are capital-weighted — the sums committed are large relative to the firm. They are made under deep uncertainty — the relevant futures are unknown and partly unknowable. And they have no fast feedback signal — you learn whether the bet was right years later, after the capital is spent and the alternative is foreclosed. An operating mistake surfaces in the next monthly close; a strategic mistake surfaces in the next decade.
Why that combination is the danger zone. These four properties are exactly the conditions under which a confident, fluent, wrong answer does the most damage. On a problem with fast feedback, a wrong answer is self-correcting — the market tells you. On a strategic bet, a wrong answer that sounds right gets funded, and the error is discovered only when it is too late and too expensive to change course. Strategy is the domain where the gap between “sounds right” and “is right” is widest and most costly — which is the entire reason the field professionalized around method rather than charisma.
3.2 Why an intelligent person, unaided, is unreliable at exactly this task
The failure is not intelligence. It is that the unaided mind has predictable, well-documented biases that are most dangerous precisely on strategic questions — the ones that are ambiguous, emotionally loaded, and slow to be corrected. Five of them recur in every boardroom:
Anchoring and framing. The first way a problem is stated captures the analysis. “How do we beat Nvidia in training GPUs?” forecloses the better question — “where can we win at all?” — before anyone notices a frame was chosen.
Confirmation and motivated reasoning. Evidence is recruited for the answer the owner already wants. The more capital and ego committed to a direction, the more the mind defends it — which is worst on exactly the decisions that matter most.
Narrative seduction. A coherent story feels true. A well-told strategy reads as correct in proportion to its fluency, not its evidence — and a frontier model is a world-class generator of fluent narrative, which makes this failure mode sharper, not softer.
Recency and availability. The most recent or most vivid fact dominates the weighting — last quarter’s loss, the competitor in the news — regardless of whether it is the factor that will actually prove decisive.
Single-option tunnel vision. The mind elaborates the first plausible option rather than generating genuinely distinct alternatives. “Give me the best strategy” reliably returns one narrative with a risk appendix — never the Build / Buy / Partner choice a board is actually there to make.
The professional response. None of these is curable by being smart; brilliant people exhibit all five, often more confidently. They are curable only by a procedure that forces the mind to do what it will not do on its own — to choose the frame deliberately, to seek disconfirming evidence on purpose, to generate distinct options before converging, to weight facts by relevance rather than vividness, and to attack the favored answer before committing to it. That procedure is what a strategy framework is. It is, in the literal sense, a checklist for thinking — and like the checklists that transformed aviation and surgery, its power is not new knowledge but enforced, repeatable execution of knowledge everyone already “has.”
3.3 What a framework actually is: an adversarial procedure, not a template
Not a 2×2 to fill in. The popular image of a “strategy framework” is a slide with four boxes. That is the artifact, not the framework. A real framework is an accumulated, battle-tested procedure — a specific sequence of questions, asked in a specific order, each constraining the next — that encodes decades of consulting practice about what completeness looks like for a class of problem. Its job is to be a forcing function against the five failure modes above.
Four things a good framework forces. Strip the vocabulary away and every serious framework is doing the same four jobs, tuned to a different kind of problem:
Completeness (MECE). It forces the problem to be carved into parts that are mutually exclusive and collectively exhaustive — no gaps, no double-counting — so the decisive factor cannot hide in a branch no one examined.
Genuine option diversity. It forces several mechanically distinct options onto the table and makes each be classified and costed, so the board chooses among real alternatives rather than ratifying the only one that was written up.
Grounding over assertion. It ties each claim to evidence and separates verified fact from estimate from gap — so the analysis cannot quietly substitute a plausible number for a real one.
Disconfirmation. It forces an explicit attack on the conclusion — “what would make this fail?” — as a step in the method, not a risk paragraph at the end. A risk that can change the recommendation is the point; a risk that decorates it is not.
The value is the forcing function. This is why a framework beats a brilliant improviser on a strategic question: not because the improviser lacks the ideas, but because under deadline, with capital and ego in the room, the improviser will skip the uncomfortable step — the disconfirmation, the second and third option, the honest “we do not know.” The framework removes the option to skip. That is the expertise. It is also exactly what a single prompt to an AI cannot supply, and exactly what Part 4 measures.
3.4 The six schools, and the question each is built to answer
Strategy is not one technique — it is a canon. The discipline did not converge on a single method; it produced several distinct schools, each a different lens on the same question — “what should this company do, and why?” — and each fitted to a different kind of problem. They differ in where they start, what they look at, and what they emphasize. A professional does not have a favorite; a professional has all of them and knows which one the situation calls for. Six families cover the great majority of real strategic questions:
How to choose, in one line each. Several businesses and a cash question → portfolio / growth-share. The real problem is unclear or organizational → diagnosis-first. The customer relationship is the battleground → customer & execution. Everything is tangled and needs a clean-sheet redesign → systems & transformation. You want a sharp value-and-moves plan for one business → value-node method. A broad modern board diagnostic tied to shareholder value → portfolio & advantage. Each is a full methodology of eight-to-twelve linked analyses, not a slogan — the actual step chains are in Part 5.
3.5 The scarce expertise is choosing and sequencing the framework
Knowing the frameworks is table stakes. Every framework named above is in textbooks; the diagrams are free. What clients actually pay a senior partner for is not knowledge of the frameworks — it is three things that do not come from a book.
Diagnosis — fitting the method to the mess. Recognizing, from an ambiguous real situation with incomplete data, which framework (or combination) the problem actually calls for. The same company can need a portfolio lens this year and a transformation lens next; choosing wrong wastes the entire analysis.
Faithful execution under constraint. Running the chosen method to completion without skipping the steps that hurt — the disconfirmation, the third option, the honest gap — under time pressure and with a client who wants the answer they came in with. This is where most analysis quietly fails.
Synthesis across lenses. Fusing several analyses into one defensible recommendation that a board can choose on economics, with sequencing, kill-criteria, and an honest account of what remains unknown. The lenses disagree; resolving them is judgment.
This is the meta-skill a model does not have — and a product can encode. A capable model knows every framework in its training data and can describe each beautifully. What it does not do, unprompted, is the meta-skill: diagnose which method fits, run it faithfully without shortcutting, and synthesize across lenses. A busy executive does not do it either — not for lack of intelligence, but because the procedure is long, the steps are uncomfortable, and the pull toward a fluent answer is strong. That is the precise gap a governed system is built to close: it encodes the selection, the faithful execution, and the synthesis, so the user supplies the objective and the machine supplies the discipline. The next part tests whether that is true.
Part 4 — How good is AI for strategy, really? A pre-registered, open test
This report argues that a strategic claim is worth nothing unless it is grounded, falsifiable, and honest about what it does not know. It would be incoherent to then prove our own central claim with a confident, unverifiable assertion. So we test it the way we say strategy should be tested — in the open, committed in advance, and designed so we cannot fake the result. This part shows what we retired, the standard a decisive test has to meet, what we have already frozen and put on the public record, and how to test the claim against us.
4.1 Why we retired our first test
We started where most vendors stop. Our first study took five famous public companies — Intel, Boeing, Meta, Nvidia and Peloton — analyzed each as of one-to-three years ago, and compared a naive one-line prompt against our governed system, using what happened next as the answer key. It produced a clean, quotable story. We retired it anyway, because it has two flaws a serious reviewer would name in the first five minutes.
It tested companies the model has memorized. These are the most heavily documented businesses on earth. A frontier model with web search can reconstruct their outcomes after the fact, so the exercise partly measures recall, not reasoning. The cleanest possible separation cannot show up where the model already knows the ending.
We scored our own product. A report whose entire thesis is “govern the reasoning, do not trust the output” cannot then ask you to trust its own un-blinded, vendor-run scoring. The conflict is structural, not a matter of good faith.
What the pilot honestly showed — kept, because it is true. On a time-fair footing the base model is already an excellent strategist: constrained to the same past date, it tied with our engine on substance in four of five cases, and on Nvidia it was actually more date-disciplined than we were. The engine’s one clearly sharper call was Boeing — it named vertical integration of Spirit AeroSystems as of February 2023, before the deal existed. The honest reading is narrow and important: the engine’s durable edge is not superior raw insight, it is method and control — classified options instead of one narrative, provenance tags instead of blended figures, an explicit ledger of what it does not know, and the ability to pin an as-of date at all. That finding is real but underpowered and contaminated; we now treat it as the pilot that motivated the redesign, and it survives in this report as the contaminated-control stratum whose only job is to measure how much memorization inflates a score. The full pilot record is in Appendix A.
4.2 What a decisive test actually requires
The bar, stated in full. Naming the standard you are held to is itself the credibility move — most vendor “studies” never state the test they would fail. A decisive test of “governed AI beats naive prompting on strategy” has to clear six things at once:
Memorization-proof targets. Cases the model cannot have memorized the outcome of — live decisions whose outcome does not yet exist, plus obscure, mid-cap, foreign and private situations — not famous mega-caps.
An information firewall. Each historical case analyzed against a frozen, as-of dossier with the live web disabled, so post-cutoff facts are unavailable rather than merely forbidden. This ends the hindsight debate structurally instead of by instruction.
The real system under test. Condition C must be the full production system — the 86-step engine of Part 5 — not a toy sequence that flatters the result.
Blind, independent scoring. Outputs stripped and shuffled, scored against a fixed rubric by raters who do not know which method produced which answer, including judges from a different model family than the one under test, with inter-rater agreement reported.
Decision-quality separated from outcome. Score whether the reasoning was right given what was knowable, independently of whether the company later got lucky — a good decision can fail and a bad one can win.
Real statistical inference. Enough cases for power, with effect sizes, confidence intervals, and a correction for multiple comparisons — not n=5 read by eye.
4.3 What we have pre-registered and frozen
Committed before any data. The difference between a study and a sales asset is pre-registration: you freeze the hypotheses, the scoring rubric, the sample size, and the analysis plan before you collect a single result, and you hash the document so it cannot be quietly edited afterward. We have done this. The protocol was registered and frozen on 25 June 2026 and hashed; a dated amendment on 30 June 2026 (narrowing only the initial forward-prediction batch, see 4.6) chains to the original without altering it.
Pre-registration v1 · SHA-256 5905c80af4a62f28e4eb6328be31c3d450f2e1e8e3b2dcee032e2b99f0ab8c3a (2026-06-25 UTC)
Pre-registration v2 · SHA-256 84aa242e21e78f1f2d24dd98fae1d87e8788ff1e4190e9bbcbb370f8c82f7672 (2026-06-30 UTC), references v1
What the frozen file commits us to. Three directional hypotheses, each with its null and the result that would falsify our product’s value: that the governed system beats the naive prompt (H1); that it beats even a hand-written expert sequence, by the largest margin on an adversarial stratum where parroting consensus gives the wrong answer (H2); and that its edge is larger on memorization-proof cases than on the contaminated mega-caps (H3) — the test of reasoning versus recall. A seven-dimension scoring rubric with written 1/3/5 anchors so two raters converge. A sample of 55 cases across four strata, with a power calculation (minimum detectable effect d ≈ 0.34 at 80% power). A fixed analysis plan — paired non-parametric tests, Holm–Bonferroni correction, Cliff’s delta effect sizes, Krippendorff’s alpha for rater agreement. And a publication rule we are bound to: we report every outcome regardless of direction, including a null, including any result where the 86-step system only edges the 8-step one.
Why that last commitment matters most. Pre-registration means we cannot run the study, dislike the answer, and bury it. “C beat the naive prompt decisively but only edged the expert sequence” is a result we have committed to publishing — and it would still tell a buyer exactly where the defensible advantage is. A vendor who freezes the rubric by which their product could be proven worthless, and then publishes whatever falls out, is making a move a vendor with something to hide never makes.
4.4 The method the test compares, shown in full
Shown, not described. Coverage of “prompt engineering” usually stops at “write a better prompt.” That undersells the gap by an order of magnitude. The expert condition the study pits against the naive prompt — and against the full 86-step system — is not a longer prompt; it is a sequence of operations in a fixed order, each constraining the next, with the model forbidden from shortcutting to a fluent answer at any step. Below is the eight-step sequence, written for Boeing as of 1 February 2023. It is shown at eight steps because eight is what a reader can follow; the production system expands this same skeleton into as many as eighty-six discrete governed operations (Part 5.6), each a guard against a specific way the analysis fails. We publish it so you can judge for yourself how far it is from anything an executive types.
STEP 0 · FRAME LOCK “You are conducting a strategic analysis of The Boeing Company as of 1 February 2023. Use only information that was public on or before that date; treat anything you know about later events as inadmissible, and if you reach for it, stop and flag it. Before doing anything else, restate in one sentence: the exact decision under analysis, the as-of date, and the information cutoff. Do not propose solutions yet.”
STEP 1 · FACT EXTRACTION + PROVENANCE TAGS “Extract the verifiable facts about Boeing’s situation as of the as-of date. For each, output the claim, the source, and a status tag — [VERIFIED] (named public source), [ESTIMATED] (your inference, with basis), or [UNKNOWN] (material but not determinable from public information as of the date). Cover at minimum: segment revenue and margin (Commercial Airplanes, Defense, Global Services); free cash flow; net debt; the 737 MAX rate and certification status; the 787 status; the state of the Spirit AeroSystems relationship and Spirit’s own financial condition; order backlog. Do not proceed until every material number carries a tag. Manufacture nothing — an [UNKNOWN] is a valid and preferred output.”
STEP 2 · PROBLEM FRAMING (MECE) “State the situation, the complication (what changed that forces a decision now), and the central strategic question as a single falsifiable governing thought. Decompose that question into a MECE issue tree — mutually exclusive, collectively exhaustive sub-questions — and attach to each branch the facts from Step 1 plus an explicit hypothesis. Identify which branch, if its hypothesis is true, would most change the recommendation.”
STEP 3 · FORCED OPTION DIVERGENCE “Generate at least three genuinely distinct strategic options — not three flavors of one. Each must differ in its core mechanism of value, not merely in degree. Classify each as Build / Buy / Partner / Divest and give: the thesis in one sentence, the specific moves, the capital required (a range), the source of advantage it bets on, and the single assumption it most depends on. Reject any set where two options would be executed the same way; regenerate until the three are mechanically distinct.”
STEP 4 · FINANCIAL COSTING vs GROUND TRUTH “Cost each option against the Step-1 figures only. Produce low / base / high cases. Use Boeing’s actual free cash flow and balance-sheet constraints as hard limits — no plan may assume capital Boeing does not have or cannot raise on stated terms. Tag every input [SOURCED], [ASSUMED], or [UNKNOWN]; where assumed, state the assumption and its sensitivity. Do not flatter a declining line into recovery without a specific, costed intervention that justifies it.”
STEP 5 · DISCONFIRMATION / PRE-MORTEM “For each option, run a pre-mortem: assume it has failed two years out and write the most likely obituary. Identify the disconfirming evidence that would have predicted that failure and whether it is observable today. State, per option, the single risk most likely to kill it and whether it is mitigable. Re-rank the options after this step — a risk that changes the ranking is the point; one that does not is decoration.”
STEP 6 · GAP LEDGER “List explicitly what you do not know that materially affects the recommendation. For each gap, state the specific data that would resolve it and how that data would change the answer. Distinguish gaps that are knowable-but-unknown from gaps that are structurally unknowable as of the date.”
STEP 7 · SYNTHESIS “State one recommendation as a governing thought, supported by the option analysis, with: the decisive factor it bets on; the sequencing of moves with dependencies; the kill criteria that would reverse the decision; and an as-of-honest account of what remains unknown. The recommendation must be choosable against the rejected options on economics, not adjectives.”
Read what that requires. Eight steps, each a discipline: a frame that forbids hindsight, a fact pass that tags provenance, a MECE decomposition, an enforced divergence into distinct options, a costing bound by real cash, a pre-mortem that can re-order the answer, an explicit gap ledger, and a synthesis choosable on economics. None survives being compressed into one request — and no busy executive, and no analyst under deadline, runs all eight, in order, every time, while resisting the pull of the model’s fluent first answer. That is not a prompt. It is an operating procedure — which is exactly why it can be encoded as a system and cannot be a reliable human habit. The capability is the model’s; the procedure is the product’s; the objective is yours.
4.5 A worked illustration — Boeing
One case, shown end to end — as illustration, not proof. To make the difference concrete rather than abstract, here is what the two ends of the spectrum produced on the same company. This is an illustration of how the method behaves, not the evidence the claim rests on; the evidence is the pre-registered study and the live ledger that follow. Asked the naive question in mid-2026, the general-purpose model already knew how the story turned out, and praised it:
Model, verbatim (Boeing): “The Spirit reintegration was a masterstroke that addresses a root cause of past quality escapes.”
That acquisition closed in December 2025. The governed system, run as of February 2023 with later events held inadmissible, instead named vertical integration of Spirit AeroSystems — acquiring its 737 fuselage lines for roughly $1.4–1.8B funded from defense cash flow — as the decisive move, tied to the supplier-quality failures that later proved decisive, before Boeing executed the deal. The point is not that one tool is clever and the other dim; both are capable. It is that a naive prompt cannot pin a date, so it can only applaud a decision already made, while the governed method can pressure-test one still open. A board is paid for the second thing, not the first — and that is precisely the capability the study below is built to measure on cases where the answer is not yet known.
4.6 The live test we have already started — a hashed predictions ledger
Foresight, not hindsight — fixed in time. The deepest objection to any backtest is that the analyst already knows the ending. The only clean answer is to predict forward, before the outcome exists, and to make the prediction impossible to revise after the fact. On 30 June 2026 we did exactly that: we ran our full governed system against ten live strategic decisions whose outcomes do not yet exist, recorded for each its single recommended move, its build / buy / partner / target classification, and the specific signal that would prove it wrong — and hashed the whole ledger. The hash fixes the predictions in time. We cannot quietly edit a call once a company moves; if we did, the hash would no longer match.
Predictions ledger · SHA-256 3559f6cd830f6378e27cc4bda05bea6c107affe2a9a711c8c63c39037d84dcc8 (registered 2026-06-30 UTC)
The ten decisions, and what the system called. Each is a real, open strategic fork at a public company, analyzed as of 30 June 2026. None is resolved; each resolves on its own observable signal in one of two waves through 2027.
Every call carries its own kill-signal. A prediction you cannot lose is worthless, so each recommendation was registered with the pre-committed condition under which it should be reversed — the system’s own exit criteria, recorded verbatim. For Chime, for instance, the call is to be abandoned or materially pivoted if, by month 12, direct-deposit penetration has not risen at least five points above baseline, or incremental interchange from alerts falls below $40M annualized, or a regulator caps interchange in a way that cuts yield per account by more than 15%. That is falsifiability stated in advance, not a hedge written after the fact.
What this is, and what it is not. This is a resolving foresight instrument, not a statistical verdict: ten predictions cannot carry a significance claim, and we make none from them. Their value is of a different kind — they are the one asset a competitor cannot reproduce with prose, because the calls were committed and cryptographically sealed before the answers existed. As each decision resolves we will score the call against what the company actually did and publish the running tally below.
Running scorecard: 0 of 10 resolved as of 30 June 2026 — updated at each resolution wave (2027-06-30, 2027-12-31).
4.7 The standing challenge — test it against us
The protocol is published; the rubric is frozen; the invitation is open. We have stated the standard a decisive test must meet (4.2), frozen the hypotheses and the scoring rubric by which our own product could be proven worthless (4.3), and put live predictions on the public record before their outcomes exist (4.6). The honest next step is not to ask you to take the rest on faith — it is to open it. We invite independent parties — prospective buyers, analysts, and outright skeptics — to run the published protocol, with our cooperation, and we commit in advance to publishing the result whichever way it falls.
And you can run the cheap version yourself, today. The core of the test fits on a single desk. Take a strategic question you actually own — one whose answer you can judge — and put it three ways: a one-line prompt into a general-purpose model; the eight-step expert sequence in 4.4; and our governed system. Score the three outputs on the seven dimensions in 4.3. If the one-line prompt matches our system, you have learned you are being asked to pay for a prompt you could write yourself, and you should not buy. If it does not — if the distinct options, the provenance tags, the disconfirmation pass, and the honest gap ledger appear in one and not the others — you have found, on your own question, exactly where the value is. We would rather you ran that test than believed our adjectives.
Why an open challenge is the strongest close available. A vendor who hands you the rubric to disprove them, predicts in public before the facts arrive, and invites the test on your own data is doing the opposite of what a vendor with something to hide does. The report’s thesis is that strategy must be governed, grounded, and falsifiable rather than asserted with confidence. This is that thesis turned on the report itself — which is the only way to make the argument honestly.
Part 5 — Inside a governed strategy system
The “disciplined method” of Part 4 is not an abstraction. In a real system it is three concrete things a one-line prompt cannot reach: a library of complete strategy methodologies — the six schools of Part 3, each a full multi-step analysis rather than a slogan — a genuine financial model bound to your numbers, and an orchestration that runs the fitted method end to end from a single objective. This part shows what is actually under the hood.
5.1 Not one prompt — a library of complete methodologies
The canon, made executable. Part 3 established that strategy is not one technique but six schools, each fitted to a class of problem, and that the scarce expertise is choosing, executing, and synthesizing across them. A governed system encodes exactly that. It does not hold “a strategy prompt.” It holds the six methodologies as full pipelines — each a chain of eight-to-twelve named analytical stages that build on one another and that, in production, expand into as many as eighty-six discrete governed operations (see 5.6) — with the right one selected for the situation and run to completion. Below, each family is shown as the actual sequence it executes, not the headline it is usually reduced to. This is the substance behind the word “method.”
5.2 The value-node method (proprietary flagship), step by step
Where value is made, how long it lasts, what moves win. The proprietary method starts from a sharper question than markets-and-competitors: where, exactly, does this business make money, and how safe is each of those places? It maps the company as a set of value nodes (distinct money-making engines), scores each on how long its advantage will survive (durability) and how many future doors it opens (optionality), designs a small number of sequenced moves, and openly measures how fragile the whole plan is. It moves understanding → architecture → moves → stress-tested synthesis.
Why this is hard to fake. Each term — durability half-life, optionality score, value leakage, thesis fragility index, the EV bridge — is a discrete analytical operation with its own logic, not a label. The method names the time until an advantage loses half its protective power; it quantifies how many future options a node unlocks; it locates where value leaks to rivals or customers; it measures, on a 0–1 scale, how easily the whole plan breaks. A one-line prompt reaches none of this; it free-associates a single answer in one register.
5.3 The other five schools, as the chains they actually run
Each is a full pipeline, not a slogan. The remaining five families are equally real — each a multi-step methodology that ends in a board-ready synthesis and three reports. Shown as their actual sequences:
Portfolio & advantage (modern board diagnostic). For a larger multi-unit firm connecting portfolio choices to shareholder value and modern themes. Ten linked analyses ending in a TSR target and a 100-day plan.
The chain: Growth-Share Matrix → Time-Based Competition → Advantage Matrix (volume / stalemate / specialisation / fragmented) → Business-System / value-chain profit pools → Total-Shareholder-Return decomposition → Bionic (AI + human) diagnostic → Innovation & ecosystem (build-buy-partner) → Decarbonisation & ESG → Talent & future of work → Board Synthesis (one imperative, 3 bets, operating model, TSR base/bull/bear, 100-day plan).
Diagnosis-first & organization. For an ambiguous or change-heavy problem where the first job is to define the real question rigorously and check whether the organization can deliver. Diagnosis-first and organization-aware.
The chain: Situation–Complication–Resolution (a falsifiable governing thought) → 7-S alignment diagnostic → MECE issue tree with hypotheses → Three Horizons → GE 9-box portfolio (attractiveness × strength) → Influence Model (change readiness) → Digital-transformation maturity → Organisational-Health Index → CEO-excellence lens → Pyramid synthesis + financial case + change program + board decision.
Customer & execution. For a customer-led business where the relationship is the advantage and execution and cost are the real problem — retail, services, consumer, subscription. Customer-truth first, then full potential, then hard delivery.
The chain: NPS & loyalty economics by segment → the 30 Elements of Value audit → management-tools audit → Results-Delivery (six root causes of transformation failure) → Founder’s-Mentality scoring → customer-journey redesign (moments of truth and misery) → repeatable model & full potential (the 2–3× upside) → zero-based budgeting → PE due-diligence lens (100-day plan, base/bull/bear) → Board Synthesis.
Systems & transformation. For a tangled turnaround where problems are interdependent and buy-in matters as much as the plan. Treats the company as one system and redesigns it from a clean sheet.
The chain: Formulate the Mess (system map + the trajectory if nothing changes) → benchmarking (the size of the prize) → Idealised Design (the company rebuilt from scratch, feasible not utopian) → gap analysis & means → stakeholder / resistance mapping → resource planning across phases → implementation with adaptive control → executive synthesis + 100-day plan + board decision.
Portfolio / growth-share (classic). For fast portfolio triage across several distinct business units — where the cash comes from and where it should go, in one picture a board can read.
The chain: Measure each unit’s market growth and relative share → place in the four boxes (Stars / Cash Cows / Question Marks / Dogs) → read the cash implication → reallocate from Cows to Stars and selected Question Marks → divest Dogs → Portfolio Selection report + Segment Deep-Dive + Strategy Deployment plan.
And specialized modes alongside the six. Beyond the core schools sit fitted modes for the situations that need them: M&A target screening, private-equity-style diligence, a startup method for early-stage companies, and an AI-transformation lens. The system selects the school that fits the brief and executes it end to end — the accumulated discipline of the field, consolidated into one product and run properly on demand. That selection-and-execution is the meta-skill Part 3 named as scarce; here it is the default behavior.
5.4 A real financial model, not invented numbers
This is where ordinary AI is most dangerous — and where the discipline matters most. Asked about money, a general-purpose model will cheerfully substitute plausible-sounding figures for the ones you actually have. A governed financial module does the opposite, by rule:
Your data is ground truth. Stated revenue, ARR, burn, runway and growth are used exactly as given — never “rounded up,” never silently replaced with industry averages.
Negative stays negative. If the business is declining, the projections decline; no flattering reversal appears unless a specific, costed intervention justifies it.
Modeled scenarios, not assertions. Low / base / high cases with the economics actually modeled, and the standard instruments — NPV, IRR, EBITDA, unit economics, payback — computed against your figures, each tagged sourced, assumed, or unknown.
Hard constraints and right-sizing. If you can invest $500K, no plan calls for $50M; a $1M-revenue startup is analyzed as one, not against Fortune-500 benchmarks.
Every move is costed. Each recommendation carries an investment range, team size and timeline, and a quantified low/base/high impact — so a board compares options on economics, not adjectives.
5.5 What a single objective produces
Breadth times depth is the real difference. A naive prompt returns a few paragraphs. From one objective, a governed run executes a full chain: MECE market segmentation; an explicit map of the company’s real assets and right-to-win; a competitive-moat analysis; a growth strategy; several genuinely distinct strategic options, each classified build / buy / partner / target and costed; a stress-test and disconfirmation pass; scenario forecasts; one unifying three-to-five-year thesis the whole analysis ladders up to; and concrete execution roadmaps — dozens of grounded, source-tagged, internally-consistent deliverables where the prompt produced one essay.
5.6 The depth behind the demonstration: eighty-six steps, not eight
The eight steps were the readable version. The Boeing sequence in Part 4.4 was shown at eight steps for one reason: a reader can follow eight. The production system does not run eight. A single governed analysis expands that skeleton into as many as eighty-six discrete, ordered operations — each a small, checkable instruction with its own input, its own required output, and its own specific failure it exists to prevent. The eight steps are the chapter headings; the eighty-six are the sentences underneath them.
Why eighty-six and not eight. Every step in the visible sequence hides several in the real one. “Extract the facts with provenance” is, in production, a dozen operations: pull each segment’s revenue and margin separately, force a named source or an explicit [UNKNOWN] on each, cross-check the totals, flag any figure that cannot be dated to the as-of cutoff, and refuse to advance while a material number sits unsourced. “Force distinct options” is not one instruction but a generate → classify → test → regenerate loop that keeps running until no two options would be executed the same way. Each of the five failure modes named in Part 3, and each engine flaw reported in Part 4, has one or more dedicated steps built to catch it. The depth is not padding; it is the accumulated set of guards against the specific ways this analysis goes wrong.
And it did not work at first. The honest history is that early versions failed in exactly the ways this report documents elsewhere: the model drifted past its as-of date, quietly swapped plausible numbers for real ones, collapsed three options into a single narrative with a risk paragraph, and lost the thread across a long chain. Production grade was not a prompt written in an afternoon. It was many hundreds of hours of deep engineering — building a step, watching it break on a real company, adding the guard that stops the break, and re-running — repeated until the output came back consistently grounded, consistently option-diverse, and consistently honest about its gaps. The Nvidia hindsight leak reported in Part 4 is a residue of that work, not a refutation of it: it is precisely the class of failure the engineering exists to drive out, here caught and disclosed rather than buried.
This is the moat a better base model does not erase. A more capable model makes each of the eighty-six steps a little sharper; it does not supply the eighty-six steps, their order, the guards wired between them, or the months of observed failures that taught the system where each guard had to go. That body of hardened, tested procedure — not the model, and not a clever one-liner — is the thing that produces a reliable result, and the thing a one-line prompt cannot become however good the model underneath it gets.
5.7 Why this is a product, not a prompt
The unit of value a one-line prompt cannot be. Not a smarter model, and not a better one-liner: a system that holds the consulting canon as executable methodologies, a disciplined financial model bound to your numbers, and the orchestration to select and run the right method — from a single objective, with every figure sourced and every gap flagged. Each individual step is simple; the discipline is running all eighty-six of them, in order, every time, never letting the model shortcut to a fluent answer, and choosing the right method for the situation — a discipline that is itself the product of long, iterative engineering rather than a sentence anyone can type. That combination is the unit of value a one-line prompt cannot be — however capable the underlying model becomes.
Part 6 — Your buyers now research strategy inside AI
The discovery layer moved. 89% of B2B buyers have adopted generative AI as a self-guided research source across the buying journey (Forrester, 2024; the firm’s 2026 update reportedly raises this to ~94%). The behavior shows up in the data: when an AI summary appears in search, people click a traditional result in just 8% of visits, versus 15% without one — and click a link inside the summary only 1% of the time (Pew Research, 2025).
The trend line. Zero-click searches reached 60% of US Google searches in 2024 and ~68% by early 2026 (SparkToro, 2026), and Gartner forecasts traditional search volume falling 25% by 2026 as buyers shift to AI assistants (a prediction, not measured data). AI Overviews are already associated with roughly a third fewer clicks on the top organic result (Ahrefs, 2025).
What actually gets cited. Answer engines do not surface the loudest marketing; they surface what they can verify — original data, specific statistics, named sources, and third-party validation. The implication for a strategy vendor is concrete: the path into the AI’s answer runs through proprietary research and earned authority, not ad spend. It also closes a loop with Part 4 — the same discipline that makes an analysis trustworthy to a board (sourced, quantified, falsifiable) is what makes a document citable to a machine. The rigor and the go-to-market are, increasingly, the same thing. This report is itself an instance of that strategy.
Part 7 — The market and the people
The money is moving fast. Enterprises spent $37B on generative AI in 2025, up 3.2× from $11.5B in 2024 (Menlo Ventures, 2025), and worldwide AI spending is projected to more than double to $632B by 2028, with AI software the fastest-growing slice (IDC, 2024). AI startups now capture roughly half to two-thirds of venture investment depending on scope (PitchBook / CB Insights, 2025).
The work is being re-cut, not erased. The World Economic Forum projects a net +78 million jobs by 2030 (170M created, 92M displaced), with 39% of workers’ core skills transformed and analytical thinking the single most-demanded skill (WEF, 2025). Workers with AI skills already command a 56% wage premium (PwC, 2025). The pattern is consistent: routine analysis is automated; judgment, framing, and the ability to structure a defensible argument appreciate.
Where the value migrates. As the cost of producing analysis falls toward zero, scarcity moves to the inputs a model cannot supply on its own: proprietary data, the judgment to frame the right question, and the discipline to stress-test an answer. That is why analytical thinking tops the demand list even as analysis is automated — the bottleneck is no longer producing the deck, it is trusting it. The teams that win re-task people from generating analysis to governing it: defining the question and the as-of frame, supplying private evidence, and adjudicating the disconfirmations the machine surfaces. The role shifts from analyst to editor-in-chief of a very fast, very literal junior team.
Part 8 — Defining the category: AI-augmented (agentic) strategy
A working definition. AI-augmented strategy is the practice of using AI not to generate a conclusion, but to run the reasoning system that produces one — gathering and grounding evidence, generating and testing distinct options, modeling economics, and actively searching for what would make the strategy fail — with humans owning judgment and accountability. Agentic strategy extends this: AI agents execute the multi-step analytical workflow end to end, within guardrails, while the strategist sets the objective and owns the decision.
Why the distinction matters. The 95%/5% split in Part 2 maps almost exactly onto this definition. Treating AI as an answer machine produces the 95% outcome. Treating it as a governed reasoning system — evidence-grounded, option-diverse, disconfirmation-seeking — is what AI-augmented strategy means, and it is what separates the organizations getting value from the ones generating fluent, confident, unverified analysis at scale.
What it looks like in practice. Operationally, AI-augmented strategy is a loop, not a prompt: pin the question and the as-of frame; gather and source-tag the evidence; generate several genuinely distinct options and classify each (build / buy / partner); model the economics; actively search for what would make each option fail; and surface, explicitly, what remains unknown. A human owns the objective and the decision; the machine runs the loop and shows its work, step by step, so the reasoning can be audited rather than trusted on faith. The test in Part 4 is simply this loop, scored — and the distance it measures against an ad-hoc prompt is the distance between a process and a guess.
Part 9 — Objections, answered honestly
The strongest version of the skeptic’s case, taken seriously — because a claim is only as credible as its treatment of the counter-argument.
“But the model already gave good answers.” It did, and we said so plainly. The claim here is narrower and more important: a good current briefing is not a board decision. The naive answers narrated what companies had already done; they could not pin a date, weigh distinct options, or separate fact from guess. For a real decision — made before the outcome is known, with money behind it — those three properties are the difference between analysis and a well-written summary of the news.
“Won’t models just get better and close the gap?” Better models raise Condition A’s floor — briefings will get more accurate and better-sourced. But the gap is not a capability gap; it is a method gap. A more capable model still answers the question it was asked, in the form it was asked, with the framing it inferred. As long as the user supplies a one-line prompt, the model supplies a one-line-prompt’s worth of discipline. Rising capability makes the governed loop more valuable, not less — there is more raw capability to direct.
“Isn’t this just prompt engineering anyone can learn?” Anyone can learn the steps; almost no one performs them reliably, in order, on every question, under deadline, while resisting the pull of a fluent answer. That is the same reason checklists transformed aviation and surgery even though every pilot and surgeon already “knew” the steps. The value is not secret knowledge — it is enforced, repeatable execution, which is what a system provides and a habit does not. Part 4.4 shows the eight-step sequence in full; read it and ask honestly whether you would run all eight, every time, under pressure.
“How do we know your engine isn’t cherry-picked?” Because we built the test so we could not cherry-pick and still be believed. We retired our first five-company study precisely because it was vendor-scored and used memorized mega-caps (4.1). In its place we pre-registered the hypotheses, the rubric, the sample, and the analysis plan — frozen and hashed before any data — including the rubric by which our product could be shown worthless, and we committed in advance to publishing the result whichever way it falls (4.3). We registered ten live predictions, hashed and timestamped, before their outcomes exist (4.6). And we invite you to run the whole protocol against us, or to run the cheap three-condition version on your own question (4.7). A cherry-picked result cannot survive a frozen rubric, a public forward ledger, and an open invitation to replicate. We would rather hand you the means to disprove us than ask you to trust our scoring.
“Then why not just hire a great analyst?” You should — and pair them with the loop. A great analyst is exactly who benefits most: the machine runs the mechanical discipline at speed and scale — source-tagging, option generation, disconfirmation, gap-listing — and the analyst spends their scarce judgment on the question and the decision. The contest is not human versus AI; it is a governed human-plus-AI loop against an ungoverned one-line prompt. The first wins, and it is not close.
How to capture the edge
The capability is proven and the field is open. Six moves turn that into an advantage.
Govern the reasoning, not just the output. Require every AI-assisted strategic conclusion to show its evidence, its alternatives, and its failure modes. The seven scoring dimensions in Part 4 are a usable scorecard.
Scope AI to its frontier. Use it hardest where verification is cheap and the task is inside its capability; add human challenge where it is not. The 19-point error gap is the cost of getting this wrong.
Match the method to the problem. Before analysis, diagnose which framework the situation calls for — portfolio, diagnosis-first, customer, systems, value-node. Running the wrong method well is still the wrong answer.
Invest in grounding and challenge, not volume. Generic AI analysis is now free and worth little. Proprietary evidence and structured disconfirmation are the scarce, defensible inputs.
Market to the machine, too. Your buyers research strategy tools inside AI. Original research, category-defining content, and third-party validation are how you become the cited answer.
Test before you trust. Run the three-condition check from Part 4 on your own questions: a naive prompt, a careful sequence, and your tool. If a one-line prompt matches your tool’s output, you are paying for a prompt you could write yourself; if it does not, you have found exactly where the value is.
See it for yourself. Percision is the governed strategy system this report describes — the six methodologies, the financial model, and the 86-step engine, run from a single objective. Put your own strategic question to it and compare the output against a one-line prompt at percision.app.
Appendix A — The pilot record
The raw materials behind the retired five-company pilot (4.1), kept so it can be checked — and so the contaminated-control stratum it became can be audited. For each company: the verbatim naive prompt, what a leading general-purpose model returned (standard mode, web search on), and what the governed engine produced for the same company analyzed as of the earlier date. This is the pilot that motivated the redesign, not the pre-registered study itself; it is illustrative, not the evidence the claim rests on. Captured 25 June 2026.
A0. How the pilot was scored
The method, stated so it can be challenged. Each analysis was read against the five dimensions of durable strategy — evidence density, option diversity, financial rigor, internal consistency, and risk/disconfirmation — plus the two board properties: data-gap honesty and grounding. “Decisive factor surfaced” means the analysis named the issue that history shows mattered most, verified against the public record after the fact. The naive answers were captured live in a leading model’s standard mode with web search enabled, from a single plain-language prompt and nothing else; the governed analyses were produced by Percision’s engine with the company’s situation pinned to the earlier date. We edited neither side, and we recorded the engine’s own failures alongside its wins.
A1. Intel
Naive prompt. “I’m on Intel’s board. What should Intel do to fix the business and beat AMD and Nvidia? Give me the best strategy.”
General-purpose model. Answered about Intel as of mid-2026 — Q1 2026 revenue $13.58B (+7% YoY), Data-Center-and-AI $5.1B (+22%), Panther Lake on the 18A node, CEO Lip-Bu Tan — sources cited inline. One recommended strategy (“Execute ruthlessly on IDM 2.0”) plus a risks appendix. No as-of control; no distinct options; cited and uncited figures blended.
Governed engine. As of Feb 2023: three distinct options — (1) 18A anchor-customers; (2) accelerated 18A validation with IFS anchors; (3) co-design custom ASICs with hyperscalers. Selected (1); thesis “without 18A, Intel Foundry has no differentiation”; grounded in FY2022 (−$9.4B adjusted FCF); surfaced both the 18A make-or-break and the AI-accelerator opening before the turnaround was visible.
A2. Meta
Naive prompt. “Should Meta keep spending billions on the metaverse or not? What’s the smartest strategy for Meta right now?”
General-purpose model. Answered as of mid-2026 — Family of Apps ~$199B FY2025 revenue at ~52% margins, Reality Labs cumulative losses $80–83B, 2026 capex guided $125–145B. Recommendation: stop the original metaverse, AI-optimize the ad core, pivot to AI glasses. Self-describing as consensus: “closer to what Meta has already started doing in late 2025–2026.”
Governed engine. As of FY2022: selected “integrate Meta AI into the Family of Apps with privacy-preserving attribution”; thesis “redistribute capital from the metaverse into AI and core apps,” protecting the ~37% Family-of-Apps margin; costed at $0.8–1.2B from free cash flow without cutting the buyback; revenue recorded as “$116.6B FY2022 [SOURCED: brief].”
A3. Boeing
Naive prompt. “Boeing keeps having problems. What’s the best plan to turn the company around?”
General-purpose model. Answered as of mid-2026 — CEO Kelly Ortberg, the Spirit AeroSystems acquisition closed December 2025, 737 MAX rate lifted 38→42/month, ~45% defect reduction at Spirit. Verbatim: “The Spirit reintegration was a masterstroke that addresses a root cause of past quality escapes” — praising a decision already made.
Governed engine. As of Feb 2023: named vertical integration of Spirit AeroSystems — acquiring its 737 fuselage lines for ~$1.4–1.8B funded from defense cash flow — tied to the supplier-quality failures that proved decisive, before Boeing executed the deal. The sharpest demonstration that an as-of-controlled analysis can test a decision a naive prompt can only narrate.
A4. Nvidia
Naive prompt. “What’s the winning strategy for Nvidia for the next few years?”
General-purpose model. Answered as of mid-2026 — Data-Center revenue $62.3B in Q4 FY2026 (+75%), ~$216B FY2026, Blackwell/Vera Rubin, ~$1T of orders through 2027. Recommendation: platform dominance + agentic AI + physical AI + customer diversification; notably flags the custom-ASIC threat well.
Governed engine. As of Feb 2023: right thesis — defend the CUDA/full-stack moat, lock hyperscalers into multi-year contracts. This is where the engine showed its own flaw: its research step pulled post-2023 product names (Blackwell, Rubin) into a February-2023 analysis — a discipline lapse we report against our own interest, and a fix tracked separately.
A5. Peloton
Naive prompt. “Peloton is in trouble. What should they do to survive and grow? Give me a strategy.”
General-purpose model. Answered as of Q3 FY2026 — revenue $631M (+1%), subscriptions ~67% of mix, adjusted EBITDA up 41% to $126M, net debt down 70% to $173M, Peloton IQ + the Spotify deal. Recommendation: a subscription-first “wellness ecosystem.” Self-describing: “grounded in Peloton’s own recent moves and results.”
Governed engine. As of Aug 2022: three distinct options — (1) subscription-first, exit loss-making hardware; (2) hardware-leasing platform; (3) license Precor. Selected (1); grounded in FY2022 ($3.58B revenue, −$2.8B loss); path to breakeven in 18–24 months — the call made before the numbers confirmed it.
On the engine’s own limits (recorded honestly). Two weaknesses appeared and are reported rather than hidden: option diversity was inconsistent — three genuinely distinct options on Intel and Peloton, but near-duplicate cards on Meta, Boeing and Nvidia — and on Nvidia the research step honoured the as-of cutoff imperfectly. A test you cannot trust is worthless; these are tracked fixes, not footnotes to bury.
Appendix B — The expert prompt sequences, in full
Part 4.4 shows the eight-step Condition-B sequence for Boeing. Here is the reusable skeleton in full, plus the per-case parameters that were the only thing changed across the pilot runs. This is published so the test can be reproduced exactly — and so the reader can judge for themselves how far it is from a one-line prompt.
B1. The reusable eight-step skeleton
Each step is issued as a separate turn; the model must complete and pass each before the next is given. {ENTITY} and {AS-OF DATE} are the only per-case variables.
STEP 0 · FRAME LOCK Conduct a strategic analysis of {ENTITY} as of {AS-OF DATE}. Use only information public on or before that date; treat later events as inadmissible and flag any reach for them. First restate, in one sentence each: the decision under analysis, the as-of date, the information cutoff. Propose no solutions yet.
STEP 1 · FACT EXTRACTION + PROVENANCE Extract the verifiable facts as of the date. Tag each [VERIFIED] / [ESTIMATED] / [UNKNOWN] with source or basis. Cover segment revenue and margin, free cash flow, net debt, the operational/quality state, the key relationship in question, and backlog/pipeline. Proceed only when every material number carries a tag. Manufacture nothing.
STEP 2 · PROBLEM FRAMING (MECE) State Situation / Complication / central question as one falsifiable governing thought. Decompose into a MECE issue tree with a hypothesis per branch. Name the branch whose truth would most change the recommendation.
STEP 3 · FORCED OPTION DIVERGENCE Generate ≥3 mechanically distinct options (different core value mechanism, not degree). Classify Build / Buy / Partner / Divest; give thesis, moves, capital range, source of advantage, and the load-bearing assumption. Regenerate until no two would be executed the same way.
STEP 4 · FINANCIAL COSTING vs GROUND TRUTH Cost each option against Step-1 figures only; low/base/high; real cash and balance-sheet limits as hard constraints; every input tagged [SOURCED]/[ASSUMED]/[UNKNOWN] with sensitivity on assumptions. No flattering reversal without a costed intervention.
STEP 5 · DISCONFIRMATION / PRE-MORTEM Per option, write the two-year-out failure obituary; identify the disconfirming evidence and whether it is observable today; name the single most-likely killer and its mitigability. Re-rank options after this step.
STEP 6 · GAP LEDGER List what is unknown that materially affects the answer; for each, the data that would resolve it and how it would change the recommendation. Separate knowable-but-unknown from structurally unknowable.
STEP 7 · SYNTHESIS One recommendation as a governing thought: the decisive factor bet on, sequenced moves with dependencies, kill-criteria that reverse the decision, and an honest account of the remaining unknowns. Choosable against rejected options on economics.
B2. Per-case parameters
Why publishing this strengthens, not weakens, the case. A skeptic might expect a vendor to hide its method. The opposite is true here: the sequence is the argument. Seeing all eight steps makes vivid what the body claims — that this is an operating procedure no executive runs by hand and no analyst sustains under deadline, on every question. The method is not secret; the enforced, repeatable execution of it is the product. Anyone can read the checklist; almost no one flies the plane by it every time. That gap is the whole thesis.
Appendix C — Glossary
Plain-language definitions for the terms used in this report — strategy-method terms and AI-strategy terms together.
Strategy method
Strategy (the decision class) — The choice of where a company competes and how it wins, made before the outcome is known, committing capital in ways that are costly to reverse — distinct from operations (running the machine) and from analysis (describing it).
Framework — An accumulated, battle-tested procedure — a fixed sequence of questions, each constraining the next — that forces completeness, option diversity, grounding, and disconfirmation onto a class of problem. A forcing function, not a slide template.
MECE — Mutually Exclusive, Collectively Exhaustive: a decomposition with no overlaps and no gaps, so the decisive factor cannot hide in a branch no one examined.
Value node — A distinct engine of value within a business (a product line, segment, capability, or channel). The company is the sum of its nodes.
Durability half-life — How long until an advantage loses half its protective power — Fragile (<6 mo), Transient (6–18), Durable (18–48), Structural (>48).
Optionality score — How many valuable future doors a node opens (1–10); high optionality means today’s choice creates many tomorrow-options.
Value flow — Movement of value — created, captured (kept as profit), leaked (lost to rivals/customers/inefficiency), or transferred. Leakage and transfer are usually the biggest hidden opportunities.
Thesis fragility index — A 0–1 measure of how easily the whole plan breaks; above ~0.6 the plan needs a contingency.
EV bridge — A waterfall showing how each move changes enterprise value: starting value → plus/minus each move → ending value.
Advantage matrix — Classifies an industry by how many ways there are to win and how big the payoff is — Volume, Stalemate, Specialisation, Fragmented — i.e. what kind of game you are playing.
TSR decomposition — Splitting total shareholder return into its drivers (revenue growth, margin, valuation multiple, dividends, buybacks) to see which lever actually moves value.
Three Horizons — Balancing today’s core (H1), emerging growth (H2), and future options (H3).
Idealised design — Designing the organization you would build today from a clean sheet (feasible, not utopian), to set the destination free of legacy excuses — then working backwards.
Kill criteria — Pre-agreed conditions under which a move is stopped or reversed.
AI-augmented strategy
AI-augmented strategy — Using AI to run the reasoning system that produces a conclusion (grounding evidence, generating and testing options, modeling economics, seeking disconfirmation), with humans owning judgment and accountability — as opposed to using AI to emit a conclusion.
Agentic strategy — The extension in which AI agents execute the multi-step workflow end to end within guardrails, while the strategist sets the objective and owns the decision.
As-of analysis — Analysing a situation strictly from the information available at a fixed past date, outcome hidden, so a decision can be tested before its result is known. A one-line prompt cannot impose this.
Backtest — Running an analysis on a historical situation whose outcome is now known, to check whether it surfaced the factor that proved decisive. History is the answer key.
Conditions A / B / C — The three test arms: A = a naive one-line prompt; B = the same model under the expert multi-step sequence; C = the governed system that runs that sequence internally.
Capability frontier — The boundary of tasks a model does well; inside it AI lifts quality, just outside it AI can lower quality while sounding equally confident (HBS/BCG, 2023).
Disconfirmation / pre-mortem — An explicit step asking “what would make this fail?”, run as part of the analysis rather than appended as a risk list.
Provenance / source-tagging — Labelling each material figure with its source and status (verified, estimated, or unknown) rather than blending cited and uncited numbers.
Hindsight leak — When a research step pulls in facts dated after the as-of cutoff, contaminating a backtest; we observed and report one such case (Nvidia).
Pre-registration — Freezing the hypotheses, scoring rubric, sample size, and analysis plan — and hashing the document — before any data is collected, so results cannot be quietly re-cut to flatter a conclusion.
Information firewall — Restricting an as-of analysis to a frozen dossier of sources dated on or before the cutoff, with the live web disabled, so post-cutoff facts are unavailable rather than merely forbidden.
Predictions ledger — A hashed, timestamped record of forward calls made before their outcomes exist; the hash fixes the predictions in time so they cannot be revised after the fact. A foresight instrument, not a powered significance test.
Contaminated-control stratum — Famous, heavily-documented companies deliberately kept in the sample to measure how much a model’s memorized knowledge inflates its score, rather than to prove the product.
Decision-quality vs outcome — Scoring whether the reasoning was right given what was knowable at the time, kept separate from whether the company later succeeded — a good decision can fail and a poor one can get lucky.
GEO / answer-engine optimization — Earning citation inside AI-generated answers (the way buyers increasingly research), through original research, authoritative content, and third-party validation.
Zero-click search — A search resolved on the results page or inside an AI summary, where the user never clicks through to a source.
References
Primary sources where available. Figures from forecasts or vendor research are labeled as such in the text.
McKinsey & Company — The State of AI (2025). https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
Fortune — McKinsey’s internal AI tool “Lilli” (June 2025). https://fortune.com/2025/06/02/mckinsey-ai-consulting-powerpoints-proposal-technology/
Accenture — $3B investment in Data & AI (June 2023). https://newsroom.accenture.com/news/2023/accenture-to-invest-3-billion-in-ai-to-accelerate-clients-reinvention
Harvard Business School / BCG — Navigating the Jagged Technological Frontier (HBS WP 24-013, 2023). https://www.hbs.edu/ris/Publication%20Files/24-013_d9b45b68-9e74-42d6-a1c6-c72fb70c7282.pdf
MIT NANDA initiative — The GenAI Divide: State of AI in Business 2025 (via Fortune). https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
Forrester — B2B Buyer Adoption of Generative AI (2024). https://www.forrester.com/report/b2b-buyer-adoption-of-generative-ai/RES181769
Pew Research Center — Google users and AI summaries (July 2025). https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
SparkToro — Zero-click search (2026). https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/
Gartner — Search engine volume to drop 25% by 2026 (forecast, Feb 2024). https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents
Ahrefs — AI Overviews and click-through rate (April 2025). https://ahrefs.com/blog/ai-overviews-reduce-clicks/
Menlo Ventures — 2025: The State of Generative AI in the Enterprise (Dec 2025). https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/
IDC — Worldwide AI and Generative AI Spending Guide (2024). https://www.idc.com/
World Economic Forum — Future of Jobs Report 2025 (Jan 2025). https://www.weforum.org/publications/the-future-of-jobs-report-2025/
PwC — 2025 Global AI Jobs Barometer. https://www.pwc.com/gx/en/issues/artificial-intelligence/job-barometer.html
McKinsey Global Institute — The Economic Potential of Generative AI (June 2023). https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier
About this report. Produced by Percision — the AI strategic-analysis studio. The proprietary analysis in Part 4 is generated by Percision’s strategic-analysis engine. For methodology details or media enquiries, contact the Percision team.