Ask for Attacks, Not Feedback
AI has two evaluation modes, and only one is biased toward telling you what you want to hear. Stop asking for review; ask it to break a specific commitment.
In the summer of 2002, the U.S. military ran Millennium Challenge, a wargame reported to have cost around a quarter of a billion dollars, pitting a technologically dominant Blue force against a scrappy Red force led by a retired Marine lieutenant general, Paul Van Riper. Van Riper’s assignment was to lose convincingly enough to validate Blue’s doctrine. Instead he attacked. Anticipating that Blue would monitor his electronic communications, he issued orders by motorcycle courier and coded signals; anticipating Blue’s naval dominance, he swarmed its fleet with small boats and cruise missiles in a preemptive strike. In the exercise’s opening, he was judged to have sunk or disabled much of Blue’s fleet, sixteen ships by widely cited accounts. The response was telling: the game was halted, the ships notionally refloated, and play resumed under constraints on the Red force. Van Riper and several journalists called the rerun scripted toward a Blue win; officials countered that the constraints were needed to test later phases, though the after-action report itself conceded the opposing force had been “constrained to the point where the end state was scripted.” Either way, the most valuable finding of a quarter-billion-dollar exercise came from the one participant told to break the plan, and the institution’s instinct was to walk it back.
That instinct is the one this essay is about, because you are now carrying an evaluator in your pocket that will do either thing on command (validate your plan or break it) and almost everyone asks it for the wrong one. AI has two distinct evaluation capabilities, and only one is biased toward the happy path: ask it to “review” and you get agreeable, comprehensive-looking feedback that finds little; ask it to break a specific, stated commitment and you get genuine adversarial search. The single highest-leverage change in how you use it is to stop asking for review and start asking for attacks.
What follows: why the review reflex produces coverage theater, why a named commitment is what unlocks the machine’s adversarial mode, and how to turn every successful attack into a permanent defense.
The strongest version of “just have the AI review it”
The reasonable-sounding practice is to paste your plan, your contract, your model, your campaign into an AI and ask it to review: to check your work, flag issues, offer feedback. Concede that this is not worthless. The model will catch typos, surface an omission or two, and produce a tidy list of considerations that reads like diligence. For genuinely rough drafts, that pass has value, and asking is better than not asking. If your work is bad, review will tell you it is bad in reassuringly organized prose.
The trouble is that “review this” invokes the mode of the agreeable consultant, and the agreeable consultant is trained to be helpful, balanced, and reluctant to tell you your plan is a disaster. Ask for feedback on a pricing scheme and you will get a competent, hedged survey (strengths, some risks to consider, a few suggestions) that carefully avoids the one sentence that matters: here is exactly how a competitor arbitrages this, and here is the customer segment that walks. The review reads as thorough precisely because it covers everything shallowly, which is the signature of coverage theater. It gives you the feeling of having been challenged without the substance of it, and that feeling is more dangerous than no review at all, because it retires the very doubt that should have kept you sharp.
So the honest version of the steelman is this: review catches surface defects and is worth a pass, but it runs in the mode structurally biased against finding the failure that actually sinks you.
The category error
The mistake is treating “review this” and “break this” as the same request phrased two ways. They are different instructions that invoke different searches, and the difference is the whole game. “Review” points the model at the space of reasonable-sounding commentary and it returns the most typical, agreeable member of that space. “Find the specific way this fails” points it at the space of adversarial strategies against a stated target, and (because it has ingested vast troves of how plans, contracts, pricing schemes, and comp plans have actually been gamed and broken) it returns something genuinely useful: an attack. The capability was there the whole time. The prompt decides which of the two you get, and most people reflexively choose the one that flatters them.
The unlock is a named commitment plus an instruction to defeat it. Not “review this comp plan” but “here is the payout formula; find the sales behavior that maximizes commission while minimizing company value.” Not “any thoughts on this contract?” but “you are the counterparty’s counsel; find every reading of this clause that favors your client.” Not “review the forecast” but “find the demand scenario this model assumes away and show me the quarter where it breaks.” Each of these gives the machine a concrete thing to attack and a concrete adversary to attack as, and the output changes character completely: from a hedged survey to a specific, actionable failure. This is the same instinct behind the pre-mortem that the psychologist Gary Klein described in Harvard Business Review in 2007: rather than asking a team to critique a plan, tell them the plan has already failed catastrophically a year from now, and have them explain why. Assuming the failure as given is what frees people (and models) to go looking for it. Prospective doubt is polite. Retrospective certainty is ruthless.
And attacks, unlike feedback, compound. A piece of review is consumed and forgotten. A successful attack is a discovered failure mode, and a discovered failure mode can be written down and made permanent: converted into a clause your contract library now always includes, a stress case your model now always runs, an item your launch checklist now always checks. This is the business-side twin of a point the engineering essays in this house have made about evaluation: that the way you turn a product from fragile to durable is to convert every failure you find into a standing test that runs forever after. The attack is how you find the failure. The standing check is how you make sure it never returns.
Where the tooling comes in
Structurally, this means keeping a dedicated adversary in your process rather than hoping the author remembers to attack their own work. A multi-model workspace like AODex makes that cheap to institutionalize: configure a persona whose entire instruction is to break a stated commitment (the counterparty’s lawyer, the arbitraging competitor, the churning customer) and run every important artifact past it before you commit, using a different model than the one that drafted the work so its blind spots don’t match the author’s. And the security layer makes the “attacks become standing checks” loop literal: AOSentry, the governance layer within AOCore, ships a jailbreak-and-injection detection guardrail precisely because the discipline of turning discovered attacks into permanent, automated checks is how you stay ahead of adversaries who keep inventing new ones. The tooling doesn’t invent the attack. It makes running one cheap and keeping one permanent.
What to do
Delete the word “review” from how you use AI on anything that matters, and replace it with a named commitment and an order to break it. State what you are committing to (this price, this clause, this forecast, this plan) name the adversary who benefits from its failure, and instruct the machine to find the failure as that adversary. Then do the thing the wargame’s referees refused to do: when the attack succeeds, do not reset the board and proceed to the outcome you wanted. Write the finding down, generalize it into its failure class, and turn it into a standing check that every future plan of that kind must survive.
The agreeable consultant will tell you your plan is fine. The adversary will tell you how it dies. Only one of them is worth the tokens. Ask for attacks, not feedback.
The Constraint Is the Business
- 1. Nothing Your Business Does Is Novel
- 2. Closable vs. Open: Which Business Problems AI Can Actually Finish
- 3. The Scorecard Is the Moat
- 4. The Brief Was Always the Expensive Part
- 5. Why "Human Review" Fails, and What Adversarial Oversight Looks Like
- 6. Irreversibility: The Only Sane Way to Allocate Executive Attention
- 7. Ask for Attacks, Not Feedback