Users Don't Want Magic. They Want Scaffolding.

The biggest factor in LLM adoption isn't model quality. It's whether the experience produces calibrated trust, and the research is clear enough to design to.

Part 3: Users Don't Want Magic. They Want Scaffolding.THE MODEL ISN'T EVERYTHINGAOCYBERPART THREEUsers Don't Want Magic.They Want Scaffolding.FRAMEWORK · THE TRUST SCAFFOLDING CHECKLISTJUSTIN DONNARUMAAOCYBER.AI

Users Don’t Want Magic. They Want Scaffolding.

In late February 2024, Klarna announced that its OpenAI-powered customer service assistant had handled 2.3 million conversations in its first month (roughly two-thirds of the company’s customer service chats) and was doing the equivalent work of 700 full-time agents. CEO Sebastian Siemiatkowski called it a breakthrough in customer interaction. The press release was unambiguous. The economics looked transformative. The deployment was, by every metric the company publicly disclosed, a success.

By mid-2025, the same CEO was publicly recalibrating. Bloomberg quoted Siemiatkowski saying that “cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.” Customer complaints about the lack of human support had piled up. The company began rehiring on the human side. The reversal was not because the model got worse. The model got better over the same period. What failed was not the AI; it was the relationship between the AI and the people it was supposed to serve. Klarna had built a system optimized for confident answers and discovered, after the trust ran out, that confidence without recourse is a different product than what their customers had been willing to pay for.

The Klarna arc is the most-documented public example of a pattern that plays out quietly in dozens of enterprise AI deployments at smaller scale. The gap between AI products that get adopted and AI products that die in production is now, in 2026, almost entirely a workflow design problem rather than a model quality problem. The single biggest factor (bigger than which provider you’re on, bigger than how well your retrieval is tuned, bigger than your latency numbers) is whether the user experience produces calibrated trust. If users over-trust your system, they get burned by the first confident wrong answer, and you lose them. If users under-trust it, they never get value, and you lose them. Both failures look identical from the dashboard: low adoption, declining usage, eventual abandonment. Neither has anything to do with the model.

Calibrated trust is a UX problem. It’s a workflow design problem. It’s a process problem. The research on what builds and breaks that trust is now mature enough to design against, which is what this essay is about.

The strongest version of “users want magic”

The magic-first view has a defensible version that deserves a fair hearing before I dismantle it.

The case for designing AI products as “magic” (fluent, confident, unhedged, answer-shaped) comes from real evidence about what users say they want when you ask them. Surveyed users do, consistently, prefer confident answers over hedged ones. Focus groups complain when AI systems say “I’m not sure.” Demos with caveat-heavy outputs poll worse than demos with declarative ones. There is a defensible reading of all this that says: people came to AI to get answers, not to get more questions, and every hedge you add to the output makes the product feel worse.

There is also a deeper version of the argument. The model is going to keep getting better. Some of the hedges users complain about are the model being overcautious: refusing to commit to claims it could safely make, refusing to answer questions that have correct answers. As models get smarter and better-calibrated, those false hedges go away, and confidence becomes appropriate. The “design for magic” view says: stop building scaffolding around model output, because the model is going to internalize the right level of confidence on its own. The scaffolding is a temporary crutch.

There’s partial truth in this. Some hedging is overcautious. Some confidence is appropriate. Frontier models in 2026 are better calibrated than 2024 models were, and they will keep improving. I am not arguing for hedge-everything UX. I am arguing for something more specific.

The part that doesn’t survive contact with the data is the assumption that user-stated preference for confident answers maps onto user behavior with confident-but-wrong answers. It does not. Anthropic’s published research on sycophancy in language models documents what should be uncomfortable reading for any product team designing for “friendly and confident”: models trained to be agreeable systematically produce worse answers when their agreement is wrong, and the agreeable tone correlates with reduced user trust over time rather than the other way around. The model that agrees with you reads as untrustworthy. The model that maintains neutral demeanor and occasionally disagrees reads as authentic and gains trust.

That result is the opposite of what most product teams optimize for. Most product teams have been told, often by the same surveyed users who shelve the products, that the system should be agreeable and confident. Then the survey respondents shelve the products and the post-mortems blame the model.

A separate study published at FAccT 2024, “I’m Not Sure, But…” (Kim et al., arXiv:2405.00623), looked at the effect of first-person uncertainty expressions on user behavior in a pre-registered N=404 experiment. When the model said “I’m not sure, but I think the answer is X,” users were measurably less likely to over-rely on incorrect answers compared to the same answer presented without the hedge. Importantly, this effect is not about the user discounting all answers; it’s about the user paying attention to the hedge as a signal that this particular answer warrants checking. The hedge is information, and a well-calibrated user reads it as information. The hedge that the focus group hated in the abstract turns out to be the hedge that prevents the product from blowing up in their faces in practice.

Put plainly: the users want magic argument is correct about what users say and wrong about what users need. Designing for stated preference produces products that demo well and die in deployment. Designing for calibrated trust produces products that feel slightly less impressive in the demo and stay running for years.

Why fluency is dangerous

There is one more piece of the trust research worth pulling out because it changes what you optimize for.

A 2025 paper, “Adjust for Trust” (Srinivasan & Thomason, arXiv:2502.13321), published at the 31st International Conference on Intelligent User Interfaces, made the trust-calibration problem operational with measured numbers. In their medical-diagnosis study, doctors accepted 26% of AI misdiagnoses when their trust in the system was high, compared with 8% when their trust was lower. The same study showed doctors rejected correct AI diagnoses 68% of the time when their trust was low, against 40% when it was higher. Both error modes (over-reliance and under-reliance) are functions of miscalibrated trust, not of model accuracy. The more fluent the output, the less the user checks. The more authoritative the prose, the less the user notices when the prose is wrong.

This is not a model-level finding. It’s a human-cognition finding. Even when the model is perfectly calibrated internally, the way its output is presented to the human reader changes how the human reader engages with it. Beautiful prose suppresses skepticism. Confident framing suppresses skepticism. Pretty UI suppresses skepticism. The polish that makes your AI product feel professional is the same polish that prevents users from noticing when it’s wrong.

This is the core tension in AI UX design. The things that make the product feel like a good product (fluent language, confident tone, clean presentation) are the same things that make the product behave badly when its output is wrong. You cannot design your way out of this tension by making the model better. The tension lives in the relationship between human cognition and authoritative-sounding language, and it requires explicit workflow design to manage.

A trust scaffolding checklist

The checklist I now run against AI product designs before they ship has six items. Each one is binary. Score yourself; below four out of six and your adoption ceiling is structural, not addressable through more model spend.

The Trust Scaffolding ChecklistA six-item binary checklist for AI product UX. Score yourself before shipping. Below four of six and the product's adoption ceiling is structural, not addressable through more model spend. Calibrated trust is the leading factor in LLM product adoption.The Trust Scaffolding ChecklistSix binary questions. Score yourself before shipping. Below 4/6 and the adoption ceiling is structural.1Every claim has a source linkNot "sources available on request." Every assertion has a citation in the user's primary path,clickable, leading to the document the model actually saw.2The product can say "I don't know", and means itAbstention is a real path through the system. Explicit conditions trigger it. The messageis honest about what happened, not a generic "I cannot help with that."3High-stakes actions require confirmation, not consentSend-email, write-to-DB, file-ticket, move-money: every one surfaces to a human beforeit executes. Install-time consent is not action-time consent.4Users can see what was retrieved before they see the answerIn the same screen the answer renders in, not three clicks away. Even if most usersnever look, the act of having it there is what makes the system feel honest.5Confidence is surfaced when low, not hidden when highAsymmetric hedging. The user learns over time that an unhedged answer is trustworthy anda hedged one needs checking. Constant hedging is noise; never hedging is danger.6One-click "this is wrong" path that goes somewhereThe flag routes to a place an eval pipeline or a human reads. The act of providing thepath increases trust. The signal it captures makes the next version better.ADOPTION CEILING BY SCORE0–1CONFIDENCE TRAPUsers will over-trust untilthe first confident wronganswer. Then they leave.No model upgrade fixes this.2–3PARTIAL TRUSTAdoption stalls in pilot.Power users get value;casual users churn.Ceiling is structural.4–6CALIBRATED TRUSTUsers build calibratedreliance over time.Confident answers trusted;hedges read as signal.CHI 2026: sycophancy reduces perceived authenticity. Fluent prose suppresses user vigilance.The model is rarely the bottleneck on adoption. The UX almost always is.aocyber.ai · AOSentry · AODex
The Trust Scaffolding Checklist

One. Every claim has a source link. Not “the system can show sources if you ask.” Not “sources are available in the audit log.” Every assertion the system makes has a citation rendered in the user’s primary path, clickable, leading to the document the model actually saw. This is the single most adoption-positive UX choice in AI product design, and it is the one that the product I mentioned in the opening had and the one that died did not. RAG with citations is not a feature. It is the default state from which you sometimes, with justification, depart.

Two. The product can say “I don’t know” and means it. Abstention has to be a real path through the system, not a fallback the model invokes ten percent of the time when it gives up. There need to be explicit conditions (retrieval confidence below a threshold, conflicting sources, questions outside the system’s domain) that route to an abstention response rather than a synthesis. The abstention itself has to be honest: not “I’m sorry, I can’t help with that,” but “I don’t have a confident answer because the documents I have access to don’t directly address this; here’s what I did find, here’s where I would look next.”

Three. High-stakes actions require confirmation, not just consent. If your agent is going to send an email, write to a database, file a ticket, or move money, the action has to surface to a human before it executes. Consent at install time is not consent at action time. The friction of confirmation is the price you pay for the trust necessary to scale automation. Skipping it is the fastest known way to destroy a deployment.

Four. Users can see what the model retrieved before they see the answer. This is subtle but it matters: in the same screen the answer renders in, the user can glance at the retrieved context that produced it. Not in a separate “transparency view” three clicks away. In the path. Even if 95% of users never look at it, the 5% who do are your power users, your auditors, your trust establishers, and the act of having it there is what makes the system feel honest even to the users who don’t open it.

Five. Confidence is surfaced when low, not hidden when high. Most teams either always show confidence (annoying) or never show it (dangerous). The pattern that produces calibrated trust is asymmetric: show uncertainty when it matters, stay quiet when it doesn’t. The user learns over time that an unhedged answer can be trusted at face value and a hedged one warrants a check. The hedge becomes a meaningful signal precisely because it isn’t on every output.

Six. There’s a one-click “this is wrong” path that goes somewhere. A button, a flag, a thumbs-down, and crucially, those flags route to a place where someone (or some eval pipeline) reads them. The act of providing the path does two things. It tells users that the system knows it can be wrong, which paradoxically increases trust. It also feeds the eval pipeline that makes the next version less wrong. The teams that build this path before they need it are the teams whose products keep getting better. The teams that bolt it on six months in have already lost the data window where they could have improved fastest.

Six binary questions. Most enterprise AI products score zero to two. The ones that score four-plus are the ones that survive the pilot-to-production cliff.

Where this lives architecturally

A trust scaffolding UX is downstream of an architecture that captures the evidence the UX needs to display. Source citations require knowing what the model actually saw. Abstention requires knowing the retrieval confidence at query time. The confirmation step requires intercepting the action before it executes. The “what did the model retrieve” view requires having logged it. The wrong-button feedback requires routing the user’s signal to a place that can use it.

None of those are model capabilities. All of them are gateway capabilities, or, more precisely, capabilities of the architectural layer that sits between the model and the application surface. If every model call passes through one place that captures the retrieved context, the confidence scores, the assembled prompt, the completion, and the user response, then trust scaffolding is something you turn on. If those calls go directly from your application to a provider API, trust scaffolding is something you build from scratch every time, and most teams don’t.

This is the case I keep making for AOCore on the build path: not that the gateway is magical, but that the gateway is where the receipts live, and trust scaffolding without receipts is performance art. “Cited sources” only mean anything if you can trace the citation back to a document the model actually saw, and that traceability is an architectural property, not a frontend feature.

For organizations that aren’t going to staff a workflow design team to learn these lessons by hand, AODex is the version of this UX implemented as a product. RAG answers ship with source citations by default. Knowledge bases are auditable. The system can say it doesn’t know and means it. PII tokenization happens before any data reaches a model provider, and the no-training-on-customer-data guarantee is in the architecture, not a contract clause. The “this is wrong” path exists and goes somewhere. The buyer doesn’t have to design any of this; the design choices have already been made, and they have been made in the direction the research in this essay points.

The product team I opened with (the one that shelved their pilot) lost on a fifteen-minute design argument that they did not have the framework to win. The other team won by accident; the engineer who pushed for citations had read the research. The decision was made by one person on one Friday afternoon, and it determined the fate of the deployment.

You can leave that decision to chance, or you can run the checklist. The model is the same either way. The product is not.

← Back to Blog