If the demo performs fine, why could it still be agent washing?
Because demo scenarios are almost always a fixed path the vendor designed and rehearsed themselves — they test whether a system can walk through a routine it has already practiced, not whether it can handle a situation it hasn't seen before. TheAgentCompany, the benchmark Carnegie Mellon developed with Salesforce, quantifies this exact gap: in a simulated company environment requiring autonomous judgment and cross-system actions, even the top-performing model completed only about 30% of tasks autonomously.
That means a smooth demo only proves "the rehearsed path works" — it says nothing about the more important question of what happens with a different scenario.
Why is cross-session memory one of the key signals for telling real agents from fake ones?
Multi-step autonomous tasks usually span multiple interactions — a plan started today might need adjusting tomorrow based on new information, which requires the system to remember what was previously discussed and decided. If every interaction starts from zero, the system is fundamentally a stateless Q&A tool, and no matter how it's packaged as an "agent," it can't genuinely sustain a task across time.
The test is straightforward: end one conversation, start a new one, and directly ask "where did we leave off." A system with real memory will give specific content; a stateless tool usually gives itself away, either failing to answer at all or offering a vague, generic response.
How does Gartner's 40% project cancellation figure directly relate to agent washing?
Gartner forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs and unclear business value. This figure connects directly to agent washing because many projects were never actually verified during procurement to be genuinely autonomous systems — buyers evaluated demos and marketing language, and only after deployment, when they discovered the system needed manual approval at every step, did they realize they'd purchased a repackaged automation script. By then, the project's cost structure could no longer pay off, and cancellation became the only option.
This is also why writing tests like "can it handle a situation it hasn't seen before" into procurement acceptance criteria matters more than watching a demo — it moves the point of verification earlier, before the money is spent, rather than discovering the gap after deployment.
If I've already signed a contract and later suspect agent washing, what should I do?
Start with the four tests described above (handling unfamiliar scenarios, cross-session memory, whether the vendor claims full autonomy, and capability differences versus a chatbot), and document the results concretely — this becomes the evidentiary basis for either pushing the vendor to improve or renegotiating contract terms. If the contract already specified concrete acceptance criteria (such as "must autonomously complete N steps"), a suspected gap can be argued directly against those terms; if only demo performance was accepted at signing, this is harder to argue on contractual grounds alone, but it still serves as leverage at renewal or when negotiating add-ons.
The more fundamental fix is to turn this gap into an internal procurement checklist — running the four tests before signing any future system, moving the point of verification firmly ahead of the money being spent.
When every pitch deck says "autonomous," "self-directed," or "AI employee," those words have nearly stopped telling you anything. Agent washing describes exactly this phenomenon: rebranding chatbots or rule-based automation as autonomous agents in marketing, when genuine autonomous reasoning isn't actually there. The problem is that simply hearing a vendor claim "we're a real agent" tells you nothing — you need a few ways to actually test it.
Demos are the most misleading part of the sales process precisely because they run through a fixed path the vendor designed and rehearsed themselves. TheAgentCompany, a benchmark developed by Carnegie Mellon University in collaboration with Salesforce, tests leading models inside a simulated company environment — complete with CTO, HR, and engineering roles — requiring agents to autonomously complete tasks using internal chat tools, company handbooks, and websites. The result: even the top-performing model completed only about 30% of office tasks autonomously, with most models scoring well below 35%. That gap is exactly the divide between "performs fine in a rehearsed demo" and "handles the uncertainty of a real situation" — and demos usually only show you the former.
Instead of asking a vendor "are you a real agent," ask a few concrete questions you can test live or verify with a free trial:
First, can it handle a situation that wasn't scripted in advance? Slightly alter the input from the demo — a task type that wasn't demonstrated, or a request containing conflicting information — and see whether the system gets stuck, loops, or responds reasonably. Rule-based automation typically fails outright or produces an unrelated response when it hits an input outside its script.
Second, does it retain memory across sessions? End one interaction and start a new one, then ask whether the system remembers what was discussed or decided previously. If every interaction starts from zero, that's usually a sign of a stateless tool rather than an agent with persistent context.
Third, does the vendor insist the system is "fully autonomous with zero human involvement in high-stakes decisions"? This is one of the clearest red flags. No system operating in an enterprise context today is genuinely fully autonomous without human oversight — a vendor claiming otherwise usually means the definition has been deliberately loosened, or the vendor underestimates the system's actual limitations.
Fourth, what can this system do that a well-configured chatbot can't? If the answer is vague, or only points to degree-level differences like "answers questions more accurately" rather than capability-level differences like "can autonomously plan multi-step tasks, call tools, and adapt based on results," it's likely a relabeled existing tool.
Gartner has already flagged agent washing as a major due-diligence risk in enterprise AI procurement, and forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs and unclear business value — a wave of cancellations that traces substantially back to procurement decisions that never tested the packaging in the first place. In practice, writing the four questions above into contract acceptance criteria carries far more weight than a verbal vendor promise — for example, explicitly requiring that "the system must autonomously complete at least N steps on a task type it was not pre-configured for," backed by a defined acceptance test scenario, rather than accepting whatever the demo day happened to show.
If you're a buyer, the direct cost of being misled by agent washing isn't "the product doesn't work well" — it's misallocated resources: you expected a system that reduces human involvement, but instead you're manually approving every critical step, gaining no time savings while taking on the extra overhead of maintaining a system that looks smart but is actually heavily dependent on people. If you're an investor, the gap shows up directly in the business-model assumptions you're underwriting — a company that repackages Level 1 or 2 automation as a Level 3 or 4 autonomous system likely has its cost structure, moat, and scaling path all built on a mistaken capability assumption. Before you sign or invest, running the four test questions above will usually tell you more than reading the entire pitch deck.