Robot scrubbing corrupted data and code to clean and optimize it

How CIOs can tell real AI agents from ‘agent washing’



Many AI products labeled as AI agents are little more than dressed-up chatbots or repackaged assistants, a practice called "agent washing." In a report published in May, Gartner warned IT leaders should be careful not to "mistake vendor positioning for true autonomy." But sorting the real from the fake is often easier said than done.

Real agentic capability is expensive to build. "Maybe 10 real players in the world do," said Oleksii Glib, founder and CEO at Acropolium, an international software development and technology consulting company.

"Almost everything else on the market is a wrapper around those players’ models, tools and keys. That’s exactly why so much gets rebranded as an ‘agent’: the real thing is expensive, so it’s cheaper to fake it," Glib said.

With products increasingly claiming agentic capabilities, how can CIOs identify the real thing?

Agents vs. other automations

One way is to focus on the particulars beyond the label and outside of the demo.

Related:Tips for successfully exiting AI vendor contracts

"When I’m looking at an agent platform, I don’t start by asking whether it’s autonomous. Every vendor can make a demo look autonomous," said Karthik Karunanithi, a solutions architect at IBM.

The first step in assessing an agent platform is to understand what type of automation the product actually delivers. Vendors often blur the lines between chatbots, assistants, robotic process automation (RPA) and agents, Karunanithi said, even though each serves a different purpose and provides a different level of autonomy.

The confusion over what is and isn’t an agent is understandable.

"Many products being called agents today are simply chatbots with a nicer interface, RPA workflows with a language model attached, or assistants that can draft an answer but cannot independently carry work through to an outcome," said Kevin Surace, CEO of TokenCore, an enterprise biometric identity-assurance company.

Automation categories

Each of these technologies serves a separate purpose and has distinct characteristics.

RPA, for example, follows a predetermined script: click here, copy this field, submit that form. It breaks when the process changes, explained Surace. Giving an RPA bot the ability to chat with you does not change its base functions.

By comparison, an AI agent has a goal, chooses its tools and often its data sources, and can determine and execute a series of actions on its own accord.

"A real agent can reason through variation. It can encounter missing data, conflicting information, a new system state or an exception, then determine what to do next based on the business objective rather than a rigid sequence of rules," Surace explained.

Related:The CIO hot seat: How to lead AI without becoming the scapegoat

Chatbots, on the other hand, have one job: to respond to user input in a single turn, explained Aman Mahapatra, chief AI strategy officer at Tribeca Softech, a technology consulting and venture/revenue acceleration firm.

"The user decides what to ask next, no state persists meaningfully between turns, and deviation from the script produces a fallback response or a handoff to a human," Mahapatra said.

The distinction comes down to how much independent decision-making the system can perform. In short:

  • RPA executes predefined, rule-based tasks with no adaptation.

  • Chatbots and assistants respond to user requests, sometimes chaining a few scripted actions, but waiting for a prompt at each turn.

  • AI agents are given a goal and pursue it with some level of autonomy over the intermediate steps, which are typically planning, selecting tools, adapting to changing conditions and executing multi-step work. They often perform this work with human checkpoints rather than full independence.

For CIOs evaluating vendor claims, those differences matter more than whether a product is labeled "agent."

A caveat: As brilliant and snazzy as agents can be, that doesn’t mean you actually need one. Even if a product has genuine agentic capabilities, CIOs should still ask whether its autonomy justifies the added cost and complexity of managing it.

Related:Your AI vendor is now a single point of failure

Glib said that in most cases, he recommends ignoring agents altogether. The key reason: "They just take the AI tokens, build a better interface and sell it at a higher price. It has no value for the buyer because you essentially can do it yourself on top of an LLM," he explained. If any vendor is set on convincing him otherwise, Glib said, they would have to "clearly articulate the additional value that they bring."

Telltale signs

If you can run only one test to separate the categories, let it be whether the system can handle a goal it has not seen before by composing its own tools or whether it fails outside its trained scripts.

"If composition under novel goals works, it is an agent. If not, it is one of the other three categories, regardless of what the vendor calls it," Mahapatra said.

But performance on a single test is only part of the evaluation. CIOs should also examine how the product is built and whether it can consistently deliver outcomes in production.

"Look hard at how a vendor has hardened their solution," advised Prasad Narasimhan Sulur, chief business officer at Certinia, a provider of enterprise resource planning software based on the Salesforce Platform.

"There’s a meaningful difference between a product that wraps an LLM in a thin interface and one with a clear architecture for data inputs, business rules and deterministic outputs in the workflows that require them," Sulur said. "The former may perform well in a demo and degrade quickly in production. The latter can sustain reliability as models and operating conditions change."

Transparency is another hallmark of an AI agent.

In the insurance industry, for example, "we are always wary of buying into terminology before we’ve understood what it really is," said Rhys Collins, managing director of #Total Systems, a specialized insurance software provider.

Whether a vendor calls something an AI agent, an assistant or an automation platform isn’t the most important question, according to Collins. What matters is understanding exactly what decisions the technology can make on its own and how those decisions can be audited. "In regulated industries, the level of transparency is often more important than the technology."

That transparency is key for any CIO in any industry to correctly assess whether an automated product is agentic. Ask the vendor to provide a trail of the agent’s reasoning and thought process. The vendor should be able to show the tools the agent called, which tools came back, how it used those tools, why the agent chose each next step and how it verifies that an action actually took place.

"If a vendor can’t show the decision trail and can’t explain how outcomes are verified, you’re probably looking at a workflow with an agent label attached to it. That’s like putting a spoiler on a sedan and calling it a race car," Karunanithi said.

Ultimately, CIOs should look for the right level of autonomy, backed by the oversight and accountability their organization requires.

"The goal is not to buy the most autonomous system possible. The goal is to deploy useful autonomy with accountable control," Surace said.

Original Post>

Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.

Leave a Reply