When can AI agents meet customer expectations? What are the right approaches, today and tomorrow?

We observe that Customer Services (insourced and outsourced – BPO) often display studies showing that “customers prefer to talk to humans rather than interact with AI robots”, but these Customer Services are nevertheless offering more AI services.

Customers do prefer to interact with humans, yet they buy products and services from sellers who strive to reduce the human cost by offering AI support solutions.


But is it economical to entrust all customer interactions to AI agents? Does the customer really win, does he really have a choice?

It is easy to understand that it is difficult to convince a customer to pay a premium at the time of purchase to be able to interact with a human in case of possible support needs,

To stay competitive, you need to find ways to make customer service operations less expensive without degrading quality by using AI-powered tools to boost the productivity of human agents.

What are the limits to consider?


In the case of basic, simple, repetitive interactions that do not involve major risks such as:

  • Financial risks
  • Reputational degradation
  • Risk of non-compliance
  • Legal risks

These interactions can be handled by AI agents.


On the other hand, during interactions that require:

  • Empathy to reduce customer stress.
  • Clarify the customer’s description of the problem.
  • Solve complex multi-source problems.
  • Consider various combinations of solutions with choices to be made
  • An understanding of the compliance issues at stake (financial, health-related, etc.) with a critical impact for the customer and the company, in terms of safety, security, etc.

Then humans can interact in a much cheaper and more efficient way.


What are the constraints to make humans efficient?

However, we must be aware that this efficiency at a controlled cost requires that effective tools be provided to agents:

  • Access to up-to-date processes.
  • Access to technical references.
  • Access to complex diagnostic methods (decision trees, etc.).
  • Climbing protocols.
  • Clear and smooth clearing protocols.
  • Contributing to the daily improvement of these tools by providing meaningful feedback.

When AI works… And when it doesn’t

Let’s be fair: there are cases where AI brings real value to customer services.

  • Instant answers to frequently asked simple and non-English questions,
  • 24/7 availability.
  • Ability to handle spikes in demand.
  • applications in several languages.

For well-defined routine tasks, automation can be really helpful.

The problem arises when you try to apply the same logic to the entire spectrum of customer service. Exceptional situations, complex problems, frustrated customers, or sensitive complaints require empathy, flexibility, and human judgment – precisely what AI can’t always deliver today.

A March 2025 report by McKinsey & Co. showed that 71% of companies are now using generative AI in at least one business function, but adoption is significantly lower in highly regulated industries: 63% in healthcare and 65% in financial services, precisely where errors have the most impact.


A proposal for balance

Maybe the question shouldn’t be “AI yes or no?”, but “how much AI and where?”. A smart hybrid approach would be:

  • Use AI to filter and categorize initial queries, recognizing their limitations
  • Truly automate simple, repetitive tasks where errors have little impact.
  • Implement human verification systems for critical responses (keeping humans in the loop, as recommended by Red Hat)
  • Facilitate quick (instantaneous?) access to a human when the situation requires it or when the AI expresses uncertainty (knowing that one of the problems of AI agents is to admit that they are not able to answer correctly… see our previous publications)
  • Include clear warnings about when information is AI-generated
  • Train human agents to work with AI tools that empower them, not replace them
  • Measuring success from the customer’s perspective in the foreground, not the company’s, success for the customer inevitably leads to the company’s success. The opposite is rarely true…
  • Be transparent about when customers interact with AI and when with humans.
  • Trust the critical thinking of your teams by involving them directly in the improvement of solutions provided to customers.

We can summarize this long list by recommending that your teams be provided with tools that use AI and manage your company’s knowledge while improving it with a Knowledge Management System (KMS) platform


The inconvenient question

Ultimately, a tricky thought remains: are companies adopting AI in customer service because it’s the best solution for their users, or because it seems like the solution to protect profitability? Is it a real innovation or an optimization disguised as modernity?

And do they do so with full knowledge of the real technical limitations of these systems? The data suggests not. When GPT-4 and GPT-4 Turbo, the most accurate models available, hallucinate 3% of the time; when advanced reasoning models such as O3 and O4-mini hallucinate 33% and 48% of the time, respectively; when OpenAI’s largest and most expensive model needs to be retired after just 4 months; When courts start holding companies accountable for the false information provided by their chatbots, it all suggests that the industry is trying to run before it learns to walk.

The answer likely varies from company to company, but the deafening silence on customer satisfaction studies, abandonment rates in automated systems, the number of users desperately seeking the “talk to a human” option, and the real, documented technical failures of LLMs, suggests that we may not be asking the right questions.

Technology is a tool, not a goal. And a tool is only useful if it solves the real problems of the real people who use it, and if it works reliably in the real world, not just in lab benchmarks.

As long as decisions about AI in customer service are made in the boardroom by looking at spreadsheets and business presentations from technology vendors, rather than talking to real customers and with an honest understanding of documented technical limitations, we will continue to see implementations that prioritize business efficiency over the human experience.

And perhaps most worryingly, we will continue to see companies surprised when their AI systems fail, when their customers become frustrated, and when they discover that short-term savings can be very costly when they measure up to the loss of customer reputation, trust, and loyalty.

Checkmate for Customer Service? When Knowing the Rules Is No Longer Enough!

Comparing a game of chess with a customer service interaction may seem unexpected at first. Yet, when you look closely at the structure and progression of both, the analogy becomes surprisingly insightful.

While their objectives differ radically—checkmate vs. customer satisfaction and problem resolution—both follow a similar escalation in complexity as the interaction unfolds.


Conceptual Parallels

Chess GameCall Center InteractionMeaning / Analogy
Openness strategyCall opening / reception.Set the tone and take control from the start.
Tactical combinationHandling of objections.Quick thinking to turn things around.
Late-game accuracyClosing of the call.Ensure resolution and satisfaction before finishing.
SacrificeOffer compensation or a gesture of goodwill.Short-term loss for long-term gain (loyalty or retention).
CheckmateCustomer satisfaction and resolutionAchieve the desired result in an efficient manner.
Critical errorCommunication error / violation of rulesA costly mistake that affects the results.
Pat (blocking)Deadlock / escalationNeither side achieved its goal.
Time pressureHigh call volume periodsDecisions under pressure; Compromise between efficiency and precision.

Now that the parallel between the 2 activities is clarified, the behavior of the AI on these interactions becomes interesting to observe:

An illustrative example of the limitations of LLMs (large language models) comes from documented experiments with chess.

 In March 2024, Chess.com held a showdown between ChatGPT and Google’s Gemini, where both systems could perfectly explain the rules of chess when asked directly, but then violated those same rules repeatedly during the game. Both bots constantly attempted to make illegal moves, and when they were informed of the error, they continued to come up with invalid moves.

Nikola Greb, an NLP data scientist and former ELO 2000+ junior chess champion, played several games against ChatGPT-4 in January 2024 and documented that the model played “like a grandmaster” in the opening first moves, but deteriorated significantly as the game progressed. ChatGPT-4 began to hallucinate, coming up with impossible movements even after being warned. Greb concluded that the overall rating of the system was below 1500, and observed something crucial: “No implicit rule learning has taken place – ChatGPT-4 still hallucinates at chess, and continues to hallucinate after the warning about hallucination. This is something that cannot happen to a human.

This disconnect between what an LLM can “say” and what it can “do” reveals a fundamental limitation: they don’t have real mental models of the world. In the context of customer service, this means that a bot can perfectly recite company policy but apply it incorrectly in specific situations, or it can explain how a product works without being able to diagnose a problem with it.

The Chatbot Chess Tournament 2025

In January 2025, a chatbot chess tournament aired on the GothamChess channel pitted professional chess engine Stockfish against seven generative AI chatbots, including ChatGPT, Google’s Gemini, and X’s Grok. The results were exactly what you would expect when language models try to play chess: decent opening moves followed by increasingly chaotic attempts to circumvent the laws of the game. The Snapchat chatbot decided that the pawns could move sideways like a tower, and when the error was reported, it repeatedly refused to continue saying “I’m sorry. I can’t engage in such a conversation. Let’s keep our conversation respectful.”

The problem of memory and context

LLMs have strict memory limits. While newer models offer wider windows of context, they still treat each conversation as relatively isolated. This means they can “forget” crucial information provided at the beginning of a long conversation, forcing customers to repeat themselves.

In one of the following articles, we will see how to avoid putting the customer in failure while making the best use of the undeniable capabilities of AI…