Customer service is where customers expect to receive a clear answer to any question they ask. Just one incorrect AI response to a customer can turn into a major public scandal. A simple screenshot on social media, a complaint in reviews, an escalation to second-line support, and your reputation takes a hit that outlasts the conversation that caused it. AI customer service promises speed and scale, but the cost of a mistake here is higher than with any in-house tool.
Most procurement decisions regarding AI customer service are made in reverse order: the team first selects a channel (chat, voice, agent assist) and a vendor to support it, and only then (if at all) checks whether the knowledge base for that channel can provide accurate answers. Choosing a channel is the easy part. But the knowledge base behind it is what truly determines whether the customer receives helpful assistance or a hallucination.
That’s why today we’ll break down what AI for customer service entails, how these tools work mechanically, and why most implementations fail to meet expectations. We’ll also explore how best to evaluate vendor choices and what a truly effective implementation plan looks like for your first few months.
What counts as “AI customer service” today?
If you look at the customer service AI market today, you’ll realize that it’s not just one tool but four distinct categories. However, when evaluating vendors, business owners often confuse these categories with one another. That’s why we suggest taking a closer look to understand what to choose and how they differ:
- Chatbots / conversational AI. This is an interface through which the customer communicates directly via text, answering questions and performing simple actions without an agent’s involvement.
- Agent assist. This is an “assistant.” It does not replace an agent but provides them with pre-written responses, links to policies, and real-time prompts during a call or chat, leaving the final decision to the human agent. This reduces the time an agent spends searching for an answer, but adding an AI layer introduces certain risks.
- Voice AI. This is the voice equivalent of a chatbot, capable of carrying on a conversation, understanding the intent behind speech, and either resolving the issue or, if the Voice AI does not understand or does not know how to resolve the customer’s issue, transferring the call to an agent with the full context (meaning the customer does not have to repeat the information).
- Agentic AI. This is autonomous issue resolution: the system doesn’t just respond; it acts (processes a return, updates account information) without human intervention at every step. That said, modern systems can also conduct empathetic conversations and understand a customer’s intent based on intonation.
To make it easier to compare and understand what each option does, we’ve put together a small table. This will help you clearly see which option is the best fit for your specific business:
| Category | What It Does | Who Interacts | Level of Autonomy |
| Chatbots / conversational AI | Answers customer questions in text format | Customer directly | Medium: can resolve a simple question end to end |
| Agent assist | Suggests answers and policies to the agent in real time | Agent, not the customer | Low: final decision stays with the human |
| Voice AI | Conducts a voice dialogue, understands speech | Customer directly | Medium-high: depends on scenario complexity |
| Agentic AI | Executes the action end to end, autonomously | Customer or system, no agent at each step | High: minimal human involvement |
All four categories have one thing in common: each is only as good as the knowledge on which it relies. A beautiful chatbot interface won’t save the day if the knowledge base contains contradictory or outdated return policies (or any other information related to your business).
That is precisely why it makes sense to evaluate AI for customer service not by channel, but by the knowledge layer underlying it. AI customer support and AI for customer experience are, in essence, two different perspectives on the same task: the first looks at it from the perspective of the support tool, while the second looks at it from the perspective of the experience the customer ultimately receives. Both perspectives rest on the same foundation.
When a team evaluates customer service AI, it naturally assesses vendors from different categories. At the same time, you may not even consider that your comparison might address different problems. For example, agent assist reduces the workload on the existing team, whereas agentic AI changes the very model of how inquiries are handled. Mixing up these categories during the evaluation phase is a common reason why a purchased tool fails to solve the original problem for which it was selected.
How AI customer service tools actually work
Regardless of the channel, the mechanism is the same: a request is received → the system determines the intent → finds relevant knowledge → generates a response or action → either resolves the issue or escalates it to a human.
But, of course, examples always make things clearer. Let’s imagine a customer contacts you about a delayed delivery. They write in the chat: “Where’s my order? It was supposed to arrive yesterday.” The system recognizes the intent (a delivery status inquiry), consults the knowledge base for the current policy on delays and the status of the specific order, generates a response, and either resolves the issue or forwards it to an agent if the situation is atypical (for example, if the delay exceeds the standard compensation threshold).
If another customer asks the same question but over the phone, the process is essentially the same. The only difference is that the chat is replaced by speech recognition. The same principle applies to a conversation with an agent during a call via agent assist; the result is simply not sent directly to the customer but is instead provided to the agent as a prompt.
It is precisely the retrieval step (searching for relevant knowledge) that breaks down first when the knowledge base is cluttered, regardless of the channel. If the delivery delay policy exists in three versions (one current, two outdated), voice AI, the chatbot, and agent assist are equally likely to select the wrong version. This is because all three channels rely on the same knowledge layer.
This is a key observation for any team evaluating customer service AI. Discussions about “which model is better” or “which voice engine sounds more natural” often distract from a far more practical question: where does the system actually get the facts it uses to form its response? A model may be flawless in terms of the naturalness of its speech. But all of that is forgotten as soon as the model starts to hallucinate, with an uncontrolled retrieval layer. The same applies to AI customer support tools focused on text-based channels: the naturalness of the wording does not guarantee the accuracy of the content.
Why most AI customer service deployments underdeliver
If you’ve already tried implementing AI in your support service, you’ve likely encountered the most common problems. For example, if we’re talking about a chatbot, it most likely repeats an outdated return policy (or something else that’s out of date) because the new version exists in the database but isn’t marked as the single source of truth.
Voice AI incorrectly announces a rate that changed a month ago – again, because the voice model reads out whatever it finds, not what’s currently valid. Agent assist suggests an article to the agent in the middle of a call, an article that has been officially replaced, and the agent, trusting the system, repeats the error aloud to the customer (or doesn’t repeat it if the agent has been with the company for a long time and knows this is the wrong version).
In none of these cases is the cause the model or the channel. The cause lies in the knowledge on which the model bases its response.
This is a systemic problem, not isolated incidents. According to Shelf’s corporate content analysis, 94% of files in a typical enterprise knowledge base contain at least one issue that affects the quality of GenAI responses.
If you think this is just our opinion, you’re mistaken. After all, according to Gartner data, by 2029, about 80% of customer service organizations will be using generative AI in one form or another. In other words, adoption is practically universal, and the question that remains open for most of them is not “whether to use AI,” but “what knowledge it runs on.”
Closing this knowledge gap must happen before or in parallel with selecting a channel and a vendor, not after the chatbot has already been launched and started responding to real customers. Learn more about how to systematically build this discipline.
This also explains why the AI customer support industry as a whole is seeing a gap between the widespread adoption of AI and actual results: organizations are connecting models to their support channels en masse. But far from all of them are simultaneously investing in ensuring that the knowledge behind these channels is manageable. Consequently, we see a gap between the number of companies that “use AI” in name only and the number that actually see fewer escalations and higher CSAT.
AI customer service use cases by channel
Chat-based deflection
Today without AI: A customer messages the chat, waits in line for a live agent, and receives a response after several minutes, even to a simple question like “where is my order?”
With AI: A chatbot instantly answers common questions such as order status, return policies, and business hours. It only forwards cases to an agent that truly require human judgment.
Voice AI Replacing IVR
Today, without AI: Customers navigate an annoying IVR menu (“press 1 for…”). Often, customers can’t even find the option they need and end up waiting for a live agent anyway. This just wastes time and adds to their frustration.
With AI: Voice AI understands natural speech from the very first sentence, eliminates the need to navigate through menus, and either resolves the issue immediately or transfers the call to the appropriate agent with the call context already established.
Agent Assist for Live Calls
Today, without AI: The agent manually searches for the right knowledge base article in the middle of a call, switching between tabs while the customer waits on the line. Moreover, this is most often done by keyword, which isn’t always convenient.
With AI: The agent sees relevant prompts and precise policy wording right during the conversation, without having to search manually.
Multilingual 24/7 support
Today without AI: Multilingual support requires either a staff of native-speaking agents for each time zone or limiting operating hours for lower-priority languages.
With AI: The same governed knowledge layer handles inquiries in different languages around the clock, without the need to manually duplicate content for each language.
Post-interaction QA and insights
Today without AI: The quality of calls and chats is checked manually on a random sample of a small percentage of interactions.
With AI: the system analyzes 100% of interactions, identifying patterns (which topics most often lead to escalation, where agents most often provide inaccurate answers), and uses this to update the knowledge base itself.
This scenario is often underestimated during the initial evaluation of AI customer support vendors because it doesn’t create a visible customer experience, unlike a chatbot or voice AI. But it is precisely this that closes the loop: without post-interaction QA, the governed knowledge layer degrades over time just as the original knowledge base did. It just happens a little more slowly because no one sees exactly where new gaps are accumulating.
A telling example of how this works in practice: Lowell Nordics, a financial services provider specializing in debt management, faced three challenges:
- Low-quality information used by agents
- Increasing wait times on hold
- The lack of a single source of truth, which caused processes across different departments to diverge.
After implementing Shelf and integrating agent assist with their Genesys Cloud, the company increased first-contact resolution from 83% to nearly 90%, reduced average handling time by 25% (from 8 to 6 minutes), and boosted its Net Promoter Score by 111%, from 26.1 to 55.2.
These five scenarios are not competing use cases from which you must choose just one. In practice, a mature AI for customer experience strategy combines multiple channels simultaneously, drawing on a single, governed knowledge layer: a chatbot handles routine inquiries, agent assist supports agents on complex calls, and post-interaction QA continuously highlights where the knowledge base needs updating.
Benefits and ROI
When governance is set up correctly, the impact of AI customer service can be measured in specific operational metrics:
- Deflection rate. The percentage of inquiries resolved without a human agent; this reduces the support team’s workload on routine requests.
- First-contact resolution. The percentage of issues resolved on the first contact, without follow-up calls or emails from the same customer.
- CSAT. Customer satisfaction depends directly on the accuracy and speed of the response, not just on the presence of AI itself.
- Average handle time. A reduction in handling time because the agent (or AI) does not have to search for the necessary information manually.
- Time-to-competency for new agents. The governed knowledge layer reduces the time it takes for a new agent to start working independently, because they don’t need to memorize where to find the current version of each policy.
A real-world example: A Fortune 50 healthcare company violated its SLA for first-contact resolution due to content ROT. This resulted in the agent assist system consistently providing poor responses, and user trust in the system plummeted. After implementing Shelf, the company eliminated content ROT in over 60% of documents, achieved a 225% increase in content utilization, and reached an FCR rate above 95%, a historic high for the company, thereby resolving the contract violation with a key client.
Important: None of these metrics are related to staff reductions. AI customer service and AI customer support function here as team enhancers. They relieve agents of routine tasks rather than completely replacing their roles.
The value of these metrics is evident only when considered in conjunction with one another, not in isolation. A high deflection rate coupled with a low CSAT indicates that the system is closing cases formally but not resolving the customer’s problem. In other words, it’s simply masking unresolved issues with a faster response. This is precisely why the evaluation of customer service AI should be based on a set of metrics, rather than a single figure that’s easiest to showcase in a presentation to senior management.
What to Look for When Evaluating AI Customer Service Vendors
When selecting a vendor, you should focus on five criteria that remain relevant regardless of the channel or solution category:
- Governed knowledge foundation. The platform must be based on a governed, verifiable knowledge layer. It should not simply connect the model to existing content in its current state.
- Control over hallucinations and accuracy. The ability to track exactly where each response comes from, and a mechanism that prevents outdated or contradictory content from making it into the final response to the customer.
- Omnichannel delivery from a single knowledge layer. Chat, voice, and agent assist should all draw from the same governed source. These should not be three separate, independently maintained databases that inevitably diverge over time.
- Integration without a complete migration from scratch. A realistic implementation path does not require halting current operations or migrating all content to a new system from scratch.
- Observability is the foundation of AI responses. The ability to see the exact source of each response is critical for both quality control and subsequent compliance audits.
Shelf is built precisely on these criteria. It’s not just another chatbot, but an AI customer service platform built on a foundation of managed knowledge and data. This doesn’t mean that alternative approaches have no right to exist. A team just starting on its journey into AI for customer experience can begin with a simpler solution. But it’s worth checking the criteria listed above when selecting any vendor, including those who may ultimately turn out to be Shelf’s competitors. Learn more about Shelf’s approach to a governed knowledge foundation.
In practice, these five criteria are rarely verified during a typical vendor demo. Demos usually showcase how quickly and naturally a response sounds, rather than where exactly that response comes from. Teams evaluating AI for customer service vendors should deliberately ask for this level of detail: request that the vendor show the trace of a specific response back to the source document, rather than just the final result in the interface. This is especially important for teams building an AI for customer experience strategy for several years ahead, rather than just for the current budget cycle.
Common Concerns and How to Address Them
Accuracy and hallucinations. This is the most common concern, and it is justified only to the extent that governance is lacking. It is addressed not by a more powerful model, but by a governed knowledge layer where outdated and contradictory content is identified before the model has a chance to rely on it.
Loss of human contact. The governed approach does not replace agents; rather, it relieves them of routine tasks, leaving more time for complex, emotionally charged interactions where human involvement is critical. That said, it’s also important to remember that properly configured agents can handle such conversations correctly by relying on clear rules and guardrails, rather than just template responses.
Security and compliance. Governed platforms ensure access control and traceability of the source of each response. Compliance with content requirements is also important, and this is a matter of platform architecture.
Integration efforts. A realistic implementation path does not require a “remove-and-replace” of the entire current infrastructure. The governed knowledge layer is integrated into existing systems in phases.
Cost. The expenses associated with preparing knowledge before AI begins to deliver value often serve as a reason to postpone the initiative. However, it is not necessary to complete the entire audit and preparation process to derive initial benefits from a limited pilot.
All five concerns have one thing in common: none of them is solved by choosing a more advanced model or a trendier channel. A team that builds AI for customer experience around a governed knowledge layer from the very beginning eliminates most of these risks before they have a chance to materialize into an actual incident with a customer. A team that puts off addressing this issue until the moment when it has to explain to management why the voice AI announced the wrong rate ends up solving the same problem, but at a much higher cost.
How to evaluate and roll out AI customer service
Unfortunately, nothing works perfectly without the right structure. And to achieve success, you must strictly follow the rules, step by step, without skipping a single one:
- Step 1. Conduct an audit of your existing support content for gaps, duplicates, and outdated policies. And you need to do this even before you choose a channel or vendor.
- Step 2. Start with a single controlled channel or use case, where an incorrect response is inexpensive to catch and easy to fix. If you launch all channels at once and an error occurs, it will cost you dearly. And we’re talking not only about reputational damage but also financial losses.
- Step 3. Establish escalation rules and “human-in-the-loop” points. Determine in advance which categories of questions the AI resolves on its own and which are always escalated to an agent.
- Step 4. Establish a baseline for CSAT, FCR, and AHT before launch. This will allow you to demonstrate a specific improvement after implementation, rather than just a general sense of improvement.
- Step 5. Expand your reach channel by channel, building on the same governed knowledge layer. There’s no need to build a separate, isolated knowledge base for each channel. Everything should work as an integrated system.
A common mistake along this path is launching multiple channels simultaneously without a unified, governed foundation underlying them. This creates a risk that a step-by-step approach is designed to prevent. Knowledge desynchronization between channels occurs faster than the team can notice, and the first noticeable errors undermine trust in the initiative as a whole.
It’s telling that teams that successfully scale AI for customer service beyond a single pilot channel almost always follow this path in exactly this sequence. The reverse order – covering as many channels as possible first and establishing governance later – almost never leads to the desired result in practice. This is because rolling back channels that have already been launched is far more difficult, both politically and operationally, than laying the foundation in advance.
Conclusion
A channel is a visible choice: chat, voice, agent assist, or agentic AI. The knowledge behind that channel is the real solution that determines whether a customer receives an accurate answer or a confidently phrased error message.
AI for customer service will be only as reliable as the governed knowledge layer underlying it. Organizations that first build this foundation and then select a channel and vendor achieve predictable results. Organizations that do it the other way around end up with hastily generated “hallucinations” across their entire customer base.
This holds regardless of the specific scenario with which the implementation of AI for customer service begins (whether it’s a chatbot, agent assist, or full-fledged agentic AI). The channel determines the form of interaction with the customer. The knowledge base underlying it determines whether that interaction will be helpful.
If you want to assess how ready your current knowledge base is to work with AI across any support channel – talk to a Shelf expert about how to build a governed foundation tailored to your customer service strategy.
FAQ
AI customer service is the use of artificial intelligence (chatbots, voice AI, agent assist, or agentic AI) to handle customer inquiries and support agents. It does not replace the support team but expands its capabilities by relying on a governed knowledge layer as the foundation for accuracy.
A chatbot interacts directly with the customer and can resolve an issue without human intervention. Agent assist works differently: it suggests responses and relevant policies to a live agent in real time, but the final decision and communication with the customer remains the responsibility of the human agent.
No. Customer service AI enhances the support team’s capabilities by relieving agents of the routine burden of searching for information, but decisions regarding complex, emotionally charged, or non-standard cases still require human involvement.
Accuracy depends directly on the quality of the knowledge base underlying the system, not just on the model itself. Governed platforms, which control for duplicates and outdated content and provide explicit traceability of the answer’s source, demonstrate significantly more consistent accuracy than tools connected to an unmanaged knowledge base.
Security depends on the architecture of the specific platform, not on the concept of AI in customer support itself. Governed solutions provide access control, traceability of the source of each response, and compliance with regulatory requirements at the content level. These are the factors to consider when evaluating any customer service AI solution, regardless of the company’s size.