Skip to content
@ Retention Marketers
Explainer

Do AI agents improve customer retention?

Trials show that AI agents can change engagement and service speed. Neither result alone proves customers stay longer. Here is what the evidence supports and how to test retention lift.

5 min read Published Updated How this was sourced

AI agents can improve parts of the customer experience, but the available studies here do not establish a general improvement in customer retention. One agentic marketing trial found more days of app engagement than the existing messaging process. A separate service trial found shorter chats but lower ratings for the chats eligible for AI handling. Neither result tells us how many additional customers later bought again or renewed because of the agent.

AI agents and retention, in short. Judge an agent by the customer outcome it is meant to change. If the goal is repeat purchasing or renewal, compare the agent-enabled process with the existing process using a randomized customer holdout and a follow-up window chosen in advance. Track service failures and handoffs alongside that outcome. Faster chats, more clicks and more automated decisions answer narrower questions.

What kind of AI agent are we talking about?

An agent takes a sequence of decisions or carries out a task, rather than supplying one prediction. In a marketing setting, it might choose a message, channel and send time from options the team has approved. In service, it might interpret a customer’s issue, ask follow-up questions and carry out a standard resolution, with a person available to take over.

Those are different interventions. The agentic marketing study examined a system that assembled messages from components supplied by marketers and chose timing and channel based on customer context. The Alibaba service study examined an autonomous system that handled a limited set of standardized chats under human supervision. Neither is a test of every product sold as an “AI agent,” and the marketing study does not show that a language model wrote every message.

What does the evidence show?

The available trials show changes in app engagement and service speed. They do not establish that customers later bought again or renewed because of an agent. The clearest distinction is between an agent changing an intermediate measure and an agent retaining a customer.

StudyComparison and observed resultWhat it leaves open
Agentic marketing at one delivery appAbout 8.8 million inactive users were assigned to agentic or usual messaging for 11 months. During the seven months without new manual content or optimization, the authors reported 2.4% (±0.21) more days with engagement per user for the agent group. Their reported relative lift for days with intent and conversion events was 0.2% (±0.27). The ± figures are the paper’s uncertainty estimates.The paper reports activity and conversion days, not a separate customer-level repeat-purchase or renewal rate. Its authors work for the agent platform.
Autonomous service at AlibabaIn an August 2024 randomized deployment involving 647 workers and 680,676 chats, average chat duration fell 3.2% across all chats. Changes in seven-day same-issue repeat contacts and ratings across all chats were not statistically significant.The experiment measured service outcomes over a short period, not later purchasing or renewal. Less than 10% of chats were eligible for AI handling. One author works for Alibaba, which ran the experiment on its platform.

In the marketing study, the control group continued to receive the app’s usual messages. For four months, marketers actively supplied content and optimized the agent; for the next seven, the agent operated without new manual content or optimization. The engagement result during that latter phase is relevant to reactivation. The much smaller reported estimate for intent and conversion days, relative to its stated uncertainty, does not establish a customer-retention lift. The vendor affiliation and single anonymous app also limit how far to carry the result.

The Alibaba study shows why service speed needs its own quality check. For chats eligible for AI handling, average duration fell 16.8%, but customer ratings fell 0.412 points on a five-point scale; the change in seven-day same-issue repeat contacts was not statistically significant. Only 17% of all chats received a rating, so ratings describe the responding subset. A seven-day contact about the same problem is a service measure, not a measure of whether that customer remained a buyer.

The researchers also compared escalation types. In their matched, non-randomized subgroup comparison, chats escalated after the system detected customer frustration had a six-percentage-point higher seven-day same-issue repeat-contact rate and a 0.928-point lower rating than comparable fully human-handled chats. That pattern makes delayed recovery worth monitoring; it does not prove that escalation timing alone caused the difference.

How would you test whether an agent retains customers?

First, specify the action the agent will change. Is it choosing reactivation messages, handling routine service issues, or both? Each creates a different comparison. Then keep an eligible customer group on the current process while another group receives the agent-enabled process. Random assignment and a persistent holdout follow the design used in the marketing study; the comparison is agent versus your existing process, rather than agent versus no contact.

Choose the customer outcome before launch: for example, repeat purchasing or renewal over a window that fits the buying cycle. Record the result for everyone assigned to each group, including people the agent never contacts or whose chat goes to a human. Otherwise, the analysis can select the easy cases the agent chose to handle.

QuestionMeasureWhat it can establish
Is the agent doing work?Decisions delivered, chats handled and human handoffsAdoption and operating load
Is the immediate experience improving?Engagement, chat duration, same-issue repeat contacts, complaints and ratingsResponse and service quality within the measured window
Are more customers staying?Repeat purchasing or renewal in the assigned groups over the chosen windowIncremental retention for this agent policy and customer group

The second row is a safeguard for the third. If a service agent reduces handling time while customers return with the same unresolved issue, that is a reason to inspect the workflow even before a retention result is available. If engagement rises but the purchase difference is too uncertain, report the engagement gain and leave the retention claim open.

An AI agent can make a retention program more responsive or a service team faster. Whether it retains more customers is a separate, testable claim about what customers do afterward.

Sources

  1. Jeunen, Hanna and Wheeler — Sustained Impact of Agentic Personalisation in Marketing · 9 Apr 2026
  2. Wang et al. — Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba's Customer Service Operations · 1 Jun 2026
Keep reading