Skip to content
@ Retention Marketers
Explainer

What does AI actually do in email marketing?

AI gets applied at six points in an email programme, and the vendor documentation is precise about the mechanical ones and silent about the decisions. What the help pages promise, what they refuse, and where the automation stops.

14 min read Published Updated How this was sourced →

Sending an email to a list resolves into a stack of separable decisions. Six of them are where the software clusters. What to send. Who gets it. What it says. What it looks like. When it goes. What happens next.

“AI email marketing” is a label stretched across all six. It is worth pulling apart, because they are not equally automated.

Vendor help documentation is precise, quantified and occasionally blunt about the mechanical jobs. On decisions it goes quiet. That pattern holds across platforms, which makes it useful when you are evaluating any of them.

Which jobs does the automation actually reach?

The jobWhat is documentedHow specific
When it goesSend-time models and testsExact thresholds, published windows
What it saysSubject-line and body generationDocumented down to which inputs each tool reads
What it looks likeBrand asset import, section generationA daily cap, unsupported blocks, elements left unfinished
Who gets itPlain-language segment buildingAvailable; accuracy handed back to you
What happens nextFlow generation from a descriptionTwo vendors, two different boundaries
What to sendNothing

The last row is the finding. Across these five platforms’ documentation, no feature decides what a campaign should be about.

Where does the automation stop inside a flow?

Brevo draws the line explicitly rather than leaving it to be inferred.

Aura generates an automation from a description, and the documentation lists what it configures for you: “Wait and delay steps: duration and unit set automatically from your described timing” · “Trigger and entry settings” · “Step labels” · “A/B split percentages” · “Update contact attribute steps” · “Add to list and remove from list steps” · “Webhook steps” · “SMS steps”.

Then it names the exclusions:

“Email step content and segmentation filters are not yet generated automatically. After Aura creates your workflow, you need to add your email templates and configure any conditions or filters manually before activating.” — Brevo Help Center, Create an automation with Aura

Read the two lists against each other. Everything in the first is plumbing: how long to wait, what to call a step, how to split a test, which webhook to fire.

Both exclusions are judgment. What the email says, and which customers should receive it.

That is not a criticism. Wiring a flow by hand is tedious and worth automating. But it locates the boundary, and the boundary sits earlier than generating an automation from a description sounds like it should.

A second vendor draws it somewhere else

The boundary is not standard. One platform generates exactly what Brevo excludes. Klaviyo’s Composer writes flow message content: “Composer will first reflect the structure of the flow that you are requesting and then generate the content for each of the messages within that new flow.” It builds audiences too — “Create a segment by describing the audience you want in plain language.”

So one vendor calls email content and segmentation the manual part, and another generates both. That split is worth carrying into any evaluation, because the phrase AI-built automations does not tell you which half of the work arrives finished.

Where the second vendor does stop

Composer publishes its own boundary, and it sits at execution rather than at content. It “recommends, drafts, audits, and QAs”, and “does not send, publish, schedule, or change your account settings on its own”, “does not apply fixes from an audit for you”, and “does not replace your judgment”. The drafting is automated. Approving it is not.

Omnisend’s AI documentation lists nine features across content, forms, segments and analysis, and no flow-generation feature appears among them.

What do the design tools generate?

Design AI is documented mostly as asset application rather than composition.

Omnisend’s brand tooling “applies your brand logo, fonts, and color palette automatically to email templates”, and needs a connected store to pull them. Of its brand kit, Mailchimp says “We’ll use AI and the assets in your brand kit to quickly generate stylish, on-brand graphics and layouts for your emails, social posts, automation flows, and more.” Applying that kit automatically to template previews is the part gated to a paid tier — “available for accounts with an Essentials plan or higher.”

The limits are the informative part

Klaviyo’s Email AI needs “a paid Klaviyo account” and generates a section from a description — its own example is “Create a sale reminder with urgency to order before midnight, and include an image at the top and a button right below”.

Then it states what will not arrive finished: “Certain elements (like product blocks, images, links, and buttons) are not fully configured by the generation tool”. Other block types, including tables, are “not supported.” Usage caps at 99 generated sections a day.

So the layout arrives and its commercially important parts — the product block, the link, the button — still need configuring. The tool covers the mechanics and stops at the part carrying the offer.

When it goes: thresholds you can check before buying

Here the documentation gives figures rather than descriptions.

Klaviyo’s Smart Send Time publishes its method instead of keeping it opaque: “Instead of determining your business’ send time through hidden formulas, Klaviyo uses a robust testing framework”.

It publishes a floor — “you must be able to send campaigns to an audience of 12,000 people or greater” — plus a table of how many test campaigns each list size needs. The same vendor separately ships a per-profile model gated on plan tier rather than volume. One category name, two mechanisms.

Five products, five different names

The category shares no vocabulary: Send Time Optimization at Mailchimp, Optimize send time for individual contacts at HubSpot, Send at best time at Brevo, Smart Send Time and Personalized Send Time at Klaviyo. Underneath, each reads different data.

Mailchimp pools behaviour across its whole customer base and says so: “Mailchimp customers send a lot of email to a lot of people, which gives us data about individual engagement patterns” — gated to “the Standard plan or higher.” HubSpot reads each contact’s own last 90 days, a Beta feature requiring a “Marketing Hub Enterprise subscription”. Brevo starts from your account’s own averages and sharpens as you send.

What happens when the data runs out

Three of the four per-contact products publish a fallback. No two are alike.

HubSpot sends contacts “without enough engagement data… at the beginning of the sending range” — a fixed time, not an optimised one. Brevo falls back twice: to your own past campaigns if you have any, and for an account that has never sent, to “the Brevo average”, drawn from “emails sent by other users”. Klaviyo also falls back twice, but stays inside your account — “send-time patterns from similar profiles in your account”, then “what’s generally best for your account.”

Mailchimp declines the job: “If there’s not enough data for Mailchimp to find an optimal time, click Cancel and choose another option.”

Which makes the labels less informative than they look. Brevo’s cold start reaches for the same cross-customer pool Mailchimp draws on — though Mailchimp still needs “enough data from your sent emails” to run.

That matters on a growing list. Every new subscriber arrives with no history, so some share of every send runs on a fallback rather than that person’s behaviour — or, on one platform, not at all.

What has to be true before predictive features work

Predictive scoring — churn risk, lifetime value, expected next order — gates on order history rather than list size.

Klaviyo publishes its gate as four conditions: at least 500 customers who have placed an order, an ecommerce integration or API sending placed-order events, 180 days of order history with orders in the last 30, and some customers with three or more orders. It is precise that the first is not a list-size number: “This does not refer to total profiles, but rather the number of people who have actually made an order with your business.”

Together those describe a business trading six months with repeat purchasing.

Why list size is the wrong number here

A store with 40,000 subscribers and 300 purchasers does not qualify. List size is the obvious number to quote when asking whether a tool suits you. This gate does not read it.

What the scores actually are

Klaviyo states its predictions “work best when averaged over many customers and are not expected to be exact for any single individual”, illustrating with a predicted order count of 1.43 — a figure that cannot describe a person.

Omnisend labels customers “at risk” using a published calendar for stores with “fewer than 100 returning customers or less than 5–7% returning customers”: “Last purchase 90–180 days ago.” That is a date subtraction, and a reasonable one. It is a rule rather than a model, which is worth establishing before assuming a score was trained on anything.

Who gets it, and what it says

Neither publishes a volume threshold, and the gates fall by job rather than by vendor. Omnisend puts its copy tools on “all Omnisend plans (including Free)” and its segment builder behind a connected store; Klaviyo’s needs a paid account. Both jobs hand something back to you.

Klaviyo’s segment builder “accepts natural language inputs… and converts them into a Klaviyo segment.” Its own examples show the scope: “Engaged profiles last 30 days”. These map to conditions that already existed. The tool writes the filter, not the strategy.

It also attaches a caveat: “you are still responsible for the final segment definition.” A tool that turns your sentence into conditions can turn it into the wrong ones, and the result looks equally confident.

Does the copy tool read your data?

The answers differ, and not neatly by vendor. Two subject-line tools are documented as taking a description. But Omnisend’s separate body-copy tool “generates texts in your brand voice based on prompts and past campaign content.”

Two generation features, one help page, different inputs.

Does AI-written copy perform better?

Mostly unmeasured. These vendors’ documentation carries no controlled tests.

One exception is worth reading, because the result does not flatter the vendor who published it. GetResponse compared more than 16,000 emails sent through its platform in 2024:

Metric — 16,000+ GetResponse emails, 2024, observationalGenerated with AIWritten from scratch
Open rate37.37%41.05%
Click-through rate9.44%8.46%
Click-to-open rate25.25%20.62%
Unsubscribe0.16%0.14%

Read whole, it is mixed. AI-generated emails earned more clicks per delivered email and more per open. Fewer opened them. Marginally more unsubscribed.

GetResponse states the caveat itself: treat the data “only as a sample and not as a direct guide to making decisions in your strategy.” It is also observational — users chose which emails to generate — so not a controlled test.

The job the documentation never automates

Set the six side by side and one has no entry.

Deciding what a campaign should be about — which offer, aimed at which moment in a customer’s relationship with you, judged against what the last one did — sits upstream of everything these tools generate. Segment builders translate a description you have already formed. Flow generators, whether or not they write the messages, start from a goal you supplied. Klaviyo’s subject-line assistant asks you to “describe the context”. Composer starts from you “describing your goal in plain language.” In the help documentation, every one takes a decision as its input.

The framing is inverted

None of this is hidden — it is in the exclusion lists above, and in the requirement running through them that you arrive with a prompt. But it inverts the pitch, and the pitch is worth quoting rather than characterising. Omnisend’s AI page: “Omnisend AI writes your copy, suggests subject lines, builds segments, recommends products, and predicts which customers are about to churn. You review, adjust, and send.” Klaviyo’s: “K:AI handles who to reach, what to say, and when to send it—personalizing every message with the right products.”

The documentation says AI formats, times and wires them, then waits for you to say what they are for.

One vendor sells the sixth job and documents the opposite

Klaviyo publishes both halves. Its marketing page promises “AI that creates, resolves, and optimizes— autonomously”, says Composer will “uncover opportunities”, and describes agents that “don’t wait to be told what to do” and work “without prompting.” Its help centre, quoted earlier, has the same product starting from a goal you supply.

Both pages are current and both are the vendor’s own — which is the argument for reading the help centre before the demo rather than after.

Which parts of the work have actually been measured?

Two surveys, six years apart, and the answer is the same both times: everything except the deciding.

One survey put hours on it, in 2017. A sample of 3,500 marketers ranked six email production tasks by average time: graphics and design 4.1, coding and development 3.8, copywriting 3.0, data pulls 2.4, testing 2.3, analytics 2.1. Copywriting — a job this generation of tools targets directly — came third by that measure, behind design and code. A reader poll on the same page, 43 responses, put copywriting first instead; two instruments, two answers, and the smaller one is nowhere near large enough to overturn the larger.

What the analyst publishing the figures noticed is the part that lasted. They saw what the list omitted, and asked in print:

“But do you notice a task that is missing? Where is the planning task? All these tasks are important, but how do you know what to write or design for the email?”

Those hours are from 2017 and should not be read as current. The durable finding is structural: a survey built to time email production timed six tasks and never counted the deciding.

What the newer research measured

Six years later the instrument changed but the omission did not. Litmus’s 2023 State of Email Workflows report, surveying “over 440 email marketers worldwide”, timed the cycle rather than the tasks: “The email production cycle is one week for 21% of email marketing teams. But for 62% of email marketing teams, it takes two weeks or more to build an email.” Weeks per email, not hours per task, so it does not update the 2017 figures.

Its blocker question bears more directly. Ranked, the obstacles are “Building (41%), Designing 40%, and Testing (39%)”, then “Collecting feedback (35%), Content creation (34%), and Getting buy in from all stakeholders (32%).”

Six blockers, and every one is downstream of a decision already taken. Choosing what a campaign is for appears in neither list, just as planning did not appear in 2017.

What that leaves unmeasured

Neither survey counted planning, so the time it takes has no published figure.

What the vendors sell against is the other half — the tasks both surveys did count. Omnisend promises “Write and personalize faster”; Klaviyo offers to “Go from prompt to campaign in minutes” because “What used to take hours can now take minutes.” Writing, designing and building are exactly where those hours were measured. The pitch lands on the half that has numbers.

Which makes it a question about your own programme. The surveys say what the industry counts, not what your week goes on.

Reading any AI claim in this category

Four questions, all answerable from documentation before a trial.

  1. Which of the six jobs does it touch? Timing, copy, design, audience, flow logic, or the decision. The first five have documented features. Claims about the sixth deserve scrutiny.
  2. What does the vendor say it will not do? An exclusion list carries more information than a feature list, and a vendor publishing one is easier to evaluate than one that does not.
  3. What has to be true before it runs? A published threshold — 12,000 recipients, 500 purchasers, a connected store, an Enterprise tier — tells you when a feature works. “Enough data” does not.
  4. Is it a model or a rule? A trained probability and a 90-to-180-day date subtraction can carry the same label.

Where this leaves a smaller programme

Several of these features are gated on volume or plan tier, so a small store will find parts of the category unavailable, and an AI feature list is a weak tiebreaker at that size.

The more useful reading is where your hours go, not which platform to buy. These tools have gone furthest into the jobs that were already mechanical, and stop — by their own documentation — where a judgement starts. If most of a week goes into deciding what the next campaign should be, that is the half the documentation stops short of.

All articles