Personalisation with an agent in the loop
A wrong service answer reaches one customer and gets corrected in the next turn. A wrong campaign reaches the whole segment, and there is no next turn.
Marketing & EngagementA service agent gives a wrong answer. One customer reads it, says that does not sound right, and the next turn of the conversation corrects the record. The cost is a moment of confusion and a slightly longer chat.
A campaign goes out carrying generated content that is wrong. The whole segment reads it, none of them asked a question, and there is no next turn. The correction is a second send, which is itself a campaign, and it arrives in the same inbox as the mistake.
That asymmetry is the design problem. Not model quality, which improves on somebody else's schedule and is not yours to control. Distribution, which is entirely yours.
The difference is distribution, not accuracy
Two questions get conflated whenever generated content is proposed for marketing. Whether the model is good enough, and what happens when it is wrong. Only the second is an architecture question, and it is the one that gets less attention in the room.
A conversational agent operates inside a correction loop. The customer is present, engaged and able to push back, and the system gets another turn to repair the answer. Marketing content has none of that. It is composed once, approved once if at all, and pushed to people who were doing something else at the time. The recipient has no way to say that looks wrong to you, and by the time one of them does, everybody has already received it.
This changes what review is for. In service, review is a sampling exercise after the fact: read a share of conversations, find patterns, improve grounding, and the next conversation is better. In marketing the same after-the-fact sample is a post mortem, because the content has already reached its audience. The gate has to sit before send, or it is not a gate.
Which does not mean every piece of generated content needs a person in front of it. It means the gate goes where the content cannot be taken back, and that is a question to answer per channel rather than a policy to set for the whole programme.
It is a Data 360 problem before it is a content problem
The first thing that goes wrong in personalisation programmes is not the content. It is that the content is correct and addressed to a version of the customer who no longer exists.
Personalisation is a claim about a person. Every generated line that says anything specific, an interest, a recent purchase, a stage in a journey, a product they already own, is a claim sourced from data. The generation step sits downstream of a resolved profile and will faithfully render whatever that profile says. Confidently addressing someone as a prospect three weeks after they became a customer is not a content failure, and no amount of review will catch it, because the copy reads perfectly.
Data 360, the current name for what many teams still call Data Cloud, is where this is decided. The unification questions are the same ones an agent faces, and we set them out in full in how data becomes grounding. Two of them behave differently in a marketing context.
Identity resolution errors distribute. A false merge in a service conversation produces one wrong answer to one customer, discovered in a conversation and correctable in the same conversation. The same false merge in a campaign puts one person's purchase history into another person's inbox, in writing, with a timestamp. It is the same defect with a very different blast radius, and the second version is a privacy incident rather than a data quality ticket.
Freshness stops being a technical preference. A profile refreshed nightly personalises on yesterday, which is harmless for a monthly newsletter and actively damaging for a message triggered by a purchase, a cancellation or a complaint. The uncomfortable version of this is that the moments most worth personalising are exactly the ones with the least tolerance for staleness.
The practical sequence follows from that. Resolution quality and a refresh cadence agreed per use case come first, and the content layer comes second. Programmes that run it the other way round produce impressive demonstrations on clean sample data and stall the week they meet the real profile estate.
Reversibility, channel by channel
The useful way to place the gate is to ask one question per channel: after this goes out, how long does it take to change it, and how many people have seen it by then.
| Channel | Reversible after send | Practical correction window | Gate required |
|---|---|---|---|
| No | None. A correction is a second send | Human approval before dispatch | |
| SMS and push | No | None, and each message carries a cost | Human approval before dispatch |
| Web and app content | Yes | Minutes, for everyone not yet served | Enforced rules live, sampled review after |
| In-session recommendations | Yes | Immediate, per session | Enforced rules live, sampled review after |
| Direct mail | No | Weeks, and the artefact stays in the house | Human approval, plus a proof step |
| Ad creative | Partly | Hours, subject to platform review | Human approval on first publish |
That table is the argument in one place. Where the correction window is zero, a person signs. Where content can be changed before most of the audience has seen the next impression, rules can run unattended and review becomes a sample taken afterwards.
Notice that the ordering has nothing to do with how sophisticated the personalisation is. An on-site recommendation module doing something genuinely clever with real-time behaviour needs a lighter gate than a plain templated email, because one of them can be switched off in a minute and the other cannot be unsent. Teams routinely get this backwards, applying the heaviest governance to the newest technique rather than to the least forgiving channel.
Where the approval gate should sit
Given a channel that needs a human gate, the next question is what the human approves, and this is where most programmes design themselves into something unworkable.
The instinct is per-message approval: every generated variant read before it goes. That is defensible for a handful of variants and collapses the moment personalisation does what it was bought to do. A campaign with meaningful variation produces more assets than any review team can read, and what happens next is predictable. Review becomes a formality, someone approves in batches without reading, and the organisation now holds an approval record documenting an inspection that did not occur. That is worse than no gate, because it converts an unmanaged risk into a documented control that does not work.
The alternative is to approve the pattern rather than the instance. A person reviews the message structure, the claims it is permitted to make, the offer, the data it may reference and the boundaries of the variation. The generated instances then have to sit inside an approved envelope, and what the reviewer signs is the envelope.
That moves the review effort to where it is finite. The number of message patterns in a campaign is small and knowable. The number of instances is neither. It also changes the reviewer's question from does this one read well, which is answerable and not very useful at scale, to could anything inside this envelope be wrong, which is the question that actually protects you.
Two cases still warrant per-instance sign-off, and they are worth naming so the exception stays an exception: anything making a regulated claim, and anything going to a segment where a single recipient can escalate it publicly. In both, the asset count is usually small enough that per-instance review is affordable, which is generally true wherever it is genuinely required.
Guardrails belong in rules, not in reviewers
A reviewer catches what they happen to notice. A rule catches what it was written for, every time, including on the send that goes out at six on a Friday when the reviewer is a person who has already read forty variants.
Brand and compliance constraints belong in the second category. Stated as a general principle rather than as a claim about any particular product feature: an approval process that depends on a human remembering a constraint fails at the rate humans forget things, which is not zero and rises with volume.
Three categories are worth expressing as enforced rules rather than as review criteria.
Claims that require substantiation. Pricing, availability, performance, comparative statements, anything with a regulatory footing in your industry. Constrain these at the source. If a message may only reference approved claim text, the variation cannot invent a new claim, because it has nothing to invent from.
Fields that may not be referenced. Not every accurate fact about a customer is one they want repeated back to them. Health-adjacent inferences, financial difficulty signals, anything derived rather than declared. The control belongs at the data layer, where the marketing use case simply does not get access to the field. That is a permissions question and it sits with the rest of your access model, which we set out in the security model behind an agent.
Suppression and consent. The rules about who must not receive a message are older than any of this and they do not change because content generation entered the picture. What does change is that a personalisation layer introduces fresh ways to accidentally re-include somebody, which makes suppression worth re-testing rather than assuming.
The distinction worth holding on to is that rules constrain what is possible and review inspects what happened. Anything you would be embarrassed to explain to a regulator belongs in the first category. Agentforce Marketing, the umbrella brand for what was Salesforce Marketing Cloud, spans the current Marketing Cloud Next platform and the earlier Marketing Cloud Engagement, and the journey and automation mechanics that carry these rules are the ones your team already operates. The operational discipline behind them is in marketing automation that survives contact with the org, and Salesforce sets out the platform view in its marketing automation guide.
What the reviewer is actually for
If rules do the constraining, the human is not there to check spelling.
The reviewer's job is the judgement nobody has written down: whether this message is appropriate right now, given something happening in the world, in the market or to this customer that no rule anticipated. A perfectly compliant, on-brand, accurately personalised message can still be the wrong message to send on a particular morning, and no rule set contains that.
The reframing matters for staffing, because it is a different skill. Reading forty variants for tone is a task that degrades with volume and produces rubber stamps. Deciding whether a campaign should go at all, and whether the envelope it sits in is still right, is a task that improves with seniority and does not scale with the number of variants.
Measuring personalisation rather than segment differences
The common measurement failure is easy to spot once it has been named. A personalised campaign outperforms a generic one, the team concludes that personalisation worked, and what actually happened is that the personalised version went to a segment that was going to respond better anyway, because targeting and personalisation were built at the same time and changed together.
Separating them means holding the audience constant and varying only the content. The same segment, the same send window, personalised content against the control, and the difference attributed to content rather than to who received it. Anything less is a comparison between two audiences wearing the language of a content test.
Three further measures are worth having in place before anything generated goes out at volume.
Complaint and unsubscribe rate at variant level, not just campaign level. A healthy aggregate can hide one variant performing badly enough to be a problem, and campaign-level reporting will never surface it.
Rule interventions. How often a guardrail blocked or altered something. Zero is not a clean bill of health, it usually means the rules are not reaching the content. A steady non-zero rate is what a working control looks like.
Correction sends. Count them, and treat each one as an incident with a cause rather than as a nuisance. The rate at which you have to apologise is the honest measure of whether the gate is in the right place.
What to build first
The sequence that works is roughly the inverse of the one that feels exciting.
Start on a reversible channel. On-site or in-app content lets you run generated personalisation under enforced rules against real traffic with a correction window measured in minutes. Everything you learn there, where the rules are too loose, where the data is thinner than anyone assumed, where the output is subtly off in a way nobody predicted, is learned at a survivable cost.
Move to email second, with pattern-level approval, once the envelope has been written down and tested somewhere it could be adjusted. The temptation is to start with email because it is the channel with the most mature programme and the clearest numbers. Start there last, precisely because it is the channel that cannot be taken back.
Underneath both, the data work is the long pole. Agentforce capability is configured in days. Getting resolution clean enough that a message addressed to a person reaches the right person, with a refresh cadence matched to the trigger that fires it, is measured in months and is the actual project.
The teams who ship personalisation that lasts are not the ones with the best content. They are the ones who worked out, before anything went out, which of their channels could forgive them.



