The question of whether AI may send replies independently is often posed as a matter of principle. It isn’t one. It’s a configuration question, and the answer differs per type of message.
Three levels, not an on-off switch
In practice, three modes work side by side:
- Assistant. AI writes a draft, the employee reviews and sends. Always a human in between.
- Suggestion with a threshold. AI sends by itself, but only if confidence is above a limit and the topic is on the approved list.
- Fully automatic. For a small, sharply defined set of questions where the answer is objectively fixed.
The mistake organisations make is choosing between all or nothing. The gain lies in the distinction: let the category determine which mode applies.
Where automatic sending is defensible
Two conditions. First: the answer follows from data, not from a judgement — where is my parcel, what’s the status of my return, what are your opening hours. Second: the damage from a mistake is limited and repairable.
With a wrong delivery date, you send a correction. With wrong advice about a medical device or an unwarranted promise about a refund, it’s different.
Where a human needs to look
- Complaints and emotion. Not because AI can’t phrase it, but because the customer wants someone to have read it.
- Everything involving money. Goodwill, discounts, credits.
- Exceptions. Precisely where the pattern the model relies on is missing.
- Legal and safety. Speaks for itself.
The confidence score: useful, but no guarantee
A confidence percentage indicates how well the question matches what the system knows. That’s useful for setting a threshold, but it’s not a truth meter. A model can give the wrong answer with high confidence if it’s working from outdated information.
So treat the score as a first sieve, not the only one. Combine it with category and with checks afterwards.
Spot checks are not optional
Anyone who lets AI send automatically without watching only finds out something went wrong when a customer complains. Read back a weekly sample of what was sent automatically. Not out of distrust, but because that’s how you discover which categories you can add — and which need to go back.
What to define up front
- Which categories may go automatic, and who decides that
- At which confidence threshold
- Who does the spot check and how often
- What happens in case of doubt: better to pass to a human than to guess
- Whether the customer can see that a reply was created automatically
That last one is a choice with consequences. Transparency builds trust, but can also create the impression that the customer isn’t being taken seriously. There’s no universally right answer; there is an answer that fits your customers.
Grow step by step
Start with assisting, measure how much the employee still adjusts, and only expand where that percentage is low. That figure is the best gauge there is: when hardly anything is being adjusted any more, that category is ripe for automatic sending.
That’s how we’ve set up human control: you determine per category what happens. Get in touch if you want to spar about what fits your organisation.
