One morning Skylar Romines opened a cold email that led with a reaction to her most recent LinkedIn post. The post was about losing her best friend. The email used that grief as the icebreaker for a sales pitch, a lead-gen tool with a free audit attached, and her reaction went out to roughly eleven thousand people: “Of all the LinkedIn and outbound sales ick I’ve seen, this must be the worst. Yikes.” The uncomfortable part for anyone shipping AI outbound is that the tool didn’t malfunction. It did precisely the job it was built to do.
That’s the thing to sit with, because AI SDR mistakes like this one aren’t rare glitches you can patch. They’re the predictable output of a system that generates volume without judgment. The AI was told to react to the most recent post from the prospect and personalized a message around it, as designed. What was missing was the one thing the tool has no way to supply on its own: a person between the draft and the send who could look at that particular email and say no.
The AI SDR Did Exactly What It Was Built To Do
An AI SDR is optimized for three things, and none of them is taste. It’s optimized for relevance, so it scrapes the freshest signal about a prospect. It’s optimized for personalization, so it weaves that signal into the opener. And it’s optimized for throughput, so it does this a thousand times before lunch. Point that machine at someone whose latest post is a funeral notice and it will build you a warm, personal, on-brand-sounding message about a death, because a death is relevant and personal and recent.
Lucas Synnott, who flagged the incident on LinkedIn, put the gap plainly: AI “still lacks the judgment to recognize when something technically relevant is completely inappropriate to mention.” A human SDR writing fifty emails a day catches that in a quarter of a second, before their hand ever reaches the send button. They don’t need a rule that says “no grief in cold emails.” They just know. A model optimizing for a personalized hook has no internal representation of “this is someone’s worst week.” The recent-post signal is usually gold but once in a while it’s a eulogy, and the tool can’t tell the difference.
The Same Blind Spot, Quieter: Your Sending Domain
The grief email is the loud version of this failure, the kind that ends up in a screenshot with thousands of views. The same missing judgment produces a quieter failure you’ll never see on your timeline, and it hits the sending domain your whole company relies on.
Volume without a check erodes your sender reputation across thousands of ordinary-looking messages. Google’s sender guidelines classify anyone sending more than 5,000 messages a day to Gmail as a bulk sender and hold them to a spam complaint rate below 0.30%, which works out to three complaints per thousand delivered. Since May 2025, Outlook has required bulk senders to authenticate with SPF, DKIM, and DMARC and rejects non-compliant mail outright. Yahoo enforces the same 0.30% ceiling Google does. None of those systems reads the quality of any single email. They read the pattern your domain produces, and an AI SDR pushing sends without a check on what leaves the building is the exact pattern they’re built to catch.
So the blind spot cuts two ways from one root cause. The tasteless send costs you a brand blowup in an afternoon. The unreviewed volume costs you a burned domain over a few weeks, and domains are slow and expensive to repair. Both come from the same place: nobody looked before it went out.
Why an Approval Queue Alone Doesn’t Fix This
The obvious response is to put a human in front of every send. That’s the right instinct for the risky messages, but “review every email before it goes” doesn’t survive contact with real volume. Ask someone to approve four hundred outbound drafts a day and by draft forty they’re tapping approve on reflex. A queue you rubber-stamp is review with an extra step and none of the protection.
So blanket approval fails at both ends. Approve nothing and you’re back to a firehose of unread drafts, one of which references a stranger’s grief. Approve everything by hand and you either drown or start rubber-stamping, which means the grief email sails through anyway on a tired afternoon. The setup that holds up puts a human on the drafts that need one, lets the safe drafts through without a tap, and can tell those two groups apart.
Confidence Scoring: Send the Routine, Stop the Risky
This is where confidence scoring does the work an approval queue can’t. Instead of treating every send as equally risky, the system scores each proposed email against the ones you’ve already approved and rejected. A send that matches a pattern you’ve reliably greenlit, a verified recipient at a warmed domain, personalization that references a straightforward business signal, a length and tone you’ve approved before, carries a high confidence score and goes out on its own. Anything outside that pattern drops in confidence and routes to a person first.
The grief email is the archetypal low-confidence send. The source signal is unusual and emotionally loaded, the stakes if it’s wrong are enormous, and the system has never watched you approve anything like it. That’s exactly the kind of draft that should stop and wait for a human, and the kind a flat approve-everything queue lets through. How confidence scoring works covers the mechanism in more depth, but the effect on your outbound is direct: the messages most likely to embarrass you are the ones a person looks at, and the volume that goes out unattended is volume that matched a pattern you already trust.
The effect compounds in the direction you want, too. Early on you review more, because the system hasn’t seen enough of your judgment to predict it. As it accumulates decisions, its confidence on the boring, repeatable sends climbs and those stop coming to you. That’s the shape of the supervised-to-autonomous ladder any AI acting on your behalf should climb. Confidence is scored on every run, so specific patterns clear on their own only as they prove out, and the weird edge cases never do, so the long tail of drafts most likely to torch your reputation keeps landing in front of a human for as long as the workflow runs.
Set the Gate Before You’re the Screenshot
If you’re standing up AI outbound now, you get to build this gate before you need it, which beats the position of the ops lead re-warming a domain or watching their company’s name attached to a viral cautionary tale. Starting supervised isn’t a tax you carry forever. It’s how the system learns where your bar is, so it can start clearing the routine sends itself while the risky ones keep coming to you.
In Rills, approvals are always free. You’re charged only for high-value actions like the AI drafting and the send itself, so putting a human in front of the messages that could damage your brand or your domain never costs you anything beyond the seconds it takes to review one. Try a live demo and swipe through a pending send yourself, or read how action credit pricing works if you want the pricing detail first. Either way the point holds: the sends most likely to make you the next screenshot are the ones that wait for a person.
Common questions
Can an AI SDR send inappropriate messages?
Yes, and it happens when the tool treats any recent signal as fair game for personalization. An AI SDR scores a prospect's latest post as relevant context and works it into a pitch, with no sense of whether that post announced a product launch or the death of a loved one. Without a human check on the risky drafts, it ships whatever it generates.
Should AI send cold emails without human approval?
Not until a pattern has earned it. New or unproven send types should wait for a human, while message types the system has watched you approve many times can start going through on their own. The edge cases, the drafts most likely to embarrass you, should always route to a person.
How do you add guardrails to an AI SDR?
Put a checkpoint between the draft and the send, and make it selective. Authenticate every send with SPF, DKIM, and DMARC, keep your lists clean, and route the drafts most likely to draw a complaint to a person while trusted patterns go out on their own. Confidence scoring is what decides which is which.
Can AI personalization backfire?
Yes. The same feature that makes outbound feel researched can surface something the recipient never wanted referenced, like a personal loss, or expose broken merge fields that read as obviously automated. Personalization pulled from a prospect's most recent post is only ever as appropriate as that post happened to be.
Ready to automate your workflows?
AI proposes the action, you approve it, and the record shows who signed off.