DashboardSign inStart your trial

Product

Which AI SDR Sends Need Human Approval, and Which Don't

Gmail starts enforcing at a 0.3% spam rate and wants you under 0.1%. That ceiling decides which AI SDR sends get gated and which ones you let run.

7 min read

The argument about AI SDR approval workflows usually gets framed as a trust question, as though the decision were about how much you believe the model. It isn’t, and treating it that way produces the two failure modes you see most often: teams that gate every send and abandon the queue within a month, and teams that gate nothing and find out through Postmaster Tools.

The useful framing is a budget. Google defines bulk senders as anyone sending more than 5,000 messages a day to Gmail accounts, and its guidance for them is direct: keep spam rates below 0.10% and avoid ever reaching 0.30%. The 0.3% figure is where enforcement starts, not where you want to live. Yahoo publishes the same 0.3% ceiling. That number is the constraint the whole approval question sits inside, because a domain that crosses it stops delivering regardless of how good the copy was.

At 5,000 sends a day, 0.1% is five complaints. You do not have room for a category of message that reliably annoys people, and you do not have the reviewer hours to read all 5,000. So the question is which fraction of sends carries most of the complaint risk, and how to route that fraction to a person without routing the rest.

Three properties that predict a complaint

Sends that generate complaints tend to share at least one of three properties, and they’re all checkable before the message leaves.

The send is new. A template that has gone out 400 times with no complaints has told you something. Its 401st send is boring in the way you want. A template on its first run, a new segment, a new sending domain, or a message type the workflow has never produced before has no history behind it, and the cost of finding out through recipients is a permanent mark on a domain you paid to warm up.

Personalization comes from a source you don’t control. This is the property behind the worst AI SDR failures. When a tool scrapes a prospect’s most recent post and works it into the opening line, the tone of that message is decided by whatever the prospect happened to post, and the model has no way to tell a product launch from a death in the family. That’s precisely how an AI SDR pitched a lead-gen tool off a post about losing a friend, a message that reached roughly eleven thousand people as a screenshot. Any send whose personalization pulls from an external feed rather than from your CRM should be treated as high-variance by default.

The message makes a claim. Pricing, timelines, integrations, availability, or anything comparative about a competitor. These are the sends where being wrong costs more than being ignored, because a prospect who acts on an invented price is a support problem and possibly a legal one. A model that has been handed your positioning doc will produce confident sentences about all of it, and confidence is not the signal you want here.

Sends with none of those three properties can run. A follow-up on an existing thread, using a template with history, personalizing from CRM fields you own, making no claim beyond what’s on your pricing page, does not need a human standing behind it. Gating it teaches your reviewers that the queue is full of things they don’t need to read, and a reviewer who has learned that will miss the one that matters.

The counterargument is right about queues

There’s a real objection to all of this, and it comes from people who do RevOps for a living rather than from vendors. Common Room’s RevOps playbook makes the case that centralized approval is the wrong instrument: requests arrive faster than anyone can process them, people build spreadsheets and side channels to route around the bottleneck, and the official system ends up describing a process nobody follows. Their prescription is to build the system so the right decisions are obvious and the wrong ones are hard to make by accident, with autonomy graduated by permission rather than gated by a queue.

That’s correct as a critique of gate-everything, and it’s the reason the three properties above exist as a filter rather than a policy of universal review. Where it needs a qualifier is the class of action that can’t be un-done by better permissioning. You can scope what an SDR manager is allowed to configure. You cannot scope a message that has already landed in a prospect’s inbox, and the domain reputation damage from a bad batch outlives whatever permission model produced it. Permissioning governs configuration; a gate governs the individual outbound action that leaves your company. Most stacks need both, doing different jobs.

The practical synthesis is that autonomy should be set per segment, which is what most serious AI SDR tooling already supports. Long-tail accounts running proven templates go out unattended. Named accounts, or anything where a single bad message costs a relationship you spent quarters building, wait for a person. Reply.io ships this as a mode on its Jason AI SDR, where co-pilot mode drafts and holds for review while automatic mode sends, and its own AI policy is blunt that choosing automatic doesn’t move responsibility for the content off the customer. That last part is worth reading twice if you’re the person whose name is on the sending domain.

Where the queue lives decides whether it gets read

A review step that requires opening a laptop, finding the right tab, and reading a table of pending sends is a review step that gets done for two weeks. The failure is behavioral rather than technical, and it shows up as batch approval: forty pending messages, one click, nobody read any of them. At that point the gate is worse than not having one, because the audit trail says a human approved each of these and that record is now false.

Keeping the reviewer’s decision small is the whole design problem. They need the recipient, the account, the actual draft text, and whatever signal triggered the send, in one screen, on the device they already have in their hand. Anything more and the review becomes a task they schedule; anything less and they’re approving blind. We’ve written separately about why approvals belong on a phone, and the argument holds harder for outbound than for anything else, because outbound queues fill up during the hours nobody is at a desk.

Making the gated fraction shrink

The three properties are a starting sort, not a permanent configuration. A template that was new in March has history by June, and continuing to route it to a human is just the earlier decision left running.

This is what confidence scoring is for. Each proposed send gets scored against what happened to similar sends before it, and sends above your threshold clear on their own while the uncertain ones queue. The effect over a quarter is that the reviewer’s queue narrows toward the messages that are genuinely unusual, which is where their attention was worth spending in the first place. The sort improves as the record accumulates, which is the opposite of how a static approval rule behaves.

What you want at the end of it is a small queue of odd sends, a large volume of proven ones going out untouched, and a decision record that shows who approved what when someone eventually asks. The ten-action version of the gating test covers the same reasoning for workflows that aren’t outbound.

Approvals on Rills are always free, so nothing about your bill pushes you to gate less than you should. Try a demo and hold a send before it goes out.

Common questions

Should AI SDR emails be reviewed before sending?

Some of them. Reviewing every send recreates the manual work the tool was bought to remove, and reviewing none of them puts your sending domain at risk. The workable split is to gate sends that are new, that pull personalization from an uncontrolled source, or that make a claim about pricing, timelines, or a competitor, and let repetitive sends of a proven template go out unattended.

What is approval mode in an AI SDR?

It's a setting where AI-generated messages queue for a human to review, edit, reject, or approve before they reach the recipient. Reply.io's Jason AI calls this co-pilot mode. Most vendors recommend starting there and loosening later, and their terms are usually explicit that running in automatic mode does not move responsibility for the content off you.

Who owns AI SDR guardrails, RevOps or sales?

RevOps generally owns the guardrails themselves, meaning the prompts, the policies, and where approvals sit, while sales leadership sets how much autonomy each segment gets. The split matters because the two decisions have different failure modes: a bad guardrail leaks a message, a bad autonomy setting buries a manager in a queue nobody reads.

Does an approval queue slow down outbound?

It does if you gate everything. A queue that grows faster than anyone can work it gets cleared in bulk without reading, which is oversight on paper and nothing in practice. Keeping the queue small enough to actually read is the constraint that should drive what gets gated.

Ready to automate your workflows?

AI proposes the action, you approve it, and the record shows who signed off.

14-DAY TRIAL · NO CREDIT CARD · APPROVALS ARE FREE