Monitoring Cold Email Deliverability: Detecting Spam
Why the warmup score is misleading and how to measure real inbox placement: bounce and reply by recipient ESP, ESP matching, thresholds, and an action ladder.
Most “deliverability” dashboards show you a proxy, not the truth. No mail provider tells the sender “this email went to spam.” That’s exactly why good campaigns often die quietly — deliverability tips over, and nobody notices until replies stop coming. This guide is the monitoring perspective: not how to set up deliverability (that’s covered in the deliverability fundamentals article), but how to recognize that it’s tipping — in time.
The 7 things most teams get wrong
- You don’t know whether you’re in the inbox. Bounces, opens, and replies don’t measure placement.
- The warmup score is a vanity metric. 98-100 points while real prospects never reply — the most common false sense of security in cold email.
- The biggest lever isn’t copy or warmup, it’s ESP matching. Gmail→Gmail, Microsoft→Microsoft.
- Aggregated metrics hide the problem. A healthy overall bounce rate can mask a Microsoft collapse.
- Don’t trust your tool’s provider labels. Detect the real ESP from the DNS MX records, not from the label.
- Microsoft is a different, harder game — lower volume ceilings, stricter filters, slowest recovery.
- Monitor daily and act automatically on thresholds. A Monday problem becomes a burned domain by Friday if a human has to notice it first.
”Are we in spam?” — harder to answer than you’d think
The uncomfortable truth: bounces, opens, and replies say nothing about placement. A bounce is a rejection at the door, not spam foldering. Opens are unreliable (no-pixel best practice + Apple Mail Privacy). What’s left:
| Signal | What it proves | Detects the spam folder? |
|---|---|---|
| Bounce rate (by recipient ESP) | Hard rejection at the door | Strongest early indicator — but rejection, not junk |
| Reply drop-off vs. baseline | People aren’t seeing the email | Best practical inbox proxy for cold email |
| Warmup-pool spam rate | Placement in the warmup network | ❌ Vanity — not your real audience |
| Seed/placement test | Inbox vs. spam via a seed panel | ✅ Direct — but synthetic |
| Google Postmaster | Real Gmail complaint rate + reputation | ✅ Direct — Gmail only |
| Microsoft SNDS | Outlook reputation | ✅ Direct — IP-based, often unavailable on managed mailboxes |
Takeaway: The free signals you already have (bounce + reply drop-off, split by ESP) are an excellent early-warning system. If you want to know rather than infer, you need a placement test or Postmaster.
The warmup-score trap
Warmup scores measure placement within the warmup pool. This number can show 98-100 while your real Microsoft prospects see nothing. Never auto-pause or auto-throttle based on the warmup score alone. Log it, alert on it — but treat it as the vanity metric it is. Trust bounce-by-ESP and reply trends instead.
ESP matching only works if both sides exist
“Google→Google, Microsoft→Microsoft” is the most effective deliverability move there is. But it has a prerequisite teams overlook: you need Microsoft sending mailboxes to route Microsoft recipients to. If your fleet is 100% Google, ESP matching does nothing for your Microsoft segment — there’s nothing to match.
The subtle trap after that: how the mailboxes are connected determines whether the matching actually takes effect. Mailboxes connected via native Microsoft OAuth are recognized as Microsoft. Mailboxes bolted on via generic IMAP/SMTP often land in the “other” bucket — you then own Microsoft mailboxes that are never used as Microsoft. Always check the connection type.
Split everything by recipient ESP
Your sequencer’s analytics are aggregated. Aggregation hides exactly the failure mode that kills German/EU B2B cold email: Gmail placement fine, Microsoft placement collapsing. Build the split yourself — map every recipient domain to its real provider (a static map for the common ones, MX lookup for the rest) and calculate bounce/reply per ESP. This one view turns an invisible problem into an obvious one.
And: classify the ESP from the MX records, not from the tool’s label. We’ve found managed mailboxes whose provider codes didn’t match the vendor’s own documentation. DNS doesn’t lie.
The numbers that matter (threshold cheat sheet)
| Metric | Target | Warning | Act |
|---|---|---|---|
| Inbox placement | >80% | <80% | <60% → investigate |
| Bounce rate (campaign) | ~1% | >1% | >2% → pause |
| Bounce rate (domain) | <2% | >2% | >4% → stop sending, rotate to a healing domain |
| Microsoft segment bounce | <3% | >3% | route to Microsoft senders / pause segment |
| Spam complaints | <0.1% | >0.05% | >0.1% → pause (Google/Yahoo limit: 0.3%) |
| Human reply rate | >10% | <5% | drop >50% vs. average → investigate |
Sending limits per provider (cold/day per inbox)
| Provider | Max cold/day | Note |
|---|---|---|
| Google Workspace | ~15 | Warmup ratio ~1:1.75 |
| Microsoft 365 | ~10 | Strictest filter, slowest recovery |
| Azure tenant (50 mailboxes/tenant) | ~5 max | Exceeding this risks a tenant-wide burn (4-6 weeks recovery) |
Scale via more domains/inboxes, never via more volume per box. Rule of thumb: 8 domains × 5 inboxes × ~20/day, not 2 inboxes × 250/day.
The staged action ladder
Automate the safe, gate the destructive:
- Alert — Slack notice with entity + metric + fix.
- Throttle — automatically lower the daily limit toward policy.
- Pause — automatically stop the campaign/mailbox at danger thresholds.
- Rotate — swap a burned domain for a healing domain — gated by a human.
Run steps 1-3 automatically with a full audit log; require a human for rotation and for resuming anything. Speed where it’s safe — a hand on the wheel where it isn’t. (The same human-in-the-loop principle as in Multi-Channel Outbound.)
The 10-point pre-flight checklist
- SPF: one record, ≤10 lookups, ends in
~all - DKIM signed and passing on every domain
- DMARC:
p=none→quarantine(4-8 weeks) →reject(4+ months) - One-click unsubscribe (RFC 8058)
- Never send cold from the primary company domain
- 2x-domains rule: half sending, half healing, rotate monthly
- ≤15/day Google, ≤10/day Microsoft, per inbox
- Sender ESP mix roughly matches the lead ESP mix
- <100 words, plain text, no tracking pixels, no links in first contact
- Daily monitoring with automatic thresholds — no human staring at dashboards
Treat any provider’s “95% inbox guarantee” as marketing, not a metric. We break down how many emails one meeting ultimately costs here: Cold Emails per B2B Meeting.
CegTec runs cold-email-deliverability monitoring as part of the outbound system — bounce/reply split by ESP, automatic threshold actions with human-in-the-loop for domain rotation. The full threshold config + monitoring blueprint is available as a lead magnet.
Start your free trial · 4 weeks free, no credit card. Prefer to see it running first? Book a demo.
Common questions
How do I know if my cold emails are landing in spam?
No mail provider tells the sender 'that went to spam.' The most reliable freely available early indicators are bounce rate and reply drop-off — each split by recipient provider (ESP). If you want to know for sure, you need a seed/placement test or Google Postmaster (Gmail only).
Is a high warmup score a good sign?
No. The warmup score only measures placement within the warmup network — accounts emailing each other. It can show 98-100 while real Microsoft recipients never see your email. Never throttle or pause based on warmup score alone.
What is ESP matching and why is it the biggest lever?
Gmail senders to Gmail recipients, Microsoft to Microsoft. This match is the single most effective deliverability lever there is. A prerequisite many overlook: you need Microsoft sending mailboxes, otherwise your Microsoft segment runs into a void — there's nothing to match.
Why is Microsoft/Outlook harder than Gmail?
Microsoft filters more strictly, has lower volume ceilings, and the slowest recovery after reputation damage. In EU/DACH B2B, the Microsoft share is high — which is why the Microsoft segment deserves its own track, not a footnote.
How do you scale cold email volume safely?
Via more domains and mailboxes, never via more volume per mailbox. Rule of thumb: 8 domains × 5 inboxes × ~20 emails/day instead of 2 inboxes × 250/day. Per inbox: ≤15/day Google, ≤10/day Microsoft.