Cold Email Deliverability Forensics in 2026: Diagnosing Spam Placement

Yananai A. ChiwutaPublished ·12 min readUpdated
Cold Email Deliverability Forensics in 2026: Diagnosing Spam Placement

TL;DR

  • First identify the failure: not accepted, accepted but filtered, accepted but not seen, or seen but not answered. A sequencer's “sent” status and an open pixel do not tell you which happened.
  • Freeze the affected cohort before changing DNS, copy, volume and mailbox provider at once. Preserve the message, full headers, SMTP response, recipient domain and sending identity.
  • Check authentication and alignment on the actual received message, then compare provider errors, complaint signals, bounces and controlled placement evidence by domain and mailbox.
  • Make one change, send a small matched test, and wait for provider data to update. If complaints, hard bounces or authentication failures rise, pause that source and fix the cause before resuming.

What are you calling a delivery failure?

“Deliverability” often hides four different outcomes. Keep them separate in the incident record:

  1. Connection or policy rejection: the receiving mail server refuses or defers the message. You should have a delivery-status notification or SMTP response such as a 4xx temporary deferral or 5xx permanent rejection.
  2. Accepted, then filtered: the receiver accepts the SMTP transaction but puts the message in spam, quarantine, another tab or a rule-controlled folder. Acceptance is not inbox placement.
  3. Delivered but not observed: the message reached a folder or mailbox that the recipient did not see. Tracking pixels can be blocked, proxied or preloaded; no open event is not proof of non-delivery.
  4. Seen but no reply: the sender has a message-quality, audience, timing or offer problem. Replacing DNS records will not make an irrelevant message useful.

Google's Postmaster Tools dashboards show selected spam, authentication and delivery-error data for messages to personal Gmail accounts, but the data is not real-time and may be absent at low volumes. This is one receiver's view. It does not report every Google Workspace mailbox, Microsoft 365 tenant, Yahoo inbox or corporate security gateway. Treat “no data” as no usable observation, not a clean bill of health.


The forensic sequence

Follow the sequence in order. Changing DNS, provider, sending rate and copy simultaneously destroys the comparison that would tell you what fixed the issue.

Stage Evidence to save If it fails
Delivery status Message ID, recipient domain, timestamp, SMTP response, bounce class Resolve auth, recipient, policy, connection or rate-limit cause first
Authentication Received headers with SPF, DKIM and DMARC results and domains Correct the real source or alignment; do not add duplicate records blindly
Receiver signal Postmaster/provider dashboard; Google's spam rate is inbox-delivered mail manually marked spam, not all mail landing in spam Isolate the affected domain/IP or campaign; a low complaint rate can coexist with automatic spam placement
Placement Controlled seed account and mailbox-provider cohort Use as a comparison, not proof of universal placement
Engagement Human replies and qualified outcomes by matched cohort Investigate relevance, offer and list selection separately

1. Preserve the evidence and define the cohort

Pick a time window around the decline: for example, the last seven days versus the prior four weeks. Export campaign events before cleaning or restarting anything. For every message, retain the sending domain and mailbox, recipient domain, campaign and variant, send time, accepted/deferred/bounced state, exact SMTP response where available, and subsequent complaint, reply or unsubscribe. Keep the original raw message and full headers from a controlled recipient.

Split results by receiving provider and by sending identity. A 4% bounce rate averaged across four providers can hide one recipient domain rejecting 16% while the other three are normal. Likewise, a single degraded mailbox can be concealed by a campaign-wide average. Group on consistent denominators: attempts, accepted messages, hard bounces, complaints and actual human replies. Exclude duplicates and test addresses consistently in both periods.

Write down the change history: new DNS or tracking records, mailbox password/OAuth reset, new sending provider or IP, campaign ramp, list import, copy/domain change, a change in recipient mix, or a new forwarding/security layer. Record dates and affected identities. Do not infer causation just because the last change happened near the first bad day.


2. Separate rejection from inbox placement

Start with actual message events. “Sent” in the sequencer can mean queued or handed to its provider. Look for the final receiver response or delivery-status notification. Categorise at least:

  • Hard bounce / permanent 5xx: invalid recipient, policy block, unauthorised source or another permanent condition. Save the complete enhanced status code and text; do not retry a bad address blindly.
  • Soft bounce / 4xx: temporary deferral or rate limit. Check whether retries succeed and how quickly. A repeated deferral across one provider may indicate a volume or reputation issue, but read the response before reducing every mailbox.
  • Accepted / no bounce: the receiver accepted the message. This still says nothing definitive about inbox versus spam placement.
  • No final event: investigate logs and provider reporting. A missing event is not evidence of successful delivery.

Do not use open rate as a placement meter. Google's sender guidance says it does not track open rates and cannot verify third-party open accuracy. Privacy proxies and image blocking further weaken the signal. Prefer receiver responses, provider telemetry, controlled mailbox checks and human replies, each with its limitation stated.


3. Verify SPF, DKIM and DMARC on the received message

Check a message delivered to a mailbox you control. Use its original “Show original” or raw-source view. Record the Authentication-Results values and the domains shown for SPF, DKIM (d=) and DMARC. Then compare with DNS. A website-level lookup alone cannot prove that the campaign's actual message used the expected envelope sender or DKIM selector.

SPF authorises sending hosts for a domain used in the SMTP envelope. The SPF standard, RFC 7208, explains the DNS policy and evaluation. Confirm there is one valid SPF policy for the relevant domain and that all legitimate senders are inventoried. Adding a second v=spf1 TXT record does not combine policies; it can cause a permanent error. Also check the DNS lookup limit and any third-party sender the business forgot about.

DKIM signs selected message content with a private key; the receiver retrieves its public key from DNS. RFC 6376 defines the signature and verification. Check that the d= signing domain is the intended one, the selector in s= exists, the key record is complete, and the received message reports dkim=pass. A test message sent by a different application may use another selector and prove nothing about this campaign.

DMARC checks alignment between the visible From: domain and an authenticated SPF or DKIM identity, then publishes the domain owner's policy. The current core specification is RFC 9989, which obsoletes RFC 7489. Aggregate reporting is specified separately in RFC 9990. Record the actual header.from, SPF domain and DKIM d= value. A message can pass SPF for a vendor's bounce domain yet fail DMARC alignment with your visible From domain. Review reports before tightening a policy: first identify every authorised source so legitimate CRM, invoicing and support mail is not accidentally quarantined.

Use the Google Admin Toolbox Check MX as a quick DNS check for MX/SPF and, when you know the selector, DKIM. Cross-check with your DNS host and the received headers. A pass is a configuration observation, not proof of good sender reputation, wanted content or inbox placement.


4. Find the failing provider, domain or mailbox

Make a small incident matrix before taking action:

Segment Attempts Accepted Hard bounce Temporary deferral Complaint signal Human replies
Gmail consumer recipients Postmaster Tools where available
Microsoft-hosted recipients Provider logs / recipient feedback
Other corporate domains Gateway response / admin report
Sending domain or mailbox Provider or campaign event

Do not compare a Gmail-only Postmaster graph to all campaign attempts. Match the denominator and date range as closely as possible. Google's Spam Rate dashboard is not the share of all sent messages automatically delivered to spam. It measures messages delivered to recipients' Gmail inboxes and then manually marked as spam; mail that Gmail automatically sends to spam can make the displayed rate look low because fewer messages reach the inbox where users can mark them. A higher rate is evidence of more user complaints among inbox-delivered mail, while a low rate does not prove good placement. Google's current guidance says keep the user-reported rate below 0.10% and avoid reaching 0.30% or higher. This is Gmail guidance, not a universal safe threshold or an endorsement of unsolicited mail. Any sharp rise is a reason to inspect targeting, recipient expectations, list source and opt-out handling.

Look at the failure's shape. One mailbox failing after a credential reset points first to that connection or identity. One provider deferring while others accept points to provider-specific policy or reputation. All recipients bouncing after a list import suggests data quality. Authentication failure across every provider suggests a sender setup problem. Acceptance across providers with low response but normal human placement checks makes message relevance a stronger hypothesis. These are diagnostic leads, not conclusions until evidence supports them.


5. Run a controlled placement test

Seed tests can tell you how a specific message appears in a set of test accounts at one moment. They cannot guarantee placement in every real recipient's mailbox: seed networks may differ in history, provider mix and user behaviour. Pair them with provider telemetry and actual responses.

Create a small test with one factor changed. Keep recipient mix, sending domain, mailbox, time window, message body and volume as similar as possible. For example, if comparing a new tracking domain, randomise eligible recipients across old and new tracking configuration while holding copy and sender constant. If the change affects all production senders at once, there is no control group; use a low-risk pilot identity and label the result as directional.

Capture the full header, folder observed, provider response, delivery delay, test date and message variant. Repeat on Gmail, Microsoft and another recipient type if those matter to the audience. Do not send repeated tests to the same handful of inboxes at campaign scale; that creates its own traffic pattern. Avoid a long checklist of “spam words” as if one phrase determines placement. Authentication, recipient expectations, complaints, sending behaviour, content and receiver-specific filtering interact.


6. Change one cause and define a safe restart

Use the evidence to choose a narrow fix. Examples: correct a missing DKIM selector; remove a duplicate SPF record; stop a stale or invalid list source; repair a reply/unsubscribe suppression path; reduce a provider-specific burst after repeated deferrals; or revise an offer that recipients are reporting as unwanted. Do not rotate domains or IPs to outrun a reputation problem. That moves the issue and obscures the root cause.

Before resuming, write explicit stop conditions for the pilot: authentication failure, a sudden hard-bounce spike, provider deferrals that persist after retry, any material complaint increase, or suppression not propagating. Pause the affected identity or segment, not automatically every unrelated business message. Have an owner inspect the raw response, repair and re-test. Set a conservative restart volume appropriate to the mailbox and recipient expectations; there is no universally safe “emails per day” number.

Allow monitoring data to catch up. Google says its dashboards are not real time and usually update within 24 hours, sometimes longer. Do not make three successive changes because yesterday's graph has not moved. Record the hypothesis, the one change, the comparison cohort, the observation window and the decision to continue, hold or roll back.

Use the Cold Email Deliverability Checklist to record the sender, authentication and monitoring checks before restarting a campaign.


Worked incident: a Gmail-only decline

The following figures are constructed to show the method; they are not a measured campaign. An agency sees a fall in replies from 4.2% to 1.1% after adding a new tracking subdomain. The sequencer says 98% sent. That timing makes the tracking change a hypothesis, not a demonstrated cause. The team is tempted to replace all sending domains.

First, export equivalent seven-day periods for the current and prior periods and split by recipient provider. Suppose Gmail consumer messages were SMTP-accepted at the same rate as before, while Postmaster shows a higher inbox-delivered, user-reported complaint rate and the other providers show no meaningful change. This supports a rise in complaints among Gmail inbox deliveries; it does not measure the total share landing in spam, and it does not yet prove the tracking domain caused the reply decline.

Next, inspect historical headers from the same sequencer and sending path, recipient provider and sending identity, if retained. If those comparable historical messages are unavailable, say so: current headers can identify a present fault but cannot date its start. Separate what changed in the tracking setup from any changed Return-Path, DKIM selector or visible From domain. Suppose the new sequencer configuration has SPF pass, DKIM pass, but DMARC fail because neither authenticated identity aligns with the visible From domain; this establishes a current authentication fault, but without comparable historical headers it does not establish when the fault began or explain the decline by itself. Correct the misalignment, verify the new message header and keep that cohort paused. If historical messages from that same sending path show the same DMARC failure, do not claim this change caused the decline; keep investigating the tracking-domain hypothesis with a matched test after addressing the authentication issue.

Only after the demonstrated authentication fault is corrected, compare a small cohort with the original tracking setup against a matched cohort using the changed tracking setup, holding copy, recipients, sender and volume as constant as practical. Compare provider responses, headers, placement observations, complaints and human replies over the same period. If both cohorts share corrected authentication and only the tracking configuration differs, the result can inform the original hypothesis; it still cannot prove universal inbox placement. If DMARC was not newly broken, keep the tracking change as an unresolved hypothesis rather than treating it as the fix.


FAQ

If the server accepted my message, did it reach the inbox?

No. SMTP acceptance means the receiving system took the message for processing. It can still classify it as spam, quarantine it, route it elsewhere or place it in a folder the recipient does not see. Use provider telemetry and controlled mailbox checks to investigate placement.

Does SPF pass mean authentication is fixed?

Not by itself. SPF evaluates an envelope sender domain. DMARC also requires alignment with the visible From domain through SPF or DKIM. Check SPF, DKIM and DMARC results and domains in the actual received message, not just a DNS lookup.

How many test emails do I need?

There is no universal number. Use enough observations across the providers that matter to distinguish a repeatable change, while keeping volume low and the test controlled. Sparse provider dashboards or a few seed accounts cannot support precise placement percentages. Record missing data as missing.

Are open rates useful for deliverability diagnosis?

They are weak evidence. Pixels can be blocked or fetched by privacy systems, and Google says it does not track or validate third-party open rates. Use final delivery status, authentication, complaints, provider telemetry, human placement checks and replies instead.

Should I move to a new domain after a reputation problem?

Not as a first response. A new domain does not fix a bad list, broken authentication, unwanted content or a suppression failure. Preserve the incident evidence, identify which identity/provider is affected, correct the cause and use a controlled restart. Follow the domain owner's policies and mailbox provider requirements.

What is the difference between a seed test and a real placement result?

A seed test reports where a sample message appeared in the specific test accounts it used. It does not represent every recipient or prove that a real buyer saw or wanted the message. Treat it as one controlled observation alongside receiver telemetry and campaign outcomes.

Yananai A. Chiwuta

Author

Yananai A. Chiwuta

CEO & Co-Founder

Yananai A. Chiwuta is the CEO and Co-Founder of Forma Nôrden, where he builds managed acquisition systems for B2B companies through signal-based outbound and precision paid ad acquisition. He has built and exited two companies, most recently FunnelVision.

Celine Sky-Chiwuta

Article reviewed by

Celine Sky-Chiwuta

Co-Founder & CMO

Celine Sky-Chiwuta is the Co-Founder and CMO of Forma Nôrden, where she shapes the positioning and marketing behind the company’s managed acquisition systems. She previously served as CMO of FunnelVision through its 2025 acquisition.

Related Articles