Skip to content
Post-mortem

Post-mortem: 72-Hour Emergency Migration from SendGrid

A customer's SendGrid account got terminated on a Tuesday with three days until their largest annual campaign. We migrated them to self-hosted infrastructure in 72 hours. The decisions made under time pressure, the corners we cut, the corners we did not, and what we learned about emergency migration.

The customer called us at 11:47 AM on Tuesday October 14. Their SendGrid account had been suspended that morning. SendGrid’s email cited content policy violations related to a campaign they had sent the previous week. The customer’s largest annual campaign was scheduled for Saturday October 18. They had subscribers expecting communication, advertisers paying for the campaign, and three days to figure out how to send approximately 280K messages without an ESP relationship.

We took the call, evaluated the situation, and agreed to attempt the migration in 72 hours. The migration completed successfully Friday afternoon. The Saturday campaign sent through new infrastructure. The numbers were within 2% of historical Saturday campaign performance.

This post-mortem documents the 72-hour migration: the decisions, the corners cut, the corners not cut, and what we learned about emergency migration from mainstream ESPs.

The customer’s situation

The customer is a B2C newsletter publisher with approximately 280K subscribers across multiple content verticals. The annual Saturday campaign is their largest single campaign of the year, with ad revenue tied to subscriber engagement.

Their pre-suspension setup:

  • SendGrid Pro plan with dedicated IPs
  • Multiple sending domains under their parent brand
  • DMARC at p=quarantine, progressing toward p=reject
  • Established reputation across major receivers
  • Approximately 2.4M monthly messages total volume

The suspension was triggered by content review of the prior week’s campaign. SendGrid’s review identified specific content patterns that violated their terms of service. The customer disputed the determination but SendGrid’s policy decision stood.

The customer’s options as they called us:

Option A: appeal the SendGrid suspension. Risk: appeals processes typically take 5-14 days. The Saturday campaign would not happen.

Option B: emergency signup with another mainstream ESP. Risk: any new ESP would require warmup. The Saturday campaign would have poor deliverability through cold IPs.

Option C: emergency migration to self-hosted infrastructure. Risk: 72 hours is a compressed migration timeline. Possible delivery quality issues.

The customer chose Option C with our assistance. The decision was based on:

  • Self-hosted independence from future ESP policy decisions
  • Better long-term economics at their scale
  • Confidence in our team’s capability to execute the migration

The Saturday campaign success would validate the choice. Failure would mean substantial revenue loss for the customer and reputational damage for us.

The 72-hour timeline

The compressed timeline forced specific decisions at each phase.

Hour 0-4 (Tuesday afternoon): Assessment and commitment

We reviewed the customer’s setup, technical requirements, and migration possibilities. The assessment covered:

  • Subscriber list size and engagement profile
  • Sending domain authentication setup
  • Application integration with SendGrid
  • Content patterns and template requirements
  • Webhook and event handling needs
  • Expected Saturday campaign volume and timing

The conclusion: migration was feasible but tight. We committed.

Hour 4-12 (Tuesday evening): Infrastructure provisioning

We provisioned the customer’s dedicated infrastructure in Bulgaria:

  • Dedicated server (32GB RAM, 8 CPU, NVMe storage)
  • PowerMTA installation with managed configuration
  • 6 dedicated IPs from our clean IP pool
  • DNS preparation for new authentication records

The infrastructure went online Tuesday evening with basic functionality. The IPs were pre-flighted before assignment to ensure clean reputation.

The risk acknowledged: 6 IPs is the minimum we recommend for the customer’s volume. Normally we would have allocated 8-10 IPs for proper isolation. Time constraint required accepting the lower IP count and managing the consequences during the campaign.

Hour 12-24 (Wednesday early morning): SendGrid data export

The customer initiated data export from SendGrid before their suspension became more restrictive. The export covered:

  • Full subscriber list (280K records)
  • Suppression list (unsubscribes, bounces, complaints)
  • Campaign templates (all active templates)
  • Recent campaign performance data for reference
  • Webhook configuration details for our reference

The export worked despite the suspension. SendGrid does not block data export during suspensions; they block sending. The 8-hour window for export was adequate.

Hour 24-36 (Wednesday afternoon): Authentication setup

We configured the customer’s authentication for the new infrastructure:

  • SPF record updates authorizing new IPs
  • DKIM key generation and signing setup
  • DMARC unchanged (already at p=quarantine)
  • PTR records configured for all 6 IPs

The DNS changes propagated over the next 4-6 hours. The authentication tests confirmed all records were resolving correctly by Wednesday evening.

The risk acknowledged: 4-6 hour DNS propagation reduces the margin for any issues. Normally we would allow 24-48 hours for full DNS propagation and verification.

Hour 36-48 (Wednesday night through Thursday morning): MailWizz installation and data import

We installed MailWizz on the customer’s infrastructure:

  • Database tuning for projected campaign volume
  • Web interface configuration with customer branding
  • API setup for application integration
  • Webhook configuration for delivery events

Subscriber list import:

  • All 280K subscribers imported successfully
  • Custom field mapping required some manual work
  • 1.4% of records failed import due to malformed data
  • The failed records were documented for customer review

Suppression list import:

  • All suppression entries imported
  • Critical to import before any sending to avoid contacting suppressed addresses

The risk acknowledged: 1.4% failed import is higher than the typical 1-2% we see in standard migrations. The tight timeline meant less time for data quality review and resolution.

Hour 48-60 (Thursday afternoon): Template port and testing

Template porting:

  • 23 active templates ported from SendGrid HTML to MailWizz
  • Each template required 30-90 minutes of work
  • Some templates required visual adjustment for MailWizz’s rendering

Application integration:

  • Customer’s website integration switched from SendGrid SDK to standard SMTP
  • Code changes deployed and tested
  • Trigger emails verified working

Test sends:

  • Internal test addresses received successful sends
  • Authentication passed at all major receivers
  • Template rendering verified across email clients

The risk acknowledged: 23 templates is significant content. Some templates had nuances we could not fully test in the available time.

Hour 60-72 (Thursday night through Friday afternoon): Production warmup and verification

This was the most operationally tense phase. New dedicated IPs need warmup. The customer needed to send 280K messages Saturday. The available warmup window was approximately 18 hours.

Our compressed warmup approach:

  • Friday morning: 1,000 messages per IP to most engaged subscribers (5,000 to recently-engaged segment)
  • Friday afternoon: 5,000 messages per IP to engaged subscribers (25,000 to engagement-segmented audience)
  • Friday evening: 15,000 messages per IP for routine transactional volume (60,000 to broader engaged audience)

The Friday volume served two purposes:

  • Establish initial sending patterns on new IPs
  • Process the customer’s normal Friday transactional volume

The volume ramp was aggressive but bounded. Each step was monitored for delivery rate, bounce rate, complaint signals.

The risk acknowledged: this is roughly 4-5x faster than standard IP warmup. The new IPs entered Saturday with limited reputation history. The Saturday campaign would carry the warmup pattern forward.

Hour 72 (Saturday morning): The campaign

Saturday morning, the campaign sent:

  • 280K messages across 6 IPs over 4 hours
  • Approximately 47K messages per IP, distributed by IP based on Friday warmup outcomes
  • Real-time monitoring for delivery rates and any issues

The campaign metrics from Saturday’s data:

  • Bounce rate: 0.43% (within 0.1% of historical SendGrid baseline)
  • Complaint rate: 0.06% (within 0.02% of historical baseline)
  • Open rate: 31.2% (within 2% of historical baseline)
  • Click rate: 4.1% (within 0.3% of historical baseline)

The numbers were within acceptable range. The customer’s Saturday operation succeeded.

What we cut corners on (intentionally)

The compressed timeline forced specific decisions about what corners to cut.

Limited IP pool

We deployed 6 IPs instead of recommended 8-10. The fewer IPs concentrated more volume per IP than ideal. The risk was acceptable for the immediate campaign but the customer added more IPs in the following weeks for normal operations.

Compressed warmup

The 18-hour warmup compressed what would normally be 14-21 days of careful ramping. The IPs entered the Saturday campaign with limited reputation history. The Saturday campaign was higher-risk than it would have been with full warmup.

The risk was mitigated by sending to most-engaged subscribers first, prioritizing engagement signals over pure volume.

Single-pop deployment

Normally we deploy customers across multiple geographic pops for resilience. The 72-hour timeline allowed only single-pop deployment in Bulgaria. The customer accepted the single-pop risk.

Limited template testing

23 templates received basic testing rather than comprehensive testing. Some edge cases that comprehensive testing would have caught were untested. The risk was mitigated by sending high-volume campaigns from the most-tested templates first.

Reduced operational redundancy

Normally we have backup capacity available for incident response. The compressed timeline meant the customer’s infrastructure had less redundancy than ideal. The risk was bounded by close monitoring during the campaign window.

Limited DMARC progression validation

The customer was at DMARC p=quarantine. We did not change this during migration. The risk was that new infrastructure might produce alignment issues that the quarantine policy would surface. We accepted this risk; aggregate reports during the first week confirmed alignment was holding.

What we did not cut corners on

Despite the time pressure, we maintained certain standards.

IP pre-flight verification

Every IP was pre-flight verified against blocklists, configuration, and reputation. We did not skip the verification despite time pressure. The verification took approximately 2 hours total for 6 IPs.

Authentication configuration

SPF, DKIM, DMARC, PTR records were all configured correctly. We did not deploy IPs without proper authentication. The DNS propagation time was tight but the configuration was complete.

Suppression list import

The customer’s suppression list was imported in full before any sending. We did not send to any previously-suppressed addresses. The risk of accidentally contacting unsubscribed users was unacceptable.

Customer communication

We communicated with the customer multiple times daily during the 72 hours. Status updates, decisions, risks, alternatives all surfaced for customer awareness and input.

Monitoring infrastructure

Real-time monitoring was set up before any production sending. We did not deploy sending capacity without visibility into what was happening.

Content review

We did not blindly accept the customer’s content. We reviewed for any patterns that might trigger receiver-side issues. The original SendGrid suspension was for content issues; we wanted to ensure the new infrastructure would not face the same issues.

The review found minor adjustments to the Saturday campaign that the customer made. The adjustments addressed the specific patterns SendGrid had flagged.

What we learned

The migration produced specific lessons.

Emergency migrations are possible

72 hours is tight but feasible for prepared infrastructure. The migration succeeded because we have established processes that could compress, not because the standard process is overly slow.

Customer relationships matter

The customer trusted us enough to commit to emergency migration on a call. The trust came from prior relationship rather than from cold-evaluating us during the crisis. Customers who establish provider relationships before they need emergency capability have better options when crises happen.

Engagement-segmented warmup works

Sending to most-engaged subscribers first during compressed warmup produced acceptable reputation. The pattern generalizes: when time is constrained, prioritize signal quality over volume.

Documentation matters under pressure

Our standard migration runbook was the foundation for compressed execution. Without the documented runbook, we would have been improvising rather than executing a compressed version of a known process. The investment in documentation paid off when documentation was the only thing maintaining structure.

Monitoring catches issues fast

Real-time monitoring during the Saturday campaign would have surfaced any issues before they compounded. Fortunately, no significant issues emerged. The monitoring was still critical for our and the customer’s confidence.

Customer decisions need explicit communication

The customer’s awareness of risks (compressed warmup, limited IPs, etc.) was operational risk management. They made informed decisions about acceptable risk for their situation. The explicit risk communication was as important as the technical work.

What we changed after the migration

The post-incident review produced procedural updates.

Emergency migration runbook

We created a specific emergency migration runbook based on this experience. The runbook covers:

  • Decision criteria for accepting emergency migrations
  • Compressed timeline templates (24-hour, 48-hour, 72-hour versions)
  • Specific corner-cutting decisions with risk assessment
  • Customer communication templates for emergency context
  • Post-migration monitoring intensification

The runbook ensures future emergency migrations have structure rather than ad-hoc improvisation.

Pre-built IP pool readiness

We tightened our IP pool inventory management. The 6 IPs allocated for this migration came from our standard pool with pre-flight verification. We now maintain larger inventory specifically for emergency scenarios.

Template porting tooling

The 23 templates required significant manual work. We invested in tooling that automates template porting from common ESP formats to MailWizz format. The tooling reduces emergency migration time significantly.

Customer onboarding documentation

We updated our customer onboarding to explicitly address emergency migration scenarios. New customers know:

  • That we can support emergency migrations
  • What the typical compressed timeline looks like
  • What risks they should be aware of
  • How to engage us if their primary ESP has issues

Provider relationship intelligence

We track which mainstream ESPs have suspension patterns affecting which customer types. The intelligence helps us advise customers about provider risks and helps us anticipate potential emergency migrations.

Operational capacity planning

The emergency migration consumed approximately 80 hours of senior engineering time over 72 hours. We need to ensure operational capacity can absorb periodic emergencies without affecting routine customer work.

What this revealed about mainstream ESP risk

The incident illustrates the operational risk of mainstream ESP dependence.

Suspension risk is real

The customer was operating in compliance with their understanding of SendGrid’s terms. The content review nonetheless produced suspension. Mainstream ESP terms have interpretation latitude that the ESP exercises at their discretion.

Suspension timing is unfavorable

The customer’s suspension came three days before their largest annual campaign. The timing was operational coincidence, but the impact would have been catastrophic without alternative infrastructure.

Recovery options are limited

Without self-hosted alternative, the customer’s options were appeal (slow) or new ESP (limited by warmup). Neither addressed the Saturday campaign timeline.

Long-term implications

The customer is now committed to self-hosted infrastructure. They are paying us less than SendGrid charged them. They are not vulnerable to future suspension. They have better operational independence.

The suspension was painful but produced strategic improvement. The customer’s long-term position is stronger.

What we tell customers about this experience

The customer agreed to anonymized sharing of the experience. Other customers benefit from understanding what is possible.

For customers on mainstream ESPs:

The risk of suspension is real. Build alternative infrastructure capacity before you need it.

The cost of self-hosted is not as high as commonly thought. At meaningful scale, self-hosted is cheaper than mainstream ESP.

The benefit of self-hosted is not just cost. Strategic independence from ESP policy decisions is real value.

The complexity of self-hosted is manageable. With managed services, the customer focus stays on their business.

For customers considering emergency migration:

The 72-hour timeline is feasible but tight. Most migrations should be planned with more time.

The cost of compressed timeline includes operational risk. Inform your team and stakeholders accordingly.

The alternative (waiting through ESP issues) may be worse. Compressed migration may be the right choice in emergencies.

The post-migration cleanup takes weeks. Plan for stabilization work after the immediate crisis resolves.

What we expect over time

Looking forward:

More emergency migrations likely

Mainstream ESPs are tightening their content policies. Suspensions are happening more frequently. Customers who were compliant under previous interpretations may not be compliant under current interpretations.

We expect emergency migration requests to continue. The runbook and tooling investments position us to handle them.

Provider concentration concerns

The major mainstream ESPs (SendGrid, Mailgun, Postmark, Amazon SES) collectively cover most of the market. Customer dependence on these providers creates systemic risk that emerging providers and self-hosted alternatives can address.

The shift to alternatives is gradual but persistent. Each emergency migration adds visible support for the alternative approach.

Self-hosted ecosystem maturation

The tooling for self-hosted infrastructure continues to mature. MailWizz, Acelle Mail, Postal, and PowerMTA all received meaningful updates over the past year. The capability of self-hosted approaches continues approaching mainstream ESP capability.

Customer expectations evolve

Customers increasingly understand the trade-offs between mainstream ESP convenience and self-hosted independence. The conversations we have with customers are more sophisticated than they were three years ago.

The longer-term customer outcome

The customer is now 5+ weeks past the migration. The operation is stable. Saturday campaigns since the migration have performed comparably or better than pre-migration baseline.

The customer reports:

  • Cost savings compared to SendGrid (approximately 60% lower)
  • Better operational control
  • More direct understanding of their email infrastructure
  • Stronger team understanding of their delivery patterns

The emergency migration produced strategic improvement that planned migration would have produced more smoothly but at slower pace. The forcing function of the SendGrid suspension accelerated what was probably the right long-term decision anyway.

The honest disclaimer

72-hour emergency migrations are stressful, risky, and best avoided. The customer in this post-mortem succeeded but the success required:

  • Established provider relationship (us)
  • Customer technical capability to execute decisions quickly
  • Compressed warmup that worked but might not have
  • Operational capacity to dedicate intense effort for 72 hours
  • Some luck that no surprises emerged during the window

Other customers facing similar situations should not assume the outcome will be similar. The right preparation is having alternative infrastructure ready before the emergency, not racing to build it during the emergency.

For our customer base, we recommend:

  • Maintaining at least one alternative sending path for critical operations
  • Documenting authentication setup so it can be applied to new infrastructure
  • Maintaining suppression list exports
  • Keeping templates in portable formats
  • Having a contingency plan with timeline assumptions

The preparation work is bounded. The protection it provides against emergency scenarios is significant. The customer who needs emergency migration should have done the preparation work in advance.

The 72 hours was achievable. We do not recommend treating it as standard. The customer’s outcome was good. The path was stressful. The right preparation makes alternative paths available without requiring 72-hour heroics.

Operating email infrastructure at scale?

We run anonymous server hosting for email operators across seven jurisdictions. Crypto-paid, no-KYC, PowerMTA-tuned. Look at the catalog or talk to us.