Back to Feed
Wednesday, Sep 9, 2026, 06:00 PM

How AWS Rebuilt Its Routing Control Plane With Zero Downtime (and What SREs Can Learn)

How AWS Rebuilt Its Routing Control Plane With Zero Downtime (and What SREs Can Learn)

Rebuilding core infrastructure while keeping the engine running is the ultimate challenge for any SRE. Amazon Web Services (AWS) recently detailed its multi-year journey of completely rewriting its routing control plane—the system responsible for managing network paths across its global footprint—without causing network degradation or downtime.

The SRE Challenge: Control Plane vs. Data Plane

At the heart of AWS's success was the strict separation of the control plane (which decides where traffic should go) and the data plane (which actually forwards the packets). By keeping the data plane simple, static, and highly resilient, AWS engineers were able to methodically swap out, test, and upgrade the complex routing control plane underneath.

Key SRE takeaways from this monumental effort include:

  1. Decouple Failure Domains: Ensure your customer-facing data plane can survive independently if your control plane or configuration systems go offline.
  2. Shadow Testing and Canarying: Run the new control plane in "shadow mode" alongside the old one, comparing decisions in real-time before giving the new system execution authority.
  3. Continuous Monitoring: When modifying fundamental routing layers, observing subtle anomalies before they cascade into outages is paramount.
  4. Expect the Unexpected: Even the most meticulously planned migrations can have micro-impacts on external dependencies.

How Rabbit SaaS Helps You Maintain Operational Visibility

While your engineering team might not be rewriting global BGP routing tables today, your application relies on cloud giants like AWS to stay online. When major infrastructure transitions occur behind the scenes, you need immediate visibility:

  • CloudStatusHQ: This is your window into upstream health. During major cloud-provider migrations, CloudStatusHQ aggregates and monitors third-party vendor dependency health, alerting your DevOps team the moment AWS or other critical SaaS providers experience localized routing or API anomalies.
  • Status Navigator: If an upstream routing hiccup does impact your services, Status Navigator lets you proactively communicate with your customers via beautiful, custom-branded status pages, keeping trust high while your SREs investigate.

Building resilient systems means preparing for changes both within your codebase and deep inside your cloud provider's stack. Keep your dependencies monitored and your customers informed with Rabbit SaaS.

Rabbit SaaS - Intelligent SaaS solutions