Back to Feed
Sunday, Sep 13, 2026, 11:00 PM

Building SRE from Scratch: Navigating the First 90 Days and Scoring Easy Wins

Building SRE from Scratch: Navigating the First 90 Days and Scoring Easy Wins

Starting a new Site Reliability Engineering (SRE) department is a daunting task, particularly when you are inheriting zero existing processes. A recent viral thread on Reddit's r/sre community highlights a classic scenario: two newly hired SREs started on the exact same day, only to discover there was no established SRE function left behind. Tasked by management to define the role from scratch, they turned to the community to ask: How do you build an SRE role from the ground up?

The 30/60/90 Day Roadmap for New SRE Teams

When establishing an SRE function, the community consensus is clear: do not try to re-architect systems or build custom platform tooling on day one. Instead, focus on gathering context, establishing boundaries, and securing quick, low-friction reliability wins:

  • Days 1–30 (Discovery & Visibility): Map out the system architecture. Identify critical dependencies, background scripts, and public-facing assets. Focus on external-facing infrastructure that could cause immediate, visible failures if neglected.
  • Days 31–60 (Standardization & Quick Wins): Set up basic external monitoring and alerting. Define boundaries between SRE, DevOps, and Product Engineering to manage expectations.
  • Days 61–90 (Proactive Automation): Begin automating away the 'toil' discovered in the first 60 days, establishing Service Level Objectives (SLOs) and incident response baselines.

Accelerating Your 'Day 1' SRE Strategy with Rabbit SaaS

One of the biggest traps for a fledgling SRE team is spending weeks configuring heavy, complex APM systems. To build credibility with management early on, you need immediate, high-value coverage that prevents embarrassing outages.

Here is how the Rabbit SaaS suite helps you establish robust baselines during your first week:

  • Stop Silent Background Failures with Cron Rabbit: Developers likely have dozens of critical background cron jobs running with no visibility. Use Cron Rabbit to set up curl-ping monitoring in minutes, ensuring you are alerted before a silent database backup or billing run failure goes unnoticed.
  • Audit External Assets with Domain Audit HQ & Certificate Guardian: Avoid the most common, preventable outages on day one. Certificate Guardian proactively monitors Certificate Transparency (CT) logs and SSL/TLS expirations, while Domain Audit HQ monitors WHOIS changes, DNS, and domain expirations.
  • Centralize Vendor Visibility with CloudStatusHQ: When external APIs go down, SREs shouldn't waste hours debugging internal code. CloudStatusHQ aggregates third-party vendor status feeds in a single pane of glass.
  • Build Communication Trust with Status Navigator: Easily set up custom-branded status pages to keep internal stakeholders and customers informed during incidents, keeping engineers focused on remediation rather than status updates.

By leveraging lightweight, targeted monitoring tools early in your roadmap, a brand new SRE team can secure their infrastructure and earn organizational trust from day one.

Source Link

www.reddit.com

Read the original Reddit discussion
Rabbit SaaS - Intelligent SaaS solutions