How Sanity Cut Their Grafana Mimir Bill 48%: The Cost of Multi-Zone Bandwidth and E2 Instances
The engineering team at Sanity recently shared a highly educational breakdown of how they cut their self-hosted Grafana Mimir operational costs by 48% on Google Cloud Platform (GCP). Their experience highlights a common, silent cost driver in modern distributed systems: inter-zone network egress and suboptimal machine family selection.
The Two High-Impact Changes
-
Zonal chunks-caches (64% Reduction in Inter-Zone Bandwidth Costs): Originally, Sanity deployed a single
chunks-cachelayer (9 pods topology-spread across three zones). Because Mimir's store-gateways query the cache constantly, a store-gateway in Zone A had a 2-in-3 chance of hitting a cache pod in Zone B or C. GCP charges for inter-zone data transfer, transforming their high cache hit-rate into a massive network bill. By deploying dedicatedchunks-cacheinstances per zone and forcing local store-gateways to query their respective zone's cache, they drastically minimized cross-zone traffic. -
Migrating to ARM C4A Instances (50% vCPU Reduction & Halved Latency): The team migrated CPU-heavy components (ingesters, distributors, and rulers) off standard GCP E2 shared-core-like instances onto ARM-based C4A instances. This shift cut their overall vCPU requirements in half while simultaneously dropping read latency from ~20ms to under 10ms.
SRE Best Practices & How Rabbit SaaS Helps
Optimizing your core observability stack is essential, but complex infrastructure migrations introduce risks of downtime, misconfigurations, and silent failures. Here is how Rabbit SaaS tools help secure your infrastructure during major shifts:
- Keep Maintenance Transparent with Status Navigator: Changing caching topology and scaling down VM clusters can lead to temporary query degradation or planned downtime. Use Status Navigator to host external, custom-branded status pages, ensuring your team and external stakeholders remain updated during critical architectural updates.
- Prevent Silent Maintenance Failures with Cron Rabbit: In self-hosted Mimir configurations, critical maintenance operations—such as bucket compaction, ledger cleanups, and rule evaluations—rely on background cron jobs. If these background processes fail silently during an instance migration, it can lead to massive storage overheads. Cron Rabbit monitors these background tasks via simple curl pings, alerting your SRE team immediately if a job fails to run.
- Audit Your Underpinning Infrastructure with Domain Audit HQ: Ensure that your public-facing endpoints, DNS routing, and internal network domains are completely secure and monitored before making massive DNS-level route changes or migrating clusters across networks.
Source Link
www.reddit.com
