When we published How moving from AWS to Bare-Metal saved us $230,000 /yr. in 2023, the story travelled far beyond our usual readership. The discussion threads on Hacker News and Reddit were packed with sharp questions: did we skip Reserved Instances, how do we fail over a single rack, what about the people cost, and when is cloud still the better answer? This follow-up is our long-form reply.

Over the last twenty-four months we:

Below we tackle the recurring themes from the community feedback, complete with the numbers we use internally.

$230,000 / yr savings? That is just an engineers salary.

In the US, it is. In the rest of the world. That's 2-5x engineers salary. We used to save $230,000 / yr but now the savings have exponentially grown. We now save over $1.2M / yr and we expect this to grow, as we grow as a business.

“Why not just buy Savings Plans or Reserved Instances?”

We tried. Long answer: the maths still favoured bare metal once we priced everything in. We see a savings of over 76% if you compare our bare metal setup to AWS.

A few clarifications:

“How much did migration and ongoing ops really cost?”

We spent a week of engineers time (and that is the worst case estimate) on the initial migration, spread across SRE, platform, and database owners. Most of that time was work we needed anyway- formalising infrastructure-as-code, smoke testing charts, tightening backup policies. The incremental work that existed purely because of bare metal was roughly one week.

Ongoing run-cost looks like this:

The opportunity cost question from is fair. We track it the same way we track feature velocity: did the infra team ship less? The answer was “no”- our release cadence increased because we reclaimed few hours/month we used to spend in AWS “cost council” meetings.

“Isn’t a single rack a single point of failure?”

We have multiple racks across two different DC / providers. We:

The AWS failover cluster we mentioned in 2023 still exists. We rehearse a full cutover quarterly using the same Helm releases we ship to customers. DNS failover remains the slowest leg (resolver caches can ignore TTL), so we added Anycast ingress via BGP with our transit provider to cut traffic shifting to sub-minute.

“What about hardware lifecycle and surprise CapEx?”

We amortise servers over five years, but we sized them with 2 × AMD EPYC 9654 CPUs, 1 TB RAM, and NVMe sleds. At our current growth rate the boxes will hit CPU saturation before we hit year five. When that happens, the plan is to cascade the older gear into our regional analytics cluster (we use Posthog + Metabase for this) and buy a new batch. Thanks to the savings delta, we can refresh 40% of the fleet every 24 months and still spend less annually than the optimised AWS bill above.

We also buy extended warranties from the OEM (Supermicro) and keep three cold spares in the cage. The hardware lasts 7-8 years and not 5, but we wtill count it as 5 to be very conservative.

“Are you reinventing managed services?”

Another strong Reddit critique: why rebuild services AWS already offers? Three reasons we are comfortable with the trade:

  1. Portability is part of our product promise. OneUptime customers self-host in their own environments. Running the same open stack we ship (Postgres, Redis, ClickHouse, etc.) keeps us honest. We eun on Kubernetes and self-hosted customers run on Kubernetes as well.
  2. Tooling maturity. Two years ago we relied on Terraform + EKS + RDS. Today we run MicroK8s (Talos in the future), Argo Rollouts, OpenTelemetry Collector, and Ceph dashboards. None of that is bespoke. We do not maintain a fork of anything.
  3. Selective cloud use. We still pay AWS for Glacier backups, CloudFront for edge caching, and short-lived burst capacity for load tests. Cloud makes sense when elasticity matters; bare metal wins when baseload dominates.

Managed services are phenomenal when you are short on expertise or need features beyond commodity compute. If we were all-in on DynamoDB streams or Step Functions we would almost certainly still be on AWS.

“How do bandwidth and DoS scenarios work now?”

We committed to 5 Gbps 95th percentile across two carriers. The same traffic on AWS egress would be 8x expensive in eu-west-1. For DDoS protection we front our ingress with Cloudflare.

“Has reliability suffered?”

Short answer: No. Infact it was better than AWS (compared to recent AWS downtimes)

We have 730+ days with 99.993% measured availability and we also escaped AWS region wide downtime that happened a week ago.

“How do audits and compliance work off-cloud now?”

We stayed SOC 2 Type II and ISO 27001 certified through the transition. The biggest deltas auditors cared about:

If you are in a regulated space (HIPAA for instance), expect the paperwork to grow a little. We worked it in by leaning on the colo providers’ standard compliance packets- they slotted straight into our risk register.

“Why not stay in the cloud but switch providers?”

We priced Hetzner, OVH, Leaseweb, Equinix Metal, and AWS Outposts. The short version:

Owning the hardware also let us plan power density (we run 15 kW racks) and reuse components. For our steady-state footprint, colocation won by a long shot.

“What does day-to-day toil look like now?”

We put real numbers to it because Reddit kept us honest:

Total toil is ~14 engineer-hours/month, including prep. The AWS era had us spending similar time but on different work: chasing cost anomalies, expanding Security Hub exceptions, and mapping breaking changes in managed services. The toil moved; it did not multiply.

“Do you still use the cloud for anything substantial?”

Absolutely. Cloud still solves problems we would rather not own:

So yes, we left AWS for the base workload, but we still swipe the corporate card when elasticity or geography outweighs fixed-cost savings.

When the cloud is still the right answer

It depends on your workload. We still recommend staying put if:

Cloud-first was the right call for our first five years. Bare metal became the right call once our compute footprint, data gravity, and independence requirements stabilised.

What is next

Questions we did not cover? Let us know in the discussion threads- we are happy to keep sharing the gritty details.

Related Reading: