Multi-cloud without the multi-cloud tax
When running on both AWS and GCP is genuinely warranted, when it's cargo-culted, and what the difference costs you in practice.
"We should be multi-cloud" is one of those sentences that sounds like risk management and is often actually risk creation. Before signing up for it, it's worth being precise about which problem multi-cloud is supposed to solve, because the honest list of real reasons is much shorter than the list of reasons people say out loud.
The real reasons
- A specific capability genuinely lives on one platform. You need GCP's BigQuery for a workload AWS doesn't have an equivalent for, or you need AWS's breadth of managed services for something GCP doesn't cover as well. This is a real, defensible reason — it's a build-vs-buy decision at the platform level, not a hedge.
- A customer or regulator requires it. Some enterprise contracts and some compliance regimes genuinely require infrastructure diversity. If that's the actual requirement, multi-cloud isn't a choice, it's a spec.
- An acquisition brought a second cloud with it. This isn't a decision at all — it's an inheritance, and the real work is deciding whether to consolidate, not whether to "go multi-cloud."
The reason that doesn't hold up
"What if AWS goes down / raises prices / we get locked in" is the one that gets said most often and holds up least. A full second cloud, run well enough to actually fail over to in an emergency, isn't a hedge you get for free — it's a second production environment, with its own IAM model, its own networking primitives, its own observability stack, and its own set of edge cases your team now has to know cold on both platforms. Most organizations that adopt multi-cloud "for resilience" never actually test the failover, because testing it is expensive and scary, which means the resilience they're paying for continuously is one they've never verified they'd actually get in the moment that matters.
What the tax actually looks like
It isn't one line item — it shows up as a tax on every subsequent decision:
- Every internal tool either gets built twice, or built once against a lowest-common-denominator abstraction that's worse on both platforms than a native implementation would have been on either.
- Every hire needs either two clouds' worth of platform knowledge, or you end up with two separate platform teams that don't fully understand each other's half of the system.
- Every incident review has an extra branch: "did this happen because of the workload, or because of the abstraction layer sitting between the workload and the cloud."
What to do instead, most of the time
Pick one cloud as the default, and let a second cloud earn its way in for a specific, named capability — not as a category. That gets you the actual thing you wanted (access to the best tool for a specific job) without paying the tax on your entire platform for a resilience story you'll probably never test. If the day comes that you genuinely need a second cloud for a real reason, you'll know, because you'll be able to name the reason in one sentence instead of gesturing at "risk."