Module 11: Cloud
Read the transcript
1. Why cloud decisions are different
Host: Welcome back. This module is called Cloud, and the subtitle I keep coming back to is ‘decisions you can’t refactor away.’ If you ship bad code, you fix it in the next sprint. If you bake in the wrong network topology or the wrong IAM boundary, that mistake outlives every team that inherits it.
Guest: Right, and that’s really the whole premise. A region layout or a VPC boundary isn’t a line you change later — it’s the foundation everything else gets built on top of. By the time you notice it’s wrong, it’s already baked into every service built on top of it.
Host: So walk me through what we’re actually covering, because I know this isn’t just ‘cloud is scary, be careful.’
Guest: Four decisions, and they all reduce to the same trade — control versus operational burden. When to pay the premium for a managed service instead of self-hosting. How VPC and IAM actually contain a breach instead of just describing good intentions. What multi-region has to mean operationally — a tested recovery time and recovery point, not just a diagram with two regions on it. And cost attribution existing before the GPU bill becomes the reason the project gets killed.
2. Managed vs. self-hosted, as a default
Host: Let’s start with the first one, then, because it sounds like the simplest and it’s probably the most argued-about in every planning meeting: managed versus self-hosted. Vector database, Kubernetes control plane — where do you land by default?
Guest: Managed, by default, until a specific measured reason says otherwise. A managed vector database or a managed control plane takes patching, high-availability maintenance, and backup mechanics off your plate entirely, and you pay a real per-unit premium for that plus less control over tuning. Self-hosting gets you full control and usually lower steady-state cost at scale, but now every operational failure mode is yours to own.
Host: So what actually counts as a good enough reason to flip that default — ‘it feels expensive’ doesn’t qualify, I assume?
Guest: Right, it has to be measured, not felt. Cost at a scale you’ve actually reached, not a projection — or a tuning requirement the managed tier genuinely can’t expose, something you’ve hit in practice, not anticipated. Absent one of those two, self-hosting is just buying yourself an operations problem you didn’t need.
3. Network segmentation is only the first layer
Host: Okay, so once you’ve decided managed versus self-hosted, the next question is basically who can even reach the thing. Where does network segmentation come in?
Guest: It’s the baseline layer everyone assumes is already handled and sometimes isn’t. Public subnets are for load balancers and NAT gateways, the stuff that has to face the internet — everything else, your application tier, your database tier, belongs in a private subnet reachable only from inside the VPC. The concrete failure this prevents is exactly what it sounds like: a database or service that was never actually moved into a private subnet, sitting directly reachable from the internet, and nobody finds out until it gets scanned and exploited.
4. IAM is the real blast-radius control
Host: So say that database tier is safely tucked into a private subnet like you described. Are we done? Is the network boundary the thing that actually contains a breach?
Guest: No, and this is the part people underrate — network topology controls reachability, but IAM controls capability. If the app tier gets compromised, the question isn’t whether the attacker can see other machines on the network, it’s what that service’s role actually permits it to do: which S3 prefix, which database table, which KMS key. That’s the real blast radius.
Host: So it’s the same idea as Module 6’s tool permissions, just applied to cloud identity instead of an AI agent’s toolset.
Guest: Exactly the same principle. A per-service role scoped to exactly what that service needs means compromising it only gets you that narrow slice. But a shared, broadly-scoped role used across many services means compromising any one of them effectively compromises all of them — the network segmentation you just built doesn’t save you at that point.
5. IAM in code: a real least-privilege policy
Host: Let’s make that concrete then, because ‘scoped to exactly what it needs’ can sound vague until you see the actual policy. Walk me through what this looks like in practice for something like a model-serving role.
Guest: So the policy has two statements, and that’s the whole point — it grants s3 GetObject against one bucket path for model artifacts, and it grants DynamoDB GetItem and PutItem against one specific table for session state. Nothing else in the account is reachable through this role, no other bucket, no other table, no admin actions — so if this service gets compromised, the attacker inherits exactly those two narrow capabilities and nothing more. That’s what containment actually looks like written down, as opposed to just asserted in an architecture review.
6. When IAM fails: the shared-role trap
Host: Okay so you just showed the narrow version done right. What does it look like when a team skips that and just reuses one role everywhere?
Guest: It looks convenient right up until it doesn’t. Someone sets up a role with broad access once, attaches it to many services because writing scoped policies for each one is more work, and now the network segmentation you built earlier doesn’t matter. If the least-secure of those services gets popped, the attacker inherits everything all of them could touch — plus, if one of those services also happens to sit in a public subnet by oversight rather than design, they didn’t even need to work for that initial foothold.
7. RTO and RPO: recovery, not just redundancy
Host: Let’s move to the third axis: recovery. You’ve got a multi-region setup in the diagram, replicated database, load balancer ready to redirect traffic. Isn’t that the whole point of building it that way?
Guest: It’s the point of the diagram, not the point of the system. What actually matters are two numbers, RTO and RPO — how long you can be down, and how much data you can afford to lose, measured in time since the last successful replication. And here’s the part people skip: those numbers should come from business impact, not from whatever the architecture happens to already deliver. You don’t ask the system what it can do and call that the target.
Host: So you set the target first, based on business impact, and then build to hit it. What’s the failure mode if a team skips that and just trusts the diagram?
Guest: The failure mode is that a secondary region with a replicated database that’s never actually been failed over to isn’t a capability, it’s an assumption. Replication lag might make that RPO target physically impossible, and you genuinely don’t know until you’ve exercised the failover under realistic conditions — not a scheduled maintenance window where everyone’s already watching and ready to intervene. Until then, it’s a diagram, not a guarantee.
8. The replication mismatch that breaks RPO promises
Host: Let’s make that concrete. Walk me through the case where a team commits to a 15-minute RTO and a 1-minute RPO for a production model-serving system — where does that actually break?
Guest: The RTO is often achievable. The problem is the RPO — it requires synchronous or near-synchronous replication. If the team actually built it on asynchronous replication running 10 minutes behind, no failover mechanism on earth fixes that, because the data you’d recover to is already 10 minutes stale before the failover even starts.
Host: So the target and the replication strategy have to be chosen together, not one after the other.
Guest: Exactly — sync replication can hit that tight RPO but it costs you write latency and even availability, since a write can’t complete if the remote region is unreachable. Async keeps writes fast and local but caps your RPO at whatever the real-world lag turns out to be, not whatever number sounded reasonable in a planning doc — and you only find out which one you actually built during a real incident.
9. Proving failover actually works
Host: So you’re saying that RTO number needs to live in code, not just in a runbook. Walk me through what that actually looks like.
Guest: Right, look at the FailoverController — should_failover doesn’t fire on one bad health check, it only triggers if every check inside the RTO window came back unhealthy. That’s the whole point: the trigger is wired directly to the RTO you promised, not to some arbitrary retry count. And the lab makes you prove it — feed it checks spanning just under the RTO with one healthy blip in the middle, confirm it stays False, then feed it a full RTO window of nothing but failures and confirm it flips True. If you haven’t written that test, you don’t actually know your RTO is real, you just know it’s written down somewhere.
Host: And that same multi-region setup you’d build for failover — you’re saying it’s not purely insurance, it’s pulling double duty.
Guest: Exactly, serving users out of the nearest region cuts latency every single day, whether or not you ever failover — that’s often the stronger business case than the disaster-recovery story alone. But it cuts both ways: cross-region latency is physics, so the instant a failover sends traffic to a standby region in another geography, every user routed there gets a different latency profile, not just for the incident but for as long as they stay pinned there.
10. Cost attribution before the GPU bill becomes a crisis
Host: Let’s land on cost, then, because I think this is the one that sneaks up on teams. When does an unattributed cloud bill actually become a crisis?
Guest: The moment someone asks ‘why did spend jump’ and nobody can answer without a week of forensic digging. Tagging every resource by team or service, or better, putting each team in its own account or project, is what turns ‘the bill went up’ into ‘team X’s GPU usage went up, and here’s why.’ Without that, you can’t even start optimizing, because you don’t know what to optimize or who owns it.
Host: And that’s the bridge back to Module 4’s per-request cost tracking — you need attribution at the account level before that per-request number means anything organizationally.
Guest: Exactly, and it has to scale as more teams and more GPU-heavy workloads land in the same footprint — a tagging scheme that worked for one team erodes fast without deliberate extension. That’s really the thread through all four of these decisions: managed-versus-self-hosted, blast radius, recovery, cost — none of them are diagrams you draw once. They’re commitments you have to keep proving, or they quietly stop being true.
Not covered
The planner wanted these and found nothing in the source to support them:
- A live walkthrough of an actual multi-region failover incident with real timestamps and outcomes
- Specific cloud provider pricing comparisons for managed vs. self-hosted services
- A step-by-step tutorial for configuring VPC peering or transit gateways
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”Cloud architecture decisions are expensive to reverse — a network topology, an IAM boundary, or a region layout gets baked into every service built on top of it, unlike a code-level choice that a refactor can undo. This module covers the decisions that matter most for AI infrastructure specifically: when a managed service is worth its premium over self-hosting, how VPC and IAM boundaries actually contain a breach instead of just documenting an intention, what “multi-region” has to mean operationally rather than just architecturally, and why cost attribution has to exist before a GPU bill becomes the reason a project gets cancelled.
Mental Model
Section titled “Mental Model”Every cloud architecture decision in this module reduces to trading control for operational burden, along three largely independent axes:
- Build vs. buy — a managed service (managed vector database, managed Kubernetes control plane, managed model endpoint) trades direct control and per-unit cost for someone else owning availability, patching, and scaling mechanics.
- Blast radius — VPC segmentation and IAM scoping determine how far a single compromised credential or misconfigured service can reach before something stops it.
- Recovery, not just redundancy — multi-region existing in the architecture diagram doesn’t mean failover actually works; it means a defined, tested Recovery Time Objective (RTO) and Recovery Point Objective (RPO), and a mechanism that hits them.
Untested disaster recovery is not disaster recovery
A secondary region with a replicated database and a load balancer, that has never actually been failed over to, is an assumption, not a capability. The only way to know an RTO/RPO target is actually achievable is to have exercised the failover — under realistic conditions, not a scheduled maintenance window everyone was already watching — at least once. Everything else is a diagram, not a guarantee.
Architecture
Section titled “Architecture”flowchart TB
DNS["Global DNS / traffic manager (health-checked failover)"]
subgraph RegionA["Region A — primary"]
ALBA[Load balancer] --> AppA["App tier (private subnet)"]
AppA --> DataA[(Primary datastore)]
NATA[NAT gateway] --- AppA
end
subgraph RegionB["Region B — standby / DR"]
ALBB[Load balancer] --> AppB["App tier (private subnet)"]
AppB --> DataB[(Replica datastore)]
NATB[NAT gateway] --- AppB
end
DNS -->|active| ALBA
DNS -.failover on health-check failure.-> ALBB
DataA -.async cross-region replication.-> DataB
IAM["IAM: per-service roles, least privilege"] -.scopes access to.-> AppA
IAM -.scopes access to.-> AppBThe IAM boundary in this diagram isn’t a separate system bolted onto the network topology — it’s what actually determines the blast radius if either region’s app tier is compromised. A load balancer and a private subnet control network reachability; IAM controls what a compromised service can actually do once it’s running, which is usually the more consequential boundary in a cloud breach.
Deep Dive
Section titled “Deep Dive”Managed vs. self-hosted, applied to AI infrastructure. A managed vector database or managed Kubernetes control plane removes an entire category of operational work (patching, HA control-plane maintenance, backup mechanics) at a real per-unit cost premium and with less control over tuning. Self-hosting the same component gives full control and usually lower steady-state cost at scale, at the cost of owning every operational failure mode yourself. The right default for most teams is managed until a specific, measured reason (cost at scale, a tuning requirement the managed tier can’t expose) justifies the switch — not the reverse.
VPC segmentation. Public subnets (reachable from the internet, typically just load balancers and NAT gateways) and private subnets (application and data tiers, reachable only from inside the VPC) are the baseline network segmentation almost every cloud architecture needs. The specific mistake this prevents: an application or database tier directly reachable from the internet because it was never actually placed in a private subnet, discovered only when it’s scanned and exploited.
IAM as the real blast-radius control. A per-service IAM role scoped to exactly the resources that service needs (one S3 prefix, one DynamoDB table, one KMS key) means a compromised service can only reach what its role permits — regardless of what else exists in the account. A shared, broadly-scoped role used across many services means compromising any one of them effectively compromises all of them. This is the same least-privilege principle Module 6 applies to tool permissions, applied here to cloud identity.
RTO and RPO. Recovery Time Objective is how long you can be down; Recovery Point Objective is how much data you can afford to lose, measured in time (the gap since the last successful replication or backup). These numbers should come from the business impact of an outage, not from whatever the current architecture happens to deliver — and the architecture should then be built to hit them, including the replication lag between regions, which directly determines the achievable RPO.
Cost attribution. Cloud spend without per-team or per-service tagging is a single number nobody can act on. Tagging resources (or, more robustly, dedicated accounts/projects per team) enough to attribute spend to whoever incurred it is what turns “the cloud bill went up” into “team X’s GPU usage went up, and here’s why” — a prerequisite for any of Module 4’s per-request cost tracking to matter at an organizational level.
Research Note
Not AI-specific, but the underlying reliability and cost-optimization questions (define RTO/RPO before choosing an architecture; attribute cost before optimizing it) apply directly and are worth reading in the source rather than through a secondary summary.
Source: AWS Well-Architected Framework, Reliability and Cost Optimization pillars
Implementation
Section titled “Implementation”A least-privilege IAM policy scoping a model-serving service’s role to exactly the resources it needs — the concrete version of this module’s “IAM as blast-radius control” principle — followed by a health-check-driven failover decision, showing RTO/RPO as executable logic rather than a diagram label:
Scoped IAM policy and a failover decision that respects a defined RTO
{ "Version": "2012-10-17", "Statement": [ { "Sid": "ReadModelArtifacts", "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::model-artifacts-prod/model-server/*" }, { "Sid": "ReadWriteSessionState", "Effect": "Allow", "Action": ["dynamodb:GetItem", "dynamodb:PutItem"], "Resource": "arn:aws:dynamodb:us-east-1:123456789012:table/model-server-sessions" } ]}from __future__ import annotations
from dataclasses import dataclass, fieldfrom datetime import datetime, timedelta
@dataclassclass HealthCheck: timestamp: datetime healthy: bool
@dataclassclass FailoverController: rto: timedelta checks: list[HealthCheck] = field(default_factory=list)
def record(self, check: HealthCheck) -> None: self.checks.append(check)
def should_failover(self, now: datetime) -> bool: window_start = now - self.rto recent = [c for c in self.checks if c.timestamp >= window_start] if not recent: return False return all(not c.healthy for c in recent)The policy grants exactly two actions against exactly two resources — nothing this service doesn’t
use is reachable through this role, so compromising it doesn’t compromise anything else in the
account. should_failover only triggers once every health check inside the RTO window has failed —
a single transient failure doesn’t trigger a failover, but sustained failure for the entire RTO
window does, tying the failover trigger directly to the RTO number it’s meant to honor rather than
to an arbitrary retry count.
Measured: tying the window to the RTO guarantees missing it
The Region Failover Budget lab runs this controller and
found three things the paragraph above does not say. Detection alone takes the whole window,
and promotion, traffic shift, and loading the model onto the standby’s GPUs all come after it — so
a window equal to the RTO overruns the RTO on every standby tier, hot included. Size the window
from what the RTO leaves after those stages. if not recent: return False reads silence as
health: a region that goes dark, taking the prober with it, is never failed over. Count a probe
that was due and never arrived as a failure. And a freshly started controller fires on its first
unhealthy probe, because with no history that probe is every check in the window.
Production Example
Section titled “Production Example”A team commits to a 15-minute RTO and a 1-minute RPO for a production model-serving system. The 1-minute RPO requires synchronous or near-synchronous cross-region replication for anything that can’t tolerate a minute of data loss — asynchronous replication with a 10-minute lag, however inexpensive, structurally cannot meet that RPO no matter how fast the failover mechanism itself is. This is a common mismatch: teams pick a comfortable-sounding RTO/RPO target and then build with a replication strategy that can’t actually deliver it, discovering the gap only during a real incident.
Failure Modes
Section titled “Failure Modes”Untested failover that doesn't actually work
A standby region exists in the architecture but has never been failed over to — configuration drift between primary and standby (a missing environment variable, an IAM role that was never replicated) surfaces for the first time during a real incident, exactly when it’s most expensive to discover. This module’s Mental Model section treats a failover as unproven until exercised.
Broad IAM roles shared across services
A single IAM role reused across many services because it was easier to set up once means compromising the least-secure service using that role compromises every resource the role can reach — turning a contained incident into an account-wide one.
Replication lag exceeding the committed RPO
Asynchronous cross-region replication with real-world lag that exceeds the committed RPO means a failover recovers to a point further in the past than promised — this is a data-loss incident on top of an availability incident, and it’s silent until the RPO is actually tested against real replication behavior, not just its documented target.
Public subnet placement for application or data tiers
A service placed in a public subnet by oversight — rather than deliberate design — is directly reachable from the internet regardless of any application-layer authentication, removing the network-level containment this module’s Deep Dive section relies on as the first layer of defense, not the only one.
Trade-offs
Section titled “Trade-offs”Managed service vs. self-hosted, for AI infrastructure components
Managed removes operational ownership of availability, patching, and backups at a real cost premium and reduced tuning control. Self-hosted gives full control and typically lower cost at scale, at the cost of owning every failure mode — the right default is managed until a specific, measured reason justifies the switch.
Synchronous vs. asynchronous cross-region replication
Synchronous replication can meet a tight RPO but adds write latency and availability risk (a write can’t complete if the remote region is unreachable). Asynchronous replication keeps write latency local but caps the achievable RPO at however far behind replication actually runs — the choice has to be driven by the committed RPO, not picked independently of it.
Per-resource tagging vs. dedicated accounts for cost attribution
Tagging is cheaper to retrofit onto an existing account structure but depends on tagging discipline being enforced everywhere, which erodes over time. Separate accounts or projects per team enforce cost (and often security) boundaries structurally, at the cost of more infrastructure-as-code overhead to manage them.
Security
Section titled “Security”- IAM least privilege is the primary blast-radius control, per this module’s Deep Dive and Implementation sections — network segmentation alone does not contain what a compromised service can do once it’s already running inside the network.
- Secrets (API keys, database credentials) belong in a managed secrets store with scoped access, not in environment variables baked into a container image or committed to a repository — the same credential-scoping discipline this module applies to IAM roles applies to secrets specifically.
- Cross-region replication traffic needs to be encrypted in transit by default, not as an optional hardening step, since it routinely carries the same sensitive data the primary region’s access controls were built to protect.
Performance
Section titled “Performance”- Cross-region latency is a real, physics-bound cost that no architecture choice removes — a failover to a standby region in a different geography changes the latency profile for every user whose traffic now routes there, not just for the duration of the incident.
- VPC-internal traffic (private subnet to private subnet) is typically far lower latency than any path crossing a public load balancer — architecture that keeps service-to-service calls inside the VPC avoids paying public-path latency and cost for purely internal traffic.
- Replication lag is a performance number with a direct security and reliability consequence — it isn’t just “how fresh is the read replica,” it’s “how much data would a failover right now actually lose,” per this module’s Production Example.
Scaling
Section titled “Scaling”- Multi-region is a scaling strategy for both latency and availability, not solely a disaster- recovery mechanism — serving users from the nearest region reduces latency independent of whether a failover ever happens, which is often the stronger justification for the investment.
- IAM and network segmentation need to scale with organizational growth, not just traffic — a scoping scheme that worked for five services becomes unmanageable at fifty without a systematic convention (per-team accounts, a consistent tagging and role-naming scheme) established early.
- Cost attribution has to scale with the number of teams sharing the cloud footprint — a tagging scheme adequate for one team’s workloads needs deliberate extension, not assumption, as more teams and more GPU-heavy workloads (per Module 9) join the same account structure.
Interview Questions
Section titled “Interview Questions”How would you design multi-region disaster recovery for a production AI system, and how do you know it actually works?
Start from a business-driven RTO/RPO, choose a replication strategy (sync vs. async) that can actually hit the RPO, and then prove it by executing a real failover, not just reviewing the architecture diagram — an untested failover is an assumption, per this module’s Mental Model section, not a working capability.
Why is IAM scoping more important than VPC network segmentation for containing a breach?
Network segmentation controls reachability from outside a boundary; IAM controls what a service can do once it’s already running inside that boundary, which is exactly the situation after a service is compromised. A broadly-scoped IAM role reachable from a well-segmented network still gives an attacker who compromises that service everything the role can touch.
What's the practical difference between RTO and RPO, and why do teams commonly miss the target for one of them?
RTO bounds downtime; RPO bounds data loss, measured as replication or backup lag. Teams commonly commit to a tight RPO without checking that their actual replication mechanism (often async, for cost reasons) can deliver it — the RPO target and the replication strategy have to be chosen together, not independently, per this module’s Production Example.
How would you decide between a managed vector database and self-hosting one for a RAG system?
Default to managed unless a specific, measured reason — cost at a scale you’ve actually reached, or a tuning capability the managed tier doesn’t expose — justifies the operational burden of self-hosting. See Module 8 for what a self-hosted index actually has to get right (sharding, tenant isolation) if that switch is made.
Hands-on Lab
Section titled “Hands-on Lab”Hands-on Lab
This module’s FailoverController, measured. A seeded simulation of regions, probes, blips, and
outages sets it against a consecutive-failure trigger and a silence rule: a window tied to the
RTO spends the whole RTO on detection, a dark region is caught only by luck, and the eager
alternative fails over on blips dozens of times a day.
Read the lab documentation →
labs/region-failover-budgetproduction-shaped
Prove a failover decision against a committed RTO
Using this module’s FailoverController, write a test that feeds it a sequence of health checks
spanning slightly less than the configured RTO with one healthy check in the middle — confirm
should_failover returns False — then a sequence spanning the full RTO with every check
unhealthy — confirm it returns True. This is the same discipline this module argues for: prove
the failover trigger actually honors the RTO it’s supposed to respect, rather than assuming it
does.
References
Section titled “References”- AWS Well-Architected Framework, “Reliability Pillar” and “Cost Optimization Pillar”
- AWS, “IAM Best Practices”
- Module 4: AI Infrastructure — the per-request cost tracking pattern this module’s cost attribution section builds on organizationally.
- Module 6: MCP — the least-privilege principle this module applies to cloud IAM instead of tool permissions.
- Module 9: Model Serving — the GPU capacity and cold-start constraints this module’s scaling section references.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-07 | Initial publication. |