← Lakefront

Resilience

Watch it,
and survive it.

Lakefront reads metrics straight out of your cloud provider, alerts on what they mean, and moves traffic to a healthy region before anyone opens a ticket. There is no agent to install and no telemetry bill, because nothing is ingested.

checkout-api · Metrics
checkout-api1h24h7d
Requests
1.84M
Error rate
0.02%
P95
68ms
Replicas
6
Requests2xx
00:0012:00now
Latencyp95
00:0012:00now
p95 above 400ms for 5 minutesSlack · #checkout-oncallArmed

Metrics without an ingestion pipeline

Requests, error rates, latency percentiles, CPU and memory come from Azure Monitor and CloudWatch on demand, cached briefly and rendered in one dashboard. Your observability data never leaves your account, and you are not billed twice for it.

  • 1h, 24h and 7dThe ranges that answer “is it happening now” and “has it been happening”.
  • Alert rulesThreshold rules on any series, routed to Slack or email.
  • Runtime logsTail any service in the browser, across both providers.
checkout-api · Logs
Alert rules3 armed
p95 latency above 400ms5 minSlackArmed
5xx rate above 1%2 minSlack · emailArmed
Replicas pinned at max15 minemailArmed
Live tailrevision 0042 · 6 replicas
12:04:18 GET /api/cart 200 18ms
12:04:18 POST /api/checkout 201 64ms
12:04:19 GET /api/rates 429 upstream throttled
12:04:19 GET /healthz 200 2ms

Failover you have actually rehearsed

A protected environment keeps a warm standby in a second region, or a second cloud entirely. When probes stop answering, the standby is promoted and DNS cuts over, typically inside ninety seconds. Drills run the same path on purpose so you find out before an incident does.

  • Cross-cloud standbysAzure primary, AWS standby, if that is the risk you are hedging.
  • A real state machineProvisioning, protected, failing over, failed over, failing back.
  • Scheduled drillsProve the path works, on a cadence, with a report.
Recovery · checkout-api
Tier 2 · multi-regionProtectedRTO 90s · RPO 0
eastusprimary6 replicas · 100% trafficHealthy
westeuropestandby2 replicas · warmHealthy
us-east-1cross-cloud2 replicas · warmHealthy
Last drill14 days ago
02:14:03 primary eastus unreachable · 3 consecutive probes
02:14:11 promoting westeurope → primary
02:15:22 DNS cut over · 79s total
✓ failed over with zero dropped requests

Say it before they ask

A public status page, an incident workspace, and updates drafted from the real signals. The composer shows exactly which internal details it removed before anything is published, so an honest update never becomes an accidental disclosure.

Incidents · INC-31
Elevated latency in eastusMonitoring
We identified elevated response times affecting a subset of checkout requests in North America. Traffic has been shifted and latency is back to normal. We are monitoring.
Drafted from 12 signals3 details redactedPost publicly
status.acme.com
Checkout99.98% · 90d
Payments APIdegraded · 14m
Dashboard99.99% · 90d

Cost, from your own bill

Deploy-time estimates before you commit, and month-to-date spend after, broken down by service and provider. Lakefront adds no markup because Lakefront never touches the transaction.

Costs
This monthzero markup
Azure
$1,284
AWS
$742
Forecast
$2,190
Aug 1Aug 14
checkout-apiContainer Apps$612
ledger-workerFargate$388
checkout-dbFlexible Server$284

Deploy it into your own cloud.