βοΈ @amitmahata
1/8
System Design Framework & Fundamentals
β
4-Step Senior System Design Framework (45 mins)
1. Clarify & Scope
β’ Functional requirements
β’ Non-functional (SLA/QPS)
β’ Back-of-envelope math
(5 - 7 mins)
β’ Non-functional (SLA/QPS)
β’ Back-of-envelope math
(5 - 7 mins)
2. High-Level Flow
β’ API contract (REST/gRPC)
β’ Core DB schema entity
β’ End-to-end block diagram
(10 - 15 mins)
β’ Core DB schema entity
β’ End-to-end block diagram
(10 - 15 mins)
3. Deep Dive Core
β’ Scale bottlenecks
β’ Cache / Sharding / Queue
β’ Algorithms & data models
(15 - 20 mins)
β’ Cache / Sharding / Queue
β’ Algorithms & data models
(15 - 20 mins)
4. Bottlenecks & Scale
β’ SPOF & Fault tolerance
β’ Monitoring, SLOs, Tracing
β’ Cost & trade-offs
(5 mins)
β’ Monitoring, SLOs, Tracing
β’ Cost & trade-offs
(5 mins)
β‘
End-to-End Architecture Flow
π Client (Web/Mobile)
Anycast / GeoDNS
β Static Assets & Edge Cache
β‘ CDN Edge Server (Cloudflare/CloudFront)
Edge PoP
β Dynamic Requests
βοΈ L4/L7 Load Balancer (NGINX / ALB)
SSL / Health
β Rate Limiting & Auth
πͺ API Gateway (Routing, Circuit Breaker)
Token Bucket
β Internal Microservices
βοΈ App Services + Cache (Redis Cluster)
Cache-Aside
β Async Events / DB Writes
ποΈ Primary/Replica DB + Kafka Queue
WAL / CDC
β’
Numbers Every 8+ YoE Engineer Must Know
β±οΈ Latency Hierarchy (Orders of Magnitude)
- L1 cache reference: ~ 0.5 - 1 ns
- L2 cache reference: ~ 3 - 7 ns
- Main Memory (RAM) access: ~ 100 ns
- NVMe SSD random read: ~ 50 - 100 Β΅s (1000x RAM)
- Read 1 MB sequentially from RAM: ~ 3 Β΅s
- Read 1 MB sequentially from SSD: ~ 1 ms
- Datacenter network roundtrip (RTT): ~ 0.5 ms
- Intercontinental RTT (US β EU): ~ 150 ms
π Back-of-Envelope Quick Multipliers
β’ 1 Day = 86,400 s β 100K seconds (for quick mental math)
β’ 1 Million req/day β 12 QPS | 100M req/day β 1.2K QPS
β’ Peak Traffic multiplier: Design for 2x to 5x of average QPS!
β’ 1 Million req/day β 12 QPS | 100M req/day β 1.2K QPS
β’ Peak Traffic multiplier: Design for 2x to 5x of average QPS!
β£
Availability & SLA Tiers
Availability Tiers (The "Nines")
99.0% ("2 Nines")3.65 days downtime / year
99.9% ("3 Nines")8.76 hours downtime / year
99.99% ("4 Nines")52.6 minutes downtime / year
99.999% ("5 Nines")5.26 minutes downtime / year (Telco/Fintech)
Key Reliability Metrics
MTBF (Mean Time Between Failures)System Uptime reliability
MTTR (Mean Time to Repair)Recovery speed & automated failover
RPO (Recovery Point Objective)Max acceptable data loss
RTO (Recovery Time Objective)Max acceptable downtime
β Senior Takeaway
β Never jump straight to drawing databases or Kafka topics! First establish Read-to-Write ratio (e.g., Twitter 100:1 vs Chat 1:1), storage size per record, and whether strong consistency is non-negotiable (e.g. Payments) or eventual consistency is acceptable (e.g. Likes, News Feed).
π‘ Staff / Senior Interview Tip
Q: How do you stand out as an 8+ YoE candidate in the first 10 minutes?
A: Don't wait for the interviewer to give you requirements. Actively drive the scope: "Assuming 50M DAU, each user making 20 requests daily yields ~12,000 QPS average, peaking at 35,000 QPS. At 500 bytes per payload, we'll store 500GB/day = 180TB over 5 years. I propose starting with an API Gateway + Cache-Aside architecture to offload 90% of reads."
A: Don't wait for the interviewer to give you requirements. Actively drive the scope: "Assuming 50M DAU, each user making 20 requests daily yields ~12,000 QPS average, peaking at 35,000 QPS. At 500 bytes per payload, we'll store 500GB/day = 180TB over 5 years. I propose starting with an API Gateway + Cache-Aside architecture to offload 90% of reads."