Fundamentals
60 posts — newest first.
-
Sagas and Two-Phase Commit: What You Do When There Is No Rollback
Three services, one transaction, and no shared database. Why 2PC blocks, how sagas trade atomicity for compensation, and why compensation is not undo.
-
Percentiles, Properly: You Cannot Average a p99
Averaging percentiles across shards or minutes produces a number that means nothing. Here is the arithmetic, plus why p99 gets worse as you add services.
-
Clock Skew: Why now() Is the Least Trustworthy Function You Call
Wall clocks jump backwards, NTP corrects under you, and two machines never agree. What to use instead, and where time is load-bearing anyway.
-
Observability for Agents: What You Instrument When the System Decides for Itself
A request either worked or didn't. An agent run can succeed on every span and still be wrong, slow and expensive. What the GenAI conventions ask you to record.
-
Prompt Injection: Why Your Agent Believes the Wrong Text
An LLM sees one stream of tokens, not instructions and data. That is why prompt injection has no parser fix — and why provenance, not filtering, is the control.
-
Token Exchange: Delegation Is Not Impersonation (RFC 8693)
When a service calls another service for a user, the token it carries decides whether your audit log says who acted. RFC 8693 makes that a design choice.
-
How CDNs Actually Work: Edge Caching, TTLs, and the Invalidation Problem
A CDN is a globe-spanning cache in front of your origin. The wins are latency and offload; the hard parts are cache keys, TTLs, and knowing when to invalidate.
-
OAuth, OIDC, and JWT: Who Are You, and What Are You Allowed to Do?
Three terms people use interchangeably and shouldn't: OAuth delegates access, OIDC proves identity, JWT is the token format. Mixing them up causes real bugs.
-
Raft: How a Cluster of Machines Agrees on a Single Truth
Consensus sounds academic until etcd loses quorum and your cluster freezes. Raft explained: leader election, log replication, and why majorities are the trick.
-
Sharding and Partitioning: Splitting Data Without Splitting Your Sanity
One database eventually runs out of one machine. Sharding buys headroom — but the shard key you pick is a near-permanent decision. Here's how to choose it.
-
Write-Ahead Logging: How a Database Survives Being Killed Mid-Write
Pull the power on a database mid-transaction and it comes back intact. The trick is boring and universal: write your intent to a log before you touch the data.
-
Secrets Management: Scope, Lifetime, and the Blast Radius of a Leak
Storing secrets safely is the easy half. The damage is decided by what one credential can do, for how long, and whether you can revoke it safely.
-
Load Balancing and Health Checks: The Three States Your Checks Don't Have
Balancing decides where load goes; health checks decide where it stops. Most outages here come from a check that can't tell busy apart from broken.
-
Idempotency Keys: The Thing That Makes a Retry Safe
A timeout never tells you whether the work happened. Idempotency keys resolve that — if you get the concurrency and scope right.
-
Vector Index Trade-offs: Flat vs IVF vs HNSW, and When Quantization Pays
Exact search is O(n) and always right. Every ANN index buys speed with recall, memory, or build time. The arithmetic, compared.
-
Rate Limiting Algorithms: Token Bucket, Leaky Bucket, and Sliding Window, Compared
Five algorithms, one trade-off: burst tolerance vs memory vs precision. Plus why LLM APIs limit tokens instead of requests.
-
Retries, Backoff, Jitter, and Circuit Breakers: The Four Controls That Decide Whether You Recover or Melt Down
A retry is a load multiplier. Backoff spreads it, jitter decorrelates it, a breaker stops it — and agent loops break all three.
-
Queueing Theory for SREs: Little's Law and the Utilisation Knee
Latency doesn't rise with load, it rises with 1/(1−utilisation). Two formulas explain the p99 cliff and size every pool in your system.
-
Cascading Failure: Backpressure, Load Shedding, and Graceful Degradation
A cascading failure is a feedback loop, not a row of dominoes. Removing the trigger does not stop it — which is why the controls have to be in place beforehand.
-
Timeouts and Deadline Propagation: The Budget Nobody Sets
A timeout is a local guess; a deadline is a shared fact. The difference decides whether your system sheds doomed work or grinds on abandoned requests.
-
Progressive Delivery: Canary, Blue-Green, and the Rollback You Never Tested
Every deployment strategy answers one question: who pays for the bug? Rolling, blue-green, and canary answer differently — only one limits blast radius.
-
The Deployment Pipeline: Build Once, Promote Everywhere
If your artifact is rebuilt per environment, staging never tested what production runs. Build once, promote everywhere is the rule the rest of CD rests on.
-
Infrastructure as Code: State, Drift, and Why the Plan Is the Product
IaC is a three-way comparison: config, state, reality. Every confusing Terraform outcome is two of them disagreeing — and the plan tells you which.
-
The Reconciliation Loop: Why Kubernetes Self-Heals and Your Script Doesn't
Level-triggered control loops are the one idea behind Kubernetes, Terraform, GitOps, and every operator. See it once and half of platform work collapses.
-
What a Container Actually Is: Namespaces, cgroups, and Layers
A container is not a lightweight VM. It is one ordinary process with three kernel restrictions applied. Knowing which three explains almost every container bug.
-
Observability and incident response — the SRE basics
A primer on observability (logs, metrics, traces) and incident response (roles, severities, blameless postmortems) — the disciplines every SRE team runs.
-
Toil and the 50% rule — what it is, how to measure it, and how to kill it
A primer on toil — the manual, repetitive work that eats SRE teams. Google's six-part definition, the 50% cap, and how AI agents change the playbook.
-
SLI, SLO, SLA, and error budgets — the reliability contract explained
SLIs, SLOs, SLAs, and error budgets — the four numbers every SRE team must agree on, the math behind 'nines,' and what changes when agents burn the budget.
-
What is Site Reliability Engineering (SRE)?
A primer on Site Reliability Engineering — where it came from, how it differs from DevOps and Platform Engineering, and what changes as AI joins on-call.
-
What are vector embeddings?
A primer on vector embeddings — how meaning becomes something you can search, cluster, and compare, and the failure modes you only see in evaluation.
-
What is function calling (tool use)?
A primer on function calling — the JSON-schema contract that lets an LLM invoke your code. The request/response loop, parallel calls, and forced tools.
-
What is prompt caching?
Prompt caching cuts repeated-prompt cost 50–90% and halves latency. How prefix matching works, TTL economics by provider, and what decides your hit rate.
-
The CAP theorem in AI-native distributed systems
CAP didn't get repealed when LLMs showed up. How the C/A/P trade-offs shift when the datastore is a vector index, context graph, or retrieval layer.
-
NAS vs SAN for GPU workloads — what changed when AI showed up
File vs block was the old NAS-vs-SAN question. GPU training rewrote it. How the calculus shifts when storage has to keep an H100 cluster fed.
-
What is an AI agent? A primer for cloud engineers
A primer on AI agents — the perceive-reason-act loop, what separates an agent from a one-shot LLM call, and the classical agent types SREs now operate.
-
What is Model Context Protocol (MCP)?
A primer on Model Context Protocol — the open standard that lets AI applications talk to tools through one interface. Hosts, clients, servers, transports.
-
What is Retrieval-Augmented Generation (RAG)?
A primer on Retrieval-Augmented Generation — grounding an LLM's answer in documents you trust. Indexing, serving, and the failure modes that bite.
-
Queues and Message Brokers: The Shock Absorber of Distributed Systems
A queue decouples producers from consumers and absorbs bursts. Backpressure, at-least-once delivery, idempotency, DLQs — now in front of every LLM call.
-
ACID, BASE, and Isolation Levels: What 'Consistent' Actually Means
ACID promises correctness; BASE trades it for scale — and your default isolation level is weaker than you think. A refresher on what 'consistent' means.
-
Graph Traversal: BFS, DFS, and Why GraphRAG Is Just a Walk
BFS or DFS — queue or stack — decides everything downstream. A refresher on graph traversal, the visited set, and why GraphRAG is just a walk.
-
Recursion and the Call Stack: Elegant Until It Overflows
Every recursive call costs a stack frame, and runaway depth ends in overflow. Frames, base cases, and why agent loops share the failure mode.
-
Garbage Collection: The Convenience That Shows Up in Your p99
GC frees you from manual memory management — and occasionally freezes your program. Reachability, generational GC, and why GC tuning is a p99 problem.
-
Compilers, Interpreters, and JIT: Why Python Is Slow and PyTorch Is Fast
Compiling ahead of time, interpreting, or JIT-compiling hot paths explains why Python is slow and PyTorch is fast — torch.compile is the same idea.
-
Floating Point and Numerical Precision: Why 0.1 + 0.2 ≠ 0.3, and Why ML Cares
Floating-point errors aren't random. Why 0.1 + 0.2 ≠ 0.3, and how the same fundamentals drive the FP32 → BF16 → FP8 march behind affordable LLMs.
-
Virtual Memory and Paging: The Illusion Every Program Lives In
Virtual memory is the beautiful lie every process lives in. Page faults, thrashing, overcommit — and why paging now manages the KV cache and OOM kills.
-
The Memory Hierarchy: Why Data Locality Beats Clock Speed
Each memory level is 10–100× slower than the one above. Cache lines, locality, and why 'keep data near compute' is the biggest lever in LLM inference.
-
Deadlocks, Locks, and Race Conditions: The Bugs That Only Happen Sometimes
Races and deadlocks only fail sometimes — that's what makes them hard. Locks, atomicity, the four Coffman conditions, and why hung agents are the same bug.
-
Concurrency vs Parallelism: The Distinction That Fixes Your Throughput
Concurrency is dealing with many things at once; parallelism is doing them at once. A refresher on the GIL, async vs threads, and scaling model calls.
-
Caching and Eviction Policies: Why LRU, LFU, and FIFO Aren't the Same Bet
The eviction policy decides whether your cache works. LRU vs LFU vs FIFO, hit rates, invalidation — and how the same bets govern prompt and KV caches.
-
Bloom Filters: Saying 'Definitely Not' in a Few Kilobytes
A Bloom filter answers 'have I seen this before?' in kilobytes, never returning a false negative. How it works, and why databases and caches rely on it.
-
Consistent Hashing: How Distributed Systems Add and Remove Nodes Without Chaos
Hash-mod-N reshuffles almost everything when a node leaves; consistent hashing moves a fraction. The ring, virtual nodes, and why your CDN depends on it.
-
TLS and Public-Key Cryptography, Explained Without the Math
Every https and model API call rides on TLS. How public-key crypto solves key distribution, what a certificate proves, and the cert-expiry realities.
-
How DNS Resolution Works — the Internet's Phone Book, and Its Cache
DNS caching decides how fast failovers propagate and how far a misconfig spreads. The recursive walk from root to authoritative, and why TTLs govern deploys.
-
How TCP/IP Actually Works — and Why the Handshake Still Bites You
Every API call and model request rides on TCP/IP. The three-way handshake, flow control, and why round-trips and TIME_WAIT shape your latency.
-
B-tree vs LSM-tree: The Two Ways Databases Store Your Data
Almost every database is a read-optimized B-tree or a write-optimized LSM-tree underneath. Knowing which explains its performance and failure modes.
-
Hash Tables: The Data Structure Behind Almost Everything
The hash table sits under your cache, index, dedup, and vector store metadata. How it turns a key into O(1) access — and what keeps it fast.
-
Big-O Notation in the Age of Billion-Vector Search
Big-O still decides whether your system survives real data. A refresher on complexity, and why it governs vector search, context windows, and outages.
-
GraphQL for API Development
GraphQL is a query language for APIs, created at Facebook in 2012. What it is, how it differs from REST, and where it fits in API development.
-
The CAP theorem
Consistency, availability, partition tolerance — you get two. A walk through the CAP trade-offs, database by database.
-
The 7 layers of ISO OSI model
The OSI model's seven layers give diverse systems a standard way to communicate. What each layer does and why the model still matters.