Observability
7 posts — newest first.
-
The EU AI Act blinked — your logging requirements didn't
The EU AI Act omnibus moved high-risk deadlines to 2027–28. The logging, oversight, and inventory work is still yours — and reliability needs it anyway.
-
Every observability vendor now sells an AI SRE agent. Here's how to evaluate one.
Datadog, Dynatrace, New Relic, AWS — every incumbent now ships an AI SRE agent. A field guide for evaluating one before it touches production.
-
Tracing the agent loop: OpenTelemetry's GenAI conventions, read like an SRE
Your agent is a distributed system wearing a chat interface. OpenTelemetry's GenAI conventions make it debuggable — what v1.41 covers and what's moving.
-
The AI-native SRE stack — a 2026 reference guide
A practitioner's map of the AI-native SRE stack in 2026: six layers from telemetry to bounded remediation, and an honest read on where AI pays off.
-
Observability for AI systems — what changes when your service calls an LLM
Golden signals miss the failure that pages you: a confident, well-formed, wrong answer. What AI observability adds — context as a span, quality as a signal.
-
Harness engineering: the third phase of AI maturity
Agent = Model + Harness, and in 2026 the harness is the bottleneck. What a production-grade SRE harness contains, with a ~40-line reference implementation.
-
Observability and incident response — the SRE basics
A primer on observability (logs, metrics, traces) and incident response (roles, severities, blameless postmortems) — the disciplines every SRE team runs.