diff --git a/_TODO.md b/_TODO.md index 69fbc093c..392bb4c05 100644 --- a/_TODO.md +++ b/_TODO.md @@ -250,14 +250,19 @@ https://mermaid.js.org/config/directives.html - We need to check for short form and deep article articles where the deep-dive index.pdf has a non-featured tag like "argo-cd" only in the pdf.mdx. In those cases, we should make sure the callout for the deep dive includes the name of that non-featured (technology) tag -- deep-dive/kubernetes-pod-resource-requests-limits-qos-classes is showing the Download hero image - - Add a "Preview Special" item to our Download CTA that lets the user know the Deep Dive content can be previewed in HTML format, and offer a switch to it. - How can we handle footnotes in List components? src/content/articles/kubernetes-multi-cluster-fleet-management-configuration/pdf.mdx line 80 - Need to update Case Studies with lists and tables too +17:15:36 [200] /articles/kubernetes-pod-disruption-budget-autoscaler-node-rotation 414ms +[WebMentions] Failed to fetch mentions for https://www.webstackbuilders.com/articles/kubernetes-pod-disruption-budget-autoscaler-node-rotation. Will retry in 60s. DOMException [TimeoutError]: The operation was aborted due to timeout + at node:internal/deps/undici/undici:15445:13 + at async requestWebmentions (/home/kevin/Repos/WebstackBuilders/CorporateWebsite/astro.webstackbuilders.com/src/components/WebMentions/server/index.ts:199:20) + at async eval (/home/kevin/Repos/WebstackBuilders/CorporateWebsite/astro.webstackbuilders.com/src/components/WebMentions/index.astro:27:18) +17:15:46 [200] /_vtbot_inspection_chamber.js 1ms + ### Snippets @@ -267,7 +272,7 @@ https://mermaid.js.org/config/directives.html @@ -332,3 +337,11 @@ https://mermaid.js.org/config/directives.html classes={{ tbody: "[&_tr>td:nth-of-type(2)]:!font-normal", }} + +PDBs require thinking about both directions — protecting from disruption __and__ allowing necessary operations. A PDB that blocks all disruptions doesn't protect your service — it protects it from getting security patches, upgrades, and capacity optimization. The goal is controlled disruption, not zero disruption. @@ -68,6 +68,7 @@ The eviction flow works like this: when a drain or autoscaler attempts to evict @@ -77,15 +78,19 @@ The eviction flow works like this: when a drain or autoscaler attempts to evict Choosing the wrong mode is one of the most common causes of autoscaler deadlocks. The two PDB modes look similar but behave differently as your replica count changes. Consider a 5-replica deployment: @@ -93,7 +98,8 @@ Choosing the wrong mode is one of the most common causes of autoscaler deadlocks The math gets interesting with percentages. `minAvailable: 80%` on 5 replicas means ceil(5 × 0.8) = 4 must stay up, allowing only 1 disruption. `maxUnavailable: 25%` means floor(5 × 0.25) = 1 can be down. Same result, but the behavior diverges as you scale.
Once you've identified the blocking PDB, you have several options depending on the situation: disruptionsAllowed is zero because pods are failing health checks, fix the health issue. Once pods are healthy, the PDB will allow disruptions again.', }, { - lead: 'Scale up the deployment.', - text: 'If you have `minAvailable: 3` but only 3 replicas and one is unhealthy, add a fourth replica. Once it\'s ready, you\'ll have headroom to drain.', + lead: 'Scale up the deployment', + text: 'If you have minAvailable: 3 but only 3 replicas and one is unhealthy, add a fourth replica. Once it\'s ready, you\'ll have headroom to drain.', }, { - lead: 'Temporarily relax the PDB.', + lead: 'Temporarily relax the PDB', text: 'For planned maintenance windows, you can patch the PDB to be less strict:', }, ]} @@ -270,17 +284,21 @@ kubectl patch pdb my-service-pdb -p '{"spec":{"minAvailable":2}}' Code: Temporarily relaxing a PDB for maintenance. kubectl delete pod --force --grace-period=0. This ignores PDBs completely — the pod just disappears.', }, ]} />
maxUnavailable: 1 doesn\'t prevent disruption — it just makes the eviction use the eviction API instead of direct deletion. This can be useful for visibility and for systems that watch eviction events, but it doesn\'t provide availability protection. If you need zero downtime, run more replicas.', }, { - lead: 'Single-replica deployments', - text: 'are a judgment call. A PDB with `maxUnavailable: 1` doesn\'t prevent disruption — it just makes the eviction use the eviction API instead of direct deletion. This can be useful for visibility and for systems that watch eviction events, but it doesn\'t provide availability protection. If you need zero downtime, run more replicas.', + lead: 'Batch jobs and CronJobs', + text: 'Generally shouldn\'t have PDBs. Jobs should be restartable by design, and a PDB on a job creates operational headaches without providing real protection.', }, ]} /> @@ -392,25 +410,30 @@ The right percentage depends on your tolerance for gaps. For logging agents like Several PDB configurations look reasonable but cause problems: minAvailable: 3, nothing can ever be evicted. This is the most common PDB misconfiguration.', }, { lead: 'Percentages that round badly.', - text: '`minAvailable: 90%` with 3 replicas means ceil(2.7) = 3 pods required — blocking all evictions. Use absolute numbers for small deployments.', + text: 'minAvailable: 90% with 3 replicas means ceil(2.7) = 3 pods required — blocking all evictions. Use absolute numbers for small deployments.', }, { lead: 'Multiple overlapping PDBs.', - text: 'If a pod matches two PDBs, both must allow the disruption. This is stricter than either PDB alone and often catches teams by surprise. For example: Team A creates a PDB for `app: payments` allowing 1 disruption, Team B creates a PDB for `tier: critical` allowing 2 disruptions. Pods with both labels need both PDBs to allow eviction simultaneously — effectively the stricter of the two.', + text: 'If a pod matches two PDBs, both must allow the disruption. This is stricter than either PDB alone and often catches teams by surprise. For example: Team A creates a PDB for app: payments allowing 1 disruption, Team B creates a PDB for tier: critical allowing 2 disruptions. Pods with both labels need both PDBs to allow eviction simultaneously — effectively the stricter of the two.', }, ]} /> @@ -428,19 +451,22 @@ PDBs fail silently. A misconfigured PDB doesn't cause an immediate outage — it The kube-state-metrics project exposes PDB status as Prometheus metrics (ensure kube-state-metrics is deployed and scraped by Prometheus — it's included in most monitoring stacks like kube-prometheus-stack). The key metrics to watch: kube_poddisruptionbudget_status_pod_disruptions_allowed', + text: 'How many pods can currently be evicted', }, { - lead: 'kube_poddisruptionbudget_status_current_healthy', - text: 'pods currently passing health checks', + lead: 'kube_poddisruptionbudget_status_current_healthy', + text: 'Pods currently passing health checks', }, { - lead: 'kube_poddisruptionbudget_status_desired_healthy', - text: 'minimum required by the PDB', + lead: 'kube_poddisruptionbudget_status_desired_healthy', + text: 'Minimum required by the PDB', }, ]} /> @@ -488,7 +514,12 @@ The `PDBViolated` alert is more urgent — it means you're already below your av Use these thresholds to assess PDB health at a glance:
disruptionsAllowed', td: ['> 20% of pods', '1-20% of pods', '0'], }, { - th: 'currentHealthy vs desired', - td: ['Equal', '-', 'Below'], + th: 'currentHealthy vs desired', + td: ['Equal', '---', 'Below'], }, { th: 'Blocking duration', @@ -529,8 +560,8 @@ PDBs are contracts between workload owners and platform operators. The goal is c Configure PDBs to allow at least enough disruptions for single-node drains — that's the minimum for cluster operations to work. Monitor for PDBs that block disruptions, and have clear procedures for resolving stuck drains when they happen. -A cluster that can't be maintained isn't a reliable cluster. PDBs that block security patches, version upgrades, and capacity optimization aren't protecting availability — they're trading one kind of risk for another. The best PDB configuration is invisible: it protects your workloads during operations without anyone noticing. That 6 AM page about stuck nodes? With proper PDB configuration, it doesn't happen. - If your platform team is constantly battling PDBs during node rotations, the PDBs are misconfigured — not too loose, but too strict. + +A cluster that can't be maintained isn't a reliable cluster. PDBs that block security patches, version upgrades, and capacity optimization aren't protecting availability — they're trading one kind of risk for another. The best PDB configuration is invisible: it protects your workloads during operations without anyone noticing. That 6 AM page about stuck nodes? With proper PDB configuration, it doesn't happen. diff --git a/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/index.mdx b/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/index.mdx index 90ae73a99..530288abd 100644 --- a/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/index.mdx +++ b/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/index.mdx @@ -16,28 +16,33 @@ featured: true It's 3 AM and your pager goes off. A production cluster with 10 nodes is experiencing random pod evictions. Investigation reveals memory pressure on several nodes—60% of pods have no memory limits, so a few memory-hungry pods consumed all available RAM. But here's the twist: the pods being evicted aren't the resource hogs. They're the well-behaved ones that set requests but no limits. -Those pods followed what seemed like best practice — they declared what they needed. Yet Kubernetes marked them as "Burstable" QoS, and the eviction algorithm measures Burstable pods against their own declared requests. The pods with no resource specs? They're "BestEffort"—technically lower priority, but with no declared baseline, there's nothing to measure them against. The well-behaved pods created a measuring stick that was used against them. +Those pods followed what seemed like best practice — they declared what they needed. Yet Kubernetes marked them as "Burstable" QoS, and the eviction algorithm measures Burstable pods against their own declared requests. The pods with no resource specs? They're "BestEffort" — technically lower priority, but with no declared baseline, there's nothing to measure them against. The well-behaved pods created a measuring stick that was used against them. -The lesson most teams learn too late: QoS class determines eviction order, and you don't realize what class your pods belong to until you're debugging an outage. The answer lies in a system most teams don't know exists. +The lesson most teams learn too late: QoS class determines eviction order, and you don't realize what class your pods belong to until you're debugging an outage. The answer lies in a system most teams don't know exists. ## The QoS Contract You Didn't Know You Signed Every pod in Kubernetes gets assigned a Quality of Service class automatically. You don't set it directly — it's derived from how you configure requests and limits. This class determines who dies first when nodes run low on resources. @@ -45,7 +50,8 @@ Every pod in Kubernetes gets assigned a Quality of Service class automatically. The rules are straightforward, but the implications aren't. A pod with carefully tuned requests and limits is Burstable. A pod with only a CPU request and nothing else is also Burstable. Same QoS class, very different behavior.
@@ -152,7 +161,7 @@ spec: Code: Burstable QoS with headroom — limits higher than requests for burst capacity. -For variable workloads like web servers, use Burstable with headroom. Set limits higher than requests to allow bursting during traffic spikes while keeping baseline reservation efficient. One gotcha: for multi-container pods, the QoS is determined by __all__ containers. If your main container has requests equal to limits but your sidecar doesn't, the whole pod is Burstable. Size sidecars explicitly. +For variable workloads like web servers, use Burstable with headroom. Set limits higher than requests to allow bursting during traffic spikes while keeping baseline reservation efficient. One gotcha: for multi-container pods, the QoS is determined by _all_ containers. If your main container has requests equal to limits but your sidecar doesn't, the whole pod is Burstable. Size sidecars explicitly. The rule of thumb: set CPU limits 2-4x requests (allows bursting), set memory limits 1.5-2x requests (headroom without waste). Monitor actual usage for two weeks, then right-size based on P95 metrics. @@ -163,7 +172,10 @@ The rule of thumb: set CPU limits 2-4x requests (allows bursting), set memory li If you do nothing else, do these three things: kubectl get pods -o custom-columns="NAME:.metadata.name, QOS:.status.qosClass" across your namespaces. If you see BestEffort on anything that matters, you have work to do.', }, { lead: 'Size memory limits with 50% headroom above peak.', diff --git a/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/pdf.mdx b/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/pdf.mdx index 9c8a9f7b7..86b3211a8 100644 --- a/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/pdf.mdx +++ b/src/content/articles/kubernetes-pod-resource-requests-limits-qos-classes/pdf.mdx @@ -1,7 +1,7 @@ --- title: "Pod Sizing: Requests, Limits, and QoS Classes" description: "How Kubernetes scheduling and eviction actually work, and how to size pods to survive node pressure." -cover: "./download.jpg" +cover: "./cover.jpg" coverAlt: "Airplane seating classes showing first class (Guaranteed), business (Burstable), and economy (BestEffort) with priority boarding representing Kubernetes QoS resource tiering" author: "kevin-brown" publishDate: 2022-10-02 @@ -20,16 +20,16 @@ import resourceSizingDiagram from "./diagrams/resource-sizing-decision-tree.jpg" *[RSS]: Resident Set Size *[VPA]: Vertical Pod Autoscaler -Every pod in Kubernetes consumes CPU and memory, but __how__ you declare those resources determines where pods land, how they behave under pressure, and which pods die when nodes run low. Requests and limits aren't optional settings you can figure out later — they're the contract between your workload and the cluster. Get them wrong and you'll face wasted capacity (over-requesting), noisy neighbor problems (under-limiting), or surprise evictions (mismatched QoS classes). +Every pod in Kubernetes consumes CPU and memory, but _how_ you declare those resources determines where pods land, how they behave under pressure, and which pods die when nodes run low. Requests and limits aren't optional settings you can figure out later — they're the contract between your workload and the cluster. Get them wrong and you'll face wasted capacity (over-requesting), noisy neighbor problems (under-limiting), or surprise evictions (mismatched QoS classes). Here's a scenario I've seen multiple times: a production cluster with 10 nodes starts experiencing random pod evictions at 3 AM. Investigation reveals memory pressure on several nodes—60% of pods have no memory limits, so a few memory-hungry pods consumed all available RAM. The pods being evicted aren't the resource hogs. They're the well-behaved ones that set requests but no limits, making them "Burstable" QoS and first in line for eviction. -The lesson: QoS class determines eviction order, and most teams don't realize what class their pods belong to until they're debugging an outage. - The most common resource mistake: setting requests without limits (or vice versa). This creates Burstable QoS pods that can be evicted before pods with no resource specs at all. Understanding QoS classes is essential for production workloads. +The lesson: QoS class determines eviction order, and most teams don't realize what class their pods belong to until they're debugging an outage. + ## The Resource Model ### Requests vs Limits @@ -56,7 +56,7 @@ spec: Code: Basic resource specification with requests and limits. -The key insight: requests affect __where__ pods run, limits affect __how__ pods run once scheduled. The scheduler sums up all requests on a node and won't over-commit beyond allocatable capacity. But limits are enforced at runtime by the kernel — a pod can burst above its request (using spare capacity) until it hits its limit. This split creates flexibility: you can request 256Mi (your baseline need) but set a 512Mi limit (your peak need), letting pods burst when memory is available without reserving peak capacity on every node. +The key insight: requests affect _where_ pods run, limits affect _how_ pods run once scheduled. The scheduler sums up all requests on a node and won't over-commit beyond allocatable capacity. But limits are enforced at runtime by the kernel — a pod can burst above its request (using spare capacity) until it hits its limit. This split creates flexibility: you can request 256Mi (your baseline need) but set a 512Mi limit (your peak need), letting pods burst when memory is available without reserving peak capacity on every node. ### CPU Behavior @@ -97,7 +97,7 @@ Memory limits without headroom are time bombs. If your application normally uses The eviction angle matters too: pods without memory limits can consume all node memory, triggering node-level OOM that affects __all__ pods on the node. This is why "no limits" is dangerous — one runaway pod can take down its neighbors.
@@ -190,6 +193,8 @@ spec: Code: Burstable QoS configuration (requests differ from limits). +BestEffort is what you get when a pod spec omits resource declarations entirely. Kubernetes treats these pods as disposable — they receive whatever capacity happens to be available, with no scheduling guarantees and no eviction protection. In practice, BestEffort pods are appropriate for truly ephemeral workloads: dev/test environments, batch jobs that checkpoint their progress, or background tasks where occasional restarts are acceptable. Running production services as BestEffort is dangerous because the kubelet will terminate them before touching any pod that declared even a single resource request. + ```yaml title="qos-besteffort.yaml" # BestEffort QoS: no resources specified at all apiVersion: v1 @@ -206,13 +211,17 @@ spec: Code: BestEffort QoS configuration (no resource specifications). +The QoS class assignment is entirely mechanical — there's no annotation or field you can set to override it. Kubernetes inspects every container in the pod (including init containers for resource accounting) and applies a strict decision tree. This means a multi-container pod where one sidecar is missing its limits will drop the entire pod from Guaranteed to Burstable, even if the primary container is perfectly configured. Auditing QoS classes across a namespace with `kubectl get pods -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass` is worth doing periodically, especially after adding sidecars or updating Helm charts that modify resource blocks. + -You can check a pod's QoS class with `kubectl get pod -o yaml | grep qosClass`. If you're surprised by the result, remember that __all__ containers in the pod must meet the Guaranteed criteria — one container without limits drops the whole pod to Burstable. +You can check a pod's QoS class with `kubectl get pod -o yaml | grep qosClass`. If you're surprised by the result, remember that _all_ containers in the pod must meet the Guaranteed criteria — one container without limits drops the whole pod to Burstable. ### Eviction Behavior @@ -236,7 +245,7 @@ kubectl get events --field-selector reason=Evicted Code: Diagnosing eviction events.
@@ -492,7 +507,8 @@ Code: Checking VPA recommendations. The VPA output includes several values: `lowerBound` (minimum safe), `target` (recommended for normal operation), and `upperBound` (recommended for spikes). A practical approach: use `target` for requests and `upperBound` with a buffer for limits.
Resources are a contract. Requests are your promise of what you need. Limits are your promise of what you'll never exceed. QoS class is how Kubernetes prioritizes that contract when the cluster is under pressure. Write good contracts. + +The goal isn't perfect efficiency — it's the right tradeoff between reliability and cost. Critical services get headroom. Batch jobs run lean. Everything gets measured. diff --git a/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/index.mdx b/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/index.mdx index 8b52caa82..971560a50 100644 --- a/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/index.mdx +++ b/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/index.mdx @@ -18,16 +18,17 @@ import patternDecisionDiagram from "./diagrams/pattern-decision-tree.jpg" It's 3 AM and Vault is down. Your on-call engineer gets paged because deployments are failing — pods stuck in ContainerCreating, blocking a critical hotfix. Meanwhile, another team's services keep humming along despite the same outage. The difference isn't luck. It's how secrets get into pods. -Both teams use Vault. Both followed the documentation. But one team chose External Secrets Operator, which syncs secrets periodically and caches them as native Kubernetes secrets. The other chose the Secrets Store CSI Driver, which fetches secrets on-demand when pods start. When Vault went down, ESO's cached secrets kept working. CSI's synchronous fetches failed, and pods couldn't start. This isn't about which tool is better — it's about understanding the failure mode you've chosen __before__ it matters. +Both teams use Vault. Both followed the documentation. But one team chose External Secrets Operator, which syncs secrets periodically and caches them as native Kubernetes secrets. The other chose the Secrets Store CSI Driver, which fetches secrets on-demand when pods start. When Vault went down, ESO's cached secrets kept working. CSI's synchronous fetches failed, and pods couldn't start. This isn't about which tool is better — it's about understanding the failure mode you've chosen _before_ it matters. ## The Two Patterns That Matter -External secrets management in Kubernetes has consolidated around two dominant patterns. External Secrets Operator runs as a controller in your cluster, periodically syncing secrets from Vault (or AWS Secrets Manager, Azure Key Vault, etc.) into native Kubernetes Secret objects. The Secrets Store CSI Driver takes a different approach: it mounts secrets directly into pods as volumes, fetching them from the external manager when pods start. +External secrets management in Kubernetes has consolidated around two dominant patterns. External Secrets Operator runs as a controller in your cluster, periodically syncing secrets from Vault (or AWS Secrets Manager, Azure Key Vault, etc.) into native Kubernetes Secret objects. The Secrets Store CSI Driver takes a different approach: it mounts secrets directly into pods as volumes, fetching them from the external manager when pods start. Both work fine when your secret manager is healthy. The difference is what happens when it isn't. ESO decouples secret fetching from pod lifecycle — the controller syncs independently, and pods consume cached Kubernetes Secrets. CSI couples them tightly — pods can't start until secrets are fetched. This architectural difference determines everything about how your applications behave during an outage.
@@ -130,8 +135,8 @@ For most organizations, ESO is the right default. It's operationally simpler, Gi The mistake isn't choosing either pattern. It's not understanding which failure mode you've chosen. The team that slept through the 3 AM Vault outage didn't get lucky — they understood that ESO's cached secrets would keep their services running. The team that got paged made a valid choice too; for their payment system, blocking on fresh credentials was the right call. Their runbooks reflected it. -Whatever you choose, document it. When the next outage happens, your incident responders shouldn't be learning your secret injection architecture for the first time. - Start with ESO for the majority of workloads, with a 15-minute refresh interval, and monitor sync status with alerts on stale ExternalSecrets. Add CSI for specific high-security applications where staleness is unacceptable. This gives you operational simplicity with escape hatches for edge cases. + +Whatever you choose, document it. When the next outage happens, your incident responders shouldn't be learning your secret injection architecture for the first time. diff --git a/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/pdf.mdx b/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/pdf.mdx index 32eee2d3e..4cbce7e95 100644 --- a/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/pdf.mdx +++ b/src/content/articles/kubernetes-secrets-external-secrets-operator-csi-vault/pdf.mdx @@ -29,24 +29,27 @@ import threeSecretDiagram from "./diagrams/three-secret-injection-patterns.jpg" It's 3 AM and Vault is down. Your on-call engineer gets paged because deployments are failing — pods stuck in ContainerCreating, blocking a critical hotfix. Meanwhile, another team's services keep humming along despite the same outage. The difference isn't luck. It's how secrets get into pods. -Kubernetes secrets have a fundamental problem: they're base64-encoded, not encrypted. They sit in etcd alongside your cluster state, readable by anyone with RBAC access to the namespace. External secret managers like Vault, AWS Secrets Manager, and Azure Key Vault solve the security problem by keeping secrets outside the cluster — but they create a new problem. Your pods now depend on an external service, and that dependency has failure modes you need to understand __before__ the 3 AM page. +Kubernetes secrets have a fundamental problem: they're base64-encoded, not encrypted. They sit in etcd alongside your cluster state, readable by anyone with RBAC access to the namespace. External secret managers like Vault, AWS Secrets Manager, and Azure Key Vault solve the security problem by keeping secrets outside the cluster — but they create a new problem. Your pods now depend on an external service, and that dependency has failure modes you need to understand _before_ the 3 AM page. Three patterns dominate secret injection: @@ -55,6 +58,7 @@ Each pattern fails differently when the secret manager goes away. Consider a 30-
@@ -164,7 +170,11 @@ Even with encryption at rest, native secrets have operational limitations. There External Secrets Operator is a Kubernetes operator that syncs secrets from external managers — Vault, AWS Secrets Manager, Azure Key Vault, GCP Secret Manager — into native Kubernetes secrets. The operator runs in your cluster, periodically fetches secrets from the external source, and creates or updates Kubernetes Secret objects that your pods consume normally.
@@ -472,7 +483,7 @@ spec: ``` Code: Pod with CSI secret volume. -The pod __cannot start__ until the volume mounts successfully — which means until secrets are fetched from Vault. This is the key difference from ESO: the dependency is synchronous. +The pod _cannot start_ until the volume mounts successfully — which means until secrets are fetched from Vault. This is the key difference from ESO: the dependency is synchronous. The CSI driver optionally supports rotation. Enable it by setting `--enable-secret-rotation=true` on the driver DaemonSet and adding `rotationPollInterval` to your SecretProviderClass. When enabled, the driver periodically re-fetches secrets and updates the mounted files. Your application must detect file changes (using inotify or periodic re-reads) to pick up rotated values — the driver updates files, but it can't restart your process. @@ -480,6 +491,7 @@ The flow differs fundamentally from ESO: secrets are fetched synchronously durin @@ -687,7 +699,8 @@ The __init container__ pattern makes sense when you need behavior that ESO and C With three patterns to choose from, the decision comes down to two questions: how fresh do your secrets need to be, and what failure mode can your application tolerate?
+The table captures raw capabilities, but choosing a pattern requires weighting those capabilities against your operational reality. A team with strong GitOps practices and tolerance for minutes of staleness will land in a different place than a team running payment infrastructure where a stale database credential means failed transactions. The deciding factor is usually the failure mode: ESO fails silently with stale data, CSI fails loudly by blocking pod startup, and init containers fail however you code them to. Each is the right answer for different risk profiles. + @@ -759,7 +775,10 @@ With three patterns to choose from, the decision comes down to two questions: ho ### Recommendations by Use Case Two techniques unlock almost any legacy codebase: __characterization tests__ and __seam identification__. You don't need to understand the code to test it, and you don't need to refactor before you can write your first test. These techniques form the foundation that makes everything else possible. ## Characterization Tests: Document Before You Judge -Traditional unit tests verify that code does what it __should__ do — you write a test based on a specification, and the test fails if the code doesn't match. Characterization tests flip this: they capture what the code __actually__ does, regardless of intent. You're not testing against a spec; you're documenting observed behavior. +Traditional unit tests verify that code does what it _should_ do — you write a test based on a specification, and the test fails if the code doesn't match. Characterization tests flip this: they capture what the code _actually_ does, regardless of intent. You're not testing against a spec; you're documenting observed behavior. The distinction matters for legacy code. You don't have a spec. The original authors are gone. The code has undocumented edge cases, implicit business rules buried in conditionals, and behaviors that might be bugs or might be features — you can't tell. Characterization tests don't try to answer "is this correct?" They answer "what does this do?" and lock it in. @@ -66,18 +66,20 @@ Consider a method that sends emails. The email-sending code is deep inside a 500 ### The Four Seam Types
initialize'], }, { th: 'C# / Java', @@ -97,40 +99,40 @@ Consider a method that sends emails. The email-sending code is deep inside a 500 }, { th: 'TypeScript', - td: ['Optional parameters with nullish coalescing (`??`)'], + td: ['Optional parameters with nullish coalescing (??)'], }, { th: 'Spring Boot', - td: ['`@Autowired` with test `@Configuration` bindings'], + td: ['@Autowired with test @Configuration bindings'], }, { th: 'Laravel', - td: ['Service container with `$this->app->bind()` in tests'], + td: ['Service container with $this->app->bind() in tests'], }, { th: '.NET', - td: ['`IServiceCollection` with test service registration'], + td: ['IServiceCollection with test service registration'], }, ], }, - figure: 'Object seam techniques by language and framework.', }} /> stub_const to replace a class entirely. In TypeScript/JavaScript, Jest\'s jest.mock() intercepts imports. The enabling point is the test setup. Link seams are powerful but fragile — they couple tests to implementation details like class names.', }, { lead: 'Subclass seams', - text: 'work by extracting behavior into a protected method, then overriding it in a test subclass. This technique is underrated for legacy code because it requires minimal changes — you extract one line into a method, and suddenly you have a seam.', + text: 'Works by extracting behavior into a protected method, then overriding it in a test subclass. This technique is underrated for legacy code because it requires minimal changes — you extract one line into a method, and suddenly you have a seam.', }, { lead: 'Preprocessor seams', - text: 'apply anywhere you use environment-based branching. Rails\' `Rails.env.test?`, Laravel\'s `app()->environment(\'testing\')`, and Node\'s `process.env.NODE_ENV === \'test\'` are all effectively preprocessor seams. Use them sparingly — they litter production code with test concerns.', + text: 'Applies anywhere you use environment-based branching. Rails\' Rails.env.test?, Laravel\'s app()->environment(\'testing\'), and Node\'s process.env.NODE_ENV === \'test\' are all effectively preprocessor seams. Use them sparingly — they litter production code with test concerns.', }, ]} /> @@ -178,8 +180,6 @@ The workflow looks like this: First, identify the behavior you need to protect. With both in place, you can isolate and test without understanding the full system. The code is no longer untestable — it's testable through observation and substitution. -This is the foundation. Deeper techniques — Extract and Override for quick dependency breaking, Parameterize Constructor for clean DI patterns, Strangler Fig for system-level migration — all build on characterization tests and seams. But start here. Get your first characterization test passing. Find your first seam. The rest follows. - +This is the foundation. Deeper techniques — Extract and Override for quick dependency breaking, Parameterize Constructor for clean DI patterns, Strangler Fig for system-level migration — all build on characterization tests and seams. But start here. Get your first characterization test passing. Find your first seam. The rest follows. + ## Conclusion The myth of "untestable" code usually means "code that's hard to test with conventional techniques." Characterization tests and seams change the equation entirely — they let you observe, document, and isolate without first having to understand every line. Start with characterization tests. Run the code, capture what happens, lock it down. Don't judge whether the behavior is correct — just document it. Then find seams: the constructor parameters, the class methods, the environment flags that let you substitute behavior without editing the code you're protecting. -These foundations enable everything else: dependency breaking, incremental extraction, system-level migration. But they're also sufficient on their own to turn "untestable" into testable. The question isn't __can__ you test legacy code — it's whether the investment is worth it for code that may never change. +These foundations enable everything else: dependency breaking, incremental extraction, system-level migration. But they're also sufficient on their own to turn "untestable" into testable. The question isn't _can_ you test legacy code — it's whether the investment is worth it for code that may never change. diff --git a/src/content/articles/legacy-code-testing-characterization-tests-seams/pdf.mdx b/src/content/articles/legacy-code-testing-characterization-tests-seams/pdf.mdx index 9fd07de3a..f426f1d6b 100644 --- a/src/content/articles/legacy-code-testing-characterization-tests-seams/pdf.mdx +++ b/src/content/articles/legacy-code-testing-characterization-tests-seams/pdf.mdx @@ -34,7 +34,7 @@ The goal isn't 100% coverage — it's getting enough safety net to make changes ## Characterization Tests -Traditional unit tests verify that code does what it __should__ do — you write a test based on a specification, and the test fails if the code doesn't match. Characterization tests flip this: they capture what the code __actually__ does, regardless of intent. You're not testing against a spec; you're documenting observed behavior. +Traditional unit tests verify that code does what it _should_ do — you write a test based on a specification, and the test fails if the code doesn't match. Characterization tests flip this: they capture what the code _actually_ does, regardless of intent. You're not testing against a spec; you're documenting observed behavior. The distinction matters for legacy code. You don't have a spec. The original authors are gone. The code has undocumented edge cases, implicit business rules buried in conditionals, and behaviors that might be bugs or might be features — you can't tell. Characterization tests don't try to answer "is this correct?" They answer "what does this do?" and lock it in. @@ -284,7 +284,8 @@ Code: Types of seams. The __object seam__ is the most common and usually the cleanest. You pass a dependency through a constructor or method parameter, and the enabling point is the call site where you can pass a different implementation. Ruby makes this trivially easy with default arguments — existing callers get production behavior, tests pass doubles. The same pattern works across many different frameworks.
initialize'], }, { th: 'C# / Java', @@ -305,19 +306,19 @@ The __object seam__ is the most common and usually the cleanest. You pass a depe }, { th: 'TypeScript', - td: ['Optional parameters with nullish coalescing (`??`)'], + td: ['Optional parameters with nullish coalescing (??)'], }, { th: 'Spring Boot', - td: ['`@Autowired` with test `@Configuration` bindings'], + td: ['@Autowired with test @Configuration bindings'], }, { th: 'Laravel', - td: ['Service container with `$this->app->bind()` in tests'], + td: ['Service container with $this->app->bind() in tests'], }, { th: '.NET', - td: ['`IServiceCollection` with test service registration'], + td: ['IServiceCollection with test service registration'], }, ], }, @@ -398,12 +399,14 @@ end ``` Code: Identifying seams in legacy code. -The sensing variable pattern deserves special attention. Sometimes you can't easily intercept a dependency, but you __can__ add a field that captures what happened. It's ugly — test-specific code in production — but it's a temporary scaffold. Add the sensing variable, write your characterization tests, then refactor toward proper seams and remove the sensing code. The tests survive because they now use better seams. +The sensing variable pattern deserves special attention. Sometimes you can't easily intercept a dependency, but you _can_ add a field that captures what happened. It's ugly — test-specific code in production — but it's a temporary scaffold. Add the sensing variable, write your characterization tests, then refactor toward proper seams and remove the sensing code. The tests survive because they now use better seams. -When choosing which seam type to use, prefer object seams for long-term maintainability. They make dependencies explicit and support proper dependency injection. But when you need tests __now__ and can't change constructor signatures (maybe there are 50 call sites), link seams or subclass seams get you there faster. You can always refactor later — once you have tests. +When choosing which seam type to use, prefer object seams for long-term maintainability. They make dependencies explicit and support proper dependency injection. But when you need tests _now_ and can't change constructor signatures (maybe there are 50 call sites), link seams or subclass seams get you there faster. You can always refactor later — once you have tests. @@ -650,7 +653,11 @@ Code: Instance delegator technique. The alternative approach — injecting the validator classes themselves rather than wrapping their calls — produces cleaner code long-term. You can pass mock classes in tests that respond to `.validate` however you need. This is effectively Parameterize Constructor applied to class method dependencies.
@@ -1214,10 +1223,10 @@ Find seams where you can alter behavior without modifying code. Object seams let Break dependencies incrementally. Extract and Override gets you started with minimal change. Parameterize Constructor creates explicit dependency graphs. Instance Delegator handles static cling. Each technique trades design improvement against change risk — choose based on your context. -Use the Strangler Fig pattern for larger extractions. Wrap legacy in a facade, extract one responsibility at a time into tested components, gradually hollow out the legacy code until it's gone. Every step is deployable. No big-bang rewrites. - -Prioritize ruthlessly. Test code that's changing, code that's breaking, code that's critical. Leave stable legacy code alone — characterization tests catch accidental changes, but comprehensive testing of code that won't change is waste. - The key insight: you don't need to understand legacy code to test it. Characterization tests capture behavior you can observe. Seams let you isolate without understanding. The myth of "untestable" code usually means "code that's hard to test with conventional techniques." With these patterns, almost any code becomes testable — the question is whether the investment is worth it for code that may never change. + +Use the Strangler Fig pattern for larger extractions. Wrap legacy in a facade, extract one responsibility at a time into tested components, gradually hollow out the legacy code until it's gone. Every step is deployable. No big-bang rewrites. + +Prioritize ruthlessly. Test code that's changing, code that's breaking, code that's critical. Leave stable legacy code alone — characterization tests catch accidental changes, but comprehensive testing of code that won't change is waste. diff --git a/src/content/articles/monorepo-affected-builds-remote-caching-ci-optimization/index.mdx b/src/content/articles/monorepo-affected-builds-remote-caching-ci-optimization/index.mdx index e09957741..962f41b16 100644 --- a/src/content/articles/monorepo-affected-builds-remote-caching-ci-optimization/index.mdx +++ b/src/content/articles/monorepo-affected-builds-remote-caching-ci-optimization/index.mdx @@ -15,7 +15,7 @@ The instinct is to throw hardware at it — faster runners, more parallelism, bi The real solution is building less. -Two techniques make this possible. __Affected-based builds__ analyze the dependency graph to identify which packages need to rebuild when specific files change. __Remote caching__ stores build outputs so identical work never runs twice, regardless of which developer or CI runner needs it. Together, they transform CI from a bottleneck into a fast feedback loop — turning that 45-minute build into a 4-minute one. +Two techniques make this possible. __Affected-based builds__ analyze the dependency graph to identify which packages need to rebuild when specific files change. __Remote caching__ stores build outputs so identical work never runs twice, regardless of which developer or CI runner needs it. Together, they transform CI from a bottleneck into a fast feedback loop — turning that 45-minute build into a 4-minute one. The biggest monorepo CI mistake: trying to make full builds faster instead of building less. Parallelization and faster machines provide linear improvements. Affected builds with caching provide order-of-magnitude improvements. @@ -32,19 +32,22 @@ The impact of a change depends on where it lands in the graph. Change a widely-u The key insight for calculating affected packages is that you need to __reverse__ the dependency graph. Instead of asking "what does this package depend on," you ask "what depends on this package." The algorithm is straightforward: @@ -58,8 +61,13 @@ Affected calculation compares the current state against a __base reference__ — A PR that branched from main two weeks ago will have many more affected packages than one that branched yesterday, simply because main has moved. Long-lived feature branches accumulate affected packages. This is one reason teams prefer short-lived branches and frequent rebasing — it keeps the affected set small. origin/main', 'All changes in the PR'], }, { th: 'Push to main', - td: ['`HEAD~1` or last successful CI', 'Just the pushed commit(s)'], + td: ['HEAD~1 or last successful CI', 'Just the pushed commit(s)'], }, { th: 'Release build', @@ -83,7 +91,6 @@ A PR that branched from main two weeks ago will have many more affected packages }, ], }, - figure: 'Base reference strategies by CI scenario.', }} /> @@ -91,12 +98,12 @@ A PR that branched from main two weeks ago will have many more affected packages Affected builds reduce what needs to run, but remote caching eliminates redundant work entirely. The idea is simple: if someone already built a package with identical inputs, download their output instead of rebuilding. This works across developers, CI runners, and even different branches — anyone who's built the same code contributes to and benefits from the shared cache. + + A build cache works by hashing all inputs to a task — source files, configuration, dependency outputs, environment variables, runtime versions — into a single cache key. Before executing a task, the build system checks whether outputs for that cache key exist. Cache hit means download and skip; cache miss means run and upload. The cache key composition matters. It must include everything that affects the output: task name, package name, input file hashes, dependency output hashes, relevant environment variables, runtime versions, and command arguments. Miss any of these, and you risk cache poisoning — returning outputs that don't match what a fresh build would produce. Include too much, and you get unnecessary cache misses. - - ### Setting Up Remote Caching Both Nx and Turborepo offer straightforward remote caching setup. @@ -125,7 +132,8 @@ Code: Turborepo remote cache with signature verification. Both tools offer self-hosted options for organizations that can't use external services. Nx Cloud supports Docker/Kubernetes deployment with S3, Azure Blob, or GCS backends. Turborepo works with the community-maintained `ducktors/turborepo-remote-cache` Docker image.
The biggest monorepo CI mistake: trying to make full builds faster instead of building less. Parallelization and faster machines provide linear improvements. Affected builds with caching provide order-of-magnitude improvements. +The impact is dramatic. Cache hit rates reach 85%. Build times drop by an order of magnitude. But the improvement isn't just faster — it changes how developers work. They run CI before grabbing coffee, not before leaving for the day. They iterate in small increments instead of batching changes to minimize CI waits. The monorepo becomes an asset again. + ## Understanding Dependency Graphs Affected builds work by analyzing the dependency graph — the directed acyclic graph that captures which packages depend on which other packages. When a file changes, the build system maps that file to its package, then walks the graph to find everything that depends on that package, transitively. Understanding this graph is essential for reasoning about what will rebuild and why. @@ -123,8 +123,13 @@ For PR builds, the natural base is the target branch (usually `main`). This capt The base reference choice has practical implications. A PR that branched from main two weeks ago will have many more affected packages than one that branched yesterday, simply because main has moved. Long-lived feature branches accumulate affected packages. This is one reason teams prefer short-lived branches and frequent rebasing — it keeps the affected set small.
origin/main', 'All changes in the PR'], }, { th: 'Push to main', - td: ['`HEAD~1` or last successful CI', 'Just the pushed commit(s)'], + td: ['HEAD~1 or last successful CI', 'Just the pushed commit(s)'], }, { th: 'Release build', @@ -148,7 +153,6 @@ The base reference choice has practical implications. A PR that branched from ma }, ], }, - figure: 'Base reference strategies by CI scenario.', }} /> @@ -286,25 +290,28 @@ A build cache works by hashing all inputs to a task — source files, configurat The cache key composition matters. It must include everything that affects the output: @@ -414,7 +421,8 @@ Code: Turborepo self-hosted cache configuration. The choice between managed and self-hosted caching comes down to operational overhead versus control. Managed services (Nx Cloud, Vercel) handle infrastructure, scaling, and availability — you just configure a token. Self-hosted options require maintaining servers and storage, but keep all build artifacts within your infrastructure and avoid per-seat pricing at scale.
@@ -603,23 +612,26 @@ Remote caching is only useful if the cache actually gets hit. A cache that misse Cache invalidation happens automatically when inputs change — that's the whole point. But understanding the different invalidation triggers helps you configure inputs correctly and debug unexpected misses. package.json or change the lock file, packages that depend on that change need to rebuild. Both Nx and Turborepo handle this automatically through their global dependencies configuration.', }, { lead: 'Manual invalidation', - text: 'is occasionally necessary when something outside the tracked inputs changes — a CI environment update, a bug in the caching system, or when you need to verify that builds still work without caching. In Nx, `npx nx reset` clears the local cache. Adding `--skip-nx-cache` bypasses caching for a single run. In Turborepo, `--force` achieves the same.', + text: 'Occasionally necessary when something outside the tracked inputs changes — a CI environment update, a bug in the caching system, or when you need to verify that builds still work without caching. In Nx, npx nx reset clears the local cache. Adding --skip-nx-cache bypasses caching for a single run. In Turborepo, --force achieves the same.', }, { lead: 'Time-based expiration (TTL)', - text: 'prevents caches from growing unboundedly and ensures stale artifacts eventually get cleaned up. Most remote cache providers let you configure retention periods—7 to 30 days is typical for active projects.', + text: 'Prevents caches from growing unboundedly and ensures stale artifacts eventually get cleaned up. Most remote cache providers let you configure retention periods—7 to 30 days is typical for active projects.', }, ]} /> @@ -649,23 +661,26 @@ Code: Cache debugging and manual invalidation commands. Low cache hit rates usually trace back to a few common issues. Fixing them can take a 40% hit rate to 85% or higher. .nvmrc or .node-version for Node, Gemfile.lock with .ruby-version for Ruby, packages.lock.json for NuGet, composer.lock for PHP, requirements.txt with pinned versions (or pip-tools) for Python. Ensure CI reads these version files, and consider Docker images with fixed runtimes for complete consistency.', }, { - lead: 'Non-deterministic outputs', - text: 'poison the cache silently. If your build includes a timestamp, a random identifier, or output that varies based on execution order, the outputs will differ even with identical inputs. Look for build timestamps in generated files, unsorted imports or exports, and random IDs in bundles. The `SOURCE_DATE_EPOCH` environment variable helps with timestamp-based non-determinism.', + lead: 'Non-deterministic output', + text: 'Poisons the cache silently. If your build includes a timestamp, a random identifier, or output that varies based on execution order, the outputs will differ even with identical inputs. Look for build timestamps in generated files, unsorted imports or exports, and random IDs in bundles. The SOURCE_DATE_EPOCH environment variable helps with timestamp-based non-determinism.', }, { lead: 'Overly broad input specifications', - text: 'cause unnecessary invalidation. If your build target includes `{projectRoot}/**/*` as inputs, then changing a test file invalidates the production build cache — even though tests don\'t affect the build output. Use named input sets (like Nx\'s "production" pattern) that exclude test files, documentation, and other non-production artifacts.', + text: 'Causes unnecessary invalidation. If your build target includes {projectRoot}/**/* as inputs, then changing a test file invalidates the production build cache — even though tests don\'t affect the build output. Use named input sets (like Nx\'s "production" pattern) that exclude test files, documentation, and other non-production artifacts.', }, { lead: 'Global dependencies included incorrectly', - text: 'cause cascading invalidation. If you list `package-lock.json` in every target\'s inputs, then every dependency update invalidates every cache. Be selective: only include lock files in inputs for targets that actually care about dependency versions (usually everything, but some lint tasks might not).', + text: 'Causes cascading invalidation. If you list package-lock.json in every target\'s inputs, then every dependency update invalidates every cache. Be selective: only include lock files in inputs for targets that actually care about dependency versions (usually everything, but some lint tasks might not).', }, ]} /> @@ -673,7 +688,11 @@ Low cache hit rates usually trace back to a few common issues. Fixing them can t This is where the named inputs configuration from earlier pays off. The "production" input set we defined in `nx.json` excludes test files, stories, and test configuration. When your build target uses `"inputs": ["production", "^production"]`, changing a test file doesn't invalidate the build cache. The `^production` syntax extends this to dependencies — a dependency's test changes don't invalidate your build either.
.nvmrc'], }, { th: 'Test files in build inputs', @@ -690,7 +709,7 @@ This is where the named inputs configuration from earlier pays off. The "product }, { th: 'Timestamps in output', - td: ['Cache never hits twice', 'Set SOURCE_DATE_EPOCH'], + td: ['Cache never hits twice', 'Set SOURCE_DATE_EPOCH'], }, { th: 'Lock file in all inputs', @@ -783,7 +802,8 @@ For __main branch builds__, target under 5 minutes with an 85%+ cache hit rate. For __release builds__, targets depend on your release strategy. If you're rebuilding from a tag that's far behind main, expect lower cache hit rates and longer durations. For frequent releases close to main, you should still see good caching.
The goal isn't the fastest possible full build — it's the fastest possible feedback for typical changes. Optimize for the common case (small, focused changes) while ensuring full builds remain tractable for major changes. + +The best time to implement these optimizations is before your CI becomes a bottleneck. The second best time is now. Start with affected builds, add caching, and iterate from there. Your future self — and your team — will thank you. diff --git a/src/content/articles/mtls-certificate-rotation-service-mesh-authentication/index.mdx b/src/content/articles/mtls-certificate-rotation-service-mesh-authentication/index.mdx index 0e43789c4..b7489f7ff 100644 --- a/src/content/articles/mtls-certificate-rotation-service-mesh-authentication/index.mdx +++ b/src/content/articles/mtls-certificate-rotation-service-mesh-authentication/index.mdx @@ -22,7 +22,7 @@ A team enables Istio mTLS across their 50-service mesh. Initial rollout goes smo Recovery takes four hours because the on-call engineer has never manually rotated Istio certificates and the runbook doesn't exist yet. -If this sounds familiar, you're not alone. Enabling mTLS is a configuration change; operating it reliably is an ongoing commitment to understanding certificate lifecycles, building automation for rotation, and monitoring for expiration failures. +If this sounds familiar, you're not alone. Enabling mTLS is a configuration change; operating it reliably is an ongoing commitment to understanding certificate lifecycles, building automation for rotation, and monitoring for expiration failures. ## The Operational Reality of mTLS @@ -30,10 +30,14 @@ Standard TLS — what you use when visiting any HTTPS website — is a one-way t Mutual TLS adds a second handshake step: after the client validates the server's certificate, the server requests and validates a certificate from the client. Both parties cryptographically prove their identity before any application data flows. -The operational cost difference is significant. With standard TLS, you manage certificates for servers — maybe dozens or hundreds. With mTLS, __every__ service needs a certificate, and every service needs to validate certificates from every other service it communicates with. In a 100-service mesh, that's potentially thousands of certificate validation paths to maintain. +The operational cost difference is significant. With standard TLS, you manage certificates for servers — maybe dozens or hundreds. With mTLS, _every_ service needs a certificate, and every service needs to validate certificates from every other service it communicates with. In a 100-service mesh, that's potentially thousands of certificate validation paths to maintain.
@@ -103,6 +110,7 @@ This sequence diagram shows the rotation flow: @@ -146,7 +154,8 @@ groups: Different certificate levels need different alert thresholds:
x509: certificate has expired. Check expiration with openssl x509 -enddate -noout -in cert.pem. Fix by forcing rotation or restarting the workload to trigger certificate renewal.', }, { lead: 'Trust chain broken', - text: 'The certificate is valid but the CA that signed it isn\'t in the trust store. You\'ll see `x509: certificate signed by unknown authority`. This happens during CA rotation if trust bundles aren\'t updated before new certificates are issued.', + text: 'The certificate is valid but the CA that signed it isn\'t in the trust store. You\'ll see x509: certificate signed by unknown authority. This happens during CA rotation if trust bundles aren\'t updated before new certificates are issued.', }, { lead: 'SAN mismatch', - text: 'The certificate is valid but doesn\'t include the hostname being used. Error: `x509: certificate is valid for X, not Y`. This commonly happens when DNS names change or when certificates are issued with incomplete SAN lists.', + text: 'The certificate is valid but doesn\'t include the hostname being used. Error: x509: certificate is valid for X, not Y. This commonly happens when DNS names change or when certificates are issued with incomplete SAN lists.', }, { lead: 'Wrong key usage', - text: 'The certificate exists but wasn\'t issued for mTLS. If it only has `serverAuth` in Extended Key Usage, it can\'t be used as a client certificate. Reissue with both `serverAuth` and `clientAuth`.', + text: 'The certificate exists but wasn\'t issued for mTLS. If it only has serverAuth in Extended Key Usage, it can\'t be used as a client certificate. Reissue with both serverAuth and clientAuth.', }, ]} /> @@ -208,7 +217,11 @@ mTLS failures produce cryptic errors. The TLS handshake fails, and you get a gen I once spent two hours debugging a "connection reset by peer" that turned out to be a certificate with `serverAuth` only — no `clientAuth`. The error message mentioned nothing about key usage. The fix was a one-line change to the certificate spec, but finding it required systematically ruling out every other possibility.
@@ -60,7 +61,11 @@ This is what makes mTLS attractive for service-to-service communication — iden The operational cost difference is significant. With standard TLS, you manage certificates for servers — maybe dozens or hundreds. With mTLS, __every__ service needs a certificate, and every service needs to validate certificates from every other service it communicates with. In a 100-service mesh, that's potentially thousands of certificate validation paths to maintain.
Certificate: Subject: CN=service-a Issuer: CN=cluster-intermediate-ca @@ -118,7 +129,7 @@ Certificate: X509v3 Extended Key Usage: TLS Web Server Authentication # serverAuth TLS Web Client Authentication # clientAuth - required for mTLS -``` + For mTLS, the Extended Key Usage must include both `serverAuth` and `clientAuth`. A certificate with only `serverAuth` can't be used as a client certificate, and vice versa. This is one of the most common misconfigurations when setting up mTLS outside of a service mesh. @@ -140,6 +151,7 @@ A __three-tier hierarchy__ adds an intermediate layer between root and issuing C @@ -149,7 +161,8 @@ The tradeoff is complexity. Longer certificate chains mean more validation steps For cross-cluster mTLS, you have two options. A __shared root__ means all clusters can communicate automatically — any certificate signed by any issuing CA validates back to the same root. Simple, but a root compromise affects everything. __Federated trust__ means each cluster has its own root, and you explicitly configure which clusters trust each other by distributing trust bundles. More work to set up, but you get selective trust and blast radius containment.
spiffe://cluster.local/ns/{namespace}/sa/{service-account}', 'istiod'], }, { th: 'Linkerd', - td: ['`spiffe://identity.linkerd.cluster.local/...`', 'linkerd-identity'], + td: ['spiffe://identity.linkerd.cluster.local/...', 'linkerd-identity'], }, { th: 'Consul Connect', - td: ['`spiffe://cluster.consul/ns/{namespace}/dc/{datacenter}/svc/{service}`', 'Consul CA or Vault'], + td: ['spiffe://cluster.consul/ns/{namespace}/dc/{datacenter}/svc/{service}', 'Consul CA or Vault'], }, ], }, @@ -266,28 +280,32 @@ __CA rotation__ is harder. When you rotate an intermediate or root CA, you're ch The safe order for CA rotation: @@ -333,7 +351,12 @@ groups: Different certificate levels need different alert thresholds:
@@ -517,29 +544,33 @@ I once spent two hours debugging a "connection reset by peer" that turned out to Knowing the common failure modes helps narrow down the problem quickly. x509: certificate has expired or TLS handshake error: certificate verify failed. Check expiration with openssl x509 -enddate -noout -in cert.pem. Fix by forcing rotation or restarting the workload to trigger certificate renewal.', }, { lead: 'Trust chain broken', - text: 'The certificate is valid but the CA that signed it isn\'t in the trust store. You\'ll see `x509: certificate signed by unknown authority`. This happens during CA rotation if trust bundles aren\'t updated before new certificates are issued. Verify the chain with `openssl verify -CAfile root.pem cert.pem`.', + text: 'The certificate is valid but the CA that signed it isn\'t in the trust store. You\'ll see x509: certificate signed by unknown authority. This happens during CA rotation if trust bundles aren\'t updated before new certificates are issued. Verify the chain with openssl verify -CAfile root.pem cert.pem.', }, { lead: 'SAN mismatch', - text: 'The certificate is valid but doesn\'t include the hostname being used. Error: `x509: certificate is valid for X, not Y`. Check SANs with `openssl x509 -text -noout | grep -A1 \'Subject Alternative\'`. This commonly happens when DNS names change or when certificates are issued with incomplete SAN lists.', + text: 'The certificate is valid but doesn\'t include the hostname being used. Error: x509: certificate is valid for X, not Y. Check SANs with openssl x509 -text -noout | grep -A1 \'Subject Alternative\'. This commonly happens when DNS names change or when certificates are issued with incomplete SAN lists.', }, { lead: 'Wrong key usage', - text: 'The certificate exists but wasn\'t issued for mTLS. If it only has `serverAuth` in Extended Key Usage, it can\'t be used as a client certificate. Error: `x509: certificate specifies an incompatible key usage`. Reissue with both `serverAuth` and `clientAuth`.', + text: 'The certificate exists but wasn\'t issued for mTLS. If it only has serverAuth in Extended Key Usage, it can\'t be used as a client certificate. Error: x509: certificate specifies an incompatible key usage. Reissue with both serverAuth and clientAuth.', }, ]} />
netstat -tlnp'], }, { th: 'Connection reset', - td: ['TLS version mismatch or expired cert', '`openssl x509 -enddate`'], + td: ['TLS version mismatch or expired cert', 'openssl x509 -enddate'], }, { th: '"Unknown authority"', - td: ['Trust bundle missing CA', '`openssl verify -CAfile`'], + td: ['Trust bundle missing CA', 'openssl verify -CAfile'], }, { th: '"Valid for X, not Y"', @@ -638,9 +669,12 @@ Before touching certificates, verify the mesh is healthy. Rotating during an exi div]:!pt-3 mb-6", + }} items={[ { - text: 'Verify current certificate status — all certificates should be Ready:', + lead: 'Verify current certificate status — all certificates should be Ready:', }, ]} /> @@ -651,9 +685,13 @@ kubectl get certificates -A | grep -v 'True' div]:!pt-3 mb-6", + }} items={[ { - text: 'Confirm trust bundle is distributed:', + lead: 'Confirm trust bundle is distributed:', }, ]} /> @@ -664,9 +702,13 @@ kubectl get configmap -n istio-system istio-ca-root-cert -o yaml div]:!pt-3 mb-6", + }} items={[ { - text: 'Check for ongoing incidents (don\'t rotate during an outage).', + lead: 'Check for ongoing incidents (don\'t rotate during an outage).', }, ]} /> @@ -675,10 +717,13 @@ __Execution:__ div]:!pt-3 mb-6", + }} items={[ { lead: 'Backup current certificates', - text: '', }, ]} /> @@ -690,6 +735,7 @@ kubectl get secret -n istio-system istio-ca-secret -o yaml > ca-secret-backup.ya div]:!pt-3 mb-6", + }} items={[ { lead: 'Verify workload certificates rotated', - text: '', }, ]} /> @@ -723,6 +772,7 @@ done div]:!pt-3 mb-6", + }} items={[ { - text: 'All services report healthy', + lead: 'All services report healthy', }, { - text: 'No TLS errors in logs', + lead: 'No TLS errors in logs', }, { - text: 'Monitor error rates for 30 minutes', + lead: 'Monitor error rates for 30 minutes', }, ]} /> @@ -754,9 +808,23 @@ __Rollback__ (if mTLS failures occur): Apply backup certificates, restart istiod When mTLS fails across the mesh — multiple services reporting TLS handshake failures, error rates spiking, certificates expired — you need to restore communication first, then fix the root cause. The instinct to immediately fix the certificate issue is wrong; restoring service is the priority. That's why step 2 exists. -__Step 1: Assess scope.__ Is it all services or a subset? Check pod status and Istio telemetry. - -__Step 2: Enable permissive mode (if needed).__ This is the emergency valve — it allows plaintext traffic so services can communicate while you fix the underlying issue: + ```yaml title="emergency-permissive.yaml" apiVersion: security.istio.io/v1beta1 @@ -771,12 +839,28 @@ spec: Apply with `kubectl apply -f emergency-permissive.yaml`. -__Step 3: Diagnose root cause.__ Check certificate expiry, CA chain validity, and istiod logs. - -__Step 4: Apply fix based on diagnosis:__ + -__Step 5: Restore strict mode:__ + ```bash kubectl delete peerauthentication emergency-permissive -n istio-system @@ -801,7 +892,15 @@ kubectl delete peerauthentication emergency-permissive -n istio-system Verify all services are using mTLS before closing the incident. -__Post-incident:__ Document the timeline, add monitoring for this failure mode, update automation, schedule a post-mortem. + Switching to PERMISSIVE mode during an incident allows plaintext traffic, which bypasses mTLS security. Document this clearly in your incident timeline and switch back to STRICT as soon as the underlying issue is resolved. @@ -815,10 +914,8 @@ The trust hierarchy you choose affects everything downstream. Two-tier is simple Certificate TTLs are a tradeoff. Short-lived certificates (24 hours) limit the damage from a compromised certificate but require robust automation. Longer certificates (7 days) are more forgiving of automation failures but increase your exposure window. -The goal is automation so complete that certificate rotation becomes invisible — happening continuously in the background without human intervention or service disruption. When your certificates rotate and nobody notices, you've built a mature mTLS operation. - Start with permissive mode, add monitoring before enforcement, and run a rotation drill before you need it for real. The worst time to learn your mTLS automation is broken is during an incident. ---- +The goal is automation so complete that certificate rotation becomes invisible — happening continuously in the background without human intervention or service disruption. When your certificates rotate and nobody notices, you've built a mature mTLS operation. diff --git a/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/index.mdx b/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/index.mdx index 37ebf29af..691f9bc5d 100644 --- a/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/index.mdx +++ b/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/index.mdx @@ -22,7 +22,7 @@ Here's a scenario I've seen play out more times than I'd like. A team deploys th It turns out `proxy_read_timeout` defaults to 60 seconds, which sounds generous until you realize a few slow endpoints occasionally take 65 seconds. Database queries, external API calls, report generation — any of these can push response times past that threshold. Meanwhile, the team is chasing ghosts in their application code. -Nginx and HAProxy ship with defaults optimized for getting started quickly, not for handling production traffic. Default timeouts assume fast backends. Default buffer sizes assume small requests. When real load arrives — slow clients on mobile networks, large authentication headers, backends that occasionally need extra time — these defaults fail in ways that are hard to diagnose. +Nginx and HAProxy ship with defaults optimized for getting started quickly, not for handling production traffic. Default timeouts assume fast backends. Default buffer sizes assume small requests. When real load arrives — slow clients on mobile networks, large authentication headers, backends that occasionally need extra time — these defaults fail in ways that are hard to diagnose. The good news: two configuration areas account for most proxy-related outages. Fix your timeouts and buffers, and you'll eliminate the majority of mysterious 502s and 400s. Let's start with the more common culprit. @@ -63,7 +63,7 @@ The per-location overrides are essential. A monolithic timeout configuration for ### HAProxy's Total-Time Semantics -HAProxy's timeout model differs in one critical way: `timeout client` and `timeout server` cover the __entire__ request or response, not per-read operations. If you set `timeout server 60s` and the backend takes 30 seconds to send the first byte, then another 35 seconds to send the body, the connection times out — even though data was flowing the whole time. +HAProxy's timeout model differs in one critical way: `timeout client` and `timeout server` cover the _entire_ request or response, not per-read operations. If you set `timeout server 60s` and the backend takes 30 seconds to send the first byte, then another 35 seconds to send the body, the connection times out — even though data was flowing the whole time. ```haproxy title="haproxy-timeouts.cfg" defaults @@ -87,8 +87,13 @@ The `timeout queue` setting deserves attention. When all backend servers reach t When translating configurations between Nginx and HAProxy, the following table maps the key timeout settings. They're not exact equivalents — Nginx's per-read semantics differ from HAProxy's total-time semantics — but this helps when translating configurations.
td:nth-of-type(2)]:!font-normal", + }} content={{ + figure: 'Nginx and HAProxy timeout comparison with recommended production values.', thead: { th: ['Phase', 'Nginx', 'HAProxy', 'Recommended'], }, @@ -116,7 +121,6 @@ When translating configurations between Nginx and HAProxy, the following table m }, ], }, - figure: 'Nginx and HAProxy timeout comparison with recommended production values.', }} /> @@ -178,7 +182,11 @@ If you're seeing 400 errors in HAProxy specifically, check whether `tune.bufsize The following table summarizes buffer settings by scenario. Note that the directive names are Nginx-specific; for HAProxy, adjust `tune.bufsize` to accommodate larger headers or bodies.
td:nth-of-type(2)]:!font-normal", + }} content={{ thead: { th: ['Scenario', 'Nginx Setting', 'Recommendation'], diff --git a/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/pdf.mdx b/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/pdf.mdx index bfed7a994..d943e3afc 100644 --- a/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/pdf.mdx +++ b/src/content/articles/nginx-haproxy-reverse-proxy-production-tuning/pdf.mdx @@ -60,6 +60,7 @@ Finally, the proxy __forwards the request__, __receives the response__, and __de @@ -70,9 +71,9 @@ Nginx and HAProxy handle connections differently, and understanding these models Nginx uses an __event-driven, single-threaded worker__ model. Each worker process handles thousands of connections using non-blocking I/O, but each connection — whether from a client or to a backend — uses one slot from the `worker_connections` pool. Since a proxied request needs both a client connection and a backend connection, each request consumes at least two slots. The formula for maximum concurrent requests is roughly: -```text -max concurrent requests = (worker_processes × worker_connections) / 2 -``` + +`max concurrent requests = (worker_processes × worker_connections) / 2` + With 4 workers and 4096 connections each, you get approximately 8,192 concurrent requests. Keep-alives complicate this — idle connections still consume slots, so you may hit connection limits before CPU or memory becomes a bottleneck. @@ -236,7 +237,11 @@ The `timeout queue` setting deserves special attention. When all backend servers The following table maps Nginx and HAProxy timeout settings to each other. They're not exact equivalents — Nginx's per-read semantics differ from HAProxy's total-time semantics — but this helps when translating configurations between the two.
td:nth-of-type(2)]:!font-normal", + }} content={{ thead: { th: ['Timeout', 'Nginx', 'HAProxy', 'Recommended'], @@ -411,7 +416,11 @@ backend api The `http-request deny if { req.body_size gt 104857600 }` line is worth noting: it rejects oversized requests __before__ buffering them. Without this, HAProxy would accept a 10GB upload attempt before rejecting it — wasting bandwidth and potentially filling disk.
@@ -714,7 +724,8 @@ backend canary The canary deployment example shows a common pattern: routing 5% of traffic to a new version while 95% goes to the stable version. If the canary starts failing health checks or returning errors, you can quickly shift traffic back to stable by adjusting weights or marking the canary server as `disabled`.
on-call sustainability isn't about headcount or rotation schedules. It's about the quality of the alerts that wake people up. One metric predicts whether your on-call will burn out your team: signal-to-noise ratio. Here's how to measure it, recognize when it's killing you, and fix it before someone quits. ## What Signal-to-Noise Ratio Actually Measures -Signal-to-noise ratio is the percentage of pages that required human action. Calculate it monthly: pages that required someone to __do something__ divided by total pages. +Signal-to-noise ratio is the percentage of pages that required human action. Calculate it monthly: pages that required someone to _do something_ divided by total pages. Healthy', + color: "bg-success", }, { - text: '50-80%: concerning', + lead: '50-80%:  Concerning', + color: "bg-warning-offset", }, { - text: 'Below 50%: your alerting is broken', + lead: 'Below 50%:  Your alerting is broken', + color: "bg-danger", }, ]} /> @@ -43,22 +49,37 @@ The math is simple but brutal. A three-person team can sustainably handle maybe Here's an example from a real team audit: span]:!font-semibold [&>span]:ml-2 mb-4", + }} + size={7} items={[ { text: '60 total pages this month', + icon: "alarm-clock", + color: "page-inverse", }, { text: '35 required action', + icon: "gear", + color: "page-inverse", }, { text: '20 auto-resolved before anyone could respond', + icon: "clock", + color: "page-inverse", }, { text: '5 were false positives', + icon: "error", + color: "page-inverse", }, { text: 'SNR: 58%—concerning, needs work', + icon: "graph", + color: "page-inverse", }, ]} /> @@ -68,7 +89,8 @@ That 20 auto-resolved pages is the killer. If an alert fires and resolves before Every alert outcome falls into one of four categories, each with a target:
- ## Recognizing Burnout Before It's Too Late The insidious thing about on-call burnout is that it accumulates slowly. By the time it's obvious, someone is already job hunting. Individual burnout shows up in behavior first. Someone starts acknowledging alerts but not actually investigating them. Response times gradually increase. They snooze alerts instead of addressing them. There's resentment in handoff meetings — subtle comments about the unfairness of the rotation or the quality of alerts. + + Emotional signs follow: dread when an on-call shift approaches, anxiety about phone notifications even when off rotation, the feeling that you can never truly disconnect. Eventually physical symptoms emerge — sleep disruption that persists even off rotation, exhaustion that doesn't recover between shifts. At the team level, watch for alerts being suppressed rather than fixed, runbooks not being updated, post-incident reviews getting skipped, or transfer requests. These are all signs that people have given up on improving the system and are just trying to survive it. @@ -150,7 +173,12 @@ At the team level, watch for alerts being suppressed rather than fixed, runbooks The numbers tell a story too:
@@ -213,12 +244,12 @@ The biggest win for most teams is after-hours filtering. Not everything needs to This isn't ignoring problems — it's acknowledging that "one replica down out of three" at 2 AM doesn't justify waking someone up when the service is still functional. The on-call person can check in the morning. -If your page budget is consistently exceeded despite these efforts, there's a nuclear option: stop feature work until alerting is fixed. This sounds dramatic, but reliability debt is real debt. A team that can't sleep can't ship features either. Sometimes you need to stop digging before you can climb out. - The weekly review is the highest-leverage practice for on-call sustainability. Thirty minutes per week of deliberate improvement compounds into dramatically better on-call within a quarter. +If your page budget is consistently exceeded despite these efforts, there's a nuclear option: stop feature work until alerting is fixed. This sounds dramatic, but reliability debt is real debt. A team that can't sleep can't ship features either. Sometimes you need to stop digging before you can climb out. + ## Start Here The team I mentioned at the start didn't need a new rotation schedule or a new incident management platform. They needed fewer, better alerts. The constraint of being a small team forced discipline that larger teams often lack — when you can't spread the pain across twenty people, you have to actually fix the problems. @@ -241,15 +272,19 @@ Start this week: div]:!pt-2", + }} items={[ { - text: 'Calculate your signal-to-noise ratio. If it\'s below 80%, you have work to do.', + lead: 'Calculate your signal-to-noise ratio. If it\'s below 80%, you have work to do.', }, { - text: 'Schedule your first weekly review during the next on-call handoff.', + lead: 'Schedule your first weekly review during the next on-call handoff.', }, { - text: 'Pick one high-volume alert and either tune it, automate it, or delete it.', + lead: 'Pick one high-volume alert and either tune it, automate it, or delete it.', }, ]} /> diff --git a/src/content/articles/on-call-rotation-small-teams-sustainable-coverage/pdf.mdx b/src/content/articles/on-call-rotation-small-teams-sustainable-coverage/pdf.mdx index cfa8ca2a1..7ab21402a 100644 --- a/src/content/articles/on-call-rotation-small-teams-sustainable-coverage/pdf.mdx +++ b/src/content/articles/on-call-rotation-small-teams-sustainable-coverage/pdf.mdx @@ -24,12 +24,12 @@ I've seen this play out the hard way. A three-person team copies the on-call set They rebuilt from scratch. Daily rotations instead of weekly, so a bad night didn't compound into a bad week. Aggressive alert suppression for anything that auto-resolved within five minutes. A hard policy that nothing non-critical pages outside business hours. Pages dropped to 2-3 per week. The team started sleeping again. -The lesson: small team on-call isn't about copying enterprise playbooks with fewer people. It's about ruthless prioritization of what actually needs human attention at 3 AM versus what can wait until morning. When you have three people, you can't afford to wake someone up for something that could wait 8 hours. - The most dangerous on-call mistake for small teams is treating every alert as equally urgent. Not everything is a 3 AM problem — and pretending otherwise burns out your team in weeks. +The lesson: small team on-call isn't about copying enterprise playbooks with fewer people. It's about ruthless prioritization of what actually needs human attention at 3 AM versus what can wait until morning. When you have three people, you can't afford to wake someone up for something that could wait 8 hours. + ## Rotation Design for Small Teams ### Rotation Length Tradeoffs @@ -37,19 +37,22 @@ The most dangerous on-call mistake for small teams is treating every alert as eq The rotation length question seems simple — weekly or daily?—but the answer depends on your page volume and how predictable it is. @@ -59,7 +62,8 @@ The hybrid approach combines these: full coverage during business hours, but aft For a three-person team, I recommend this structure regardless of rotation length:
div]:!pt-2 mb-6", + }} items={[ { - text: 'Borrow someone from an adjacent team', + lead: 'Borrow someone from an adjacent team', }, { - text: 'Hire a contractor for acknowledge-and-escalate coverage', + lead: 'Hire a contractor for acknowledge-and-escalate coverage', }, { - text: 'Temporarily reduce alerting sensitivity and accept slower response times', + lead: 'Temporarily reduce alerting sensitivity and accept slower response times', }, ]} /> @@ -128,23 +136,23 @@ Alert quality makes or breaks small team on-call. You can have perfect rotation Every alert needs a severity level, and that level determines whether it pages, when it pages, and how fast you need to respond. For small teams, I use four levels: @@ -158,22 +166,37 @@ Signal-to-noise ratio is the percentage of alerts that actually required human a Here's an example from a real team audit: span]:!font-semibold [&>span]:ml-2 mb-4", + }} + size={7} items={[ { text: '60 total pages this month', + icon: "alarm-clock", + color: "page-inverse", }, { text: '35 required action', + icon: "gear", + color: "page-inverse", }, { text: '20 auto-resolved before anyone could respond', + icon: "clock", + color: "page-inverse", }, { text: '5 were false positives', + icon: "error", + color: "page-inverse", }, { text: 'SNR: 58%—concerning, needs work', + icon: "graph", + color: "page-inverse", }, ]} /> @@ -184,21 +207,25 @@ Every week during on-call handoff, review every page from the past week and ask div]:!pt-2 mb-6", + }} items={[ { - text: 'Did this require human action?', + lead: 'Did this require human action?', }, { - text: 'Could it have been prevented?', + lead: 'Could it have been prevented?', }, { - text: 'Could it have waited until morning?', + lead: 'Could it have waited until morning?', }, { - text: 'Was the runbook sufficient?', + lead: 'Was the runbook sufficient?', }, { - text: 'Should this alert exist at all?', + lead: 'Should this alert exist at all?', }, ]} /> @@ -210,7 +237,8 @@ The most common noise sources are predictable. __Flapping alerts__ fire and reso [^hysteresis]: Hysteresis means the alert threshold differs depending on direction. An alert might fire when CPU exceeds 90% but only resolve when it drops below 80%. This 10% gap prevents the alert from flapping when CPU hovers around a single threshold value.
@@ -278,7 +309,8 @@ There are a few escalation anti-patterns to avoid. __Paging everyone simultaneou Time-based routing is what makes small team on-call sustainable. The idea is simple: during business hours, page for P1 and P2 incidents. After hours, page only for P1.
div]:!pt-2 mb-6", + }} items={[ { - text: 'What this alert means in plain English', + lead: 'What this alert means in plain English', }, { - text: 'Why we alert on this (what\'s the user impact)', + lead: 'Why we alert on this (what\'s the user impact)', }, { - text: 'The first three things to check', + lead: 'The first three things to check', }, { - text: 'How to mitigate if known', + lead: 'How to mitigate if known', }, { - text: 'When to escalate', + lead: 'When to escalate', }, { - text: 'Who owns this system', + lead: 'Who owns this system', }, ]} /> @@ -394,23 +433,26 @@ The most important section is "First Response"—the first three things to check Common runbook anti-patterns: @@ -434,7 +476,12 @@ At the team level, watch for increasing alert suppression (people silencing thin The numbers tell a story too. These thresholds are rough guidelines, but they're useful:
@@ -546,7 +597,10 @@ Every alert that can be auto-remediated should be. Human attention is the scarce Common candidates for auto-remediation: +Raw numbers alone don't tell the story. A team averaging four pages per week looks healthy by the volume metric, but if three of those four are false positives that wake someone at 4 AM, the on-call experience is miserable. Each metric below adds a dimension that the table above can't capture — the _why_ behind the numbers and the corrective action each one points to. + @@ -645,7 +703,10 @@ __Incident management data__ requires an exporter. For PagerDuty, the [pagerduty Run the exporter as a sidecar or standalone service that scrapes your incident management API on a schedule (typically every 60 seconds). The exporter authenticates with a read-only API token and exposes metrics on a `/metrics` endpoint that Prometheus scrapes. div]:!pt-2 mb-6", + }} items={[ { - text: 'Calculate your current signal-to-noise ratio. What percentage of pages required action?', + lead: 'Calculate your current signal-to-noise ratio. What percentage of pages required action?', }, { - text: 'Implement after-hours filtering: P1 pages anytime, P2 during business hours only', + lead: 'Implement after-hours filtering: P1 pages anytime, P2 during business hours only', }, { - text: 'Schedule your first weekly review during the next on-call handoff', + lead: 'Schedule your first weekly review during the next on-call handoff', }, { - text: 'Pick one high-volume alert and either tune it, automate it, or delete it', + lead: 'Pick one high-volume alert and either tune it, automate it, or delete it', }, { - text: 'Set up basic metrics: page volume per person per week, auto-resolve rate', + lead: 'Set up basic metrics: page volume per person per week, auto-resolve rate', }, ]} /> diff --git a/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/index.mdx b/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/index.mdx index a8b4d926b..be3f7cd69 100644 --- a/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/index.mdx +++ b/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/index.mdx @@ -15,14 +15,14 @@ featured: true A developer opens a pull request for a Terraform change. Sixteen hours later, a security review rejects it: the S3 bucket lacks encryption. The developer fixes it, waits another day for re-review. This cycle — write, wait, reject, fix, wait — drains velocity and breeds resentment toward security processes. -Shift-left advocates say to check policies earlier. But "earlier" often means CI, which still means waiting for pipelines after pushing code. Real shift-left means __before the commit__ — policy checks that run in seconds during `git commit`, catching violations while the context is fresh and the fix is trivial. - -OPA (Open Policy Agent) and Conftest make this possible. OPA is a general-purpose policy engine that evaluates structured data against policies written in Rego. Conftest wraps OPA with ergonomic defaults for infrastructure files — it parses Terraform, Kubernetes YAML, and Dockerfiles into JSON that OPA can evaluate. Together, they provide fast, local policy enforcement that doesn't require cloud credentials or pipeline execution. +Shift-left advocates say to check policies earlier. But "earlier" often means CI, which still means waiting for pipelines after pushing code. Real shift-left means _before the commit_ — policy checks that run in seconds during `git commit`, catching violations while the context is fresh and the fix is trivial. Policy adoption paradox: comprehensive policies with slow feedback get disabled. Minimal policies with fast feedback get expanded. Start with five critical policies — encryption enabled, no public access, resource limits, no privileged containers, no `:latest` tags — that run in under two seconds. Get adoption first, then add coverage. +OPA (Open Policy Agent) and Conftest make this possible. OPA is a general-purpose policy engine that evaluates structured data against policies written in Rego. Conftest wraps OPA with ergonomic defaults for infrastructure files — it parses Terraform, Kubernetes YAML, and Dockerfiles into JSON that OPA can evaluate. Together, they provide fast, local policy enforcement that doesn't require cloud credentials or pipeline execution. + ## Pre-commit Hooks for Instant Feedback Pre-commit is where shift-left becomes real. Conftest integrates with the pre-commit framework to run policies against staged files before they're committed. The key is speed — pre-commit hooks that take more than a few seconds get disabled. @@ -57,8 +57,10 @@ Pre-commit only runs against staged files by default, which keeps evaluation fas The pre-commit framework shown above is language-agnostic, but each ecosystem has its canonical approach:
pip install pre-commit && pre-commit install'], }, { th: 'Node.js', - td: ['Husky + lint-staged', '`npx husky init`'], + td: ['Husky + lint-staged', 'npx husky init'], }, { th: 'Ruby/Rails', - td: ['Overcommit', '`gem install overcommit && overcommit --install`'], + td: ['Overcommit', 'gem install overcommit && overcommit --install'], }, { th: 'Go', - td: ['Lefthook', '`go install github.com/evilmartians/lefthook`'], + td: ['Lefthook', 'go install github.com/evilmartians/lefthook'], }, ], }, - figure: 'Pre-commit tools by ecosystem.', }} /> @@ -95,15 +96,18 @@ Speed is non-negotiable: if policy checks take more than a few seconds, develope Four principles separate policies that get adopted from policies that get bypassed: deny_unauthorized_image, deny_missing_resources, deny_missing_labels. When a violation fires, developers know exactly what to fix.', }, { lead: 'Actionable messages.', - text: '"Policy violation" tells the developer nothing. "Container missing resource limits" is better. "Container \'nginx\' missing resource limits. Add `spec.containers[].resources.limits.cpu` and `memory`" is what they actually need. Include the resource name, the violation, and the fix path.', + text: '"Policy violation" tells the developer nothing. "Container missing resource limits" is better. "Container \'nginx\' missing resource limits. Add spec.containers[].resources.limits.cpu and memory" is what they actually need. Include the resource name, the violation, and the fix path.', }, { lead: 'Minimal false positives.', @@ -130,24 +134,61 @@ Examples: "Deployment 'api-server': Missing required label 'team'. Fix: Add meta Organize policies by technology, then resource type, then concern. This structure makes policies discoverable and enables selective evaluation: -```text -policies/ -├── kubernetes/ -│ ├── pods/ -│ │ ├── privileged.rego -│ │ ├── resources.rego -│ │ └── images.rego -│ └── common/ -│ └── helpers.rego -├── terraform/ -│ └── aws/ -│ ├── s3.rego -│ ├── iam.rego -│ └── security_groups.rego -└── data/ - └── allowed_registries.json -``` -Code: Policy directory structure. + Conftest uses the `--namespace` flag to selectively evaluate policies. Run only Kubernetes policies on YAML files with `conftest test deployment.yaml --namespace kubernetes`. Run only Terraform AWS policies with `conftest test tfplan.json --namespace terraform.aws`. This is how namespace organization pays off — you can run subsets of policies based on context. diff --git a/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/pdf.mdx b/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/pdf.mdx index d00f35269..f974084ad 100644 --- a/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/pdf.mdx +++ b/src/content/articles/opa-conftest-policy-as-code-infrastructure-guardrails/pdf.mdx @@ -23,7 +23,7 @@ import conftestDiagram from "./diagrams/conftest-evaluation-flow.jpg" A developer opens a pull request for a Terraform change. Sixteen hours later, a security review rejects it: the S3 bucket lacks encryption. The developer fixes it, waits another day for re-review. This cycle — write, wait, reject, fix, wait — drains velocity and breeds resentment toward security processes. -Shift-left advocates say to check policies earlier. But "earlier" often means CI, which still means waiting for pipelines after pushing code. Real shift-left means __before the commit__ — policy checks that run in seconds during `git add`, catching violations while the context is fresh and the fix is trivial. +Shift-left advocates say to check policies earlier. But "earlier" often means CI, which still means waiting for pipelines after pushing code. Real shift-left means _before the commit_ — policy checks that run in seconds during `git add`, catching violations while the context is fresh and the fix is trivial. This article covers building infrastructure guardrails with OPA (Open Policy Agent) and Conftest. The focus is on speed: pre-commit hooks that run in under two seconds, CI checks that parallelize across hundreds of policies, and feedback loops that make compliance a developer experience issue rather than a security bottleneck. @@ -42,7 +42,10 @@ OPA is a general-purpose policy engine that decouples policy decisions from poli Three deployment modes serve different use cases: not input.spec.selector'], }, { th: 'Allowed values', - td: ['Whitelist valid options', '`input.type in {"ClusterIP", "NodePort"}`'], + td: ['Whitelist valid options', 'input.type in {"ClusterIP", "NodePort"}'], }, { th: 'Regex matching', - td: ['Pattern validation', '`regex.match("^[a-z][a-z0-9-]*$", name)`'], + td: ['Pattern validation', 'regex.match("^[a-z][a-z0-9-]*$", name)'], }, { th: 'Numeric constraints', - td: ['Range validation', '`input.replicas > 0; input.replicas <= 10`'], + td: ['Range validation', 'input.replicas > 0; input.replicas <= 10'], }, { th: 'Cross-resource', - td: ['Reference related objects', '`data.namespaces[input.metadata.namespace]`'], + td: ['Reference related objects', 'data.namespaces[input.metadata.namespace]'], }, { th: 'Conditional', - td: ['Environment-specific rules', '`is_production; ... production-only rule ...`'], + td: ['Environment-specific rules', 'is_production; ... production-only rule ...'], }, ], }, @@ -217,15 +221,18 @@ Rego's learning curve is real but bounded. Most infrastructure policies use a sm Four principles separate policies that get adopted from policies that get bypassed: deny_unauthorized_image, deny_missing_resources, deny_missing_labels. When a violation fires, developers know exactly what to fix.', }, { lead: 'Actionable messages.', - text: '"Policy violation" tells the developer nothing. "Container missing resource limits" is better. "Container \'nginx\' missing resource limits. Add `spec.containers[].resources.limits.cpu` and `memory`" is what they actually need. Include the resource name, the violation, and the fix path.', + text: '"Policy violation" tells the developer nothing. "Container missing resource limits" is better. "Container \'nginx\' missing resource limits. Add spec.containers[].resources.limits.cpu and memory" is what they actually need. Include the resource name, the violation, and the fix path.', }, { lead: 'Minimal false positives.', @@ -240,7 +247,7 @@ Four principles separate policies that get adopted from policies that get bypass Structure violation messages consistently: -```text +```markdown [Resource Type] [Resource Name]: [Violation]. Fix: [Specific remediation] ``` @@ -249,13 +256,19 @@ Code: Violation message template. Examples: Deployment \'api-server\': Missing required label \'team\'. Fix: Add metadata.labels.team with your team name.', + color: "border-danger", }, { - text: "`Pod 'worker': Container 'app' uses image from unauthorized registry 'docker.io'. Fix: Use images from 'gcr.io/company-project' or 'artifactory.company.com'.`", + lead: 'Pod \'worker\': Container \'app\' uses image from unauthorized registry \'docker.io\'. Fix: Use images from \'gcr.io/company-project\' or \'artifactory.company.com\'.', + color: "border-danger", }, ]} /> @@ -264,40 +277,102 @@ Examples: Organize policies by technology, then resource type, then concern. This structure makes policies discoverable and enables selective evaluation — run only Terraform policies on Terraform files. -```text -policies/ -├── kubernetes/ -│ ├── pods/ -│ │ ├── privileged.rego -│ │ ├── resources.rego -│ │ └── images.rego -│ ├── deployments/ -│ │ ├── replicas.rego -│ │ └── labels.rego -│ ├── services/ -│ │ └── types.rego -│ └── common/ -│ ├── labels.rego -│ └── helpers.rego -├── terraform/ -│ ├── aws/ -│ │ ├── s3.rego -│ │ ├── iam.rego -│ │ └── security_groups.rego -│ └── common/ -│ └── tags.rego -├── docker/ -│ ├── base_images.rego -│ └── best_practices.rego -├── data/ -│ ├── allowed_registries.json -│ ├── required_labels.json -│ └── team_mappings.json -└── lib/ - ├── helpers.rego - └── constants.rego -``` -Code: Policy directory structure. + Rego packages map to this directory structure. Use namespacing to group related policies and enable selective evaluation: @@ -534,8 +609,10 @@ Pre-commit only runs against staged files by default, which keeps evaluation fas The [pre-commit framework](https://pre-commit.com/) shown above is language-agnostic and works across ecosystems, but each stack has its canonical approach:
pip install pre-commit && pre-commit install'], }, { th: 'Node.js', - td: ['Husky + lint-staged', '`npx husky init` (works for any project with package.json)'], + td: ['Husky + lint-staged', 'npx husky init'], }, { th: 'Ruby/Rails', - td: ['Overcommit', '`gem install overcommit && overcommit --install`'], + td: ['Overcommit', 'gem install overcommit && overcommit --install'], }, { th: 'PHP/Laravel', - td: ['GrumPHP', '`composer require --dev phpro/grumphp`'], + td: ['GrumPHP', 'composer require --dev phpro/grumphp'], }, { th: '.NET/C#', - td: ['Husky.Net', '`dotnet tool install husky`'], + td: ['Husky.Net', 'dotnet tool install husky'], }, { th: 'Go', - td: ['pre-commit or lefthook', '`go install github.com/evilmartians/lefthook`'], + td: ['pre-commit or lefthook', 'go install github.com/evilmartians/lefthook'], }, ], }, - figure: 'Pre-commit tools by ecosystem.', }} /> @@ -651,7 +727,7 @@ __Branch protection is required for enforcement.__ Without it, developers can me Other platforms have equivalent mechanisms:
terraform plan)'], }, { th: 'Credentials', @@ -943,7 +1027,11 @@ HCL parsing is best for syntax checks and naming conventions where speed matters HCL parsing (`conftest test *.tf --parser hcl2`) catches static violations without cloud access. Plan JSON evaluation sees the complete resolved configuration. Use both at different stages:
div]:!pt-2 mb-6", + }} items={[ { - text: 'Developer adds exception annotation with reason', + lead: 'Developer adds exception annotation with reason', }, { - text: 'PR triggers policy check — warns about exception usage', + lead: 'PR triggers policy check — warns about exception usage', }, { - text: 'Security team reviews exception request', + lead: 'Security team reviews exception request', }, { - text: 'If approved, merge with exception and expiration date', + lead: 'If approved, merge with exception and expiration date', }, { - text: 'Exception logged in audit system', + lead: 'Exception logged in audit system', }, ]} /> @@ -1323,7 +1416,8 @@ Code: Exception annotation format. Different exception types warrant different governance levels:
Start with five critical policies that run in under two seconds. Get adoption. Add coverage. The fastest path to comprehensive guardrails runs through developer trust. + +The goal is policies that developers trust: fast enough to not slow them down, accurate enough to not cry wolf, and flexible enough to handle real-world complexity. The measure of success isn't how many violations you block — it's how few violations reach production combined with how little friction developers experience. Both matter. diff --git a/src/content/articles/openapi-spec-documentation-sdk-generation-validation/index.mdx b/src/content/articles/openapi-spec-documentation-sdk-generation-validation/index.mdx index 2f5cf9ad7..d403c2cb2 100644 --- a/src/content/articles/openapi-spec-documentation-sdk-generation-validation/index.mdx +++ b/src/content/articles/openapi-spec-documentation-sdk-generation-validation/index.mdx @@ -13,7 +13,7 @@ featured: true Most teams I work with treat OpenAPI specs as __output__ — something you generate from existing code and call "documentation." That's archaeology, not specification. Meanwhile, their hand-written docs in Confluence are perpetually out of date, their manually coded SDKs break when the API changes, and validation logic differs between the gateway and the backend. -The real value of OpenAPI emerges when you flip the relationship. The spec becomes the source of truth that drives your code, not the other way around. When the spec drives everything — validation middleware, CI checks, generated clients — consistency becomes automatic. The spec can't disagree with your validation because your validation __reads__ the spec. Documentation can't drift because it's generated from the same source. +The real value of OpenAPI emerges when you flip the relationship. The spec becomes the source of truth that drives your code, not the other way around. When the spec drives everything — validation middleware, CI checks, generated clients — consistency becomes automatic. The spec can't disagree with your validation because your validation _reads_ the spec. Documentation can't drift because it's generated from the same source. That's the theory. Here's how it works in practice, starting with the two areas that pay off immediately: request validation and CI automation. @@ -21,7 +21,7 @@ That's the theory. Here's how it works in practice, starting with the two areas Validation is where OpenAPI specs prove their worth. Instead of writing validation logic by hand — checking types, verifying formats, ensuring required fields exist — you derive it directly from the spec. The same schema that defines your API contract also enforces it at runtime. -The key is understanding what schema validation actually does. It handles __structural__ correctness: rejecting requests where `quantity` is a string instead of an integer, where `email` doesn't match email format, where required fields are missing. It won't reject a request where `customerId` is a valid UUID that doesn't exist in your database. That's business validation, and you still need it. +The key is understanding what schema validation actually does. It handles _structural_ correctness: rejecting requests where `quantity` is a string instead of an integer, where `email` doesn't match email format, where required fields are missing. It won't reject a request where `customerId` is a valid UUID that doesn't exist in your database. That's business validation, and you still need it. In Node.js/Express, `express-openapi-validator` is the standard choice. Point it at your spec file, and it automatically validates request bodies, query parameters, path parameters, and headers against your schemas: @@ -54,10 +54,11 @@ Code: Express middleware that validates requests against your OpenAPI spec. The `validateResponses: true` option deserves attention. Enable it in development and staging — it catches cases where your implementation returns data that doesn't match your spec. That's an early warning that spec and code have drifted apart. You might disable it in production for performance, but keeping it on during development catches bugs before you ship. -Other frameworks have equivalent solutions. If you're in Python, FastAPI handles this automatically — your Pydantic type hints __are__ your validation schema, and FastAPI generates an OpenAPI spec from them. You get validation and documentation from the same source without any additional configuration. Rails has `committee` (Rack middleware that validates against a spec file), Laravel has `spectator`, Go has `kin-openapi` for Echo, Chi, and Gin. The pattern is the same across all of them: point middleware at your spec, let it reject malformed requests before they hit your business logic. +Other frameworks have equivalent solutions. If you're in Python, FastAPI handles this automatically — your Pydantic type hints _are_ your validation schema, and FastAPI generates an OpenAPI spec from them. You get validation and documentation from the same source without any additional configuration. Rails has `committee` (Rack middleware that validates against a spec file), Laravel has `spectator`, Go has `kin-openapi` for Echo, Chi, and Gin. The pattern is the same across all of them: point middleware at your spec, let it reject malformed requests before they hit your business logic.
express-openapi-validator', 'Middleware validates against spec file'], }, { th: 'FastAPI/Python', @@ -74,11 +75,11 @@ Other frameworks have equivalent solutions. If you're in Python, FastAPI handles }, { th: 'Rails', - td: ['`committee`', 'Rack middleware validates against spec'], + td: ['committee', 'Rack middleware validates against spec'], }, { th: 'Go', - td: ['`kin-openapi`', 'Middleware for Echo, Chi, Gin'], + td: ['kin-openapi', 'Middleware for Echo, Chi, Gin'], }, ], }, diff --git a/src/content/articles/openapi-spec-documentation-sdk-generation-validation/pdf.mdx b/src/content/articles/openapi-spec-documentation-sdk-generation-validation/pdf.mdx index 9a73de751..0d380337b 100644 --- a/src/content/articles/openapi-spec-documentation-sdk-generation-validation/pdf.mdx +++ b/src/content/articles/openapi-spec-documentation-sdk-generation-validation/pdf.mdx @@ -32,19 +32,19 @@ If you've worked with OpenAPI specs before, you know they can get unwieldy fast. An OpenAPI spec has four top-level sections that matter: info for metadata', }, { - text: '`servers` for environment URLs', + lead: 'servers for environment URLs', }, { - text: '`paths` for your actual endpoints', + lead: 'paths for your actual endpoints', }, { - text: '`components` for reusable definitions', + lead: 'components for reusable definitions', }, ]} /> @@ -184,8 +184,9 @@ Code: Using allOf to compose shared base properties into entity schemas. For polymorphic types — where a field can be one of several shapes — use `oneOf` with a discriminator. Payment methods are the classic example: credit cards have `last4` and `expiryMonth`, bank transfers have `routingNumber` and `accountNumber`, PayPal has `email`. The discriminator tells parsers which schema applies based on a `type` field.
allOf composition', td: ['Multiple entities share identical base fields (audit timestamps, IDs)'], }, { - th: '`oneOf` polymorphism', + th: 'oneOf polymorphism', td: ['A field can be one of several distinct types with different structures'], }, ], @@ -261,7 +262,7 @@ The workflow looks like this: design the API in your spec editor, review with st You'll want a decent editor for spec-first work. Here are the main options:
@@ -443,9 +447,13 @@ responses: ``` Code: Named examples for different order states. +These individual techniques — rich descriptions, multiple examples, error scenarios — compound when applied consistently across your spec. The table below summarizes which documentation elements to prioritize and what each should contain, so you can audit an existing spec or build a new one with the right coverage from the start. +
@@ -490,25 +497,35 @@ The good news is that SDK generation has matured significantly. With the right g The SDK generator landscape breaks down into three categories: comprehensive multi-language tools, language-specific tools with better output, and types-only generators. openapi-typescript', + text: 'Takes a different approach: it generates only TypeScript types, no runtime code. You get interfaces for your request/response bodies and use your own HTTP client. This is the lightest-weight option and works well when you want type safety without adopting a generated SDK wholesale.', }, ]} /> +These tools solve different problems, so the right choice depends less on raw feature count and more on who will consume the generated code. A frontend-heavy team may prioritize React Query integration and ergonomic TypeScript output, while a platform team supporting multiple languages may accept less idiomatic clients in exchange for broader coverage. The comparison below makes those tradeoffs explicit. +
committee', 'Rack middleware that validates against OpenAPI specs'], }, { th: 'Laravel', - td: ['`spectator`', 'Request/response validation with PHPUnit integration'], + td: ['spectator', 'Request/response validation with PHPUnit integration'], }, { th: '.NET', - td: ['`Microsoft.AspNetCore.OpenApi`', 'Built-in validation in ASP.NET Core 9+; earlier versions use `NSwag`'], + td: ['Microsoft.AspNetCore.OpenApi', 'Built-in validation in ASP.NET Core 9+; earlier versions use NSwag'], }, { th: 'Go', - td: ['`kin-openapi`', 'Middleware for Echo, Chi, Gin; validates requests and responses'], + td: ['kin-openapi', 'Middleware for Echo, Chi, Gin; validates requests and responses'], }, { th: 'Spring Boot', - td: ['`springdoc-openapi`', 'Code-first with automatic spec generation and validation'], + td: ['springdoc-openapi', 'Code-first with automatic spec generation and validation'], }, ], }, @@ -714,7 +731,8 @@ Code: Error handler that transforms technical validation errors into readable me It's important to understand the boundaries of schema validation:
@@ -881,15 +909,18 @@ Run it in PR checks with `--fail-on-incompatible` to block merges that would bre Contract tests verify your implementation actually matches your spec. Dredd and Schemathesis are the main tools: diff --git a/src/content/articles/opentelemetry-span-design-granularity-overhead/index.mdx b/src/content/articles/opentelemetry-span-design-granularity-overhead/index.mdx index b5607dddc..e32755dd4 100644 --- a/src/content/articles/opentelemetry-span-design-granularity-overhead/index.mdx +++ b/src/content/articles/opentelemetry-span-design-granularity-overhead/index.mdx @@ -15,22 +15,26 @@ A team I worked with instrumented a new service with spans for every function ca They refactored to instrument only service boundaries, significant I/O operations, and error paths. Instrumented service boundaries, significant I/O operations, and error paths. Span count dropped to 15-20 per request. Traces became readable. The critical path was obvious at a glance. Storage costs dropped 90%. Debugging time went from minutes to seconds. -The lesson: granularity without readability is noise, not observability. More spans don't mean better visibility — they often mean worse. The goal is __enough__ spans to debug problems, not so many that you create new ones. +The lesson: granularity without readability is noise, not observability. More spans don't mean better visibility — they often mean worse. The goal is _enough_ spans to debug problems, not so many that you create new ones. ## What to Instrument The question "should this operation have a span?" comes up constantly. Here's how I think about it.
@@ -173,7 +187,10 @@ A trace is only useful if you can read it. I've seen traces that technically con Most readability problems fall into a few common anti-patterns: startActiveSpan.', }, { lead: 'Missing gaps.', @@ -189,7 +206,7 @@ Most readability problems fall into a few common anti-patterns: }, { lead: 'Cryptic names.', - text: 'Span names like "span," "operation," "handler," or "process" that don\'t explain what\'s happening. Auto-instrumentation often produces these. The fix is customizing span names to follow the pattern "operation resource" (`HTTP GET /api/orders`, `db.query orders`).', + text: 'Span names like "span," "operation," "handler," or "process" that don\'t explain what\'s happening. Auto-instrumentation often produces these. The fix is customizing span names to follow the pattern "operation resource" (HTTP GET /api/orders, db.query orders).', }, { lead: 'Attribute explosion.', @@ -205,7 +222,7 @@ A good trace has 10-30 spans per request, 3-5 levels of nesting, clear names tha Naming is particularly important for quick scanning. Span names should answer "what operation on what resource?" without requiring you to read the code. The OpenTelemetry semantic conventions provide a good starting point, and the table below shows common patterns and pitfalls.
HTTP GET /api/orders/{id}', 'handler', 'No indication of method or resource'], }, { th: 'Database', - td: ['`db.query orders`', '`SELECT * FROM...`', 'SQL syntax noise, potential sensitive data'], + td: ['db.query orders', 'SELECT * FROM...', 'SQL syntax noise, potential sensitive data'], }, { th: 'Cache', - td: ['`cache.get`', '`redis`', 'Technology not operation — what did you do?'], + td: ['cache.get', 'redis', 'Technology not operation — what did you do?'], }, { th: 'Business logic', - td: ['`order.validate`', '`validateOrder`', 'Function name leaks implementation detail'], + td: ['order.validate', 'validateOrder', 'Function name leaks implementation detail'], }, { th: 'External API', - td: ['`HTTP POST`', '`API call`', 'No method, no way to distinguish calls'], + td: ['HTTP POST', 'API call', 'No method, no way to distinguish calls'], }, ], }, @@ -256,8 +273,8 @@ Span design is an engineering tradeoff: visibility versus overhead, granularity Instrument service boundaries, I/O operations, and significant business logic. Use events for milestones within spans. Use attributes for metadata that helps filtering and debugging. The best traces answer three questions: what happened, where time was spent, and what failed. -Audit your current traces: pick a typical request, count the spans, and ask whether you could identify the slow operation in under 10 seconds. If not, you've found your first refactoring target. - Start with minimal instrumentation — auto-instrumentation plus key business operations — then add spans only when you can't debug a specific problem. You can always add granularity; removing it requires code changes. Let debugging needs drive instrumentation, not the quest for "complete" visibility. + +Audit your current traces: pick a typical request, count the spans, and ask whether you could identify the slow operation in under 10 seconds. If not, you've found your first refactoring target. diff --git a/src/content/articles/opentelemetry-span-design-granularity-overhead/pdf.mdx b/src/content/articles/opentelemetry-span-design-granularity-overhead/pdf.mdx index 50aa3b27e..72c2fc007 100644 --- a/src/content/articles/opentelemetry-span-design-granularity-overhead/pdf.mdx +++ b/src/content/articles/opentelemetry-span-design-granularity-overhead/pdf.mdx @@ -15,7 +15,6 @@ import overInstrumentedDiagram from "./diagrams/over-instrumented-trace-wall-of- import readableTraceDiagram from "./diagrams/readable-trace-waterfall-with-clear-hierarchy.jpg" import requestDiagram from "./diagrams/request-instrumentation-sequence.jpg" -*[DB]: Database *[gRPC]: gRPC Remote Procedure Calls *[OTel]: OpenTelemetry *[OTLP]: OpenTelemetry Protocol @@ -25,18 +24,18 @@ import requestDiagram from "./diagrams/request-instrumentation-sequence.jpg" More spans provide more visibility. That's the intuition, anyway. But each span has costs: CPU overhead for creation, memory for attributes, network bandwidth for export, storage in your backend, and cognitive load when you're actually trying to read the trace. A request that creates 500 spans might have excellent granularity, but the trace waterfall becomes a solid block of color — no white space, no visible hierarchy. Debugging means scrolling through hundreds of spans looking for the slow one. Storage costs explode. The instrumentation itself becomes a performance concern. -Span design requires judgment, not just enthusiasm for visibility. The goal is __enough__ spans to debug problems, not so many that you create new ones. +Span design requires judgment, not just enthusiasm for visibility. The goal is _enough_ spans to debug problems, not so many that you create new ones. I learned this the hard way. A team I worked with instrumented a new service with spans for every function call, database query, cache lookup, and external request. Thorough, right? A typical request generated 200+ spans. The trace backend showed a waterfall of solid color — no gaps, no obvious structure. Finding the slow operation meant scrolling through hundreds of spans, mentally filtering out the noise. Storage costs tripled in a month. -They refactored. Instrumented service boundaries, significant I/O operations, and error paths. Span count dropped to 15-20 per request. Traces became readable. The critical path was obvious at a glance. Storage costs dropped 90%. Debugging time went from minutes to seconds. - -The lesson: granularity without readability is noise, not observability. - The span count that's "right" depends on your debugging needs. A payment service might need fine-grained spans to audit every step. A high-throughput cache might need minimal spans to avoid overhead. There's no universal number — but there are universal principles. +They refactored. Instrumented service boundaries, significant I/O operations, and error paths. Span count dropped to 15-20 per request. Traces became readable. The critical path was obvious at a glance. Storage costs dropped 90%. Debugging time went from minutes to seconds. + +The lesson: granularity without readability is noise, not observability. + ## Span Fundamentals Before diving into granularity decisions, it helps to understand what you're actually creating when you start a span. Every span carries overhead, and that overhead scales with how much data you attach to it. @@ -46,15 +45,15 @@ Before diving into granularity decisions, it helps to understand what you're act A span has six core components: trace_id (16 bytes) is shared across every span in the trace — it\'s what lets your backend stitch together spans from different services. The span_id (8 bytes) uniquely identifies this span. The parent_span_id links to the parent span, creating the hierarchy that becomes your waterfall visualization.', }, { lead: 'Naming', - text: 'Tells you what operation this span represents. The name should describe the operation (`HTTP GET /api/orders/{orderId}`), and the kind indicates the span\'s role in the request flow: `SERVER` for handling incoming requests, `CLIENT` for outgoing calls, `PRODUCER` and `CONSUMER` for async messaging, `INTERNAL` for operations that don\'t cross network boundaries.', + text: 'Tells you what operation this span represents. The name should describe the operation (HTTP GET /api/orders/{orderId}), and the kind indicates the span\'s role in the request flow: SERVER for handling incoming requests, CLIENT for outgoing calls, PRODUCER and CONSUMER for async messaging, INTERNAL for operations that don\'t cross network boundaries.', }, { lead: 'Timing', @@ -62,15 +61,15 @@ A span has six core components: }, { lead: 'Status', - text: 'Records the outcome: `UNSET` (no status set), `OK` (operation succeeded), or `ERROR` (operation failed). For errors, you can include a message describing what went wrong.', + text: 'Records the outcome: UNSET (no status set), OK (operation succeeded), or ERROR (operation failed). For errors, you can include a message describing what went wrong.', }, { lead: 'Attributes', - text: 'Key-value pairs that add context. Some follow OpenTelemetry semantic conventions (`http.method`, `db.system`), others are custom to your domain (`order.id`, `customer.tier`). Attributes are where most of your per-span overhead comes from — each attribute requires memory allocation and serialization.', + text: 'Key-value pairs that add context. Some follow OpenTelemetry semantic conventions (http.method, db.system), others are custom to your domain (order.id, customer.tier). Attributes are where most of your per-span overhead comes from — each attribute requires memory allocation and serialization.', }, { lead: 'Events', - text: 'Timestamped logs within the span\'s lifetime. Use them for milestones: `cache.miss`, `retry.attempt`, `validation.failed`. Events are lighter than child spans but still have overhead.', + text: 'Timestamped logs within the span\'s lifetime. Use them for milestones: cache.miss, retry.attempt, validation.failed. Events are lighter than child spans but still have overhead.', }, { lead: 'Links', @@ -200,15 +199,19 @@ Span hierarchy creates the waterfall visualization. Parent-child relationships s The question "should this operation have a span?" comes up constantly. Here's how I think about it.
@@ -526,7 +551,10 @@ Head sampling reduces instrumentation overhead but might drop interesting traces A trace is only useful if you can read it. I've seen traces that technically contain all the information needed to debug a problem, but the information is buried in noise. Here are the anti-patterns to avoid. HTTP GET /api/orders, db.query orders).', }, { lead: 'Attribute explosion.', @@ -557,6 +585,7 @@ Compare this over-instrumented trace to the readable one shown earlier: @@ -568,33 +597,37 @@ The over-instrumented trace has 15 spans where 4 would suffice. You can't see at Span names should answer "what operation on what resource?" without requiring you to read the code. The OpenTelemetry semantic conventions provide a good starting point. HTTP {method} {route}: HTTP GET /api/orders/{orderId}. The route should be parameterized (with {orderId}, not the actual ID) to keep cardinality low.', }, { lead: 'For HTTP clients:', - text: 'Use `HTTP {method}` with the target service in an attribute: `HTTP GET` with `peer.service: payments-api`.', + text: 'Use HTTP {method} with the target service in an attribute: HTTP GET with peer.service: payments-api.', }, { lead: 'For databases:', - text: 'Use `db.{operation} {table}`: `db.query orders`, `db.insert users`. Don\'t include the full SQL query in the name — that goes in an attribute if needed.', + text: 'Use db.{operation} {table}: db.query orders, db.insert users. Don\'t include the full SQL query in the name — that goes in an attribute if needed.', }, { lead: 'For caches:', - text: 'Use `cache.{operation}`: `cache.get`, `cache.set`. Include the key pattern in an attribute.', + text: 'Use cache.{operation}: cache.get, cache.set. Include the key pattern in an attribute.', }, { lead: 'For internal operations:', - text: 'Use `{domain}.{operation}`: `order.validate`, `payment.process`, `inventory.reserve`.', + text: 'Use {domain}.{operation}: order.validate, payment.process, inventory.reserve.', }, ]} />
HTTP GET /api/orders/{id', 'handler', 'No indication of method or resource'], }, { th: 'Database', - td: ['`db.query orders`', '`SELECT * FROM...`', 'SQL syntax noise, potential sensitive data'], + td: ['db.query orders', 'SELECT * FROM...', 'SQL syntax noise, potential sensitive data'], }, { th: 'Cache', - td: ['`cache.get`', '`redis`', 'Technology not operation — what did you do?'], + td: ['cache.get', 'redis', 'Technology not operation — what did you do?'], }, { th: 'Business logic', - td: ['`order.validate`', '`validateOrder`', 'Function name leaks implementation detail'], + td: ['order.validate', 'validateOrder', 'Function name leaks implementation detail'], }, { th: 'External API', - td: ['`HTTP POST`', '`API call`', 'No method, no way to distinguish calls'], + td: ['HTTP POST', 'API call', 'No method, no way to distinguish calls'], }, ], }, @@ -638,7 +671,10 @@ Span names should answer "what operation on what resource?" at a glance. A devel Attributes carry the metadata that makes traces useful. Choose them carefully — every attribute consumes storage, affects query performance, and either helps or clutters your debugging experience. @@ -862,7 +901,7 @@ The sequence diagram shows how these pieces fit together. The root span encompas Batch processing is where span proliferation gets dangerous. Creating a span per item in a 10,000-item batch generates 10,000 spans — overwhelming your collector, bloating storage, and producing unreadable waterfalls. { Code: Single span for a batch with events for failures. { Code: Chunked spans for large batches. Start with minimal instrumentation — auto-instrumentation plus key business operations — then add spans only when you can't debug a specific problem. You can always add granularity; removing it requires code changes. Let debugging needs drive instrumentation, not the quest for "complete" visibility. + +The best traces answer three questions: what happened, where time was spent, and what failed. The worst traces are walls of noise that hide the signal you need. + +Start by auditing your current traces: pick a typical request, count the spans, and ask whether you could identify the slow operation in under 10 seconds. If not, you've found your first refactoring target. diff --git a/src/content/articles/performance-testing-load-models-benchmark-accuracy/index.mdx b/src/content/articles/performance-testing-load-models-benchmark-accuracy/index.mdx index 5f7f94722..52c4f9630 100644 --- a/src/content/articles/performance-testing-load-models-benchmark-accuracy/index.mdx +++ b/src/content/articles/performance-testing-load-models-benchmark-accuracy/index.mdx @@ -18,11 +18,11 @@ import fourRequestsDiagram from "./diagrams/four-requests-averaging-100ms-but-on The post-mortem was awkward. A team had spent three weeks building a performance test suite for their new API gateway. The benchmark showed 50,000 RPS with 2ms P99 latency. Leadership signed off on the deployment. Production fell over at 5,000 RPS. -The engineers weren't incompetent — they were victims of performance testing's hidden traps. Their benchmark measured __something__, just not anything useful for predicting production behavior. The load generator backed off when the system struggled (coordinated omission). The dashboards showed averages that hid catastrophic tail latency. The test environment bore no resemblance to production. +The engineers weren't incompetent — they were victims of performance testing's hidden traps. Their benchmark measured _something_, just not anything useful for predicting production behavior. The load generator backed off when the system struggled (coordinated omission). The dashboards showed averages that hid catastrophic tail latency. The test environment bore no resemblance to production. This pattern repeats constantly. Teams run benchmarks, get impressive numbers, deploy with confidence, and watch production burn. The numbers were accurate; they just didn't answer the right question. -Two mistakes cause most of the damage: __coordinated omission__ and __misusing averages__. Understanding these will save you from benchmarks that create false confidence before production incidents. +Two mistakes cause most of the damage: _coordinated omission_ and _misusing averages_. Understanding these will save you from benchmarks that create false confidence before production incidents. ## The Coordinated Omission Trap @@ -86,7 +86,8 @@ Averages collapse under outliers in both directions. A single 10-second timeout __Percentiles__ tell a more honest story:
@@ -134,7 +136,7 @@ Report percentiles, not averages. Show distributions, not just summary statistic Coordinated omission and average-worship are the most common problems, but they're not the only ones. A few other factors regularly invalidate benchmarks: -Each of these deserves deeper treatment than space allows here. The point is that performance testing has many failure modes, and getting impressive numbers is easy — getting __meaningful__ numbers requires understanding all the ways benchmarks can mislead. +Each of these deserves deeper treatment than space allows here. The point is that performance testing has many failure modes, and getting impressive numbers is easy — getting _meaningful_ numbers requires understanding all the ways benchmarks can mislead. The most common performance testing mistake: measuring throughput without realistic load patterns. Synthetic load (constant RPS, uniform distribution) produces numbers that have no relationship to production performance. +The lesson: a benchmark is a model of reality. If the model is wrong, the predictions are worthless. A "fast" benchmark that doesn't reflect reality is worse than no benchmark — it creates false confidence that leads to production incidents. + +Getting useful numbers requires realistic load models that match production traffic, proper warmup to reach steady state, environment parity with production infrastructure, and statistical rigor to interpret results correctly. This article covers each of those requirements. + ## Load Model Fundamentals -A load model defines __what__ you're testing. It's the specification that turns "test our API" into something concrete and reproducible. Without a well-defined load model, you're just throwing requests at a server and hoping the numbers mean something. +A load model defines _what_ you're testing. It's the specification that turns "test our API" into something concrete and reproducible. Without a well-defined load model, you're just throwing requests at a server and hoping the numbers mean something. Before diving into load model design, a quick note on tooling. I use k6 for load testing, integrated into CI/CD pipelines. k6 scripts are JavaScript/TypeScript, which makes them easy to version control and review alongside application code. For alternatives, Locust (Python) and Gatling (Scala) are solid choices with their own ecosystems. Artillery is another JavaScript option with good AWS integration. @@ -52,29 +52,33 @@ The bigger question is where to run these tests. Quick regression checks (2-5 mi A complete load model has four components: throughput (how much load), distribution (when requests arrive), workload mix (what operations), and user behavior (how users interact). @@ -182,8 +186,10 @@ Code: Test harness with weighted endpoint selection and realistic think time. Different load patterns serve different purposes. Choose based on what you're trying to learn.
@@ -228,7 +233,10 @@ A cold system behaves nothing like a warm one. If you measure performance during Several systems need time to reach optimal performance: Never include warmup data in your results. Tag phases separately and apply thresholds only to the measurement phase. Warmup latencies can be 10-100x higher than steady-state and will completely skew your statistics. +How long should warmup run? The honest answer is "until metrics stabilize." Watch P99 latency over time — when it stops improving and variance drops below 5%, you've reached steady state. For simple services, 30-60 seconds often suffices. Complex services with many code paths and large caches may need 2-5 minutes. When in doubt, run longer and look at the metrics. + ## Statistical Rigor Most performance reports show averages. Averages lie. @@ -324,7 +332,11 @@ An average latency of 100ms might mean all requests completed in roughly 100ms. Averages also collapse under outliers. A single 10-second timeout in a thousand requests shifts the average dramatically, even though 99.9% of users had a good experience. Conversely, if that timeout represents a real failure mode that will hit 1% of users in production, the average masks it. @@ -451,30 +463,35 @@ The gap between test and production environments creates systematic errors that Four categories of factors determine whether your test environment will produce representative results:
@@ -511,7 +527,7 @@ Four categories of factors determine whether your test environment will produce Perfect parity is expensive — a full production replica for testing might double your infrastructure costs. The goal is sufficient parity: matching the factors that actually affect your specific workload's performance characteristics. The question isn't "how fast is my system?" but "how will my system behave under production conditions?" Build tests that answer the second question, even if the numbers are less impressive than synthetic benchmarks. + +What does "production-equivalent infrastructure" cost? For a typical web service, expect to spend $500-2,000/month on a dedicated performance testing environment — matching your production instance types, database size, and cache configuration. This sounds expensive until you compare it to the cost of a performance-related outage: lost revenue, engineering time debugging in production, and customer trust. A single incident that a proper benchmark would have caught easily justifies a year of testing infrastructure. + +The cost of a misleading benchmark — false confidence followed by production incidents — far exceeds the cost of building proper infrastructure. A "fast" benchmark that doesn't reflect reality is worse than no benchmark at all. diff --git a/src/content/articles/platform-architecture-control-plane-data-plane-separation/index.mdx b/src/content/articles/platform-architecture-control-plane-data-plane-separation/index.mdx index 9c136c22a..ac4e41927 100644 --- a/src/content/articles/platform-architecture-control-plane-data-plane-separation/index.mdx +++ b/src/content/articles/platform-architecture-control-plane-data-plane-separation/index.mdx @@ -16,9 +16,9 @@ import controlPlaneDiagram from "./diagrams/control-plane-pushes-desired-state-t *[RBAC]: Role-Based Access Control *[VPC]: Virtual Private Cloud -A platform team I worked with hit a wall at fifty teams. Their internal Kubernetes platform had grown organically — API server, controllers, etcd, worker nodes all running together because it was simpler that way. Then Monday mornings started hurting. Everyone deploying at once slowed the API server enough that running workloads couldn't get their service endpoints updated. A bug in a custom controller caused repeated panics that prevented all controllers from reconciling — deployments wouldn't scale, services wouldn't update, affecting __all__ tenants. Control plane upgrades required scheduling maintenance windows across every team. +A platform team I worked with hit a wall at fifty teams. Their internal Kubernetes platform had grown organically — API server, controllers, etcd, worker nodes all running together because it was simpler that way. Then Monday mornings started hurting. Everyone deploying at once slowed the API server enough that running workloads couldn't get their service endpoints updated. A bug in a custom controller caused repeated panics that prevented all controllers from reconciling — deployments wouldn't scale, services wouldn't update, affecting _all_ tenants. Control plane upgrades required scheduling maintenance windows across every team. -The fix wasn't more hardware. It was architectural: dedicated control plane cluster, separate data plane clusters per environment, GitOps for configuration sync. The pattern that made this possible comes from networking, where routers have long separated the control plane (where routing decisions happen) from the data plane (where packets actually flow). Platform engineering borrowed this separation because it solves the same fundamental problem — scaling decision-making independently from execution. +The fix wasn't more hardware. It was architectural: dedicated control plane cluster, separate data plane clusters per environment, GitOps for configuration sync. The pattern that made this possible comes from networking, where routers have long separated the control plane (where routing decisions happen) from the data plane (where packets actually flow). Platform engineering borrowed this separation because it solves the same fundamental problem — scaling decision-making independently from execution. ## The Separation That Scales @@ -31,7 +31,7 @@ The __data plane__ is where actual work happens. It executes workloads, routes t The characteristics differ in ways that matter for architecture:
@@ -77,7 +79,10 @@ Control plane and data plane separation exists to serve multi-tenancy. Without m The fundamental question: how much isolation do tenants need, and what are you willing to pay for it?
The most common platform architecture mistake: building for single-tenant simplicity, then retrofitting multi-tenancy. Design separation into your abstractions from the start, even if you deploy everything together initially.