Search
20 results for “kubernetes”
Search results
Kubernetes Fundamentals
Learn Kubernetes through Pods, Deployments, Services, probes, resources, rollouts, and an official-docs-based troubleshooting workflow.
Kubernetes Security
How RBAC's additive model, Pod Security Standards, and NetworkPolicy fit together, and why each surprises people used to simpler permissions.
Kubernetes OOMKilled Troubleshooting Guide
A practical Kubernetes OOMKilled troubleshooting guide using kubectl describe, previous logs, events, metrics, requests, limits, QoS, and node memory pressure.
What does the Kubernetes control loop actually do?
Every Kubernetes controller (Deployment, ReplicaSet, etc.) runs a reconciliation loop: it continuously compares the desired state (what you declared in a manifest, stored in etcd) against the observed actual state of the cluster, and takes action to close any gap. If you declared 3 replicas and only 2 Pods are running, the ReplicaSet controller creates one more. This is the same declarative, converge-toward-desired-state model as Terraform, but running continuously and automatically rather than on-demand.
How does Kubernetes RBAC decide whether a request is allowed, and can you write an explicit deny rule?
RBAC is default-deny and purely additive: a request is denied unless some Role/RoleBinding or ClusterRole/ClusterRoleBinding explicitly grants it, and there is no such thing as an explicit deny rule. Every applicable binding's permissions are unioned together, so restricting access means removing a grant (or removing the subject from a binding), not adding a deny statement on top of an existing grant, an approach that works cleanly for grant-based access but has no mechanism for "allow everything except X" within RBAC itself.
GitOps Principles
Why treating Git as the single source of truth for cluster state changes how deployments, rollbacks, and audits actually work.
Observability Fundamentals
How logs, metrics, and traces answer different questions, why monitoring is not debuggability, and the cardinality trap that breaks metrics systems.
What makes a workflow "GitOps" rather than just "we deploy from CI"?
The defining property is a pull-based reconciliation loop, not just that Git triggers a deploy. A GitOps agent (Argo CD, Flux) runs inside the cluster and continuously compares the live state against what's declared in a Git repository, pulling and applying any drift, with or without a new commit. A CI pipeline that runs `kubectl apply` on push is push-based: it changes things once, on trigger, and has no ongoing awareness of whether the cluster later drifts from that state. GitOps closes that loop continuously and treats Git, not the cluster, as the source of truth.
How does GitOps make rollbacks different from a traditional deployment rollback?
In a traditional deploy, rolling back means re-running a deployment process with an older artifact reference, a distinct operation from a normal deploy. In GitOps, a rollback is just a Git revert: since the desired cluster state is fully described by the repository at any commit, reverting to a previous commit and letting the reconciliation loop pick it up produces the previous cluster state through the exact same mechanism as any other change. There is no separate "rollback pipeline" to maintain or that can itself have bugs.
Why does GitOps improve auditability compared to engineers running kubectl or terraform apply directly?
Every change to cluster state has to go through a Git commit, which means it inherits Git's existing history, authorship, and (if branch protection is configured) pull-request review, automatically. Direct `kubectl apply` access leaves no equivalent trail: two changes with the same effect are indistinguishable, there's no required review step, and reconstructing "who changed what and why" after an incident means digging through cluster event logs instead of reading a linear, reviewed commit history.
What is the difference between a Pod, a Deployment, and a Service?
A Pod is the smallest deployable unit, one or more containers that share network and storage, scheduled together onto a node. Pods are ephemeral and disposable; Kubernetes recreates them freely and their IPs change every time. A Deployment manages a set of identical Pods (a ReplicaSet under the hood), handling rolling updates, rollbacks, and keeping the desired replica count running even as individual Pods die. A Service gives that changing set of Pods a stable network identity, a fixed virtual IP and DNS name, so other things in the cluster don't need to track individual Pod IPs, which change constantly.
Why can't you rely on a Pod's IP address for service discovery?
Pods are ephemeral by design, Kubernetes kills and recreates them constantly (failed health checks, node drains, rolling deployments, autoscaling), and every new Pod gets a brand-new IP address. Hardcoding or caching a Pod IP breaks the moment that Pod is replaced. A Service solves this by providing a stable virtual IP and DNS name that always routes to whichever Pods currently match its label selector, regardless of how many times the underlying Pods have been replaced.
What is the difference between readiness, liveness, and startup probes?
Readiness decides whether a Pod should receive Service traffic; a failed readiness probe removes the Pod from eligible backends without restarting it. Liveness decides whether a stuck container should be restarted. A startup probe protects slow-starting containers by delaying readiness and liveness checks until startup succeeds. Reusing one strict check for all three can create restart loops or route traffic too early.
How would you troubleshoot a Service that exists but returns no response?
Work from the application outward: confirm the selected Pods are Ready and serving on the expected container port, compare the Service selector with Pod labels, inspect EndpointSlices to verify Kubernetes discovered backends, confirm port and targetPort, then test Service DNS and IP from inside the cluster. An empty EndpointSlice usually points to a selector/readiness mismatch; healthy endpoints with failed DNS or routing move the investigation to cluster networking.
What is the difference between the Baseline and Restricted Pod Security Standards levels, and why are they cumulative?
Baseline blocks the most well-known container privilege-escalation paths, privileged containers, host namespaces, hostPath volumes, dangerous Linux capabilities, while still allowing a fairly permissive pod spec otherwise. Restricted inherits every Baseline rule and adds real hardening on top: it requires running as non-root, forbids privilege escalation outright, requires a restricted seccomp profile, and requires dropping all Linux capabilities except NET_BIND_SERVICE. A read-only root filesystem is not part of either standard, it's a separate hardening measure some organizations layer on as their own policy, on top of, not as part of, Restricted. They're cumulative by design, Restricted is Baseline plus more, so a workload that passes Restricted automatically satisfies Baseline too, and a cluster can apply different levels per namespace based on how much a given workload can be trusted.
If no NetworkPolicy exists in a namespace, what traffic is allowed between pods, and what changes the moment one NetworkPolicy is applied?
With no NetworkPolicy at all, pods are non-isolated: every pod can send and receive traffic from any other pod, with no restriction in either direction. The moment any NetworkPolicy selects a pod for a given direction (ingress or egress), that pod becomes isolated for that direction specifically, and only the traffic explicitly allowed by an applicable policy's rules gets through from then on; unrelated pods elsewhere in the cluster that no policy selects remain fully open. This is why introducing NetworkPolicy incrementally, rather than all at once, tends to break things: the first policy applied to a namespace can silently cut off traffic nobody had previously needed to declare.
What is the difference between monitoring and observability?
Monitoring means watching a predefined set of signals for known failure modes, dashboards and alerts built around questions you already knew to ask ("is CPU above 80%?"). Observability is a property of a system: how well you can answer new, previously-unasked questions about its internal state using only its external outputs (logs, metrics, traces), without shipping new code. Monitoring tells you something is wrong; observability is what lets you figure out why, including for failure modes nobody anticipated when the dashboards were built.
What distinct question does each of logs, metrics, and traces answer?
Metrics answer "what is happening, in aggregate, over time", cheap to store, good for dashboards and alerting thresholds, but they lose individual event detail. Logs answer "what exactly happened in this specific event", full detail but expensive to store and search at scale. Traces answer "where did time go across this one request as it moved through multiple services"; they reconstruct causality and latency across service boundaries that neither logs nor metrics show on their own. A mature observability setup uses all three together, correlated by shared identifiers like a request or trace ID.
What is metric cardinality, and why can it break a monitoring system?
Cardinality is the number of unique label/tag combinations a metric can have. A metric like `http_requests_total{user_id=...}` has cardinality equal to the number of distinct users, potentially millions, because most metrics backends store a separate time series per unique label combination. High-cardinality labels cause a combinatorial explosion in stored time series, which can degrade or crash a metrics backend entirely. The fix is keeping metric labels low-cardinality (route, status code, method) and pushing genuinely high-cardinality data (user IDs, request IDs) into logs or traces instead, where it belongs.
Production-Ready AKS GitOps with Terraform and ArgoCD
The DevOps Project That Finally Made Kubernetes, GitOps, and Terraform Click