Cloud Tech

Search

30 results for “troubleshooting

Search results

Blog

Terraform Troubleshooting Guide: Fix the Errors Engineers Actually Hit

Diagnose Terraform initialization, validation, provider, authentication, state, drift, import, replacement, timeout, and CI failures with a safe workflow.

Blog

Kubernetes OOMKilled Troubleshooting Guide

A practical Kubernetes OOMKilled troubleshooting guide using kubectl describe, previous logs, events, metrics, requests, limits, QoS, and node memory pressure.

DevOps

AWS Fundamentals

How AWS accounts, regions, identity, and core services fit together, with a practical CLI and access-denied troubleshooting workflow.

DevOps

Azure Fundamentals

How Azure management groups, subscriptions, resource groups, and Resource Manager fit together, with practical deployment troubleshooting.

DevOps

Cloud IAM Fundamentals

How identities, roles, policies, scopes, and temporary credentials map across AWS, Azure, and Google Cloud, with practical access troubleshooting.

DevOps

Docker Fundamentals

How images, containers, volumes, and networks fit together in Docker's runtime model, and the shift from installing software to running images.

DevOps

Git Fundamentals

How Git moves changes through the working tree, index, and history, including safe recovery from conflicts, rewrites, and mistaken commits.

DevOps

GitOps Principles

Why treating Git as the single source of truth for cluster state changes how deployments, rollbacks, and audits actually work.

DevOps

Kubernetes Fundamentals

Learn Kubernetes through Pods, Deployments, Services, probes, resources, rollouts, and an official-docs-based troubleshooting workflow.

Security

Kubernetes Security

How RBAC's additive model, Pod Security Standards, and NetworkPolicy fit together, and why each surprises people used to simpler permissions.

DevOps

Microsoft Entra ID (formerly Azure AD)

How Microsoft Entra tenants, app registrations, service principals, managed identities, and Azure RBAC govern human and workload access.

DevOps

Observability Fundamentals

How logs, metrics, and traces answer different questions, why monitoring is not debuggability, and the cardinality trap that breaks metrics systems.

DevOps

Terraform Basics

How Terraform's state model, providers, and plan/apply workflow make infrastructure reviewable like code, plus the state pitfalls that trip up teams.

AWS Fundamentals

What is the fundamental unit of isolation in AWS, and how does that differ from a single resource-group boundary in Azure?

In AWS, the account itself is the fundamental security and billing isolation boundary, every resource lives inside exactly one account, and account-level separation is what actually contains blast radius (a compromised credential in one account cannot directly touch resources in another). This differs from Azure, where a single subscription can contain many resource groups as an additional lifecycle boundary beneath it. AWS has no equivalent nested container inside an account for "delete everything in this group together," which is why multi-account strategies (via AWS Organizations) do the job that resource groups partly do in Azure, at the account level instead of a sub-account level.

AWS Fundamentals

What is a Service Control Policy (SCP), and what is the one thing it does not do?

An SCP is a policy attached to an AWS Organizations root, organizational unit, or member account that defines the maximum available permissions for every identity in that account, including that account's own administrators and its root user. What an SCP does not do is grant any permission by itself, it only sets a ceiling; an identity still needs an actual IAM allow (from an identity-based or resource-based policy) within that ceiling to do anything. An SCP with no matching IAM allow underneath it results in access denied, not access granted, which is the most common misunderstanding of how SCPs work. One exception worth knowing: SCPs never apply to the organization's management account itself, only to member accounts.

AWS Fundamentals

Why would an organization use multiple AWS accounts instead of one account holding all resources?

Separate accounts per environment (production, staging, development) or per team give a hard isolation boundary that a single account with tags or naming conventions cannot: a mistake or compromised credential in a development account cannot reach production resources at all, rather than merely being restricted by IAM policy within the same account. It also gives cleaner cost attribution (billing rolls up per account), independent service quotas, and a natural blast-radius limit for security incidents. AWS Organizations, and patterns built on top of it like a landing zone, exist specifically to make many accounts manageable, centralized billing, centralized logging, and org-wide SCPs, without losing that isolation.

Azure Fundamentals

What is the difference between a resource group and a subscription in Azure?

A subscription is a billing and access-management boundary; it's tied to an agreement with Microsoft, has its own spending limits and quotas, and is typically the unit organizations use to separate environments (production vs. non-production) or business units. A resource group is a logical container inside a subscription that groups related resources (a VM, its disks, its network interface) that share the same lifecycle, created and deleted together. Deleting a resource group deletes everything in it, which makes resource groups the practical unit of "this is one deployable thing," while subscriptions are the practical unit of "this is one billing and governance boundary."

Azure Fundamentals

What is Azure Resource Manager (ARM) and why does every Azure operation go through it?

ARM is the deployment and management layer that every Azure operation, whether from the Portal, CLI, PowerShell, or an ARM/Bicep template, ultimately goes through. It provides a consistent API surface, handles authentication and authorization checks against Azure RBAC, and is what enables declarative deployment (submit a template describing desired resources, ARM figures out what to create/update). Because every path converges on ARM, access control and activity logging are consistent regardless of which tool was used to make a change.

Azure Fundamentals

How do management groups extend governance above the subscription level?

Management groups let an organization apply policies (via Azure Policy) and role assignments (via Azure RBAC) across multiple subscriptions at once, instead of configuring each subscription independently. They form a hierarchy above subscriptions, a root management group can contain child management groups (e.g., by department or environment type), each containing multiple subscriptions, so a single policy assignment at the right level of that hierarchy can enforce a rule (like "no public IP addresses" or "must use approved regions") across every subscription beneath it.

Cloud IAM Fundamentals

What is the difference between authentication and authorization in a cloud IAM context?

Authentication answers "who is making this request", verifying an identity via credentials, a token, or a federated login. Authorization answers "is this identity allowed to do this specific action on this specific resource", evaluated after authentication succeeds, by checking the identity's attached policies against the requested action. A request can be perfectly authenticated (the caller genuinely is who they claim) and still be denied, because authorization is a separate check against what that identity is actually permitted to do.

Cloud IAM Fundamentals

What does the principle of least privilege mean in practice, and why is it hard to maintain over time?

Least privilege means granting an identity only the specific permissions it needs to do its job, nothing broader "to be safe" or "to save time." It's hard to maintain because permissions tend to accumulate, someone gets a broad role to unblock a one-time task and it's never revoked, or a service starts with wildcard permissions during initial development and nobody narrows them before shipping. Maintaining least privilege requires ongoing review (access audits, unused-permission detection), not just a careful initial setup, because the natural drift over time is always toward more access, not less.

Cloud IAM Fundamentals

What is the difference between a role and a policy in most cloud IAM systems?

The word "role" is provider-specific. In AWS, an IAM role is an assumable principal with policies attached. In Azure RBAC and Google Cloud IAM, a role is primarily a reusable collection of permissions; a role assignment or IAM policy binding grants that role to a principal at a scope. Always reduce the model to four questions: which principal, which permissions, on which resource scope, under which conditions. Translating the word "role" literally between providers causes dangerous design mistakes.

Docker Fundamentals

What is the difference between a Docker image and a Docker container?

An image is a read-only, layered filesystem snapshot plus metadata (entrypoint, exposed ports, environment); it never changes once built and can be shared through a registry. A container is a running (or stopped) instance of an image: Docker adds a thin writable layer on top of the image's read-only layers and starts a process inside an isolated namespace. You can start many independent containers from the same image, each with its own writable layer and state, the same way many processes can run from the same binary on a normal OS.

Docker Fundamentals

Why does data written inside a container disappear when the container is removed?

Anything a container writes lands in its own writable layer, which is deleted along with the container by `docker rm`. That is deliberate; it is what makes containers disposable and reproducible, a fresh container from the same image always starts from the same known state. Anything that actually needs to survive a container's lifecycle (a database's data files, uploaded assets) has to live outside that writable layer, in a named volume or a bind mount, which Docker mounts into the container at a chosen path but manages independently of the container itself.

Docker Fundamentals

Why is publishing a port with `-p 8080:80` different from the container just "having" port 80?

A container's ports exist only on its own private network namespace by default; nothing on the host or outside can reach them until Docker explicitly forwards a host port to it. `-p 8080:80` tells Docker's network layer to forward the host's port 8080 to port 80 inside the container's namespace, host port first, container port second. Leaving a port `EXPOSE`d in a Dockerfile only records metadata/documentation, it has no effect on connectivity at all: another container on the same Docker network can already reach any port the first container is listening on, EXPOSE or not. Publishing to the host is the one thing that always requires an explicit `-p`.

Git Fundamentals

What is the difference between the working directory, the staging area, and a commit in Git?

The working directory is the actual files on disk, whatever state you've left them in. The staging area (the "index") is a snapshot of exactly what will go into the next commit; `git add` copies changes from the working directory into it, one file or hunk at a time, which is why you can commit only part of what you've changed. A commit is a permanent, immutable snapshot of the staging area at the moment you ran `git commit`, plus a pointer to its parent commit, which is what forms the project's history graph. Understanding that staging is a separate, explicit step - not just "what's changed" - explains why `git status` shows both staged and unstaged changes for the same file.

Git Fundamentals

What is the practical difference between `git merge` and `git rebase`?

Both bring one branch's commits into another, but they produce different history shapes. `git merge` creates a new merge commit with two parents, preserving exactly how the branches diverged and came back together; nothing is rewritten, which is why merge is safe on shared/public branches. `git rebase` replays your branch's commits one by one on top of the target branch's tip, producing a linear history with no merge commit, but every replayed commit gets a new hash. That rewriting is why rebase should be avoided on branches other people have already pulled; their history and yours will diverge as soon as they fetch the rewritten commits.

Git Fundamentals

What is the difference between `git reset` and `git revert`, and when should you use each?

git reset moves the current branch pointer (and optionally the staging area and working directory) to a different commit, effectively rewriting history as if the reset-past commits never happened on this branch - fine for commits that only exist locally and haven't been pushed. git revert creates a brand new commit that applies the inverse of a previous commit's changes, leaving history intact and additive. Because it doesn't rewrite anything, revert is the safe choice for undoing a commit that's already been pushed and pulled by others; reset --hard on shared history causes exactly the same divergence problem as a rebase on a shared branch.

GitOps Principles

What makes a workflow "GitOps" rather than just "we deploy from CI"?

The defining property is a pull-based reconciliation loop, not just that Git triggers a deploy. A GitOps agent (Argo CD, Flux) runs inside the cluster and continuously compares the live state against what's declared in a Git repository, pulling and applying any drift, with or without a new commit. A CI pipeline that runs `kubectl apply` on push is push-based: it changes things once, on trigger, and has no ongoing awareness of whether the cluster later drifts from that state. GitOps closes that loop continuously and treats Git, not the cluster, as the source of truth.

GitOps Principles

How does GitOps make rollbacks different from a traditional deployment rollback?

In a traditional deploy, rolling back means re-running a deployment process with an older artifact reference, a distinct operation from a normal deploy. In GitOps, a rollback is just a Git revert: since the desired cluster state is fully described by the repository at any commit, reverting to a previous commit and letting the reconciliation loop pick it up produces the previous cluster state through the exact same mechanism as any other change. There is no separate "rollback pipeline" to maintain or that can itself have bugs.

Search results for “troubleshooting” | Cloud Tech by Victor