Infrastructure & Platform

DevOps Engineer

Master Linux, Docker, Kubernetes, Terraform, CI/CD pipelines, and cloud observability.

Progression:0%
0 / 9 Milestones Mastered
L1

1. Beginner

System administration and terminal mastery.

Advanced Linux Internals

essential

Kernel namespaces, cgroups, systemd, iptables/nftables, storage volume types.

Namespaces & CgroupsProcess schedulingNetworking socket statesSystemd unit files
L2

2. Foundation

Networking protocols and version controlled infrastructure.

Cloud Networking & Protocols

essential

CIDR blocks, VPC subnets, NAT gateways, route tables, DNS, SSL/TLS termination.

VPC architectureSubnettingBGP & AnycastTLS mutual auth (mTLS)
L3

3. Core Skills

Containers and automated build pipelines.

Docker & Container Runtimes

essential

OCI spec, layered caching, rootless containers, multi-arch builds, security scanning.

Layer optimizationMulti-stage buildsTrivy vulnerability scanDistroless images

CI/CD Pipeline Automation

essential

GitHub Actions, GitLab CI, semantic release, artifact registries, canary rollouts.

Pipeline syntaxSecret managementMatrix buildsAutomated rollbacks
L4

4. Framework

Container orchestration with Kubernetes and Infrastructure as Code.

Kubernetes (K8s) Production Architecture

essential

Pods, Deployments, Services, Ingress, ConfigMaps, Secrets, HPA, Helm charts.

Control plane vs Worker nodesService meshesHorizontal Pod AutoscalerHelm templating

Terraform & OpenTofu (IaC)

essential

HCL syntax, state management, remote backends with S3/DynamoDB, module patterns.

State locksTerraform modulesDrift detectionPolicy as code (OPA)
L5

5. Real Projects

Deploy and monitor realistic cloud infrastructure.

Production EKS/GKE GitOps Deployment

essential

ArgoCD gitops controller, automated certificate manager, ingress-nginx.

ArgoCDCert-ManagerPrometheus Operator
L6

6. Advanced

Site Reliability Engineering (SRE) and Observability.

Prometheus, Grafana & OpenTelemetry

essential

Metrics, logs, traces (LGTM stack), golden signals, alerting rules, SLO/SLA.

PromQL queriesDistributed tracingError budgetsAlertmanager routing
L7

7. Interview Prep

Incident post-mortems, failure scenarios, and architectural defense.

Incident Response & Troubleshooting

essential

OOMKilled pods, high CPU throttle, DNS latency spikes, network partition triage.

Post-mortem cultureRoot cause analysisZero-downtime migrations

Hands-on Portfolio Projects for this Roadmap

Multi-region Highly Available Terraform VPC on AWS
GitOps Continuous Delivery with ArgoCD & Kubernetes
Full OpenTelemetry Observability Dashboard with Alerts