Archtin
All articles
DevOpsAIProjects20 min read

5 DevOps + AI + Cloud Projects That Actually Stand Out

Portfolio-grade DevOps projects combining AI, Kubernetes, CI/CD, observability and cloud—with architecture, step-by-step build guides and open-source GitHub repos to learn from.

Why most DevOps projects fail to impress

“Deployed a Docker container to AWS” proves you followed a tutorial. Interviewers want evidence that you can reason about failure, safety, cost and trade-offs. The five projects below each combine AI with real operational plumbing—CI/CD, Kubernetes, observability, cloud APIs—so you end up with something you can defend for 45 minutes, not just demo for 2.

The rule for every project
The AI is the smallest part. The engineering around it—data collection, guardrails, approvals, evaluation and rollback—is what makes the project credible.

Shared foundation for all five

Build this base once and reuse it. It turns each project into an incremental extension rather than a fresh start.

  1. A demo workload: deploy a multi-service app such as Google's Online Boutique or the OpenTelemetry demo so you have realistic traffic and failures.
  2. A local cluster: kind or k3d locally; move to EKS/GKE/AKS later with Terraform.
  3. Observability: kube-prometheus-stack (Prometheus, Alertmanager, Grafana) plus Loki for logs and optionally Tempo for traces.
  4. An LLM interface: one small Python service that wraps model calls, enforces a JSON output schema and logs every prompt/response for evaluation.
  5. A failure injector: scripts (or Chaos Mesh) that break things on purpose—bad config, memory leaks, a dependency outage—so you can test end-to-end.
GitHub repoWhat to learn from it
GoogleCloudPlatform/microservices-demo11-service demo app to deploy and break
open-telemetry/opentelemetry-demoDemo app already instrumented with metrics, logs and traces
prometheus-community/helm-chartskube-prometheus-stack chart for monitoring
grafana/lokiLog aggregation queried by your AI tools
chaos-mesh/chaos-meshControlled failure injection on Kubernetes

1. AI-powered DevOps incident copilot

An agent that receives an alert, gathers evidence from metrics, logs and traces, and produces a probable root cause, a suggested fix and an incident summary.

1Alert fires2Agent receives3Query telemetry4Correlate5Root cause6Summary + fix
Alert-driven investigation: the agent gathers evidence before reasoning about cause.

How to build it

  1. Configure Alertmanager to send a webhook to a FastAPI endpoint.
  2. Define read-only tools: query_prometheus(promql), query_loki(logql), get_pod_events(ns, pod), recent_deployments().
  3. Orchestrate with LangGraph: plan → call tools → summarize evidence → hypothesize → verify one more query → report.
  4. Force structured output: {cause, confidence, evidence[], suggested_actions[]}.
  5. Post the report to Slack and store it; add a 👍/👎 feedback loop.
  6. Evaluate: inject 10 known failures with Chaos Mesh and measure how often the top hypothesis is correct.
Interview talking point
Why read-only tools? Because an investigator that can't change anything can't make an incident worse. Explain how you limited query time ranges and token budgets.
GitHub repoWhat to learn from it
robusta-dev/holmesgptOpen-source AI agent for investigating alerts—closest reference
k8sgpt-ai/k8sgptScans clusters and explains issues with LLMs
robusta-dev/robustaAlert enrichment and automation on Kubernetes
langchain-ai/langgraphAgent orchestration framework

2. AI-powered CI/CD pipeline

A pipeline that adds AI review and risk scoring alongside standard build, test and security gates before deploying to Kubernetes.

1Git push2Build + test3Security scan4AI review5Risk score6Deploy
AI is an extra gate, not a replacement for tests and scanners.

How to build it

  1. GitHub Actions workflow: lint, unit tests, Docker build, Trivy image scan, SonarQube or Semgrep.
  2. A job that sends the PR diff plus test output to an LLM and posts review comments via the GitHub API.
  3. Compute a risk score from objective signals (files touched, migrations, auth code, coverage drop) plus the AI's assessment.
  4. Low risk → auto-deploy with Argo CD; high risk → require manual approval via GitHub environments.
  5. Generate release notes from merged PRs; on failed tests, have the AI explain the failure.
GitHub repoWhat to learn from it
qodo-ai/pr-agentOpen-source AI PR reviewer you can study or extend
aquasecurity/trivyContainer and IaC vulnerability scanning
semgrep/semgrepStatic analysis rules in CI
argoproj/argo-cdGitOps deployment to Kubernetes
argoproj/argo-rolloutsCanary releases with automated analysis

For the underlying flow, see How Code Flows: Local to Production.

3. Self-healing Kubernetes

An agent that detects failures such as CrashLoopBackOff, analyzes logs and metrics, chooses a remediation and executes it—only through a policy gate.

1Failure2Alert3AI proposes4Policy check5Execute6Verify
AI proposes; policy decides; a narrow executor acts; verification closes the loop.

How to build it

  1. Start with a fixed action catalogue: restart pod, rollback deployment, scale replicas, increase memory limit.
  2. The agent outputs one action plus justification; a policy layer (OPA or simple rules) checks namespace, blast radius and rate limits.
  3. The executor uses a Kubernetes ServiceAccount whose RBAC permits only those actions.
  4. After acting, verify recovery (readiness, error rate) and auto-revert or escalate if it didn't help.
  5. Add dry-run mode and an audit log of every proposal and decision.
Important
Never give the AI unrestricted production access. AI → proposed action → policy check → execution.
GitHub repoWhat to learn from it
k8sgpt-ai/k8sgpt-operatorRuns K8sGPT analysis continuously inside the cluster
open-policy-agent/gatekeeperPolicy enforcement for Kubernetes
kyverno/kyvernoKubernetes-native policy engine
kubernetes-client/pythonOfficial client for your executor

Understand the probes you'll be reasoning about in 5 DevOps Concepts Developers Get Wrong.

4. AI cloud cost optimizer

A system that analyzes billing and utilization data to find idle, oversized or unused resources, then proposes Terraform changes that a human approves.

1Billing + metrics2Find waste3AI explains4Terraform plan5Human approves6Apply
Deterministic analysis finds candidates; AI explains and drafts the change; humans approve.

How to build it

  1. Pull Cost Explorer data and CloudWatch utilization (CPU, network) for 14–30 days.
  2. Use plain rules first: CPU < 10% p95, unattached EBS volumes, old snapshots, idle load balancers.
  3. Send candidates to the LLM to rank, explain trade-offs and estimate savings.
  4. Generate a Terraform diff in a pull request; run terraform plan and Infracost in CI.
  5. Track realized savings after merges on a dashboard.
GitHub repoWhat to learn from it
opencost/opencostKubernetes cost allocation
infracost/infracostCost estimates for Terraform in pull requests
cloud-custodian/cloud-custodianRules engine for finding and cleaning cloud waste
hashicorp/terraformInfrastructure as code

5. AI-powered cloud security engine

A pipeline that ingests cloud logs, applies detection rules, and uses AI to correlate, classify and explain alerts—grounded in your own runbooks via RAG.

1Cloud logs2OpenSearch3Detect4AI analysis5Classify6Alert
Detection rules stay deterministic; AI adds context, correlation and explanation.

How to build it

  1. Ship CloudTrail (and Falco runtime events) into OpenSearch.
  2. Write detection rules: root login, IAM policy changes, public S3 buckets, access from unusual regions.
  3. For each match, fetch related events in a time window and ask the LLM to correlate and explain.
  4. Index security policies and runbooks into a vector store; retrieve them so recommendations cite your rules.
  5. Measure precision on a labeled set of benign and simulated malicious events.
GitHub repoWhat to learn from it
prowler-cloud/prowlerOpen-source AWS/Azure/GCP security posture checks
falcosecurity/falcoRuntime threat detection for containers
wazuh/wazuhOpen-source SIEM/XDR platform
opensearch-project/OpenSearchLog storage and search

Which one should you build?

Don't build all five. Pick one, finish it properly, and write about it.

Your goalBuild thisCore skills shown
AI + DevOpsIncident copilotAI agents, observability, Kubernetes
CI/CDAI deployment pipelineGitHub Actions, DevSecOps, GitOps
KubernetesSelf-healing infrastructureControllers, RBAC, policy, reliability
CloudCost optimizerFinOps, Terraform, cloud APIs
SecurityCloud security engineSIEM, detection, RAG

Presenting it on your resume

  • Don't write “Built an AI chatbot.” Write: “Built an AI agent that analyzes production telemetry and assists incident remediation on Kubernetes; correct root cause in 8/10 injected failures.”
  • Include an architecture diagram, a 2-minute demo video and a README section on guardrails and limitations.
  • Be ready to discuss: what happens when the model is wrong, how you prevented unsafe actions, and cost per investigation.

Practice explaining these designs with the system design learning path.

Keep reading

Suggested next articles based on this one.

Design it, don't just read it.

Practise LLD and system design problems with structured rubrics and AI feedback.

Start practising free