Roles · Cloud Engineer · Lead

What a Lead } should know

43 core skills, 58 in total. Expectations per skill, and what changes at the next level.

This page lists what a Lead } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Cloud & Infrastructure.

43core skills
15additional skills
10skill areas
100%at Advanced or Expert
Assess myself as Lead Full role matrix

Core skills for a Lead

Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.

Programming Fundamentals · 3

Evaluates algorithmic complexity of cloud solutions: cost of resource traversal through API, pagination when working with thousands of instances, Lambda cold start optimization. Defines performance budgets for IaC pipelines and automates degradation detection during infrastructure scaling.

Establishes Infrastructure as Code quality standards: Terraform linting (tflint, checkov), Helm chart validation, policy-as-code through OPA/Sentinel. Configures quality gates in CI/CD for infrastructure changes and balances deployment speed with security.

Defines standards for cloud infrastructure configuration storage and structuring: hierarchical Terraform state files, resource dependency DAGs, graph models for network topologies. Reviews data structure choices for CloudFormation stacks and Helm charts.

Backend Development · 1

Defines architectural decisions for S3 / Object Storage at the product level. Establishes standards. Conducts design review and defines technical roadmap.

Cloud & Infrastructure · 18

Ansible Expert

Defines Ansible automation standards and best practices for cloud infrastructure teams across the organization. Establishes role and collection development guidelines, testing frameworks with Molecule, and Ansible Tower/AWX governance for self-service cloud provisioning. Conducts architecture reviews of automation codebases and drives adoption of standardized cloud provisioning patterns through reusable Ansible collections.

AWS Expert

Defines organizational AWS strategy: account vending machine, service control policies, centralized networking (Transit Gateway Hub). Manages enterprise support, cost allocation, budgets and alerts. Introduces FinOps practices and optimizes costs at the company level.

Defines CDN and edge computing strategy across multi-cloud environments. Establishes caching standards, edge security policies, and FinOps optimization for content delivery. Conducts architecture reviews for global content distribution.

Defines infrastructure strategy with Container Security Scanning. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

Docker Expert

Defines containerization standards for the organization: golden base images, security hardening guidelines, image scanning policies. Introduces automation for base image updates and patch management for container runtime.

Defines infrastructure strategy with Google Cloud Platform. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

Helm Expert

Defines Helm chart standards and library charts for organization-wide Kubernetes deployments. Establishes chart review processes, security scanning, and versioning governance. Drives adoption of GitOps workflows with Helm and ArgoCD or Flux.

Defines multi-cluster Kubernetes strategy across cloud providers. Establishes GitOps workflows with ArgoCD/Flux, implements service mesh policies, and optimizes cluster costs with FinOps practices.

Defines Kubernetes standards for the organization: cluster provisioning through IaC, security baseline (CIS benchmarks), cluster health monitoring. Introduces platform engineering approach — internal developer platform based on Kubernetes.

Defines load balancing strategy for cloud platform: multi-region traffic management, failover automation, and cost-optimized balancer selection. Establishes IaC standards for load balancer provisioning and conducts architecture reviews.

Defines infrastructure strategy with Microsoft Azure. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

Defines organizational networking strategy: centralized networking account, IP address management (IPAM), segmentation standards. Introduces service mesh for microservices, zero-trust networking. Manages network changes through IaC with automated compliance.

Pulumi Expert

Defines cloud infrastructure strategy with Pulumi: establishes IaC standards, module registry governance, and multi-environment promotion workflows. Conducts architecture reviews for cross-team Pulumi adoption. Optimizes FinOps through automated cost policies and resource lifecycle management.

Defines infrastructure strategy with Serverless Containers. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

Defines serverless-first infrastructure strategy across the organization. Establishes IaC standards for function deployment, monitoring, and cost governance. Conducts architecture reviews of serverless solutions and optimizes FinOps for compute spend.

Terraform Expert

Defines organizational IaC standards: repository structure, naming conventions, tagging strategy, module registry. Introduces policy-as-code through Sentinel/OPA, drift detection, cost estimation (Infracost). Trains teams on Terraform best practices.

Defines VPN and network isolation strategy for cloud infrastructure: establishes connectivity standards, transit architecture patterns, and zero-trust network policies. Conducts architecture reviews for multi-cloud network designs. Optimizes network costs through traffic engineering and interconnect planning.

Yandex Cloud Expert

Defines infrastructure strategy with Yandex Cloud. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

DevOps & CI/CD · 4

ArgoCD Expert

Defines ArgoCD strategy and GitOps standards for the organization's cloud infrastructure fleet. Establishes multi-tenant ArgoCD governance with project structures, RBAC policies, and approval workflows. Conducts architecture reviews of ArgoCD configurations and drives adoption of declarative infrastructure delivery patterns across cloud engineering teams.

Defines DevOps strategy with Blue/Green Deployment. Establishes CI/CD standards. Introduces platform engineering approaches.

Defines DevOps strategy with Canary Deployment. Establishes CI/CD standards. Introduces platform engineering approaches.

Defines CI/CD strategy for all infrastructure repositories: standard workflow templates, self-hosted runners in private networks, secrets management through OIDC. Introduces governance — required reviews, environment protection rules, audit trail for infrastructure changes.

Security · 2

Defines cloud platform security strategy: security baseline for new accounts, incident response runbooks, vulnerability management program. Introduces CSPM (Cloud Security Posture Management), trains teams on secure-by-default approach. Manages security exceptions and risk acceptance.

Defines secrets management strategy for the organization: Vault vs AWS Secrets Manager vs GCP Secret Manager, namespace hierarchy for multi-tenancy, emergency break-glass procedures. Introduces compliance controls and automated audit of secrets access.

AI-Assisted Development · 1

Defines AI tool usage policy for infrastructure development: approved tools, security review of AI-generated IaC, CI/CD integration for automated review. Evaluates ROI and introduces best practices combining AI generation with human review.

Architecture & System Design · 4

Defines product architectural strategy with Capacity Planning. Establishes architecture guidelines. Conducts architecture review.

Defines product architectural strategy with Disaster Recovery Design. Establishes architecture guidelines. Conducts architecture review.

Defines architectural standards for high-load cloud-native systems: reference architectures, performance budgets, scalability review checklist. Conducts architecture review, identifies bottlenecks through load testing and designs capacity planning processes.

Defines organizational architectural standards: technology radar, reference architectures for common scenarios (REST API, event processing, data pipeline). Establishes architecture review process and trains teams on cloud-native design patterns.

Observability & Monitoring · 6

ELK Stack Expert

Defines organizational logging strategy: centralized logging account, compliance requirements (audit trails, retention), cost optimization. Introduces logging standards for all cloud workloads, configures alerting for anomalies and security events.

Defines on-call strategy for the cloud organization: follow-the-sun rotation, tier-1/tier-2 escalation, incident commander role. Introduces incident management process (ITIL/SRE), blameless postmortems, reliability metrics (MTTA, MTTR). Manages on-call load balancing and burnout prevention.

OpenTelemetry Expert

Defines observability strategy based on OpenTelemetry: unified collection pipeline, vendor-agnostic instrumentation, standards for all teams. Manages telemetry data costs through sampling strategies and tiered storage. Introduces SLO-based alerting based on traces.

Defines metrics strategy for cloud platform: Prometheus vs CloudWatch vs Datadog, golden signals framework, standard dashboards for each workload type. Introduces SLO-based monitoring, budget alerts and automated capacity recommendations based on metrics.

Defines organizational SLO culture: SLO review process, error budget governance, SLA negotiations with clients. Introduces tooling (Sloth, Google SLO Generator) and standards for all cloud services. Balances reliability requirements and delivery speed based on error budgets.

Defines organizational logging standards: mandatory fields, log levels policy, sensitive data handling. Introduces automated log validation in CI/CD, standards compliance monitoring. Trains teams on best practices and reviews logging configurations.

Version Control & Collaboration · 3

Code Review Expert

Defines review process for infrastructure changes: review checklist, mandatory approvals by severity (production = 2+ approvals), automated policy checks. Builds a culture of quality review — knowledge sharing, not gate-keeping. Balances speed and safety.

Defines infrastructure documentation standards: ADR template, runbook structure, service catalog format. Introduces docs-as-code workflow — PR-based updates, automated freshness checks, documentation coverage metrics. Trains teams on documentation culture.

Git Advanced Expert

Defines IaC repository management strategy: repo topology, access control, branch protection rules, mandatory reviews for production changes. Introduces GitOps workflow — infrastructure changes only through PR with automated plan and manual approve for apply.

Documentation · 1

Defines runbook strategy for the organization: coverage requirements (each service — minimum 5 runbooks), review process, freshness policy. Introduces self-healing runbooks — automated execution on specific alerts. Links runbooks with SLOs and incident severity levels.

Additional skills

Not assessed by the team, but part of the self-assessment and the development plan.

API DocumentationAsync ProgrammingChatGPT / ClaudeCursor IDEDesign PatternsGraphQL DesignIntegration TestingMultithreadingOOP & SOLID PrinciplesPostgreSQLPrompt Engineering for CodeRedisREST API DesignSecure Coding PracticesUnit Testing

What changes at Principal

0 skills get a higher expectation or become core when moving from Lead to Principal. The biggest jumps first.

See the Principal page →
Run this with your whole team
Self-assessment plus manager and peer reviews against the same matrix, gap analysis and next-level readiness for every engineer. Team Pro is free for 14 days; individual tools stay free forever.
Start a team trial (14 days free) Send to my manager

} in the open competency matrix: 58 skills across 5 levels. The matrix is free for individuals and stays free.