Reviews algorithmic solutions in infrastructure tools: analyzing provisioning script complexity, optimizing service dependency graph traversal. Evaluates deployment planning and container placement algorithms, implements automation performance standards for fleets of hundreds of servers.
Roles · Infrastructure Engineer · Lead
What a Lead } should know
45 core skills, 60 in total. Expectations per skill, and what changes at the next level.
This page lists what a Lead } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Database Management.
Core skills for a Lead
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 4
Implements asynchronous execution patterns in infrastructure automation: parallel deployment to server groups, concurrent health checks, non-blocking metrics collection. Standardizes timeout handling and retry strategies for cloud API operations that may take minutes.
Implements infrastructure code quality practices: tflint and checkov for Terraform, ansible-lint for playbooks, shellcheck for bash scripts. Configures CI pipelines with automatic IaC validation, defines code review standards for infrastructure modules and merge policies.
Standardizes infrastructure data modeling approaches: host inventory files, service YAML configs, DNS record trees. Reviews data structure choices in IaC modules — Terraform resource mappings, nested Ansible variables, Helm chart configuration hierarchies.
Backend Development · 3
Designs Kafka infrastructure as a platform service for development teams: cluster deployment standards, retention policies, authorization schemas through ACL. Implements GitOps approach to Kafka topic management and defines SLO for broker throughput and latency.
Designs internal platform service architecture on Python: self-service API for infrastructure management, webhook servers for GitOps pipelines, admin panels. Defines infrastructure backend development standards, reviews integrations with Terraform Cloud API and Vault.
Designs Redis as a managed platform service: cluster deployment standards, backup policies, sharding strategies. Defines SLO for Redis infrastructure, implements self-service instance provisioning and ensures monitoring through Prometheus exporters.
Database Management · 3
Defines data management strategy at product level. Establishes Backup and Disaster Recovery standards. Conducts data schema and scaling strategy reviews.
Designs PostgreSQL as a platform service for developers: self-service instance creation through Terraform, configuration standards for different workload profiles, update automation. Defines SLO for database availability and performance, reviews team architectural decisions.
Defines data management strategy at product level. Establishes Replication and High Availability standards. Conducts data schema and scaling strategy reviews.
API & Integration · 1
Defines API standards for internal platform services: unified error format, naming conventions, versioning, authentication through service accounts. Reviews API interfaces of team infrastructure tools and ensures compatibility with corporate API gateway.
Cloud & Infrastructure · 17
Defines Ansible automation standards for infrastructure engineering teams across the organization. Establishes infrastructure-as-code practices with Ansible including role development guidelines, testing frameworks with Molecule, and Tower/AWX governance for change management and self-service provisioning. Conducts architecture reviews of infrastructure automation and drives standardization of server provisioning, configuration management, and lifecycle automation patterns.
Defines AWS infrastructure standards for the organization: landing zone architecture, account vending through Control Tower, tagging standards for FinOps. Reviews team AWS architectures against Well-Architected Framework, implements self-service through Service Catalog and defines SLO for AWS services.
Architects global CDN topology with edge computing for latency-critical workloads. Establishes IaC standards for CDN provisioning and cache governance across teams. Drives FinOps practices to optimize CDN spend and bandwidth costs.
Defines container security standards for the organization: image admission policies for production, vulnerability remediation SLA by severity, zero-day CVE handling processes. Implements shift-left security through development-stage scanning and reviews container workload security architecture.
Defines infrastructure strategy with Crossplane. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.
Defines containerization standards for the organization: base image policies with automatic updates, scanning pipelines, tagging strategies. Reviews team Dockerfiles for best practices compliance, manages registry infrastructure with geo-replication and retention policies.
Defines infrastructure strategy with Envoy Proxy. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.
Architects Helm-based deployment strategy across multi-cluster Kubernetes environments. Establishes IaC standards for chart development, testing, and promotion workflows. Optimizes Helm release management for large-scale infrastructure with FinOps practices.
Defines infrastructure strategy with Istio Service Mesh. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.
Defines Kubernetes platform strategy including cluster provisioning automation, security hardening standards, and disaster recovery procedures. Conducts architecture reviews for workload placement and resource optimization.
Defines Kubernetes platform standards for the organization: multi-tenant cluster architecture, namespace policies, deployment standards through GitOps. Reviews team Kubernetes manifests, designs self-service abstractions on top of K8s and defines SLO for control plane and workload availability.
Defines load balancing infrastructure strategy: balancer technology selection (hardware vs software), HA configuration standards, and capacity planning for traffic growth. Establishes monitoring, alerting, and incident response procedures for load balancer failures.
Defines organizational infrastructure network standards: IP addressing and CIDR planning for hundreds of VPCs, firewall rule standards, DNS naming policies. Reviews team network architectures, designs network traffic observability and defines SLO for network availability.
Defines infrastructure-as-code strategy with Pulumi across the organization: establishes coding standards, testing frameworks, and component library governance. Conducts architecture reviews for complex multi-cloud deployments. Drives FinOps optimization through policy enforcement and cost allocation tagging.
Defines platform-wide serverless strategy including standards for function design, deployment patterns, and operational runbooks. Establishes FinOps practices for serverless cost optimization. Reviews cross-team serverless architectures for reliability and security.
Defines Terraform standards for the organization: repository structure, resource naming conventions, state separation strategy. Implements Terraform Cloud/Enterprise for teamwork, reviews modules for security and efficiency, designs drift detection and automatic remediation process.
Defines network security strategy for VPN and isolation across the organization: establishes encryption standards, network segmentation policies, and connectivity governance for hybrid environments. Conducts architecture reviews for complex multi-site deployments. Drives adoption of zero-trust network architectures.
DevOps & CI/CD · 1
Standardizes infrastructure-as-code through GitOps: defines repository structure and module strategy, creates infrastructure scaffolding tools, implements policy compliance in infrastructure PR review. Designs infrastructure change management process.
Testing & QA · 1
Defines infrastructure resilience strategy: designs multi-region failover architecture validated through chaos, creates infrastructure chaos suite for continuous verification. Standardizes DR procedures and ensures RTO/RPO compliance through regular testing.
Security · 5
Defines cloud security standards for the organization: baseline security controls for each account type, IAM role standards, data encryption policies. Reviews team security architectures, implements security guardrails through SCPs and Terraform modules, defines SLO for vulnerability time-to-remediate.
Defines Kubernetes security standards for the organization: Kyverno/OPA policies for all clusters, image admission standards, security review process for Helm charts. Implements security-as-code approach, reviews team RBAC matrices and designs incident response process for Kubernetes incidents.
Defines organizational network security standards: segmentation policies for all environments, TLS configuration standards, firewall change management processes. Reviews team network architectures for zero-trust compliance, implements automated network policy auditing and defines SLO for security patching.
Defines application security standards at infrastructure level: WAF policies for all public endpoints, security header standards, vulnerability response process. Reviews team security configurations and implements continuous security testing in infrastructure pipelines.
Defines secrets management standards for the organization: Vault namespace architecture for multi-tenant, secret rotation and TTL policies, team onboarding process. Reviews Vault policies and auth configurations, designs self-service portal for managing secrets and certificates.
AI-Assisted Development · 1
Defines AI assistant usage standards for the infrastructure team: guidelines for Copilot with IaC code, review rules for AI-generated Terraform modules on security compliance. Implements custom instructions for infrastructure context and trains team on effective prompt patterns for DevOps tasks.
Architecture & System Design · 2
Defines capacity planning process for the organization: regular resource utilization reviews, forecast models based on business metrics, quarterly infrastructure budgeting. Implements automatic capacity monitoring through Prometheus alerts and coordinates planning with development teams for new launches.
Defines DR standards for organizational infrastructure: service classification by criticality (Tier 1-4), standard DR patterns for each tier, regular DR drills. Reviews team DR plans, implements automated failover testing and coordinates quarterly disaster recovery exercises.
Observability & Monitoring · 4
Defines distributed tracing standards for the organization: mandatory instrumentation for all services, span naming and attribute convention standards, sampling policies. Implements OTel Operator for Kubernetes, reviews team configurations and designs self-service onboarding for the observability stack.
Defines monitoring standards for the organization: mandatory metrics (RED/USE methods), standard dashboards for each service type, alert creation process with runbooks. Implements Grafana-as-a-service for self-service monitoring, reviews team alerting rules and defines SLO for monitoring infrastructure.
Defines SLO standards for all infrastructure: standard SLIs for each component class (compute, storage, network, DB), SLO negotiation process with teams. Implements SLO-driven prioritization for engineering work, reviews team SLOs and coordinates error budget policy with product management.
Defines logging standards for the organization: unified structured log format, mandatory fields (trace_id, service, environment), retention policies by data class. Reviews team logging configurations, designs self-service log access and defines SLO for log ingestion latency and availability.
Version Control & Collaboration · 3
Defines code review standards for the infrastructure team: checklists for different change types (network, compute, security), review time SLA, escalation process for critical changes. Implements automated review through Danger.js and custom bots, coordinates cross-team review for shared infrastructure.
Defines infrastructure documentation standards for the organization: mandatory sections for each Terraform module, runbook and playbook templates, ADR standards. Implements docs-as-code pipeline with review process, reviews team documentation for completeness and coordinates on-call handbook creation.
Defines Git standards for the infrastructure organization: unified branching strategy for IaC repositories, commit message standards for automated changelog, protected branch policies. Implements GitOps approach (ArgoCD, Flux) for Kubernetes and Terraform, reviews team git workflows and defines disaster recovery for git repositories.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
What changes at Principal
0 skills get a higher expectation or become core when moving from Lead to Principal. The biggest jumps first.
} in the open competency matrix: 60 skills across 5 levels. The matrix is free for individuals and stays free.