Applies algorithmic thinking to infrastructure management: load balancing algorithms for traffic distribution, capacity planning algorithms for resource forecasting, network routing optimization algorithms. Designs efficient provisioning algorithms for infrastructure automation with dependency resolution.
Roles · Infrastructure Engineer · Senior
What a Senior } should know
45 core skills, 60 in total. Expectations per skill, and what changes at the next level.
This page lists what a Senior } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Database Management.
Core skills for a Senior
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 4
Designs async architectures for infrastructure automation: concurrent multi-host management, async configuration deployment at scale, non-blocking infrastructure monitoring. Mentors team on async patterns for reliable infrastructure operations.
Designs code quality standards for infrastructure automation: Ansible playbook structure, Pulumi/Terraform module conventions, infrastructure testing frameworks. Refactors complex provisioning scripts into declarative, idempotent modules. Establishes review practices for blast radius minimization and rollback safety.
Selects optimal data structures for infrastructure management: graph structures for network topology modeling, tree structures for DNS and LDAP hierarchies, time-series data models for capacity planning. Optimizes configuration data structures for rapid provisioning and rollback. Designs efficient data models for infrastructure asset tracking and dependency mapping.
Backend Development · 3
Administers Kafka clusters as part of infrastructure: configuring replication, monitoring consumer lag, optimizing partition strategy. Automates topic management through Terraform provider, ensures broker fault tolerance and configures alerts on performance degradation.
Develops internal APIs and web interfaces for infrastructure services on Flask/FastAPI: self-service provisioning portals, API wrappers over Terraform and Ansible, monitoring dashboards. Integrates Python backends with cloud SDKs and configuration management systems.
Manages Redis infrastructure: configuring Sentinel and Redis Cluster, memory optimization through eviction policies, operation latency monitoring. Automates Redis deployment through Ansible/Terraform, configures data persistence and memory exhaustion alerts.
Database Management · 3
Designs multi-region disaster recovery architectures with automated failover and failback procedures. Implements infrastructure-as-code for backup topology ensuring consistency across environments. Optimizes RTO/RPO targets through tiered backup strategies and conducts chaos engineering for DR validation.
Designs fault-tolerant PostgreSQL infrastructure: Patroni configuration for automatic failover, PgBouncer for connection pooling, WAL archiving to S3. Optimizes OS kernel and filesystem parameters for PostgreSQL, plans capacity for IOPS and storage.
Designs infrastructure for multi-region replicated database clusters: network topology for low-latency replication, storage architecture for consistent performance, and automated DR failover. Optimizes infrastructure costs while meeting RPO/RTO SLAs. Mentors team on HA infrastructure patterns.
API & Integration · 1
Designs REST API for internal infrastructure services: resource provisioning API, self-service portal endpoints, GitOps webhooks. Integrates multiple APIs (cloud providers, Vault, Terraform) into unified automation workflow with proper error handling and idempotency.
Cloud & Infrastructure · 17
Designs Ansible automation architecture for enterprise infrastructure management across data centers and cloud environments. Implements advanced patterns including custom connection plugins for network devices, callback plugins for ITSM integration, and Tower/AWX workflows for multi-stage infrastructure provisioning. Optimizes Ansible for large-scale infrastructure operations through performance tuning, custom inventory plugins, and integration with configuration management databases for compliance tracking.
Designs production-grade AWS architecture: multi-AZ and multi-region deployment, Transit Gateway for network connectivity, Organization with SCPs for multi-account strategy. Optimizes costs through Reserved Instances and Savings Plans, configures AWS Config for compliance and GuardDuty for security monitoring.
Designs infrastructure solutions with CDN and Edge Computing. Optimizes cost and performance. Implements best practices and security hardening.
Designs comprehensive container security system: runtime scanning through Falco, Kubernetes admission controller for image verification, SBOM generation through Syft. Configures automatic base image patching and integrates scanning results with SIEM system.
Designs multi-cloud Crossplane Compositions with cost-optimized resource patches. Implements ProviderConfig rotation for credential security. Builds Composition pipelines with validation webhooks and drift detection.
Designs corporate Docker infrastructure: hardened base image standards, multi-arch builds, Docker BuildKit configuration with S3 caching. Optimizes containerd runtime configuration, configures storage drivers for high loads and integration with corporate PKI for image signing.
Designs Envoy fleet configurations with advanced load balancing (ring hash, Maglev) and circuit breaking. Implements rate limiting with external gRPC services and custom Lua/Wasm filters. Hardens mTLS settings with SDS integration and certificate rotation.
Designs infrastructure solutions with Helm. Optimizes cost and performance. Implements best practices and security hardening.
Designs multi-cluster Istio deployments with advanced traffic management. Implements mTLS policies and authorization rules for zero-trust networking. Optimizes Envoy proxy resource consumption and latency overhead.
Designs infrastructure solutions with Kubernetes Advanced. Optimizes cost and performance. Implements best practices and security hardening.
Designs production-grade Kubernetes infrastructure: cluster deployment through kubeadm/kOps/EKS, CNI configuration (Calico, Cilium), etcd setup for fault tolerance. Optimizes scheduler, configures PodDisruptionBudget and topologySpreadConstraints, designs cluster upgrade strategy.
Designs load balancing architecture for high-traffic systems: ECMP routing, DPDK-accelerated proxying, and global server load balancing (GSLB). Optimizes for minimal latency with connection pooling, keep-alive tuning, and hardware offloading.
Designs complex network architectures: Transit Gateway for hub-and-spoke topology, VPC peering between regions, Direct Connect/Interconnect for hybrid scenarios. Optimizes traffic routing, configures advanced DNS (GeoDNS, failover routing) and designs network segmentation for compliance.
Designs enterprise-grade infrastructure solutions with Pulumi: multi-account architectures, custom resource providers, and Automation API for self-service platforms. Optimizes cloud costs through right-sizing automation and resource lifecycle policies. Implements CrossGuard policies and security hardening.
Designs production-grade serverless infrastructure with multi-region failover, DLQ handling, and observability. Optimizes function performance and cost through right-sizing and provisioned concurrency. Implements security hardening and least-privilege IAM.
Designs modular Terraform architecture: reusable modules for networks, clusters, databases, workspace strategy for environments. Configures Terragrunt for DRY configurations, implements policy-as-code through OPA/Sentinel, optimizes plan/apply time for large state files.
Designs enterprise VPN and network isolation architectures: multi-region mesh topologies, zero-trust network access with ZTNA gateways, and hybrid cloud connectivity with dedicated interconnects. Optimizes throughput and latency for high-bandwidth tunnels. Implements security hardening with certificate pinning and MFA integration.
DevOps & CI/CD · 1
Designs infrastructure GitOps: configures Crossplane for Kubernetes-native infrastructure management, implements drift detection and reconciliation for infrastructure state. Designs multi-account/multi-region infrastructure through GitOps with proper state management.
Testing & QA · 1
Designs infrastructure resilience testing: creates automated DR drills, tests backup/restore procedures under load, implements region failover experiments. Configures infrastructure monitoring for chaos impact detection and automatic rollback.
Security · 5
Designs cloud infrastructure security architecture: multi-account strategy with security hub, centralized logging through CloudTrail + S3 + Athena, GuardDuty for threat detection. Implements CSPM (Cloud Security Posture Management), configures automatic remediation through Lambda and designs cross-account access patterns.
Designs comprehensive Kubernetes security: admission controllers (OPA Gatekeeper, Kyverno) for policy enforcement, runtime security through Falco, network segmentation through Cilium NetworkPolicy. Configures audit logging, encrypts secrets at rest through KMS provider and designs workload identity for cloud services.
Designs enterprise-grade network security: micro-segmentation through Cilium/Calico NetworkPolicy, mTLS for service-to-service communication, DDoS protection through AWS Shield/CloudFlare. Implements network detection and response (NDR), configures deep packet inspection and designs secure connectivity for hybrid cloud.
Designs infrastructure protection against OWASP vulnerabilities: multi-layer WAF with custom rules, rate limiting at ALB and nginx level, SSRF protection through network segmentation. Implements automatic DAST scanning for infrastructure services and configures ModSecurity with OWASP Core Rule Set.
Designs production-grade Vault infrastructure: HA cluster with Raft storage, auto-unseal through AWS KMS, disaster recovery through replication. Configures dynamic secrets for all databases and cloud providers, implements Vault Agent for transparent secret injection into Kubernetes pods.
AI-Assisted Development · 1
Applies AI tools for complex infrastructure tasks: generating boilerplate for Terraform providers, refactoring Ansible roles, creating unit tests for IaC through Terratest. Knows Copilot limitations for HCL/YAML and adjusts prompts considering cloud API specifics and network configurations.
Architecture & System Design · 2
Conducts capacity planning for infrastructure components: load forecasting based on historical metrics, Kubernetes cluster growth planning, IOPS calculation for storage subsystem. Builds resource consumption models (CPU, RAM, network bandwidth) and defines scaling points for critical services.
Designs disaster recovery for critical infrastructure: multi-AZ architecture with automatic failover, backup strategy with cross-region replication, recovery runbooks. Configures automatic DR testing through chaos engineering (Chaos Monkey, Litmus), defines RPO/RTO for each component.
Observability & Monitoring · 4
Designs production-grade OTel architecture: multi-layer collector topology (agent → gateway → backend), tail-based sampling for intelligent filtering, batching and retries. Configures correlation between traces/metrics/logs, optimizes collector resource consumption and integrates OTel with service mesh telemetry.
Designs scalable Prometheus infrastructure: federation for multi-cluster monitoring, Thanos/Cortex for long-term storage and global query, highly available Alertmanager. Optimizes metric cardinality, configures recording rules for query performance and designs Grafana dashboards-as-code through Jsonnet.
Designs SLO framework for infrastructure platform: cascading SLOs from infrastructure to services, composite SLIs for complex systems, automated error budget calculation. Implements SLO-as-code through Sloth or OpenSLO, configures automated incident creation on breach and integrates SLO with capacity planning.
Designs production-grade logging infrastructure: multi-tenant ELK clusters, Loki with S3 backend for cost-effective storage, Kafka as buffer for peak loads. Optimizes storage costs through tiered storage, configures correlation ID for log-based tracing and designs alerts for log pattern anomalies.
Version Control & Collaboration · 3
Conducts architectural reviews of infrastructure code: evaluating Terraform module design, analyzing networking topology changes, verifying DR capability of configurations. Automates review through Atlantis/Spacelift for terraform plan in PRs, implements automated policy checks and mentors junior engineers through constructive feedback.
Designs infrastructure documentation system: auto-generation from IaC code (terraform-docs, ansible-doc), live architecture diagrams through Mermaid/D2, service catalog integration. Implements documentation testing (broken links, schema validation), creates postmortem templates and designs knowledge base for on-call engineers.
Designs Git workflow for infrastructure repositories: monorepo vs polyrepo strategy for Terraform, GitOps branching model for Kubernetes manifests, automated release pipeline. Configures CODEOWNERS for critical configurations, git-crypt for secret encryption and optimizes work with large state files.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
What changes at Lead
60 skills get a higher expectation or become core when moving from Senior to Lead. The biggest jumps first.
- Algorithms & Complexity: Advanced → Expert
- Ansible: Advanced → Expert
- Apache Kafka: Advanced → Expert
- Async Programming: Advanced → Expert
- AWS: Advanced → Expert
- Backup & Disaster Recovery: Advanced → Expert
- Capacity Planning: Advanced → Expert
- CDN & Edge Computing: Advanced → Expert
- Chaos Engineering: Advanced → Expert
- Cloud Security: Advanced → Expert
} in the open competency matrix: 60 skills across 5 levels. The matrix is free for individuals and stays free.