Defines algorithmic efficiency strategy for the entire platform infrastructure: load balancing algorithms between availability zones, optimal bin-packing for pod placement, traffic routing in service mesh. Makes architectural decisions on scheduler logic and autoscaling optimization at thousands-of-nodes scale.
Roles · Infrastructure Engineer · Principal
What a Principal } should know
43 core skills, 60 in total. Expectations per skill, and what changes at the next level.
This page lists what a Principal } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Database Management.
Core skills for a Principal
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 4
Architecturally designs platform asynchronous processes: event-driven autoscaling by metrics, parallel infrastructure provisioning pipelines, concurrent webhook event processing from AWS/GCP. Defines graceful degradation and circuit breaker strategy for distributed infrastructure operations.
Shapes IaC quality culture at company level: Terraform module standards, Ansible collection templates, Helm chart archetypes. Defines infrastructure technical debt metrics, legacy configuration refactoring strategy and fitness functions for platform code evolution control.
Designs data models for CMDB and service catalog at organizational level: network topology graphs, microservice dependency trees, cloud resource hierarchies. Defines Terraform state file structures and metadata schemas for managing thousands of infrastructure objects in multi-cloud environments.
Backend Development · 3
Defines event-streaming infrastructure strategy at company level: choosing between Kafka, Pulsar, Redpanda, multi-region replication, capacity planning. Designs fault-tolerant cluster architecture, DR scenarios for messaging infrastructure and monitoring standards for dozens of clusters.
Defines internal Developer Platform API and portal strategy: IDP (Internal Developer Platform) architecture on Python, API gateway for infrastructure operations, service catalog integration. Makes technology stack decisions for platform services across the organization.
Defines organizational caching infrastructure strategy: choosing between managed Redis (ElastiCache, MemoryDB) and self-hosted, designing multi-region replication. Plans Redis infrastructure capacity and cost, makes caching architecture decisions for hundreds of services.
Database Management · 3
Defines database backup and recovery strategy at organizational level: RPO/RTO for different data classes, multi-region backup storage, automatic recovery testing. Designs cross-cloud DR architecture and makes decisions on storage cost vs recovery speed.
Defines organizational PostgreSQL infrastructure strategy: choosing between self-hosted Patroni, managed RDS and Aurora, designing multi-region architecture. Plans major version migrations, makes decisions on sharding (Citus) and horizontal scaling for petabyte-scale data volumes.
Shapes data replication strategy for all infrastructure: multi-region async/sync replication, logical replication for version migrations, CDC pipelines for analytics. Defines consistency and lag tolerance standards for different services, designs active-active database cluster architecture.
API & Integration · 1
Designs Internal Developer Platform API strategy: unified API layer over multi-cloud infrastructure, integration standards between platform services, API for programmatic management of all infrastructure. Defines API archetypes for IaC operations and deprecation policies for platform APIs.
Cloud & Infrastructure · 17
Shapes the organization's infrastructure automation strategy with Ansible for hybrid and multi-cloud configuration management at enterprise scale. Drives innovation in infrastructure automation including event-driven infrastructure response, self-healing patterns with ansible-rulebook, and integration with modern platform engineering approaches. Influences infrastructure automation community through Ansible collection development, upstream contributions, and thought leadership on datacenter-scale configuration management.
Shapes organizational cloud strategy on AWS: multi-account architecture for hundreds of accounts, FinOps practices for optimizing million-dollar budgets, Enterprise Support and TAM interaction. Makes decisions on AWS vs multi-cloud, designs cross-region DR strategy and defines roadmap for new AWS service adoption.
Defines CDN and edge infrastructure strategy for the organization: choosing between CloudFront, Cloudflare, Fastly, multi-CDN architecture with failover. Designs edge computing strategy (Cloudflare Workers, Lambda@Edge), defines caching standards and purge policies for global content delivery.
Shapes container supply chain security strategy at company level: SLSA compliance, Sigstore for signing and verification, NIST/NVD database integration. Defines zero-trust architecture for container workloads and compliance standards (SOC2, PCI DSS) for container infrastructure.
Shapes Crossplane strategy as infrastructure control plane: designing Composite Resource Definitions (XRD) for cloud resource abstraction, provider configuration architecture for multi-cloud. Defines Kubernetes-native IaC vs Terraform approach, composition template standards and integration with internal developer portal.
Shapes container platform strategy: choosing container runtime (Docker, containerd, CRI-O), multi-tenant registry architecture, supply chain security standards through Sigstore/Cosign. Defines Docker-to-alternative runtime migration roadmap, designs secure build pipeline for all development teams.
Defines Envoy strategy as universal data plane for the organization: sidecar and gateway deployment architecture, xDS control plane integration, custom WASM filter development. Designs observability standards through Envoy access logs and tracing, makes decisions on Envoy vs nginx vs HAProxy.
Defines Kubernetes application packaging strategy for the organization: library chart architecture, templating standards, umbrella chart management for complex systems. Designs chart repository infrastructure with OCI support, defines versioning processes and automatic Helm chart testing through ct and helm-unittest.
Shapes service mesh strategy at company level: choosing between Istio, Linkerd and Cilium mesh, multi-cluster mesh architecture with east-west gateway. Defines mTLS communication standards, traffic management through VirtualService/DestinationRule and designs observability stack on top of mesh telemetry data.
Defines advanced Kubernetes pattern architecture for the organization: custom operators with controller-runtime, CRD design for internal abstractions, multi-cluster federation through Liqo or Admiralty. Designs Kubernetes extension strategy through admission webhooks, scheduler extenders and custom CNI plugins.
Shapes company Kubernetes platform strategy: choosing between managed (EKS, GKE, AKS) and self-hosted, designing multi-cluster architecture with federation. Defines workload migration roadmap, standards for edge/IoT clusters, cost optimization strategy and FinOps for container infrastructure.
Defines organizational load balancing strategy: L4/L7 balancing architecture, Global Server Load Balancing between regions, choosing between cloud ALB/NLB and self-hosted HAProxy/Envoy. Designs health check standards, graceful draining and circuit breaking for critical services.
Shapes organizational network strategy: global network architecture across multiple regions and clouds, backbone design, IPv6 migration strategy. Makes decisions on SD-WAN, SASE architecture and zero-trust networking, defines roadmap for network automation through intent-based approaches.
Defines Pulumi usage strategy in multi-cloud infrastructure: language selection (TypeScript, Python, Go) for IaC, shared component architecture and policy-as-code through CrossGuard. Designs Pulumi CI/CD integration, state management standards and makes decisions on Pulumi vs Terraform for different use cases.
Defines organizational serverless infrastructure strategy: choosing between AWS Lambda, Cloud Functions, Knative for internal platforms. Designs event-driven infrastructure automation architecture through serverless functions, defines cold start optimization standards and cost boundaries for serverless workloads.
Shapes organizational IaC strategy: Terraform platform architecture for hundreds of modules and dozens of teams, provider versioning standards, multi-cloud abstractions. Makes decisions on Terraform version migrations, defines OpenTofu transition roadmap and designs self-service IaC for developers.
Shapes organizational VPN infrastructure strategy: site-to-site VPN for hybrid cloud, client VPN for remote access, WireGuard vs IPSec vs OpenVPN. Designs zero-trust VPN alternatives through BeyondCorp approach, defines network access architecture for multi-cloud and on-premise environments.
Security · 5
Shapes company cloud security strategy: Security Operations Center architecture for cloud, compliance framework (SOC2, ISO27001, PCI DSS), vendor risk management. Defines roadmap for CNAPP, zero-trust architecture and cloud-native SIEM, coordinates with auditors and regulators.
Shapes Kubernetes security strategy at company level: zero-trust architecture within clusters through service mesh mTLS, compliance framework (CIS, NSA hardening guide), multi-tenant isolation. Defines roadmap for confidential computing, eBPF-based security and designs security posture management for dozens of clusters.
Shapes company network security strategy: SASE/SSE architecture, zero-trust network access replacing traditional VPN, eBPF-based observability for security. Defines network automation roadmap with security-first approach, coordinates network incident response processes with SOC and designs DR scenarios.
Shapes application security strategy at infrastructure level for the entire organization: defense-in-depth architecture, WAF-as-code standards through Terraform, bug bounty program integration. Defines RASP and runtime protection implementation roadmap, coordinates with AppSec team on unified OWASP protection approach.
Shapes secrets management strategy at company level: Vault Enterprise architecture with namespace isolation, cross-region replication, SOC2/PCI compliance. Defines zero-trust secrets roadmap (SPIFFE/SPIRE), ephemeral credential standards and makes decisions on Vault vs cloud-native secret managers.
AI-Assisted Development · 1
Shapes AI-assisted Infrastructure Engineering strategy: evaluating Copilot Enterprise vs alternatives for IaC, integrating AI into platform engineering workflows, custom AI models for generating infrastructure patterns. Defines governance for AI-generated code in regulated environments and measures ROI from AI tools.
Architecture & System Design · 2
Shapes capacity management strategy at company level: long-term infrastructure investment planning, TCO models for cloud vs on-premise, FinOps practices for cost optimization. Defines predictive scaling approach, negotiates enterprise contracts with cloud providers and plans capacity for 3-5 year horizon.
Shapes Business Continuity and Disaster Recovery strategy for the company: active-active multi-region architecture, DR for multi-cloud environments, compliance with regulatory DR requirements. Defines DR infrastructure investments, designs full region loss scenarios and coordinates DR strategy with company leadership.
Observability & Monitoring · 4
Shapes unified observability strategy through OpenTelemetry: vendor-agnostic observability pipeline architecture, migration from proprietary agents to OTel, custom instrumentation standards. Defines roadmap for OTel-based eBPF observability, telemetry data cost models and observability platform governance.
Shapes metrics and monitoring strategy at company level: architecture for millions of time-series (Mimir, Victoria Metrics), multi-region observability, monitoring stack FinOps. Defines unified observability platform roadmap, custom metrics standards for business observability and cost models for monitoring infrastructure.
Shapes SLO-driven infrastructure strategy for the company: framework for defining business-critical SLIs, SLA strategy for external and internal customers, SLO integration with FinOps. Defines SLO approach for emerging technologies (AI/ML inference, edge), C-level reporting standards and SLO correlation with business metrics.
Shapes company-level observability through logs strategy: platform selection (ELK vs Loki vs Datadog), architecture for petabyte-scale log volumes, storage compliance requirements. Defines roadmap for AI-driven log analysis, audit logging standards for security and cost optimization for logging infrastructure.
Version Control & Collaboration · 3
Shapes infrastructure code review culture at organizational level: review standards for different change classes (Tier 1 — network core, Tier 4 — dev environment), change management integration. Defines automated governance architecture through policy-as-code, review process metrics and peer review approach for Principal level.
Shapes knowledge management strategy for company infrastructure: Internal Developer Portal architecture with auto-generated documentation, Backstage integration with IaC repositories. Defines technical writing standards for infrastructure, AI-assisted documentation roadmap and documentation coverage metrics for compliance.
Shapes version control strategy for all company infrastructure: Git platform architecture (GitHub Enterprise, GitLab), standards for thousands of IaC repositories, inner source model for Terraform modules. Defines GitOps delivery governance, compliance audit trail through git history and change management process integration.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
} in the open competency matrix: 60 skills across 5 levels. The matrix is free for individuals and stays free.