Designs data processing strategies at cloud platform scale: optimal autoscaling algorithms, load balancing considering latency and cost, efficient DNS sharding and cross-region traffic routing strategies.
Roles · Cloud Engineer · Principal
What a Principal } should know
43 core skills, 58 in total. Expectations per skill, and what changes at the next level.
This page lists what a Principal } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Cloud & Infrastructure.
Core skills for a Principal
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 3
Defines organizational IaC quality strategy: unified Terraform module standards, golden templates for cloud services, automated compliance. Manages infrastructure technical debt and plans legacy configuration migration.
Designs data models for cloud platform management: CMDB, service catalog, cost allocation trees. Makes strategic decisions on state storage — S3 backend vs Terraform Cloud, structuring multi-account AWS Organizations hierarchies.
Backend Development · 1
Defines S3 / Object Storage strategy at company level. Evaluates new technologies and approaches. Establishes enterprise standards and reference architectures.
Cloud & Infrastructure · 18
Shapes the organization's infrastructure automation strategy with Ansible as a core tool for multi-cloud configuration management. Drives innovation in Ansible-based cloud automation including event-driven automation with ansible-rulebook, GitOps integration patterns, and self-healing infrastructure playbooks. Influences Ansible community practices through contributions to cloud provider collections and conference presentations on enterprise-scale cloud automation.
Shapes enterprise-level cloud strategy: multi-region disaster recovery, hybrid cloud (Outposts, Direct Connect), compliance frameworks (SOC2, HIPAA, PCI DSS). Designs reference architectures and makes strategic decisions on cloud-native vs lift-and-shift.
Defines CDN and edge computing strategy: CloudFront/Cloud CDN for static and API acceleration, Lambda@Edge/CloudFlare Workers for edge logic. Designs caching architecture — cache invalidation, origin shield, multi-CDN failover for global applications.
Defines platform-level container security strategy: Trivy/Snyk integration in CI/CD, admission controllers in Kubernetes for blocking vulnerable images, runtime protection through Falco/Aqua. Establishes CVE remediation SLA and automates compliance reporting.
Shapes container platform strategy: choosing between ECS/EKS/Fargate/Cloud Run, standards for multi-tenant environments, governance for container registries. Designs supply chain security architecture for container artifacts.
Defines GCP strategy in multi-cloud architecture: BigQuery for analytics, GKE Autopilot for Kubernetes, Anthos for hybrid cloud. Designs GCP integration with AWS/Azure, evaluates GCP advantages for ML/AI workloads and data-intensive applications.
Designs organizational Helm ecosystem architecture: umbrella charts for complex applications, library charts for reuse, Helmfile for release management. Defines versioning standards, template functions and values hierarchies for multi-cluster environments.
Shapes Kubernetes platform evolution strategy: multi-cluster federation, service mesh (Istio/Linkerd), GitOps through ArgoCD/Flux. Designs custom operators, admission webhooks, scheduler extensions. Defines multi-tenancy architecture and resource governance.
Shapes company-level container orchestration strategy: multi-cluster topology, cross-region failover, hybrid cloud deployment. Evaluates alternatives (Nomad, ECS) and makes strategic decisions on Kubernetes platform evolution.
Defines platform-level load balancing strategy: L4 (NLB) vs L7 (ALB) vs Global (GCLB), multi-region load balancing with GeoDNS, service mesh for internal traffic. Designs zero-downtime deployment patterns and traffic management architecture for canary/blue-green releases.
Shapes Azure strategy for enterprise: Azure Landing Zones, Azure Arc for hybrid, Azure DevOps for CI/CD. Designs integration with Active Directory/Entra ID, defines Azure vs AWS use cases — compliance, Microsoft ecosystem, government cloud.
Shapes platform-level network architecture: global backbone, multi-region connectivity, cloud interconnect strategy. Designs solutions for edge computing, CDN integration, hybrid DNS. Defines network security standards and Zero Trust Architecture.
Evaluates Pulumi as strategic alternative to Terraform: advantages of real programming languages for complex logic, Automation API for self-service platforms, Policy as Code through CrossGuard. Defines use cases for Pulumi vs Terraform in multi-cloud architecture.
Defines serverless containers strategy for the organization: AWS Fargate vs Google Cloud Run vs Azure Container Apps. Designs architecture for burst workloads, event-driven scaling, cold start optimization. Establishes cost model and guidelines for choosing between serverless and dedicated compute.
Shapes serverless architecture strategy: Lambda/Cloud Functions for event processing, API Gateway integration, Step Functions for orchestration. Designs patterns — fan-out/fan-in, saga, circuit breaker. Defines serverless applicability boundaries and migration paths.
Shapes company-level Infrastructure as Code strategy: multi-cloud provisioning, Terraform Enterprise/Cloud vs self-hosted, modular platform for self-service. Designs state management architecture for hundreds of stacks and dozens of teams.
Shapes organizational VPN strategy: Site-to-Site VPN vs Direct Connect/ExpressRoute, client VPN for remote access, mesh VPN between clouds. Designs high-availability VPN with BGP, failover and monitoring. Defines migration path to Zero Trust Network Access (ZTNA).
Defines Yandex Cloud strategy for the Russian market: Managed Kubernetes, Object Storage, Cloud Functions. Evaluates compliance with FZ-152, integration with local infrastructure, migration paths between Yandex Cloud and international providers.
DevOps & CI/CD · 4
Shapes the organization's GitOps strategy around ArgoCD for enterprise-scale cloud platform delivery. Drives architectural decisions on multi-cluster ArgoCD federation, disaster recovery patterns, and integration with cloud-native control planes. Influences industry practices through contributions to ArgoCD ecosystem and GitOps community standards.
Shapes blue-green deployment strategy for cloud infrastructure: Route 53 weighted routing, ALB target group switching, ECS blue-green through CodeDeploy. Designs rollback automation, database migration strategy and smoke testing for zero-downtime infrastructure updates.
Defines platform-level canary deployment strategy: progressive delivery through Flagger/Argo Rollouts, weighted traffic shifting through service mesh, automated rollback based on SLO metrics. Designs observability pipeline for canary analysis and integration with cloud-native load balancers.
Shapes infrastructure change automation platform: GitOps with GitHub Actions as orchestrator, integration with Terraform Cloud/Spacelift, multi-cloud deployment pipelines. Designs self-service developer platform with guardrails for safe resource provisioning.
Security · 2
Shapes enterprise-level cloud security strategy: Zero Trust Architecture, cloud-native SIEM (CloudTrail Lake, Chronicle), supply chain security. Defines compliance frameworks (SOC2, ISO 27001, PCI DSS), designs security reference architecture for multi-cloud.
Shapes enterprise-level secrets management platform: multi-region Vault with disaster recovery, cross-cloud secrets synchronization, zero-trust secret distribution. Defines cryptographic standards, key management strategy and HSM integration for critical workloads.
AI-Assisted Development · 1
Shapes AI-assisted Infrastructure Engineering strategy: custom AI models for generating IaC per organizational standards, AI-driven anomaly detection in infrastructure, predictive scaling. Defines AI applicability boundaries and governance framework for AI-driven infrastructure changes.
Architecture & System Design · 4
Defines capacity planning strategy for cloud platform: predictive scaling based on ML models, reserved capacity management, burst-to-cloud scenarios. Designs FinOps processes — unit economics, capacity forecasting, rightsizing recommendations. Balances cost and availability at enterprise level.
Shapes organizational DR strategy: multi-region active-active vs pilot light vs warm standby, RPO/RTO matrix by criticality. Designs automated failover through Route 53 health checks and cross-region replication. Organizes regular DR drills and gameday exercises.
Shapes cloud platform scalability strategy: cell-based architecture, multi-region active-active with conflict resolution, edge computing for latency-sensitive workloads. Designs platform abstractions for automatic scaling and cost-effective peak load handling.
Shapes cloud platform architectural strategy: domain-driven cloud architecture, platform engineering vision, build vs buy framework. Designs cross-cutting concerns — observability, security, networking — as platform capabilities and defines technology governance.
Observability & Monitoring · 6
Shapes enterprise-level log management strategy: platform selection (OpenSearch vs Datadog vs Splunk), unified logging for multi-cloud, log-based security analytics. Designs data pipeline for processing petabytes of logs with cost-effective storage and compliance.
Shapes enterprise-level incident management framework: unified incident process for multi-cloud, automated incident response (AWS Systems Manager, PagerDuty Rundeck), AIOps for anomaly detection. Defines organizational resilience strategy and chaos engineering program.
Shapes enterprise observability platform based on OpenTelemetry: multi-cloud telemetry aggregation, custom semantic conventions for cloud resources, auto-instrumentation platform. Designs AI-driven analysis based on telemetry data and defines observability maturity framework.
Shapes enterprise-level observability platform: unified metrics across clouds (Grafana Cloud/Datadog), custom metric taxonomies, FinOps dashboards. Designs architecture for processing millions of time series with cost-effective storage and sub-second query latency.
Shapes reliability engineering strategy: platform-wide SLO framework, business-aligned reliability targets, cost of reliability analysis. Designs SLO platform for automated tracking of hundreds of services, defines reliability investment priorities at organizational level.
Shapes platform-level logging governance: schema registry for log formats, automated compliance validation, cross-service correlation framework. Designs structured log integration with tracing and metrics for unified observability stack.
Version Control & Collaboration · 3
Shapes platform-level review governance: automated compliance validation, risk-based review tiers, architecture review board for major changes. Designs tooling for automated IaC review — policy engines, cost estimation, security scanning integrated into review workflow.
Shapes enterprise-level documentation strategy: unified service catalog, automated architecture documentation from IaC, knowledge base for operational procedures. Designs internal developer platform with self-service docs, API reference and interactive architecture explorer.
Shapes Git strategy for infrastructure platform: mono-repo governance for hundreds of modules, automated dependency management, release management for shared IaC libraries. Designs developer experience for self-service infrastructure provisioning through Git-based workflows.
Documentation · 1
Shapes enterprise-level operational knowledge management: AI-assisted runbook generation, automated validation through chaos engineering, runbook-as-code in Git. Designs operational knowledge management platform with versioning, testing and continuous improvement.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
} in the open competency matrix: 58 skills across 5 levels. The matrix is free for individuals and stays free.