Roles · Cloud Engineer · Principal

What a Principal } should know

43 core skills, 58 in total. Expectations per skill, and what changes at the next level.

This page lists what a Principal } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Cloud & Infrastructure.

43core skills
15additional skills
10skill areas
100%at Advanced or Expert
Assess myself as Principal Full role matrix

Core skills for a Principal

Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.

Programming Fundamentals · 3

Designs data processing strategies at cloud platform scale: optimal autoscaling algorithms, load balancing considering latency and cost, efficient DNS sharding and cross-region traffic routing strategies.

Defines organizational IaC quality strategy: unified Terraform module standards, golden templates for cloud services, automated compliance. Manages infrastructure technical debt and plans legacy configuration migration.

Designs data models for cloud platform management: CMDB, service catalog, cost allocation trees. Makes strategic decisions on state storage — S3 backend vs Terraform Cloud, structuring multi-account AWS Organizations hierarchies.

Backend Development · 1

Defines S3 / Object Storage strategy at company level. Evaluates new technologies and approaches. Establishes enterprise standards and reference architectures.

Cloud & Infrastructure · 18

Ansible Expert

Shapes the organization's infrastructure automation strategy with Ansible as a core tool for multi-cloud configuration management. Drives innovation in Ansible-based cloud automation including event-driven automation with ansible-rulebook, GitOps integration patterns, and self-healing infrastructure playbooks. Influences Ansible community practices through contributions to cloud provider collections and conference presentations on enterprise-scale cloud automation.

AWS Expert

Shapes enterprise-level cloud strategy: multi-region disaster recovery, hybrid cloud (Outposts, Direct Connect), compliance frameworks (SOC2, HIPAA, PCI DSS). Designs reference architectures and makes strategic decisions on cloud-native vs lift-and-shift.

Defines CDN and edge computing strategy: CloudFront/Cloud CDN for static and API acceleration, Lambda@Edge/CloudFlare Workers for edge logic. Designs caching architecture — cache invalidation, origin shield, multi-CDN failover for global applications.

Defines platform-level container security strategy: Trivy/Snyk integration in CI/CD, admission controllers in Kubernetes for blocking vulnerable images, runtime protection through Falco/Aqua. Establishes CVE remediation SLA and automates compliance reporting.

Docker Expert

Shapes container platform strategy: choosing between ECS/EKS/Fargate/Cloud Run, standards for multi-tenant environments, governance for container registries. Designs supply chain security architecture for container artifacts.

Defines GCP strategy in multi-cloud architecture: BigQuery for analytics, GKE Autopilot for Kubernetes, Anthos for hybrid cloud. Designs GCP integration with AWS/Azure, evaluates GCP advantages for ML/AI workloads and data-intensive applications.

Helm Expert

Designs organizational Helm ecosystem architecture: umbrella charts for complex applications, library charts for reuse, Helmfile for release management. Defines versioning standards, template functions and values hierarchies for multi-cluster environments.

Shapes Kubernetes platform evolution strategy: multi-cluster federation, service mesh (Istio/Linkerd), GitOps through ArgoCD/Flux. Designs custom operators, admission webhooks, scheduler extensions. Defines multi-tenancy architecture and resource governance.

Shapes company-level container orchestration strategy: multi-cluster topology, cross-region failover, hybrid cloud deployment. Evaluates alternatives (Nomad, ECS) and makes strategic decisions on Kubernetes platform evolution.

Defines platform-level load balancing strategy: L4 (NLB) vs L7 (ALB) vs Global (GCLB), multi-region load balancing with GeoDNS, service mesh for internal traffic. Designs zero-downtime deployment patterns and traffic management architecture for canary/blue-green releases.

Shapes Azure strategy for enterprise: Azure Landing Zones, Azure Arc for hybrid, Azure DevOps for CI/CD. Designs integration with Active Directory/Entra ID, defines Azure vs AWS use cases — compliance, Microsoft ecosystem, government cloud.

Shapes platform-level network architecture: global backbone, multi-region connectivity, cloud interconnect strategy. Designs solutions for edge computing, CDN integration, hybrid DNS. Defines network security standards and Zero Trust Architecture.

Pulumi Expert

Evaluates Pulumi as strategic alternative to Terraform: advantages of real programming languages for complex logic, Automation API for self-service platforms, Policy as Code through CrossGuard. Defines use cases for Pulumi vs Terraform in multi-cloud architecture.

Defines serverless containers strategy for the organization: AWS Fargate vs Google Cloud Run vs Azure Container Apps. Designs architecture for burst workloads, event-driven scaling, cold start optimization. Establishes cost model and guidelines for choosing between serverless and dedicated compute.

Shapes serverless architecture strategy: Lambda/Cloud Functions for event processing, API Gateway integration, Step Functions for orchestration. Designs patterns — fan-out/fan-in, saga, circuit breaker. Defines serverless applicability boundaries and migration paths.

Terraform Expert

Shapes company-level Infrastructure as Code strategy: multi-cloud provisioning, Terraform Enterprise/Cloud vs self-hosted, modular platform for self-service. Designs state management architecture for hundreds of stacks and dozens of teams.

Shapes organizational VPN strategy: Site-to-Site VPN vs Direct Connect/ExpressRoute, client VPN for remote access, mesh VPN between clouds. Designs high-availability VPN with BGP, failover and monitoring. Defines migration path to Zero Trust Network Access (ZTNA).

Yandex Cloud Expert

Defines Yandex Cloud strategy for the Russian market: Managed Kubernetes, Object Storage, Cloud Functions. Evaluates compliance with FZ-152, integration with local infrastructure, migration paths between Yandex Cloud and international providers.

DevOps & CI/CD · 4

ArgoCD Expert

Shapes the organization's GitOps strategy around ArgoCD for enterprise-scale cloud platform delivery. Drives architectural decisions on multi-cluster ArgoCD federation, disaster recovery patterns, and integration with cloud-native control planes. Influences industry practices through contributions to ArgoCD ecosystem and GitOps community standards.

Shapes blue-green deployment strategy for cloud infrastructure: Route 53 weighted routing, ALB target group switching, ECS blue-green through CodeDeploy. Designs rollback automation, database migration strategy and smoke testing for zero-downtime infrastructure updates.

Defines platform-level canary deployment strategy: progressive delivery through Flagger/Argo Rollouts, weighted traffic shifting through service mesh, automated rollback based on SLO metrics. Designs observability pipeline for canary analysis and integration with cloud-native load balancers.

Shapes infrastructure change automation platform: GitOps with GitHub Actions as orchestrator, integration with Terraform Cloud/Spacelift, multi-cloud deployment pipelines. Designs self-service developer platform with guardrails for safe resource provisioning.

Security · 2

Shapes enterprise-level cloud security strategy: Zero Trust Architecture, cloud-native SIEM (CloudTrail Lake, Chronicle), supply chain security. Defines compliance frameworks (SOC2, ISO 27001, PCI DSS), designs security reference architecture for multi-cloud.

Shapes enterprise-level secrets management platform: multi-region Vault with disaster recovery, cross-cloud secrets synchronization, zero-trust secret distribution. Defines cryptographic standards, key management strategy and HSM integration for critical workloads.

AI-Assisted Development · 1

Shapes AI-assisted Infrastructure Engineering strategy: custom AI models for generating IaC per organizational standards, AI-driven anomaly detection in infrastructure, predictive scaling. Defines AI applicability boundaries and governance framework for AI-driven infrastructure changes.

Architecture & System Design · 4

Defines capacity planning strategy for cloud platform: predictive scaling based on ML models, reserved capacity management, burst-to-cloud scenarios. Designs FinOps processes — unit economics, capacity forecasting, rightsizing recommendations. Balances cost and availability at enterprise level.

Shapes organizational DR strategy: multi-region active-active vs pilot light vs warm standby, RPO/RTO matrix by criticality. Designs automated failover through Route 53 health checks and cross-region replication. Organizes regular DR drills and gameday exercises.

Shapes cloud platform scalability strategy: cell-based architecture, multi-region active-active with conflict resolution, edge computing for latency-sensitive workloads. Designs platform abstractions for automatic scaling and cost-effective peak load handling.

Shapes cloud platform architectural strategy: domain-driven cloud architecture, platform engineering vision, build vs buy framework. Designs cross-cutting concerns — observability, security, networking — as platform capabilities and defines technology governance.

Observability & Monitoring · 6

ELK Stack Expert

Shapes enterprise-level log management strategy: platform selection (OpenSearch vs Datadog vs Splunk), unified logging for multi-cloud, log-based security analytics. Designs data pipeline for processing petabytes of logs with cost-effective storage and compliance.

Shapes enterprise-level incident management framework: unified incident process for multi-cloud, automated incident response (AWS Systems Manager, PagerDuty Rundeck), AIOps for anomaly detection. Defines organizational resilience strategy and chaos engineering program.

OpenTelemetry Expert

Shapes enterprise observability platform based on OpenTelemetry: multi-cloud telemetry aggregation, custom semantic conventions for cloud resources, auto-instrumentation platform. Designs AI-driven analysis based on telemetry data and defines observability maturity framework.

Shapes enterprise-level observability platform: unified metrics across clouds (Grafana Cloud/Datadog), custom metric taxonomies, FinOps dashboards. Designs architecture for processing millions of time series with cost-effective storage and sub-second query latency.

Shapes reliability engineering strategy: platform-wide SLO framework, business-aligned reliability targets, cost of reliability analysis. Designs SLO platform for automated tracking of hundreds of services, defines reliability investment priorities at organizational level.

Shapes platform-level logging governance: schema registry for log formats, automated compliance validation, cross-service correlation framework. Designs structured log integration with tracing and metrics for unified observability stack.

Version Control & Collaboration · 3

Code Review Expert

Shapes platform-level review governance: automated compliance validation, risk-based review tiers, architecture review board for major changes. Designs tooling for automated IaC review — policy engines, cost estimation, security scanning integrated into review workflow.

Shapes enterprise-level documentation strategy: unified service catalog, automated architecture documentation from IaC, knowledge base for operational procedures. Designs internal developer platform with self-service docs, API reference and interactive architecture explorer.

Git Advanced Expert

Shapes Git strategy for infrastructure platform: mono-repo governance for hundreds of modules, automated dependency management, release management for shared IaC libraries. Designs developer experience for self-service infrastructure provisioning through Git-based workflows.

Documentation · 1

Shapes enterprise-level operational knowledge management: AI-assisted runbook generation, automated validation through chaos engineering, runbook-as-code in Git. Designs operational knowledge management platform with versioning, testing and continuous improvement.

Additional skills

Not assessed by the team, but part of the self-assessment and the development plan.

API DocumentationAsync ProgrammingChatGPT / ClaudeCursor IDEDesign PatternsGraphQL DesignIntegration TestingMultithreadingOOP & SOLID PrinciplesPostgreSQLPrompt Engineering for CodeRedisREST API DesignSecure Coding PracticesUnit Testing
Run this with your whole team
Self-assessment plus manager and peer reviews against the same matrix, gap analysis and next-level readiness for every engineer. Team Pro is free for 14 days; individual tools stay free forever.
Start a team trial (14 days free) Send to my manager

} in the open competency matrix: 58 skills across 5 levels. The matrix is free for individuals and stays free.