Evaluates algorithmic complexity of platform solutions: scheduling algorithms for K8s, routing in service mesh, index strategies for log storage. Leads optimization of IDP critical path components. Conducts design reviews focused on scalability and computational efficiency.
Roles · Platform Engineer · Lead
What a Lead } should know
47 core skills, 63 in total. Expectations per skill, and what changes at the next level.
This page lists what a Lead } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Cloud & Infrastructure, DevOps & CI/CD.
Core skills for a Lead
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 4
Defines code quality standards for platform services: linting rules, testing requirements, documentation standards. Implements automated quality gates in CI/CD: coverage thresholds, static analysis, dependency scanning. Creates quality-focused culture through code review practices and pair programming.
Applies data structure knowledge for designing platform components: B-tree for storage engine, bloom filters for cache, CRDT for distributed state. Leads optimal data structure selection for high-performance platform services considering memory and CPU trade-offs.
Implements architectural patterns for platform services: Operator pattern for K8s, Sidecar/Ambassador for service mesh, Circuit Breaker for resilience. Leads design pattern standardization through golden path templates. Conducts pattern-focused design reviews.
Cloud & Infrastructure · 18
Defines Ansible automation standards for platform engineering teams, establishing infrastructure-as-code practices for the internal developer platform. Creates governance frameworks for platform Ansible collections, testing requirements, and Tower/AWX workflow standards for platform operations. Conducts architecture reviews of platform automation and drives adoption of self-service infrastructure provisioning patterns across the organization.
Defines organizational AWS strategy: Well-Architected review, cost governance, security baseline. Leads AWS migration or optimization of existing infrastructure. Designs multi-region DR strategy. Manages AWS Enterprise Support and TAM interactions.
Defines organizational edge strategy: edge computing vs centralized, latency budget, cost optimization. Leads edge platform adoption (Cloudflare Workers, Deno Deploy) for distributed workloads. Designs global traffic management.
Defines corporate container security framework: SBOM generation, SLSA compliance, supply chain attestation. Coordinates shift-left approach adoption with security team. Manages vulnerability management program for all platform services with MTTR metrics.
Defines infrastructure strategy with Crossplane. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines corporate containerization strategy: image governance, supply chain security (SBOM, SLSA). Manages multi-registry architecture with geo-replication. Implements admission controllers for image validation before deployment. Trains teams on containerization best practices for the platform.
Defines infrastructure strategy with Envoy Proxy. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines infrastructure strategy with Google Cloud Platform. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines packaging and distribution strategy for internal platform: Helm library charts, umbrella charts for environments. Manages chart lifecycle: versioning, deprecation, migration. Creates self-service component catalog based on Helm for development teams.
Defines infrastructure strategy with Istio Service Mesh. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines multi-cluster Kubernetes strategy for organization: federation, fleet management (Rancher/Tanzu). Leads platform API development through CRDs and controllers. Designs cluster lifecycle management with automatic upgrades and disaster recovery.
Defines organizational Kubernetes strategy: managed vs self-hosted, version policy, upgrade process. Leads capacity planning and cost optimization for clusters. Creates internal standards and best practices for Kubernetes deployment across all teams.
Defines organizational traffic management strategy: API gateway, service mesh, load balancing as unified architecture. Leads progressive delivery adoption through traffic shifting. Designs multi-region active-active with intelligent failover for critical services.
Defines organizational network strategy: zero-trust architecture, network segmentation, compliance. Leads network automation adoption: NetDevOps, infrastructure-as-code for networking. Designs multi-region networking considering latency and data residency.
Defines infrastructure strategy with Pulumi. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines infrastructure strategy with Serverless Containers. Establishes IaC standards. Conducts architecture reviews. Optimizes FinOps.
Defines organizational serverless strategy: use cases vs containers, cost model, vendor lock-in analysis. Leads serverless platform creation with self-service for teams. Designs multi-region failover and disaster recovery for serverless workloads.
Defines organizational IaC strategy: Terraform Enterprise/Cloud, module governance, state management. Leads internal module registry development. Designs blast radius isolation through workspace strategy. Implements drift detection and automated remediation.
DevOps & CI/CD · 8
Defines ArgoCD platform strategy as the foundation of the organization's GitOps-driven developer experience. Establishes ArgoCD governance standards including multi-tenant project hierarchies, platform team RBAC models, and self-service onboarding workflows for development teams. Drives architectural decisions on ArgoCD extensibility and integration with the broader platform toolchain.
Defines zero-downtime deployment strategy: when to use blue-green vs canary vs rolling update. Leads standardization of deployment strategies on the platform. Designs cost-efficient dual-environment approach (spot instances, autoscaling). Creates deployment reliability metrics.
Defines organizational progressive delivery strategy: canary + feature flags + observability as a unified process. Leads automated canary analysis adoption across all services. Designs blast radius management through graduated rollout policies and automatic halts.
Defines organizational feature management strategy: governance, naming conventions, audit trail. Leads progressive delivery adoption through flags for all teams. Designs impact analytics and automatic rollback based on metrics when SLI degrades.
Defines organizational CI/CD strategy on GitHub Actions: governance, compliance checks, security scanning pipeline. Leads migration from legacy CI to Actions. Designs runner infrastructure with autoscaling (ARC). Optimizes costs and performance at enterprise level.
Defines organizational GitLab CI/CD strategy: instance vs SaaS, runner infrastructure, license optimization. Leads internal component library creation. Designs compliance framework with automatic audit trails. Integrates GitLab with IDP for end-to-end developer workflow.
Standardizes GitOps at platform level: defines platform team's GitOps workflow for infrastructure changes, creates golden path templates, designs policy-as-code enforcement via OPA/Kyverno integrated into GitOps. Ensures platform self-healing through GitOps reconciliation.
Defines DevOps strategy with Progressive Delivery. Establishes CI/CD standards. Implements platform engineering approaches.
Testing & QA · 1
Standardizes chaos engineering at platform level: designs automated resilience scoring infrastructure, creates chaos experiment marketplace for reuse. Defines platform-level chaos: testing platform components themselves (control plane, etcd, ingress).
Security · 1
Defines organizational secrets management strategy: Vault Enterprise features, namespaces for BU isolation, audit compliance. Leads zero-trust secrets adoption: transit encryption, tokenization. Designs DR strategy for Vault and key management governance process.
AI-Assisted Development · 1
Defines organizational GitHub Copilot strategy: Enterprise deployment, usage policies, security review process. Leads AI-assisted development ROI measurement. Creates governance framework for AI-generated infrastructure code. Integrates Copilot into developer experience metrics.
Architecture & System Design · 4
Defines organizational capacity management strategy: FinOps practices, showback/chargeback model, reserved capacity vs spot. Leads quarterly capacity planning review with technical and business stakeholders. Designs multi-region capacity strategy with failover reserves.
Defines organizational DR strategy: tiered RPO/RTO by service criticality, budget allocation, compliance requirements. Leads game days and tabletop exercises for DR plans. Designs organizational DR governance with regular review and improvement cycles.
Defines organizational scaling strategy: SLA tiers, capacity tiers, performance budgets. Leads performance engineering team. Designs multi-region architecture for low-latency globally distributed platform. Creates performance governance and review process.
Defines architectural standards and principles for the platform: API-first design, modular architecture, domain boundaries. Leads Architecture Guild and review process. Creates Technology Radar for the organization. Designs evolutionary architecture for long-term IDP development.
Observability & Monitoring · 8
Defines organizational metrics strategy: golden signals for each tier, cardinality budget, cost allocation. Leads observability standards adoption. Designs metric-driven decision framework: automated scaling, deployment decisions, capacity planning based on custom metrics.
Defines centralized logging strategy: ELK vs alternatives (Loki, Datadog), cost-performance balance. Leads capacity planning for logging infrastructure. Designs compliance-compliant log retention with audit trail. Creates observability standards and SLA for log platform.
Adopts Grafana Loki as cost-effective logging solution for the platform: multi-tenant configuration, retention policies. Designs label strategy for optimal query performance. Integrates with Grafana for unified observability (logs + metrics + traces in single UI).
Defines organizational incident management strategy: on-call expectations, compensation, burnout prevention. Leads MTTR improvement through automation and tooling. Designs cross-team incident coordination for complex incidents. Creates incident readiness program.
Defines OpenTelemetry-based observability strategy: vendor-neutral collection, backend strategy, cost optimization. Leads organizational migration to OTel. Designs observability data pipeline with routing, sampling, and aggregation. Creates OTel maturity roadmap.
Defines metrics platform strategy: build vs buy, cardinality management, cost optimization. Leads observability-as-code adoption across all teams. Designs SLO-based monitoring with automated alerting. Creates observability maturity model for the organization.
Defines organizational SLO strategy: tiered SLO targets, error budget governance, SLA management process. Leads SRE practice adoption through SLO framework. Designs organizational error budget policy: freeze deployments, allocate engineering time on depletion.
Defines organizational logging standards: compliance requirements (PII masking), retention policies, access control. Leads logging best practices adoption through golden paths. Designs log analytics pipeline for business insights and incident investigation.
Version Control & Collaboration · 2
Establishes organizational code review culture: cross-team reviews for shared platforms, architecture review for major changes. Leads CODEOWNERS and required approvals adoption for critical infrastructure. Designs review workflows balancing velocity and quality.
Defines organizational Git strategy: repository structure, branching model, code ownership for all teams. Leads migration between Git platforms (GitLab/GitHub). Designs inner source model with cross-team contribution workflows. Creates Git governance and audit process.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
What changes at Principal
0 skills get a higher expectation or become core when moving from Lead to Principal. The biggest jumps first.
} in the open competency matrix: 63 skills across 5 levels. The matrix is free for individuals and stays free.