Shapes algorithmic optimization strategy across the entire data platform. Designs algorithms for internal DBA tools: automatic execution plan regression detection, statistical workload pattern analysis, predictive capacity planning.
Roles · Database Engineer / DBA · Principal
What a Principal } should know
40 core skills, 55 in total. Expectations per skill, and what changes at the next level.
This page lists what a Principal } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Database Management, Cloud & Infrastructure.
Core skills for a Principal
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 3
Defines quality strategy for the entire database layer: testing pyramid for DBs (unit → integration → load), automated schema validation, compliance checking for data governance. Establishes database reliability engineering culture.
Designs data models for the entire platform: sharding strategies considering data locality, time-series storage schemes for DB metrics, hierarchical structures for RBAC. Defines storage approaches considering retention policies and query performance.
Database Management · 16
Defines Cassandra strategy for the data platform: partition key design for DBA metrics, consistency levels for different use cases, multi-DC replication. Evaluates Cassandra vs ScyllaDB for DB monitoring time-series workloads.
Shapes organizational backup/recovery strategy: RPO/RTO by tier, automated backup verification (restore testing), cross-region backup replication. Defines DR playbooks, conducts regular DR drills. Evaluates backup solutions: native vs Percona XtraBackup vs cloud-native.
Shapes analytics platform strategy: ClickHouse vs TimescaleDB vs InfluxDB for database observability. Designs real-time analytics architecture for monitoring the entire organization's database fleet.
Evaluates CockroachDB for distributed SQL workloads: automatic sharding, multi-region deployment, PostgreSQL compatibility. Defines scenarios where CockroachDB is preferred over classic RDBMS for geo-distributed applications.
Shapes organizational data modeling strategy: domain-driven data design, data mesh principles, enterprise data model. Defines modeling standards for different DBMSes and use cases, including multi-model approaches.
Shapes indexing strategy for the entire data platform: automated index efficiency analysis, ML-based index recommendations, index maintenance standards. Defines indexing approaches for different DBMSes in the organization.
Shapes organizational migration strategy: tooling standards (gh-ost vs pt-osc vs native), cross-database migration patterns, automated migration risk assessment. Defines governance for schema changes across the entire data platform.
Evaluates DynamoDB as an alternative for specific workloads: session storage, metadata catalog, configuration management. Defines when DynamoDB vs relational DBs are optimal. Establishes standards for DynamoDB integration into the data platform.
Defines MongoDB strategy for the data platform: use cases for document storage (configurations, audit logs), sharding strategy, integration with relational DBs. Evaluates MongoDB Atlas vs self-hosted for different requirements.
Shapes company-wide MySQL strategy: MySQL vs Aurora vs Vitess for different workloads, multi-region replication standards, cost optimization. Defines MySQL platform development roadmap and evaluates alternative solutions.
Evaluates graph databases for DBA tasks: visualizing database dependencies, mapping data lineage, analyzing migration impact. Defines when Neo4j complements relational DBs in infrastructure management.
Shapes organizational PostgreSQL strategy: Patroni vs Stolon vs Citus for different patterns, cross-region DR, RTO/RPO by tier. Defines PostgreSQL platform standards, evaluates managed vs self-hosted approaches.
Shapes organizational query optimization strategy: automated query analysis platform, ML-driven optimization recommendations, cross-database query performance standards. Defines investments in query performance tooling.
Defines replication strategy: sync vs async vs semi-sync for different tiers, multi-region replication topology, conflict resolution for multi-master. Establishes replication monitoring standards, lag SLA, and automated failover procedures.
Defines transaction management strategy: isolation levels for different workloads (OLTP vs analytics), distributed transactions (2PC, Saga), optimistic vs pessimistic locking guidelines. Establishes data consistency standards for the entire platform.
Shapes MySQL horizontal scaling strategy via Vitess: VSchema design, sharding strategy, migration from monolithic MySQL. Defines when Vitess is justified vs native MySQL sharding or switching to other DBMSes.
Cloud & Infrastructure · 3
Shapes organizational cloud database strategy: multi-cloud vs AWS-only, managed vs self-hosted, Aurora vs RDS vs EC2. Defines cloud database governance, security compliance, and architectural patterns for the enterprise data tier.
Shapes container strategy for the data platform: Docker vs Kubernetes operators for stateful workloads, storage orchestration (OpenEBS, Portworx). Defines when containerized databases are justified vs dedicated infrastructure.
Shapes IaC strategy for the data platform: Terraform vs Crossplane vs database operators, self-service database provisioning. Defines governance for database infrastructure changes, including approval workflows for production.
Security · 1
Shapes database security strategy via secrets management: zero-trust database access, dynamic credentials for all tiers, encryption key management. Defines compliance requirements and audit standards for database credentials.
AI-Assisted Development · 1
Shapes AI-assisted database management strategy: automated query tuning, AI-driven capacity planning, predictive maintenance. Evaluates AI/ML tools for database operations: autonomous databases, self-healing infrastructure.
Architecture & System Design · 2
Shapes organizational capacity planning strategy: automated capacity forecasting for all DBMSes, cost optimization models, multi-year planning. Defines investment strategy for database infrastructure.
Shapes organizational disaster recovery strategy: multi-region active-active vs active-passive, cross-cloud DR, RTO/RPO frameworks. Defines DR governance, compliance requirements, and investments in database resilience.
Observability & Monitoring · 6
Shapes data platform metrics strategy: custom business metrics via database, cardinality management, metrics for capacity planning and cost attribution. Defines the framework for organizational database performance KPIs.
Shapes log analytics strategy for the data platform: ELK vs Loki vs ClickHouse for database logs, cost optimization at high volume. Defines unified logging architecture for the entire database fleet.
Shapes incident management strategy for the data platform: automated incident response, AI-assisted diagnostics, cross-database impact analysis. Defines on-call sustainability and investments in automation for database operations.
Shapes database observability strategy: Prometheus federation for fleet monitoring, Thanos/Mimir for long-term storage, unified dashboards. Defines metrics strategy for the entire organizational data platform.
Shapes organizational SLO strategy for the data tier: SLO framework covering all DBMSes, SLA for internal database-as-a-service, SLO-driven investment decisions. Defines reliability culture for database engineering.
Shapes observability strategy through structured logging: unified log schema for all DBMSes, automated anomaly detection based on log patterns, compliance logging. Defines investments in log infrastructure.
Version Control & Collaboration · 2
Shapes review culture for database engineering: cross-team review standards, automated quality gates for database changes, architectural review board for data model decisions. Defines database change governance.
Shapes version control strategy for database-as-code: GitOps for database changes, automated deployment pipelines, schema versioning across environments. Defines governance for database code repositories.
Documentation · 2
Shapes documentation strategy for the data platform: automated data catalog (DataHub, Apache Atlas), living documentation for database architecture, compliance documentation. Defines knowledge management for database engineering.
Shapes organizational runbook strategy: self-healing databases through automated runbooks, AI-assisted troubleshooting, runbook marketplace for cross-team sharing. Defines investments in operational automation.
Performance Engineering · 4
Defines Benchmarking Tools strategy at organizational level. Shapes enterprise approaches. Mentors leads and architects.
Shapes performance engineering strategy for the data platform: automated performance analysis, hardware selection criteria (CPU architecture for database workloads), performance budgets. Defines investments in database performance tooling.
Defines Database Performance Tuning strategy at organizational level. Shapes enterprise approaches. Mentors leads and architects.
Defines I/O and Disk Profiling strategy at organizational level. Shapes enterprise approaches. Mentors leads and architects.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
} in the open competency matrix: 55 skills across 5 levels. The matrix is free for individuals and stays free.