Evaluates computational complexity of SQL transformations and dbt models on large data volumes. Optimizes aggregation algorithms and window functions in analytical queries. Implements query plan profiling practices in the data warehouse.
Roles · Analytics Engineer · Lead
What a Lead } should know
37 core skills, 53 in total. Expectations per skill, and what changes at the next level.
This page lists what a Lead } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Database Management.
Core skills for a Lead
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 3
Establishes dbt project quality standards: sqlfluff for SQL linting, pre-commit hooks for YAML validation, mandatory dbt tests in CI/CD. Implements code review processes focused on model performance and readability.
Defines data modeling standards in dbt: choosing between nested structures and normalized tables, ARRAY/MAP types in analytical queries. Reviews data structure decisions in the context of BI dashboard performance.
Backend Development · 4
Defines the strategy for integrating streaming data into the analytics platform: Kafka topics as sources for near-real-time dashboards, CDC streams for incremental models, schema evolution policies for analytics contracts.
Defines the search functionality architecture for the analytics platform: Elasticsearch for data discovery, autocomplete for metric and model names, fuzzy search by descriptions and tags in the data catalog.
Defines the data services architecture: APIs for metadata catalog, webhook integrations with BI tools, self-service endpoints for analysts. Standardizes approaches to serializing analytical data through APIs.
Defines the analytics platform caching strategy: Redis for hot metrics and KPIs, invalidation policies on model recalculation. Implements caching at the semantic layer for accelerating BI queries.
Database Management · 6
Defines the ClickHouse usage strategy within the analytics platform: as an engine for real-time dashboards, integration with dbt via the ClickHouse adapter. Implements modeling standards and TTL policies for storage management.
Defines data modeling standards for the organization: naming conventions, dbt project layering (staging/intermediate/marts), mandatory tests and documentation. Implements data mesh approaches with domain-oriented models.
Defines indexing standards for the analytics platform. Implements automated query performance monitoring and optimization recommendations. Trains the team on reading query plans in the warehouse context.
Defines the data management strategy at the product level. Establishes database migration standards. Conducts reviews of data schemas and scaling strategies.
Defines the PostgreSQL usage strategy in the analytics stack: read replicas for BI queries, materialized views for intermediate aggregations. Implements naming and schema documentation standards for the data catalog.
Defines performance standards for analytical models: SLA on execution time, compute budgets. Implements automated performance testing for dbt models and alerting on query speed degradation.
API & Integration · 2
Defines integration documentation standards: templates for API source descriptions, mandatory fields in data contracts, automated validation of documentation matching actual schemas.
Defines API integration standards for the analytics stack: choosing between custom extractors and managed tools (Fivetran, Airbyte), SLA for API data freshness. Implements data contracts for internal APIs as analytics sources.
Cloud & Infrastructure · 2
Defines the AWS strategy for organizational analytics: choosing between Redshift and Athena for different workloads, cost management through reserved capacity and serverless. Implements tagging and cost allocation for analytics resources.
Defines containerization standards for the analytics stack: base images for dbt projects, templates for data extraction jobs, security policies for images. Implements a container registry and versioning for analytics tools.
DevOps & CI/CD · 1
Defines CI/CD standards for organizational analytics projects: reusable workflows for dbt, automated data quality gates, cost estimation in PRs. Implements a GitOps approach for managing warehouse resources and permissions.
Testing & QA · 2
Defines integration testing standards: dbt source freshness checks, cross-model consistency tests, automated BI dashboard validation. Implements contract testing between data producers and analytical models.
Defines analytical model testing standards: mandatory test coverage for the mart layer, automated regression suites for KPI metrics. Implements test environments with representative data samples for model validation.
Data Engineering · 12
Defines orchestration strategy for the analytics pipeline: Airflow for coordinating dbt runs, sensors for upstream data dependencies. Implements DAG design standards: idempotency, retry policies, SLA monitoring for analytical models.
Defines the organization's BI strategy: tool selection and standardization (Looker vs Tableau vs Metabase), governance for metrics and dashboards, self-service analytics for business users. Implements a semantic layer for unified metric definitions.
Implements modern orchestration for analytics: Dagster assets as native integration with dbt models, software-defined assets for Python transformations. Defines standards for observable, testable pipelines with built-in data quality checks.
Implements a data catalog for the analytics platform: integrating dbt docs with an enterprise catalog (Atlan, DataHub, Alation), automated metadata enrichment. Defines standards for classifying and tagging analytical models.
Implements data contracts on the analytics platform: schema contracts in dbt, SLA for freshness and quality between source teams and analytics engineering. Defines processes for contract change management.
Defines the data lake architecture for the analytics platform: medallion approach (bronze/silver/gold), storage format selection (Parquet, Delta, Iceberg). Implements partitioning and retention standards for cost optimization.
Implements data lineage for the analytics platform: dbt lineage graph, column-level lineage, BI integration for end-to-end tracking. Uses lineage for impact analysis when changing models and during migrations.
Defines the organization's data quality standards: quality SLA per layer, mandatory tests for production models, processes for responding to quality incidents. Implements a data observability platform (Monte Carlo, Elementary).
Defines the analytics warehouse architecture: Kimball vs Data Vault vs One Big Table for different domains, clustering and partitioning strategy. Implements Snowflake/BigQuery best practice standards for dbt projects.
Defines the organization's dbt project architecture: modular package structure, cross-project references, shared macros library. Implements standards for incremental models, snapshot strategies, and dbt environment management (dev/staging/prod).
Defines standards for Python vs SQL usage on the analytics platform: when pandas/polars is justified over dbt, templates for Python models in dbt. Implements best practices for reproducible data preparation notebooks.
Defines organizational SQL transformation standards: coding style guide, mandatory patterns (surrogate keys, audit columns), dbt macros and packages library. Implements automated SQL review and performance benchmarking for critical models.
AI-Assisted Development · 1
Defines standards for using AI tools in the analytics engineering team: guidelines for Copilot in dbt projects, review processes for AI-generated SQL. Implements AI-assisted documentation and testing of analytical models.
Observability & Monitoring · 1
Defines observability standards for the analytics platform: dashboards for dbt pipeline health, SLA monitoring for data freshness and quality. Implements incident management processes for data quality issues with structured runbooks.
Version Control & Collaboration · 3
Defines code review standards for analytics engineering: checklists for dbt PRs (model, tests, docs, performance), automated checks in CI, review SLA. Builds a culture of quality through constructive feedback and knowledge sharing.
Defines documentation standards for the analytics organization: mandatory descriptions for all production models, templates for new models and packages. Implements metrics and SLA for documentation coverage, automatic reminders for undocumented models.
Defines version control standards for the analytics organization: mono-repo vs multi-repo for dbt projects, semantic versioning for dbt packages, release management. Implements Git-based governance for production models.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
What changes at Principal
0 skills get a higher expectation or become core when moving from Lead to Principal. The biggest jumps first.
} in the open competency matrix: 53 skills across 5 levels. The matrix is free for individuals and stays free.