Designs algorithmically efficient ML systems. Optimizes end-to-end pipelines considering computational complexity. Applies approximate algorithms for large-scale data (HyperLogLog, Count-Min Sketch).
Roles · ML Engineer · Senior
What a Senior } should know
26 core skills, 58 in total. Expectations per skill, and what changes at the next level.
This page lists what a Senior } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, API & Integration.
Core skills for a Senior
Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.
Programming Fundamentals · 4
Designs ML code quality standards. Introduces pre-commit hooks for ML projects. Configures CI for ML (tests, linting, data validation). Creates cookiecutter templates for ML projects.
Designs custom data structures for ML pipelines. Works with PyTorch Dataset/DataLoader. Optimizes data layouts for GPU computations. Uses memory-mapped files for large datasets.
Designs ML frameworks using SOLID principles. Creates extensible abstractions for training loops, model registry, experiment tracking. Applies composition over inheritance for ML components.
Backend Development · 1
Designs production-ready ML API: rate limiting, prediction caching, A/B testing through routing. Optimizes API for high-throughput inference. Uses middleware for logging and monitoring.
API & Integration · 1
Designs production ML API: A/B testing through routing, model fallback, caching. Optimizes API for high-throughput inference. Implements streaming responses for generative models.
Cloud & Infrastructure · 2
Designs Docker strategy for ML workloads. Optimizes build cache for ML images. Creates base images for ML team. Configures GPU scheduling in Docker.
Designs Kubernetes deployment strategy for ML. Configures GPU scheduling and node affinity. Optimizes resource utilization for ML workloads. Uses Kubernetes Operators for ML.
Testing & QA · 1
Designs testing strategy for ML systems. Tests model behavior (regression tests, invariance tests). Configures automated model quality tests in CI.
Data Engineering · 4
Designs Spark-based ML pipelines for production. Optimizes Spark for ML workloads: memory tuning, shuffle optimization. Integrates Spark with ML platform (MLflow, feature store).
Designs data quality framework for ML. Integrates data validation into ML pipeline. Configures alerting on data anomalies. Defines data quality SLAs for ML.
Designs data processing pipelines for ML. Chooses pandas vs Polars vs Spark for different scales. Optimizes memory usage for large datasets. Writes reusable feature transformers.
Designs ETL architecture for ML data pipeline. Optimizes ETL for large data volumes. Configures data quality checks in ETL. Integrates ETL with feature store.
Machine Learning & AI · 9
Designs ML systems on scikit-learn for production. Creates custom transformers and estimators. Optimizes pipeline performance. Integrates sklearn with MLflow for tracking and serving.
Designs experiment tracking infrastructure. Automates experiment analysis. Integrates tracking with CI/CD for automated model promotion.
Designs feature store architecture. Optimizes materialization for large volumes. Configures streaming feature computation. Ensures feature consistency between training and serving.
Designs production gradient boosting systems. Optimizes inference speed (model pruning, quantization). Builds ensembles from multiple gradient boosting models. Integrates with feature store and model serving.
Designs ML pipeline architecture. Optimizes pipeline execution (caching, parallel steps). Configures CI/CD for pipeline deployment. Ensures reproducibility.
Designs MLflow infrastructure for the team. Configures MLflow on Kubernetes. Integrates MLflow with CI/CD for automated model promotion. Creates custom model flavors.
Designs model monitoring architecture. Configures custom monitoring for specific ML tasks. Integrates monitoring with ML pipeline for closed-loop retraining.
Designs model serving architecture. Optimizes throughput (batching, GPU scheduling). Configures autoscaling for ML serving. Implements model fallback and canary deployment.
Designs custom architectures and training frameworks. Optimizes inference: ONNX export, TensorRT. Configures distributed training (DDP, FSDP). Works with PyTorch Lightning for production training.
Observability & Monitoring · 2
Designs logging strategy: retention, sampling, costs. Configures centralized logging (EFK/Loki/Datadog). Optimizes log volume and cost. Integrates logs with traces (via OpenTelemetry). Configures log-based alerting.
Designs observability stack for the service. Creates SLI/SLO dashboards. Writes complex PromQL (rate, histogram_quantile, recording rules). Configures Alertmanager with routing and silencing. Optimizes metric cardinality. Integrates with PagerDuty/OpsGenie.
Version Control & Collaboration · 2
Reviews ML system architecture. Verifies production readiness of ML code. Mentors through code review. Establishes ML code review checklist.
Designs Git workflow for ML team. Configures Git hooks for ML code quality. Integrates DVC with CI/CD. Manages ML monorepo/multirepo strategy.
Additional skills
Not assessed by the team, but part of the self-assessment and the development plan.
What changes at Lead
37 skills get a higher expectation or become core when moving from Senior to Lead. The biggest jumps first.
- Apache Spark: Advanced → Expert
- Classical ML (scikit-learn): Advanced → Expert
- Data Quality: Advanced → Expert
- Experiment Tracking: Advanced → Expert
- Feature Stores: Advanced → Expert
- Gradient Boosting: Advanced → Expert
- ML Pipelines: Advanced → Expert
- MLflow: Advanced → Expert
- Model Monitoring: Advanced → Expert
- Model Serving: Advanced → Expert
} in the open competency matrix: 58 skills across 5 levels. The matrix is free for individuals and stays free.