Roles · ML Engineer · Mid-level

What a Mid-level } should know

26 core skills, 58 in total. Expectations per skill, and what changes at the next level.

This page lists what a Mid-level } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, API & Integration.

26core skills
32additional skills
9skill areas
2%at Advanced or Expert
Assess myself as Mid-level Full role matrix

Core skills for a Mid-level

Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.

Programming Fundamentals · 4

Evaluates algorithm complexity of data processing in ML pipelines. Understands memory/speed trade-offs in feature engineering. Optimizes batch operations considering computational complexity.

Applies type hints in ML code. Uses mypy for static analysis. Writes unit tests for data processing and model evaluation. Organizes ML code into modules (data, features, models, evaluation).

Data Structures Working

Effectively uses data structures for ML: sparse matrices, ordered structures, heaps. Works with pandas MultiIndex and categorical data. Optimizes dataset memory footprint.

Applies OOP for structuring ML code: abstract classes for models, strategies for feature engineering. Uses patterns for ML component reuse. Writes custom sklearn transformers.

Backend Development · 1

Designs ML API with FastAPI: async endpoints, pydantic validation, batch prediction. Implements health checks and model versioning in API. Integrates API with model registry.

API & Integration · 1

REST API Design Working

Designs RESTful ML API: batch prediction, model versioning, health checks. Documents ML API with OpenAPI. Implements pagination for prediction results. Handles model errors.

Cloud & Infrastructure · 2

Docker Working

Creates optimized Docker images for ML: multi-stage for training and serving, CUDA-based images. Configures GPU access in Docker. Uses .dockerignore for ML artifacts.

Kubernetes Core Working

Deploys ML services in Kubernetes. Configures resource limits for CPU/GPU workloads. Uses ConfigMaps for model configuration. Configures HPA for ML serving autoscaling.

Testing & QA · 1

Unit Testing Working

Writes comprehensive tests for ML: data validation, model prediction format, edge cases. Uses fixtures for ML test data. Tests pipeline components in isolation.

Data Engineering · 4

Apache Spark Working

Uses PySpark for large-scale feature engineering. Optimizes Spark jobs (partitioning, caching, broadcast joins). Uses Spark ML for distributed model training.

Data Quality Working

Uses Great Expectations/Soda for data validation. Configures automated data quality checks in ML pipeline. Monitors data drift before retraining.

Pandas / Polars Working

Optimizes pandas code for ML: vectorized operations, category dtype, chunked reading. Uses Polars for faster processing. Writes efficient feature engineering pipelines.

SQL-based ETL Working

Designs SQL ETL for feature computation. Uses dbt for ML feature transformation. Writes incremental ETL for updating training data. Automates through Airflow.

Machine Learning & AI · 9

Designs sklearn Pipelines for production. Performs feature selection (SelectKBest, RFE). Configures hyperparameter tuning (GridSearchCV, RandomizedSearchCV, Optuna). Handles imbalanced data (SMOTE, class_weight).

Designs experiment tracking workflow. Organizes experiments by projects/tasks. Configures hyperparameter sweeps (Optuna, W&B Sweeps). Analyzes results for decision making.

Feature Stores Working

Configures Feast for the project. Defines feature definitions (entities, feature views). Configures materialization for online store. Integrates feature store with training pipeline.

Performs hyperparameter tuning for gradient boosting (learning_rate, max_depth, n_estimators, regularization). Handles categorical features (CatBoost native, target encoding). Configures early stopping and cross-validation. Analyzes SHAP values.

ML Pipelines Working

Designs ML pipelines with Kubeflow/Airflow. Configures parameterized pipelines for different models. Automates retraining with data quality checks. Implements pipeline testing.

MLflow Working

Designs MLflow workflow: experiment naming, run tags, artifact storage. Uses Model Registry for versioning. Configures autologging for sklearn/PyTorch. Writes custom MLflow Plugins.

Configures data drift detection (Evidently, NannyML). Monitors feature distributions. Configures alerting on model degradation. Implements automated retraining trigger.

Model Serving Working

Uses model serving frameworks: Triton, BentoML, Seldon. Configures batch and real-time inference. Optimizes inference latency (ONNX, model optimization). Configures A/B testing for models.

PyTorch Working

Designs custom models in PyTorch. Configures training loop: optimizer, scheduler, early stopping. Uses transfer learning (fine-tuning pretrained models). Logs experiments in MLflow/W&B.

Observability & Monitoring · 2

Writes structured logs in JSON format. Adds correlation IDs for tracing. Uses proper log levels. Configures log aggregation (EFK/Loki). Does not log sensitive data (PII, passwords).

Adds custom metrics to application (counter, gauge, histogram). Writes PromQL queries for dashboards. Creates Grafana dashboards. Configures basic alerts (high error rate, high latency).

Version Control & Collaboration · 2

Code Review Working

Reviews ML code: checks for data leakage, feature correctness, model evaluation. Gives constructive feedback. Verifies experiment reproducibility.

Git Advanced Working

Uses DVC for version control of data and models. Organizes ML code by branches: experiments, features, releases. Resolves conflicts in ML configurations.

Additional skills

Not assessed by the team, but part of the self-assessment and the development plan.

API DocumentationAsync ProgrammingAWSCursor IDEData Modeling & Schema DesignDesign PatternsE2E TestingGitHub Actions / GitLab CIGraphQL DesignIntegration TestingJWT / OAuth2 / OIDCMemory ManagementMultithreadingNetwork FundamentalsOpenTelemetryOWASP & Application SecurityPostgreSQLSecure Coding PracticesSLI / SLO / SLASystem Design FundamentalsTerraformType Safety & Type SystemsWebSocket API DesignApache KafkaChatGPT / ClaudeDatabase IndexingGitHub CopilotgRPC & Protocol BuffersPrompt Engineering for CodeQuery OptimizationRedisTask Queues

What changes at Senior

58 skills get a higher expectation or become core when moving from Mid-level to Senior. The biggest jumps first.

See the Senior page →
Run this with your whole team
Self-assessment plus manager and peer reviews against the same matrix, gap analysis and next-level readiness for every engineer. Team Pro is free for 14 days; individual tools stay free forever.
Start a team trial (14 days free) Send to my manager

} in the open competency matrix: 58 skills across 5 levels. The matrix is free for individuals and stays free.