Roles · LLM Engineer · Lead

What a Lead } should know

31 core skills, 80 in total. Expectations per skill, and what changes at the next level.

This page lists what a Lead } is expected to know and do, skill by skill. Core skills are the ones a manager and peers assess in a review cycle; the rest count only in self-assessment. Main areas: Programming Fundamentals, Backend Development, Database Management.

31core skills
49additional skills
9skill areas
100%at Advanced or Expert
Assess myself as Lead Full role matrix

Core skills for a Lead

Grouped by area. The label on the right is the expected depth: Awareness, Working, Advanced or Expert.

Programming Fundamentals · 2

Defines async programming standards for LLM team: streaming response architecture guidelines, concurrent inference design reviews, async RAG pipeline patterns. Establishes best practices for async patterns in LLM serving systems.

Defines OOP/SOLID standards for LLM team: chain/agent class hierarchy guidelines, model provider interface contracts, retrieval module architecture patterns. Conducts reviews of abstraction design in LLM pipeline components.

Backend Development · 5

Defines Elasticsearch/OpenSearch architectural decisions at product level. Establishes standards. Conducts design reviews and defines technical roadmap.

Defines architecture for LLM serving APIs and prompt management platforms using FastAPI. Establishes standards for streaming response endpoints, token usage tracking, and guardrail middleware. Conducts design reviews and defines the roadmap for LLM gateway infrastructure.

Redis Expert

Defines Redis architectural decisions at product level. Establishes standards. Conducts design reviews and defines technical roadmap.

Defines S3/Object Storage architectural decisions at product level. Establishes standards. Conducts design reviews and defines technical roadmap.

Task Queues Expert

Defines Task Queues architectural decisions at product level. Establishes standards. Conducts design reviews and defines technical roadmap.

Database Management · 1

PostgreSQL Expert

Defines PostgreSQL data strategy for LLM products: establishes standards for pgvector usage and RAG data architecture, designs database patterns for conversation management at scale, drives adoption of PostgreSQL best practices for LLM application data infrastructure.

API & Integration · 3

Defines API strategy at product level. Establishes design standards. Conducts API design reviews. Coordinates cross-team API interaction.

Defines API strategy for LLM services at product level: LLM inference API standards, prompt management API governance, usage metering endpoint policies. Conducts API architecture reviews for LLM platform services and establishes LLM API design practices.

Defines API strategy at product level. Establishes design standards. Conducts API design reviews. Coordinates cross-team API interaction.

Cloud & Infrastructure · 2

AWS Expert

Defines AWS infrastructure strategy for LLM serving platforms spanning GPU clusters, Bedrock integration, and model deployment automation. Establishes IaC standards for GPU instance management, model artifact lifecycle, and inference scaling policies. Conducts architecture reviews optimizing LLM serving costs and coordinates FinOps for GPU compute spending.

Defines infrastructure strategy with Kubernetes Core. Establishes IaC standards. Conducts architecture review. Optimizes FinOps.

Testing & QA · 1

Unit Testing Expert

Defines testing strategy at product level for LLM systems: prompt evaluation testing standards, chain/agent testing governance, LLM output quality testing frameworks. Establishes shift-left testing culture with automated quality gates for LLM pipeline deployments.

Data Engineering · 1

Defines data engineering strategy. Establishes data platform. Coordinates data teams. Optimizes data mesh/data fabric approaches.

Machine Learning & AI · 12

Defines agent framework architecture standards and evaluation methodologies for the organization's LLM engineering teams. Establishes best practices for agent testing, safety guardrails, cost management, and production monitoring across agent-based systems. Drives architectural decisions on framework selection, custom versus off-the-shelf agent infrastructure, and integration patterns with existing ML platform services.

Defines experiment tracking strategy for LLM product teams: standardizes tracking of prompt engineering iterations, model evaluations, and cost/quality trade-offs; builds team dashboards connecting experiment outcomes to product metrics and release decisions

Leads the LLM platform team, defining architecture standards for inference, fine-tuning, and evaluation infrastructure. Coordinates cross-team prompt engineering governance and token budget allocation. Drives vendor evaluation for foundation models, balancing capability, cost, and compliance requirements across the engineering organization.

Defines evaluation standards for the LLM team. Establishes model evaluation guidelines, regression testing, benchmark management. Coordinates human evaluation processes and quality assurance.

Defines fine-tuning strategy for the LLM team. Establishes best practices for data preparation, training configuration, evaluation. Coordinates fine-tuning experiments and model selection process.

MLflow Expert

Sets MLflow strategy for LLM teams, standardizing tracking across fine-tuning, RLHF, and prompt engineering workflows. Establishes model registry policies for LLM versioning with adapter compatibility matrices. Reviews tracking practices for data lineage compliance and model card generation.

Defines model monitoring strategy for LLM products at team level. Establishes evaluation frameworks combining automated metrics, human review workflows, and safety monitoring. Sets standards for cost tracking, latency budgets, and quality gates across prompt versions. Conducts reviews of monitoring coverage for all LLM-powered features.

Model Serving Expert

Defines LLM serving strategy for the organization. Establishes inference cost management policies, serving SLA targets, and GPU infrastructure governance. Evaluates serving frameworks (vLLM, TGI, TensorRT-LLM). Conducts architecture reviews for LLM infrastructure. Drives adoption of cost-efficient LLM serving patterns.

PyTorch Expert

Defines PyTorch-based LLM training standards: PEFT strategy selection, distributed training configs, evaluation benchmarks. Evaluates tools (vLLM, TensorRT-LLM) for production inference. Reviews fine-tuning architectures. Establishes best practices for reproducible LLM experiments.

Defines RAG Architecture strategy at team/product level. Establishes standards and best practices. Conducts reviews.

Defines LLM platform strategy: model selection criteria, fine-tuning pipelines, and serving infrastructure standards. Establishes evaluation benchmarks, safety guardrails, and cost optimization practices. Reviews architectural decisions for RAG systems, agent frameworks, and multi-model orchestration.

Defines vector database strategy for the LLM platform. Establishes guidelines for vector DB selection, schema design, indexing strategy, monitoring. Coordinates migration and upgrades.

AI-Assisted Development · 4

Defines AI agent development standards, safety frameworks, and infrastructure architecture for the organization's LLM engineering practice. Establishes agent evaluation methodologies, production deployment requirements, and cost optimization strategies across agent-powered systems. Drives architectural decisions on agent infrastructure including runtime environments, tool marketplace governance, and cross-team agent capability reuse patterns.

Defines commercial LLM API usage strategy for the organization. Establishes guidelines for model selection, rate limit management, quality and cost monitoring at platform level.

Defines Model Context Protocol strategy at team/product level. Establishes standards and best practices. Conducts reviews.

Defines prompt engineering culture for the LLM team. Establishes prompt design patterns, review processes, prompt development tooling. Coordinates prompt optimization as a discipline.

Additional skills

Not assessed by the team, but part of the self-assessment and the development plan.

Algorithms & ComplexityAPI DocumentationCode Quality & RefactoringCode ReviewCursor IDEData StructuresDatabase IndexingDeep LearningDesign PatternsDistributed TrainingDockerEmbeddings & Vector DBGit AdvancedGitHub Actions / GitLab CIGitHub CopilotGPU ProgrammingGraphQL DesignIntegration TestingKubernetes OrchestrationLLM AlignmentLLM DeploymentLLM Prompt EngineeringLLM SafetyLLM ScalingML DeploymentML Experiment TrackingML Model EvaluationML PipelinesMultithreadingNatural Language ProcessingNetwork FundamentalsNeural Network ArchitecturesOpenTelemetryOWASP & Application SecurityPrometheus & GrafanaPython ProgrammingRAGRLHF TechniquesSecure Coding PracticesStructured LoggingSystem Design FundamentalsTensorFlow / PyTorchTerraformTokenizationTransfer LearningTransformer ArchitectureType Safety & Type SystemsUnit TestingvLLM Inference

What changes at Principal

0 skills get a higher expectation or become core when moving from Lead to Principal. The biggest jumps first.

See the Principal page →
Run this with your whole team
Self-assessment plus manager and peer reviews against the same matrix, gap analysis and next-level readiness for every engineer. Team Pro is free for 14 days; individual tools stay free forever.
Start a team trial (14 days free) Send to my manager

} in the open competency matrix: 80 skills across 5 levels. The matrix is free for individuals and stays free.