The second round of the screen lives inside this section, the final gate ahead of any
interview going on offer. A recruiter takes more time at this point, and even so the chair
you sit in now carries roughly 95% of the outcome.
That fits: nothing demonstrates your shipped production work as plainly as the seat you're sitting in this quarter. To pull a yes, the block has to land each entry from the
ML Engineer role profile, one bullet per area named under Domain
Expertise. And each bullet has to come off something you actually held in production,
never a ticket that brushed past your queue.
1
Model Training & Development
You train serious models at scale, on budget. Training cost and offline lift are numbers a hiring
manager can sanity-check, so real figures beat "trained deep learning models". Talk about how
you used distributed training and disciplined tuning, in PyTorch with Ray, to lift the offline metric
while cutting the training bill.
Techniques
Deep learning & transformers
Gradient boosting
Distributed training
Hyperparameter tuning
Tools
PyTorch, TensorFlow
Hugging Face Transformers
Ray, Horovod
Metrics
Models trained at scale
Offline metric lift
Training cost cut
2
Model Serving & Inference
You get models answering real requests, fast. A model nobody can call is shelfware, so hiring managers
want serving numbers, not another checkpoint in a bucket. Lay out how you used request batching and
quantization, served through Triton with ONNX, to hold p99 latency while cutting inference cost per
request.
Techniques
Online vs batch inference
Request batching
Quantization & distillation
A/B model routing
Tools
Triton, TorchServe
KServe, BentoML
ONNX, TensorRT
Metrics
Latency (p95 / p99)
Requests per second
Inference cost per request
3
Feature Engineering & Pipelines
You match features between training and serving. Train-serve skew is the classic silent model killer, so
hiring managers want proof you engineered it away, not that you got lucky. Show them how you used a
feature store with backfills and replays, on Feast with Spark and Kafka, to keep features fresh and cut
skew incidents.
Techniques
Train-serve consistency
Online vs offline features
Backfills & replays
Embedding pipelines
Tools
Feast, Tecton
Spark, Beam
Kafka, Flink
Metrics
Features in production
Feature freshness
Skew incidents cut
4
MLOps & Deployment
You ship models like software: versioned, canaried, reversible. Hiring managers look here to see whether
you ship models like software, or whether every launch is a hand-carried event. Point out how you used
canary deploys and CI/CD for models, through MLflow and ArgoCD, to deploy weekly and cut iteration time.
Techniques
Model registry & versioning
Canary & shadow deploys
CI/CD for models
Automated retraining
Tools
MLflow / W&B
SageMaker / Vertex AI
GitHub Actions, ArgoCD
Metrics
Deploys per week
Iteration time cut
Rollback MTTR
5
Monitoring, Drift & Reliability
You know your model is still right after launch day. Two things ride on it for a hiring manager: drift
caught before users feel it, and latency budgets that hold under load. Walk them through how you used
drift detection and shadow scoring, in Evidently with Grafana, to hit your SLOs and catch drift
incidents early.
Techniques
Data & concept drift
Latency & error budgets
Shadow scoring
Alerting & runbooks
Tools
Prometheus, Grafana
Evidently, WhyLabs
Datadog, Sentry
Metrics
SLO hit rate
Drift incidents caught
On-call MTTR
6
ML Infrastructure & Compute
You feed the GPUs without burning the budget. Compute efficiency pays back across every team's
runs, so owning it tells a hiring manager you treat GPUs as money, not magic. Mention how you used GPU
scheduling and mixed-precision training, on Kubernetes with Ray, to push utilization up and cut
time-to-train.
Techniques
GPU scheduling
Mixed-precision training
Multi-node distributed
Cost attribution
Tools
Kubernetes, Kubeflow
Ray, SLURM
Terraform, Pulumi
Metrics
GPU utilization
Training $ saved
Time-to-train cut
7
Cross-Functional Collaboration
You move research models into production with the team. Companies keep the ML engineers who get
researchers' work shipped, not the ones who gatekeep it, so hiring managers watch for it. Spell out
how you used launch reviews and clean research handoffs, agreed in writing, to ship models jointly and
cut the handoff time.
Techniques
Research-to-prod handoffs
Model launch reviews
SLO negotiation
Office hours
Tools
Notion, Confluence
Slack, Linear
Jira, GitHub
Metrics
Models shipped jointly
Handoff time cut
Squads supported
8
Tooling & Workflow
You make every model reproducible by anyone on the team. Hiring managers read tests and experiment
tracking as engineering maturity, not overhead, and it shows fast in an interview. Tell them how you
used ML unit tests and tracked experiments, with pytest and Docker, to keep pipelines reproducible and
ramp new teammates quickly.
Techniques
Reproducible environments
Unit & integration tests for ML
Code review for model PRs
Experiment tracking
Tools
Git, GitHub
Docker, Poetry, uv
pytest, Great Expectations
Metrics
Repos maintained
Pipeline reproducibility rate
Onboarding ramp time cut