Round two of the screen plays out inside this section, the final gate before any interview is on
the table. A recruiter genuinely slows the pace here, and even at that, your current role still
drives roughly 95% of the result.
That tracks: nothing demonstrates what you can run in production today better than the role you
are sitting in right now. To earn a "yes", the section must hit every part of the
Data Engineer role profile, one bullet per area listed under Domain Expertise. And
each bullet has to land on something you genuinely held in production, not on a ticket that
crossed your queue.
1
Pipeline Development & ETL/ELT
You move data from messy sources into the warehouse, reliably. Half-finished loads quietly poison
everything downstream, so hiring managers want pipelines built to rerun safely, not scripts that work
when nothing goes wrong. Talk about how you used incremental loads and idempotent design, in Spark and
dbt, to push more volume per day while cutting pipeline runtime.
Techniques
Batch ETL / ELT
Incremental loads
Idempotent design
Schema evolution
Tools
Spark, dbt
pandas
Apache Beam
Metrics
Pipelines in production
Volume per day
Runtime cut
2
Data Modeling & Warehousing
You shape raw data into models people can query. A model with a reasoned grain reads as engineering to a
hiring manager; "built some tables" does not. Show them how you used Kimball-style dimensions
and slowly changing dimensions, on Snowflake or BigQuery, to cut query latency and storage at the same
time.
Techniques
Dimensional / Kimball
Star & snowflake schema
SCDs
Iceberg / Delta tables
Tools
Snowflake, BigQuery
Databricks
Redshift
Metrics
Tables modeled
Query latency cut
Storage saved
3
Orchestration & Scheduling
You run pipelines on schedule, with retries and backfills. A pipeline that dies overnight is where
orchestration gets hard, so hiring managers want proof your DAGs recover on their own, not that you
"used Airflow". Point out how you used solid DAG design and automated retries, in Airflow or
Dagster, to bring the failure rate down and shrink backfill time.
Techniques
DAG design
Retries & backfills
Sensors & triggers
Cross-DAG dependencies
Tools
Airflow
Dagster
Prefect
Metrics
DAGs in production
Failure rate down
Backfill time cut
4
Streaming & Real-Time Data
You process streaming data while it's still fresh, without double-counting. Events per second and
end-to-end lag are numbers anyone can check, so real streaming figures carry more weight than
"worked with Kafka". Mention how you used CDC through Debezium and windowed aggregations in
Flink to scale throughput while holding end-to-end lag low.
Techniques
CDC
Event ingestion
Windowed aggregations
Exactly-once semantics
Tools
Kafka, Kinesis
Flink, Spark Streaming
Debezium
Metrics
Events per second
End-to-end lag
Throughput scaled
5
Performance & Cost Optimization
You cut warehouse runtime and cost together. Hiring managers care because a slow, wasteful warehouse
burns money on every run, and the cost is easy to measure. Walk them through how you used partition
tuning and query-plan analysis, in the Spark UI and Snowflake Query Profile, to cut real dollars and
hours off the monthly run.
Techniques
Partition & cluster tuning
Cost attribution
Query plan analysis
Caching & materialization
Tools
Spark UI
Snowflake Query Profile
dbt threads
Metrics
Cost cut ($)
Runtime cut (h)
$/TB processed
6
Data Quality & Observability
You catch bad data before your stakeholders do. Two things ride on it for a hiring manager: tests that
catch breakage upstream, and freshness SLAs that actually hold. Lay out how you used schema tests and
freshness SLAs, with Great Expectations and Monte Carlo, to catch incidents upstream and keep the SLA
hit rate high.
Techniques
Schema tests
Row-count assertions
Freshness SLAs
Anomaly detection
Tools
Great Expectations
dbt tests
Monte Carlo, Soda
Metrics
SLA hits %
Incidents caught upstream
MTTR
7
Cloud Infrastructure & DevOps
You make the data platform repeatable with code, not clicks. Repeatable infrastructure keeps every
engineer out of the weeds, so owning it tells a hiring manager you think past your own pipelines. Spell
out how you used Terraform and CI/CD for data, containerized with Docker, to ship platform changes more
often and remove the manual steps.
Techniques
Infrastructure as code
CI/CD for data
Containerization
Secrets management
Tools
Terraform
AWS, GCP, Azure
Docker, GitHub Actions
Metrics
Deploy frequency
Environments managed
Manual steps removed
8
Cross-Functional Collaboration
You serve many teams without becoming their bottleneck. Companies lean on the engineers other teams can
build against, not the ones guarding a black box, so hiring managers look for it. Tell them how you used
schema contracts and honest incident response, tracked through Jira and reviews, to serve more teams
while shifting the on-call load down.
Techniques
Schema contracts
Incident response
Stakeholder reviews
Roadmap planning
Tools
Jira, Linear
Slack
Notion docs
Metrics
Teams served
Contracts signed
On-call load shifted