This is where the second pass actually plays out, the last gate before an interview hits your
inbox. The recruiter slows down right here, and even then your current role still drives
around 95% of the decision.
Makes sense: nothing tells a hiring team what you can run in production right now the way your
current job does. To clear that "yes", this section has to walk the full
SRE role profile, one bullet per slot you listed in Domain
Expertise above. Every bullet has to come off something you actually held in production,
not a Jira card that wandered past your queue.
1
SLO & Error Budget Engineering
You turn "reliable" into SLOs with real teeth. Hiring managers read a working error-budget
policy behind it, not "we aim for five nines", so this is where an SRE stands apart. Talk
about how you used SLI selection and burn-rate alerting, in Prometheus with Nobl9, to get services onto
SLOs and defend the error budget.
Techniques
SLI selection
Burn-rate alerting
Error-budget policy
Tier classification
Tools
Prometheus
Sloth, Nobl9
Grafana SLO panels
Metrics
SLO hit rate
Services on SLO
Error budget defended
2
Observability & Tracing
You find what's broken in minutes, not hours. Hiring managers look here to see whether you can find
the cause fast, or whether an outage turns into a guessing game. Show them how you used distributed
tracing and RED metrics, in OpenTelemetry and Grafana Tempo, to widen tracing coverage and cut
time-to-diagnose.
Techniques
RED & USE metrics
Distributed tracing
Structured logging
Cardinality control
Tools
Prometheus, Grafana
OpenTelemetry, Tempo
Datadog, Honeycomb
Metrics
Tracing coverage
Alert noise reduced
Time-to-diagnose cut
3
Incident Response & Command
You lead the response when production goes down. A calm, well-run incident is the difference between a
short blip and an all-day outage, so hiring managers weigh it heavily. Point out how you ran incident
command with a tight severity model, coordinated through incident.io, to cut MTTR and your P0 count.
Techniques
Incident command
Severity model
Comms cadence
On-call rotation
Tools
PagerDuty, Opsgenie
incident.io, FireHydrant
Statuspage
Metrics
MTTR cut
P0 incidents reduced
Time-to-detect
4
Postmortems & Reliability Improvement
You drive an incident to fixes that stick. Repeat incidents come from postmortems that die in a doc, so
hiring managers want to see the fixes you drove to done. Mention how you used blameless postmortems and
tracked actions, in incident.io retros and Jira, to bring the repeat-incident rate down.
Techniques
Blameless postmortems
Action tracking
Reliability bets
Trend analysis
Tools
Notion, Confluence
Jira, Linear
incident.io retros
Metrics
Actions closed
Repeat-incident rate down
SLO regressions caught
5
Capacity Planning & Performance
You size the system to hold when traffic spikes. Peak load and cost-per-request are numbers a hiring
manager can check, so a real capacity result beats "handled scale". Walk them through how you
used load and soak testing with capacity modeling, in k6 and pprof, to carry peak load while cutting
cost per request.
Techniques
Load & soak testing
Capacity modeling
Headroom planning
Performance profiling
Tools
k6, Locust
JMeter
pprof, perf
Metrics
Peak load handled
Latency at peak
Cost per RPS cut
6
Toil Reduction & Automation
You automate the repetitive work off on-call. Two things ride on it for a hiring manager: hours handed
back to the team, and a pager that only fires when it matters. Lay out how you used self-healing
automation and alert hygiene, scripted in Python and Ansible, to cut toil hours and pages per shift.
Techniques
Toil measurement
Self-healing automation
Runbook codification
Alert hygiene
Tools
Python, Go
Ansible, Terraform
Rundeck, StackStorm
Metrics
Toil hours cut
Pages per shift down
Self-heal rate
7
Chaos Engineering & Resilience Testing
You break the system on purpose before it breaks itself. Resilience only counts if you've tested
it, so hiring managers want proof you ran real failure injection, not a DR doc nobody rehearsed. Spell
out how you used failure injection and game days, with Chaos Mesh and AWS FIS, to close failure modes
and hold DR RTO.
Techniques
Failure injection
Game days
DR drills
Hypothesis testing
Tools
Chaos Mesh, LitmusChaos
Gremlin
AWS FIS
Metrics
Failure modes closed
Game days run
DR RTO held
8
Tooling & Workflow
You make reliability everyone's job, not just yours. Companies keep the SREs who raise the whole
org's reliability bar, not the ones who firefight alone, so hiring managers look for it. Tell them
how you used SLO-as-code and production-readiness reviews, in Git and Backstage, to standardize
reliability and cut on-call ramp.
Techniques
SLO as code
Production-readiness reviews
Runbook libraries
On-call shadowing
Tools
Git, GitHub
Python, Go, Bash
Backstage
Metrics
SLOs as code
Runbooks maintained
On-call ramp cut