Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Anna Price
Site Reliability Engineer with experience in SLOs, error budgets, incident response, and recoverable infrastructure. Practical work includes Kubernetes, Terraform, Prometheus, SLO.
Experience
NVIDIA
Santa Clara · United States
Site Reliability Engineer
Mar 2022 - now
- Built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
- Rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
- Built a Go probe for the customer-facing request path and a Python recovery check for restored records. Used both in a failover exercise to detect healthy processes serving incomplete data; updated the runbook and verified the correction in a repeat drill.
- Reconstructed an incident timeline from traces, deploy events and SLO alerts; converted the triggering dependency failure into a recovery drill with measured restoration time.
- Owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
- Built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.
Datadog
New York · United States
Site Reliability Engineer
Jan 2019 - Feb 2022
- Led an incident response review using request traces and a timestamped timeline, assigning owners to retry limits and queue backpressure fixes. Closed the failure path with a replay test rather than only updating the runbook.
- Reviewed a proposed change against the existing interface contract, reproduced a reviewer’s edge case, and added the missing test or example before merging the change.
- Investigated a failure reported by a downstream team using the exact version and input they supplied. Reduced it to a reproducible test, delivered the fix with compatibility notes, and confirmed the original case passed before closing the report.
Selected project
Site Reliability Engineer — independent case study
Project owner
Feb 2024 - Jun 2024
- Replaced manual production changes with reviewed Terraform plans, drift checks, and staged rollout gates
- Built an isolated fixture with normal, malformed and interrupted inputs; recorded the expected response before running the implementation so failures were distinguishable from a changed assumption.
- Packaged the reproduction steps, configuration and regression tests with the project; checked the instructions from a clean environment and documented the unsupported cases instead of claiming production readiness.
- Owned the independent regression fixture using isolated inputs, resolved a failing edge case and released a reproducible test package with expected outputs.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science
Sep 2013 - Jun 2017
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
Kubernetes · Terraform · Prometheus · SLO · Incident response
Certifications
AWS Certified Solutions Architect – Associate
Amazon Web Services
Jun 2024
Publications
- Published an independent technical note using the project’s before-and-after traces, explaining the failed assumption, regression coverage and cases the implementation does not support.

