Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Stephanie Woods
Site Reliability Engineer Intern with supervised experience in SLOs, error budgets, incident response, and recoverable infrastructure. Practical work includes Kubernetes, Terraform, Prometheus, SLO.
Experience
NVIDIA
Santa Clara · United States
Site Reliability Engineer Intern
Jun 2026 - Aug 2026
- Under supervision, reconstructed an incident timeline from traces, deploy events and SLO alerts; converted the triggering dependency failure into a recovery drill with measured restoration time.
- The team defined availability and latency SLOs for a payment API and implemented multi-window burn-rate alerts in Prometheus; I supported research preparation, validation, and documentation under mentor review. The team replaced noisy threshold alerts with pages tied to customer-visible failure; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
- The team rehearsed a regional failover for Kubernetes services, testing database recovery points and rollback criteria before shifting traffic; I supported research preparation, validation, and documentation under mentor review. The team exposed a dependency that was missing from the recovery runbook; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
- Under supervision, owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
- Under supervision, built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.
Selected project
Site Reliability Engineer — independent case study
Intern project team member
Sep 2025 - May 2026
- In a mentor-reviewed simulation, built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
- As a second supervised exercise, rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
- Under supervision, the team replaced manual production changes with reviewed Terraform plans, drift checks, and staged rollout gates; I supported research preparation, validation, and documentation under mentor review
- Built an isolated fixture with normal, malformed and interrupted inputs; recorded the expected response before running the implementation so failures were distinguishable from a changed assumption.
- Packaged the reproduction steps, configuration and regression tests with the project; checked the instructions from a clean environment and documented the unsupported cases instead of claiming production readiness.
- Owned the independent regression fixture using isolated inputs, resolved a failing edge case and released a reproducible test package with expected outputs.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science — in progress
Sep 2023 - Jun 2027
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
Kubernetes · Terraform · Prometheus · SLO · Incident response
Certifications
AWS Certified Solutions Architect – Associate
Amazon Web Services
Dec 2025
Publications
- Published an independent technical note using the project’s before-and-after traces, explaining the failed assumption, regression coverage and cases the implementation does not support.

