Skip to content

Site Reliability Engineer resume example

See how a complete Site Reliability Engineer resume organizes experience, projects, education, and skills.

NVIDIAPosting location: Santa Clara · United StatesIndependent fictional example, not company material.

Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.

Anna Price

Site Reliability Engineer · Santa Clara · United States

Site Reliability Engineer with experience in SLOs, error budgets, incident response, and recoverable infrastructure. Practical work includes Kubernetes, Terraform, Prometheus, SLO.

Experience

NVIDIA

Santa Clara · United States

Site Reliability Engineer

Mar 2022 - now

  • Built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
  • Rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
  • Built a Go probe for the customer-facing request path and a Python recovery check for restored records. Used both in a failover exercise to detect healthy processes serving incomplete data; updated the runbook and verified the correction in a repeat drill.
  • Reconstructed an incident timeline from traces, deploy events and SLO alerts; converted the triggering dependency failure into a recovery drill with measured restoration time.
  • Owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
  • Built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.

Datadog

New York · United States

Site Reliability Engineer

Jan 2019 - Feb 2022

  • Led an incident response review using request traces and a timestamped timeline, assigning owners to retry limits and queue backpressure fixes. Closed the failure path with a replay test rather than only updating the runbook.
  • Reviewed a proposed change against the existing interface contract, reproduced a reviewer’s edge case, and added the missing test or example before merging the change.
  • Investigated a failure reported by a downstream team using the exact version and input they supplied. Reduced it to a reproducible test, delivered the fix with compatibility notes, and confirmed the original case passed before closing the report.

Selected project

Site Reliability Engineer — independent case study

Project owner

Feb 2024 - Jun 2024

  • Replaced manual production changes with reviewed Terraform plans, drift checks, and staged rollout gates
  • Built an isolated fixture with normal, malformed and interrupted inputs; recorded the expected response before running the implementation so failures were distinguishable from a changed assumption.
  • Packaged the reproduction steps, configuration and regression tests with the project; checked the instructions from a clean environment and documented the unsupported cases instead of claiming production readiness.
  • Owned the independent regression fixture using isolated inputs, resolved a failing edge case and released a reproducible test package with expected outputs.

Education

University of Washington

Seattle, Washington · United States

B.S. Computer Science

Sep 2013 - Jun 2017

Relevant coursework: Algorithms, operating systems, databases, computer networks

Skills

Role expertise

Kubernetes · Terraform · Prometheus · SLO · Incident response

Certifications

AWS Certified Solutions Architect – Associate

Amazon Web Services

Jun 2024

Publications

  • Published an independent technical note using the project’s before-and-after traces, explaining the failed assumption, regression coverage and cases the implementation does not support.

How this Site Reliability Engineer resume addresses the posting

See which posting requirements are supported by specific work.

NVIDIA

Senior Site Reliability Engineer, AIOPs

Santa Clara · United StatesSenior

View original job posting

This resume is a fictional example. Check the original posting for current details.

  • In the posting
      • Kubernetes
      • observability
      ATS keywords with experience

    “Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.”

    “Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.”

    In this resume
    Built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
    Site Reliability Engineer · NVIDIA
  • In the posting
      • Terraform
      ATS keywords with experience

    “Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.”

    In this resume
    Rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
    Site Reliability Engineer · NVIDIA
  • In the posting
      • on-call
      • Grafana
      • Prometheus
      ATS keywords with experience

    “Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.”

    “Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.”

    In this resume
    Owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
    Site Reliability Engineer · NVIDIA
  • In the posting
      • distributed systems
      • monitoring
      No relevant experience found

    “BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.”

    “Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.”

    In this resume

    No supporting experience found. Do not add this keyword unless your own work supports it.

View 4 more matched requirements
  • In the posting
      • Linux
      • CI/CD
      • debugging
      ATS keywords with experience

    “Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.”

    “Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.”

    “Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.”

    In this resume
    Built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.
    Site Reliability Engineer · NVIDIA
  • In the posting
      • Python
      ATS keywords with experience

    “Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.”

    In this resume
    Built a Go probe for the customer-facing request path and a Python recovery check for restored records. Used both in a failover exercise to detect healthy processes serving incomplete data; updated the runbook and verified the correction in a repeat drill.
    Site Reliability Engineer · NVIDIA
References2 sourcesReviewed

Start with this example. Finish with your experience.

Open it in the editor and replace every role, project, and result with your real experience.