Skip to content

Internship Site Reliability Engineer resume example

See how a complete Internship Site Reliability Engineer resume organizes experience, projects, education, and skills.

NVIDIAPosting location: Santa Clara · United StatesIndependent fictional example, not company material.

Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.

Stephanie Woods

Site Reliability Engineer Intern · Santa Clara · United States

Site Reliability Engineer Intern with supervised experience in SLOs, error budgets, incident response, and recoverable infrastructure. Practical work includes Kubernetes, Terraform, Prometheus, SLO.

Experience

NVIDIA

Santa Clara · United States

Site Reliability Engineer Intern

Jun 2026 - Aug 2026

  • Under supervision, reconstructed an incident timeline from traces, deploy events and SLO alerts; converted the triggering dependency failure into a recovery drill with measured restoration time.
  • The team defined availability and latency SLOs for a payment API and implemented multi-window burn-rate alerts in Prometheus; I supported research preparation, validation, and documentation under mentor review. The team replaced noisy threshold alerts with pages tied to customer-visible failure; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
  • The team rehearsed a regional failover for Kubernetes services, testing database recovery points and rollback criteria before shifting traffic; I supported research preparation, validation, and documentation under mentor review. The team exposed a dependency that was missing from the recovery runbook; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
  • Under supervision, owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
  • Under supervision, built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.

Selected project

Site Reliability Engineer — independent case study

Intern project team member

Sep 2025 - May 2026

  • In a mentor-reviewed simulation, built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
  • As a second supervised exercise, rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
  • Under supervision, the team replaced manual production changes with reviewed Terraform plans, drift checks, and staged rollout gates; I supported research preparation, validation, and documentation under mentor review
  • Built an isolated fixture with normal, malformed and interrupted inputs; recorded the expected response before running the implementation so failures were distinguishable from a changed assumption.
  • Packaged the reproduction steps, configuration and regression tests with the project; checked the instructions from a clean environment and documented the unsupported cases instead of claiming production readiness.
  • Owned the independent regression fixture using isolated inputs, resolved a failing edge case and released a reproducible test package with expected outputs.

Education

University of Washington

Seattle, Washington · United States

B.S. Computer Science — in progress

Sep 2023 - Jun 2027

Relevant coursework: Algorithms, operating systems, databases, computer networks

Skills

Role expertise

Kubernetes · Terraform · Prometheus · SLO · Incident response

Certifications

AWS Certified Solutions Architect – Associate

Amazon Web Services

Dec 2025

Publications

  • Published an independent technical note using the project’s before-and-after traces, explaining the failed assumption, regression coverage and cases the implementation does not support.

How this Site Reliability Engineer resume addresses the posting

See which posting requirements are supported by specific work.

NVIDIA

Senior Site Reliability Engineer, AIOPs

Santa Clara · United StatesSenior

View original job posting

This resume is a fictional example. Check the original posting for current details.

  • In the posting
      • Kubernetes
      ATS keywords with experience

    “Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.”

    In this resume
    The team rehearsed a regional failover for Kubernetes services, testing database recovery points and rollback criteria before shifting traffic; I supported research preparation, validation, and documentation under mentor review. The team exposed a dependency that was missing from the recovery runbook; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
    Site Reliability Engineer Intern · NVIDIA
  • In the posting
      • Terraform
      • Python
      • CI/CD
      • debugging
      ATS keywords with experience

    “Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.”

    “Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.”

    “Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.”

    In this resume
    Under supervision, built Python debugging checks using Linux logs and Terraform state, reduced incident triage from 32 to 14 minutes and released tested CI/CD recovery instructions.
    Site Reliability Engineer Intern · NVIDIA
  • In the posting
      • observability
      ATS keywords with experience

    “Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.”

    In this resume
    In a mentor-reviewed simulation, built observability for Kubernetes services with Grafana burn-rate dashboards and on-call runbooks; tied each page to an SLO and a tested mitigation rather than a raw CPU threshold.
    Site Reliability Engineer — independent case study
  • In the posting
      • distributed systems
      • monitoring
      No relevant experience found

    “BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.”

    “Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.”

    In this resume

    No supporting experience found. Do not add this keyword unless your own work supports it.

View 4 more matched requirements
  • In the posting
      • on-call
      • Grafana
      ATS keywords with experience

    “Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.”

    “Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.”

    In this resume
    Under supervision, owned an on-call recovery drill using Prometheus alerts and Grafana traces, resolved a Kubernetes dependency failure and completed rollback within the agreed SLO window.
    Site Reliability Engineer Intern · NVIDIA
  • In the posting
      • Linux
      ATS keywords with experience

    “Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.”

    In this resume
    As a second supervised exercise, rehearsed Linux node replacement from reviewed Terraform plans, checking workload drain, storage mounts, and recovery time; recorded a mount dependency that had been missing from the recovery procedure.
    Site Reliability Engineer — independent case study
  • In the posting
      • Prometheus
      ATS keywords with experience

    “Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.”

    In this resume
    The team defined availability and latency SLOs for a payment API and implemented multi-window burn-rate alerts in Prometheus; I supported research preparation, validation, and documentation under mentor review. The team replaced noisy threshold alerts with pages tied to customer-visible failure; confirmed with my mentor that my contribution was limited to the assigned support and validation work.
    Site Reliability Engineer Intern · NVIDIA
References2 sourcesReviewed

Start with this example. Finish with your experience.

Open it in the editor and replace every role, project, and result with your real experience.