Site Reliability Engineer CV Example
Professional site reliability engineer CV example with recruiter-tested sections you can customise in minutes.
Example only
All personal details and career history in this example are fictional. Customise this CV with your own real information before applying.
Resume preview

Skills
Certifications
Google Professional Cloud DevOps Engineer
Google
2023-06-01
Linux Foundation Site Reliability Engineering Certificate
Linux
2022-06-01
Languages
English (Native)
Grace Murphy
Site Reliability Engineer
8 years of experience supporting highly available cloud services and improving operational reliability through automation.
Experience
Senior Site Reliability Engineer
Uptime Foundry, Full-time
2020-03-01 – Present
Seattle, WA, US
Reliability engineering and incident management — SLOs/error budgets, auto-remediation playbooks, and incident command.
- Led reliability engineering and incident management for critical systems
- Defined service-level objectives and error budgets for 25+ microservices, improving reliability prioritization during planning cycles.
- Built auto-remediation playbooks triggered by Prometheus alerts to resolve common saturation incidents without pager escalation.
- Reduced p1 incident duration by introducing structured incident command and timeline capture tooling.
Site Reliability Engineer
Pulse Transit Cloud, Full-time
2018-02-01 – 2020-02-01
Toronto, ON, CA
High availability and observability — multi-region failover drills, Kubernetes autoscaling tuning, and RED/USE metrics.
- Maintained high availability and observability for customer-facing APIs
- Implemented multi-region failover drills and chaos scenarios to validate dependency timeouts and fallback logic.
- Tuned Kubernetes HPA and VPA settings based on workload profiles, reducing CPU throttling events during spikes.
- Instrumented RED and USE metrics with dashboard templates adopted by all product teams.
Junior Site Reliability Engineer
Granite SignalOps, Full-time
2016-01-01 – 2018-01-01
Austin, TX, US
On-call operations and reliability automation — synthetic checks, post-incident verification, and log-indexing improvements.
- Supported on-call operations and reliability automation tasks
- Created synthetic checks for critical customer journeys and integrated them with pager duty routing rules.
- Wrote post-incident follow-up scripts to verify remediation tasks and stale alert cleanup.
- Improved log indexing strategy in OpenSearch, lowering investigation time for cross-service incidents.
Education
Computer Engineering, Bachelor of Science
North Carolina State University
2014-09-01 – 2018-06-01
Bachelor's-level training in computer engineering at North Carolina State University. Completed coursework in Linux, Kubernetes, Prometheus, capstone projects, and collaborative assignments that prepared for professional roles.
CV writing guide
- Demonstrate mature incident response and learning culture practices.
- Show measurable reliability outcomes tied to SLOs and MTTR.
- Include examples of automation that removes repetitive operational toil.
- Google Professional Cloud DevOps Engineer
- Linux Foundation Site Reliability Engineering Certificate
- Certified Kubernetes Administrator (CKA)