← All job descriptionsEngineering

Site Reliability Engineer Job Description

A complete site reliability engineer job description template, plus the ATS keywords teams screen for and how much this role really overlaps with DevOps.

SeniorSoftwareHybrid

Site Reliability Engineering started as a specific discipline at Google: run production the way you'd run a software project, with explicit reliability targets (SLOs), an error budget that gives engineering teams permission to ship risk up to a defined limit, and a staffing model where the team writing the reliability tooling is the same team that gets paged when it breaks. That is a real, coherent set of practices, and where it exists, it changes what the job actually is.

Outside companies with that culture built out, "SRE" quite often just means "DevOps, renamed" — the title reads better on a job posting in a market where "reliability engineer" sounds more senior than "ops engineer," without a corresponding change in what error budgets or SLOs actually govern day to day. Neither version is a lesser job. They are different jobs that happen to share a title, and the gap between them is usually invisible until you're three weeks in.

The clearest signal in a posting is the ratio between building and operating. A job that talks about writing the tooling, contributing to the services you're reliability engineering for, and owning reliability as a design input from the start is doing SRE in the fuller sense, and will usually expect real software engineering skill in the interview loop. A job that is mostly dashboards, runbooks and an on-call rotation for systems built by someone else is a legitimate and necessary role, but it is closer to production operations than to the discipline SRE originally named — and the coding bar in the interview will usually tell you which one you're in before the offer does.

Sample job description — not a live opening

Verrantine · Portland, OR

Full-time · Hybrid

$145,000 – $190,000

About the role

Verrantine is hiring a Senior Site Reliability Engineer to join the Reliability team, which owns uptime and performance for the production systems the rest of engineering ships into. You'll split your time between writing software — reliability tooling, automation, and contributions to the services you support — and the operational work of keeping a system with real customer traffic healthy under load.

We run a genuine SLO practice: each service has an agreed error budget, and when a team spends it, shipping speed for that service slows down until reliability recovers. You'll be one of the people who makes that tradeoff real rather than aspirational, which means this role has teeth other "reliability" titles sometimes don't.

What you'll do

  • Define and maintain SLOs and error budgets for the services our Reliability team supports, in partnership with the teams that own them.
  • Build and maintain the observability stack — metrics, logging, tracing — that makes production behavior legible to everyone, not just to you.
  • Write software: automation, internal tooling, and direct contributions to services when reliability work means changing the code, not just the dashboard.
  • Lead incident response for major production issues, and own the postmortem through to shipped follow-up work.
  • Participate in an on-call rotation, roughly one week in six, with a clear escalation path so it doesn't fall on one person by default.
  • Run capacity planning ahead of predictable load events, not as a reaction to one that already happened.
  • Push back on launches that skip reliability review, and make that pushback about the data, not the title.
  • Mentor engineers on other teams in operating their own services well, so reliability work doesn't bottleneck on your team.

What we're looking for

  • Five or more years in an SRE, DevOps, or production/platform engineering role, including direct on-call ownership.
  • Strong software engineering ability in at least one of Python or Go — this role writes production code, not only scripts.
  • Hands-on experience with Kubernetes in production, including debugging it under real incident pressure.
  • Practical experience with an observability stack (Prometheus, Grafana, Datadog or similar) beyond just viewing dashboards someone else built.
  • Experience leading incident response, including writing a postmortem people other than you find useful.
  • Comfort with the mechanics of reliability targets — what an SLO and an error budget actually constrain.

Nice to have

  • Experience defining SLOs from scratch at a company that didn't have them before you arrived.
  • Familiarity with chaos engineering practices — deliberately introducing failure to find weaknesses before an incident does.
  • Distributed systems experience: consensus, replication, or the failure modes of a system with more than one datacenter.
  • A track record of reducing toil — measurably cutting the manual, repetitive work a system used to require.
  • Experience with cost-aware infrastructure decisions, not just performance-aware ones.

Benefits

  • Medical, dental and vision coverage, with employee premiums covered in full.
  • 401(k) with a 4% company match, vested immediately.
  • Hybrid schedule: three days a week in the Portland office.
  • $2,000 annual learning budget, usable on conferences, courses or books.
  • Twenty days of paid time off plus company holidays.

Salary range

As posted for this sample role. Real pay varies by employer, location and experience.

$145,000$190,000/ yr

What gets you noticed

ATS keywords for this role

The applicant tracking system (ATS) — the recruiting software a hiring team searches and filters applicants with — will screen for these. Weight shows how central each one is to this specific posting.

Required and central (4)

Kubernetesincident responseon-call rotationSLOs and error budgets

Important (9)

PythonGoTerraformobservabilityPrometheuspostmortem analysisLinux systems administrationAWSdistributed systems

Mentioned in passing (5)

Grafanacapacity planningchaos engineeringCI/CDcross-functional communication

Frequently asked questions

Is Site Reliability Engineer just DevOps Engineer with a fancier title?

Sometimes, yes — and it's worth finding out which before you accept an offer. Where a company has built a genuine SRE practice (formal SLOs, error budgets that actually constrain shipping speed, a real software-engineering bar in the interview), the job is meaningfully different from DevOps. Where it hasn't, "SRE" is often just the more attractive title for the same operations work. The interview loop is usually the tell: a heavy coding bar suggests the former.

How much of the job is actually writing code?

It varies more than the title implies. At companies with a mature SRE practice, expect real software engineering — building tooling, contributing to the services you support, sometimes a coding bar close to a backend engineering interview. At companies using the title more loosely, expect scripting and configuration rather than software development. Ask directly in the interview what percentage of a typical week is spent writing code versus operating existing systems.

What's an error budget, and do I actually need to understand one?

An error budget is the amount of unreliability a service is allowed before its team has to stop shipping features and fix reliability instead — a way of making "how reliable is reliable enough" a number instead of an argument. You'll need to understand the concept for almost any SRE interview, but whether you'll use it for real depends on whether the company has actually operationalized it, which is worth asking about directly.

Do I need prior on-call experience to break into SRE?

It helps, but it's not the only path in. Backend, platform, and infrastructure engineers move into SRE regularly, usually by demonstrating they understand production systems and can reason about failure, even if their on-call experience so far has been informal. What matters more than a title on your resume is being able to talk concretely about an incident you helped resolve.

Tailor it to a real posting

This was a sample. Your resume should be tailored to the real thing.

Rezi Ninja reads an actual job posting and rewrites your resume to match it, with a Ninja Score so you know it lands before you hit send. Free to start, with the AI usage included.

We use privacy-conscious analytics to see how the site is used — no ads, no selling data. Read our Privacy Policy.