Skip to content

Advertisement

DevOps Society

DevOps Engineer vs SRE vs Platform Engineer: What's the Difference?

DevOps engineer, SRE, and platform engineer compared: daily work, ownership boundaries, skills, salary bands, and which role fits your team and career next.

Share

Published your local timeupdated

DevOps Engineer vs SRE vs Platform Engineer: What's the Difference?

Open three job ads with three different titles and you'll usually find the same eight bullets: Kubernetes, Terraform, a CI system, an observability stack, on-call rotation, "partner with product teams". The tidy definitions everyone repeats — DevOps is a culture, SRE is reliability, platform engineering builds an internal product — are true in the abstract and useless when you're deciding which req to open or which offer to accept. Sorting out DevOps engineer vs SRE vs platform engineer needs a test that survives contact with real organisations, and definitions don't.

Two questions do. What is the role accountable for when things go wrong, and who is its customer? Everything below is built on those.

Where the three titles actually came from

DevOps was a correction, not a job. The vocabulary crystallised around 2009 — John Allspaw and Paul Hammond's Velocity talk on how Flickr got dev and ops to cooperate, and Patrick Debois's first devopsdays in Ghent that autumn. The argument was organisational: separating the people who write software from the people who run it creates an incentive conflict, and that conflict causes slow, dangerous releases. Nothing in the original material describes a person. "DevOps engineer" is what happened when recruiters needed a searchable word for the automation work the movement implied.

SRE predates the term DevOps. Ben Treynor Sloss stood up the first site reliability engineering team at Google in 2003, and his framing is deliberately provocative: hand operations to software engineers and watch what they refuse to do by hand. The discipline went public with the 2016 O'Reilly book and its 2018 workbook, and it arrived with something DevOps never had — a specification. SLOs, error budgets, a cap on toil, blameless postmortems. Google's workbook puts the relationship neatly: class SRE implements interface DevOps. DevOps says what should be true; SRE is one opinionated implementation with numbers attached.

Platform engineering is the most recent and the most reactive. "You build it, you run it" worked, and it also dumped an enormous surface area onto product teams: clusters, IAM trust policies, pipeline YAML, tracing config, image signing, on top of their actual domain. Team Topologies gave the field its vocabulary for that problem — cognitive load, the platform team as one of four fundamental team types, and X-as-a-Service as the interaction mode a platform should offer. Our guide to building an internal developer platform goes deeper than this comparison can.

Timeline of the three disciplines emerging, each labelled with the problem it was created to solve

DevOps engineer vs SRE vs platform engineer: the test that works

Skills overlap almost completely. All three write Terraform, all three know a cloud well, all three argue about Helm. Skills are worthless as a discriminator. What differs is what you're on the hook for, and who gets to be disappointed in you.

An SRE's customer is the reliability of a running service. When it's slow or down, that's the SRE's problem regardless of who wrote the code. A platform engineer's customer is other engineers: when a product team can't ship without filing a ticket, that's the platform engineer's problem regardless of whether anything is down. A DevOps engineer, in most postings, is accountable for infrastructure and delivery automation — pipelines, provisioning, environments — and the customer is the engineering organisation in general, which is a polite way of saying sometimes nobody in particular.

Role

Customer

Primary artifact

Success metric

On-call posture

SRE

The service's reliability, on behalf of its users

SLOs, error budget policies, postmortems, capacity models, targeted resilience fixes

SLO attainment and error budget burn; toil kept under control; repeat-incident rate

Primary or shared-primary for production services; the pager is the job

Platform engineer

Other engineers in the company

Golden paths, service templates, internal APIs and controllers, CLIs, paved-road modules, docs

Adoption of the paved road; lead time from repo to prod; number of tickets not filed

On-call for the platform itself (control plane, CI, registry) — not for tenant services

DevOps engineer

The engineering org broadly; often whoever asks loudest

IaC modules, CI/CD pipelines, cloud accounts and networking, environment automation

Deployment frequency and lead time; environments that exist and work; cost and toil reduction

Highly variable — sometimes infra-only, often the de facto catch-all pager

That table is the article in one screen. The interesting part is where it breaks down, which is roughly everywhere below a thousand engineers.

Primary artifacts, and what they tell you

Ask a candidate or a hiring manager what the role produces in a quarter. The answer is more diagnostic than any question about etcd.

An SRE's quarter produces documents and targeted engineering: SLOs a product manager actually agreed to, an error budget policy with teeth, postmortems whose action items landed, and code that removed a recurring failure mode. An SRE who ends the quarter with a pile of Terraform modules and no SLOs has been doing platform work under an SRE title.

A platform engineer's quarter produces something other engineers use voluntarily: a scaffolder template, a capability behind a self-service API, a migration that removed a manual step from thirty repos. The test is adoption. If teams route around it, it failed, and architectural elegance doesn't rescue that.

A DevOps engineer's quarter produces infrastructure and pipeline change: new accounts, new modules, a CI migration, a security gate. That last one drifts fast — wire SAST, dependency scanning and image signing into pipelines and you're doing the work in our walkthrough of shifting security into the delivery pipeline, whatever the title says. Cost drifts the same way; plenty of DevOps engineers now own tagging hygiene and rightsizing, which is managing cloud spend as an engineering concern under another name.

Success metrics, and the one SRE artifact nobody else produces

SRE is the only one of the three with a documented, portable measurement system, which is much of why the title travels well between companies.

An SLO is a target for a service level indicator over a window. The error budget is its inverse: the unreliability you have deliberately decided to tolerate. Product management sets the objective, monitoring measures reality, and the gap between them is the failure budget left for the period.

# slo/checkout-api.yaml — lives in the service repo, owned with the service
apiVersion: openslo/v1
kind: SLO
metadata:
  name: checkout-api-availability
spec:
  service: checkout-api
  indicator:
    spec:
      ratioMetric:
        good:  { query: 'sum(rate(http_requests_total{job="checkout-api",code!~"5.."}[1m]))' }
        total: { query: 'sum(rate(http_requests_total{job="checkout-api"}[1m]))' }
  objectives:
    - target: 0.999        # 0.1% of requests may fail over the window
  timeWindow:
    - duration: 30d
      isRolling: true      # rolling: a bad Tuesday keeps costing you for 30 days

That file alone changes nothing. The artifact that changes behaviour is the policy attached to it, which is the part teams skip:

# error-budget-policy.yaml — agreed with the product owner, in advance, in writing
service: checkout-api
thresholds:
  - remaining: "< 50%"   # half the budget gone before the window is half over
    action: "Reliability work takes the top slot in the next sprint."
  - remaining: "< 25%"
    action: "No new feature flags enabled in production. Deploys continue."
  - remaining: "<= 0%"
    action: "Feature releases frozen; only reliability fixes and rollbacks ship."
override: "VP Engineering, in writing, recorded in the incident channel."

The override line matters more than the thresholds. A policy with no documented escape hatch gets ignored the first time a launch date collides with it, and once ignored it's dead.

The second SRE number is toil: work tied to running a service that is manual, repetitive, automatable, devoid of enduring value and scales linearly with the service. Google's guardrail is a 50% cap, with at least half an SRE's time going to engineering that reduces future toil. Google admits the cap is aspirational, and outside Google most teams quietly blow through it. Treat it as a tripwire, not a KPI. A team consistently over 50% toil isn't an SRE team, it's an ops team with a better title, and it will lose people.

Platform teams have no equivalent standard, which is their weak point. The usable proxies are adoption per capability, time from empty repo to running production service, and the share of infrastructure requests arriving as a ticket rather than a self-service action. Pick two and publish them.

On-call is the sharpest line

If you get one question, ask what wakes this person up.

SRE carries production for services. The pager isn't incidental; it's the feedback loop the discipline is built around. Mature SRE organisations treat that support as conditional — a team handed an unreliable service with no authority to fix it can, in principle, hand the pager back. Few companies outside the big platforms honour that, and its absence is worth probing in an interview.

Platform engineers carry the platform. If CI, the internal control plane, the registry or the shared ingress is down, that's a platform incident. A tenant service crash-looping on a bad migration is not. The moment platform engineers get paged for tenant services, the platform team has become a shared ops team and will stop shipping product work within two quarters.

DevOps engineers are the variable one, and that variance is the most useful signal in a job ad. Ask which rotation the role joins, how many people are in it, and what the last three pages were about. "The infra rotation, one week in six, mostly pipeline and cluster issues" is a real job. "We're pretty flexible about on-call" means there is no rotation and the answer is you.

The same production incident escalating differently under SRE, platform and generalist models

What a normal week looks like

An SRE's week has a spine of review: burn-rate dashboards, a postmortem's follow-ups, a production readiness review for something about to launch, a capacity conversation, then a block of engineering on whichever failure mode showed up twice. Much of it is negotiating with people who don't report to them.

A platform engineer's week looks like product engineering: conversations with consuming teams, a design doc, a chunk of Go or TypeScript, a module release, and a depressing amount of documentation and migration support. The hardest part isn't technical. It's saying no to bespoke requests without losing the relationship.

A DevOps engineer's week is the most fragmented — a broken pipeline, a new environment, an access request, a Terraform refactor, a cost spike, an hour of someone else's build problem. That fragmentation is the job's defining characteristic and its main career risk.

The same title, at three company sizes

Under about 30 engineers

One infrastructure person, maybe two, doing all three jobs. The contract usually says "DevOps engineer" and that's honest: at this size there is no platform to product-manage and no service stable enough for a meaningful SLO. Don't hire an SRE here. Hire the strongest generalist you can and give them explicit permission to name which of the three jobs they're dropping this quarter.

Roughly 30 to 300

This is where the distinction pays for itself and where most orgs get it wrong. The usual mistake is splitting into named specialists too early: a two-person "SRE team" really doing ticket-driven ops, plus a one-person "platform team" building a portal nobody opens. What works is keeping one infrastructure group, making the paved road its product, and introducing SLOs for the services that carry revenue before you introduce the SRE title.

Beyond a thousand

Here the roles genuinely separate and Team Topologies' vocabulary starts describing real teams: platform teams offering capabilities as a service to stream-aligned teams, SREs assigned to specific high-criticality services. At this scale the DevOps engineer vs SRE vs platform engineer split is real, and "DevOps engineer" usually means one of two things — a cloud infrastructure specialist in a central group, or a rebranded sysadmin role maintaining legacy estate. Respectable work, but a different career.

When a posting is asking for three jobs at once

A large share of "DevOps engineer" ads describe three jobs. You'll recognise the shape: build and own the CI/CD platform, carry production on-call for all services, and drive developer experience and self-service adoption. Cluster administration and security tooling get added on top.

Occasionally that's an honest early-stage generalist role. More often it signals one of three things. Nobody has decided what the role is for, so the ad is a wish list assembled from other companies' ads. Or the previous person did all of it and burned out, and the req backfills a workload that was never sustainable. Or — most common — the organisation wants platform outcomes without funding a platform team, so someone with an infrastructure title is expected to produce developer-experience improvements in the gaps between incidents. That never works: interrupt-driven work always beats project work, and the platform half of the job silently becomes zero.

Ask directly: if all three are on fire in the same week, which do I let burn? An organisation that has thought about the role has an answer. One that hasn't will tell you it depends.

If you're choosing a path

Pick on the kind of feedback you want, because that's what differs day to day.

Choose SRE if you like measurable systems problems, are comfortable with production pressure, and want a documented body of practice behind your skill set. It's the most portable of the three titles and the most legible outside engineering. The cost is the pager and constant organisational negotiation.

Choose platform engineering if you enjoy building things other engineers use and can tolerate slow, adoption-shaped feedback instead of an incident's immediacy. It keeps you closest to software engineering. The cost is that your users can ignore you, and sometimes will.

Choose the DevOps engineer path deliberately rather than by default. It's a fast way to see a lot of infrastructure, especially early on or at a small company, but it's also the title most at risk of becoming a maintenance role. Pick employers where it means infrastructure automation with real engineering time, and be honest with yourself if two years in you've written no code that outlived a quarter.

Decision tree for which role to hire first based on what is currently hurting

If you're writing the job description

Write the accountability line before the skills list. One sentence: this person is accountable for X, and their customer is Y. If you can't write it, you're not ready to open the req.

Then scope it honestly. Name the rotation and its size. Name the one metric the role is judged on in year one. Name what it is explicitly not responsible for — the hardest line to write, and the one candidates will trust you for.

Hire in the order the pain arrives. If production is unreliable and you don't know how unreliable, you want SLOs before you want an SRE team, and one experienced SRE can establish them. If delivery is slow because every environment is a ticket, you want platform work; hiring an SRE for it produces a frustrated SRE. If you have no infrastructure function at all, hire the generalist and don't pretend otherwise in the ad.

One last check before you post: read the ad back and count the distinct products the role owns. If it's more than one, cut it or raise the level and the budget. The tooling list, the years of experience and the cloud certifications all matter far less than whether you and your future hire agree about what happens at 3am. More on structuring infrastructure teams sits in our engineering leadership coverage.

Advertisement

Follow DevOps Society on LinkedIn

Practical infrastructure engineering in your feed.

Follow

Written by

DevOpsSociety Editorial Team

Editorial Team

The DevOpsSociety Editorial Team covers DevOps, cloud infrastructure, Kubernetes, AI infrastructure, platform engineering, cybersecurity, FinOps, and modern engineering practices. We publish practical insights, technical guides, architecture analysis, and research for engineers and technology leaders.

More from DevOpsSociety →
The Infrastructure Briefing

Get the infrastructure briefing.

Practical DevOps, cloud, AI infrastructure and engineering insights, delivered weekly. Read by engineers and engineering leaders.

No spam. Unsubscribe anytime.

DevOps Engineer vs SRE vs Platform Engineer: What's the Difference?, DevOps Society