Applied Methods
~The MetaEngineeringSite Reliability Engineer

Site Reliability Engineer

Engineers in this role maintain the reliability and performance of AI infrastructure at scale, spending their days on incident response, automation, and observability across distributed systems that power AI workloads. They differ from software engineers by focusing on operational excellence and system resilience rather than feature development, and from DevOps roles by owning broader platform-level reliability goals. These teams typically sit within infrastructure or platform organizations, partnering closely with product engineering teams to ensure AI services remain fast, secure, and always available across multiple regions.

$ titles --canonical
Site Reliability EngineerSenior SREStaff SREProduction EngineerReliability EngineerInfrastructure SRE
Open Jobs114
Companies Hiring40
$ expectations --role site-reliability-engineer

Measured across 113 of 114 open postings.

21%
expect AI in the role's own work
2% state it as a requirement
4%
work directly with customers
16%
manage people
Senior
most common level
37% of open postings
BY LEVEL

This role is advertised at 3 levels, so a single figure for the role would describe none of them. Experience and pay are the midpoints for each level on its own.

LevelShareMedian yearsMedian pay
Mid27%(30)4—
Senior37%(42)5$250k
Staff / Principal29%(33)8$280k

A dash means too few postings stated it to report a midpoint. Most companies do not publish a salary band, so pay is indicative rather than a market rate. 3 levels with fewer than 10 open postings are not shown.

WHAT THEY ASK FOR, VERBATIM

“Familiarity with AI coding tools is expected; experience as an SRE, with development frameworks like multi-agent workflows”

Vapi · Member of Technical Staff, Agentic Release Engineer

“Real experience running LLM applications, including tracing, evals, and prompt and cache mechanics”

LangChain · Software Engineer, Agent Systems (GTM Engineering)

“Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.”

Lambda · Senior HPC Engineer - Fleet Engineering

“supporting rapidly growing AI workloads”

Harvey · Senior Software Engineer, Production Engineering
$ barriers
18%
advertised as remote
of postings that state a work mode
18%
state a degree requirement
6%
need a security clearance
91%
still advertised a month later
6 points slower than the board

Requirements are a share of every open posting, so a role missing from this list is one where almost nobody asks. Work mode is different: many postings never say, so that figure counts only the ones that do. A posting stops being advertised when it is filled, cancelled or reorganised, so read the last figure as how long these stay on the market, not as time to hire.

$02

Skills

What companies are looking for in this role.

$ skills --core

Incident response and reliability

92%

Monitoring and observability

88%

Infrastructure automation and IaC

72%

Cloud infrastructure operations

71%

Distributed systems architecture

65%

Systems performance optimization

43%

CI/CD and release automation

29%

Developer platform engineering

12%

Cloud and infrastructure security

11%

Network engineering and operations

9%

Technical issue diagnosis

9%
$ skills --emerging

AI and GPU infrastructure operations

36%
$ skills --soft

Technical team leadership and mentoring

17%
$03

Technology

The tools and technologies that define this role.

$ tech --language
Pythonhigh
Gomoderate
Bashlow
C/C++low
Javalow
Rustlow
$ tech --platform
Kuberneteshigh
AWSmoderate
Azuremoderate
Dockermoderate
Google Cloud Platformmoderate
Linuxmoderate
Datadoglow
PostgreSQLlow
$ tech --tool
Grafanamoderate
Prometheusmoderate
Terraformmoderate
Ansiblelow
Argo CDlow
CloudFormationlow
GitHub Actionslow
Helmlow
Pulumilow
$ tech --concept
CI/CDlow
GPUlow
Observabilitylow
$04

Open Jobs

114 open Site Reliability Engineer jobs across 40 companies.

Latent Health1w
Site Reliability Engineer
San Francisco·Engineering
Nscale2w
Senior Operational Engineer
US·Engineering
MongoDB2w
Cloud Operations Engineer
Bengaluru·Engineering
Inferact2w
Member of Technical Staff, Production Site Reliability Engineer
San Francisco·Engineering
Vapi2w
Member of Technical Staff, Agentic Release Engineer
San Francisco·Engineering
CoreWeave3w
Operations Engineer, MetalDev
New York, NY/ Livingston, NJ·Engineering
Replit3w
Staff Site Reliability Engineer
Remote - United States·Engineering
Replit3w
Senior Site Reliability Engineer
Remote - US·Engineering
Vannevar Labs3w
Site Reliability Engineer
San Diego, California·Engineering
Nscale3w
Operational Data & Observability Engineer
US·Engineering
Legora3w
Site Reliability Engineer - Platform Team
Stockholm HQ·Engineering
Legora3w
Senior Site Reliability Engineer - Platform Team
Stockholm HQ·Engineering
Vapi4w
Member of Technical Staff, Infrastructure Engineer
San Francisco·Engineering
Anthropic1mo
Staff+ Software Engineer, ML Sampling Path
San Francisco, CA·Engineering
Nebius1mo
Site Reliability Engineer (SRE) - Early Talent
Amsterdam, Netherlands·Engineering
LangChain1mo
Software Engineer, Agent Systems (GTM Engineering)
San Francisco, CA·Engineering
Waymo1mo
Senior Staff Site Reliability Engineer, Lead
Mountain View, CA, USA·Engineering
Anthropic1mo
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco, CA·Engineering
DataHub1mo
DevOps
Bengaluru, Karnataka, India·Engineering
Shield AI1mo
Senior Staff Lead Site Reliability Engineer (R5803)
San Diego, California·Engineering