Principal Site Reliability Engineer
About The Role
As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs both in our own cloud and inside customer-controlled environments — and you will hold the organization to that standard.
We run dedicated, isolated environments per customer, which makes repeatability and automation the central engineering problem rather than an afterthought. Depth of judgment about reliability engineering matters far more here than experience with any particular cloud, orchestrator, or observability vendor.
This is an individual contributor role with organization-level influence.
Core Responsibilities
Own the reliability architecture of the platform: deployment topology, failure domains, capacity strategy, and the automation that makes environments reproducible.
Define service level objectives with product and engineering leadership, and drive the work needed to meet them.
Set the standard for observability — dashboards, logs, metrics, tracing, and alerting — so that issues are detected before customers report them.
Lead major incidents, run blameless postmortems, and make sure the corrective work actually lands.
Contribute to design and architecture across infrastructure and applications, with automation, performance, reliability, and security as first-class concerns.
Drive infrastructure lifecycle at scale: provisioning, upgrades, and decommissioning across many isolated environments.
Ensure infrastructure and applications meet or exceed enterprise compliance requirements, and design identity and access controls across platforms and services.
Partner with enterprise customers on custom infrastructure requirements, translating their constraints into repeatable patterns rather than one-off work.
Raise the bar through code and design review, and mentor SREs and product engineers on reliability practice.
Improve and maintain infrastructure and process documentation.
Participate in and help evolve the on-call rotation, including how the team balances operational load against project work.
Qualifications and Experience
10+ years in infrastructure, SRE, or platform engineering, including deep experience operating large-scale distributed systems in production.
Expertise designing, analyzing, and troubleshooting distributed systems, with a track record of reliability decisions that held up under growth.
Deep experience with at least one major public cloud provider, and the ability to reason across providers rather than within one.
Strong command of container orchestration: cluster operation, workload scheduling, networking, and the failure modes that come with them.
Fluency with infrastructure as code, configuration management, and CI/CD pipeline design.
Strong scripting and automation ability, and comfort reading and debugging application code in the languages your services are written in.
Experience defining observability strategy — instrumentation, query languages, dashboards, and alert design that minimizes noise.
Demonstrated ability to influence without authority and align multiple teams behind a technical direction.
Experience operating under enterprise security and compliance frameworks.
Solid understanding of Unix/Linux operating systems and networking fundamentals.
About the Company
More jobs at MetaRouter
-
Senior Software Engineer - Full Stack
Remote · Full-time · Aug 19, 2026