Search Results for "sre-incidents-and-monitoring"

Found 437 jobs

Staff Site Reliability Engineer

replit Remote (United States) Remote, Full time

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the role:

Join our Site Reliability E...

Posted: Apr 8, 2026 15 views
site reliability engineer sre staff engineer kubernetes docker gcp python go terraform pulumi monitoring observability incident management devops infrastructure remote united states
Apply Now

Staff Software Engineer, AI Reliability Engineering

anthropic London, UK Hybrid

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working...

Posted: Apr 7, 2026 8 views
staff software engineer AI reliability engineering SRE distributed systems infrastructure monitoring observability high-availability incident response ML hardware GPUs TPUs reliability Anthropic London
Apply Now

Lead Site Reliability Engineer

gleanwork San Francisco Bay Area Hybrid

About Glean:

Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With...

Posted: Apr 7, 2026 7 views
lead site reliability engineer sre cloud docker kubernetes terraform aws azure gcp hybrid palo alto automation monitoring incident management
Apply Now

Senior Platform Engineer

attio London; United Kingdom Hybrid

Attio is the CRM built for the AI era. Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ve...

Posted: Apr 7, 2026 10 views
senior platform engineer devops sre aws gcp azure docker kubernetes typescript go python rust terraform pulumi ci/cd monitoring logging tracing infrastructure automation cloud containerization observability
Apply Now

Site Reliability Engineer

workos United States Remote, Full time

About WorkOS 🚀

WorkOS builds modern developer tools and APIs that make it easy for companies to become Enterprise Ready. Our platform powers authentication, identity, authorization, and other critical infrastructure that developers need to securely scale their products to large organizations.

We...

Posted: Apr 8, 2026 6 views
site reliability engineer sre aws typescript kubernetes prometheus grafana datadog opentelemetry monitoring alerting incident response cloud infrastructure reliability performance observability
Apply Now

Senior Platform Engineer

attio Poland; Germany; Ireland; Portugal Remote, Full time

Attio is the CRM built for the AI era. Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ve...

Posted: Apr 7, 2026 8 views
Senior Platform Engineer Platform Product Engineer DevOps SRE Site Reliability Engineering AWS GCP Azure Docker Kubernetes Terraform Pulumi CI/CD Infrastructure as Code IaC Monitoring Logging Tracing Typescript Go Python Rust Automation Cloud Infrastructure Containerisation
Apply Now

Site Reliability Engineer

runpod Remote, USA Remote

Runpod is the foundational platform for developers to build and run custom AI systems that scale. With over 500,000 developers worldwide and an annual recurring revenue run rate exceeding $120M, Runpod operates at the intersection of developer velocity and production-scale AI. Founded in 2022, we’ve...

Posted: Apr 21, 2026 7 views
site reliability engineer sre reliability engineering production engineering linux networking containers distributed systems sli slo incident response postmortem scripting python go bash monitoring alerting prometheus grafana ci/cd gpu ai ml infrastructure as code remote usa
Apply Now

DevOps/SRE Engineer

chronicle-labs Remote - APAC, Asia, Australia Full-Time, Remote

About Chronicle Labs

Chronicle Protocol is a cutting-edge decentralized Oracle solution delivering secure, transparent, and verifiable real-time data. With over $10 billion in collateral secured for major DeFi ecosystems, we empower institutions, builders, and tokenized asset issuers with unmatc...

Posted: Nov 15, 2025 8 views
AWS CI/CD Cloud DevOps Engineer GitOps Kubernetes Linux SRE Web3
Apply Now

Lead Site Reliability Engineer (Remote)

livepeer Remote Remote

About Livepeer:

Livepeer is on a mission to build the world’s open video infrastructure. Founded in 2017, it is the world’s first open-source protocol for decentralized video streaming, built on Ethereum. The project has empowered developers to create scalable, cost-effective, and censorship-re...

Posted: Jun 12, 2025 20 views
Ansible AWS CI/CD DevOps Docker Engineer Ethereum GCP Grafana Kubernetes Linux Prometheus SRE Terraform Web3
Apply Now

Tech Lead, Site Reliability Engineering (SRE)

edge-node Remote Full-Time

At Edge & Node, we’re focused on building The Graph, a decentralized protocol for accessing and organizing the world’s knowledge and information. Subgraphs, a core technology developed by Edge & Node to access blockchain data, are widely used across web3 to power decentralized applications.

We’re a...

Posted: Apr 14, 2025 13 views
Cloud DevOps Engineer GCP Kubernetes SRE Tech Lead
Apply Now
Page 1 of 44 Next