Search Results for "site-reliability-engineering-lead"

Found 3075 jobs

Staff Site Reliability Engineer

replit Remote (United States) Remote, Full time

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the role:

Join our Site Reliability Engineer…

Posted: Apr 8, 2026 15 views
site reliability engineer sre staff engineer kubernetes docker gcp python go terraform pulumi monitoring observability incident management devops infrastructure remote united states
Apply Now

Tech Lead, Site Reliability Engineering (SRE)

edge-node Remote Full-Time

At Edge & Node, we’re focused on building The Graph, a decentralized protocol for accessing and organizing the world’s knowledge and information. Subgraphs, a core technology developed by Edge & Node to access blockchain data, are widely used across web3 to power decentralized applications.

We’re a tig…

Posted: Apr 14, 2025 13 views
Cloud DevOps Engineer GCP Kubernetes SRE Tech Lead
Apply Now

Senior DevOps Engineer

olo Belfast, Northern Ireland, Remote Remote

Olo is a leading SaaS platform accelerating digital transformation in the restaurant industry, by helping customers deliver more personalised and profitable guest experiences. As a result, our digital ordering, payment, and guest engagement solutions enable brands to do more with less and make every…

Posted: Apr 7, 2026 19 views
devops senior aws kubernetes eks terraform ci/cd gitops linux bash python .net argo cd flux github actions prometheus kafka postgres soc2 pci
Apply Now

Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)

deepgram USA | Remote Remote

Company Overview

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice…

Posted: Apr 7, 2026 10 views
Site Reliability Engineer SRE AI ML Infrastructure Kubernetes AWS Terraform Slurm GPU HPC DevOps Platform Engineering Python Go Bash CI/CD Hybrid Cloud
Apply Now

Senior Platform Engineer

attio London; United Kingdom Hybrid

Attio is the CRM built for the AI era.Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ven…

Posted: Apr 7, 2026 11 views
senior platform engineer devops sre aws gcp azure docker kubernetes typescript go python rust terraform pulumi ci/cd monitoring logging tracing infrastructure automation cloud containerization observability
Apply Now

Software Engineer, Site Reliability (SRE)

sierra San Francisco, CA Full time

About us

  • At Sierra, we’re creating a platform to help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco, with growing offices in Atlanta, New York, London, Paris, Madrid, Munich, Singapore, Japan, and Sydney.
  • We are guided by…
Posted: Apr 15, 2026 9 views
software engineer site reliability engineer sre infrastructure terraform aws observability devops ci/cd llm ai saas cloud
Apply Now

Senior Platform Engineer

attio Poland; Germany; Ireland; Portugal Remote, Full time

Attio is the CRM built for the AI era.Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ven…

Posted: Apr 7, 2026 8 views
Senior Platform Engineer Platform Product Engineer DevOps SRE Site Reliability Engineering AWS GCP Azure Docker Kubernetes Terraform Pulumi CI/CD Infrastructure as Code IaC Monitoring Logging Tracing Typescript Go Python Rust Automation Cloud Infrastructure Containerisation
Apply Now

Staff Software Engineer, AI Reliability Engineering

anthropic London, UK Hybrid

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working toge…

Posted: Apr 7, 2026 8 views
staff software engineer AI reliability engineering SRE distributed systems infrastructure monitoring observability high-availability incident response ML hardware GPUs TPUs reliability Anthropic London
Apply Now

Senior Site Reliability Engineer, Platform & Cloud FinOps (100% Remote - Toronto)

hopper Toronto - Remote Remote, Full time

About the job

We are looking for a senior site reliability engineer to join the Cloud FinOps team at Hopper. We manage a large infrastructure in Google Cloud that is used by hundreds of engineers to provide a first class experience to millions of end users around the world.

You are passionate about au…

Posted: Apr 15, 2026 8 views
senior site reliability engineer sre platform cloud finops google cloud kubernetes istio datadog devops remote toronto canada
Apply Now

Senior Software Engineer, Agent Orchestration

decagon New York City On-site

About Decagon

Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.

Our technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that po…

Posted: Apr 10, 2026 8 views
senior software engineer agent orchestration python typescript distributed systems production systems workflow engines ai agents model driven applications engineering on-site new york city
Apply Now
Page 1 of 308 Next