Search Results for "sre-ai-engineer"

Found 1714 jobs

Staff Software Engineer, AI Reliability Engineering

anthropic London, UK Hybrid

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working toge…

Posted: Apr 7, 2026 8 views
staff software engineer AI reliability engineering SRE distributed systems infrastructure monitoring observability high-availability incident response ML hardware GPUs TPUs reliability Anthropic London
Apply Now

Senior Site Reliability Engineer, Platform & Cloud FinOps (100% Remote - Toronto)

hopper Toronto - Remote Remote, Full time

About the job

We are looking for a senior site reliability engineer to join the Cloud FinOps team at Hopper. We manage a large infrastructure in Google Cloud that is used by hundreds of engineers to provide a first class experience to millions of end users around the world.

You are passionate about au…

Posted: Apr 15, 2026 8 views
senior site reliability engineer sre platform cloud finops google cloud kubernetes istio datadog devops remote toronto canada
Apply Now

Site Reliability Engineer (USA Only - 100% Remote)

close USA - Remote Remote, Full time

About Us

Close is a bootstrapped, profitable, 100% remote, ~100 person team of thoughtful individuals who prioritize taking ownership and making a meaningful impact. We’re eager to make a product our customers fall in love with over and over again.

We 💛 small scaling businesses. Since 2013, we’ve been…

Posted: Apr 7, 2026 8 views
site reliability engineer sre infrastructure aws terraform kubernetes ansible mongodb postgresql elasticsearch python flask cicd remote usa
Apply Now

Site Reliability Engineer

runpod Remote, USA Remote

Runpod is the foundational platform for developers to build and run custom AI systems that scale. With over 500,000 developers worldwide and an annual recurring revenue run rate exceeding $120M, Runpod operates at the intersection of developer velocity and production-scale AI. Founded in 2022, we’ve…

Posted: Apr 21, 2026 7 views
site reliability engineer sre reliability engineering production engineering linux networking containers distributed systems sli slo incident response postmortem scripting python go bash monitoring alerting prometheus grafana ci/cd gpu ai ml infrastructure as code remote usa
Apply Now

Forward Deployed Engineer - US

dash0 United States - remote Remote, Full time

About Dash0

Join Dash0 and help us define the future of observability. We are OpenTelemetry-native, building a delightful, simple, and AI-centric platform that eliminates vendor lock-in and meaningless toil. Shape a product that developers love—all with transparent pricing and cost-control built in.

T…

Posted: Apr 17, 2026 7 views
forward deployed engineer observability opentelemetry kubernetes sre platform engineer distributed systems go python java typescript remote customer success
Apply Now

DevOps Engineer

zowie Poland Remote, Full time

About Zowie:

At Zowie, we’re revolutionizing how businesses interact with their customers. We’re creating a future where AI Agents handle 100% of customer interactions - delivering instant, personalized, and exceptional experiences.

We believe AI Agents represent the next major technological shift, an…

Posted: Apr 7, 2026 7 views
devops platform engineer sre aws gcp iac ci/cd mongodb postgres elasticsearch cassandra redis remote poland
Apply Now

Lead Site Reliability Engineer

gleanwork San Francisco Bay Area Hybrid

About Glean:

Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With ov…

Posted: Apr 7, 2026 7 views
lead site reliability engineer sre cloud docker kubernetes terraform aws azure gcp hybrid palo alto automation monitoring incident management
Apply Now

Senior DevOps Engineer

zipline Canada Remote, Full time

About Zipline

Zipline is a well-funded, rapidly growing SaaS company transforming how frontline teams work. We empower the world’s leading brands across retail, healthcare, logistics, and beyond to connect, align, and inspire their employees- from headquarters to the front lines. Our customers consis…

Posted: Apr 15, 2026 7 views
Senior DevOps Engineer DevOps SRE AWS Terraform Kubernetes Docker CI/CD Infrastructure as Code Cloud Remote Ruby SaaS Platform Reliability
Apply Now

Site Reliability Engineer

workos United States Remote, Full time

About WorkOS 🚀

WorkOS builds modern developer tools and APIs that make it easy for companies to become Enterprise Ready. Our platform powers authentication, identity, authorization, and other critical infrastructure that developers need to securely scale their products to large organizations.

We recen…

Posted: Apr 8, 2026 6 views
site reliability engineer sre aws typescript kubernetes prometheus grafana datadog opentelemetry monitoring alerting incident response cloud infrastructure reliability performance observability
Apply Now

Senior Site Reliability Engineer, Platform & Cloud FinOps (100% Remote - USA Central & EST)

hopper Boston - Remote; Austin - Remote Remote, Full time

About the job

We are looking for a senior site reliability engineer to join the Cloud FinOps team at Hopper. We manage a large infrastructure in Google Cloud that is used by hundreds of engineers to provide a first class experience to millions of end users around the world.

You are passionate about au…

Posted: Apr 7, 2026 6 views
senior site reliability engineer sre platform cloud finops google cloud kubernetes istio datadog devops remote usa
Apply Now
Previous Page 2 of 172 Next