Search Results for "site-reliability-engineer-ai-first-platform"

Found 1121 jobs

Site Reliability Engineer

runpod Remote, USA Remote

Runpod is the foundational platform for developers to build and run custom AI systems that scale. With over 500,000 developers worldwide and an annual recurring revenue run rate exceeding $120M, Runpod operates at the intersection of developer velocity and production-scale AI. Founded in 2022, we’ve…

Posted: Apr 21, 2026 7 views
site reliability engineer sre reliability engineering production engineering linux networking containers distributed systems sli slo incident response postmortem scripting python go bash monitoring alerting prometheus grafana ci/cd gpu ai ml infrastructure as code remote usa
Apply Now

Lead Site Reliability Engineer

gleanwork San Francisco Bay Area Hybrid

About Glean:

Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With ov…

Posted: Apr 7, 2026 7 views
lead site reliability engineer sre cloud docker kubernetes terraform aws azure gcp hybrid palo alto automation monitoring incident management
Apply Now

Senior DevOps Engineer

zipline Canada Remote, Full time

About Zipline

Zipline is a well-funded, rapidly growing SaaS company transforming how frontline teams work. We empower the world’s leading brands across retail, healthcare, logistics, and beyond to connect, align, and inspire their employees- from headquarters to the front lines. Our customers consis…

Posted: Apr 15, 2026 7 views
Senior DevOps Engineer DevOps SRE AWS Terraform Kubernetes Docker CI/CD Infrastructure as Code Cloud Remote Ruby SaaS Platform Reliability
Apply Now

Senior Platform Engineer

attio London; United Kingdom Hybrid

Attio is the CRM built for the AI era.Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ven…

Posted: Apr 7, 2026 11 views
senior platform engineer devops sre aws gcp azure docker kubernetes typescript go python rust terraform pulumi ci/cd monitoring logging tracing infrastructure automation cloud containerization observability
Apply Now

Senior Site Reliability Engineer, Platform & Cloud FinOps (100% Remote - USA Central & EST)

hopper Boston - Remote; Austin - Remote Remote, Full time

About the job

We are looking for a senior site reliability engineer to join the Cloud FinOps team at Hopper. We manage a large infrastructure in Google Cloud that is used by hundreds of engineers to provide a first class experience to millions of end users around the world.

You are passionate about au…

Posted: Apr 7, 2026 6 views
senior site reliability engineer sre platform cloud finops google cloud kubernetes istio datadog devops remote usa
Apply Now

Senior Backend Engineer, MCP, RAG and Fine-Tuning (100% Remote)

hopper Boston - Remote; Austin - Remote; Dallas - Remote; Denver - Remote; Virginia - Remote Remote, Full time

About the job

Did you know Hopper's technology powers some of the largest travel portals in the world? Our Hopper Technology Solutions (HTS) platform drives travel purchases for major global brands — from banks to airlines to fintech leaders. This platform is essentially a configurable, multi-tenant…

Posted: Apr 7, 2026 19 views
Senior Backend Engineer MCP RAG Fine-Tuning Scala GCP React TypeScript LLM AI distributed systems remote travel fintech
Apply Now

Site Reliability Engineer

workos United States Remote, Full time

About WorkOS 🚀

WorkOS builds modern developer tools and APIs that make it easy for companies to become Enterprise Ready. Our platform powers authentication, identity, authorization, and other critical infrastructure that developers need to securely scale their products to large organizations.

We recen…

Posted: Apr 8, 2026 6 views
site reliability engineer sre aws typescript kubernetes prometheus grafana datadog opentelemetry monitoring alerting incident response cloud infrastructure reliability performance observability
Apply Now

Staff+ Site Reliability Engineer, Safeguards ML Infra

anthropic San Francisco, CA | Seattle, WA | New York City, NY Remote-Friendly (Travel-Required)

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working toge…

Posted: Aug 11, 2026 4 views
site reliability engineer safeguards ML infra production deployment canary Python AWS GCP on-call automation
Apply Now

Engineering Manager | Remote | Europe

n8n Berlin Office; Albania; Austria; Belgium; Bosnia; Bulgaria; Croatia; Czech Republic; Denmark; Estonia; Finland; France; Germany; Greece; Hungary; Ireland; Italy; Kosovo; Latvia; Lithuania; London Office; Moldova; Montenegro; Netherlands; Norway; Poland; Portugal; Romania; Serbia; Slovakia; Slovenia; Spain; Sweden; Ukraine; United Kingdom Remote, Full time

The AI orchestration of your wildest imagination.

n8n is the open workflow orchestration platform built for the new era of AI. We give technical teams the freedom of code with the speed of no-code, so they can automate faster, smarter, and without limits. Backed by a fiercely inventive community and…

Posted: Apr 7, 2026 7 views
engineering manager remote europe typescript node.js ai saas open source workflow automation leadership team management software engineering high-growth
Apply Now

Site Reliability Engineer (SRE)

Freeplay Boulder, CO Full time

The Opportunity

We're hiring an experienced Site Reliability Engineer to own the reliability of the Freeplay platform and drive success for our most advanced enterprise customers. In this role, you will bridge the gap between core infrastructure engineering and high-stakes customer deployments. You w…

Posted: Apr 7, 2026 3 views
Site Reliability Engineer SRE Kubernetes Terraform PostgreSQL Elasticsearch AWS GCP Azure Helm Replicated KOTS Datadog NATS JetStream VPC IAM Infrastructure as Code BYOC Enterprise AI ML
Apply Now
Previous Page 2 of 113 Next