Search Results for "-lead-site-reliability-engineer"

Found 3280 jobs

Lead Site Reliability Engineer

gleanwork San Francisco Bay Area Hybrid

About Glean:

Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With ov…

Posted: Apr 7, 2026 7 views
lead site reliability engineer sre cloud docker kubernetes terraform aws azure gcp hybrid palo alto automation monitoring incident management
Apply Now

Site Reliability Engineer (SRE), Platform Engineering Team

cubist Remote Remote, Full time

Overview:

Most of the world’s digital infrastructure has decades of reliability engineering behind it. Web3, by contrast, is newer and still catching up to the standards of modern production systems. At Cubist, we’re laying the groundwork the industry needs: secure, high-assurance infrastructure for…

Posted: Apr 7, 2026 7 views
Site Reliability Engineer SRE Platform Engineering AWS CDK Observability OpenTelemetry Rust TypeScript Go WebAssembly Kubernetes Terraform Ansible Cloud Infrastructure CI/CD Security Web3 Blockchain
Apply Now

Senior Backend Engineer

proton-ai Latin America Only Full-Time

About Us:

The wholesale distribution industry is ready for a revolution, and Proton is leading the charge. The world relies on distributors to sell nearly every physical product, but despite its massive contribution to the global economy, this industry has been left behind in terms of technology. Pro…

Posted: Mar 18, 2025 79 views
backend python go elasticsearch redis sql kubernetes airflow machine learning ai vue.js nuxt docker github actions
Apply Now

Site Reliability Engineer

asymmetric.re Remote - AMER Remote

Asymmetric Research:

Asymmetric Research ("AR") is a boutique security venture focused on deep partnerships with L1/L2 blockchains and DeFi protocols in an effort to keep them safe. We specialize in four core domains of web3 security: research, engineering, incident response, and infrastructure servi…

Posted: Apr 7, 2026 6 views
site reliability engineer sre devops infrastructure linux load balancer haproxy ansible chef puppet saltstack go python rust ci/cd grafana loki prometheus alertmanager nomad kubernetes blockchain bitcoin ethereum solana cosmos remote full time
Apply Now

Senior Site Reliability Engineer, Platform & Cloud FinOps (100% Remote - USA Central & EST)

hopper Boston - Remote; Austin - Remote Remote, Full time

About the job

We are looking for a senior site reliability engineer to join the Cloud FinOps team at Hopper. We manage a large infrastructure in Google Cloud that is used by hundreds of engineers to provide a first class experience to millions of end users around the world.

You are passionate about au…

Posted: Apr 7, 2026 6 views
senior site reliability engineer sre platform cloud finops google cloud kubernetes istio datadog devops remote usa
Apply Now

Software Engineer, Observability

airtable San Francisco, CA; New York, NY; Remote (Seattle, WA only) Hybrid/Remote

Airtable is the no-code app platform that empowers people closest to the work to accelerate their most critical business processes. More than 500,000 organizations, including 80% of the Fortune 100, rely on Airtable to transform how work gets done.

The Observability team at Airtable ensures that our…

Posted: Apr 7, 2026 5 views
software engineer observability logging metrics tracing prometheus grafana datadog opentelemetry elk stack kubernetes distributed systems llm ai infrastructure sre site reliability
Apply Now

Senior Software Engineer, Infrastructure

decagon New York City Full time

About Decagon

Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.

Our technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that po…

Posted: Apr 7, 2026 14 views
Senior Software Engineer Infrastructure Decagon conversational AI platform networking data ML serving developer platform real-time voice SLOs low-latency production infrastructure Kubernetes GCP AWS Azure Terraform GitOps CI/CD observability OpenTelemetry Prometheus Grafana Datadog on-prem air-gapped customer-managed deployments equity New York City on-site
Apply Now

Staff Software Engineer, AI Reliability Engineering

anthropic London, UK Hybrid

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working toge…

Posted: Apr 7, 2026 8 views
staff software engineer AI reliability engineering SRE distributed systems infrastructure monitoring observability high-availability incident response ML hardware GPUs TPUs reliability Anthropic London
Apply Now

Senior Platform Engineer

attio London; United Kingdom Hybrid

Attio is the CRM built for the AI era.Designed for the most ambitious go-to-market teams, it gives companies the power to understand every customer, automate at scale, and build their go-to-market motion exactly as they need. We've raised $116M from some of the world's best investors: GV (Google Ven…

Posted: Apr 7, 2026 11 views
senior platform engineer devops sre aws gcp azure docker kubernetes typescript go python rust terraform pulumi ci/cd monitoring logging tracing infrastructure automation cloud containerization observability
Apply Now

Staff+ Site Reliability Engineer, Safeguards ML Infra

anthropic San Francisco, CA | Seattle, WA | New York City, NY Remote-Friendly (Travel-Required)

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working toge…

Posted: Aug 11, 2026 4 views
site reliability engineer safeguards ML infra production deployment canary Python AWS GCP on-call automation
Apply Now
Previous Page 2 of 328 Next