SRE/Senior SRE
✨ AI Summary
Saviynt, an AI-powered identity security platform protecting Fortune 500 companies, is hiring a Senior SRE Engineer in Bengaluru. You'll own reliability for AWS and Kubernetes (EKS) infrastructure, build self-healing systems with AI agents, and use LLM APIs like OpenAI for alert triage and incident automation. Requires 5+ years in SRE/DevOps/Platform Engineering, strong coding in Python or Go, and production-grade Kubernetes experience; bonus for LLM/AI tooling exposure.
We’re a fast-moving AI Security Company building AI-native infrastructure and applications powered by LLMs and autonomous agents. Our stack is deeply integrated with AWS, Kubernetes, and OpenAI-based systems, and we’re rethinking reliability in a world where software can reason, adapt, and self-heal.
We’re hiring a Senior SRE Engineer to own reliability across our cloud-native and AI-driven platform. You’ll work at the intersection of distributed systems, Kubernetes operations, and LLM-powered automation, building systems that don’t just scale—but think and fix themselves.
WHAT YOU BRING
- 5+ years in SRE / DevOps / Platform Engineering.
- Strong hands-on experience with:
- AWS infrastructure at scale
- Kubernetes (production-grade clusters)
- Proven ability to debug complex distributed systems under pressure.
- Strong coding skills (Python or Go)—you build internal platforms and tools.
- Experience implementing monitoring, alerting, and incident management systems.
- Experience working with LLM APIs such as the OpenAI API.
- Familiarity with agent frameworks like:
- LangChain
- AutoGen
- Built or experimented with:
- AI agents for DevOps / SRE workflows
- Retrieval-Augmented Generation (RAG) systems
- Vector databases (Pinecone, Weaviate, etc.)
- Exposure to AIOps or intelligent automation systems.
Bonus (AI / LLM Focus)
WHAT YOU WILL BE DOING
- Own uptime, reliability, and performance of services running on AWS + Kubernetes (EKS).
- Design and implement self-healing infrastructure using automation and AI agents.
- Build LLM-powered operational tooling using APIs such as the OpenAI API for:
- Intelligent alert triage
- Incident summarization
- Root cause analysis
- Runbook automation
- Manage and scale Kubernetes workloads:
- Deployments, autoscaling, resource optimization
- Cluster reliability and cost efficiency
- Build and evolve observability systems:
- Metrics (Prometheus), dashboards (Grafana)
- Logs (ELK / OpenSearch)
- Tracing (OpenTelemetry)
- Define and enforce SLOs, SLAs, and error budgets tied to business metrics.
- Automate infrastructure using Terraform and CI/CD pipelines.
- Lead incident response, postmortems, and continuous reliability improvements.
- Introduce chaos engineering practices to proactively test system resilience.
About the Company
More jobs at Saviynt
-
Associate Principal Engineer - Threat Researcher
Bengaluru · full_time · Sep 29, 2026
-
Sr.Manager, Software Engineering
Bengaluru · full_time · Sep 28, 2026
-
Distinguished Software Engineer, Data Platform
Milpitas, California · full_time · Aug 19, 2026
-
Staff Solutions Engineer
Japan · full_time · Aug 19, 2026
-
Principal Site Reliability Engineer, Google Cloud
Atlanta · full_time · Aug 19, 2026