Senior Site Reliability Engineer
We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.
Responsibilities
- Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
- Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
- Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
- Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
- Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
- Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
- Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
- Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect
Requirements
- 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
- Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
- Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
- Knowledge of FinOps basics for AI workloads
- Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code
Benefits
We gather like-minded people:
- Top tech minds driving innovation in AI, cloud and digital platform modernization
- Supportive team and agile, startup-like culture
- Hybrid by design mode and opportunity to work remotely within Poland
- Chance to work abroad for up to 60 days annually
- Business-driven relocation opportunities
We provide growth opportunities:
- Career development programs
- Thought leadership, mentoring, soft skills and well-being programs
- Certification (Anthropic, Gemini, GCP, Azure, AWS)
- English classes
We cover it all:
- Stable pay
- Participation in the Employee Stock Purchase Plan with a 15% discount
- Benefits package (health insurance, multisport, shopping vouchers)
- Referral bonuses up to $2,000
- Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
- Corporate, social and well-being events
Please, note:
- Benefits listed above are available to employees only
- We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
- We will reach out to selected candidates exclusively
About the Company
More jobs at EPAM Systems
-
Lead Software Engineer - Java8, Spring, REST API, Microservices
· · Aug 19, 2026
-
Senior Java Software Developer – Java 8, Microservices, ReactJS, JUnit
· · Aug 19, 2026
-
Lead / Senior Software Developer - .NET
· · Aug 19, 2026
-
Software Engineer with Anthropic Claude
· · Aug 19, 2026
-
Senior Software Engineer - Python with GenAI, LLM
· · Aug 19, 2026