Senior Site Reliability Engineer

Company: EPAM Systems
Location: Poland
Type: remote
Posted: Aug 19, 2026
Views: 0

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.

Responsibilities

  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect

Requirements

  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
  • Knowledge of FinOps basics for AI workloads
  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code

Benefits

We gather like-minded people:

  • Top tech minds driving innovation in AI, cloud and digital platform modernization
  • Supportive team and agile, startup-like culture
  • Hybrid by design mode and opportunity to work remotely within Poland
  • Chance to work abroad for up to 60 days annually
  • Business-driven relocation opportunities

We provide growth opportunities:

  • Career development programs
  • Thought leadership, mentoring, soft skills and well-being programs
  • Certification (Anthropic, Gemini, GCP, Azure, AWS)
  • English classes

We cover it all:

  • Stable pay
  • Participation in the Employee Stock Purchase Plan with a 15% discount
  • Benefits package (health insurance, multisport, shopping vouchers)
  • Referral bonuses up to $2,000
  • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
  • Corporate, social and well-being events

Please, note:

  • Benefits listed above are available to employees only
  • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
  • We will reach out to selected candidates exclusively

About the Company

Name: EPAM Systems

No detailed information available about this company.