Lead Site Reliability Engineer

Company: EPAM Systems
Location: Mexico, Colombia, Brazil, Argentina
Type: remote
Posted: Oct 7, 2026
Views: 0

We are seeking a Lead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate DevOps maturity for high-impact services. You will design scalable automation, improve CI/CD and release practices, and lead rapid incident response.

Responsibilities

  • Design reliability strategies and SRE practices for business-critical infrastructure
  • Build automation and tooling in Python to improve stability, consistency, and operational leverage
  • Develop and maintain CI/CD workflows and source control practices using GitLab
  • Lead incident response during on-call rotations and restore service for business-critical issues
  • Improve release management processes to support enterprise-scale delivery
  • Harden cloud infrastructure across networking, compute, security, and IAM controls
  • Implement configuration automation to reduce manual work and prevent drift
  • Operate and troubleshoot Kubernetes-based workloads and developer-facing platform usage
  • Partner with engineering stakeholders to prioritize reliability work and manage change safely
  • Assess systemic risks and drive corrective actions to prevent recurring incidents

Requirements

  • 5+ years of site reliability engineering or DevOps experience in cloud environments
  • Hands-on experience with a leading cloud provider, with practical work across Amazon Web Services and Microsoft Azure
  • Leadership ability to guide technical direction and take ownership of critical infrastructure outcomes
  • Project delivery experience improving DevOps tools, processes, and engineering maturity at scale
  • Deep CI/CD knowledge across pipelines, source control, and release management
  • Strong Kubernetes skills with practical usage as a developer
  • Advanced Python programming skills for automation and tooling
  • Enterprise-scale release management experience supporting complex systems
  • Solid infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
  • Strong analytical skills to diagnose complex issues and identify high-leverage solutions
  • Effective incident response skills, including on-call ownership and rapid restoration of service
  • Upper-Intermediate English proficiency (B2, Upper-Intermediate)

Nice to have

  • Amazon Web Services certification or proven advanced AWS operational experience
  • Microsoft Azure certification or proven advanced Azure operational experience
  • AI Architecture experience for reliability-focused platform design
  • AI Solution Engineering experience integrating AI-enabled capabilities into operations
  • Gen AI Solutions Development experience for operational intelligence and automation use cases

Benefits

  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

About the Company

Name: EPAM Systems

No detailed information available about this company.

More jobs at EPAM Systems