Lead Site Reliability Engineer
Company:
EPAM Systems
Location:
Mexico, Colombia, Brazil, Argentina
Type:
remote
Posted:
Oct 7, 2026
Views:
0
We are seeking a Lead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate DevOps maturity for high-impact services. You will design scalable automation, improve CI/CD and release practices, and lead rapid incident response.
Responsibilities
- Design reliability strategies and SRE practices for business-critical infrastructure
- Build automation and tooling in Python to improve stability, consistency, and operational leverage
- Develop and maintain CI/CD workflows and source control practices using GitLab
- Lead incident response during on-call rotations and restore service for business-critical issues
- Improve release management processes to support enterprise-scale delivery
- Harden cloud infrastructure across networking, compute, security, and IAM controls
- Implement configuration automation to reduce manual work and prevent drift
- Operate and troubleshoot Kubernetes-based workloads and developer-facing platform usage
- Partner with engineering stakeholders to prioritize reliability work and manage change safely
- Assess systemic risks and drive corrective actions to prevent recurring incidents
Requirements
- 5+ years of site reliability engineering or DevOps experience in cloud environments
- Hands-on experience with a leading cloud provider, with practical work across Amazon Web Services and Microsoft Azure
- Leadership ability to guide technical direction and take ownership of critical infrastructure outcomes
- Project delivery experience improving DevOps tools, processes, and engineering maturity at scale
- Deep CI/CD knowledge across pipelines, source control, and release management
- Strong Kubernetes skills with practical usage as a developer
- Advanced Python programming skills for automation and tooling
- Enterprise-scale release management experience supporting complex systems
- Solid infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
- Strong analytical skills to diagnose complex issues and identify high-leverage solutions
- Effective incident response skills, including on-call ownership and rapid restoration of service
- Upper-Intermediate English proficiency (B2, Upper-Intermediate)
Nice to have
- Amazon Web Services certification or proven advanced AWS operational experience
- Microsoft Azure certification or proven advanced Azure operational experience
- AI Architecture experience for reliability-focused platform design
- AI Solution Engineering experience integrating AI-enabled capabilities into operations
- Gen AI Solutions Development experience for operational intelligence and automation use cases
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
About the Company
More jobs at EPAM Systems
-
Chief Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Senior Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Junior Site Reliability Engineer
Mexico · remote · Oct 7, 2026
-
Java Developer - Big Data
Poland · remote · Oct 7, 2026
-
Lead AI Software Engineer
Mexico · remote · Oct 7, 2026