Chief Site Reliability Engineer
Company:
EPAM Systems
Location:
Mexico, Colombia, Brazil, Argentina
Type:
remote
Posted:
Oct 7, 2026
Views:
0
We are seeking a Chief Site Reliability Engineer to strengthen reliability and DevOps maturity across mission-critical platforms in a fast-changing environment. You will drive resilient cloud and Kubernetes operations, improve CI/CD and release processes, and lead rapid incident response.
Responsibilities
- Lead reliability strategy for critical infrastructure to enable rapid business change
- Design resilient cloud architectures and operational patterns across AWS and Azure
- Build and improve CI/CD pipelines, source control practices, and release workflows
- Automate infrastructure and operational tasks using Python to reduce toil and risk
- Harden platform foundations across networking, compute, security, IAM, and configuration automation
- Operate and evolve Kubernetes usage patterns to improve stability and delivery speed
- Coordinate on-call response and resolve business-critical incidents under time pressure
- Investigate systemic issues, perform root-cause analysis, and drive corrective actions
- Partner with engineering teams to deliver high-leverage solutions over quick fixes
- Define and track reliability metrics and operational controls to measure maturity
Requirements
- Extensive site reliability engineering experience (7+ years) supporting critical infrastructure
- Strong cloud platform experience (7+ years) with leading providers, including Amazon Web Services and Microsoft Azure
- Proven leadership skills to set direction, influence stakeholders, and raise engineering standards
- Enterprise-scale release management experience delivering reliable software delivery processes
- Deep CI/CD expertise across pipelines, source control, and infrastructure automation
- Advanced Python programming skills to build tooling and automation
- Solid Kubernetes experience as a developer working with clusters and workloads
- Hands-on DevSecOps platform experience with GitLab preferred
- Strong analytical skills for complex problem-solving and strategic decision-making
- Upper-Intermediate English proficiency (B2, Upper-Intermediate)
Nice to have
- Amazon Web Services certification or equivalent hands-on expertise
- Microsoft Azure certification or equivalent hands-on expertise
- AI Architecture experience applied to reliability and operational decision-making
- AI Solution Engineering experience supporting platform automation and operations
- Gen AI Solutions Development exposure for operational tooling or incident workflows
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
About the Company
More jobs at EPAM Systems
-
Senior Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Lead Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Junior Site Reliability Engineer
Mexico · remote · Oct 7, 2026
-
Java Developer - Big Data
Poland · remote · Oct 7, 2026
-
Lead AI Software Engineer
Mexico · remote · Oct 7, 2026