Senior Site Reliability Engineer
Company:
EPAM Systems
Location:
Mexico, Colombia, Brazil, Argentina
Type:
remote
Posted:
Oct 7, 2026
Views:
0
We are seeking a Senior Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will drive improvements across CI/CD, cloud platforms, and operational readiness while solving high-impact production issues.
Responsibilities
- Design reliability strategies for critical infrastructure and services
- Build and maintain CI/CD pipelines and release workflows using GitLab
- Automate infrastructure operations and tooling with Python
- Operate cloud environments across AWS and Azure to meet availability goals
- Harden and standardize infrastructure practices across networking, security, IAM, and compute
- Coordinate incident response during on-call shifts and restore service quickly
- Implement Kubernetes-based deployment and operational patterns to improve stability
- Analyze reliability signals and root causes to prevent repeat incidents
- Improve DevOps processes and engineering capabilities to enable faster change
- Partner with stakeholders to prioritize resilience work over short-term fixes
Requirements
- 3+ years of site reliability engineering experience in cloud environments
- Proven leadership ability to influence reliability practices across teams
- Enterprise-scale release management experience supporting frequent deployments
- Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure
- Advanced Python programming skills for automation and tooling
- Solid Kubernetes skills using clusters as a developer
- Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms
- Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
- Strong analytical skills for complex problem solving under pressure
- Upper-Intermediate English proficiency (B2)
- Reliable on-call readiness to assess and resolve business-critical issues
Nice to have
- Amazon Web Services expertise, including design patterns for resilient systems
- Microsoft Azure expertise, including governance and operational best practices
- AI Architecture experience applied to platform reliability and automation
- AI Solution Engineering experience for production-grade AI-enabled operations
- Gen AI Solutions Development experience focused on operational use cases
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
About the Company
More jobs at EPAM Systems
-
Chief Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Lead Site Reliability Engineer
Mexico, Colombia, Brazil, Argentina · remote · Oct 7, 2026
-
Junior Site Reliability Engineer
Mexico · remote · Oct 7, 2026
-
Java Developer - Big Data
Poland · remote · Oct 7, 2026
-
Lead AI Software Engineer
Mexico · remote · Oct 7, 2026