Site Reliability Engineer
✨ AI Summary
EPAM Systems, a global IT services company, is hiring a Site Reliability Engineer in Argentina, Chile, or Colombia. The role focuses on cloud infrastructure, CI/CD, observability, and incident response. Tech stack includes Terraform/CloudFormation, AWS/Azure/GCP, Docker/Kubernetes, and Python/Bash/Go/Rust. Requires 2+ years of SRE/DevOps experience and advanced English (C1). Benefits include international projects, healthcare, paid time off, and learning resources.
We are seeking a proactive Site Reliability Engineer to strengthen the reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response—apply now to help reduce downtime and enable fast, safe releases.
Responsibilities
- Design and maintain cloud infrastructure using Infrastructure as Code practices
- Build and optimize CI/CD pipelines to automate deployments and operational workflows
- Implement logging, monitoring, and alerting to improve observability and reliability
- Define and track Service Level Objectives and Service Level Indicators with clear reporting
- Respond to production incidents and drive rapid service restoration
- Lead blameless post-mortems to identify root causes and prevent recurrence
- Partner with engineers to improve performance, scalability, and capacity planning
- Automate repetitive operational tasks to reduce toil and operational risk
- Harden production environments to improve resilience and safe change practices
Requirements
- 2+ years of experience in site reliability engineering, DevOps, or systems administration
- Hands-on experience with Infrastructure as Code using Terraform or CloudFormation
- Hands-on experience building and improving CI/CD pipelines for automated deployments
- Strong troubleshooting and incident response leadership skills in production environments
- Solid project skills to coordinate reliability work with software development teams
- Proficiency in scripting or programming with Python, Bash, Go, or Rust
- Cloud platform experience with AWS, Azure, or GCP
- Containerization experience with Docker and Kubernetes
- Deep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS)
- Strong communication skills with a reliability mindset focused on automation and reducing toil
- Advanced English proficiency (C1, Advanced)
Nice to have
- Experience with Prometheus, Grafana, or Datadog
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
About the Company
More jobs at EPAM Systems
-
Lead Data DevOps
Mexico, Colombia · remote · Aug 20, 2026
-
Senior Security & Test Engineer, A2A
Armenia, Georgia, Kazakhstan, Kyrgyzstan, Uzbekistan · remote · Aug 20, 2026
-
Senior Data DevOps
Mexico, Colombia · remote · Aug 20, 2026
-
Senior Machine Learning Engineer
Viet Nam · remote · Aug 19, 2026
-
Senior Full Stack Java Engineer (TypeScript)
Viet Nam · remote · Aug 19, 2026