Senior DevOps Engineer
We are seeking a Senior DevOps Engineer to join an AI Workbench Platform team focused on operationalizing domain foundation models. These are large models trained on log, seismic, drilling, and production data, analogous to general-purpose large language models but specialized for oil & gas subsurface and production domains. While these models exist today and are maturing, there is currently no platform to commercialize them, enable internal teams (geo units, business units, data scientists) to use them at scale, or allow external customers to interactively consume or fine-tune them. In this role, you will help design and build the infrastructure foundation that brings these powerful domain models to production.
Responsibilities
- Design, build, and maintain scalable infrastructure to support the AI Workbench Platform and its underlying domain foundation models
- Implement and manage Kubernetes clusters with multi-GPU scheduling capabilities to support large-scale model training and inference workloads
- Develop and maintain Infrastructure as Code using Terraform to provision cloud resources reliably and reproducibly
- Package, deploy, and manage applications using Helm and Kustomize across multiple environments
- Collaborate with data scientists, ML engineers, and business units to enable seamless model consumption and fine-tuning workflows at scale
- Ensure platform reliability, scalability, and security for both internal teams and external customers
- Optimize resource utilization and cost efficiency across GPU-intensive workloads
- Establish CI/CD pipelines and automation to accelerate platform delivery and model deployment
- Monitor system performance and troubleshoot production issues to maintain high availability
- Contribute to platform architecture decisions and best practices for MLOps at enterprise scale
Requirements
- 3+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles
- Expertise in Kubernetes with proven experience in multi-GPU scheduling for AI/ML workloads
- Proficiency in Terraform for Infrastructure as Code and cloud resource management
- Skills in Helm and Kustomize for Kubernetes application packaging and configuration management
- Background in building and operating production-grade platforms that support large-scale, distributed workloads
- Understanding of MLOps principles and infrastructure requirements for training and serving large foundation models
- Capability to collaborate cross-functionally with data scientists, ML engineers, and business stakeholders
- Excellent command of written and spoken English (B2+ level)
Nice to have
- Prior experience with LightOps infrastructure
- Familiarity with on-premises infrastructure environments
- Knowledge of High-Performance Computing (HPC) systems and workloads
Benefits
We connect like-minded people
- Delivering innovative solutions to industry leaders, making a global impact
- Enjoyable working environment, whether it is the vibrant office or the comfort of your home
- Opportunity to work abroad for up to two months per year
- Relocation opportunities within our offices in 55+ countries
- Corporate and social events
We invest in your growth
- Leadership development, career advising, soft skills and well-being programs
- Certifications, including GCP, Azure and AWS
- Unlimited access to EPAM's internal learning database
- Free English classes with certified teachers
We cover it all
- Participation in the Employee Stock Purchase Plan
- Monetary bonuses for engaging in the referral program
- Comprehensive medical & family care package
- Four trust days per year for personal needs
- Discounts for fitness clubs
- Benefits package (hotels, restaurants, stores and services)
We connect like-minded people
- Delivering innovative solutions to industry leaders, making a global impact
- Enjoyable working environment, whether it is the vibrant office or the comfort of your own home
- Opportunity to work abroad for up to two months per year
- Relocation opportunities within our offices in 55+ countries
- Corporate and social events
We invest in your growth
- Leadership development, career advising, soft skills and well-being programs
- Certifications, including GCP, Azure and AWS
- Unlimited access to EPAM's internal learning database
- Free English classes with certified teachers
We cover it all
- Participation in the Employee Stock Purchase Plan
- Monetary bonuses for engaging in the referral program
- Comprehensive medical & family care package
- Five trust days per year (sick leave without a medical certificate)
- Benefits package (sports activities, a variety of stores and services)
About the Company
More jobs at EPAM Systems
-
Unreal Engine UI Engineer (C++)
Argentina, Colombia, Mexico · remote · Sep 28, 2026
-
Senior Full Stack Engineer (.NET, React, Azure)
Ukraine · remote · Sep 28, 2026
-
Frontend Engineer (Medical Imaging Viewer)
Argentina, Brazil, Chile, Colombia, Mexico · remote · Sep 28, 2026
-
Senior/Lead Security Engineer - HIPPA
· · Sep 28, 2026
-
Senior Security Engineer
Argentina, Brazil, Chile, Colombia · remote · Sep 28, 2026