Lead Platform Engineer/Architect - HPC, Kubernetes

Company: EPAM Systems
Location: USA
Type: remote
Posted: Oct 8, 2026
Views: 0

Join a high-growth infrastructure team operating Kubernetes platforms across multiple cloud providers at massive scale. You'll build the systems that power thousands of GPUs, where your code and configurations directly protect thousands of GPU-hours from costly failures.

EPAM is where tech talent thrives—building groundbreaking solutions, advancing your skills through world-class learning platforms, and working alongside a global community of problem-solvers to make the future real.

Req# 1103823066

Responsibilities

  • Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth
  • Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, GCP, OCI, and additional cloud providers
  • Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads
  • Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews
  • Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale
  • Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery

Requirements

  • 10+ years of experience in infrastructure engineering, cloud platforms, or high-performance computing environments
  • Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
  • Advanced Python skills with a track record of building production-grade tools, not just scripts; experience with Go, Rust, or C++ is a strong plus
  • Daily proficiency in Terraform for writing and reviewing infrastructure as code
  • Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre
  • Strong site reliability engineering background with experience building monitoring, alerting, and incident response practices
  • CKA, CKS certificates are highly preferred

Benefits

  • Medical, Dental and Vision Insurance (Subsidized)
  • Health Savings Account
  • Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)
  • Short-Term and Long-Term Disability (Company Provided)
  • Life and AD&D Insurance (Company Provided)
  • Employee Assistance Program
  • Unlimited access to LinkedIn learning solutions
  • Matched 401(k) Retirement Savings Plan
  • Paid Time Off – the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)
  • Paid Holidays - nine (9) total per year
  • Legal Plan and Identity Theft Protection
  • Accident Insurance
  • Employee Discounts
  • Pet Insurance
  • Employee Stock Purchase Program
  • If otherwise eligible, participation in the discretionary annual bonus program
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program

About the Company

Name: EPAM Systems

No detailed information available about this company.

More jobs at EPAM Systems