Staff Software Engineer, HPC
✨ AI Summary
Zoox, an autonomous vehicle company, is hiring a Staff Software Engineer for its HPC infrastructure in Foster City, CA. The role focuses on modernizing a platform built on Ray.io, SLURM, and Kubernetes, with a strong emphasis on reliability, scalability, and developer velocity. Requires deep experience designing and operating large-scale distributed systems in production, proficiency in Python, and a track record of shipping reliable infrastructure. Bonus for ML workload exposure and scaling Kubernetes or SLURM beyond 10k nodes.
In this role, you will:
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
Lead multi-quarter, cross team initiatives that drive org-wide improvements
Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
Mentor junior engineers, guiding them through their career development
Qualifications
Experience designing and operating large-scale distributed systems in production
Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
Experience with Kubernetes, particularly for heterogeneous workloads
Experience with cloud infrastructure on AWS or similar providers
Track record of shipping and operating reliable, highly available scalable infrastructure
Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
Proficiency with Python
Bonus Qualifications
Exposure to machine learning workloads (training, inference, data generation)
Experience with Kubernetes or SLURM at scale (>10k+ nodes)
Experience with SLURM workload manager and advanced scheduling policies
Background in algorithmic optimization or operations research
Experience building developer tools and platforms used by large engineering organizations
About the Company
More jobs at zoox
-
Senior Software Engineer, Continuous Deployment
Foster City, CA · full_time · Sep 30, 2026
-
Software Engineer - Embedded Linux Operating Systems
Foster City, CA · full_time · Sep 30, 2026
-
Software Engineer - Planner GPU Compute
Boston, MA · full_time · Sep 28, 2026
-
Senior Software Engineer - Planner GPU Compute
Boston, MA · full_time · Sep 25, 2026
-
Sr. Electrical Integration Engineer - Diagnostics
Foster City, CA · full_time · Sep 25, 2026