Software Engineer- AI/ML, Amazon Neuron Training
✨ AI Summary
Amazon is hiring a Software Engineer for its Annapurna Labs team to build distributed training infrastructure for AWS Neuron and Trainium accelerators. The role focuses on large-scale pretraining and reinforcement learning using PyTorch, JAX, and high-performance computing techniques. Key responsibilities include implementing parallelism strategies, profiling workloads, and optimizing the ML compiler and runtime stack. No specific years of experience or salary range are listed, but the team emphasizes a strong work-life balance culture.
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.
The Distributed Training team enables the training of a wide range of models, from large-scale pretraining through post-training and reinforcement learning, on AWS's custom ML accelerators. As more customer workloads shift toward RLHF, PPO/GRPO, and other fine-tuning methods, we are building the distributed training infrastructure, parallelism techniques, numerics, and high-performance kernels that these methods depend on. As part of the broader Neuron organization, we work across frameworks, kernels, compiler, runtime, and collectives — a true hardware and software co-design in practice. We not only optimize current performance but also contribute to future architecture designs, since the gaps we characterize today become requirements for the next generation of Trainium.
We are looking for software engineers to help build and fine tune these distributed training solutions. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, where you will help shape the direction of AI acceleration technology.
Key job responsibilities
You’ll implement and tune components of our distributed training stack for large-scale training, post-training, and reinforcement learning workloads on the latest Trainium instances, working across PyTorch and the Neuron software stack. You'll contribute to parallelism strategies such as data, tensor, and pipeline parallelism and apply reduced-precision formats under the guidance of senior team members. You'll profile workloads to help determine whether a bottleneck sits in compute, memory, collectives, or host overhead, and work with compiler, runtime, and collectives engineers to help land the fix. You'll build and maintain internal tooling, benchmarks, and tests that keep the team's performance work reproducible, and take on increasing ownership as you grow in the role.
About the team
Inclusive Team Culture
Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.
Work/Life Balance
Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.
About the Company
More jobs at Amazon
-
Mission Operations Systems Engineer, Optical Inter-Satellite Link
Redmond, Washington, USA · · Sep 22, 2026
-
Delivery Consultant - Agentic AI Full Stack Developer, AWS Professional Services
Melbourne, Victoria, AUS · · Sep 22, 2026
-
Sr. Software Engineer- AI/ML, Amazon Neuron Training
Cupertino, California, USA · · Sep 22, 2026
-
Security Engineer , Region Services
Sydney, New South Wales, AUS · · Sep 22, 2026
-
Sr. Solutions Architect, Annapurna ML
Seattle, Washington, USA · · Sep 22, 2026