Machine Learning Engineer - Multimodal Intelligence
Imagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. Multifaceted, amazing people and inspiring, innovative technologies are the norm here. The people who work here have reinvented entire industries with all Apple Hardware products. The same passion for innovation that goes into our products also applies to our practices, strengthening our commitment to leave the world better than we found it. Join us in this truly exciting era of Artificial Intelligence to help deliver the next groundbreaking Apple products and experiences!
The Multimodal Intelligence team builds and ships the Computer Vision and Machine Learning systems behind Apple Intelligence — spanning data collection and curation, training and fine-tuning, evaluation, optimization, and on-device deployment. Our team has an established track record of delivering features that combine Apple's sensing hardware with large foundation models, including Visual Intelligence and the on-device foundation models that power text and visual understanding across iPhone, iPad, Mac, and Apple Vision Pro. We are focused on building experiences where a device can see, read, and reason about the world around it — privately, responsively, and on-device wherever possible.
We are looking for a Machine Learning Engineer to build the pipelines, infrastructure, and production systems that turn multimodal foundation models into shipping Apple Intelligence features. You will own end-to-end model delivery: building and scaling data curation and training pipelines, fine-tuning and optimizing large multimodal models for on-device and hybrid execution, standing up reproducible evaluation and regression testing for text and visual understanding, and hardening promising approaches into robust, maintainable production systems under real latency, memory, power, and privacy constraints.
You will work closely with modeling, platform, hardware, and product engineering teams across Apple — taking future hardware design and product needs into account as you make implementation decisions — and you will have the opportunity to collaborate broadly to deliver the best possible products.
Minimum Qualifications
- Experience in deep learning with demonstrated work in at least one area of multimodal systems (e.g., vision, language, video, audio, etc.)
- Proficiency in Python and in a modern deep learning framework such as PyTorch or JAX
- Experience with rapid prototyping, reproduction, and validation of research ideas
- Ability to work in a collaborative environment
- Ability to communicate the results of analyses in a clear and effective manner
- BS and a minimum of 3 years of relevant industry experience
Preferred Qualifications
- Master's or PhD, or equivalent practical experience, in Computer Science, Computer Vision, Machine Learning, or related technical field
- Deep expertise in multimodal foundation models, with a focus on practical applications
- Track record of translating research into practical applications either through published work or industry experience
- Strong applied research experience in at least one major area of model development (data curation, pre-training, fine-tuning, alignment, or evaluation), particularly as it applies to multimodal systems
- Experience with large-scale training pipelines, including working with large datasets and scaling models across distributed systems
- Experience bridging research ideas with production constraints
About the Company
More jobs at Apple
-
Software Engineering Manager (Watch Faces) - Watch Software
Cupertino · · Sep 15, 2026
-
Senior Machine Learning Engineer
Cupertino · · Sep 15, 2026
-
Design for Test Engineer
Austin · · Sep 15, 2026
-
Senior III-IV Integration Engineer
San Francisco Bay Area · · Sep 15, 2026
-
Sr. Software Engineer, Infrastructure Services (Data Plane)
San Francisco Bay Area · · Sep 15, 2026