Principal Engineer, Apple Cloud Object Storage (ACOS), Storage Infrastructure

Company: Apple
Location: Seattle
Type:
Posted: Sep 16, 2026
Views: 0

✨ AI Summary

Apple is hiring a Principal Engineer for its exabyte-scale, geo-distributed object storage service (ACOS) that powers iCloud and other major products. The role focuses on storage efficiency, integrity, performance, and capacity planning to support high-density drives and AI workloads. Tech stack involves erasure coding, distributed systems, and heterogeneous storage media (CMR, HAMR, SMR). Requires 15+ years in software development with 10+ years in large-scale distributed storage systems and deep expertise in encoding layers and durability modeling.

ACOS is Apple's geo-distributed, exabyte-scale object storage service. It serves billions of requests per day for iCloud, Apple Music, Apple TV, Apple Maps, and Apple's internal data and analytics platforms. It provides 11-nines of durability and 4-nines of availability.

The ACOS Storage Infrastructure team owns how data is encoded, placed, verified, repaired, and reclaimed inside a storage cluster. This includes erasure and replicated encoding schemes, data placement and rebalancing optimization, compaction, scrubbing and repair, durability modeling, performance tiering, and the supply chain that keeps available capacity in front of customer demand. In practice, this team determines what storage costs, how it survives in the face of software and hardware failures, and how fast it is served.

This is a role for an engineer who has designed, shipped, and operated the data path of a storage system at exabyte scale.

The next big challenge here is to support the industry's transition to much denser drives. This means capacity per drive is growing far faster than IOPS per drive, and maintenance operations — sealing, compaction, scrubbing, repair, rebalancing, cross-region replication — consumes that increasingly scarce IOPS budget. This is made harder by the growing demand for high-performance storage tiers by Data and AI workloads

You will own the technical agenda for solving that, across four areas:

  • Storage Efficiency. Drive cost per usable byte down further: reduce replication factor without conceding the durability bar, reclaim storage overheads for compaction and other processes, and make placement maximize both space and IOPS utilization. Land support for denser and heterogeneous drive generations (CMR, HAMR, SMR) without letting IOPS-per-TB become the binding constraint on the fleet.
  • Storage Integrity. Guarantee 11-nines of durability in production in the face of software and hardware failures. Own detection and repair end to end: balance detection latency against the IOPS cost of scrubbers and ensure correctness of the encode/decode/repair/rebalance paths.
  • Storage Performance. Deliver differentiated performance tiers. That means improving tail latency and IOPS for the standard performance tier, while also introducing high-performance SSD-backed tiers for data and AI workloads.
  • Capacity Planning and Management. Make supply closely track customer demand. Own demand forecasting, drive-generation transitions, capacity on-boarding and decommissioning, quota and rate-limit commitments, and the release value when the system is oversubscribed.

Your influence will extend well beyond ACOS. You will be the person Apple's storage organization, its hardware engineering partners, and its executives rely on for judgment on encoding, durability, and fleet efficiency.

Minimum Qualifications

  • 15+ years in software development, with 10+ years designing, building, and operating large-scale distributed storage systems.
  • BS, MS, or PhD in Computer Science or a related field. A PhD in distributed systems or storage is a strong plus.
  • Proven track record of delivering storage systems at ~exabyte scale. Your past work should demonstrate depth in erasure coding, replication, data placement, repair and reconstruction, consistency models, and consensus protocols (e.g., Paxos, Raft).
  • Deep, hands-on expertise in the encoding layer specifically: Reed-Solomon and Local Reconstruction Codes, wide-stripe and geo-distributed coding schemes, and the ability to reason quantitatively about the resulting durability, availability, repair-bandwidth, and IOPS trade-offs.
  • Familiarity with the architecture of industry-leading storage systems (e.g., Google Colossus, Ceph/RADOS, Azure Blob Storage, Amazon S3, Meta f4/Tectonic, HDFS).
  • Experience with the modern drive landscape and its implications: QLC and dense flash, SMR, HAMR, heterogeneous fleets, declining IOPS per TB, and disaggregated storage and compute.
  • Experience designing storage for high-performance workloads, particularly AI/ML or large-scale analytics. Familiarity with the I/O patterns of training pipelines, data processing engines, and high-transaction databases.
  • Operational credibility: you have carried a pager for a storage system at scale, and your architectural judgment is shaped by having debugged data loss and availability incidents.
  • Exceptional ability to influence without authority, with a history of driving major technical decisions and aligning senior engineering leaders across a large organization.
  • Outstanding communication skills — able to distill trade-offs for an executive audience and to hold a deep technical debate with principal-level engineers.

Preferred Qualifications

  • Experience building out a credible, funded, multi-year path to a materially lower cost per usable byte.
  • In-depth understanding of maintenance overhead per usable byte trending down as drive density goes up — support denser drives.
  • Experience with durability guarantees measured and guaranteed in production.
  • Knowledge of performance tiers in production that support AI/ML and data teams.
  • Understanding of capacity that tracks demand within a tight band.

About the Company

Name: Apple

No detailed information available about this company.

More jobs at Apple