Senior Data Engineer, Selling Partner Agentic Interfaces Data Products
Own the architecture of a data product that makes AI-powered commerce measurable, trustworthy, and improvable for millions of sellers worldwide. Join SP-AI Data Products as a founding technical leader who will define how an entire organization observes, governs, and learns from every AI-driven interaction at scale, built entirely on AWS large-scale data processing infrastructure.
We're at an inflection point. AI agents are replacing traditional seller workflows, generating 7.5 million interactions annually across 50+ internal product teams, with thousands of Developers building on the ecosystem. Every one of these innovations creates a measurement obligation, and right now there's no unified infrastructure to fulfill it. You'll change that. You're joining a small team transforming from traditional reporting into a production data product organization, and this role determines what that product becomes. You'll design and build on AWS services including Kinesis for real-time streaming ingestion, EMR and Spark for large-scale distributed data processing, Glue for ETL orchestration, Redshift and Athena for analytical workloads, S3 and Lake Formation for governed storage, and Lambda and Step Functions for event-driven pipeline automation.
What you'll own:
- Design and govern the canonical event schema: the unified measurement format that makes every AI interaction across every surface produce a comparable, correlated record. You define what gets measured and how.
- Own end-to-end data architecture across telemetry layers (user engagement, action execution, domain response), ensuring cross-layer correlation through session-level tracing
- Build and operate streaming and batch ingestion pipelines on AWS with production SLAs, serving real-time observability for leadership, risk teams, applied scientists, and product teams simultaneously
- Architect tenantized data products that enable 50+ teams to onboard once and receive self-serve metrics (adoption, quality, risk, impact) without building custom pipelines
- Drive data engineering standards across the organization: schema governance, naming conventions, data quality, operational excellence. You set the bar others build to.
- Make architectural trade-offs that balance short-term delivery against long-term scalability: tiered storage, build-vs-buy decisions, schema evolution across dozens of consumers
- Identify and resolve systemic architecture deficiencies, proposing and leading cross-team initiatives that unblock innovation for adjacent teams
- Decompose complex, ambiguous problems into parallel workstreams executable by you and others, then reassemble them into cohesive solutions
- Elevate the engineering team through mentorship and technical leadership. Your presence makes the team stronger, but the team doesn't require your presence to succeed.
Why you'll love this role:
- Founding-team impact: You're defining the architecture an entire organization builds on, not inheriting legacy systems
- Breadth of influence: Your schema decisions, quality standards, and architectural patterns are consumed by risk teams, scientists, product managers, and leadership across the business
- Technical depth at scale: Streaming infrastructure, cross-surface correlation, schema governance for 50+ consumers, production SLAs on AWS. Hard, consequential engineering.
- Career-defining scope: Cross-cutting schema design and multi-team domain onboarding at this scale is the kind of work that shapes what comes next
If you've built large-scale data products on AWS that other teams depend on, thrive in ambiguity, and want to define how an organization measures AI-driven commerce at scale, we'd love to talk.
Key job responsibilities
- Own the design, implementation, and evolution of large-scale data architecture on AWS (Kinesis, EMR, Spark, Glue, Redshift, Athena, S3, Lake Formation), providing system-wide technical guidance and ensuring all data products meet production-grade reliability, scalability, and security standards
- Write exemplary, production-quality code in Java, Scala, and Python to build streaming and batch data pipelines that process terabytes of telemetry data daily across distributed systems, with proper testing, lineage documentation, and quality controls. SQL proficiency is expected as a baseline.
- Design and maintain the canonical event schema and standardized, reusable data products following data mesh governance principles that eliminate metric inconsistencies and enable self-service analytics for 50+ domain teams
- Build and operate ETL and ELT pipelines using AWS-native services (Glue, EMR, Step Functions, Lambda) and JVM-based frameworks (Spark on Scala/Java) that transform raw telemetry into trusted, AI-ready datasets with defined SLAs for freshness and completeness
- Architect real-time observability infrastructure and analytical workloads (Kinesis, Redshift, Athena, QuickSight) that provide actionable insights on user engagement, operational health, risk detection, quality monitoring, and compliance tracking
- Collaborate with and influence peer Software Development Engineers, Applied Scientists, Product Managers, and Business Intelligence Engineers to define requirements, validate data quality, and align data product design with business and technical strategy
- Define and enforce data governance standards across the organization including schema governance, Fine-Grained Access Control, data classification, naming conventions, and lineage tracking across all data products
- Identify systemic data quality issues and architecture deficiencies, drive root-cause resolution, and automate manual processes to improve operational efficiency and eliminate recurring failures
- Document data products, architectural decisions, and design patterns clearly to ensure ease of use, extensibility, and maintainability by team members and downstream consumers across the organization
- Lead cross-team initiatives to enable GenAI capabilities through Amazon Quick Suite, defining how AI surfaces consume governed data while maintaining robust security, access controls, and auditability at scale
A day in the life
Your primary focus is owning the data architecture that powers AI-driven commerce for millions of sellers. Most of your day is spent writing Java, Scala, and Python code: building Spark pipelines on EMR, designing streaming ingestion on Kinesis, and optimizing how terabytes of interaction data flow through AWS infrastructure into governed, queryable products.
You'll be based in Bengaluru, working alongside a team split between Bengaluru and Seattle. Your mornings give you uninterrupted build time while Seattle is offline. You use this window for deep architectural work: writing design documents, pushing complex pipeline code, and leaving detailed code review feedback so your Seattle teammates wake up unblocked. During overlap hours, you join focused design reviews and cross-team discussions where your job is to bring clarity, shape schema decisions, and drive alignment across engineering teams consuming your data products.
Beyond the core technical work, you're the person domain teams come to when they need guidance on how to emit telemetry that conforms to the canonical schema. You're investigating why a streaming pipeline breached its SLA and fixing the root cause before it recurs. You're mentoring engineers through pull requests that teach long-term maintainability, not just correctness. Some days bring unexpected production issues that need fast diagnosis. Other days you're heads-down on a multi-week design for the next generation of tenantized data products that will serve 50+ teams.
The Bengaluru-Seattle structure means you operate with high autonomy. Decisions don't wait for handoffs across timezones. You own outcomes end-to-end, and the team trusts your judgment to move fast and get it right.
About the team
SP-AI Data Products is a small, high-autonomy team that builds the data infrastructure powering AI-driven commerce for millions of Amazon Selling Partners worldwide. Our mission is straightforward: make every AI agent interaction measurable, governable, and improvable through trusted, self-service data products. We exist so that product teams can ship with confidence, risk teams can detect abuse in real time, scientists can train better models, and leadership can see exactly how the business is performing without waiting for someone to compile a report.
We're a team of software engineers, data engineers, and business intelligence engineers split between Bengaluru and Seattle. We operate like a startup inside Amazon: small enough that every person's work is visible and consequential, but connected to a business growing at 7x year-over-year that serves 50+ domain teams, thousands third-party developers, and millions of sellers. Our customers are internal: the engineering teams building AI agents, the risk and trust teams protecting the ecosystem, the product managers tracking adoption, and the scientists improving model quality. We build for all of them simultaneously through governed, tenantized data products rather than one-off reports.
Our culture is built on ownership and craft. We write production-grade code in Java, Scala, and Python. We design schemas that dozens of teams consume. We hold each other to high engineering standards through rigorous code and design reviews, and we invest in mentorship because raising the bar across the team matters more than any individual contribution. We're transparent about what we are: a lean team executing an ambitious transformation from reactive reporting into a production data product organization built entirely on AWS. If you want to work somewhere your architectural decisions directly shape how an organization operates, where you'll have real autonomy to drive outcomes, and where the problems are genuinely hard and unsolved, this is the team.
About the Company
More jobs at Amazon
-
Security Engineer, AWS AppSec
Vancouver, British Columbia, CAN · · Aug 29, 2026
-
Senior Software Development Engineer, AWS SecDevOps
Seattle, Washington, USA · · Aug 29, 2026
-
Data Scientist II, Worldwide Design Engineering - Data Science
Bellevue, Washington, USA · Full-time · Aug 29, 2026
-
Solutions Architect
Oslo, Oslo, NOR · · Aug 28, 2026
-
Test Engineer, Amazon Leo
Kirkland, Washington, USA · · Aug 28, 2026