Lead Data Software Engineer

Company: EPAM Systems
Location: Argentina, Brazil, Chile, Colombia, Mexico
Type: remote
Posted: Sep 16, 2026
Views: 0

We are seeking a Lead Data Software Engineer to design reusable data-sharing adapters across a cloud lakehouse and external analytics platforms with governed access. You will partner with engineers to deliver modular integrations and dependable pipelines, while showcasing AI-assisted development in day-to-day work.

Responsibilities

  • Architect a UniForm lakehouse write layer with dual-format metadata (Delta and Iceberg) to serve multiple consumers
  • Create and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational datasets
  • Deliver CDC pipelines with Kafka to enable real-time and near-real-time lakehouse updates
  • Establish dependency-aware bookkeeping and data lineage tracking patterns across data pipelines
  • Develop modular, version-controlled adapter code that can be reused for new data source integrations
  • Set up Snowflake external table definitions and enable governed access using Horizon catalog metadata
  • Build and certify Delta Sharing adapters to support zero-copy data sharing for Databricks consumers
  • Configure Delta Sharing endpoints, registrations, and sharing agreement management
  • Test end-to-end freshness, sharing latency, and SLA compliance for external data sharing flows
  • Create connector registry entries, RBAC, and tenant-scoped authorization for all data-out paths
  • Add metering hooks aligned with billing requirements for governed external data flows
  • Document integration patterns and operational procedures for reuse across additional data products

Requirements

  • Proven data software engineering experience (5+ years) using Python to build and maintain data pipelines
  • Hands-on experience with Google Cloud BigQuery, including advanced usage and performance optimization
  • Solid background with data lakehouse table formats, including Apache Iceberg and Delta Lake
  • Demonstrated experience integrating Databricks and applying governed data access patterns
  • Practical knowledge of Snowflake, including external table access to lakehouse data
  • Strong understanding of Kafka and CDC patterns for real-time and near-real-time ingestion
  • Deep architecture skills in data lake design, modular adapter development, and version control practices
  • Working proficiency with AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
  • Excellent documentation skills for reusable integration patterns and operational runbooks
  • Upper-Intermediate English proficiency (B2) for technical collaboration and written communication

Nice to have

  • Apache Spark experience to validate shared reads and pipeline patterns
  • Databricks Unity Catalog experience for governed metadata and access control
  • Delta Lake expertise in sharing and interoperability patterns
  • Gen AI Assisted Development experience with measurable workflow improvements
  • Snowflake Horizon Catalog experience for metadata governance and access controls

Benefits

  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

About the Company

Name: EPAM Systems

No detailed information available about this company.

More jobs at EPAM Systems