Lead Data Software Engineer
Company:
EPAM Systems
Location:
Argentina, Brazil, Chile, Colombia, Mexico
Type:
remote
Posted:
Sep 16, 2026
Views:
0
We are seeking a Lead Data Software Engineer to design reusable data-sharing adapters across a cloud lakehouse and external analytics platforms with governed access. You will partner with engineers to deliver modular integrations and dependable pipelines, while showcasing AI-assisted development in day-to-day work.
Responsibilities
- Architect a UniForm lakehouse write layer with dual-format metadata (Delta and Iceberg) to serve multiple consumers
- Create and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational datasets
- Deliver CDC pipelines with Kafka to enable real-time and near-real-time lakehouse updates
- Establish dependency-aware bookkeeping and data lineage tracking patterns across data pipelines
- Develop modular, version-controlled adapter code that can be reused for new data source integrations
- Set up Snowflake external table definitions and enable governed access using Horizon catalog metadata
- Build and certify Delta Sharing adapters to support zero-copy data sharing for Databricks consumers
- Configure Delta Sharing endpoints, registrations, and sharing agreement management
- Test end-to-end freshness, sharing latency, and SLA compliance for external data sharing flows
- Create connector registry entries, RBAC, and tenant-scoped authorization for all data-out paths
- Add metering hooks aligned with billing requirements for governed external data flows
- Document integration patterns and operational procedures for reuse across additional data products
Requirements
- Proven data software engineering experience (5+ years) using Python to build and maintain data pipelines
- Hands-on experience with Google Cloud BigQuery, including advanced usage and performance optimization
- Solid background with data lakehouse table formats, including Apache Iceberg and Delta Lake
- Demonstrated experience integrating Databricks and applying governed data access patterns
- Practical knowledge of Snowflake, including external table access to lakehouse data
- Strong understanding of Kafka and CDC patterns for real-time and near-real-time ingestion
- Deep architecture skills in data lake design, modular adapter development, and version control practices
- Working proficiency with AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
- Excellent documentation skills for reusable integration patterns and operational runbooks
- Upper-Intermediate English proficiency (B2) for technical collaboration and written communication
Nice to have
- Apache Spark experience to validate shared reads and pipeline patterns
- Databricks Unity Catalog experience for governed metadata and access control
- Delta Lake expertise in sharing and interoperability patterns
- Gen AI Assisted Development experience with measurable workflow improvements
- Snowflake Horizon Catalog experience for metadata governance and access controls
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
About the Company
More jobs at EPAM Systems
-
Senior Data Software Engineer
Colombia · remote · Sep 16, 2026
-
Senior DevOps Engineer
Argentina, Brazil, Chile, Colombia, Mexico · remote · Sep 16, 2026
-
Lead DevOps Engineer
Argentina, Brazil, Chile, Colombia, Mexico · remote · Sep 16, 2026
-
Senior AI Software Engineer
Singapore · remote · Sep 16, 2026
-
Senior Java Developer
Argentina, Brazil, Colombia · remote · Sep 16, 2026