Design, build, and maintain robust pipelines feeding both DWH and production processes
Own pipeline orchestration in Apache Airflow end to end, including DAG design, monitoring, alerting, and production reliability
Design the data models behind our core domains and the contracts that downstream consumers depend on
Partner with business analysts, PMs, and business stakeholders to translate product and business questions into stable, qualified, and well-modeled data
Work with a variety of data sources including product events, application data, APIs, operational databases, logs, and third-party systems
Debug and solve data quality, pipeline, performance, and reliability issues
Use and create AI agents and LLM-based tooling to improve development, testing, documentation, and root cause analysis
Be part of the team and adopt effective ways of using AI tools in the engineering workflow.
Requirements: 3+ years in data engineering or a similar data-focused engineering role
Strong hands-on Python and SQL skills, with the ability to write clean, performant, and maintainable code
Proven track record of building and running code in large scale production environments
Deep, hands-on Apache Airflow experience, including ownership of production DAGs at scale
Deep data modeling expertise and the ability to design clean, scalable data assets that hold up as the business changes
Strong business and product orientation. You care about what the metric means, not just whether the job succeeded
Experience working with PMs, analysts, and business stakeholders to understand requirements and build the right data solutions
Practical experience using AI coding agents and LLM-based tools as part of your daily engineering workflow
Experience working with large and varied data sources including product events, application data, APIs, operational databases, logs, and third-party systems
Strong debugging and problem-solving skills around data quality, pipeline failures, performance, and reliability
Nice to have
Experience running data workloads in containerized or Kubernetes-based environments
Experience with large-scale or distributed processing using Spark / PySpark
Experience with query engines such as Trino
Experience with CI/CD and software engineering best practices
Experience building internal tooling or agents on top of LLMs.
This position is open to all candidates.