We are looking for a Senior Backend Engineer to join our Data Export Team. In this role, you will lead the architecture and development of backend microservices that drive our entire distributed export ETL infrastructure. You will build the intelligent service layer responsible for orchestrating, executing, and scaling high-throughput batch and real-time data transformations.
RESPONSIBILITIES:
Service Ownership: Own backend microservices end-to-end-from architecture and API design to production deployment, observability, and continuous execution tuning.
ETL Engine & Performance: Drive performance, throughput, and sub-second reliability across our distributed ETL infrastructure, using data, profiling, and operational metrics to guide architectural decisions.
Pipeline Infrastructure: Build, orchestrate, and scale high-throughput batch and real-time streaming data pipelines supporting analytics, export workflows, and UI platform integrations.
Cross-Functional Collaboration: Partner closely with Product, Data Science, Analytics, and cross domain RnD teams to turn complex data movement requirements into seamless backend integrations.
Technical Excellence: Raise the engineering bar across the team through design reviews, robust code reviews, asynchronous system architecture discussions, and technical mentorship.
Requirements: Backend Expertise: 5+ years of backend engineering experience designing, operating, and scaling distributed microservices.
Data, ETLs & Databases: Hands-on experience architecting data-intensive systems, and high-throughput batch or real-time pipelines. Strong proficiency across diverse database technologies (relational, analytical warehouses like Postgres, Databricks, and Snowflake) with a track record of optimizing complex queries and data storage at scale.
Engineering Fundamentals: Strong foundation in core system design, data structures, clean API design, automated testing, and production observability (metrics, tracing, logging).
Cloud & Infrastructure: Excellent proficiency with cloud environments (AWS or GCP) and Infrastructure-as-Code (Terraform) for managing distributed execution environments.
AI Productivity: Demonstrated use of AI tools to work more efficiently, whether professionally or personally, and a curiosity for finding new ways to apply them. Comfort integrating generative AI into day-to-day workflows to boost productivity, quality, and output.
Education: BSc in Computer Science or equivalent practical experience.
NICE TO HAVE:
Python Services & Scripting: Strong proficiency with Python as a core backend language-building clean, modular microservice APIs (using frameworks like FastAPI or Flask), writing asynchronous tasks, and wrapping data processing logic into testable Python packages.
Distributed Processing with PySpark: Hands-on experience writing PySpark batch and streaming jobs to transform multi-terabyte datasets. Ability to optimize Spark execution plans, handle data skew, manage memory allocation, and minimize expensive shuffle operations across cluster nodes.
This position is open to all candidates.