The Data-Nexus team, a central part of the R&D department , is seeking a versatile Senior Data Engineer to take part in leading data vision using a lake-house architecture, and building high-scale solutions for handling and consuming the different data assets of the company.
In this role youll go beyond traditional boundaries, taking on full product responsibilities - from conceptualization and architecture to design, maintenance, and continuous development. You will collaborate closely with stakeholders across the company to ensure our solutions effectively meet their needs.
The ideal candidate will have strong data engineering capabilities, along with software development skills including understanding of designing and managing AWS cloud infrastructure and strong knowledge of different data architectures, methods and tools.
Responsibilities:
Design and development of scalable, reliable, and secure data infrastructure and build ETL pipelines that handle diverse clinical data for research. Write production SQL, Spark jobs and craft schemas that evolve gracefully as research and production questions change.
Automate releases with CI/CD and Infrastructure as Code.
Optimise throughput, latency and cloud cost to meet research timelines at large scale.
Develop and maintain high-quality, scalable, and efficient code while fostering a culture of continuous improvement through regular retrospectives and knowledge sharing.
Take ownership of the full development lifecycle, including requirement gathering, design, implementation, testing, deployment, and ongoing maintenance.
Collaborate with cross-functional teams, including data scientists, data analysts, product managers, regulatory teams, and other developers, to drive the development of new tools and features that support mission.
Explore and adopt new technologies and frameworks that can enhance the capabilities of the Data-Nexus team and the overall data projects.
Requirements: BSc/MSc in Computer Science, Engineering, or a related field.
8+ years building data or backend systems in Python or a similar coding language, with a strong focus on cloud data infrastructure and scalable systems.
Strong command of SQL and a track record of pragmatic schema design.
Deep understanding of data modelling, ETLs and streaming technologies, including hands-on experience with tools like big-data tools like AWS Kinesis / Kafka, Spark.
Familiarity with modern lakehouse / warehouse tech, like Databricks, Delta Lake, Iceberg, Snowflake, Redshift.
Strong understanding of distributed systems, microservices architecture, containerization, and CI/CD pipelines.
Proficiency in Infrastructure as Code (IaC) tools, like Terraform or AWS CDK.
Experience with containerization and orchestration tools - Docker, Kubernetes (K8s).
Ability to take full product ownership from ideation to delivery, ensuring alignment with business objectives.
Familiarity with agile development methodologies.
Excellent communication skills with the ability to work effectively across teams and a customer-oriented mindset. Clear English communication is required.
Advantages:
Deep expertise in Databricks (Delta Lake, Unity Catalog, DLT).
Knowledge of the medical tech world, including EHR (HL7, FHIR) data and imaging data.
Prior work in regulated domains (healthcare, fintech, aviation), experience with enforcing and managing compliance requirements like HIPAA or FedRAMP.
This position is open to all candidates.