You will be a senior engineer on the team that owns real-time data platform. The platform turns operational events (orders, driver locations, geofence transitions, shifts, MQTT session events) into live analytics, routing inputs, and warehoused data across Postgres/TimescaleDB, ClickHouse, and BigQuery. You will own the platform end-to-end: design, implementation, deployment, and production operation of the streaming jobs, CDC pipelines, Kafka Connect bridges, and downstream sinks that move this data.
What Will You Do
Designing, implementing, and operating stateful Apache Flink streaming pipelines: keyed state, windowing, watermarks, timers, side outputs, custom sources/sinks.
Designing, developing, deploying, and maintaining Change Data Capture (CDC) pipelines using Flink CDC or Debezium, moving operational database changes into the streaming platform with correct snapshot/incremental handling, schema evolution, and downstream idempotency.
Building and operating Kafka and Kafka Connect pipelines: topic and partition design, source/sink connectors in distributed mode, schema and converter management, and dead-letter routing.
Owning the path from Kafka → Flink → Postgres / TimescaleDB / ClickHouse / BigQuery, including schema design, idempotency strategy, batch tuning, and observability.
Diagnosing and fixing production issues: checkpoint failures, backpressure, state growth, sink slowness, autoscaler oscillation, restart loops.
Hardening the platform: delivery guarantees, watermarks, DLQ handling, schema migrations, alerting coverage.
Code reviews and mentoring mid-level engineers.
Contributing to the deployment side: Helm charts, ArgoCD applications, GKE configuration, Grafana dashboards, Prometheus alert rules.
Influencing direction: state backend choices, schema migrations, when a pipeline needs to be split or rebuilt.
Requirements: Bachelors degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.
5+ years of production Java experience, with fluency in modern Java (records, sealed types, switch expressions, streams). Comfortable with Maven, annotation processors (Lombok, MapStruct), and the JVM runtime model (heap behavior, GC, classloading, shaded jars).
Hands-on experience with streaming data platforms: Apache Flink, Kafka Streams, Spark Structured Streaming, or equivalents. You have built and operated stateful streaming pipelines in production.
Hands-on Apache Flink production experience: keyed state, RocksDB vs heap state backends, savepoints, the Flink Kubernetes Operator, autoscaler tuning.
Apache Kafka beyond producer/consumer basics: delivery guarantees, transactional producers, consumer groups, offset semantics, schema registries.
Kafka Connect: you have designed, deployed, and operated source/sink connectors in distributed mode. Comfortable with worker configuration, schema/converter handling, and dead-letter queue routing.
Production CDC experience: you have designed, developed, deployed, and maintained Change Data Capture pipelines using Flink CDC or Debezium. You understand how CDC interacts with the source databases WAL / binlog, the operational considerations (initial snapshot vs incremental, replica slots, DDL / schema evolution, backfill, downstream idempotency), and the failure modes (lag spikes, slot bloat, schema drift, dropped events).
This position is open to all candidates.