we are looking for a hands-on Data Platform Engineer to operate and evolve our enterprise Kafka ecosystem, which serves as the companys central data streaming platform. The role is critical to our data infrastructure and supports a major ongoing migration to the cloud, while covering production operations, troubleshooting, automation, maintenance, and reliability across high-volume, low-latency, 24/7 on-premises and cloud environments.
Responsibilities:
Operate and monitor Kafka environments across Confluent Cloud, Confluent Platform, and Kubernetes/CFK deployments, including availability, consumer lag, replication, throughput, latency, storage, and capacity.
Troubleshoot production incidents, participate in on-call, perform root-cause analysis, and improve resilience, failover, and disaster-recovery readiness.
Manage topics, partitions, retention, ACLs, consumer groups, Kafka Connect connectors, and Schema Registry.
Deploy, configure, upgrade, patch, and scale platform components and connectors while maintaining secure access and configuration consistency.
Automate platform operations and provisioning using scripting, infrastructure-as-code, and CI/CD.
Work with DevOps, SRE, Infrastructure, Security, Architecture, Data, and application teams to onboard workloads and continuously improve platform reliability and governance.
Requirements: 2+ years of hands-on experience with Apache Kafka with strong knowledge of topics, partitions, replication, producers, consumers, and consumer groups.
Hands-on experience with Confluent Platform and/or Confluent Cloud, including Kafka Connect and Schema Registry.
Experience supporting production systems on Linux, including monitoring, logs, networking, certificates, authentication, and access control.
Scripting and automation experience using Python, Bash, PowerShell, or similar tools, plus familiarity with Git and CI/CD.
Strong ownership, troubleshooting, and communication skills, with a willingness to participate in an on-call rotation.
Preferred Qualifications:
Experience with Microsoft Azure, Kubernetes / Confluent for Kubernetes (CFK), Helm, or Terraform.
Experience with the Kafka Connect ecosystem, including Debezium plugins, SMTs, Replicator, and source/sink connectors.
Familiarity with observability, SRE practices, SLIs/SLOs, incident management, and disaster recovery.
Experience in payments, fintech, banking, or another regulated environment; Confluent/Kafka certification is a plus.
This position is open to all candidates.