we are looking for a Senior DevOps Engineer.
As Senior DevOps Engineer, you will play a critical role in shaping our infrastructure, deployment pipelines, and operational foundations. This is a hands-on, high-ownership role with real influence over how the company builds, deploys, and operates its systems in production.
Youll be building infrastructure that supports work with frontier AI labs such as Google, Anthropic, and OpenAI, and meets the scale, security, and reliability standards this ecosystem requires. This includes architecting ephemeral, on-demand environments, and spinning up and tearing down multi-service deployments (VPCs, compute, serverless components, and more) as needed.
You will work closely with research, engineering, security, and leadership to design and maintain scalable infrastructure, improve system reliability, and enable fast, safe, and secure product delivery.
Key Responsibilities:
Own and evolve cloud infrastructure, with a strong focus on scalability, security, and reliability
Design, build, and maintain end-to-end CI/CD pipelines and deployment workflows
Lead and continuously improve containerized environments (Docker, Kubernetes)
Build and maintain Infrastructure as Code and automation frameworks (e.g., Terraform)
Own production readiness: availability, monitoring, logging, alerting, and incident response
Work closely with engineering and research teams to support development, testing, and production needs
Drive improvements in system performance, reliability, and security posture
Take part in - and often lead architectural decisions and long-term infrastructure planning
Act as a key operational owner in a fast-growing startup, setting best practices and standards
Requirements: 5+ years of hands-on experience as a DevOps / Infrastructure Engineer in production environments
Strong, hands-on experience with AWS (architecture, networking, security, and cost awareness)
Proven experience with containerization and orchestration (Docker, Kubernetes) in production
Hands-on experience designing and maintaining CI/CD pipelines (e.g., GitHub Actions or similar)
Strong experience with Infrastructure as Code and automation (e.g., Terraform)
Ability to take end-to-end ownership and operate independently in an early-stage, high-ambiguity environment
Nice to Have:
Experience working closely with research or security teams in production environments
Background supporting data-heavy, distributed, or high-scale systems
Familiarity with monitoring, observability, and alerting tools (e.g., Prometheus, Grafana, Datadog)
This position is open to all candidates.