Required Senior Software Engineer, Infrastructure, Cloud
About the job
Our software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to our needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.
As a Platform Engineer on this team, you will build, operate, and scale the mission-critical cloud infrastructure that powers Chronicle Security Orchestration, Automation and Response (SOAR) at scale. This is a heavy infrastructure operations and automation role. You will take autonomous ownership of complex operational challenges-from administering advanced Kubernetes (GKE) clusters to building sophisticated observability pipelines-ensuring the highest level of Enterprise-grade reliability and security.
You will not simply implement others' ideas, you will be expected to scope operational challenges, write robust Infrastructure as Code (IaC), and transform manual toil into software-driven platforms in close partnership with product development teams.
Responsibilities
Own all aspects of your immediate infrastructure area, leading the design and maintenance of platform tooling, deployment pipelines, and cloud systems that enhance the reliability and performance of Chronicle SOAR.
Administer, operate, and troubleshoot advanced Kubernetes (GKE) clusters and onboard new infrastructure features at scale.
Develop and manage infrastructure configurations and automation policies using Go to ensure secure, consistent, and auditable management of GCP resources.
Design and implement sophisticated monitoring, logging, and tracing solutions (e.g., Prometheus, Grafana) and manage Service Level Objectives (SLOs)/Service Level Indicators (SLIs).
Act as a point of contact for cross-functional partners. Analyze past incidents and proactively develop software solutions to prevent recurrence. Build automation to reduce manual toil and accelerate incident resolution.
Requirements: Minimum qualifications:
Bachelors degree or equivalent practical experience.
5 years of experience with software development in one or more programming languages.
3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
Experience with Cloud compute platforms like Kubernetes and Cloud functions.
Preferred qualifications:
Master's degree or PhD in Computer Science, or a related technical field.
2 years of experience in a technical leadership role.
Experience designing and operating modern Observability stacks (Prometheus, Grafana, Datadog, ELK) and managing SLO/SLI frameworks.
Experience with container orchestration technologies (e.g., Kubernetes, GKE).
Understanding of Infrastructure as code tools (Terraform, Ansible, etc.).
This position is open to all candidates.