דרושים » מחשבים ורשתות » Senior Site Reliability Engineer (Cortex)

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
22/06/2026
משרה זו סומנה ע"י המעסיק כלא אקטואלית יותר
מיקום המשרה: תל אביב יפו
סוג משרה: משרה מלאה
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769584
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - youll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization.
Your Impact:
Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
Drive end-to-end troubleshooting across complex, distributed systems with high context switching
Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
Develop and maintain automation and tooling (primarily in Python)
Gain deep understanding of system architecture and interconnected services
Contribute to a culture of operational excellence in a high-scale, high-availability environment
On call responsibilities:
Daytime hours (12:00-20:00)
Occasional weekends and holidays (rotation-based).
Requirements:
Your experience:
5+ years of experience in SRE roles in production environments at scale
Strong hands-on experience with Kubernetes and Terraform
Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
Proficiency in Python for scripting and automation
Strong troubleshooting and problem-solving skills with a passion for incident handling
Ability to work in fast-paced environments with high context switching
Highly responsive, proactive, and ownership-driven
Strong collaboration and communication skills
Curious mindset and eagerness to learn.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779543
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
17/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a DevOps Engineer to join our engineering team. The ideal candidate has strong hands-on experience with cloud-native infrastructure, a GitOps mindset, and the ability to independently research and learn new tools and techniques in a fast-moving, large-scale Kubernetes environment.
Responsibilities
Design, operate, and troubleshoot Kubernetes clusters (EKS/AKS) at scale
Manage application delivery and infrastructure using GitOps tools (ArgoCD/Flux)
Build and maintain Helm charts and Kustomize overlays for multi-environment deployments
Provision and manage cloud infrastructure using Crossplane and/or Terraform
Own and optimize CI/CD pipelines (GitHub Actions, GitLab CI) for build, test, and deployment workflows
Maintain and extend observability stacks (Prometheus, Grafana, alerting rules, dashboards)
Write automation scripts and tooling in Python, Go, or Bash to streamline operations
Support AWS infrastructure across multiple accounts/regions (networking, IAM, compute, storage)
Participate in on-call rotation, troubleshoot production incidents, and drive root-cause analysis
Collaborate with platform, security, and application engineering teams on infrastructure design and reliability improvements.
Requirements:
2-3 years of experience in DevOps, SRE, platform engineering, or a related role
Solid working knowledge of Kubernetes and Helm in production environments
Experience with Crossplane and/or Terraform for infrastructure as code
Proficiency with AWS services (EKS, IAM, VPC, networking, compute)
Hands-on experience with GitOps tools such as ArgoCD or Flux
Scripting ability in Python, Go, or Bash for automation and tooling
Experience with Prometheus and Grafana for monitoring and alerting
Practical experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI)
Strong self-learning ability, capable of independently researching unfamiliar technologies, reading documentation, and applying findings without handholding
Solid troubleshooting skills across networking, compute, and distributed systems
Nice to Have
Programming proficiency in Go or Python (beyond scripting)
Experience with service mesh technologies (Istio or similar)
Exposure to multi-cloud environments.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8785730
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
29/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for an experienced SRE Team Lead to drive the reliability, observability, and automation practices across our private cloud infrastructure and operations. In this role, you will lead a team of site reliability engineers, own the engineering roadmap for monitoring and automation, and act as a key liaison between development, operations, and platform teams. You bring at least 3-4 years of hands-on people management experience and a deep technical background in SRE or DevOps disciplines.


What will you do?

Leadership & Team Management

Lead, mentor, and grow a team of SREs, providing technical direction, career development guidance, and day-to-day management.

Own the team roadmap for reliability, observability, and automation initiatives - prioritizing work, removing blockers, and driving delivery.

Conduct regular 1:1s, performance reviews, and hiring processes to build and sustain a high-performing team.

Foster a culture of operational excellence, blameless post-mortems, and continuous improvement.

Act as an escalation point for complex incidents and reliability issues, leading post-incident reviews and ensuring follow-through on action items.


Automation & Infrastructure

Design, develop, and maintain automation tools to support infrastructure and operations teams at scale.

Manage pipelines and infrastructure workflows using Jenkins, Ansible, Python, and Bash.

Drive the adoption of infrastructure-as-code practices across the organization.

Collaborate with system engineers to improve scalability, performance, and fault tolerance of critical systems.


Monitoring & Observability

Build and extend monitoring and alerting systems using Grafana, the ELK (Elastic) stack, Zabbix, and custom scripts.

Implement and enforce observability best practices to ensure full visibility into systems, applications, and infrastructure.

Define and track SLIs, SLOs, and error budgets across key services.

Partner with development teams to embed observability earlier in the software development lifecycle.


Database & Platform Support

Support monitoring and infrastructure integration for databases including MongoDB and PostgreSQL.

Maintain documentation and champion knowledge sharing around automation, monitoring, and reliability practices.
Requirements:
Experience & Leadership:

3-4+ years of experience in a people management or team lead capacity within SRE, DevOps, or infrastructure engineering.

5-8+ years of overall experience in SRE, DevOps, or infrastructure automation roles.

Proven track record of building, coaching, and retaining high-performing engineering teams.

Experience owning an engineering roadmap and driving cross-functional reliability initiatives.


Technical Skills :

Strong scripting skills in Python and Bash; comfortable building and maintaining production-grade automation.

Hands-on experience with infrastructure automation tools, particularly Ansible.

Solid experience with monitoring and observability platforms - ELK stack, Grafana, and Zabbix.

Good understanding of CI/CD pipelines and related tooling, including Jenkins.

Familiarity with managing and monitoring MongoDB and PostgreSQL in a production environment.

Comfortable working in Linux-based environments.

Excellent problem-solving skills and strong written and verbal communication.


Ability to support the following:

Experience with cloud providers - AWS, GCP, or Azure.

Exposure to containerization technologies such as Docker and Kubernetes.

Familiarity with infrastructure provisioning using Terraform.

Experience introducing SRE practices (SLOs, error budgets, chaos engineering) at an organizational level.

Exposure and experience with migrating/ building AI tools to improve process.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8760168
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced DevOps Engineer to join our Engineering team and play a key role in building and operating our cloud-native platform. The ideal candidate will have hands-on experience managing production environments at scale, driving cloud transformation initiatives, and supporting the of enterprise systems from on-premises deployments to modern SaaS and cloud-native architectures. You will be responsible for designing, automating, and maintaining scalable infrastructure, CI/CD pipelines, and deployment processes that support both our core products and emerging AI-driven capabilities. Working closely with Engineering, QA, Product, and AI teams, you will help ensure the reliability, security, and performance of our services while driving operational excellence and continuous improvement across our technology stack.
Responsibilities:
Design, implement, and maintain CI/CD pipelines.
Manage and optimize cloud infrastructure across AWS, Azure, and/or GCP.
Develop and maintain Infrastructure as Code using Terraform.
Manage Kubernetes-based environments and GitOps deployment workflows using Argo CD and Kustomize.
Lead and support the migration of enterprise applications and infrastructure from on-premises
environments to scalable SaaS and cloud-native architectures.
Establish, maintain, and continuously improve production environments, ensuring high availability, security, scalability, and operational excellence.
Demonstrate strong production ownership, including incident management, root cause analysis, capacity planning, and performance optimization.
Collaborate with Engineering, QA, Product, and AI teams.
Support the deployment, operation, and monitoring of AI and Generative AI services.
Build and maintain monitoring, logging, and alerting systems.
Troubleshoot and resolve infrastructure, deployment, and production issues.
Requirements:
5+ years of experience as a DevOps Engineer or similar
infrastructure-focused role.
Hands-on experience with Azure, Aws, or GCP.
Experience with CI/CD tools such as Jenkins, GitHub Actions, or similar platforms.
Strong knowledge of Terraform and Infrastructure as Code practices.
Experience with Docker, Kubernetes, Argo CD, and Kustomize.
Experience designing and operating production-grade Kubernetes saas platforms or enterprise environments.
Experience with monitoring and observability tools such as Prometheus, Grafana, and ELK.
Strong troubleshooting, analytical, and communication skills.
B.Sc. in Computer Science, Computer Engineering, Information Systems, or a related field (or equivalent practical experience).
Nice to have:
Experience supporting AI, Machine Learning, or Generative AI workloads, including familiarity with MLOps concepts, AI deployment platforms, or cloud-based AI services.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796409
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
05/08/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
Your Career:
Own and continuously improve AWS production infrastructure for scalability, reliability, security, performance, and cost.
Run and evolve Kubernetes environments that support fast, safe product delivery.
Drive developer velocity and production safety through better CI/CD pipelines, release workflows, deployment visibility, and GitOps practices.
Improve observability and incident response - reduce alert noise and raise signal quality.
Design and ship AI-assisted operational agents that change how engineers work - triaging monitoring alerts, summarizing incidents, proposing fixes, onboarding new services, answering questions and requests. This is a core part of the role, not a side project.
Build automation and self-service tooling that removes manual work from provisioning, monitoring, incident response, and developer workflows.
Analyze operational data across incidents, alerts, deployments, infra health, and cost to find reliability gaps, inefficiencies, and automation opportunities.
Partner with engineering, security, product, and leadership to remove bottlenecks and support safe production growth.
Evaluate and introduce new tools and AI-assisted approaches, balancing innovation with reliability, cost, and operational simplicity.
Your Impact:
You'll help scale production systems, improve deployment velocity and reliability, reduce operational overhead, and build automation and AI workflows that help engineering teams move faster and operate more efficiently.
This role is a strong fit for someone who enjoys ownership, collaboration, and operational innovation.
Requirements:
Your Experience:
4+ years operating production infrastructure in AWS.
Deep hands-on experience with Kubernetes, Helm, ArgoCD, Terraform, and CI/CD.
Strong experience with observability and alerting in Datadog or comparable platforms.
Solid grounding in Linux, networking, cloud security, and reliability best practices.
Strong scripting skills in Python and Bash.
Proven ability to own platform projects end-to-end, from design through production operation and ongoing improvement.
Strong troubleshooting across distributed systems, Kubernetes, CI/CD, and live incidents.
Collaborative mindset - comfortable working across engineering, security, product, and leadership.
Comfort in a fast-paced, high-ownership environment where priorities shift but production quality doesn't.
Genuine interest in applying AI, automation, and intelligent workflows to operational work.
Key qualities
Ownership-driven - You take responsibility for the systems you build and operate, from design through production support and continuous improvement.
Collaboration - You work effectively across engineering, security, product, and leadership to align priorities and drive shared outcomes.
Developer experience focus - You are committed to reducing friction for engineering teams through thoughtful automation, self-service workflows, and reliable internal tooling.
Innovation balanced with pragmatism - You actively explore new approaches, particularly in AI-assisted operations, while weighing them against reliability, maintainability, and operational simplicity.
Security mindset - You design and build with least privilege, auditability, and production safety as foundational principles rather than afterthoughts.
Clear communication - You articulate infrastructure, reliability, cost, and security tradeoffs precisely to both technical and non-technical stakeholders.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8769987
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
30/07/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior DevOps Engineer , youll play a critical role in scaling and evolving our infrastructure as we grow. Youll work alongside experienced engineers to drive automation, optimize cloud operations, and ensure our systems are secure, resilient, and high-performing. This role is perfect for a hands-on engineer who thrives in fast-paced environments and wants to shape infrastructure practices at a product-focused security startup.



What Youll Do

Own and Evolve Infrastructure: Design, deploy, and operate infrastructure on AWS using Kubernetes to orchestrate containerized services.
Build Tools and Automate Everything: Streamline internal workflows with smart tooling, configuration management, and monitoring systems.
Lead CI/CD Improvements: Define and refine deployment pipelines to support rapid, reliable releases and cross-team agility.
Strengthen Edge Security: Manage WAFs, gateways, and load balancers to balance strong protection with great user experience.
Drive Reliability: Lead incident detection, troubleshooting, and automated recovery-improving uptime and system robustness continuously.
Requirements:
6+ years in DevOps or infrastructure engineering roles.
Deep experience with Kubernetes, Helm, containerization, and cloud.
Strong background in AWS (bonus: experience with GCP or Azure).
Proficiency in CI/CD platforms (GitHub Actions, GitLab, Jenkins, etc.).
Solid understanding of networking, distributed systems, and both SQL and NoSQL databases.
Linux power user with strong scripting skills (e.g., Bash, Python).
Hands-on with infrastructure-as-code tools like Terraform.
Great communicator with a security mindset and a bias for automation.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8762093
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior Site Reliability Engineer at our company 911, you'll own the infrastructure that keeps our platform reliable, scalable, and secure - work that directly supports mission-critical 911 systems used by public safety agencies. You'll drive infrastructure-as-code practices across AWS, lead observability efforts through Datadog, and bring modern AI-assisted engineering approaches into how the team builds and operates.
What You'll Do
Own and evolve AWS infrastructure using Infrastructure-as-Code (Terraform / Terragrunt)
Architect and scale AWS environments
Deploy, scale, and manage containerized workloads using Kubernetes and Docker; contribute to HA/DR architecture and platform strategy
Lead deployment and release processes using Argo (reference JD also names Bitbucket, Jenkins as part of the CI/CD toolset).
Define and enforce SLOs, SLIs, and error budgets; drive toil reduction across the platform
Drive full utilization of Datadog for monitoring, dashboards, and alerting across the platform (reference JD also names Prometheus, Grafana as potential observability tooling)
Build self-service internal developer platforms that empower teams to ship faster.
Take end-to-end ownership of infrastructure projects - define success criteria, execute, and measure outcomes.
Partner cross-functionally with engineering teams (e.g., network engineering, Dev owners) on long-term technical planning.
Bring AI-assisted engineering practices (e.g., Claude, MCP integrations) into daily workflows to improve team efficiency
Document work and provide cross-training to peers.
Resolve JIRA tickets across Cloud, CI/CD, deployments, and monitoring.
Requirements:
At least 6 years of experience as a DevOps/SRE engineer in a cloud environment
Hands-on, production-level AWS experience.
Hands-on production experience with Kubernetes and containerization
Experience with Terraform/Terragrunt (or similar Infrastructure-as-Code tools) - required
Strong Bash scripting skills
Deep understanding of SRE principles: SLOs, SLIs, error budgets, toil reduction, blameless post-mortems
Strong incident management / on-call experience
Solid understanding of APIs, microservices, and distributed systems
Demonstrated experience leading a project end-to-end, from defining success criteria through delivery and measurement
Communicates effectively across teams and can drive long-term technical planning
Practical experience with AI-assisted engineering tools (e.g., Claude, Cursor) and MCP-style integrations is a strong plus
Experience building AI/ML infrastructure (model deployment, inference pipelines)-plus.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8796929
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior SRE Engineer, you will be a key player in ensuring the reliability, scalability, and performance of our critical IT infrastructure. You will leverage SRE principles and an automation-first mindset to build and maintain resilient hybrid cloud environments. This role is ideal for a candidate who thrives in a fast-paced, innovative setting and is passionate about solving complex challenges with cutting-edge technology.
Key Responsibilities
Provision, configure, and support resilient hybrid cloud deployment architectures using an Infrastructure-as-Code framework.
Proactively collaborate with development teams to ensure new applications are production-ready, scalable, and reliable from inception.
Develop and maintain tools and frameworks to automate operational tasks, including deployment, monitoring, and recovery.
Conduct thorough root cause analysis of production issues and implement preventative measures to improve system resilience, demonstrating strong problem-solving skills.
Manage CI/CD platforms, Linux infrastructure, and contribute to capacity planning and operational runbooks.
Design and implement proactive service monitoring, alerting, and trend analysis to maintain service availability and performance SLAs.
Participate in an on-call rotation to support critical applications and services, responding to and resolving incidents efficiently.
Contribute to comprehensive documentation related to infrastructure design, deployment, and operational procedures.
Requirements:
Your Expereience:
6+ years of Devops engineering experience on mission-critical, enterprise-level systems in a hybrid (both cloud and on-prem) environment.
3+ years of hands-on experience with cloud environments, preferably Google Cloud Platform (GCP).
Expertise in configuration management and Infrastructure-as-Code using frameworks such as Terraform and Ansible.
Strong programming/scripting knowledge in languages like Python, Bash, or Go for infrastructure automation.
Demonstrated experience with CI/CD pipelines (e.g., GitHub, Jenkins, Artifactory) and a strong foundation in Linux/Unix administration.
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
Preferred Qualifications
Experience with containerization and orchestration technologies, particularly Kubernetes.
Hands-on experience with monitoring and observability tools such as Datadog, Grafana, or Prometheus.
Understanding of networking principles including firewalls, load balancers, and complex network designs.
A curious and positive mindset with a passion for applied learning and challenging existing processes for continuous improvement.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
04/08/2026
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
Were looking for a highly skilled Infrastructure Engineer to join our team and own the scaling, management, and automation of our platforms distributed environments. If youre excited about building high-scale distributed systems and solving deep DevOps and infrastructure challenges, lets talk.

What Youll Do:

Own and scale infrastructure - Design, build, and optimize the backbone of our observability platform, ensuring seamless deployment across hundreds of distributed environments.
Solve complex scalability challenges - Tackle unique problems in multi-cluster Kubernetes environments, multi-cloud setups, and high-ingestion observability pipelines.
Manage data at scale - Build and optimize configurable data pipelines, ensuring efficient ingestion, storage, and querying of large volumes of observability data with resilience, consistency, and analytical capabilities.
Automate everything - Develop infrastructure as code, improve CI/CD processes, and automate environment provisioning for reliability and efficiency.
Enhance system reliability - Design robust monitoring, alerting, and self-healing mechanisms for a high-scale production environment.
Collaborate cross-functionally - Work closely with backend engineers, product teams, and customers to design scalable, developer-friendly infrastructure.
Adopt and implement cutting-edge technologies - Continuously evaluate and introduce new tools and frameworks to improve scalability, performance, and cost efficiency.
Improve deployment efficiency - Optimize Helm charts, Kubernetes operators, and Terraform configurations to streamline environment creation and lifecycle management.
Requirements:
5+ years of experience in DevOps, SRE, or Infrastructure Engineering roles.
Strong expertise in Kubernetes, Terraform, Helm, and cloud environments (AWS, GCP, or Azure).
Experience with scalable observability stacks (e.g., ClickHouse, VictoriaMetrics, OpenTelemetry) is a huge plus.
Deep understanding of distributed systems, networking, and containerized workloads.
Proficiency in at least one programming language (Go, Python, or similar) for automation and tooling
Passion for building scalable, reliable, and efficient infrastructure.
A problem-solving mindset with the ability to tackle complex technical challenges independently.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8768205
סגור
שירות זה פתוח ללקוחות VIP בלבד