דרושים » תוכנה » Software Development Engineer (AWS ML), Machine Learning Israel (MLIL) - FLOW sub-team (Fleet Lifecy

משרות על המפה
 
בדיקת קורות חיים
VIP
הפוך ללקוח VIP
רגע, משהו חסר!
נשאר לך להשלים רק עוד פרט אחד:
 
שירות זה פתוח ללקוחות VIP בלבד
AllJObs VIP
כל החברות >
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL FLOW team is looking for a Software Development Engineer to design and build automation, tooling, and monitoring systems for our next-generation ML accelerator servers. We build production software to validate, initialize, monitor, and qualify these servers - from first silicon through fleet-scale deployment. Our work spans hardware diagnostics, manufacturing test automation, CI/CD pipelines, operational dashboards, and data-driven fleet health monitoring.

Key job responsibilities
Design and develop software infrastructure - automation frameworks, deployment systems, and test orchestration platforms that run at scale across manufacturing and production environments.
Work cross-functionally with Hardware, Manufacturing, and EC2 teams to automate coordinated software delivery and qualification workflows.
Debug and root-cause hardware/software interaction failures using systematic data analysis and automation-assisted triage.
Build and own CI/CD pipelines end-to-end: from code commit through build, test, deploy, and production validation - driving fast, reliable software delivery for hardware teams.
Create data pipelines and analytics systems (ETL, aggregation, real-time reporting) that transform raw hardware test results into actionable engineering insights.
Develop monitoring dashboards, alerting systems, and data visualization tools for fleet health, yield tracking, and performance benchmarking.
Own features end-to-end: from design through implementation, testing, deployment, and operational excellence.
Requirements:
Basic Qualifications
- 3+ years of software development engineer or related occupational experience.
- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent.
- Experience using Linux, demonstrating proficiency with associated tools or languages.
- Can work proactively and independently, meet deadlines, and deliver on projects and tasks.
- Knowledge of software engineering best practices across the development life cycle, including agile methodologies, coding standards, code reviews, source management, build processes, testing, and operations.

Preferred Qualifications
- Experience building monitoring dashboards and data visualization (Grafana, CloudWatch, QuickSight, or similar).
- Experience with data pipelines, ETL, or analytics (S3, Athena, Spark, or similar).
- Experience with systems programming languages (C, C++, Rust).
- Familiarity with computer architecture concepts (PCIe, memory hierarchy, power management).
- Advantage: experience with hardware bring-up, ASIC/FPGA validation, or manufacturing test development.
This position is open to all candidates.
 
Hide
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774254
סגור
שירות זה פתוח ללקוחות VIP בלבד
משרות דומות שיכולות לעניין אותך
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
22/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Senior Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need a senior IC who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures (decoder-only, mixture-of-experts) end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
- Mentor engineers, drive design reviews, and raise the engineering bar across the team.
Requirements:
Basic Qualifications
- Bachelor's degree in computer science or equivalent
- 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
- Knowledge of computer architecture, operating systems, and parallel computing
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with CUDA kernels or ML/low-level kernels.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Experience with hardware simulation environments and model validation workflows.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8749429
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
3 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications:
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications:
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8776998
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration.
Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models.
This is a ground-up effort with rapidly evolving hardware and software. We need an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities
- Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
- Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
- Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
- Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
- Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
- Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
- Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
Requirements:
Basic Qualifications
- Bachelor's degree or equivalent.
- 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- Knowledge of computer architecture, operating systems, and parallel computing.
- Strong proficiency in C/C++.
- Strong Linux systems knowledge.
- Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
- Proven track record of owning and delivering complex software features end-to-end.

Preferred Qualifications
- Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware.
- Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
- Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming.
- Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8774268
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and Perplexity kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use internal our CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Requirements:
Basic Qualifications
- 5+ years of non-internship professional software development experience.
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
- 3+ years as a mentor, tech lead or leading engineering teams.
- 3+years experience in SW/HW Co-Design.

Preferred Qualifications
- Bachelor's degree in computer science or equivalent.
- Experience creating automated dashboards and visualization (such as Grafana).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748498
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking a Staff Systems Software Engineer to design and build the foundational infrastructure that powers products used by billions of people worldwide. In this role, you will architect and implement large-scale distributed systems, low-level platform components, and high-performance services that underpin our core product stack. You will drive technical strategy across system reliability, performance, and scalability, partnering closely with product, infrastructure, and data engineering teams to deliver systems that operate at global scale with high availability and efficiency.
Software Engineer, Systems Responsibilities
Architect and implement large-scale distributed systems and platform services that support high-throughput, low-latency workloads across our product infrastructure
Lead the technical design of systems components including storage layers, compute pipelines, networking abstractions, and service orchestration frameworks
Identify and resolve systemic performance bottlenecks through instrumentation, profiling, and targeted optimization across the full systems stack
Define and enforce service level objectives for owned systems, building dashboards, alerting pipelines, and runbooks to reduce mean time to mitigation during incidents
Drive reliability improvements by reducing failure surface, designing resilient rollout strategies, and leading regular resiliency and overload testing exercises
Collaborate with cross-functional partners across product engineering, infrastructure, and data science to align system architecture with evolving product and business requirements
Establish and evolve coding standards, architectural patterns, and engineering best practices for systems development across the broader organization
Leverage AI-assisted development workflows to accelerate design iteration, code generation, and systems analysis, applying sound judgment on when to rely on AI versus deep systems expertise
Mentor other engineers on systems design principles, debugging methodologies, and production operations, and contribute to onboarding programs for new team members
Lead incident retrospectives, identify root causes of complex production failures, and drive implementation of systemic improvements to prevent recurrence
Requirements:
Minimum Qualifications
8+ years of experience designing and implementing large-scale distributed systems, platform infrastructure, or systems software in production environments
Experience leading major technical initiatives end-to-end, including architecture design, cross-team coordination, staged rollout, and post-launch reliability ownership
Experience debugging complex, non-reproducible systems issues including concurrency bugs, memory management failures, and distributed consistency problems
Experience defining service level objectives, building observability infrastructure, and driving reliability improvements across production systems
Experience communicating technical architecture decisions and trade-offs in writing to both engineering and non-engineering stakeholders
Preferred Qualifications
Experience with systems programming languages such as C, C++, or Rust in the context of high-performance or low-latency infrastructure
Experience building or improving developer tooling, automation frameworks, or internal platforms that measurably improve engineering efficiency across teams
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews).
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8777005
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high-performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems-ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.
Key Responsibilities
Scale Distributed Training: Design and optimize infrastructure for training and fine-tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Requirements:
Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands-On) working with cloud environments.
Model Lifecycle Engineering: Hands-on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine-tuning, optimization, and high-throughput production serving.
Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron-LM, or cloud-native training infrastructure).
Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
Python Expertise: Expert-level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
AI Tooling & Development: Proficient in leveraging day-to-day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
Strong GCP ecosystem experience.
Background in data science or deep learning workflows.
Cybersecurity domain knowledge.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8781454
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
5 ימים
Job Type: Full Time
We are looking for an outstanding Senior Networking Software Architect to join the NIC/DPU Software and Firmware Architecture group. In this role, you will help define the next generation of our datacenter and AI networking platforms, with focus on DPU management, QoS, performance, telemetry, and software architecture across stacks. You will work closely with hardware designers, firmware/kernel driver teams, system engineers, validation, product management and customers. The role spans early architecture definition, pre-silicon design, bring-up, and production readiness for large-scale AI and cloud datacenter deployments.

What Youll Be Doing:

Own software and system architecture for next-generation DPU management, QoS, performance, telemetry, and observability features.

Define end-to-end control and management flows across DOCA, host drivers, embedded firmware, BMC, management controllers and external management systems.

Specify telemetry and observability requirements, including counters, logs, traces, events, health monitoring, debug data, and streaming telemetry.

Define management interfaces and APIs for configuration, provisioning, lifecycle operations, diagnostics, and field serviceability.

Write clear architecture specifications, interface definitions, flow diagrams, and design documents for software, firmware, and system teams.

Partner with R&D teams to translate high-level architecture into implementable designs and guide features through development, validation, silicon bring-up, and production.

Analyze system performance bottlenecks, interoperability issues, telemetry gaps, and customer-reported issues, then feed learnings into future architecture.

Collaborate with system and cluster architects to ensure NIC/DPU features fit end-to-end AI datacenter and cloud networking designs.
Requirements:
What We Need To See:

B.Sc. or M.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or equivalent experience.

9+ years of experience in networking, system software, embedded software, firmware, or datacenter infrastructure.

Proven experience in software architecture, or technical leadership roles.

Deep understanding of networking concepts and protocols such as Ethernet, TCP/IP, RDMA/RoCE, congestion control, QoS, virtualization overlays, and traffic management.

Strong background with DPUs, SmartNICs, or other high-performance networking devices.

Experience with system management, provisioning, monitoring, telemetry, diagnostics, or lifecycle-management flows.

Familiarity with management protocols and frameworks such as Redfish, PLDM, MCTP, IPMI, gNMI, SNMP, Netconf, REST, or gRPC-based APIs.

Ability to lead cross-functional architecture discussions across software, firmware, hardware, validation, product, and customer-facing teams.

Excellent written and verbal communication skills, including the ability to create clear architecture documents and present trade-offs.


Ways To Stand Out From The Crowd:

Experience defining software architecture for DPU products, including management, telemetry, QoS, performance, security, virtualization, or offload features.

Hands-on background with Linux networking, device drivers, firmware, embedded Linux, BMC software, DOCA, DPDK, OVS or Kubernetes networking.

Experience with performance counters, profiling tools, eBPF, Prometheus, Grafana, dashboards, heat maps, or large-scale telemetry systems.

Experience in defining and developing GAI-based analysis tools to extract insights from telemetry data and streams.

Background in RAS, diagnosability, serviceability, field failure analysis, production debug, or customer escalation handling.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8773378
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
חברה חסויה
Location: Tel Aviv-Yafo
Job Type: Full Time
As a Senior SRE Engineer, you will be a key player in ensuring the reliability, scalability, and performance of our critical IT infrastructure. You will leverage SRE principles and an automation-first mindset to build and maintain resilient hybrid cloud environments. This role is ideal for a candidate who thrives in a fast-paced, innovative setting and is passionate about solving complex challenges with cutting-edge technology.
Key Responsibilities
Provision, configure, and support resilient hybrid cloud deployment architectures using an Infrastructure-as-Code framework.
Proactively collaborate with development teams to ensure new applications are production-ready, scalable, and reliable from inception.
Develop and maintain tools and frameworks to automate operational tasks, including deployment, monitoring, and recovery.
Conduct thorough root cause analysis of production issues and implement preventative measures to improve system resilience, demonstrating strong problem-solving skills.
Manage CI/CD platforms, Linux infrastructure, and contribute to capacity planning and operational runbooks.
Design and implement proactive service monitoring, alerting, and trend analysis to maintain service availability and performance SLAs.
Participate in an on-call rotation to support critical applications and services, responding to and resolving incidents efficiently.
Contribute to comprehensive documentation related to infrastructure design, deployment, and operational procedures.
Requirements:
Your Expereience:
6+ years of Devops engineering experience on mission-critical, enterprise-level systems in a hybrid (both cloud and on-prem) environment.
3+ years of hands-on experience with cloud environments, preferably Google Cloud Platform (GCP).
Expertise in configuration management and Infrastructure-as-Code using frameworks such as Terraform and Ansible.
Strong programming/scripting knowledge in languages like Python, Bash, or Go for infrastructure automation.
Demonstrated experience with CI/CD pipelines (e.g., GitHub, Jenkins, Artifactory) and a strong foundation in Linux/Unix administration.
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
Preferred Qualifications
Experience with containerization and orchestration technologies, particularly Kubernetes.
Hands-on experience with monitoring and observability tools such as Datadog, Grafana, or Prometheus.
Understanding of networking principles including firewalls, load balancers, and complex network designs.
A curious and positive mindset with a passion for applied learning and challenging existing processes for continuous improvement.
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8779502
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
28/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are looking for a Technical Lead to drive the architectural direction and engineering excellence of this group. This is a senior, deeply hands-on role for a technology leader who can own the technical roadmap, mentor a team of elite engineers, and build the infrastructure that challenges platform to its theoretical limits.
What You'll Lead:
Define and own the technical architecture of the group's distributed testing and reliability platform - designing for massive scale, real-world workload simulation, and adversarial failure injection
Lead effort involving multiple engineers, setting technical standards, running architecture reviews, driving design decisions, and mentoring engineers to grow
Build the systems that orchestrate millions of concurrent IO operations, inject chaos at the infrastructure layer (latency, packet loss, hardware failures), and expose the hardest-to-find race conditions and consistency bugs
Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
Drive observability and reliability engineering across the group - building telemetry pipelines that track P99 latency, jitter, and system health, turning quality into a quantitative discipline
Collaborate deeply with Core R&D, Storage Kernel, and Infrastructure teams - translating architectural knowledge into targeted reliability strategies
Establish engineering practices - design docs, production-grade code reviews, testing philosophy, and cross-team technical alignment
Requirements:
Strong software engineering background with 6+ years of hands-on Python development experience is required. The ability to read, debug, and reason about C++, Rust, or Go is a significant advantage
Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system behavior under stress
Background in one or more of: storage systems, networking (TCP/IP, RDMA), cloud infrastructure, database internals, or high-performance backend systems
Experience building large-scale infrastructure platforms, internal developer platforms, or reliability engineering systems
Leadership:
Proven track record leading complex technical initiatives from architecture through delivery
Experience mentoring and growing engineers - raising the technical bar of a team, not just directing work
Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
Comfortable operating at both the strategic and hands-on level - you write code, review designs, and shape roadmaps
Previous experience in people management roles - Advantage
Mindset:
You approach quality through the lens of Site Reliability Engineering: you care about MTTD, observability, and building self-healing systems
You have a "hacker" instinct - you don't just find bugs; you find the architectural flaws that allowed them to exist
You are an early adopter of AI tools and excited about applying LLMs and generative AI to accelerate engineering velocity
Big Advantages
Experience with storage systems, file systems, or high-performance distributed environments
Background in chaos engineering, fault injection, or simulation systems
Familiarity with observability tooling and performance engineering at scale
Experience building testing or reliability platforms as first-class engineering products
Prior experience as a Team Lead in a high-growth infrastructure company
This position is open to all candidates.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8757543
סגור
שירות זה פתוח ללקוחות VIP בלבד
סגור
דיווח על תוכן לא הולם או מפלה
מה השם שלך?
תיאור
שליחה
סגור
v נשלח
תודה על שיתוף הפעולה
מודים לך שלקחת חלק בשיפור התוכן שלנו :)
21/07/2026
Location: Tel Aviv-Yafo
Job Type: Full Time
We are seeking an experienced Senior Delivery Consultant - Modernization with deep expertise in Artificial Intelligence to join AWS Professional Services (ProServe). This role combines strategic architectural vision with hands-on technical leadership to deliver innovative AI solutions that drive customer success and business transformation across diverse industries and use cases.

Key job responsibilities
* Architecture & Design: Design and architect end-to-end AI-powered application solutions aligned with customer business objectives and technical requirements.
* Define application architecture patterns, standards, and best practices for AI/ML integration on AWS.
* Create technical roadmaps for customer AI application development and modernization initiatives
* Evaluate and recommend AWS AI/ML services and technologies including our Bedrock, SageMaker, and generative AI solutions
* Design data pipelines and ETL processes to support AI model training and inference using AWS services
* Customer Engagement & Consulting:
Lead customer engagements from discovery through implementation, serving as trusted technical advisor
* Conduct AI readiness assessments and develop adoption strategies tailored to customer maturity levels
* Facilitate architecture workshops and design sessions with customer stakeholders
* Deliver Well-Architected reviews focused on AI/ML workloads
* Build strong relationships with customer technical teams and executive leadership
* Guide customers in constructing AI processes aligned with AWS best practices
* Technical Leadership: Lead cross-functional teams in implementing AI solutions from concept to production
* Provide technical guidance on AI model integration, deployment strategies, and optimization on AWS
* Conduct architecture reviews ensuring solutions meet scalability, performance, security, and cost-efficiency requirements
* Mentor customer teams and junior ProServe consultants on AI best practices and AWS technologies
* Collaborate with data scientists, ML engineers, and software developers to translate AI models into production applications
* AI Solution Development: Design architectures for generative AI applications including RAG (Retrieval-Augmented Generation) systems, chatbots, and intelligent agents using our Bedrock
* Architect real-time and batch AI inference pipelines with appropriate monitoring and observability
* Implement MLOps practices using SageMaker for model versioning, deployment automation, and continuous improvement
* Design solutions for responsible AI including bias detection, explainability, and governance frameworks
* Optimize AI application performance, cost, and resource utilization across AWS services
Knowledge Sharing & Thought Leadership
* Develop reusable assets, reference architectures, and best practice documentation
* Contribute to AWS ProServe knowledge base and customer-facing content.
דרישות:
Basic Qualifications
- 10+ years of experience in application architecture and software development.
- 5+ years of hands-on experience with AI/ML technologies and frameworks (TensorFlow, PyTorch, scikit-learn, Hugging Face).
- Deep expertise in AWS cloud platform with focus on AI/ML services (SageMaker, Bedrock, Comprehend, Rekognition, etc.).
- Proficiency in programming languages such as Python, Java, or similar.
- Strong knowledge of generative AI technologies including LLMs, prompt engineering, fine-tuning, and RAG architectures.
- Understanding of various AI domains: NLP, computer vision, recommendation systems, predictive analytics.
- Willingness to travel to customer sites as needed.

Preferred Qualifications
- AWS Certified Machine Learning Specialty or AI Practitioner or Generative AI - Associate.
- Contributions to open-source AI projects or published research.
- Experience with responsible AI frameworks, governance practices, and compliance requirements.
- Prior experience in ProServe, consulting, or systems integration roles המשרה מיועדת לנשים ולגברים כאחד.
 
Show more...
הגשת מועמדותהגש מועמדות
עדכון קורות החיים לפני שליחה
עדכון קורות החיים לפני שליחה
8748479
סגור
שירות זה פתוח ללקוחות VIP בלבד