In this role, you will architect and implement large-scale distributed systems, low-level platform components, and high-performance services that underpin core product stack. You will drive technical strategy across system reliability, performance, and scalability, partnering closely with product, infrastructure, and data engineering teams to deliver systems that operate at global scale with high availability and efficiency.
Software Engineer, Systems Responsibilities
Architect and implement large-scale distributed systems and platform services that support high-throughput, low-latency workloads across Meta's product infrastructure
Lead the technical design of systems components including storage layers, compute pipelines, networking abstractions, and service orchestration frameworks
Identify and resolve systemic performance bottlenecks through instrumentation, profiling, and targeted optimization across the full systems stack
Define and enforce service level objectives for owned systems, building dashboards, alerting pipelines, and runbooks to reduce mean time to mitigation during incidents
Drive reliability improvements by reducing failure surface, designing resilient rollout strategies, and leading regular resiliency and overload testing exercises
Collaborate with cross-functional partners across product engineering, infrastructure, and data science to align system architecture with evolving product and business requirements
Establish and evolve coding standards, architectural patterns, and engineering best practices for systems development across the broader organization
Leverage AI-assisted development workflows to accelerate design iteration, code generation, and systems analysis, applying sound judgment on when to rely on AI versus deep systems expertise
Mentor other engineers on systems design principles, debugging methodologies, and production operations, and contribute to onboarding programs for new team members
Lead incident retrospectives, identify root causes of complex production failures, and drive implementation of systemic improvements to prevent recurrence
Requirements: 8+ years of experience designing and implementing large-scale distributed systems, platform infrastructure, or systems software in production environments
Experience leading major technical initiatives end-to-end, including architecture design, cross-team coordination, staged rollout, and post-launch reliability ownership
Experience debugging complex, non-reproducible systems issues including concurrency bugs, memory management failures, and distributed consistency problems
Experience defining service level objectives, building observability infrastructure, and driving reliability improvements across production systems
Experience communicating technical architecture decisions and trade-offs in writing to both engineering and non-engineering stakeholders
This position is open to all candidates.