We are looking for a proactive and skilled Software Infrastructure Developer with a strong Agentic AI focus to join our team. In this role, you will design and build the infrastructure, tooling, and AI-powered automation that keep our engineering systems reliable, scalable, and intelligent.
You will be responsible for embedding GenAI and agentic capabilities into our internal developer platforms and CI/CD workflows - creating frameworks, libraries, and multi-agent systems that accelerate automation, improve system stability, and unlock new levels of engineering productivity. This is a hands-on role for someone who thrives at the intersection of infrastructure engineering and applied AI.
The ideal candidate brings deep technical expertise in Python (FastAPI) and/or TypeScript, Kubernetes, CI/CD (Jenkins, GitHub Actions), containerization, and cloud infrastructure - combined with hands-on experience building LLM-powered applications, agents, MCP servers, and RAG pipelines.
What you'll be doing:
AI Platform Development: Design and build internal AI platforms, reusable SDKs, libraries, tools, and MCP servers for agentic applications.
Agentic Workflow Development: Develop single- and multi-agent workflows covering planning, tool selection, delegation, state and memory management, execution, retries, fallbacks, error recovery, and human-in-the-loop approval.
AI-Powered Infrastructure & CI/CD: Integrate AI capabilities into Kubernetes infrastructure and Jenkins/GitHub Actions pipelines to improve development, testing, deployment, and operational workflows.
AI Observability & Evaluation: Implement agent tracing, prompt and tool-call logging, latency and error monitoring, token and cost tracking, regression testing, and failure analysis.
AI Security & Governance: Apply guardrails, access controls, auditability, and governance practices to ensure agents operate safely and reliably.
End-to-End Platform Ownership: Own AI platform capabilities from architecture through production and build reusable solutions that accelerate adoption across engineering teams.
Cross-Functional Collaboration: Collaborate with engineering and business teams to identify high-impact AI opportunities and translate them into scalable technical solutions.
AI Research & Innovation: Evaluate emerging AI frameworks, models, and cloud technologies and promote their adoption where they provide measurable value.
Requirements: Core Engineering:
Minimum of 5 years of hands-on development experience with Python and/or TypeScript.
Experience building asynchronous backend services and APIs using FastAPI or an equivalent modern framework.
Strong software-engineering, analytical, troubleshooting, and cross-functional collaboration skills.
Ability to build reliable, reusable internal platforms that improve developer productivity and accelerate technology adoption.
Infrastructure & DevOps:
Strong experience with Kubernetes and Helm, including deployment, networking, resource management, autoscaling, troubleshooting, chart development, templating, versioning, and release management.
Strong experience with Docker, Jenkins, GitHub Actions, and AWS and/or Azure.
Familiarity with Infrastructure as Code, GitOps, and Git-based collaborative workflows.
AI & GenAI:
AI Experience - Mandatory: Minimum of 2 years of hands-on production experience building LLM-powered applications, AI agents, multi-agent systems, or internal AI platforms.
Hands-on experience with LangChain, LangGraph, Claude Agent SDK, and OpenAI and/or Anthropic APIs.
Strong understanding of agent architecture, including planning, tool selection, delegation, state and memory management, retries, fallbacks, error recovery, and human-in-the-loop workflows.
Experience with prompt and context engineering, tool/function calling, and structured outputs.
Experience building and integrating MCP servers and clients, including tool schemas, permissions, reliable execution, and error handling.
This position is open to all candidates.