our company already runs agents in production. Our WAF Copilot takes a freshly disclosed CVE and turns it into a validated, provider-specific WAF rule - a state-driven agent graph that researches the vulnerability, maps the attack surface, composes the rule, then attacks its own output with an independent bypass judge and a false-positive prober before a human is ever offered a deploy. When a weakness survives, the system downgrades its own recommendation and records what it could not prove.
That last part is the thesis of this role.
An agent that touches production security must prove it works and declare what it couldn't prove. The security industry is about to be flooded with agentic claims nobody can verify. We intend to be the company that made verification legible - and that starts with holding our own agents to a standard we're willing to publish.
We're looking for an AI Tech Lead to own that standard across three surfaces:
The platform - Today each agent flow is close to a bespoke implementation. You'll turn our hard-won patterns into shared components, conventions, and infrastructure so the next agent is a week of work rather than a quarter - with evaluation, observability, and cost control built in rather than bolted on.
Enablement - our company's advantage compounds only if the whole company is AI-fluent, not just R&D. You'll raise that fluency everywhere - engineering, research, product, GTM - through tooling, patterns, and teaching.
The voice - You'll publish the methodology: how we benchmark agentic security output, how we model residual risk, what we learned failing. This is a category-defining position and we want it argued in public.
This is a hands-on lead role with no direct reports. Your authority comes from the quality of what you build and how clearly you explain it.
What You'll Do
Define and build our company's agent framework - shared components, orchestration patterns, tool interfaces, and conventions that make "how do I build an agent here" a five-minute answer.
Make evaluation a gate, not an afterthought: trajectory testing, golden datasets, offline replay, and per-dimension scoring that runs in CI. Land the release benchmark so promote / hold / rollback is a number, not a vibe.
Own agent observability end to end - the execution graph, tool invocations, intermediate reasoning, latency per step, and quality drift - including the external Temporal-orchestrated flows that are hardest to introspect today.
Own the economics: provider abstraction, model routing by task complexity, small-model substitution where it holds, cost attribution per flow. Thousands of CVEs through a labeling agent is a budget line, not a detail.
Drive latency and determinism in our production agent flows - tighter loops, early stopping, deterministic state machines over prose-in-prompt orchestration.
Raise the company's AI fluency. Build the internal tooling, skills, and playbooks that let every team - not just R&D - work AI-natively, and teach the judgment for when not to.
Partner with security research, product, and engineering leadership so agent capability and product roadmap actually converge.
Requirements: You've shipped agentic systems to production - real orchestration, tool use, structured outputs, and the failure modes that only appear at scale. Not "I've called an LLM API."
You've built the evaluation discipline, not just consumed it: trajectory tests, golden datasets, regression gates, offline replay. "It seems better" is not a metric, and you have opinions about what is.
Deep backend and distributed-systems engineering. Strong Python, and comfort with workflow orchestration (Temporal or equivalent), streaming, and cloud-native infrastructure. Agent platforms are systems problems wearing an AI hat.
Fluency across the modern agent stack - LangChain/LangGraph-style frameworks, multi-provider routing, structured output contracts, prompt and context engineering.
This position is open to all candidates.