we are looking for a Senior AI Engineer - Exploration & Prototyping.
That is this role. You take an open question, run a focused spike, and come back with numbers and a recommendation. You read the source of the frameworks you evaluate rather than trusting their marketing. You build prototypes to settle arguments.
You do not own a subsystem and you do not ship to customers, which is exactly what protects the work: exploration inside a delivery team always loses to the sprint. You sit alongside the platform, research, and forward-deployed teams, you borrow their context freely, and your output is evidence they can act on.
You will be trusted with real influence early. The recommendations you write become the architecture other people build against, so the bar is not a working demo but a defensible conclusion, including the ones that say no.
What Youll Do:
Run technical spikes that close open decisions, covering agent orchestration frameworks, real-time transport, memory protocols, agent interoperability standards, LLM selection and routing, evaluation harnesses, and the production library and stack choices underneath all of it.
Build prototypes to de-risk, standing up something real quickly, proving or disproving the thing in question, and moving on without becoming attached to the code.
Read and evaluate unfamiliar codebases, going into the source of a candidate framework to find out whether it can actually support what we need rather than what its documentation implies.
Design the measurements that make a decision defensible, building the harness, running the comparison, and reporting latency, cost, and failure behavior honestly.
Own build-versus-adopt recommendations for platform infrastructure, frameworks, and libraries, including a clear statement of what it would cost to be wrong.
Write the recommendation down. Every spike ends in a short, decisive document another engineer can act on, with the evidence, the rejected options, and the reasoning behind the call
Hand off cleanly, transferring what you learned to the team that will own the capability in production, and staying available while they pick it up.
Track the landscape across agentic infrastructure, real-time frameworks, and adjacent AI tooling, and bring forward the things that genuinely change what we can build.
Requirements: B.Sc. in Computer Science (or equivalent technical field), mandatory.
7+ years of industry experience in software, ML, or research engineering roles, with real ownership of production systems.
Genuine technical breadth. You have worked across backend services, runtime, and infrastructure, and you are comfortable close to ML systems without needing to own the models. You can hold several unfamiliar domains at once.
Strong Python skills, and the ability to get something real running quickly.
A track record of technical evaluations that led to decisions, where you compared real options, produced evidence, and the organization acted on the result.
Evidence over intuition. You have designed benchmarks or measuremet harnesses, and you can describe a time you were convinced something would work and the numbers said otherwise.
Experience with real-time, streaming, or latency-sensitive systems.
Hands-on experience with LLMs and agentic systems, including orchestration, tool calling, and how these systems behave and fail in production.
Comfortable working as an individual contributor without a team, self-directed, and able to finish. Exploration that never lands is the failure mode of this role.
Experience in a fast-moving SaaS company and in cloud environments (AWS, GCP, or Azure).
This position is open to all candidates.