We are looking for a highly motivated AI Research Student for our Robotics & Simulation AI team. You will be passionate about data science, AI quality measurement, and the intersection of AI with high-accuracy industrial domains. This role is ideal for someone who grows with challenging existing assumptions, defining rigorous evaluation methodologies, and pushing AI systems toward measurable reliability. You will combine strong data science and research skills with hands-on experimentation in a production AI product that controls robotic manufacturing simulations.
Key Responsibilities
AI Quality & Stability Research
Define and implement metrics to measure the stability, accuracy, and reliability of our production AI Copilot (PSCoPilot)
Design evaluation frameworks for AI agents operating in safety-critical robotics and simulation environments
Build observability and tracing pipelines using platforms like Langfuse to monitor Copilot performance, token usage, latency, and tool-calling success rates
Analyze failure modes, hallucination patterns, and tool-calling accuracy across 48+ MCP-exposed robotic operations
Develop automated regression and benchmark pipelines to track Copilot quality over time
Data Science & Experimentation
Conduct statistical analysis on Copilot interactions (tool selection accuracy, collision detection correctness, path planning success rates)
Research and apply techniques from AI evaluation literature (LLM-as-judge, human-in-the-loop validation, domain-specific benchmarks)
Leverage LangChain and similar orchestration frameworks to prototype evaluation chains, structured agent workflows, and comparison experiments
Explore how multi-layer agent architectures (fast triage + deep execution) affect accuracy and cost trade-offs
Challenge the team with data-driven insights on where the Copilot underperforms and why
Domain AI Research
Investigate how to improve AI reasoning for robotics tasks requiring high precision (welding, clearance checking, reachability analysis)
Research grounding techniques that reduce errors in domains with zero tolerance for incorrect outputs
Explore and apply new AI technologies (GenAI tools, AI coding assistants, agents, MCP-based solutions) to enhance Copilot capabilities
Validate technical feasibility through hands-on experimentation with our production MCP bridge and two-layer streaming agent engine
Communication & Documentation
Document findings, benchmarks, and technical recommendations
Present results to technical and business partners
Collaborate with the team to translate research insights into production improvements
Requirements: Currently pursuing a Master's degree or equivalent experience in AI, Data Science, Machine Learning, Computer Science, Statistics, or a related field with minimum 2 years until graduation
Strong background in data science, statistical analysis, and experiment design
Proficiency in Python with experience using LLM orchestration platforms (LangChain, AWS Bedrock, Azure OpenAI)
Familiarity with LLM evaluation techniques, prompt engineering, or AI agent architectures
Experience with or strong interest in defining quality metrics for AI systems
Fast learner with strong self-motivation and ability to challenge technical decisions with data
Proficiency in English
This position is open to all candidates.