we are looking for a hands-on technical leader to build and operate a secure local large-language-model platform for the company. The platform will allow engineering and business teams to use generative AI with proprietary source code, product documentation, technical standards, test artifacts, support knowledge, and other approved internal data while keeping sensitive information within company-controlled environments.
This is a senior individual-contributor role spanning applied LLM engineering, platform architecture, search and data pipelines, security, and production operations. You will turn promising prototypes into a dependable internal capability: selecting and optimizing open-weight models, building permission-aware retrieval, creating reusable APIs and tools, integrating with existing engineering workflows, and establishing objective ways to measure quality, safety, latency, capacity, and business value.
The successful candidate will understand that a useful enterprise LLM is more than a model and a chat interface. It requires trustworthy source grounding, strong access controls, repeatable evaluation, careful tool permissions, observable production services, and an operating model that keeps data, indexes, prompts, models, and dependencies current. You will make pragmatic build-versus-buy decisions and choose the simplest approach-search, retrieval-augmented generation (RAG), prompting, workflow automation, or model adaptation-that meets each use case.
Initial use cases may include engineering knowledge discovery, source-code understanding, troubleshooting assistance, technical-document Q&A and summarization, test and log analysis, and drafting structured engineering artifacts. The platform should be extensible to additional approved use cases as needs and model capabilities evolve.
Requirements: BSc or MSc in Computer Science, Computer Engineering, Electrical Engineering, Data Science, or a related field, or equivalent practical experience.
Typically 7+ years of hands-on experience in production software, ML platform, search, data, or infrastructure engineering, including meaningful recent experience shipping LLM-powered systems; exceptional candidates with equivalent depth are welcome.
Strong Python engineering skills and experience designing maintainable APIs, services, libraries, and data pipelines. Experience with Go, Java, or C/C++ is an advantage.
Strong understanding of transformer-based language models and production inference, including tokenization, context management, batching, KV caching, parallelism, quantization, structured output, tool calling, and common model failure modes.
Demonstrated experience building production RAG or enterprise-search systems using embeddings, vector and/or lexical search, metadata filtering, reranking, source attribution, and systematic retrieval evaluation.
Experience defining task-specific LLM evaluations using representtive datasets, strong baselines, domain-expert review, automated metrics, human feedback, error analysis, and regression thresholds.
Experience deploying and operating containerized services on Linux using Docker and Kubernetes or an equivalent orchestration environment.
Practical experience with GPU-backed model serving, performance profiling, capacity planning, monitoring, and reliability engineering.
Strong knowledge of distributed-system fundamentals, authentication and authorization, API security, secrets handling, encryption, auditability, and data lifecycle controls.
Experience with Git, automated testing, CI/CD, infrastructure as code, observability, and production incident response.
Sound technical judgment about quality, security, maintainability, hardware efficiency, and total cost-not just model benchmark scores.
Ability to lead an ambiguous, cross-functional initiative, explain complex AI behavior in plain language, and help other teams ship safely on a shared platform.
This position is open to all candidates.