Operate and continuously improve a production ML/AI platform in an air-gapped environment.
Build and maintain infrastructure supporting AI development, evaluation, deployment, monitoring, and model lifecycle management.
Deploy and manage open-weight and self-hosted LLMs (Llama, Mistral, Gemma, Phi, embedding models, rerankers, and domain-specific models).
Optimize inference performance, GPU utilization, and multi-model serving.
Design secure offline workflows for importing, validating, scanning, and mirroring software, models, datasets, and container images.
Maintain private registries and offline installation and upgrade mechanisms.
Build and operate GitOps, CI/CD pipelines, observability, logging, metrics, tracing, GPU monitoring, SLOs,
Requirements: 5+ years of experience in Software Engineering, Platform Engineering, DevOps, Infrastructure, data Engineering, or ML Engineering.
Hands-on experience building or operating production ML/AI platforms.
Strong experience with Kubernetes (or OpenShift/Rancher), Linux, containers, networking, Storage, and production environments.
Strong scripting skills and experience working with cross-functional engineering teams.
Experience with model serving frameworks such as vLLM, Triton Inference Server, KServe, BentoML, TorchServe, or Ray Serve.
Experience with ML lifecycle platforms including MLflow, Kubeflow, Metaflow, Airflow, or Argo Workflows.
Experience with RAG infrastructure and vector databases such as pgvector, OpenSearch/Elasticsearch, MongoDB Vector Search, Milvus, Weaviate, or Qdrant.
This position is open to all candidates.