We are looking for a highly skilled Senior LLMOps/MLOps Engineer with strong expertise in LLM inferencing, model hosting, and serving Large Language Models (LLMs) at scale. The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building, troubleshooting, and optimizing production Al systems. Experience in MLOps platforms and scalable Al infrastructure is essential.
Must-Have Skills
10 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
Strong proficiency in Python and software engineering best practices.
Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
Strong expertise in LLM Inferencing and Model Hosting using VLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and Continuous/Dynamic Batching.
