We are looking for a Lead Machine Learning Engineer with strong hands-on experience across traditional machine learning infrastructure, MLOps, cloud platforms, and modern generative/agentic AI systems. The ideal candidate should have a strong foundation in building production-grade ML platforms and infrastructure and should also have significant recent experience working with LLMs, RAG, AI agents, agentic workflows, and LLMOps. This is a hands-on technical leadership role where you will own the architecture, development, deployment, scalability, reliability, and observability of ML and AI systems in production.
Responsibilities:
- Lead the design and development of scalable ML infrastructure and production ML platforms.
- Own the complete ML lifecycle, including data pipelines, model training, experimentation, deployment, serving, monitoring, and retraining.
- Design and implement robust MLOps/LLMOps pipelines for production AI systems.
- Build and maintain CI/CD/CT pipelines for ML models, LLM applications, prompts, and AI agents.
- Architect cloud-native ML infrastructure using AWS/Azure/GCP, Docker, Kubernetes, Terraform, and related technologies.
- Build scalable model-serving and inference infrastructure with a focus on performance, reliability, latency, and cost optimization.
- Lead the development and deployment of generative AI applications using LLMs, RAG, embeddings, vector databases, and tool calling.
- Design and build agentic AI systems, including autonomous agents, multi-agent workflows, tool use, memory, orchestration, and enterprise integrations.
- Work with frameworks such as LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, or equivalent technologies.
- Establish observability for ML and AI systems covering model performance, latency, token usage, cost, quality, agent traces, failures, and drift.
- Implement model and AI evaluation frameworks, including offline evaluation, regression testing, LLM evaluation, human-in-the-loop evaluation, and production monitoring.
- Design security and governance mechanisms for AI systems, including guardrails, access control, data security, prompt-injection protection, and auditability.
- Collaborate closely with data science, backend, data engineering, DevOps, product, and platform teams.
- Define engineering standards, conduct architecture/code reviews, and mentor other ML/AI engineers.
- Evaluate emerging ML, GenAI, Agentic AI, and MLOps technologies and identify opportunities to bring them into production.
Requirements:
- We are specifically looking for someone who can bridge the gap between:
- Traditional ML Infrastructure, MLOps Cloud/Platform Engineering, LLM Engineering, and Agentic AI.
- The ideal candidate is not only an LLM/GenAI developer but not purely an MLOps engineer. They should have a strong engineering foundation in production ML infrastructure and the ability to build and scale modern GenAI and Agentic AI systems.
- Education: Bachelor's or master's degree in computer science, engineering, mathematics, artificial intelligence, or a related technical field preferred.
- Key Competencies: ML Infrastructure, MLOps, ML Engineering, Python, Kubernetes, Docker, Cloud, CI/CD, Terraform, MLflow, Model Serving, LLMs, RAG, Vector Databases, LLMOps, Agentic AI, LangGraph, LangChain, AI Agents, Tool Calling, AI Observability, Model Evaluation, Distributed Systems.
Traditional ML / ML Engineering:
- 5+ years of experience in machine learning engineering, ML infrastructure, MLOps, or a closely related engineering role.
- Strong understanding of traditional machine learning and deep learning concepts.
- Hands-on experience building production ML systems and ML pipelines.
- Experience with model lifecycle management, model registry, experiment tracking, model deployment, monitoring, and rollback.
- Strong Python programming and software engineering fundamentals.
- Experience with ML frameworks such as PyTorch, TensorFlow, Scikit-learn, or equivalent.
ML Infrastructure & MLOps:
- Strong hands-on experience with MLOps and ML infrastructure.
- Experience building production-grade training and inference pipelines.
- Strong knowledge of Docker and Kubernetes.
- Experience with CI/CD, GitHub Actions, Jenkins, GitLab CI, ArgoCD, or similar tools.
- Experience with Terraform or other Infrastructure-as-Code tools.
- Experience with one or more major cloud platforms: AWS, Azure, or GCP.
- Experience with tools such as MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, or equivalent.
- Strong understanding of monitoring, logging, observability, scalability, reliability, and distributed systems.
Generative AI / LLM:
- Strong hands-on experience building production-grade LLM applications.
- Experience with RAG architectures, embeddings, vector databases, prompt engineering, function/tool calling, and LLM evaluation.
- Experience working with models such as OpenAI, Claude, Gemini, Llama, Mistral, or equivalent.
- Understanding of LLM inference, model serving, context management, latency, token optimization, and cost optimization.
- Experience with LLMOps, including prompt/model versioning, evaluation, monitoring, and deployment.
Agentic AI:
- Strong practical experience building AI agents and agentic AI workflows.
- Experience designing multi-step and/or multi-agent systems.
- Hands-on experience with tool calling, agent orchestration, memory, planning, reasoning workflows, and external API/system integrations.
- Experience with frameworks such as LangGraph, LangChain, LlamaIndex, AutoGen, Semantic Kernel, CrewAI, or equivalent.
- Understanding of AgentOps, agent observability, evaluation, guardrails, and production reliability.
- Exposure to emerging agent architecture patterns such as MCP and agent-to-agent communication is a plus.
Preferred Qualifications:
- Experience architecting AI/ML platforms from the ground up.
- Experience managing GPU-based workloads, distributed inference, or high-throughput model serving.
- Experience with vLLM, TGI, Ray, Triton, or similar inference technologies.
- Experience with Kafka, Spark/PySpark, Airflow, or other distributed data/workflow technologies.
- Experience with feature stores, vector databases, knowledge graphs, or real-time ML systems.
- Strong understanding of AI security, governance, responsible AI, and production risk management.
- Experience leading technical projects and mentoring ML/AI engineers.

