This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Platform Engineer based in India.
As an AI Platform Engineer, you will join a newly established team building a central AI platform from the ground up.
You will develop the infrastructure, services, and tooling that enable multiple product teams to deliver AI capabilities reliably at scale.
Your work will span AI platform engineering, RAG, LLMOps, observability, security, automation, and production operations.
You will make hands-on technical decisions that directly shape how AI systems are built, deployed, evaluated, and operated.
The role offers substantial ownership, with your work moving rapidly from design and implementation into production environments.
You will collaborate closely with senior engineers and product teams in a distributed, engineering-driven environment.
This is an opportunity to help establish the foundations of an AI platform while solving complex challenges around reliability, quality, scalability, and cost.
Accountabilities:
- Design, build, deploy, and operate scalable AI platform services with clear API contracts, versioning, service-level objectives, and production-readiness standards.
- Develop production-grade Python services, shared libraries, SDKs, and integration patterns for consumption by multiple product engineering teams.
- Build and evolve a multi-provider model gateway with routing, fallback, retry, rate limiting, cost attribution, and budget enforcement capabilities.
- Develop RAG-as-a-service capabilities covering data ingestion, chunking, embeddings, retrieval, hybrid search, and retrieval-quality measurement.
- Establish prompt management capabilities including templates, versioning, evaluation, rollout controls, and tenant-specific customization.
- Build guardrails and content-safety mechanisms such as input filtering, output validation, PII redaction, and controlled tool execution.
- Develop agent and tool-use patterns suitable for reliable production workflows.
- Build observability capabilities covering prompt and response tracing, latency, errors, cost per request, quality signals, drift detection, and audit logging.
- Design and maintain evaluation infrastructure using benchmark datasets, offline evaluations, LLM-as-judge approaches, human review workflows, and regression testing.
- Build and operate model deployment pipelines, including safe deployment and rollback processes for fine-tuned models where applicable.
- Establish alerting and SLO frameworks for AI services, treating quality regressions as a first-class operational concern.
- Participate in on-call rotations, create incident response runbooks, and continuously improve operational procedures following incidents.
- Collaborate directly with product engineering teams to onboard AI features, provide integration guidance, and troubleshoot platform issues.
- Contribute to architecture decisions, document technical tradeoffs, and constructively challenge approaches when appropriate.
- Evaluate third-party technologies and vendors through structured comparisons, making cost, quality, security, and reliability tradeoffs explicit.
- Strengthen AI security across platform services, including PII protection, tenant isolation, prompt-injection defenses, and auditability.
- Identify opportunities to improve scalability, reliability, developer experience, and operational efficiency across the platform.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
- At least 5 years of professional software engineering experience with a strong production track record.
- Approximately 6+ years of experience building production distributed systems, ideally including internal developer platforms, infrastructure platforms, or API gateways at scale.
- At least 5 years of experience in MLOps, LLMOps, ML platform engineering, or a comparable combination of DevOps and machine-learning engineering at production scale.
- At least 2 years of hands-on production experience with LLM-based systems, including areas such as prompt engineering, RAG, evaluation, or LLM infrastructure.
- Strong Python development skills plus proficiency in Go or Java, with experience writing production-quality code rather than primarily notebooks or scripts.
- Strong understanding of asynchronous programming, backpressure, rate limiting, distributed systems, and production service design.
- Practical experience with LLM observability and evaluation platforms such as LangSmith, Langfuse, Braintrust, Arize, or comparable tools.
- Working knowledge of LLM evaluation methodologies, including benchmark design, LLM-as-judge approaches, regression testing, and human review workflows.
- Familiarity with modern LLM technologies, including OpenAI or Anthropic APIs, an orchestration framework such as LangChain or LlamaIndex, vector databases, and LLM observability tooling.
- Experience designing and operating multi-tenant systems with strong data and service isolation guarantees.
- Strong cloud-native experience with AWS or Azure, including Kubernetes, service mesh technologies, infrastructure-as-code, Terraform, and CI/CD.
- Experience deploying model updates safely using techniques such as canary releases, shadow evaluation, feature controls, and rollback triggers.
- Understanding of the full ML lifecycle, including training pipelines, model serving, monitoring, evaluation, and cost management.
- Strong understanding of AI security fundamentals, particularly PII handling, tenant isolation, prompt-injection risks, and audit logging.
- Strong written and verbal communication skills, with the ability to document technical decisions clearly in an asynchronous, distributed environment.
- Fluent English communication skills.
- Preferred experience in procure-to-pay, ERP integration, accounts payable, procurement, finance, operations software, or related domains.
- Experience working at a product company, B2B SaaS organization, or internal platform team is advantageous.
- Contributions to open-source AI/ML infrastructure projects are a plus.
- Production experience with agent frameworks such as LangGraph, AutoGen, CrewAI, or custom orchestration is desirable.
- Experience as part of a founding platform team responsible for launching a new service used by multiple internal customers is highly valuable.
- Fully remote opportunity available across India.
- Work schedule aligned to 11:00 AM-8:00 PM.
- Opportunity to join a newly formed AI platform team and help build its initial platform from the ground up.
- Significant technical ownership and the opportunity to influence foundational architecture and engineering standards.
- Hands-on exposure to modern LLM, RAG, AI agent, MLOps, LLMOps, cloud, observability, and platform technologies.
- Opportunity to work on production AI systems used across multiple product lines.
- Close collaboration with experienced engineers and product development teams across a distributed environment.
- Broad technical scope spanning applied AI, backend engineering, cloud infrastructure, security, reliability, and developer experience.
- Opportunity to establish best practices for AI reliability, evaluation, observability, security, and cost management.
- Fast-moving environment with a strong emphasis on engineering quality, ownership, learning, and measurable production impact.

