As the AI Platform Architect at ConveGenius, you'll design and scale the AI infrastructure that powers model training, fine-tuning, inference, agentic workflows, and orchestration across our learning applications. You should understand how LLMs and foundation models are actually trained, not just served or fine-tuned, and have worked on AI safety and security model design. Experience with benchmarking across AI systems and hands-on work with multi-agent orchestration frameworks is preferred. This role is part of a 0-to-1 vertical build. The person joining should be comfortable with ambiguity, high ownership, rapid iteration, hands-on execution, team building, and periods of high-intensity work during the early setup phase.
Responsibilities:
- Design and own end-to-end AI platform architecture: LLM serving, RAG pipelines, agentic orchestration, and model lifecycle management.
- Build scalable LLM inference infrastructure capable of serving millions of concurrent learners across diverse devices and connectivity conditions.
- Define and enforce AI safety, governance, and evaluation standards across all AI systems deployed in production.
- Drive RAG and agentic AI architectures: retrieval design, tool use, multi-agent frameworks, and orchestration patterns.
- Lead fine-tuning pipeline design: LoRA, QLoRA, RLHF, DPO from data preparation to production serving.
- Architect MLOps practices: CI/CD for models, experiment tracking, deployment versioning, and rollback mechanisms.
- Manage GPU cluster infrastructure: compute allocation, distributed training setup, and cost engineering.
- Drive AI unit-economics by setting up token-tracking, model drift metrics, and distributed tracing via OpenTelemetry, LangSmith, or Databricks.
- Collaborate with product, data, and engineering teams to translate AI platform capabilities into learner- facing features.
- Mentor and grow a cross-functional team of ML engineers, platform engineers, and AI researchers.
- Design and operate multi-tenant AI platform infrastructure ensuring isolation, fairness, and cost attribution across teams.
- Incorporate Indic language NLP requirements into platform design: multilingual embedding support, code- switching handling, and Indic model serving.
Requirements:
- Strong experience in designing and scaling AI/LLM platforms.
- Deep understanding of Generative AI, LLMs, and AI system architecture.
- Hands-on experience with model training, fine-tuning, and deployment (PyTorch, LoRA, RLHF, DeepSpeed, Megatron-LM).
- Experience with inference frameworks such as vLLM, Triton, or TGI.
- Strong knowledge of RAG, Graph-RAG, vector databases, and agent orchestration.
- Understanding of distributed systems, cloud infrastructure, scalability, and security.
- Experience building production-grade, multi-tenant AI applications.
- Exposure to Indian-language NLP is preferred.
- Deep understanding of how LLMs and foundation models are trained (not just fine-tuned): model training pipelines, distributed training, dataset construction; hands-on experience with AI safety and security model design; proficiency in AI benchmarking methodologies and multi-agent orchestration frameworks.
- Working hands-on experience with AWS (Bedrock, SageMaker), Azure AI, or GCP; Docker, Kubernetes (EKS/GKE), etc is a plus.

