IDFC First Bank
We are looking for an experienced AI Engineer who can design, build, and deploy end-to-end AI systems with a strong emphasis on scalable GenAI platforms, Retrieval-Augmented Generation (RAG) pipelines, copilots, and applied AI solutions that drive measurable business impact. This role goes beyond traditional model training. You will architect and implement efficient large language model (LLM) fine-tuning techniques such as LoRA, QLoRA, and Unsloth, and build robust LLMOps pipelines to automate model lifecycle management from training to deployment and monitoring. The ideal candidate brings a strong blend of AI engineering, MLOps, and cloud infrastructure expertise and is comfortable working with both LLMs and applied AI use cases, including NLP, computer vision, and multimodal AI. You will collaborate cross-functionally with business, product, and engineering teams to deliver secure, compliant, and enterprise-grade AI-driven capabilities at scale, enabling smarter banking operations, enhanced customer experiences, and operational efficiencies.
The core responsibilities for the job include the following:
Primary:
- Build and deploy AI-powered applications (chatbots, copilots, automation flows) to enhance banking operations and customer service.
- Design and implement Retrieval-Augmented Generation (RAG) pipelines and AI agents for secure financial data access and insight generation.
- Fine-tune and optimise LLMs using Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRA, QLoRA, and Unsloth for domain-specific tasks.
- Automate end-to-end LLMOps pipelines, including training, evaluation, versioning, deployment, and monitoring.
- Develop and expose backend APIs and microservices to integrate LLMs with banking platforms and digital channels.
- Deploy scalable AI models on AWS/Azure using Docker, Kubernetes, and CI/CD pipelines while ensuring enterprise-grade security and compliance.
- Implement monitoring and observability for production LLMs, covering drift detection, hallucination tracking, and latency management.
- Ensure all AI systems adhere to financial regulations, security standards (e. g., GDPR, PCI-DSS), and internal governance policies.
Secondary:
- Collaborate with product managers, data scientists, and compliance teams to translate business needs into AI-enabled solutions.
- Create reusable AI components, SDKs, and templates to accelerate internal development and adoption of GenAI solutions.
- Contribute to model optimisation strategies such as quantisation, prompt tuning, and adapter layers.
- Support data engineering and preprocessing tasks using ETL/ELT pipelines and scalable data infrastructure (e. g., Spark, Airflow).
- Mentor junior engineers and conduct internal workshops on LLMOps, GenAI tools, and PEFT techniques like Unsloth and QLoRA.
- Assist in creating internal documentation and wikis for AI/LLM systems, model usage guidelines, and deployment standards.
- Work with the security team to assess risks related to LLM integrations and mitigate vulnerabilities in prompt injection or data leakage.
- Perform A/B testing and user feedback collection on LLM-based features to drive continuous improvements.
- Contribute to internal research and POCs around new GenAI technologies, open-source models, or fine-tuning strategies.
- Participate in vendor/tool evaluations (e. g., vector DBs, cloud LLM services) for LLMOps stack improvements.
- Help establish internal standards and best practices for PEFT workflows (e. g., LoRA config templates, quantised model packaging).
- Review and improve the latency, throughput, and cost-efficiency of deployed LLM inference pipelines.
- Track and evaluate model behaviour over time, identifying drifts, degradation, or hallucination risks in specific financial tasks.
- Support cross-functional innovation initiatives (e. g., AI for fraud prevention, risk modelling, compliance automation).
