1,438,441open jobs
85,144companies
219,560added this week
Browse all
Salary
≈ $21k – $53k per year (Estimated)
Location
In office (Delhi)
Experience
3+ years exp

First seen by Alion on Oct 9, 2026.

Overview
Company
Impact
Profile match
QubeLabs Workmate is a purpose-built Sovereign AI System for the Financial Services Industry — unifying workforce operations, intelligence, and enterprise execution at scale.

LLM/AI Model Engineer - Training, Evaluation & Benchmarking

Location : New Delhi, India (On-site)

Employment Type : Full-time

Experience : 3 - 5 years

Industry : Enterprise AI / Generative AI / Machine Learning / NLP

About QubeLabs :

QubeLabs is building the next generation of Enterprise AI Systems that transform workforce operations and intelligence for the financial services industry.

Our platform combines Conversational AI, Agentic AI, workflow automation, proprietary language models and enterprise intelligence. Our flagship product, QubeLabs Workmate, serves banks, NBFCs, MFIs, wealth management firms, insurance companies and fintechs across India and Europe.

At the core of our AI platform is the Vectro Series, our proprietary family of language models designed for enterprise and financial-services use cases.

We are looking for an LLM/AI Model Engineer to build the training pipelines, datasets and evaluation infrastructure required to continuously improve the Vectro Series.

Role Overview :

You will work closely with the Head of AI, Research Partners and engineering teams to implement and operationalize LLM training approaches and build a reliable model improvement lifecycle: Data Preparation - Training - Evaluation - Benchmarking - Error Analysis - Model Improvement.

The role combines hands-on LLM fine-tuning, dataset engineering, tokenization, evaluation framework development and model quality management.

Key Responsibilities :

1. LLM Training & Fine-tuning :

- Build and maintain pipelines for training and fine-tuning open-source foundation models for the Vectro Series.

- Implement supervised fine-tuning (SFT), parameter-efficient fine-tuning (PEFT), LoRA/QLoRA and domain adaptation techniques; support continued pre-training where required.

- Manage training configurations, checkpoints, model versions and experiment tracking.

- Optimize training workflows for computational efficiency, reproducibility and reliability.

- Translate model development approaches defined with the Head of AI and Research Partners into practical, scalable training pipelines.

2. Training Data & Tokenization :

- Build data pipelines to collect, clean, filter, normalize, deduplicate and validate training data.

- Develop instruction-tuning, SFT and domain-specific datasets for financial-services and enterprise use cases.

- Implement tokenization workflows, tokenizer configuration, vocabulary management, sequence packing, truncation and padding.

- Address data representation challenges involving financial terminology, numerical values and multilingual content.

- Maintain dataset versioning, lineage and quality controls, incorporating identified data gaps and relevant model failures into future training data.

3. Golden Dataset & Model Evaluation :

- Build and maintain QubeLabs' Golden Dataset to evaluate key model capabilities and target use cases.

- Develop automated evaluation pipelines and reusable evaluation harnesses, combining automated metrics with human assessment where appropriate.

- Evaluate accuracy, reasoning, factuality, instruction following, safety, multilingual performance and financial-domain capabilities.

- Establish consistent evaluation methodology and regression tests to measure changes across model versions.

4. Benchmarking & Error Analysis :

- Evaluate Vectro against relevant public benchmarks and proprietary QubeLabs BFSI/enterprise benchmarks.

- Compare results with relevant open-source and commercial models using consistent evaluation conditions.

- Analyze failures, identify root causes across data, training and model behavior, and translate findings into actionable improvements.

- Maintain reproducible benchmark results and reports to track model strengths, limitations and progress.

5. Model Quality & Release :

- Define model release quality gates and ensure each major release meets agreed evaluation and benchmark criteria.

- Support red-teaming, robustness and safety testing.

- Coordinate with the Head of AI, Research Partners and engineering teams to ensure validated model improvements are suitable for integration into production systems.

Experience :

- 3 - 5 years in Machine Learning, Deep Learning, NLP, Generative AI or related fields.

- Hands-on experience training or fine-tuning LLMs using open-source foundation models.

- Practical experience building training pipelines, preparing datasets and evaluating model performance.

- Experience with GPU-based training; distributed training is an advantage.

- Financial-services, multilingual AI or enterprise AI experience is preferred.

Required Skills :

- LLM & Machine Learning : Large Language Models, Transformer architectures, Generative AI, NLP, Deep Learning, SFT, PEFT, LoRA/QLoRA, fine-tuning and continued pre-training.

- Frameworks & Infrastructure : Python, PyTorch, Hugging Face Transformers, Hugging Face Tokenizers, GPU computing, experiment tracking and ML pipelines.

- Data & Tokenization : Tokenization, tokenizer configuration, text preprocessing, sequence packing, dataset construction, data quality, versioning and lineage.

- Evaluation & Benchmarking : Golden datasets, LLM evaluation, evaluation harnesses, automated and human evaluation, LLM-as-a-Judge, public and domain-specific benchmarks, regression testing and error analysis.

The role requires the ability to build both the model training pipeline and the measurement system that determines whether the model has actually improved.

What Success Looks Like :

- Reliable and reproducible training and fine-tuning pipelines for the Vectro Series.

- High-quality training datasets, tokenization workflows and a comprehensive Golden Dataset.

- Automated evaluation and benchmarking infrastructure with measurable model quality gates.

- Clear, reproducible evidence of model performance against public, proprietary BFSI and relevant competing-model benchmarks.

- A continuous improvement loop that converts evaluation findings into measurable gains across successive Vectro releases.

Key Performance Indicators :

- Training pipeline reliability and experiment throughput.

- Training data quality and coverage.

- Golden Dataset and evaluation coverage.

- Benchmark reproducibility and model performance improvement.

- Accuracy of error diagnosis and effectiveness of regression detection.

- Time required to complete training, evaluation and benchmarking cycles.

- Compliance with model release quality criteria.

Why Join QubeLabs?

- Build the proprietary Vectro Series of enterprise and financial-services language models.

- Develop model training, evaluation and benchmarking infrastructure from the ground up.

- Work closely with the Head of AI and Research Partners.

- Solve real-world AI challenges across financial services.

- Develop proprietary datasets, benchmarks and model improvement systems.

- Contribute directly to production AI systems serving customers across India and Europe.

Preferred Candidate Profile :

- B.Tech, M.Tech, MS or PhD in Computer Science, AI, ML, Mathematics or a related field.

- 3 - 5 years of relevant experience with strong practical skills in Python, PyTorch, LLM fine-tuning and training data pipelines.

- Working knowledge of tokenization, evaluation datasets, benchmarking and regression testing.

- Experience with open-source models such as Llama, Qwen, Mistral or similar is preferred.

- Ability to independently build reliable model training and evaluation systems from scratch.

QubeLabs is an equal opportunity employer and values diversity and inclusion. We encourage applications from qualified candidates regardless of background or identity.

Skills

LLM, Artificial Intelligence, LORA, NLP, Machine Learning, Deep Learning, Generative AI, Python, PyTorch

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,438,441 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Delhi
AI Engineer 19 days ago
≈ $95k – $215k per year (Estimated) • In office • Full-Time • Geneva
Python
Python
FastAPI
Databases
Qdrant
AI/ML
LangChain
Qwen
vLLM
Langfuse
Pydantic AI
TGI
Llama
RAG
DevOps
Prometheus
GitLab CI
Docker
Kubernetes
Grafana
SLI/SLO/SLA
GitLab
Apply
≈ $100k – $249k per year (Estimated) • Remote (Europe, Israel) • PhD • Israel • United Kingdom
Python
AI/ML
Post-training
Model Distillation
Machine Learning
Apply
≈ $100k – $249k per year (Estimated) • Remote (Europe) • PhD • Prague
Python
AI/ML
Post-training
Model Distillation
Machine Learning
Apply
≈ $103k – $258k per year (Estimated) • In office • 6+ years exp • Tel Aviv
Python
Databases
Neo4j
AI/ML
Embeddings
AI Agents
NER
Reranking
Knowledge Graph
Apply
≈ $101k – $251k per year (Estimated) • Remote (Europe, Israel) • Bachelor's Degree • Amsterdam • Prague
Python
AI/ML
Flash Attention
Reinforcement Learning
Quantization
AI Agents
LLM
DPO
PPO
Post-training
FSDP
Model Distillation
Reward Modeling
Machine Learning
DevOps
CI/CD
Apply
≈ $82k – $168k per year (Estimated) • In office • Full-Time • 6+ years exp • Dublin
Python
Python
FastAPI
Databases
PostgreSQL
Neo4j
pgvector
Pinecone
AI/ML
LangChain
Claude
LlamaIndex
Scikit-learn
Prompt Engineering
AI Agents
NLP
Llama
Mistral
Transformers
TensorFlow
Pandas
NumPy
PyTorch
Gemini
LLM
RAG
OpenAI
Anthropic
LLM Guardrails
DevOps
OpenShift
Azure DevOps
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
Kubernetes
Apply
SDLC AI Engineer 1 day ago
≈ $93k – $247k per year (Estimated) • In office • Full-Time • London
Python
Go
Java
Rust
TypeScript
AI/ML
Claude
Claude Code
Model Context Protocol
AI Agents
OpenAI Codex
Agentic Workflows
DevOps
CI/CD
GitHub
Management
Agile
Apply
≈ $26k – $76k per year (Estimated) • In office • Beijing
Python
Java
Scala
Databases
Apache Kafka
AI/ML
Spark
Flink
Apply
≈ $101k – $202k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Dallas
Python
MATLAB
Apply
$57k – $87k per year • In office • 1+ year exp • Bachelor's Degree • Spokane
Python
Management
Outlook
Apply
$10k – $16k per year (gross) • In office • Full-Time • 4+ years exp • Delhi
AI/ML
AI Agents
LLM
RAG
Apply
≈ $15k – $36k per year (Estimated) • Remote (India) • Delhi
JavaScript
AI/ML
LocalAI
Tokenization
LLM Evaluation
Edge AI
Frontend
React.js
Web3
Bitcoin
Management
Telegram
WhatsApp
Apply
In office • Full-Time • Delhi
Cybersecurity
ISO 27001
Apply
≈ $8k – $23k per year (Estimated) • In office • Full-Time • Delhi
Cybersecurity
ISO 27001
Apply
Actuarial 1 day ago
In office • Full-Time • Delhi
Python
SQL
Analytics
Microsoft Excel
Apply
≈ $16k – $33k per year (Estimated) • In office • 4+ years exp • Master's Degree • Delhi
Apply
See all jobs
This is one of many
1,438,441 more open roles from verified company boards, updated every day.