744,295open jobs
44,695companies
107,333added this week
Browse all
Salary
≈ $33k – $86k per year (Estimated)
Location
Hybrid (Pune, India)
Seniority
Architect · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. Citi scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Citi is one of the largest banks in the world, tracing its lineage to the City Bank of New York founded in 1812 and taking its modern shape through the 1998 merger that created Citigroup. Its most distinctive asset is a cross-border payments and treasury network unmatched by any competitor, moving trillions of dollars a day for multinational corporations, governments and other banks across roughly ninety countries. Alongside that institutional franchise it runs markets and investment banking, wealth management and a United States personal bank, and has spent recent years simplifying itself by exiting consumer operations across Asia, Europe and Latin America.

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform.This role spans the full spectrum of modern AI quality engineering - fromAgentic AI flow testing andRAG pipeline validation toAI safety, test automation, andperformance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they aresafe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Responsibilities

Agentic AI Testing

  • Design and executeend-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • Validate agent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • Testtool-use correctness - ensuring agents invoke the right tools, with correct parameters, at the right time.
  • Evaluate agent memory systems (short-term, long-term, episodic) for accuracy and context retention across sessions.
  • Validate agent handoff and delegation logic in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • Testtermination conditions, loop detection, and infinite loop prevention in autonomous agent loops.

RAG (Retrieval-Augmented Generation) Testing

  • Design comprehensive test strategies forend-to-end RAG pipelines - covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • Validateretrieval accuracy and relevance - ensuring the correct context chunks are retrieved for a given query.
  • Testembedding model quality and vector similarity thresholds across different document corpora.
  • Evaluatefaithfulness, groundedness, and answer relevance of generated responses using frameworks likeRAGAS, TruLens, DeepEval.
  • Test chunking strategies (fixed, semantic, hierarchical) for their impact on retrieval quality.
  • Validatecontext window management - ensuring retrieved context does not exceed token limits or degrade generation quality.
  • Conductend-to-end regression testing when the underlying knowledge base, embedding model, or LLM changes.

  • Testmulti-turn conversational RAG for context coherence and citation accuracy across turns.

Test Automation

  • Build and maintainautomated test harnesses for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • Developautomated evaluation pipelines integrated into CI/CD workflows for continuous model and agent validation.
  • Create data validation and data quality frameworks (using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.

  • Buildprompt regression suites to detect behavioral drift across LLM versions or prompt changes.
  • Implementdeterminism and reproducibility tests for stochastic LLM-based decisions.
  • Automatevector database validation - index integrity, embedding drift, and retrieval consistency checks.

AI Safety & Security Testing

  • Conductred-teaming and adversarial testing to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • Testoutput guardrails and content filters for unsafe, biased, toxic, or out-of-scope model behavior.
  • Validateprivilege escalation controls - ensuring agents do not exceed permitted actions or access unauthorized resources.
  • Performdata poisoning and backdoor attack simulations to assess model robustness.
  • Evaluate models forbias, fairness, and discrimination using frameworks such as AI Fairness 360 and Aequitas.
  • TestPII leakage and data privacy controls in RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
  • Conduct security testing aligned with the OWASP Top 10 for LLM Applications, including:
    • Prompt Injection (Direct & Indirect)
    • Insecure Output Handling
    • Training Data Poisoning
    • Insecure Plugin / Tool Design
    • Sensitive Information Disclosure
  • Validateconstitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
  • Collaborate with cybersecurity teams onAI-specific threat modeling and vulnerability management.
  • Maintainsafety testing playbooks and document red-team findings with severity ratings and remediation recommendations.

Performance & Reliability Testing

  • Define and executeload, stress, soak, and spike testing for AI-powered APIs, inference endpoints, and agent orchestration services.
  • Measure and optimizeend-to-end latency across RAG and agentic pipelines - from query to final response.
  • Benchmark LLM inference throughput (tokens/second) and identify bottlenecks across model serving infrastructure.
  • Testauto-scaling behavior of AI services under variable load conditions.
  • Validatecircuit breaker, retry, and fallback mechanisms in agentic and RAG systems for graceful degradation.
  • Testvector database performance - query latency, index build time, and retrieval accuracy under high concurrency.
  • Conductcost efficiency analysis - measuring token consumption, API call costs, and infrastructure spend per agent task.
  • Establish SLOs (Service Level Objectives) and SLAs for AI system availability, latency percentiles (P50, P95, P99), and error rates.
  • Collaborate with MLOps teams to set upobservability dashboards, monitoring alerts, and automated anomaly detection for production AI systems.
  • Performchaos engineering experiments to validate agent and RAG system resilience under infrastructure failures.

Domain Knowledge

  • Deep understanding ofRAG architecture patterns - naive RAG, advanced RAG, modular RAG.
  • Solid grasp ofagent design patterns: ReAct, Plan-and-Execute, Reflexion, MRKL, Mixture-of-Agents.
  • Familiarity with AI safety and alignment principles (RLHF, Constitutional AI, guardrail layers).
  • Knowledge oftoken economics, context management, and LLM cost optimization.

  • Proficiency inperformance engineering methodologies for distributed AI systems.

Preferred Qualifications

  • Experience with MCP (Model Context Protocol) or similar agentic communication standards.
  • Exposure to multi-modal agent testing (agents handling text, images, code, documents).
  • Experience in regulated industries (banking, finance, healthcare) with strict compliance requirements.
  • Familiarity with chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos).

Education

  • Bachelor’s degree in Computer Science, Engineering, or a related field.

  • Master’s degree is a plus.

Experience

  • 8+ years of experience in software or AI/ML quality engineering.
  • 3+ years of hands-on experience withRAG systems, or Agentic AI.

  • Proven experience buildingautomated test frameworks for non-deterministic AI systems.
  • Strong background inperformance testing andAI safety/security assessments.

------------------------------------------------------

Job Family Group:

Technology

------------------------------------------------------

Job Family:

Technology Quality

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
744,295 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Pune
≈ $35k – $85k per year (Estimated) • In office • 15+ years exp • Pune
Python
JavaScript
TypeScript
AI/ML
Prompt Engineering
AI Agents
LLM
Mobile
JUnit
DevOps
Azure DevOps
Azure
CI/CD
Jenkins
Git
Shift-Left
Self-Healing
Cybersecurity
Shift-Left Security
Management
Agile
QA
TestNG
Selenium
Playwright
Appium
Postman
Rest-Assured
Apply
≈ $41k – $97k per year (Estimated) • Hybrid • Full-Time • 18+ years exp • Bachelor's Degree • Bengaluru
Python
Go
Java
C++
Python
pySpark
Databases
Apache Kafka
AI/ML
Spark
Function Calling
AI Agents
Flink
Tool Use
Machine Learning
DevOps
GCP
Azure
AWS
Apply
≈ $41k – $98k per year (Estimated) • Hybrid • 15+ years exp • Bengaluru
Python
JavaScript
Databases
SAP HANA
Snowflake
Databricks
AI/ML
Model Context Protocol
AI Agents
RAG
A2A
Knowledge Graph
DevOps
GCP
Azure
AWS
Apply
Principal AI Engineer 8 hours ago
≈ $34k – $81k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Hyderabad
Python
Python
FastAPI
Asyncio
Databases
PostgreSQL
Redis
AI/ML
LangGraph
AutoGen
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
Langfuse
LangSmith
AWS Bedrock
CrewAI
LLM
OpenAI
Anthropic
LLMOps
A2A
GPT-4
Multi-Agent Systems
Tool Use
DevOps
GCP
WebSockets
Azure
CI/CD
Git
AWS
Docker
Cybersecurity
LDAP
Apply
≈ $25k – $59k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Bengaluru
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Vertex AI
Fine-tuning
AI Agents
Arize Phoenix
LangSmith
AWS Bedrock
CrewAI
Gemini
LLM
RAG
Hallucination
Semantic Search
OpenAI
Amazon SageMaker
LLMOps
AWS Bedrock AgentCore
A2A
Human-in-the-Loop
Context Engineering
Semantic Search
LLM Guardrails
Edge AI
Agentic Workflows
Multi-Agent Systems
Tool Use
Vertex AI Agent Builder
Machine Learning
DevOps
GCP
Datadog
Azure
CI/CD
AWS
FinOps
IAM
Cybersecurity
Defense in Depth
Apply
≈ $20k – $42k per year (Estimated) • Hybrid • 3+ years exp • Bachelor's Degree • Moscow
Python
SQL
Python
Flask
FastAPI
Dask
Databases
PostgreSQL
ClickHouse
FAISS
Qdrant
RabbitMQ
Apache Kafka
Greenplum
AI/ML
LangGraph
LangChain
Spark
LlamaIndex
vLLM
Fine-tuning
Scikit-learn
Prompt Engineering
AI Agents
SGLang
TensorRT
Transformers
TensorFlow
Pandas
NumPy
PyTorch
LLM
RAG
Streamlit
Triton
Hugging Face
KV Cache
Machine Learning
DevOps
Rest API
gRPC
GitLab CI
CI/CD
ArgoCD
Git
Docker
Kubernetes
Grafana
Amazon S3
Analytics
Tableau
Power BI
ETL/ELT
Superset
Management
Confluence
QA
Swagger
Apply
≈ $35k – $73k per year (Estimated) • In office • Moscow
Python
SQL
Python
pySpark
AI/ML
LangGraph
LangChain
Spark
AI Agents
LLM
RAG
Apply
≈ $37k – $92k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Vimercate
Python
AI/ML
AI Agents
RAG
Machine Learning
Apply
≈ $23k – $50k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Pune
JavaScript
Java
TypeScript
SQL
Java
Maven
Spring Boot
Hibernate
Spring MVC
Gradle
Databases
PostgreSQL
Redis
Apache Kafka
AI/ML
Copilot
AI Agents
LLM
Frontend
RxJS
Angular
React.js
DevOps
Rest API
GCP
OpenShift
GitHub Actions
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
API Gateway
Cybersecurity
SonarQube
Management
Agile
Apply
≈ $23k – $48k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Pune
Python
SQL
Python
pySpark
Databases
Snowflake
Apache Kafka
Amazon Redshift
AI/ML
Hadoop
Spark
Embeddings
AI Agents
RAG
DevOps
Terraform
GitHub Actions
CI/CD
Jenkins
Git
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon ECS
Analytics
ETL/ELT
Dimensional Modeling
Management
Agile
Apply
GEN AI Developer 2 days ago
≈ $39k – $99k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Master's Degree • India
Python
JavaScript
Node JS
AI/ML
LoRA
Fine-tuning
Embeddings
Quantization
Prompt Engineering
AI Agents
PEFT
Transformers
TensorFlow
PyTorch
RAG
Hugging Face
Management
Agile
Apply
GEN AI Developer 2 days ago
≈ $32k – $82k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Master's Degree • Gurgaon • Bengaluru
Python
JavaScript
Node JS
AI/ML
LoRA
Fine-tuning
Embeddings
Quantization
Prompt Engineering
AI Agents
PEFT
Transformers
TensorFlow
PyTorch
RAG
Hugging Face
Management
Agile
Apply
≈ $34k – $81k per year (Estimated) • Hybrid • Full-Time • 14+ years exp • Bachelor's Degree • Pune • Chennai
Databases
Neo4j
ArangoDB
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
Model Context Protocol
Fine-tuning
Embeddings
Prompt Engineering
AI Agents
NLP
NER
CrewAI
Gemini
RAG
Google ADK
Tokenization
Hybrid Search
OpenAI
Hugging Face
LLMOps
OpenAI Agents SDK
A2A
GraphRAG
Context Engineering
Knowledge Graph
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
OpenTelemetry
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
≈ $33k – $84k per year (Estimated) • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Pune • Chennai
Python
Databases
Neo4j
ArangoDB
AI/ML
LangGraph
LangChain
Claude
LlamaIndex
Model Context Protocol
Fine-tuning
Embeddings
Prompt Engineering
Function Calling
AI Agents
NLP
NER
CrewAI
Gemini
RAG
Google ADK
Tokenization
Hybrid Search
OpenAI
OpenAI Agents SDK
A2A
GraphRAG
Context Engineering
Knowledge Graph
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
DevOps
OpenTelemetry
CI/CD
AWS
Docker
Kubernetes
Apply
≈ $33k – $85k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Pune • Chennai
Python
Python
Flask
FastAPI
Django
AI/ML
RAG
OpenAI
Edge AI
DevOps
CI/CD
Git
Apply
In office • Full-Time • Pune
AI/ML
Prompt Engineering
Apply
AI Engineer 10 hours ago
Remote (India) • Full-Time • Pune
Python
Python
Flask
FastAPI
Databases
Supabase
AI/ML
LangGraph
AutoGen
LangChain
Claude
DSPy
LlamaIndex
LoRA
Model Context Protocol
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
Chain-of-Thought
AI Agents
OpenAI SDK
PEFT
Transformers
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
LLMOps
GPT-4
Agentic Workflows
Multi-Agent Systems
Vercel AI SDK
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Management
n8n
Apply
Principle Accountant 7 hours ago
≈ $24k – $53k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Pune
Management
Microsoft Office
Apply
≈ $32k – $70k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pune
JavaScript
TypeScript
SQL
PowerShell
Node JS
Apex
Node JS
Electron
Databases
SQLite
Frontend
Angular
Mobile
Ionic
Capacitor
Cordova
DevOps
Rest API
Azure DevOps
Azure
Git
Bitbucket
Windows
Apache HTTP Server
Management
UML
QA
Playwright
Apply
≈ $23k – $53k per year (Estimated) • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Pune
Apply
See all jobs
This is one of many
744,295 more open roles from verified company boards, updated every day.