658,489open jobs
38,337companies
96,793added this week
Browse all
Salary
$26k – $59k per year (Estimated)
Location
Remote/Hybrid (Pune, India)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Michelin (Compagnie Générale des Établissements Michelin SCA) is a multinational industrial manufacturing enterprise specializing in tire technology, mobility services, and high-tech materials. Headquartered in Clermont-Ferrand, France, the publicly traded company designs, manufactures, and distributes tires for passenger vehicles, commercial logistics fleets, aircraft, two-wheel transport, and heavy industrial machinery.
Functional AI Tester - GenAI

- - - - - - - - - - - -

We are seeking a Quality Assurance (QA) Engineer focused on testing Generative AI (GenAI) applications with a strong emphasis on Python-based test automation, GenAI evaluation, and ETL/data quality validation. You will design and execute end-to-end test strategies that ensure our AI solutions are accurate, reliable, safe, and compliant.

About the Role

You will be involved in QA for GenAI features including Retrieval-Augmented Generation (RAG), conversational AI and Agentic evaluations. The role centers on:

  • Systematic GenAI evaluation (qualitative and quantitative metrics)

  • ETL and data quality testing for the data flows that feed AI systems

  • Python-driven automated testing

This position is hands-on and collaborative, partnering with AI engineers, data engineers, and product teams to define measurable acceptance criteria and ship high-quality AI features.

Key Responsibilities

  • Test strategy and planning

    • Define risk-based test strategies and detailed test plans for GenAI features.

    • Establish clear acceptance criteria with stakeholders for functional, safety, and data quality aspects.

  • Python test automation

    • Build and maintain automated test suites using Python (e.g., PyTest, requests).

    • Implement reusable utilities for prompt/response validation, dataset management, and result scoring.

    • Create regression baselines and golden test sets to detect quality drift.

  • GenAI evaluation

    • Develop evaluation harnesses covering factuality, coherence, helpfulness, safety, bias, and toxicity etc.

    • Design prompt suites, scenario-based tests, and golden datasets for reproducible measurements.

    • Implement guardrail tests including prompt-injection resilience, unsafe content detection, and PII redaction checks.

    • Track quality metrics over time.

  • RAG and semantic retrieval testing

    • Verify alignment between retrieved sources and generated answers.

    • Verify adversarial tests.

    • Measure retrieval relevance, precision/recall, grounding quality, and hallucination reduction.

  • API and application testing

    • Test REST endpoints supporting GenAI features (request/response contracts, error handling, timeouts).

  • ETL and data quality validation

    • Test ingestion and transformation logic; validate schema, constraints, and field-level rules.

    • Implement data profiling, reconciliation between sources and targets, and lineage checks.

    • Verify data privacy controls, masking, and retention policies across pipelines.

  • Non-functional testing

    • Performance and load testing focused on latency, throughput, concurrency, and rate limits for LLM calls.

    • Cost-aware testing (token usage, caching effectiveness) and timeout/retry behavior validation.

    • Reliability and resilience checks including error recovery and fallback behavior.

  • Share results and insights; recommend remediation and preventive actions.

Required Qualifications

  • Experience

    • 5+ years in software QA, including test strategy, automation, and defect management.

    • 2+ years testing AI/ML or GenAI features, with hands-on evaluation design.

    • 4+ years testing ETL/data pipelines and data quality.

  • Technical skills

    • Python: Strong proficiency building automated tests and tooling (PyTest, requests, pydantic or similar).

    • API testing: REST contract testing, schema validation, negative testing.

    • GenAI evaluation: crafting prompt suites, golden datasets, rubric-based scoring, and automated evaluation pipelines.

    • RAG testing: retrieval relevance, grounding validation, chunking/indexing verification, and embedding checks.

    • ETL/data quality: schema and constraint validation, reconciliation, lineage awareness, data profiling.

  • Quality and governance

    • Understanding of LLM limitations and methods to detect/reduce hallucinations.

    • Safety and compliance testing including PII handling and prompt-injection resilience.

    • Strong analytical and debugging skills across services and data flows.

  • Soft skills

    • Excellent written and verbal communication; ability to translate quality goals into measurable criteria.

    • Collaboration with AI engineers, data engineers, and product stakeholders.

    • Organized, detail-oriented, and outcomes-focused.

Nice to Have

  • Experience with evaluation frameworks or tooling for LLMs and RAG quality measurement.

Experience creating synthetic datasets to stress specific behaviors.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
658,489 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Pune
$170k – $300k per year • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Tampa
Python
Java
Scala
Databases
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
LangGraph
LangChain
Hadoop
Spark
Flink
RAG
DevOps
GCP
Azure
CI/CD
AWS
Amazon S3
Amazon Kinesis
Management
Agile
Apply
$48k per year • In office • Full-Time • Bachelor's Degree • Ottawa
Python
JavaScript
PowerShell
C#
AI/ML
Copilot
ChatGPT
DevOps
gRPC
Git
GitHub
Management
Jira
Apply
$34k – $69k per year (Estimated) • In office • Internship • Bachelor's Degree • United States
Python
Chips/EDA
Altium Designer
Design
SolidWorks
AutoCAD
Apply
$34k – $69k per year (Estimated) • In office • Internship • Bachelor's Degree • United States
Python
Chips/EDA
Altium Designer
Design
SolidWorks
AutoCAD
Apply
$53k – $159k per year (Estimated) • In office • Internship • Dayton
Python
SQL
Marketing
GA4
Apply
$29k – $43k per year (Estimated) • In office • 1+ year exp • Salon-de-Provence
Apply
Warehouse Associate 8 hours ago
$42k per year • In office • Full-Time • Channahon
Apply
$46k – $118k per year (Estimated) • In office • Full-Time • Regensburg
Apply
$30k – $73k per year (Estimated) • Remote/Hybrid • Romania
Analytics
Microsoft Excel
Apply
$48k per year • In office • Full-Time • Piedmont
Apply
$24k – $56k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
DevOps
SLI/SLO/SLA
Apply
$12k – $28k per year (Estimated) • In office • Full-Time • 1+ year exp • Pune
Python
SQL
DevOps
Rest API
Git
Apply
$16k – $37k per year (Estimated) • In office • Full-Time • 1+ year exp • Pune
Python
Java
C#
C#
.NET
Databases
Weaviate
Pinecone
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
Chain-of-Thought
AI Agents
Semantic Kernel
CrewAI
Gemini
LLM
RAG
Semantic Search
GPT-4
Structured Outputs
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
GitOps
Git
AWS
Docker
Kubernetes
Vector
Management
ServiceNow
Apply
AI / ML Engineer 7 hours ago
$18k – $47k per year (Estimated) • In office • Full-Time • 4+ years exp • Hyderabad • Pune • Mumbai • Bengaluru
Python
AI/ML
Fine-tuning
Scikit-learn
Transformers
Pandas
NumPy
PaddlePaddle
PyTorch
PaddleOCR
Hugging Face
OCR
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Chennai • Hyderabad • Pune • Kolkata
Python
SQL
Python
Flask
Django
AI/ML
LangChain
Management
Agile
Apply
See all jobs
This is one of many
658,489 more open roles from verified company boards, updated every day.