653,119open jobs
37,986companies
91,654added this week
Browse all
Salary
$100k – $200k per year
Location
In office (Palo Alto)
Employment
Full-Time
Overview
Company
Impact
Profile match

OPPO US Research Center is seeking a full-time meticulous and innovative AI/LLM Test Engineer to join our cutting-edge AI team. In this critical role, you will evaluate the performance, reliability, and safety of Large Language Models (LLMs) in real-world product scenarios and test end-to-end generative AI solutions. Your work will directly shape how users experience AI-powered features by ensuring robustness, accuracy, and alignment with product goals. This is a unique opportunity to pioneer testing methodologies for next-generation AI systems at the forefront of technology.

We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies, execute evaluation workflows, and assist in model performance validation across diverse generative AI use cases.

This contract role is ideal for someone with hands-on experience in AI/ML evaluation, QA engineering, or data analysis who wants to deepen their exposure to generative AI systems.

Requirements

Full-time position requirement:

Core Testing & Evaluation

  • Design and execute performance tests for LLMs across diverse product use cases (e.g., chatbots, content generation etc.).
  • Develop automated test frameworks to evaluate LLM outputs for accuracy, bias, safety, and coherence.
  • Conduct end-to-end testing of integrated generative AI solutions, including APIs, data pipelines, and user interfaces.

Optimization & Validation

  • Collaborate with ML engineers to validate fine-tuned models and optimize prompts for target scenarios.
  • Analyze model failures, edge cases, and adversarial inputs to identify risks and improvement areas.
  • Benchmark LLM performance against industry standards and product-specific KPIs.

Collaboration & Quality Assurance

  • Partner with product, engineering, and research teams to define test requirements and acceptance criteria.
  • Document defects, performance metrics, and test results to drive data-driven improvements.
  • Advocate for AI ethics and safety through rigorous testing of fairness, bias mitigation, and content moderation.

Innovation & Tooling

  • Build scalable tools for synthetic test data generation, prompt variation testing, and automated evaluation workflows.
  • Stay current with advancements in generative AI testing, including red-teaming techniques and evaluation frameworks (e.g., HELM, Dynabench).
  • Propose novel testing strategies for emerging challenges (e.g., hallucinations, context drift).

Basic Qualifications:

  • Bachelor’s degree in Computer Science, Data Science, Engineering, or a related technical field, or equivalent practical experience.
  • 1+ years of experience in software testing, data science, or ML validation, with exposure to AI/ML systems.
  • Proficiency in Python and testing frameworks (e.g., PyTest, Selenium).
  • Hands-on experience evaluating LLMs in production environments (e.g., GPT, Claude, Llama, Gemini).
  • Strong analytical skills for dissecting model behavior, statistical performance, and failure modes.
  • Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling (e.g., MLflow, Weights & Biases).
  • Experience with version control (Git) and agile development methodologies.

Preferred Qualifications:

  • Master’s degree in AI, Machine Learning, or a related field.
  • Expertise in prompt engineering, LLM fine-tuning (e.g., LoRA, RLHF), or optimization techniques.
  • Experience with automated evaluation tools (e.g., LangChain, TruLens) or LLM-specific test suites.
  • Knowledge of data pipelines, SQL/NoSQL databases, and API testing (e.g., Postman).
  • Background in statistics, quantitative analysis, or data visualization for test insights.
  • Contributions to AI safety/ethics initiatives or open-source LLM evaluation projects.
  • Experience testing mobile-integrated AI solutions (Android/iOS).

Contractor position requirements:

Testing & Evaluation Support:

  • Execute pre-defined performance tests for LLMs across various tasks (e.g., summarization, Q&A, chatbot flows).
  • Run scripted evaluations to assess outputs for factuality, coherence, and safety.
  • Perform manual and automated test execution on APIs and LLM-integrated user interfaces.

Prompt & model validation:

  • Assist ML engineers in evaluating prompt variations and prompt-tuning outcomes.
  • Log and analyze failure cases, anomalies, and edge cases based on provided guidelines.

Collabration & Documentation

  • Work with QA leads, product managers, and ML engineers to understand test goals and criteria.
  • Report defects, compile evaluation summaries, and maintain testing logs.

Tooling & Antomation:

  • Use existing internal tools or frameworks to automate test runs and result collection.
  • Contribute to prompt generation, input templating, or result tagging processes.

Basic Qualifications:

  • Bachelor's degree or equivalent work experience in a technical field (e.g., Computer Science, Engineering, Data Science).
  • 6+ months experience in software QA, data labeling, LLM evaluation, or ML testing projects.
  • Basic Python proficiency, especially for data processing and automation tasks.
  • Familiarity with LLMs (e.g., GPT, Claude, Gemini) and prompt-based outputs.
  • Comfortable working with tools like Jupyter, Postman, or testing dashboards.
  • Detail-oriented with good documentation habits.

Contractor Details:

  • Duration: Long term
  • Rate: Commensurate with experience
  • Conversion Opportunity: High-performing contractors may be considered for full-time roles

Benefits

OPPO is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.

The US base salary range for this full-time position is $100,000-$200,000 + bonus + long term incentives benefits. Our salary ranges are determined by role, level, and location.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
653,119 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
In office
Python
TypeScript
DevOps
Terraform
GCP
Azure
CI/CD
Git
AWS
IAM
Cybersecurity
Okta
Zero Trust
Least Privilege
Microsoft Entra ID
Apply
$34k – $74k per year (Estimated) • In office • Full-Time • 13+ years exp • Kochi
Python
JavaScript
TypeScript
AI/ML
Model Context Protocol
AI Agents
Frontend
Lighthouse
DevOps
New Relic
CI/CD
Jenkins
Docker
Grafana
Platform Engineering
Shift-Left
GitHub
GitLab
Cybersecurity
Shift-Left Security
Management
Agile
QA
Selenium
Cypress
Playwright
Gatling
Apply
$24k – $58k per year (Estimated) • In office • Full-Time • 5+ years exp • Kochi
Python
JavaScript
Java
TypeScript
Java
Maven
Gradle
AI/ML
Copilot
Cursor
Windsurf
ChatGPT
Frontend
npm
Mobile
JUnit
DevOps
Azure DevOps
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
GitLab
Cybersecurity
OWASP ZAP
Management
Agile
Scrum
Kanban
QA
TestNG
Selenium
Cucumber
JMeter
Cypress
Playwright
Gatling
Appium
Postman
BrowserStack
Rest-Assured
k6
Apply
$76k – $106k per year • Remote/Hybrid • Full-Time • 3+ years exp • Irving
Python
SQL
Databases
Apache Kafka
AI/ML
Hadoop
Spark
QA
Sentry
Apply
$156k – $260k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Kalamazoo
Python
PowerShell
Cybersecurity
Qualys Cloud Platform
Wiz
CVSS
EPSS
KEV
SBOM
Management
ServiceNow
Apply
AI Engineer 4 months ago
$100k – $200k per year • In office • Full-Time • Master's Degree • Palo Alto
AI/ML
LangChain
Prompt Engineering
Multimodal AI
AI Agents
NLP
Gemini
LLM
RAG
OpenAI
Anthropic
Context Engineering
Tool Use
Apply
Remote/Hybrid • Contractor • Palo Alto
Apply
$150k – $270k per year (Estimated) • In office • Contractor • Bachelor's Degree • Palo Alto
Java
Java
Spring Boot
Databases
Redis
ElasticSearch
Apache Kafka
Apply
$60k – $80k per year • In office • Contractor
DevOps
SLI/SLO/SLA
Analytics
Microsoft Excel
Management
Freshdesk
Apply
$100k – $200k per year • In office • Full-Time • Master's Degree • Palo Alto
Python
Python
Flask
Django
Databases
MySQL
PostgreSQL
Redis
Firestore
Memcached
Apache Kafka
AI/ML
Vertex AI
Gemini
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Google GKE
Google Cloud Run
Apply
$132k – $198k per year • Equity • Remote/Hybrid • Full-Time • Palo Alto
Apply
$167k – $230k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Palo Alto
Python
SQL
Databases
Databricks
AI/ML
dbt
AI Agents
LLM
RAG
Streamlit
Context Engineering
LLM Guardrails
DevOps
Terraform
CI/CD
Git
Platform Engineering
GitHub
GitLab
Analytics
Tableau
Power BI
ETL/ELT
Apply
$172k – $319k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Palo Alto
Apply
$179k – $269k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Computer Vision
TensorFlow
PyTorch
Robotics
Sensor Fusion
Motion Planning
Apply
$160k – $210k per year • Equity • In office • Full-Time • 5+ years exp • Palo Alto
Apply
See all jobs
This is one of many
653,119 more open roles from verified company boards, updated every day.