386,234open jobs
10,126companies
50,599added this week
Browse all
Salary
$129k – $251k per year (Estimated)
Location
Remote/Hybrid (Birmingham, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
General Parts Group (GenPT) is an independent provider of commercial food service equipment support, offering repair, installation, and planned maintenance services. The company supplies genuine OEM replacement parts for commercial kitchens, serving restaurants, institutions, and foodservice operators across the United States. Through its network of certified technicians and distribution hubs, it helps businesses minimize equipment downtime and maintain compliance with safety standards.

Senior AI Test Automation Engineer

Summary

The Senior AI Test Automation Engineer designs, develops, maintains, and executes automated testing and evaluation solutions for both traditional software and LLM-powered applications. Operating in a forward-deployed capacity, this role works directly with users and delivery teams to capture real-world usage patterns and feedback, translating them into a continuously growing evaluation framework that validates LLM behavior against the consistent flows users actually follow. This role partners with delivery teams, QA, development, product, AI engineering, and the QA Center of Excellence (CoE) to establish scalable automation and evaluation practices, increase test and eval coverage, and integrate quality controls throughout the software delivery lifecycle ,from pre-deployment regression testing through production observability.

You must be eligible to work in the US without Visa Sponsorship.

Responsibilities

LLM Evaluation & Observability (Core)

  • Act as a forward-deployed quality engineer: engage directly with users and stakeholders to collect feedback, observe real usage patterns, and identify the consistent flows users follow through LLM-powered features.

  • Translate user feedback and production traces into curated evaluation datasets in LangSmith, and continuously expand the eval framework as new feedback, edge cases, and failure modes are discovered.

  • Design, build, and maintain offline evaluation suites (regression, benchmarking, and backtesting) that gate prompt, model, and LangGraphworkflow changes before deployment.

  • Develop and calibrate evaluators ,heuristic/code-based checks, LLM-as-judge evaluators, and pairwise comparisons ,and validate judge reliability against human review.

  • Instrument and maintain end-to-end tracing across LangGraphagents and workflows using LangSmith, ensuring trace coverage, quality, and useful metadata for debugging and analysis.

  • Manage annotation queues and human-in-the-loop feedback workflows, routing interesting or problematic production runs to reviewers and feeding results back into datasets and evaluator calibration.

  • Analyze agent trajectories and multi-step LangGraphexecutions (tool calls, state transitions, retrieval steps) to pinpoint failure points and distinguish nondeterministic LLM variance from genuine product defects.

  • Integrate eval runs into CI/CD pipelines so that dataset versions, experiments, and quality thresholds provide automated feedback on every relevant change.

  • Support the adoption of production auditing and monitoring capabilities - such as online evaluations on live traffic, quality drift detection, and alerting - to help teams detect issues in production (supportive to the role, not its core focus).

Test Automation (Core)

  • Design, develop, and implement automated test scripts for UI, API, integration, and regression testing, including deterministic E2E coverage of LLM-powered application surfaces.

  • Integrate automated tests into CI/CD pipelines to enable timely feedback and continuous quality validation.

  • Collaborate with QA, development, product, and business teams to translate requirements, acceptance criteria, and expected agent behaviors into effective automated test and eval coverage.

  • Analyze and triage automation test failures, differentiating framework or script issues from valid product defects ,including the added dimension of expected LLM nondeterminism.

  • Support test data management (including eval dataset versioning, splits, and provenance) and help identify or resolve test-environment stability issues.

  • Report on test execution results, eval experiment outcomes, automation and eval coverage, quality trends, and risks.

  • Participate in code reviews and contribute to automation and evaluation standards, reusable components, and best practices.

  • Engage with the QA CoEto align automation and AI evaluation practices with enterprise standards while contributing domain-specific feedback, lessons learned, and continuous-improvement opportunities.

Required Experience

  • 5+ years of experience in test automation engineering, software quality assurance, or a related role.

  • 1-2+ years of hands-on experience testing or evaluating LLM-powered applications, including building eval datasets, defining pass/fail criteria for nondeterministic outputs, and using LLM-as-judge or heuristic evaluators.

  • Hands-on experience with LangSmithfor tracing, evaluation, and observability of LLM applications ,including creating datasets, running experiments, and configuring evaluators.

  • Hands-on experience with LangGraph, including graph-based agent workflows, state and context management, and tool calling.

  • Hands-on experience building and maintaining automated test suites using Playwright, preferably with TypeScript.

  • Working proficiency in Python and/or TypeScript sufficient to author custom evaluators, tracing instrumentation, and test code.

  • Experience automating UI and API testing, andexperience with API testing tools or frameworks.

  • Familiarity with CI/CD tools such as Azure DevOps, Jenkins, GitHub Actions, or equivalent, including integrating eval runs as pipeline gates.

  • Working knowledge of source-code version control, including Git.

  • Understanding of Agile/Scrum delivery practices and participation in Agile ceremonies.

  • Ability to analyze requirements and user feedback, identify test and eval scenarios, and create maintainable automated coverage.

  • Strong problem-solving, troubleshooting, communication, and collaboration skills ,including comfort engaging directly with end users to gather feedback in a forward-deployed capacity.

Preferred / Nice to Have

  • Ability to understand and document end-to-end business processes, user journeys, and operational workflows.

  • Partner with business stakeholders and end users to translate process requirements into test scenarios, acceptance criteria, and LLM evaluation datasets.

  • Identify process exceptions, edge cases, dependencies, and risks that may affect application or agent behavior.

  • Validate that automated workflows and LLM-powered features produce outcomes aligned with defined business rules and user needs.

  • Use production feedback and observed user behavior to continuously refine process coverage, test automation, and evaluation frameworks.

  • Experience with online evaluations, production monitoring, and quality drift detection for LLM applications.

  • Experience with agent trajectory evaluation, RAG evaluation (retrieval relevance, groundedness, hallucination detection), or guardrails validation.

  • Familiarity with prompt engineering and prompt versioning workflows, andevaluating the impact of prompt or model changes.

  • Understanding of statistical approaches to nondeterministic testing (multiple-run sampling, confidence thresholds, summary metrics across datasets).

  • Experience authoring BDD/Gherkin scenarios using Cucumber or a similar framework.

  • Experience with contract testing tools.

  • Familiarity with SAFeand Agile Release Train (ART) practices.

  • Experience with AI-assisted engineering practices (e.g., using AI coding agents to accelerate test and eval development).

  • Experience with performance, accessibility, mobile, or security test automation.

  • Experience with test management and defect-tracking tools, such as Azure DevOps.

Not the right fit? Let us know you're interested in a future opportunity by joining our Talent Community on jobs.genpt.com or create an account to set up email alerts as new job postings become available that meet your interest!

GPC conducts its business without regard to sex, race, creed, color, religion, marital status, national origin, citizenship status, age, pregnancy, sexual orientation, gender identity or expression, genetic information, disability, military status, status as a veteran, or any other protected characteristic. GPC's policy is to recruit, hire, train, promote, assign, transfer and terminate employees based on their own ability, achievement, experience and conduct and other legitimate business reasons.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,234 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Birmingham
$34k – $74k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bengaluru
DevOps
Azure
Azure DevOps
CI/CD
GitHub
Management
Confluence
Jira
Marketing
LinkedIn
Apply
$42k – $107k per year (Estimated) • Remote • Full-Time • 3+ years exp • Guadalajara
Python
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Azure
CI/CD
GCP
Git
Analytics
ETL/ELT
Power BI
Tableau
Apply
$29k – $77k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
Java
Java
Gradle
Maven
DevOps
CI/CD
Configuration Management
Git
GitLab
GitLab CI
Jenkins
Cybersecurity
SBOM
Management
Confluence
Jira
Marketing
LinkedIn
Apply
$29k – $76k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
Java
DevOps
CI/CD
Git
GitLab
Jenkins
Management
Confluence
Jira
Marketing
LinkedIn
Apply
$160k – $175k per year • In office • Full-Time • Tempe
C++
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
GCP
IAM
Terraform
Terragrunt
Web3
Chainlink CCIP
Chainlink
Apply
$147k – $325k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Birmingham • Atlanta
AI/ML
AI Agents
DevOps
IAM
Cybersecurity
GDPR
Least Privilege
Apply
$88k – $177k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Birmingham
Apply
In office • Internship • Bachelor's Degree • Birmingham
SQL
QA
Selenium
Apply
In office • Internship • Bachelor's Degree • Birmingham
Java
TypeScript
JavaScript
Node JS
Node JS
Nest.JS
Frontend
Angular
Next.js
React.js
DevOps
CI/CD
GCP
Git
Apply
In office • Internship • Bachelor's Degree • Birmingham
Java
DevOps
CI/CD
GCP
Git
Apply
$88k – $177k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Birmingham
Apply
$135k – $244k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Birmingham • Dallas
SQL
DevOps
Incident Management
Analytics
ETL/ELT
Apply
$80k – $181k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Denver • Cleveland • Birmingham
Apply
$79k – $159k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pittsburgh • Birmingham • Cleveland
Apply
Software Engineer 1 day ago
$87k – $172k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Pittsburgh • Birmingham • Dallas
C#
C++
JavaScript
Python
Apply
See all jobs
This is one of many
386,234 more open roles from verified company boards, updated every day.