368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$33k – $73k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Tekion is an end-to-end, AI-native automotive retail platform that unifies dealership operations - DMS, CRM, digital retail, service, payments and analytics - on a single cloud-based operating system. Tekion powers 3,000+ dealerships and has processed $43B+ in transactions.

About Tekion:

Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud-native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting-edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe.

About the Role

We are looking for a highly motivated Senior SDET - AI Evaluation to join Tekion’s AI Platform team. Evaluation is the backbone of trustworthy AI: as Tekion scales from a handful of AI agents

to 100+ across Service, Sales, F&I, and Analytics, this role builds the evaluation platform and frameworks that let every ML team measure, trust, and improve the quality of AI outputs.In this role, you will be responsible for defining and building Tekion’s AI evaluation capabilities as a shared platform service. You will work closely with ML Engineers, Data Scientists, the AI Platform team, and Product Management to design evaluation datasets, automated scoring pipelines, and quality metrics that quantify the accuracy, consistency, and safety of AIgenerated outputs across the organization. You will own the systems that answer “is this model or agent good enough to ship, and is it

staying good in production?” - from offline benchmarks and LLM-as-judge pipelines to online evaluation and continuous quality monitoring. You will also use AI and LLMs to scale evaluation itself, building automated judges and synthetic datasets that expand coverage faster than manual review ever could.

What You’ll Do

  • Develop a deep understanding of Tekion’s AI agents, ML models, and the quality dimensions that matter for each business domain.

  • Design, enhance, and own Tekion’s AI evaluation infrastructure as a shared capability used across ML teams.

  • Create, curate, and maintain evaluation datasets (evals) and golden/ground-truth sets across use cases and domains.

  • Define quality metrics for AI outputs - accuracy, relevance, faithfulness/groundedness, consistency, safety, and task success.

  • Build automated scoring pipelines, including LLM-as-judge, rubric-based, and referencebased evaluation methods.

  • Validate user intents and measure response accuracy and consistency for AI-powered capabilities such as the Analytics Agent.

  • Identify hallucinations, unsafe or biased outputs, and edge cases; design targeted eval suites to catch them.

  • Build both offline evaluation (pre-release benchmarking) and online evaluation (production quality monitoring, A/B, drift detection).

  • Establish evaluation gates in CI/CD so model, prompt, or data changes are quality-checkedbefore release.

  • Develop dashboards and reporting that make AI quality visible and actionable for ML and product teams.

  • Use AI/LLMs to scale evaluation - automated judges, synthetic data generation, and eval tooling

  • Champion evaluation and responsible-AI quality best practices across the organization.

What We’re Looking For

  • 5-8 years in SDET, quality engineering, ML engineering, or data science, with hands-on experience building evaluation or measurement systems - or a strong SDET background with deep LLM/ML fluency.

  • Strong programming skills in Python, with the ability to build robust, reusable evaluation pipelines and tooling.

  • Deep understanding of ML/LLM evaluation, benchmark design, and the pitfalls of evaluating non-deterministic systems.

  • Hands-on experience with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation.

  • Solid grasp of LLM/agent concepts - prompting, RAG, embeddings, tool use - and generative failure modes (hallucination, drift, prompt sensitivity, bias).

  • Experience designing and curating datasets, including labeling/annotation strategy and data quality.

  • Strong statistical intuition for interpreting eval results and significance.

  • Excellent communication skills to translate quality signals into decisions for ML and product teams.

Nice to Have

  • Experience with eval frameworks/tools such as Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, or provider eval suites.
  • Experience building online evaluation, guardrails, or production model monitoring.
  • Familiarity with responsible AI / safety evaluation and red-teaming.
  • Experience with experiment tracking (MLflow, Weights & Biases) and A/B testing.
  • Prior work standing up evaluation as a platform capability for multiple teams

Effective 4 Aug 2026, Current Tekion Employees should apply via the Internal Job Board in Ashby

Tekion is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, victim of violence or having a family member who is a victim of violence, the intersectionality of two or more protected categories, or other applicable legally protected characteristics.

For more information on our privacy practices, please refer to our Applicant Privacy Notice here.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$117k – $177k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Chicago • Dallas
Apex
C#
Java
Python
Apex
Salesforce Data Cloud
AI/ML
Agentforce
AI Agents
Chain-of-Thought
Embeddings
Few-Shot Learning
Fine-tuning
Hallucination
LLM Guardrails
NLP
Prompt Engineering
RAG
RLHF
Semantic Search
Semantic Search
Frontend
GraphQL
DevOps
CI/CD
Git
Analytics
A/B Testing
Marketing
Salesforce
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 2 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
$96k – $134k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • New York
JavaScript
Swift
TypeScript
Java
Java
Spring Framework
Databases
Apache Kafka
PostgreSQL
AI/ML
AI Agents
Claude
Copilot
Fine-tuning
Flink
LangChain
LangGraph
Llama
LlamaIndex
Prompt Engineering
PyTorch
RAG
TensorFlow
Transformers
Devin
Hugging Face
OpenAI
Frontend
Angular
React.js
Mobile
MVC
DevOps
AWS
CI/CD
Docker
Kubernetes
OpenShift
Splunk
Vector
GitHub
Analytics
Tableau
Apply
$129k – $231k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Atlanta
SQL
AI/ML
AI Agents
Claude
Claude Code
DevOps
Azure
Design
Figma
Apply
$32k – $78k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
AWS
Azure
Incident Management
Kubernetes
Platform Engineering
SLI/SLO/SLA
Terraform
IAM
Cybersecurity
CIS Benchmarks
ISO 27001
NIST CSF
SOC 2
Apply
$166k – $249k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pleasanton
Apply
$34k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Master's Degree • Bengaluru
Apply
$197k – $246k per year • Equity • In office • Full-Time • Pleasanton
SQL
Databases
Databricks
AI/ML
AI Agents
Management
n8n
Zapier
Marketing
HubSpot
Marketo
Salesforce
Apply
$125k – $207k per year • Equity • Remote • Full-Time
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.