415,275open jobs
14,693companies
74,504added this week
Browse all
Salary
$100k per year
Location
Remote (United States)
Employment
Contractor
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a AI Evaluators: Assessing A Shopping Assistant based in United States.

This role offers an opportunity to help evaluate and improve the quality of an AI-powered digital shopping assistant.

You will analyze real-world e-commerce interactions to determine whether responses are accurate, logical, useful, and aligned with user needs.

By identifying subtle failures and weaknesses, you will provide structured feedback that directly contributes to improving model performance.

The work combines AI evaluation, quality assurance, e-commerce analysis, and structured data assessment in a practical, research-oriented environment.

You will also develop rubrics and verifiers that create consistent standards for evaluating future AI responses.

This is a sustained remote engagement suited to detail-oriented professionals who enjoy analyzing complex text interactions and shaping better AI experiences.

Accountabilities:

    • Review real user interaction traces with an AI-powered shopping assistant, carefully assessing conversations and responses within a dedicated evaluation platform.
    • Identify logical failures, factual inaccuracies, irrelevant or unhelpful responses, and poor product recommendations, including subtle issues that may negatively affect the shopping experience.
    • Analyze the quality of AI-generated responses from an e-commerce perspective, considering whether recommendations appropriately address user queries and real-world shopping needs.
    • Create structured evaluation rubrics that establish clear, repeatable criteria for judging response accuracy, helpfulness, reasoning quality, and overall usefulness.
    • Develop verifiers and other structured evaluation mechanisms that can consistently assess future responses and help surface recurring model weaknesses.
    • Contribute insights from individual evaluations to broader efforts to improve AI model behavior, response quality, and performance on real-world e-commerce scenarios.
    • Maintain a sustained evaluation workload of at least 20 hours per week while working independently and maintaining a high level of accuracy and consistency.
    • Requirements:

      • Experience in data evaluation, quality assurance, AI training, data annotation, software testing, prompt engineering, or a closely related analytical discipline.
      • Strong analytical and critical-thinking skills, with the ability to identify subtle logical errors, inaccuracies, inconsistencies, and quality issues within written AI interactions.
      • Familiarity with e-commerce search, online shopping journeys, product discovery, recommendations, and digital shopping experiences.
      • Ability to analyze complex text interactions in depth and distinguish between technically correct responses and responses that are genuinely useful to the user.
      • Experience creating structured evaluation criteria, annotation frameworks, testing methodologies, rubrics, or similar quality-assurance systems is valuable.
      • Strong attention to detail and consistency, with the ability to apply evaluation standards objectively across a high volume of interactions.
      • Ability to work independently in a remote environment, learn new evaluation tools and processes, and communicate findings clearly.
      • Availability to commit to a sustained workload of 20+ hours per week.
      • Benefits:

        • Compensation of $50 USD per hour.
        • Fully remote work, providing flexibility to complete evaluation activities from within the United States.
        • Sustained part-time engagement requiring 20+ hours per week, allowing for a consistent workload.
        • Opportunity to contribute directly to the evaluation and improvement of AI-powered shopping technology.
        • Hands-on exposure to AI evaluation, model quality assessment, e-commerce interactions, structured rubrics, and verification frameworks.
        • Opportunity to apply expertise in quality assurance, data evaluation, e-commerce, software testing, or AI training to real-world AI development.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
415,275 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Quality Engineer 3 2 hours ago
$77k – $118k per year • Remote • 7+ years exp
JavaScript
TypeScript
AI/ML
AI Agents
Claude
Copilot
Cursor
Prompt Engineering
RAG
DevOps
Azure
Azure DevOps
CI/CD
Git
GitHub
GitHub Actions
Jenkins
Management
Microsoft Teams
QA
JMeter
k6
Playwright
Apply
In office • 5+ years exp • Bachelor's Degree
JavaScript
PHP
Python
SQL
TypeScript
Databases
FAISS
Pinecone
PostgreSQL
AI/ML
AI Agents
Embeddings
Hugging Face
LangChain
LlamaIndex
LLM
OCR
OpenAI
Prompt Engineering
RAG
DevOps
AWS
Azure
Azure DevOps
CI/CD
Datadog
Docker
GCP
Git
GitHub
GitHub Actions
Jenkins
Kubernetes
Prometheus
Vector
Chips/EDA
PoC Library
Management
Outlook
Apply
In office • 5+ years exp • Bachelor's Degree
PHP
Python
SQL
AI/ML
AI Agents
ChatGPT
Copilot
Prompt Engineering
Analytics
Power BI
Management
Jira
Smartsheet
Apply
$76k – $214k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Singapore
SQL
Databases
MS SQL
Snowflake
AI/ML
AI Agents
Prompt Engineering
RAG
DevOps
CI/CD
Vector
Analytics
ETL/ELT
Apply
AI Architect 3 hours ago
$144k – $152k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree
Python
Databases
Amazon Redshift
DynamoDB
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
Claude
Human-in-the-Loop
LLM
LLM Guardrails
Prompt Engineering
RAG
Tool Use
DevOps
Amazon S3
AWS
AWS Lambda
Azure
GCP
IAM
Vector
Cybersecurity
AWS WAF
Apply
$30k – $72k per year (Estimated) • Remote • Full-Time • 10+ years exp
Apply
$59k – $141k per year (Estimated) • Remote • Full-Time • 10+ years exp
Apply
$35k – $89k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
AI/ML
Amazon SageMaker
Evidently AI
MLFlow
PyTorch
Recommender Systems
TensorFlow
DevOps
AWS
CI/CD
GitLab
GitLab CI
Grafana
Prometheus
Apply
$53k – $131k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
AI/ML
Amazon SageMaker
Evidently AI
MLFlow
PyTorch
Recommender Systems
TensorFlow
DevOps
AWS
CI/CD
GitLab
GitLab CI
Grafana
Prometheus
Apply
$11k – $83k per year (Estimated) • Remote/Hybrid • Internship
Apply
See all jobs
This is one of many
415,275 more open roles from verified company boards, updated every day.