685,924open jobs
39,773companies
97,229added this week
Browse all
Salary
$670k per year
Location
Remote (United Kingdom, Pakistan)
Overview
Company
Impact
Profile match
Manage jobs, engineers, assets, compliance, quotes and invoices in one system. Cloud-based field service software built for UK trades. Book a free demo today.

The Joblogic Story

Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 2013, we have grown from ~£500K ARR to ~£35M+ ARR and expanded our team from 11 to 500+ people.

Recently, we secured a strategic growth investment from Vista Equity Partners - a global technology investor specialising in enterprise software. This investment includes over £100 million in new primary capital and will fuel our next phase of growth by accelerating our AI-first roadmap, expanding our platform into CAFM (Computer-Aided Facilities Management) capabilities, and supporting our expansion across Europe and beyond.

With Vista’s backing, we’re transforming from a successful UK business into a global scaling SaaS rocket ship, and we’d love for you to join us on our journey to £100M ARR across international markets.

Joblogic provides software to service contractors who install and maintain the built environment. Our platform helps businesses streamline operations, improve profitability, ensure compliance, and achieve rapid growth. With over 100,000 users across industries including HVAC, plumbing, electrical maintenance, facilities management, and building fabric maintenance, we are entering a new era of intelligent automation, predictive maintenance, and data-driven decision-making for service firms.

About the Role  

We are building Joblogic’s AI Agent Platform - a multi-tenant system for designing, versioning, evaluating, and running AI agents that work across email, voice, SMS, WhatsApp, and CRM channels on behalf of our customers. The platform is built on a LangGraph runtime with retrieval over Azure AI Search, a real-time voice stack, human-in-the-loop review queues, and an evaluation harness backed by LangSmith and PromptFoo. 

We are looking for an AI Evaluation Engineer to own quality for everything we ship that has a model in it. You will define what “good” means, measure it rigorously, find out why it is not met, and drive the fixes. The role is deliberately hybrid: you design the rubrics, datasets, and judges, and you build the harnesses and release gates that run them in CI and against live production traffic. Evaluation is not a reporting function here - it is the mechanism by which agent quality improves release over release, and you own it. 

You will work closely with the engineers building agents, the data team, and product, and your work will directly determine what tens of thousands of field-service businesses experience when an agent answers on their behalf. 

What You’ll Do  

  • Own release sign-off - gate every agent version on a green regression suite, judge-scored evaluations within agreed error bars, and a red-team pass - no promotion without them. 
  • Run error analysis - review sampled production traces across chat, email, WhatsApp, and voice every week, maintain the failure taxonomy, and turn new failure modes into dataset examples within the sprint. 
  • Build datasets and rubrics - design and maintain golden datasets and grading rubrics for multi-turn, tool-using agents, promoting interesting production runs into datasets in LangSmith. 
  • Build and calibrate LLM judges - align judges to human labels, re-label a held-out set each cycle, publish agreement per rubric, and retire or retrain judges that drift. 
  • Run offline and online evaluation - regression suites in CI alongside online scoring of sampled production traffic, with clear pass/fail thresholds and drift tracking. 
  • Triage regressions to root cause - attribute failures to prompt, tool, retrieval, model upgrade, or speech provider using paired statistics rather than aggregate deltas, and hand engineers concrete fixes. 
  • Own voice quality metrics - define and gate call-outcome metrics - task success, containment, barge-in recovery, word error rate under noise - alongside latency budgets. 
  • Evaluate classical ML models - set acceptance criteria, slice-level thresholds, calibration checks, and drift alerts for the vision, speech, and predictive models the platform depends on. 
  • Run adversarial testing - probe for prompt injection through inbound email and messaging content, jailbreaks, PII leakage, and tool misuse, and add regression coverage for every finding. 
  • Run the human evaluation programme - own annotation queues, reviewer guidelines, and inter-annotator agreement as an ongoing operation rather than a one-off study, and keep human and automated scores connected. 
  • Collaborate & ship - work in a cross-functional team using tools such as Jira and Slack, write clear documentation, and ship iteratively with a strong quality bar. 

Essential Experience and Skills  

  • 3+ years in roles where evaluating models was the core of the job - ML evaluation, ML quality, applied ML, or data science with an evaluation focus. 
  • Strong Python engineering skills, with experience building test harnesses and clean, well-tested code. 
  • Practical experience evaluating LLM-powered applications or AI agents: building datasets, defining heuristic and LLM-as-judge rubrics, running evaluations, and interpreting results to improve a system. This is a core requirement. 
  • Experience grading tool-using agents on both trajectory and outcome - tool-call correctness, expected-trajectory match, end-state checks - and reporting reliability over repeated trials. 
  • Experience building LLM judges and aligning them to human labels, including agreement statistics and mitigation of position, verbosity, and self-preference bias. 
  • Solid classical ML evaluation foundations: classification, regression, and ranking metrics, cross-validation, calibration, slice-based evaluation, and drift monitoring. 
  • Statistics for small evaluation sets: paired comparisons, confidence intervals, power analysis, and minimum detectable effect - you know roughly how many examples a claim needs before you make it. 
  • Experience with RAG evaluation: faithfulness, groundedness, context precision and recall. 
  • Hands-on experience with an LLM observability and evaluation platform (LangSmith, MLflow, or equivalent): datasets, experiments, custom evaluators, feedback, and wiring evaluations into release gates. 
  • Strong data analysis skills using Pandas, NumPy, and SQL to quantify behaviour and communicate findings. 
  • Working knowledge of how agents are built - prompts, tools, retrieval, and memory - and how they interact, sufficient to root-cause a failure rather than only report it. 
  • Awareness of AI safety and adversarial risk: prompt injection, jailbreaks, data leakage, tool misuse, and responsible-AI practice. 
  • Awareness of the compliance side of evaluation: handling conversation data under UK GDPR, keeping evaluation evidence auditable, and emerging record-keeping expectations for AI systems. 
  • Committed to continuous learning, proactive problem-solving, and timely issue identification, with a keen interest in staying current with a fast-moving field. 
  • Strong communicator, experienced in collaborating with cross-functional teams using tools such as Jira and Slack. 
  • Creative and innovative thinker, consistently contributing fresh ideas and solutions in alignment with current technological trends. 

Nice to Have  

  • Experience with LangGraph and LangSmith specifically (tracing, datasets, online evaluators, judge alignment). 
  • Experience with evaluation tooling such as PromptFoo, DeepEval, RAGAS, or Inspect. 
  • Experience with red-team tooling such as PromptFoo red team, Microsoft PyRIT, or NVIDIA Garak, and familiarity with the OWASP Top 10 for LLM and Agentic Applications. 
  • Experience evaluating voice agents: call simulation, turn-taking and latency budgets, speech recognition accuracy under noise. 
  • Experience evaluating computer-vision or speech models (mAP/IoU, WER/CER) with sample-level failure mining. 
  • Experience with Databricks or AWS SageMaker and experiment tracking with MLflow. 
  • Experience designing human-in-the-loop annotation programmes and measuring inter-annotator agreement. 
  • Experience with Python web frameworks (FastAPI / Flask), pytest-native harnesses, and CI/CD. 
  • Publications, open-source contributions, or public evaluation work. 

What We Offer  

  • Professional Working environment 
  • Market Competitive Salary 
  • Life Insurance & Medical Insurance (Including Family) 
  • OPD 
  • Provident Fund 
  • Gym Facility 
  • Maximum 45 Weekly Hours (Monday-Friday) 
  • Remote Working (During Pandemic Situation) 
  • Company trip 
  • 29 Annual Leaves 
  • 8 Sick & uncapped Compassionate Leaves (As per Company Policy) 
  • Have a chance to work onsite with the UK team 
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
685,924 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$21k – $57k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Aspropyrgos
Python
SQL
PowerShell
Bash
DevOps
Windows Server
Incident Management
Apply
In office • Full-Time • Bengaluru
Python
Java
SQL
Python
Flask
FastAPI
Django
Java
Spring Boot
Hibernate
Spring MVC
Databases
MySQL
PostgreSQL
Oracle
Management
Agile
Scrum
Apply
$59k – $143k per year (Estimated) • In office • Full-Time • 2+ years exp • High School Diploma • Paris
SQL
Cybersecurity
GDPR
Analytics
Power BI
Management
Agile
Marketing
Salesforce
Apply
$21k – $53k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Hyderabad
Python
Python
FastAPI
Databases
Milvus
Pinecone
Qdrant
AI/ML
Prompt Engineering
Arize Phoenix
Langfuse
LangSmith
LLM
TruLens
LLMOps
DevOps
Rest API
CI/CD
Vector
Apply
In office • Full-Time • Bengaluru
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Frontend
Angular
Bootstrap
React.js
DevOps
Rest API
Azure DevOps
Azure
Git
Management
Agile
Scrum
Apply
$670k per year • In office • 1+ year exp • Bachelor's Degree • Lahore
Python
SQL
Databases
MySQL
PostgreSQL
DevOps
Azure
Git
Analytics
Power BI
ETL/ELT
SSIS
Azure Data Factory
Apply
MLOps Engineer 11 hours ago
$670k per year • Remote • Bachelor's Degree
Python
Bash
Python
FastAPI
Celery
Databases
PostgreSQL
Redis
Databricks
AI/ML
Ray Serve
vLLM
MLFlow
ElevenLabs
Computer Vision
AI Agents
SGLang
Arize Phoenix
Langfuse
LangSmith
TensorRT
TensorRT-LLM
LLM
Ray
Braintrust
KServe
Triton
LLMOps
LLM Guardrails
ONNX Runtime
Mobile
Twilio
DevOps
Terraform
Ansible
Helm
GitHub Actions
OpenTelemetry
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Ubuntu
Grafana
Configuration Management
Pipeline as Code
Bicep
Amazon EC2
FinOps
Amazon S3
IAM
Amazon CloudWatch
Cybersecurity
ISO 27001
OWASP Top 10
SOC 2
Least Privilege
Management
Slack
Jira
WhatsApp
QA
k6
Locust
Apply
$670k per year • In office • Birmingham
SQL
Management
Agile
Marketing
HubSpot
Apply
SQL/BI Developer 3 days ago
$670k per year • Remote • Bachelor's Degree • Lahore
SQL
C#
C#
.NET
DevOps
Azure DevOps
Azure
Management
Trello
Jira
Agile
Apply
$670k per year • In office • 1+ year exp • Birmingham
AI/ML
Copilot
Fine-tuning
AI Agents
RAG
Explainable AI
Human-in-the-Loop
Context Engineering
Analytics
A/B Testing
Apply
See all jobs
This is one of many
685,924 more open roles from verified company boards, updated every day.