1,289,116open jobs
74,730companies
206,532added this week
Browse all
Salary
≈ $82k – $216k per year (Estimated)
Location
In office (Leeds)
Employment
Contractor

Confirmed on the employer's own hiring board on Oct 7, 2026. First seen by Alion on Aug 4, 2026. CreateFuture scores A on the Alion truth index.

Overview
Company
Impact
Profile match
CreateFuture is an AI transformation and software engineering partner that builds digital products, platforms, and AI-native solutions for enterprise clients. Headquartered in Edinburgh, Scotland, the company assists organizations across sectors like fintech, energy, and government in modernizing legacy technology and scaling intelligent workflows. Its services span business strategy, cloud infrastructure, customer experience design, and data engineering to drive digital growth.

Working at CreateFuture

CreateFuture is an AI-native consulting partner where people do work that matters and are supported to do it well. We work alongside organisations such as PayPal, adidas, NatWest, FanDuel and Money Saving Expert, building digital products and services that make a difference  while always putting people first.

We’re a team of creators. We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people. We work side by side with our clients, challenging what’s not working and helping them to build the future. Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.

Our UK Benefits 

35 days leave (including bank holidays). 

  • Private medical insurance.
  • Enhanced parental and adoption leave. 
  • Financial coaching + 5% pension match.
  • 40 hours of paid learning and development.

 View our full list of UK benefits. 

CreateFuture is a Great Place to Work-Certified™ company and has won Best Workplaces UK multiple years in a row.

Join us on our journey. Let’s create tomorrow, together, today.

About the engagement

Most organisations shipping AI features have no reliable way of knowing whether those features are any good. They have vibes, a demo that went well, and a support queue. We're engaging two Senior AI Engineers to fix that properly on a major AI platform programme with a client in a high-traffic, heavily regulated consumer sector.

You'll own the evaluation and observability layer of the platform. That covers the golden datasets, the scoring pipelines, the eval gates that sit in the promotion path, and the production monitoring that tells the client when quality has quietly drifted. It's the part of AI engineering that decides whether the rest of it can be trusted, and in a regulated business it's also the part that has to stand up to scrutiny.

This is a hands-on role embedded with the client's engineers and product people, and a large part of the job is persuasion. An eval result that nobody acts on has cost the programme money and taught it nothing. You need to be the person who can say "this score is real, here's what we should do" and be believed by a team that didn't ask for your help.

Key Responsibilities

Technical Delivery & Implementation

  • Eval platform: Design and build the evaluation pipelines on a platform such as Braintrust, LangSmith, Arize or Weights & Biases, and own the decision about which one fits the problem.
  • Datasets and scoring: Curate golden datasets that represent real user behaviour. Design LLM-as-judge and human-in-the-loop scoring pipelines with a clear view of where each one is trustworthy and where it isn't.
  • Production observability: Build token-level tracing, quality signal monitoring, hallucination and refusal detection, and per-use-case cost attribution, with alerting that catches degradation before users report it.
  • Quality gates: Wire eval gates into CI/CD promotion paths so a regression blocks a release rather than being discovered in it.
  • Regression and drift: Establish the baselines, regression suites and drift detection that let the client's teams change prompts, models and retrieval without holding their breath.
  • Statistical honesty: Know the difference between a signal and noise, size your eval sets accordingly, and be straight with people when the data can't answer the question they're asking.

Client Delivery & Stakeholder Management

  • Turning data into decisions: Present eval and observability findings to engineering and product audiences in a way that leads to a decision, and follow through until something changes.
  • Setting the standard: Help the client's teams adopt eval-first habits, and challenge the "ship it and see" instinct where you find it, with evidence rather than dogma.
  • Delivery ownership: Plan and prioritise your own stream, estimate accurately, manage shifting requirements without dropping quality, and flag risk to timelines early.
  • Cost awareness: Understand the cost of what you're measuring. Eval runs and judge models aren't free, so design for a sensible ratio of cost to confidence and be able to justify that to a budget holder.
  • Fitting in fast: Work within the client's existing tooling and release process where it's sound, and make the case for change where it isn't.

Documentation & Handover

  • Documentation as you go: Leave the dataset provenance, scoring rationale and runbooks that let the client's engineers maintain and extend the eval suite without you.
  • Knowledge transfer: Bring the client's engineers along in evaluation practice, prompt and context engineering, and critical review of model output, so the discipline outlasts the engagement.
  • Exit readiness: Treat a clean handover as part of the definition of done from the first sprint, not something arranged in the final fortnight.

Skills & Experience

Core Technical Capabilities

We're looking for a mix of AI engineering (50%), software engineering (30%), data engineering (20%).

  • Python: Strong and production-grade. You write code others maintain.
  • Eval and observability platforms: Hands-on production experience with Braintrust, LangSmith, Arize, Weights & Biases or an equivalent, including the parts that didn't work well.
  • Scoring design: Demonstrable experience designing LLM-as-judge and human-in-the-loop pipelines, and calibrating them against human judgement.
  • Production AI systems: Experience operating LLM-backed features in production, covering tracing, latency, token cost, failure modes and retrieval quality.
  • CI/CD: Enough fluency with GitHub Actions, GitLab CI or similar to own an eval gate in a promotion pipeline yourself.
  • Data handling: Comfortable building the data pipelines and stores behind datasets, traces and results at volume.
  • Communication: The ability to make a technical result land with a non-technical audience, tailoring the framing to the room.

Domain & Sector Experience

  • Regulated industries: Experience where model behaviour has compliance or consumer-protection consequences, such as iGaming, financial services or health, is a strong advantage. Safer gambling and responsible-messaging contexts are directly relevant here.
  • Contract and consulting delivery: A track record of arriving on an unfamiliar programme and being productive quickly, including building credibility with engineers who didn't ask for your help.

Useful credentials

Track record matters more than certification here. If you have an eval framework, a write-up or an open-source contribution you're proud of, lead with that. The following are useful supporting evidence:

  • AWS Certified Machine Learning Engineer (Associate) or AWS Certified AI Practitioner.
  • An associate-level AWS or GCP certification for the platform fundamentals around your work.

Contract details

These are contract engagements delivering a defined scope of work on a client programme. Duration, rate and IR35 status will be confirmed with our Talent Acquisition team at first contact. The work is largely remote, with travel to the client's site or to CreateFuture offices as the programme requires.

Our hiring process

  • A call with our Talent Acquisition team, covering rate, availability and notice
  • A capability and technical interview focused on the client and project work.

We move quickly for contract roles and will be straight with you about timelines.

Inclusion at CreateFuture

We want CreateFuture to be a place where people from every background can do their best work. If you need any adjustments to the process, or there's something that would help you show us what you can do, tell us and we'll sort it.

What we’ll offer you:

We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You’ll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do.

We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed.

We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.

Our hiring process

We try to keep our hiring process clear, fair and respectful of your time. We aim to get back to everyone who applies and we will be upfront about where you are in the process.

It usually looks like this:

  • Call with our Talent Acquisition Team 
  • Role specific capability interview 

Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation. We will explain what is involved before anything happens.

Inclusion at CreateFuture 

We believe diverse teams build better workplaces and better products. We want CreateFuture to be a place where people feel able to be themselves and do their best work.

If you need any adjustments or support during the application process, just. We will do what we can to help.

We look forward to your application!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,289,116 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Leeds
≈ $108k – $225k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • London
Python
AI/ML
Weights & Biases
PyTorch Lightning
Knowledge Distillation
Computer Vision
PyTorch
Ray
Few-Shot Learning
Model Distillation
Machine Learning
DevOps
AWS
Apply
≈ $106k – $213k per year (Estimated) • Hybrid • Full-Time • London • Lisbon
Python
TypeScript
Ruby
AI/ML
Fine-tuning
VLM
TensorFlow
PyTorch
LLM
Triton
DevOps
AWS
Kubernetes
Apply
Applied AI Engineer 11 months ago
≈ $96k – $254k per year (Estimated) • In office • Full-Time • London
AI/ML
AI Agents
LLM
Context Engineering
DevOps
GitHub
Apply
≈ $134k – $263k per year (Estimated) • Remote (United Kingdom) • Full-Time • 5+ years exp • Bachelor's Degree • United Kingdom • Switzerland • Finland
Python
C++
AI/ML
Model Context Protocol
Function Calling
AI Agents
LLM
RAG
Hallucination
Agentic Workflows
Tool Use
DevOps
Rest API
Apply
$299k – $338k per year • Hybrid • Bachelor's Degree • London
AI/ML
Claude
Model Context Protocol
Multimodal AI
LLM
Anthropic
Interpretability
Apply
SDET - Tech Lead 1 hour ago
≈ $30k – $80k per year (Estimated) • In office • 6+ years exp • Bengaluru
Python
JavaScript
TypeScript
DevOps
Rest API
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
QA
TestNG
Selenium
Cypress
Playwright
Postman
Insomnia
Pytest
Apply
Hybrid • Internship • Sofia
Python
JavaScript
PowerShell
Bash
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Linux
Unix
Management
Agile
Scrum
Apply
≈ $64k – $176k per year (Estimated) • Remote (United States) • Public Trust • 3+ years exp • Bachelor's Degree
Python
JavaScript
Java
TypeScript
Java
Spring Boot
Apache Tomcat
Databases
PostgreSQL
Oracle
MariaDB
ElasticSearch
OpenSearch
Amazon Aurora
Frontend
Vue.js
DevOps
CloudFormation
GitLab CI
CI/CD
Jenkins
AWS
Docker
Kubernetes
Bitbucket
AWS Lambda
Amazon EC2
GitHub
Amazon S3
AWS Step Functions
Cybersecurity
ISO 27001
Analytics
ETL/ELT
Management
Confluence
Jira
Agile
Scrum
Kanban
Apply
≈ $84k – $196k per year (Estimated) • Hybrid • Full-Time • Auckland
DevOps
GCP
Vercel
Azure
AWS
Cloudflare
Cybersecurity
Zscaler
PCI DSS
SIEM
Apply
QA Tester 1 hour ago
≈ $73k – $166k per year (Estimated) • Remote (likely United States) • 5+ years exp
Python
Java
Mojo
DevOps
Rest API
CI/CD
SLI/SLO/SLA
Cybersecurity
ISO 27001
Management
Agile
Apply
≈ $94k – $196k per year (Estimated) • In office • PhD • Edinburgh
Python
AI/ML
Model Context Protocol
AWS Bedrock
LLM
OpenAI
Anthropic
DevOps
Kong
AWS
Platform Engineering
API Gateway
Apply
≈ $83k – $167k per year (Estimated) • In office • Edinburgh
Java
Rust
C#
C#
.NET
DevOps
GCP
Azure
CI/CD
Platform Engineering
Management
Scrum
Kanban
Apply
≈ $80k – $157k per year (Estimated) • In office • Edinburgh
JavaScript
TypeScript
Node JS
AI/ML
AI Agents
Frontend
GraphQL
Next.js
React.js
Astro
DevOps
Rest API
Terraform
AWS
Amazon S3
Apply
≈ $68k – $136k per year (Estimated) • In office • Edinburgh
AI/ML
AWS Bedrock
AWS Bedrock AgentCore
DevOps
AWS
Incident Management
Apply
≈ $83k – $151k per year (Estimated) • In office • Edinburgh
AI/ML
AWS Bedrock
AWS Bedrock AgentCore
DevOps
AWS
Incident Management
Apply
≈ $36k – $63k per year (Estimated) • Hybrid • Leeds
Analytics
Power BI
Management
Power Apps
SharePoint
Apply
≈ $105k – $194k per year (Estimated) • In office • Full-Time • Leeds
Apply
≈ $114k – $225k per year (Estimated) • In office • Full-Time • Leeds
AI/ML
AI Agents
Apply
≈ $42k – $75k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Leeds
Management
Microsoft Office
Apply
Project Manager 1 day ago
≈ $39k – $87k per year (Estimated) • Equity • Hybrid • Full-Time • Leeds
Management
Smartsheet
Jira
Agile
Waterfall
Apply
See all jobs
This is one of many
1,289,116 more open roles from verified company boards, updated every day.