368,746open jobs
9,444companies
47,506added this week
Browse all
Location
In office (Dubai)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Dyson is a global technology and engineering company founded by Sir James Dyson, renowned for pioneering bagless cyclonic vacuum cleaners. Headquartered in Singapore, the company designs and manufactures high-performance consumer appliances, including cordless vacuums, bladeless air purifiers, hair care styling tools, and high-speed hand dryers. Driven by heavy investment in hardware research, digital motors, robotics, and battery technology, Dyson continuously reinvents everyday household technology through innovative engineering.

About us

At Dyson, we’re driven by a relentless pursuit of innovation-pushing boundaries in engineering, AI, and robotics. Our new Data Intelligence team sits at the heart of this mission: shaping Dyson’s future through data. Here, we blend creativity, precision, and audacity to power intelligent products. We craft data strategies and pipelines that fuel the next generation of connected devices.

You’ll work alongside brilliant minds from Dyson global engineering team and external software/hardware partners in an environment built for exploration, discovery, delivery and impact.

About the role

We are looking for a specialized Lead Data Intelligence Machine Learning Engineer to design and implement in-house tools that automate our data labelling pipelines. Your primary goal will be to reduce our reliance on manual annotation by leveraging techniques like Active Learning, Weak Supervision, and Synthetic Data Generation. You will bridge the gap between raw data collection and model-ready datasets, ensuring high-quality labels at scale.

Key Responsibilities

  • Architect Labelling Pipelines: Design and deploy end-to-end automated labelling systems using frameworks like Snorkel, Cleanlab, or custom active learning loops.

  • Develop "Human-in-the-Loop" (HITL) Systems: Build interfaces and workflows where models pre-label data and humans only intervene on high-uncertainty samples.

  • Quality Assurance & Denoising: Implement algorithmic checks to identify and correct mislabelled or "noisy" data within existing datasets.

  • Tooling & Integration: Collaborate with software engineers to integrate labelling tools with our existing data lakes and ML training infrastructure.

  • Model Optimization: Fine-tune "teacher" models to generate high-quality pseudo-labels for "student" models.

  • Set up and maintain robust data preparation infrastructure-optimising for data quality, speed, and seamless integration with downstream MLOps pipelines.

  • Perform data visualization and in-depth analysis using advanced data and feature engineering techniques. You’ll help transform raw data into actionable insight, supporting both research and deployment.

  • Work closely with Data Scientists, Software Engineers, and Product teams to ensure high data quality and usability across products and projects.

About you

  • At least 8+ years of professional experience in Machine Learning engineering, specifically focused on data centric-AI or computer vision/NLP pipelines.

  • Proficiency in Python: Mastery of the Machine Learning stack (PyTorch or TensorFlow, NumPy, Pandas, Scikit-learn).

  • Automated Labelling Expertise: Proven experience with Weak Supervision (labelling functions) or Active Learning strategies (uncertainty sampling, diversity sampling).

  • Data Engineering: Experience with SQL and NoSQL databases, and managing large-scale unstructured data (images, text, or audio).

  • Cloud Infrastructure: Familiarity with AWS (SageMaker Ground Truth), GCP (Vertex AI), or Azure ML labelling services.

  • Version Control for Data: Experience with DVC (Data Version Control) or similar tools to track dataset iterations.

  • Hands-on expertise building auto-labelling solutions or working with large-scale data annotation workflows.

  • Advanced skills in Python (and/or other relevant languages), and experience with key ML/data science libraries (e.g. TensorFlow, PyTorch, scikit-learn, pandas).

  • Experience designing, deploying, and maintaining scalable data pipelines, including data cleansing, transformation, and storage (cloud, on-prem, or hybrid).

  • Strong background in feature engineering, data analysis, and data visualization-comfortable using tools like Jupyter, Tableau, or Power BI.

  • Great communicator who documents solutions clearly and collaborates effortlessly across technical and non-technical teams.

  • Able to balance speed and quality, stay curious about new developments, and deliver results in a fast-moving environment.

  • Bachelor’s or Master's degree in computer science, Engineering, Mathematics, Data Science, or a related field.

Dyson is an equal opportunity employer. We know that great minds don’t think alike, and it takes all kinds of minds to make our technology so unique. We welcome applications from all backgrounds and employment decisions are made without regard to race, colour, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other any other dimension of diversity.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Dubai
$23k – $58k per year (Estimated) • In office • Full-Time • 10+ years exp • Pune
PowerShell
Python
SQL
Databases
MS SQL
Oracle
PostgreSQL
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Configuration Management
GCP
Grafana
Helm
Kubernetes
OpenShift
Prometheus
Apply
$17k – $49k per year (Estimated) • Remote/Hybrid • Kaliningrad
Node JS
Python
JavaScript
Node JS
BullMQ
Fastify
Python
FastAPI
Flask
Databases
Apache Kafka
Chroma
Pinecone
PostgreSQL
Qdrant
RabbitMQ
Redis
AI/ML
Claude
Copilot
Cursor
LangChain
LlamaIndex
LLM
Model Context Protocol
Prompt Engineering
RAG
Anthropic
OpenAI Codex
Structured Outputs
Function Calling
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Gitflow
Rest API
WebSockets
GitHub
Management
Jira
Apply
$32k – $72k per year (Estimated) • In office • Full-Time • 6+ years exp • Bengaluru
C++
Java
Python
YARA
Databases
Amazon Aurora
DevOps
Azure
GCP
Kubernetes
Cybersecurity
MITRE ATT&CK
Suricata
YARA
Zeek
Apply
$12k – $30k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Mumbai
Java
AI/ML
AI Agents
Edge AI
DevOps
AWS
Azure
Apply
$97k – $220k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
AI/ML
AI Agents
Qwen
OpenAI
DevOps
AWS
Azure
Apply
$78k – $202k per year (Estimated) • In office • Bachelor's Degree • Singapore
Python
DevOps
GCP
Apply
$37k – $87k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Mexico City
DevOps
SLI/SLO/SLA
Analytics
Power BI
Tableau
Apply
$69k – $158k per year (Estimated) • In office • Full-Time • 4+ years exp • Amsterdam
Analytics
A/B Testing
Apply
In office • Full-Time • Amsterdam
Marketing
GA4
Apply
$28k – $65k per year (Estimated) • In office • 12+ years exp • Bachelor's Degree • Gurgaon
Perl
Python
DevOps
AWS
GCP
Cybersecurity
GDPR
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Dubai
Apply
Remote/Hybrid • Full-Time • 15+ years exp • Master's Degree • Dubai
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dubai
Apply
In office • Dubai
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.