368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$30k – $76k per year (Estimated)
Location
Remote/Hybrid (Bengaluru, India)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
ABBYY Timeline is a leading process mining software with advanced process discovery & task mining capabilities for end-to-end business process visibility.

Join ABBYY and be part of a team that celebrates your unique work style. With flexible work options, a supportive team, and rewards that reflect your value, you can focus on what matters most - driving your growth, while fueling ours.

Our commitment to respect, transparency, and simplicity means you can trust us to always choose to do the right thing.

As a trusted partner for purpose-built AI and intelligent automation, we solve highly complex problems for our enterprise customers and put their information to work to transform the way they do business. Over 10,000 customers trust ABBYY, including many Fortune 500 ones. You will work on further developing a portfolio already containing client names such as DHL, Johnson & Johnson, FDA, DMV, PwC, KeyBank, Spotify, and H&R BLOCK.

About the Role  

We are seeking a Senior Machine Learning Engineer - Synthetic Data & Document Understanding to own the synthetic data generation track within ABBYY’s Document AI Data team. 

This role focuses on building generative pipelines that produce high-quality, diverse, and realistic synthetic training data at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties. 

This is an ideal role for engineers who combine deep generative modeling expertise with rigorous data quality evaluation and production engineering skills. 

Key Responsibilities  

Technical Development & Innovation  

  • Design and implement pipelines that analyze real documents to inform high-fidelity synthetic data generation 
  • Build generative systems capable of producing documents across diverse formats, layouts, and domains 
  • Develop evaluation frameworks to ensure synthetic data maintains distributional fidelity and diversity 
  • Research and apply generative modeling techniques suited for document AI training 
  • Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training 
  • Partner with Modeling teams to measure the impact of synthetic data on model performance 

Project Ownership & Leadership  

  • Own the synthetic data generation track end-to-end, from architecture to quality validation 
  • Drive architectural decisions balancing quality, diversity, scale, and cost efficiency 
  • Define and maintain data quality metrics and generation dashboards 
  • Collaborate closely with annotation teams to ensure compatibility with downstream pipelines 
  • Contribute to roadmap planning alongside Principal-level leadership 

Infrastructure & Scale  

  • Build scalable pipelines capable of generating millions of synthetic training examples 
  • Implement post-processing, filtering, and validation mechanisms to remove low-quality outputs 
  • Design cost-efficient workflows balancing compute, quality, and throughput 
  • Develop monitoring systems to detect distribution shifts or quality degradation over time 
  • Collaborate with Platform teams on compute orchestration, storage, and scheduling 

Qualifications  

Education & Experience  

  • MS or PhD in Computer Science, Engineering, Mathematics, or related field 
  • 5+ years of experience in Machine Learning / AI, with focus on:  
  • Generative models 
  • Vision-Language Models (VLMs) 
  • Synthetic data systems 
  • Proven experience building and evaluating synthetic data pipelines for ML training 
  • Strong background in data quality evaluation and statistical analysis 

Technical Expertise  

  • Deep expertise in Vision-Language Models and document understanding (layout, structure, semantics) 
  • Strong knowledge of generative modeling for structured and semi-structured data 
  • Understanding of what makes synthetic data valuable:  
  • Distributional fidelity 
  • Diversity 
  • Realistic noise patterns 
  • Domain coverage 
  • Strong programming skills in Python with experience in PyTorch or similar frameworks 
  • Experience evaluating data quality via automated metrics and downstream model impact 
  • Familiarity with large-scale data pipelines, cloud environments, and experiment tracking 

Leadership & Communication  

  • Proven ability to independently own complex technical workstreams 
  • Strong collaboration across data, modeling, and platform teams 
  • Ability to clearly communicate data quality and generation trade-offs 
  • Data-driven mindset with strong attention to coverage gaps and quality signals 

Here are some of our local benefits:  

  • Comprehensive medical, accidental, and life insurance 
  • Weekly wellness sessions to support your physical and mental well-being 
  • A generous paid time off policy 

Join ABBYY, and you will:

Love how you work

  • We provide remote and hybrid working options to fit all lifestyles.
  • We use flexible hours across most of our teams to allow you to find your own definition of balance.
  • Encouraging a culture of giving, we provide two paid volunteering days off every year so you can take time to contribute to the causes you care about.
  • To ensure your family is cared for, we offer paid parental leave in all our locations.

Love whom you work with

  • We are a global team of 600+ colleagues, spread across 15 countries on four continents.
  • With colleagues representing 30+ nationalities, our workforce reflects the world.
  • Innovation and excellence run through our veins. Our teams gather the expertise which has garnered ABBYY more than 140 technology patents.
  • We are guided by the values of respect, transparency, and simplicity.
  • "Team Environment" is in the top three highest-scoring drivers of engagement across all of our departments.

Love what you work on

  • We are a company with more than 35 years of experience in the technology market;
  • Over 10,000 customers trust ABBYY, including many Fortune 500 ones, with names such as DHL, Johnson & Johnson, FDA, DMV, PwC, KeyBank, Spotify, and H&R BLOCK;
  • We have modernized the capture market by creating the first low-code/no-code IDP platform.
  • Our Machine Learning, Natural Language Processing, Computer Vision Technologies, and a marketplace built with AI, can transform any document in any process;
  • Top Analyst firms recognize ABBYY's market leadership, including Gartner, Everest PEAK Matrix ® Assessment, ISG Intelligent Automation Lens, and NelsonHall, amongst others.

ABBYY is an Equal Employment Opportunity employer that values the strength that diversity brings to the workplace. To learn more about our commitment to Diversity and Inclusion, check out the careers section on our website.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$67k – $160k per year (Estimated) • In office • Full-Time • France
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
Multimodal AI
OpenCV
PyTorch
TensorFlow
Edge AI
Apply
$120k – $250k per year • Equity 0.2–1% • In office • Full-Time • Master's Degree • Seattle
Python
AI/ML
Diffusion Models
Fine-tuning
PyTorch
Self-Supervised Learning
Apply
$20k – $56k per year (Estimated) • Equity 0–0.1% • In office • Full-Time • 3+ years exp • Master's Degree • Bengaluru
MATLAB
Python
Apply
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
$23k – $53k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Moscow
Python
DevOps
Git
QA
Pytest
Apply
Data Engineer 5 days ago
$63k – $82k per year • Remote/Hybrid • 5+ years exp • Bachelor's Degree • Budapest
Python
SQL
AI/ML
Computer Vision
Analytics
ETL/ELT
Apply
Remote/Hybrid • Bengaluru
JavaScript
Python
AI/ML
Computer Vision
OCR
DevOps
Docker
Rest API
QA
Postman
Apply
$79k – $82k per year • Remote/Hybrid • 10+ years exp • Budapest
C#
C++
SQL
C#
.NET
Databases
Azure SQL Database
AI/ML
Computer Vision
LLM
NLP
OCR
OpenAI
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
GitHub Actions
Kubernetes
GitHub
Apply
Remote/Hybrid • 10+ years exp • Belgrade
C#
C++
SQL
C#
.NET
Databases
Azure SQL Database
AI/ML
Computer Vision
LLM
NLP
OCR
OpenAI
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
GitHub Actions
Kubernetes
GitHub
Apply
Backend Developer 4 months ago
$21k – $73k per year (Estimated) • Remote/Hybrid • 2+ years exp • Budapest
C#
Node JS
SQL
TypeScript
JavaScript
C#
ASP.NET Core
AI/ML
Computer Vision
DevOps
AWS
Azure
CI/CD
Cloudflare
Docker
GCP
Git
Kubernetes
Rest API
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.