431,915open jobs
15,085companies
62,496added this week
Browse all
Salary
$24k – $50k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 4+ years exp
Overview
Company
Impact
Profile match
Humynlabs is an AI-first data intelligence platform enabling AI training and evaluation to deliver high-quality multimodal datasets with robust quality control for reliable, production-ready outcomes.

We are looking for a Senior Data Engineer to own, extend, and harden the multimodal data infrastructure that powers our AI training data supply chain. You will work directly with the Head of Data, building and maintaining production-grade pipelines that move real-world audio, video, and egocentric capture data from ingestion through AWS-scale storage, into GPU-backed technical validation, and finally to delivery-ready datasets for frontier AI lab customers. This role sits at the intersection of data engineering and ML infrastructure; you will own the full pipeline: ingestion, lake management, quality validation, and compute handoff.

The candidate will have responsibilities across the following functions:

Multimodal Ingestion and Preprocessing:

  • Build and maintain pipelines that ingest real-world audio, egocentric video, and image data at scale, covering format normalisation, chunking, metadata extraction, and landing into S3-based storage with consistent schema and partitioning.

AWS Data Lake:

  • Manage and extend the KGen data lake: Athena query optimisation, Glue crawlers and cataloguing, Apache Hudi table management, Lake Formation column-level permissions, and S3 lifecycle policies at TB-PB scale.

GPU Validation Pipeline Handoff:

  • Design and maintain the data layer that feeds GPU-backed technical validation workers, including sharding strategies, manifest generation, throughput optimisation, and I/O design so validation compute is never the bottleneck.
  • Understand how data format choices (Parquet, WebDataset, sharded archives) affect GPU-side loading performance.

Airflow DAG Management:

  • Author, debug, and monitor Airflow DAGs for scheduled processing, GPU job orchestration, and pipeline coordination across ingestion, validation, and delivery stages.

QC and Annotation Tooling:

  • Support the FastAPI-backed audio QC portal used by annotation workers; extend data validation and quality-check scripts across egocentric video and audio datasets.

Universal Data Schema (UDS):

  • Contribute to and enforce the Universal Data Schema for audio, image, and code modalities in the Humyn.
  • Labs dataset marketplace covering schema evolution, versioning, and partition strategies.

Infrastructure and Access Management:

  • Maintain AWS IAM, Lake Formation, and S3 bucket policies; manage data engineer access controls; handle cross-region data movement and vendor data sharing infrastructure.

ETL and Third-Party Integrations:

  • Build and maintain ingestion pipelines from APIs (Twitch, gaming analytics, Google Forms) into DynamoDB and PostgreSQL where required by business context.

Requirements:

  • 4+ years in a data engineering role with end-to-end pipeline ownership.
  • Strong Python async patterns, subprocess management, API clients, data processing at scale.
  • Hands-on AWS Athena, Glue, S3 DynamoDB, Lake Formation; production-grade, not just familiarity.
  • Apache Hudi or Delta Lake schema evolution, partition strategies, upsert patterns.
  • SQL proficiency able to write and optimise complex analytical queries.
  • Experience with Airflow or an equivalent workflow orchestrator.
  • Demonstrated experience with large-scale media data pipelines: audio/video format conversion, metadata extraction, chunking, egocentric or multimodal datasets (hard requirement, not a plus).
  • Understanding of how data format and I/O design affect downstream GPU compute workloads.
  • WebDataset, sharded Parquet, tfrecord, or equivalent.

Nice to Have:

  • Direct experience designing data handoff for GPU clusters (AWS Batch GPU instances, Ray, or SLURM).
  • Familiarity with ML training data formats and dataset standards used by AI labs (Hugging Face datasets, WebDataset, dataset cards).
  • Experience with rclone, large-scale file transfer, or cloud-to-cloud sync pipelines.
  • Exposure to data lineage or provenance tooling (OpenLineage, DataHub, or custom metadata schemas).
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
431,915 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
Data Architect 5 hours ago
$30k – $70k per year (Estimated) • In office • 12+ years exp • Bengaluru
Python
Java
SQL
Scala
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Spark
dbt
RAG
Knowledge Graph
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Kubernetes
Vector
Analytics
ETL/ELT
Apply
$20k – $44k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
SQL
DevOps
Rest API
QA
Postman
Apply
$10k – $23k per year (Estimated) • In office • 1+ year exp • Mumbai
Python
JavaScript
SQL
Databases
Snowflake
Management
WhatsApp
Marketing
Salesforce
Apply
SDET 2 5 hours ago
$12k – $34k per year (Estimated) • In office • 4+ years exp • Gurgaon
Python
SQL
Databases
Apache Kafka
AI/ML
MLFlow
TensorFlow
PyTorch
DevOps
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
GitHub
GitLab
QA
Selenium
Playwright
Rest-Assured
Pytest
Apply
Python Engineer 5 hours ago
$23k – $56k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Go
SQL
Python
Flask
AI/ML
LangGraph
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
AI Agents
Pandas
NumPy
LLM
RAG
Semantic Search
Semantic Search
Agentic Workflows
Multi-Agent Systems
DevOps
Rest API
GCP
CI/CD
Management
Google Sheets
Apply
In office • Bengaluru
Python
Rust
Bash
AI/ML
CUDA Toolkit
Multimodal AI
Computer Vision
PyTorch
MediaPipe
Ray
CUDA
cuDNN
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
AWS Lambda
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
AWS Step Functions
Robotics
Localization
Visual-Inertial Odometry
Apply
$23k – $50k per year (Estimated) • In office • Full-Time • PhD • Bengaluru
Apply
Remote/Hybrid • Full-Time • PhD • Bengaluru
AI/ML
Reinforcement Learning
Time Series Forecasting
Interpretability
Apply
$17k – $38k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru • Gurgaon
Management
Outlook
Apply
$91k – $151k per year • In office • Full-Time • 1+ year exp • Bachelor's Degree • Atlanta • Bengaluru • Hyderabad
Analytics
Tableau
Power BI
Microsoft Excel
Apply
In office • Bengaluru
Apply
See all jobs
This is one of many
431,915 more open roles from verified company boards, updated every day.