539,450open jobs
19,395companies
74,343added this week
Browse all
Location
In office (Bengaluru)
Overview
Company
Impact
Profile match
Humynlabs is an AI-first data intelligence platform enabling AI training and evaluation to deliver high-quality multimodal datasets with robust quality control for reliable, production-ready outcomes.

Responsibilities:

  • Productionize research prototypes. Take model code from researchers (PyTorch, CUDA, exotic dependency stacks) and deliver production pipeline stages: reproducible Docker images, pinned GPU/CUDA/cuDNN environments, clean I/O contracts, retries, idempotency, and observability. Real examples of this work from our pipeline:
  • A researcher's stereo depth-estimation model: a GPU Batch stage with a locked CUDA base image, standardised S3 input/output layout, and per-clip cost tracking.
  • A fused MediaPipe + WiLoR hand-pose prototype: a single orchestrated labelling stage with well-defined intermediate artefacts and failure-isolation per clip.
  • A monocular depth model + a stereo model: a three-stage fuse pipeline where resolution-alignment invariants are enforced by the platform, not by tribal knowledge.
  • Own the orchestration layer. Design and evolve Step Functions state machines, AWS Batch compute environments and job queues, Lambda glue, and selective stage re-execution (re-run just one stage across a fleet of clips without redoing everything).
  • Own the infrastructure as code. All of it lives in Terraform modules, per-environment stacks, ECR, IAM, networking. You'll extend and harden this, not click around a console.
  • Drive cost efficiency as a first-class feature. Spot capacity strategies, right-sizing GPU instance families, eliminating GPU idle time (we've measured it, we hunt it), storage lifecycle policies on multi-TB S3 datasets, batching strategies that keep expensive GPUs saturated.
  • You should be the person who can say what a pipeline run costs per clip and then make that number go down.
  • Make it reliable at scale. Structured logging, metrics, and alerting across stages; dead-letter handling and automatic retries for flaky clips; data-quality gates so bad inputs fail fast and loudly instead of silently poisoning downstream datasets.
  • Manage the container fleet. A dozen-plus GPU images with heavy, conflicting ML dependencies.
  • Keep builds fast, images slim, CUDA stacks consistent, and breakage (e. g., an upstream wheel disappearing from an index) fixed within hours, not weeks.
  • Move fast with researchers. Sit close to the research loop: prototype, deploy to staging, run on real fleet data, iterate on feedback, promote to production.

Requirements:

  • 5+ years in infrastructure/platform with real ownership of production systems.
  • Strong Python. Not just scripting: you write clean, tested, maintainable pipeline and tooling code that other engineers build on.
  • Deep AWS experience: compute, networking, IAM, storage and strong opinions about cost.
  • A track record of taking rough prototypes (ideally ML/research code) to production.
  • Hands-on Docker/containerization depth: you debug CUDA base-image conflicts and dependency hell without flinching; ECS/EKS or other orchestration experience.
  • Terraform (or equivalent IaC) used seriously, in a team, across environments.
  • Working GPU knowledge: what saturates a GPU, what leaves it idle, how instance choice and batching change the bill.
  • Cost-optimisation instinct: spot strategies, right-sizing, storage tiering, and the discipline to measure before and after.
  • Systems thinking and bias to ship: you'd rather run it on real data today and iterate than perfect it in isolation.

Nice to Have:

  • Experience with ML labelling/inference pipelines, video or multimodal sensor data.
  • Exposure to computer-vision workloads (depth estimation, pose estimation, SLAM/VIO).
  • EKS/Kubernetes at scale; Ray or other distributed-compute frameworks.
  • CI/CD for container-heavy repos (CodeBuild, GitHub Actions).

Our Stack:

  • Languages: Python (primary; you must be genuinely strong here), Bash; Go/Rust a plus.
  • Orchestration: AWS Step Functions, AWS Batch, Lambda.
  • Containers: Docker, ECR; multi-stage GPU image builds; container orchestration concepts (ECS/EKS experience welcome); IaC: Terraform (modules, multi-environment).
  • GPU/ML runtime: NVIDIA CUDA/cuDNN, PyTorch deployment environments, GPU instance families on AWS, spot vs. on-demand economics.
  • Data: S3 at multi-TB scale, structured artefact layouts, dataset versioning, high-throughput transfer.
  • Observability: CloudWatch logs/metrics/alarms, cost attribution and reporting.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
539,450 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
DevOps Engineer 1 hour ago
$16k – $45k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Bash
DevOps
Terraform
GCP
Helm
Azure DevOps
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
SRE
Platform Engineering
IAM
Cybersecurity
ISO 27001
SOC 2
Apply
$77k – $204k per year (Estimated) • In office • Full-Time • Sydney • Melbourne
JavaScript
TypeScript
Node JS
AI/ML
AI Agents
Frontend
GraphQL
React.js
Storybook
DevOps
Rest API
GitHub Actions
Azure
CI/CD
Git
AWS
Docker
Kubernetes
TeamCity
Octopus Deploy
GitHub
Cybersecurity
SonarQube
Management
Agile
QA
Playwright
Jest
Apply
Python Engineer - AI 2 hours ago
In office • 3+ years exp • Sydney
Python
AI/ML
AI Agents
Agentic Workflows
DevOps
CI/CD
Apply
$26k – $62k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
DevOps
CI/CD
Apply
In office • Internship • Singapore
Python
SQL
AI/ML
Anomaly Detection
Analytics
Tableau
Power BI
Apply
$24k – $50k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
SQL
Python
FastAPI
Databases
PostgreSQL
Delta Lake
DynamoDB
Apache Hudi
AI/ML
Airflow
Multimodal AI
Ray
Hugging Face
DevOps
SLURM
AWS
Amazon S3
IAM
Analytics
ETL/ELT
AWS Glue
Apply
In office • Bengaluru
Python
Rust
Bash
AI/ML
CUDA Toolkit
Multimodal AI
Computer Vision
PyTorch
MediaPipe
Ray
CUDA
cuDNN
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
AWS Lambda
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
AWS Step Functions
Robotics
Localization
Visual-Inertial Odometry
Apply
$19k – $47k per year (Estimated) • In office • Full-Time • 13+ years exp • Hyderabad • Bengaluru
DevOps
Terraform
CI/CD
AWS
Amazon S3
IAM
Cybersecurity
Least Privilege
Apply
In office • Full-Time • 13+ years exp • Master's Degree • Bengaluru
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
JavaScript
Frontend
React.js
Apply
Security Architect 1 day ago
$32k – $76k per year (Estimated) • In office • Full-Time • 15+ years exp • Bengaluru
AI/ML
Model Context Protocol
AI Agents
DevOps
GCP
Azure
AWS
Kubernetes
Platform Engineering
IAM
Cybersecurity
ISO 27001
Open Policy Agent
PCI DSS
SOC 2
GDPR
Zero Trust
Least Privilege
Threat Modeling
Apply
Prompt Engineer 1 day ago
$24k – $65k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
AI/ML
Prompt Engineering
Apply
See all jobs
This is one of many
539,450 more open roles from verified company boards, updated every day.