573,000open jobs
23,970companies
78,613added this week
Browse all
Salary
$24k – $64k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Own the infrastructure and pipelines for integrating trained ASR/TTS/Speech-LLM models into production. Focus on scalable serving, GPU optimization, monitoring, and continuous improvement of inference latency and reliability.

What You’ll Do

  • Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations.
  • Develop and manage CI/CD pipelines for model promotion (staging → production).
  • Configure autoscaling on Kubernetes (GPU pools) based on active calls or streaming sessions.
  • Build health and observability dashboards: latency, token delay, WER drift, SNR/packet loss monitors.
  • Integrate LM bias APIs, failover logic, and model switchers for fallback to larger/cloud models.
  • Implement on-device or edge inference paths for low-latency scenarios.
  • Collaborate with AI team to expose APIs for context biasing, rescoring, and diagnostics.
  • Optimize GPU/CPU utilization, cost optimization, and memory footprint for concurrent ASR/TTS/Speech LLM workloads.
  • Maintain data and model versioning pipelines with MLflow, DVC, or internal registries.

Desired Skills

  • Experience with Triton, TensorRT, Docker, Kubernetes, and GPU scheduling.
  • Familiarity with speech inference (streaming ASR, TTS pipelines).
  • Proficient in Python, Bash, and cloud services (AWS/GCP/Azure).
  • Understanding of observability stacks (Prometheus, Grafana, ELK).
  • Knowledge of DevSecOps, access policies, and PHI-safe environments.
  • Interest in inference optimization, mixed precision, and quantization.

Qualifications

  • B.Tech / M.Tech in Computer Science or related field.
  • 4-6 years of experience in backend or ML-ops; at least 1-2 years with GPU inference pipelines.
  • Proven experience deploying models to production environments with measurable latency gains.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
573,000 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$75k – $172k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
AI/ML
NCCL
InfiniBand
DevOps
Ansible
CI/CD
HPC
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Least Privilege
Apply
$61k – $157k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Wellington • Auckland
Python
Java
TypeScript
Java
Spring Framework
Spring Boot
Databases
PostgreSQL
Apache Kafka
DevOps
Rest API
Splunk
OpenShift
Dynatrace
Kong
Azure
CI/CD
Jenkins
AWS
Kubernetes
Spinnaker
Platform Engineering
Amazon EKS
Azure AKS
Management
Agile
Apply
$56k – $163k per year (Estimated) • Remote/Hybrid • Full-Time • Auckland • Wellington
Java
Java
Spring Boot
DevOps
Rest API
CI/CD
Jenkins
AWS
Kubernetes
API Gateway
Management
Agile
Apply
$52k – $130k per year (Estimated) • Remote • Full-Time
Python
Python
FastAPI
Django
Databases
Redis
RabbitMQ
Apache Kafka
DevOps
Terraform
GCP
Datadog
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Apply
Staff ML Engineer 1 day ago
$75k – $166k per year (Estimated) • Remote • Full-Time
Python
AI/ML
Multimodal AI
Computer Vision
PyTorch
Edge AI
DevOps
Terraform
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
$34k – $80k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Bengaluru
Python
AI/ML
Dify
LoRA
vLLM
MLFlow
ElevenLabs
Fine-tuning
RLHF
Quantization
Multimodal AI
Function Calling
Speech Recognition
PEFT
QLoRA
Kaldi
Transformers
PyTorch
LLM
Whisper
Hallucination
Synthetic Data
Triton
Hugging Face
DPO
SFT
NVIDIA NeMo
Text-to-Speech
Deepgram
LiveKit
KV Cache
Tool Use
DevOps
WebRTC
Apply
$37k – $93k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Bengaluru
Python
AI/ML
LoRA
MLFlow
Fine-tuning
Speech Recognition
PEFT
Kaldi
PyTorch
LLM
Whisper
Hugging Face
NVIDIA NeMo
Text-to-Speech
Deepgram
LiveKit
DevOps
WebRTC
Apply
$19k – $47k per year (Estimated) • In office • Full-Time • 13+ years exp • Hyderabad • Bengaluru
DevOps
Terraform
CI/CD
AWS
Amazon S3
IAM
Cybersecurity
Least Privilege
Apply
In office • Full-Time • 13+ years exp • Master's Degree • Bengaluru
Apply
See all jobs
This is one of many
573,000 more open roles from verified company boards, updated every day.