Salary
≈ $24k – $64k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Own the infrastructure and pipelines for integrating trained ASR/TTS/Speech-LLM models into production. Focus on scalable serving, GPU optimization, monitoring, and continuous improvement of inference latency and reliability.
What You’ll Do
- Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations.
- Develop and manage CI/CD pipelines for model promotion (staging → production).
- Configure autoscaling on Kubernetes (GPU pools) based on active calls or streaming sessions.
- Build health and observability dashboards: latency, token delay, WER drift, SNR/packet loss monitors.
- Integrate LM bias APIs, failover logic, and model switchers for fallback to larger/cloud models.
- Implement on-device or edge inference paths for low-latency scenarios.
- Collaborate with AI team to expose APIs for context biasing, rescoring, and diagnostics.
- Optimize GPU/CPU utilization, cost optimization, and memory footprint for concurrent ASR/TTS/Speech LLM workloads.
- Maintain data and model versioning pipelines with MLflow, DVC, or internal registries.
Desired Skills
- Experience with Triton, TensorRT, Docker, Kubernetes, and GPU scheduling.
- Familiarity with speech inference (streaming ASR, TTS pipelines).
- Proficient in Python, Bash, and cloud services (AWS/GCP/Azure).
- Understanding of observability stacks (Prometheus, Grafana, ELK).
- Knowledge of DevSecOps, access policies, and PHI-safe environments.
- Interest in inference optimization, mixed precision, and quantization.
Qualifications
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
573,000 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Bengaluru
Network Engineer – AI Network & Security
1 day ago
≈ $75k – $172k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Singapore
Python
AI/ML
NCCL
InfiniBand
DevOps
Ansible
CI/CD
HPC
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Least Privilege
Apply
Senior Technical Specialist
1 day ago
≈ $61k – $157k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Wellington • Auckland
Python
Java
TypeScript
Java
Spring Framework
Spring Boot
Databases
PostgreSQL
Apache Kafka
DevOps
Rest API
Splunk
OpenShift
Dynatrace
Kong
Azure
CI/CD
Jenkins
AWS
Kubernetes
Spinnaker
Platform Engineering
Amazon EKS
Azure AKS
Management
Agile
Apply
Integration Developer
1 day ago
≈ $56k – $163k per year (Estimated) • Remote/Hybrid • Full-Time • Auckland • Wellington
Java
Java
Spring Boot
DevOps
Rest API
CI/CD
Jenkins
AWS
Kubernetes
API Gateway
Management
Agile
Apply
≈ $52k – $130k per year (Estimated) • Remote • Full-Time
Python
Python
FastAPI
Django
Databases
Redis
RabbitMQ
Apache Kafka
DevOps
Terraform
GCP
Datadog
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Apply
Staff ML Engineer
1 day ago
≈ $75k – $166k per year (Estimated) • Remote • Full-Time
Python
AI/ML
Multimodal AI
Computer Vision
PyTorch
Edge AI
DevOps
Terraform
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)
10 months ago
≈ $34k – $80k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Bengaluru
Python
AI/ML
Dify
LoRA
vLLM
MLFlow
ElevenLabs
Fine-tuning
RLHF
Quantization
Multimodal AI
Function Calling
Speech Recognition
PEFT
QLoRA
Kaldi
Transformers
PyTorch
LLM
Whisper
Hallucination
Synthetic Data
Triton
Hugging Face
DPO
SFT
NVIDIA NeMo
Text-to-Speech
Deepgram
LiveKit
KV Cache
Tool Use
DevOps
WebRTC
Apply
≈ $37k – $93k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Bengaluru
Python
AI/ML
LoRA
MLFlow
Fine-tuning
Speech Recognition
PEFT
Kaldi
PyTorch
LLM
Whisper
Hugging Face
NVIDIA NeMo
Text-to-Speech
Deepgram
LiveKit
DevOps
WebRTC
Apply
≈ $19k – $47k per year (Estimated) • In office • Full-Time • 13+ years exp • Hyderabad • Bengaluru
DevOps
Terraform
CI/CD
AWS
Amazon S3
IAM
Cybersecurity
Least Privilege
Apply
Client Financial Mgmt Manager
2 days ago
In office • Full-Time • 13+ years exp • Master's Degree • Bengaluru
Apply
This is one of many
573,000 more open roles from verified company boards, updated every day.

