427,850open jobs
14,428companies
63,552added this week
Browse all
Salary
$28k – $63k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Staff · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Skit.ai builds voice artificial intelligence agents for contact centre conversations. Its systems handle collections, verification and customer service calls autonomously. The company serves financial services clients in India and the United States.

Job Title: Lead DevOps Engineer

Location: Bengaluru (100% WFO)

Job Type: Full-time

The problem

Skit.ai runs autonomous voice agents for regulated enterprises - India's largest banks and telcos, and US collections operations. Every call is a live distributed system: PSTN/SIP → media server → ASR → LLM → TTS → back, spread across three clouds and multiple vendors, with a conversational response budget measured in hundreds of milliseconds.

The platform peaks at roughly **1 million calls per hour**. Billed minutes grew **5,000x+ in eight months**. At this scale, infrastructure is not a support function - latency, cost-per-minute, and auditability are product features. When infra degrades, a customer mid-sentence hears silence.

We're hiring a Lead DevOps Engineer to own this substrate and keep it ahead of the growth curve.

What you'll own:

  • Multi-cloud substrate: Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget - p95 per hop in tens of milliseconds, not "best effort."
  • Real-time media plane: Self-hosted LiveKit and SIP infrastructure at scale. Media servers are stateful; you'll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
  • Model-serving infrastructure: GPU fleets (A100/H100/B200-class) for self-hosted ASR and open-weight LLMs - inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets - alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
  • Reliability & observability: OTel-native tracing (Grafana/Tempo stack), per-turn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You'll act as incident commander for infrastructure and write the runbooks you'd want at 3 a.m.
  • Cost engineering: Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture - you'll own the next 10x.
  • Security & compliance: Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001; working command of ISO/IEC 42001:2023 (AI management systems) - hands-on preferred, rigorous theoretical grounding acceptable. You'll face bank and telecom auditors directly, including data-residency requirements.
  • Technical leadership: Terraform-first IaC standards, production-readiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team.

Problems on our plate right now

  • Scaling stateful, self-hosted media servers past current concurrency ceilings - HPA on CPU doesn't cut it
  • Migrating LLM inference from managed APIs to self-hosted open-weight models on GPUs without breaking TTFT budgets
  • ASR, LLM, and telephony living in different clouds: interconnect topology that keeps the packet path short and private
  • Multi-region DR that satisfies bank audits without doubling spend

If these read as interesting rather than terrifying, keep reading.

Must-have

  • 6+ years hands-on cloud infrastructure; 3+ years operating multiple clouds simultaneously in production; deep expertise in at least two of AWS/GCP/Azure
  • Real-time audio/video systems in production - WebRTC, SIP/PSTN, or streaming media; you've debugged jitter, not just read about it
  • Networking depth: VPC/VNet design, load balancing, DNS, NAT; private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect); transit gateways and cross-cloud mesh
  • Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling *stateful* workloads; service mesh familiarity (Istio/Linkerd)
  • Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates; drift is a bug
  • AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
  • Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR; ElevenLabs / Azure-class TTS)
  • Security fundamentals: IAM/RBAC, secrets management (Vault or cloud-native), encryption and key rotation
  • CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline

Strong signal (nice-to-have)

  • LiveKit, pipecat, Twilio, or comparable real-time platforms; SIP trunking and PSTN integration
  • KEDA or other event-driven autoscaling used in anger
  • MLOps: model versioning, canary and A/B rollout
  • FinOps discipline: reserved/spot strategy, unit-economics reporting
  • Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert
  • ISO/IEC 42001:2023 exposure

What we're NOT looking for

  • Single-cloud depth with documentation-level knowledge of the other two
  • Tool-checklist DevOps without production AI/ML serving scars
  • "Can learn quickly" as the primary qualification - this role needs day-one production credibility
  • Anyone who has never traced a packet across a cloud boundary
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
427,850 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
In office • 10+ years exp • Bachelor's Degree
Python
Go
JavaScript
TypeScript
C++
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Google ADK
Hallucination
OpenAI
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Vector
Apply
AI, Sr Architect 12 hours ago
In office • 15+ years exp • Bachelor's Degree
Python
Go
JavaScript
C++
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Google ADK
Hallucination
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Vector
Apply
In office • 5+ years exp • Bachelor's Degree
Python
JavaScript
SQL
C++
Node JS
Python
Flask
Django
AI/ML
Pandas
NumPy
DevOps
Rest API
CI/CD
GitHub
Apply
Remote/Hybrid • Full-Time • London
Python
JavaScript
Java
TypeScript
C#
C#
.NET
Databases
Apache Kafka
Kafka
AI/ML
Copilot
Cursor
AI Agents
Devin
Frontend
Vue.js
Angular
React.js
DevOps
Rest API
WebSockets
CI/CD
Kubernetes
GitHub
Cybersecurity
Volatility
Apply
SOC Analyst - L1 1 hour ago
$53k – $134k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Kuwait City
DevOps
GCP
Azure
AWS
SLI/SLO/SLA
Cybersecurity
MITRE ATT&CK
Apply
Collections Manager 26 days ago
$18k – $41k per year (Estimated) • In office • Full-Time • 5+ years exp • Mumbai
AI/ML
LLM
Analytics
A/B Testing
Management
WhatsApp
Apply
$23k – $56k per year (Estimated) • In office • Full-Time • Bengaluru
Python
Databases
PostgreSQL
AI/ML
Speech Recognition
LLM
NVIDIA NeMo
Text-to-Speech
LiveKit
DevOps
Terraform
GCP
GitHub Actions
Jaeger
OpenTelemetry
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Chaos Engineering
Self-Healing
Incident Management
GitHub
Apply
In office • Full-Time • Bengaluru
Python
Bash
Databases
PostgreSQL
AI/ML
LLM
LiveKit
DevOps
Terraform
GCP
GitHub Actions
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Self-Healing
FinOps
GitHub
Apply
$26k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
AI/ML
Stable Diffusion
LLM
LCM
OpenAI
LLM Guardrails
NIST AI RMF
ISO 42001
DevOps
Terraform
Azure
CI/CD
AWS
Docker
Kubernetes
IAM
Cybersecurity
Snyk
OWASP ZAP
SonarQube
Trivy
Semgrep
ISO 27001
Checkov
Grype
Microsoft Defender
CIS Benchmarks
OWASP Top 10
PCI DSS
SOC 2
Least Privilege
SBOM
Syft
Microsoft Defender for Cloud
tfsec
Dockle
Cryptography
Vault
Apply
$27k – $64k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
Go
Java
Kotlin
SQL
Databases
PostgreSQL
Redis
DynamoDB
RabbitMQ
Apache Kafka
Kafka
AI/ML
Speech Recognition
LLM
Text-to-Speech
Edge AI
DevOps
gRPC
Terraform
GCP
WebRTC
Datadog
Prometheus
WebSockets
CI/CD
AWS
Kubernetes
Grafana
SLI/SLO/SLA
Apply
Data Engineer 1 hour ago
$24k – $50k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • Chennai • Bengaluru
Python
Python
pySpark
AI/ML
Spark
Analytics
ETL/ELT
Apply
In office • Full-Time • Bachelor's Degree • Bengaluru
Apply
In office • Full-Time • Bachelor's Degree • Bengaluru
Apply
See all jobs
This is one of many
427,850 more open roles from verified company boards, updated every day.