368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$25k – $63k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Skit.ai builds voice artificial intelligence agents for contact centre conversations. Its systems handle collections, verification and customer service calls autonomously. The company serves financial services clients in India and the United States.

About the Role

Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Senior Site Reliability Engineer - Voice AI Platform

Type: Full-time

Location: Bangalore

Why this role exists:

We run a voice AI platform that places and answers up to ~1 million calls per hour for regulated enterprises in banking, telecom, and collections. Unlike most SaaS, our workload is real-time and conversational: every call is a live media session where an extra few hundred milliseconds anywhere in the ASR → LLM → TTS loop is the difference between a natural exchange and a caller hanging up. Traffic is also bursty - outbound campaigns spin up huge concurrency inside narrow calling windows - and it runs across multiple clouds for resilience and data residency.

We are hiring a Senior SRE to own the reliability and performance of that system: the SLOs, the observability that makes problems visible, the capacity that absorbs campaign spikes, and the incident response that keeps regulated clients online. This is a systems-reliability role - latency, uptime, saturation, and the health of the telephony and serving path. (Model quality and evaluation live with a separate AI Observability role; you'll partner with them, not own their signals.)

If you want reliability problems that are genuinely hard - real-time media, sub-second budgets, six-figure concurrency, multi-cloud failover - this is that.

What you'll own:

  • SLOs and error budgets. Define and defend service-level objectives for availability and latency across the call path, and use error budgets to steer the balance between shipping and stability.
  • Observability. Own the metrics, tracing, and logging stack so failures surface fast and root cause is minutes not hours - distributed traces across signaling, ASR, LLM, TTS, and infra, with dashboards and alerting that page on real problems and stay quiet otherwise.
  • The real-time media path. Keep SIP signaling and RTP media healthy at scale - concurrency, jitter, packet loss, session setup - and the reliability of the components that carry them.
  • Capacity and autoscaling. Plan for peak (campaign windows that push toward the platform's concurrency ceiling), pre-warm capacity ahead of demand, and tune autoscaling so we neither drop calls nor burn money idling.
  • Incident response. Run a calm, structured on-call: triage, mitigation, clear comms to stakeholders on regulated accounts, and blameless postmortems that actually change the system.
  • Multi-cloud resilience. Design for failure across AWS, GCP, and Azure - redundancy, failover, disaster recovery, and the data-residency constraints that come with Indian banking and telecom clients.
  • Automation and toil reduction. Turn manual operations into infrastructure-as-code and self-healing systems. If you did it twice by hand, the third time is a script.

What the first your looks like:

  • First 90 days. Learn the call path end to end. Establish baseline SLIs for availability and latency, close the biggest gaps in alerting, and take a full turn in the on-call rotation.
  • By 6 months. Published SLOs with error budgets for the core services. A tracing/dashboards setup that makes the ASR→LLM→TTS latency budget visible per call. A repeatable pre-warm-and-scale playbook for campaign peaks.
  • By 12 months. Demonstrable reduction in incident frequency and time-to-mitigate. Tested multi-cloud failover for a critical path. On-call toil measurably down through automation.

What we're looking for

Must-have

  • Several years running high-availability, high-throughput production systems, including real on-call ownership.
  • Depth in at least one major cloud (AWS, GCP, or Azure) and with containers/Kubernetes.
  • Strong observability practice - metrics, distributed tracing, and logging (e.g. Prometheus/Grafana, OpenTelemetry, Tempo/Jaeger).
  • Fluency with SLIs/SLOs/error budgets and structured incident management.
  • Infrastructure-as-code (Terraform or similar) and CI/CD.
  • A programming language for automation and tooling (Python, Go, or similar) - beyond shell scripting.
  • Solid Linux systems and networking fundamentals; capacity planning and performance tuning.

Nice-to-have

  • Real-time media or VoIP experience - SIP/RTP, media servers, LiveKit, SBC/Kamailio.
  • Reliability of GPU/ML serving infrastructure.
  • Regulated-industry operations - uptime SLAs, DR, data residency.
  • Load testing and chaos engineering at scale.
  • PostgreSQL operations at scale.

Our stack:

Representative - you'll help shape it.Multi-cloud across AWS, GCP, and Azure; LiveKit/SIP for telephony; self-hosted and managed ASR (e.g. NVIDIA Parakeet / NeMo), LLMs, and TTS; Modal for ML deployment and pre-warming; PostgreSQL; Grafana/Tempo for metrics and traces; infrastructure-as-code and GitHub Actions CI/CD.

How you'll know you're succeeding:

Calls connect and stay fast even during the busiest campaign windows. Alerts mean something, and the ones that page you are worth waking up for. When something breaks, it's found and mitigated quickly and it doesn't break the same way twice. And the on-call rotation gets calmer over time, not busier, because the system increasingly heals itself.

We're an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
Data Engineer 1 day ago
In office • Full-Time • Singapore
Node JS
Python
SQL
JavaScript
Python
Beautiful Soup
Databases
Apache Kafka
MySQL
PostgreSQL
RabbitMQ
SQLite
AI/ML
Hadoop
Spark
DevOps
AWS
AWS Lambda
Azure
CI/CD
GCP
Amazon ECS
Amazon EventBridge
Amazon S3
Analytics
ETL/ELT
QA
Selenium
Apply
$170k – $318k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
In office • Full-Time • 3+ years exp • Indonesia
DevOps
AWS
Azure
GCP
IAM
Apply
$12k – $35k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Databases
MySQL
Oracle
PostgreSQL
AI/ML
AI Agents
DevOps
AIOps
Amazon EC2
AWS
CI/CD
CloudFormation
GitHub Actions
Kubernetes
Terraform
Amazon CloudWatch
Amazon S3
GitHub
IAM
Apply
$54k – $159k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Singapore
Bash
PowerShell
Python
DevOps
Ansible
AWS
Azure
Chef
Docker
Hyper-V
Incident Management
Kubernetes
Puppet
Red Hat
Terraform
VMWare
Windows Server
Apply
$18k – $81k per year (Estimated) • In office • Full-Time • Bengaluru
Bash
Python
Databases
PostgreSQL
AI/ML
LLM
LiveKit
DevOps
AWS
Azure
CI/CD
FinOps
GCP
GitHub Actions
Grafana
Kubernetes
Prometheus
Self-Healing
Terraform
GitHub
Apply
$30k – $68k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
AI/ML
LCM
LLM
Stable Diffusion
ISO 42001
LLM Guardrails
NIST AI RMF
OpenAI
DevOps
AWS
Azure
CI/CD
Docker
Kubernetes
Terraform
IAM
Cybersecurity
Checkov
CIS Benchmarks
Dockle
Grype
ISO 27001
Microsoft Defender
Microsoft Defender for Cloud
OWASP Top 10
OWASP ZAP
PCI DSS
SBOM
Semgrep
Snyk
SOC 2
SonarQube
Syft
tfsec
Trivy
Least Privilege
Cryptography
Vault
Apply
Quality Analyst 3 months ago
$15k – $34k per year (Estimated) • In office • Full-Time • 2+ years exp • Bengaluru
Apply
$26k – $68k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Go
Java
Kotlin
Python
SQL
Databases
Apache Kafka
DynamoDB
PostgreSQL
RabbitMQ
Redis
AI/ML
LLM
Edge AI
Speech Recognition
Text-to-Speech
DevOps
AWS
CI/CD
Datadog
GCP
Grafana
gRPC
Kubernetes
Prometheus
SLI/SLO/SLA
Terraform
WebRTC
WebSockets
Apply
$30k – $72k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Databases
PostgreSQL
AI/ML
LLM
DevOps
Ansible
AWS
Azure
Configuration Management
Docker
GCP
Grafana
Kubernetes
Loki
Prometheus
Terraform
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.