368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$30k – $72k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Principal · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
SambaNova Systems is an American artificial intelligence semiconductor and cloud platform company headquartered in Palo Alto, California. Founded in 2017 by Stanford professors Kunle Olukotun and Christopher Ré alongside former Oracle executive Rodrigo Liang, the company builds full-stack hardware and software infrastructure purpose-built for enterprise AI, large-scale model training, and high-speed agentic AI inference.

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

As a Cloud Platform Engineer you'll keep our AI inferencing platform reliable, fast, and scalable. Your focus is uptime, latency, and resource utilization on the inference endpoints customers depend on. The work spans monitoring, deployment, capacity planning, and incident response, and includes a shared on-call rotation covering 24/7 coverage.

Responsibilities  

Some of your responsibilities will include:

  • Owning the availability, latency, and efficiency of the production inferencing service, including change management, emergency response, and standing up AI infrastructure in new regions
  • Building and maintaining monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog covering service health, model latency and throughput, and RDU utilization
  • Leading incident response, running blameless post-mortems, and automating the toil those incidents expose
  • Finding and fixing performance bottlenecks, and designing auto-scaling policies that handle variable inference load without overspending
  • Managing cloud and on-prem infrastructure as code in Terraform and Ansible, and building the CI/CD pipelines that deploy new model versions and service updates

Required Qualifications

  • B.S. in Computer Science, Computer Engineering, or related field, or equivalent practical experience
  • 3+ years in a Site Reliability Engineering, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
  • Strong programming and scripting skills in Python, Go, or Java
  • Production experience with Docker and Kubernetes
  • Deep understanding of monitoring and observability tooling such as Prometheus, Grafana, Datadog, or the ELK Stack
  • Experience with infrastructure as code using Terraform or CloudFormation
  • Experience with CI/CD tooling such as Jenkins, GitHub Actions, or ArgoCD
  • Strong Linux system administration fundamentals

Preferred Qualifications

  • Experience in a hybrid environment spanning cloud and on-premise data center infrastructure
  • Experience supporting ML or AI inferencing services in production
  • Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs, for purposes of mapping to RDUs
  • Knowledge of model serving frameworks such as vLLM, SGLang, or Ray
  • Experience managing and tuning databases (SQL or NoSQL) and caching systems such as Redis or Memcached

Base Salary Range:

Base Pay Range

₹6,000,000—₹8,000,000 INR

Submission Guidelines

Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy

SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions

SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$20k – $49k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
Ruby
SQL
Databases
Amazon Neptune
Neo4j
AI/ML
Hallucination
LangChain
LangGraph
LLM
Model Context Protocol
Spark
AI Agents
LLM Guardrails
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
GitLab CI
Jenkins
Kubernetes
Rest API
Terraform
GitHub
GitLab
QA
Playwright
Postman
Selenium
Swagger
Apply
$41k – $103k per year (Estimated) • Remote • Full-Time • 5+ years exp
Go
Python
AI/ML
Reinforcement Learning
Edge AI
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Docker
GCP
GitHub Actions
Google GKE
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
GitHub
Apply
$20k – $51k per year (Estimated) • Remote • Full-Time • 3+ years exp
DevOps
AWS
Azure
GCP
IAM
Apply
$80k – $169k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$112k – $217k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$117k – $232k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • San Jose
Bash
C++
Python
SystemVerilog
Chips/EDA
UVM
Apply
$151k – $293k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Jose
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
CUDA
CUDA Toolkit
DeepSeek
DeepSpeed
JAX
Llama
LLM
Multimodal AI
OpenCL
PyTorch
Quantization
Qwen
TensorFlow
TensorRT
Triton
vLLM
cuDNN
Megatron-LM
Apply
$182k – $346k per year (Estimated) • In office • Full-Time • 12+ years exp • San Jose
Go
Python
Rust
AI/ML
LLM
DevOps
Helm
Kubernetes
Apply
$35k – $83k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
Go
Python
Rust
DevOps
Amazon EKS
ArgoCD
AWS
Azure
CI/CD
GCP
GitLab CI
Google GKE
Istio
Jenkins
Kubernetes
Linkerd
Service Mesh
Terraform
GitHub
GitLab
IAM
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Go
Python
Rust
AI/ML
LLM
TPU
DevOps
Helm
Kubernetes
SLI/SLO/SLA
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 2 hours ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru • Hyderabad • Chennai • Noida
Databases
Db2
IMS
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.