368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$18k – $82k per year (Estimated)
Location
Remote/Hybrid (Bengaluru, India)
Employment
Full-Time
Overview
Company
Impact
Profile match
Calix is a broadband platform and services company headquartered in San Jose, California, and founded in 1999. The company supplies access hardware, cloud software, and managed services that let broadband providers deliver and monetize fiber and wireless subscriber networks. It works mainly with regional carriers, cooperatives, and municipal utilities in North America and is listed on the New York Stock Exchange.
The Calix platform enables Communication Service Providers (CSPs) of all sizes to transform and future-proof their businesses. Through real-time data, automation, and actionable insights delivered via Calix One - our cloud-first, AI-powered platform - CSPs can simplify operations, collapse cost, and accelerate innovation. Calix One brings together the automation of everything and the experience of one, empowering customers to deliver differentiated subscriber experiences while driving acquisition, loyalty, and revenue growth. This is the Calix mission: to enable CSPs of all sizes to simplify, innovate, and grow, strengthening both their businesses and the communities they serve.

We’re at the forefront of a once in a generational change in the broadband industry. Join us as we innovate, help our customers reach their potential, and connect underserved communities with unrivaled digital experiences.

The Site Reliability Engineer ensures our production services remain highly available, scalable, and efficient on Google Cloud Platform. You will bridge the gap between development and operations by diving deep into application source code and cloud infrastructure to permanently engineer away underlying issues, rather than just patching symptoms. This role focuses on owning complex alert triage, reading and debugging code to solve root causes, GitOps application deployments via ArgoCD, and leveraging Grafana observability alongside advanced AIOps platforms to drive down operational toil.

Key Responsibilities:

  • Code-Level Alert Resolution: Act as the ultimate owner of complex alerts by investigating stack traces, reading application source code, and submitting code-level fixes alongside infrastructure adjustments to permanently resolve chronic issues.
  • Infrastructure Optimization: Diagnose and resolve deep OS and distributed system bottlenecks-including CPU throttling, memory leaks, and storage constraints-across GKE worker nodes and critical data infrastructure (e.g., Kafka, Datastream).
  • GitOps & Deployments: Deploy, roll back, and manage the lifecycle of containerized applications using ArgoCD pipeline workflows, ensuring safe and reliable release rollouts.
  • AIOps & Automation: Utilize AI-driven operations tools and build custom Python/Go automation (e.g., PagerDuty API integrations) to interpret correlated events, reduce alert noise, and automate manager/triage workflows.
  • Network Troubleshooting: Diagnose complex connectivity and latency issues across all network layers, isolating problems between cloud VPCs, Kubernetes overlays, and microservices.
  • Incident Response: Participate in on-call rotations, using Grafana dashboards, AIOps suggestions, and code-level tracing to rapidly mitigate and permanently fix production container issues.

What You'll Actually Do (Example Scenario):

  • The Alert: PagerDuty pages you for elevated consumer lag on a critical Kafka topic and latency spikes in a downstream microservice.
  • The Investigation: You use Grafana to correlate the latency with CPU throttling on specific GKE worker nodes. You drop into the command line, run top and tcpdump, and notice the application pods are churning through memory and network connections.
  • The Code Dive: You pull the Python application source code and discover a recently merged commit introduced an inefficient retry loop when failing to parse certain Kafka payloads, causing a memory leak.
  • The Fix: You temporarily scale up the GCP machine type or rollback via ArgoCD to restore service. Then, you write a Python PR to fix the exception handling, adjust the Kubernetes resource limits in the Git repo, and update your AIOps definitions to catch this specific anomaly automatically in the future.

Required Technical Skills:

  • Software Engineering: Strong proficiency in reading, debugging, and modifying software in Python or Go to fix production bugs, build APIs, and create automation scripts.
  • Cloud & Orchestration: Deep functional knowledge of deploying, scaling, and managing workloads in GKE, with a strong grasp of GCP infrastructure and machine type optimization.
  • Operating Systems: Strong foundational knowledge of Linux internals, process management, file systems, kernel parameters, and how application performance triggers OS-level constraints (e.g., OOM killers).
  • Networking: Practical troubleshooting skills in Kubernetes Networking (Pod-to-Pod communication, CNI plugins, Ingress) as well as L3-L7 mechanics (TCP/UDP, gRPC, DNS, IP routing).
  • Observability & Alerting: Deep experience with the Grafana Labs ecosystem (Grafana, Mimir/Prometheus, Loki, Tempo) and advanced incident management platforms like PagerDuty.
  • CI/CD & GitOps: Practical experience managing applications using ArgoCD and Git version control systems.
  • System & Network Utilities: Proficiency with command-line diagnostic tools (e.g., tcpdump, curl, dig, traceroute, top, iostat).
  • ML/AIOps (Nice to have): Familiarity with anomaly detection concepts, log-based ML models, and automated noise reduction.

Soft Skills & Qualifications:

  • Engineering Mindset: A relentless drive to engineer away toil and fix root causes at the architectural or code layer rather than relying on manual runbooks.
  • Autonomous Problem Solving: Ability to systematically troubleshoot complex, highly distributed microservice issues under pressure without needing a playbook.
  • Urgency & Prioritization: Strong sense of ownership and ability to prioritize alerts based on business and customer impact.
  • Communication: Clear written and verbal communication during high-stress incident responses and blameless Root Cause Analyses (RCAs).

Location:

  • India - (Flexible hybrid work model - work from Bangalore office for 20 days in a quarter)
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$170k – $318k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
Quality Engineer 9 hours ago
$10k – $42k per year (Estimated) • In office • Full-Time • 2+ years exp • Bengaluru
DevOps
CI/CD
Apply
In office • Full-Time • 3+ years exp • Indonesia
DevOps
AWS
Azure
GCP
IAM
Apply
$47k – $165k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Singapore
JavaScript
Node JS
Databases
PostgreSQL
AI/ML
AI Agents
Frontend
D3.js
Next.js
React.js
DevOps
CI/CD
Git
Apply
$27k – $58k per year (Estimated) • In office • 3+ years exp • Saratov
C++
Java
C++
Qt
Java
Gradle
Maven
DevOps
CI/CD
Apply
$32k – $69k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bengaluru
Databases
Apache Kafka
DevOps
CI/CD
Cybersecurity
Tcpdump
Wireshark
QA
Cypress
Playwright
Selenium
Apply
$44k – $95k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Java
Python
Rust
SQL
AI/ML
AI Agents
AutoGen
CrewAI
Embeddings
Google ADK
LangChain
LLM
NLP
Prompt Engineering
RAG
Reinforcement Learning
DevOps
Rest API
Apply
$44k – $96k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bengaluru
Python
AI/ML
Function Calling
Gemini
LiteLLM
LLM
RAG
Vertex AI
A2A
LLM Guardrails
AI Agents
Model Context Protocol
DevOps
CI/CD
IAM
GitHub
Design
Figma
Management
Confluence
Jira
Apply
$47k – $101k per year (Estimated) • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru
C++
Apply
$26k – $69k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
C++
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 2 hours ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru • Hyderabad • Chennai • Noida
Databases
Db2
IMS
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.