698,699open jobs
41,391companies
99,310added this week
Browse all
Salary
$15k – $43k per year (Estimated)
Location
In office (Pune)
Seniority
Middle · 4+ years exp
Overview
Company
Impact
Profile match
Gruve delivers AI-native infrastructure & managed AI cybersecurity services built for enterprise workloads — with speed, governance, and measurable outcomes.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position summary: 

L2 escalation owner for the network/DMT track and shift anchor for the NOC pod, and L2 for the PulseAI infrastructure layer. Deep RCA on the Juniper fabric and edge network devices and Cisco firewalls, non-standard change execution and vendor TAC ownership - taking handoff from the L1 NOC bench and closing out or escalating to the Network Operations Consultant (L3) - plus remediation of GPU-server, node, fabric and OpenShift cluster-networking faults within Gruve's scope and execution of firmware/driver, switch-configuration and node-lifecycle changes within maintenance windows. Security investigations and PulseAI platform-layer remediation belong to the SOC pod and are not part of this role.

Key responsibilities:

  • Act as shift anchor: own P1/P2 first response, open and run technical bridges until L3 engagement; own escalations from the NOC/DMT queue through resolution or structured L3 handoff.
  • Complex network RCA (EVPN/BGP, fabric and edge health, interface/optics issues) across the in-scope Juniper data-center devices (QFX, EX, MX) and Cisco cdFMC/FMC, FTD/SRX; correlate fabric events with OpenShift node and cluster-network symptoms to separate network-side from cluster-side faults.
  • Execute non-standard and complex changes beyond L1 authority under change control - fabric node additions, routing changes, firmware upgrades, firewall policy pushes - with senior oversight.
  • Remediate PulseAI infrastructure incidents within Gruve's scope: GPU-server firmware and driver faults (within the change process), OpenShift node and cluster-networking issues (OVN-Kubernetes/CNI, node NotReady from hardware or network causes, MachineConfig drift), front-end and RoCEv2 back-end fabric connectivity, switch configuration within the PulseAI fabric; diagnose storage capacity/health and node-hardware faults and escalate to the vendor with complete diagnostics.
  • Execute firmware and GPU driver updates, switch configuration changes and OpenShift node lifecycle operations (MachineConfig, node cordon/drain/reboot, node-pool changes) within agreed maintenance windows with the required customer notice; verify post-change node and cluster-network health with oc/kubectl.
  • Own vendor TAC cases end to end - Juniper for fabric and edge; Cisco, and Cloudflare where a network issue touches that platform; OEM, neocloud-provider and storage-vendor cases for hardware - coordinate RMA/smart-hands, deliver daily status on Premium, and manage restoration-clock pauses correctly.
  • Drive preventive maintenance: firmware/BIOS/driver currency, EOL/EOS tracking, spare posture - for network devices and PulseAI GPU and control-plane nodes.
  • Own infrastructure observability for the NOC pod: GPU-cluster, fabric and OpenShift node health dashboards, log queries and alert thresholds in Grafana; contribute to the environment validation checklist at onboarding (telemetry reachability per switch, storage throughput, access path).

Mandatory Qualifications:

  • 4-6 years NOC/network operations experience in data-center environments, including hands-on Junos (EVPN-VXLAN fabric) and/or Cisco cdFMC/FMC, FTD.
  • Demonstrated independent RCA ownership on P1/P2/P3 network incidents; incident bridge experience; comfortable running a shift independently.
  • Strong routing/switching troubleshooting (BGP, EVPN-VXLAN) and next-generation firewall operations; disciplined change-management practice.
  • Hands-on Red Hat OpenShift / Kubernetes operations in production - node lifecycle and MachineConfig, operators, cluster networking (OVN-Kubernetes/CNI), storage, oc/kubectl troubleshooting - plus Linux (RHEL) administration.
  • Working understanding of Kubernetes networking - CNI (Cilium), Services/Ingress/LoadBalancer and east-west flows - to troubleshoot fabric-to-cluster (GKE and OpenShift) connectivity end to end.
  • Working understanding of GPU-server operations: NVIDIA GPU Operator and driver stack, DCGM-class telemetry, firmware/driver update procedures, RDMA/RoCEv2 NIC health and common GPU failure modes.
  • Multi-vendor switch monitoring and troubleshooting (SNMP, syslog, streaming telemetry); out-of-band management proficiency.

Preferred Qualifications:

  • JNCIP/JNCIS-DC or CCNP, or other professional-level data-center networking certification.
  • Apstra or other intent-based networking exposure; IaC exposure; Python scripting / Ansible for operations automation.
  • GKE cluster networking exposure (VPC-native, LoadBalancer/Gateway) and NetworkPolicy troubleshooting; working knowledge of GCP.
  • Red Hat OpenShift Administration certification (EX280) or RHCSA/RHCE; CKA; exposure to AI/ML workload scheduling and GPU node pools on OpenShift.
  • GPU-cluster performance troubleshooting - RoCEv2/PFC/ECN tuning, ECMP polarisation, NCCL-visible latency/jitter - and hands-on with high-performance fabric telemetry, NVLink/NVSwitch topologies and GPU-node network profiling.
  • AI/HPC, neocloud or hyperscale data-center fabric exposure.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,699 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Pune
In office • Full-Time • New Zealand
Python
DevOps
Splunk
Red Hat
OpenShift
Helm
Kustomize
Prometheus
ArgoCD
Jenkins
Kubernetes
Grafana
Service Mesh
Bitbucket
GitHub
Linux
DNS
Cybersecurity
CyberArk
Management
Confluence
Jira
ServiceNow
Agile
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
$135k – $288k per year (Estimated) • Equity • Remote • 7+ years exp
Python
SQL
Databases
PostgreSQL
ClickHouse
RabbitMQ
AI/ML
Claude Code
AI Agents
LLM
LLM Guardrails
DevOps
Kubernetes
Analytics
ETL/ELT
Apply
$73k – $184k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Sydney
Python
AI/ML
Anomaly Detection
DevOps
Terraform
GCP
OpenTofu
Dynatrace
GitLab CI
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Amazon ECS
Cybersecurity
Wiz
Apply
In office
Databases
RabbitMQ
Apache Kafka
DevOps
Terraform
CI/CD
AWS
Kubernetes
Grafana
Amazon EKS
Amazon EC2
Amazon S3
DNS
Apply
$30k – $68k per year (Estimated) • In office • 8+ years exp • Pune
DevOps
GCP
OpenShift
Cilium
Kubernetes
Grafana
Google GKE
SLI/SLO/SLA
VPN
Cybersecurity
HashiCorp Vault
SIEM
Apply
$13k – $31k per year (Estimated) • In office • 2+ years exp • Pune
AI/ML
NVLink
DevOps
GCP
Red Hat
OpenShift
Cilium
Kubernetes
kubectl
Google GKE
SLI/SLO/SLA
Linux
BGP
Management
ITSM
Apply
$25k – $55k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Mumbai
DevOps
GCP
Azure
AWS
Incident Management
TCP/IP
VPN
Cybersecurity
Palo Alto NGFW
DLP
Management
ITIL
Apply
$125k – $180k per year • Remote • 8+ years exp • Bachelor's Degree
Cybersecurity
Zero Trust
Apply
$25k – $52k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Pune
JavaScript
Java
TypeScript
Java
Spring Boot
Databases
Apache Kafka
AI/ML
Copilot
Claude
Claude Code
Prompt Engineering
AI Agents
Frontend
Angular
DevOps
Azure
CI/CD
Jenkins
Git
Docker
Kubernetes
Cybersecurity
IBM QRadar
LDAP
SIEM
Management
Jira
Agile
Scrum
ITSM
Apply
In office • Full-Time • Pune
Apply
In office • Pune
Apply
$25k – $53k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Pune
JavaScript
DevOps
GCP
Azure
AWS
Apply
Remote/Hybrid • 4+ years exp • Pune
Python
SQL
Analytics
ETL/ELT
Informatica
Master Data Management
Apply
$9k – $18k per year (Estimated) • In office • Internship • Pune
Apply
See all jobs
This is one of many
698,699 more open roles from verified company boards, updated every day.