698,688open jobs
41,428companies
98,846added this week
Browse all
Salary
$22k – $54k per year (Estimated)
Location
In office (Pune)
Seniority
Middle · 4+ years exp
Overview
Company
Impact
Profile match
Gruve delivers AI-native infrastructure & managed AI cybersecurity services built for enterprise workloads — with speed, governance, and measurable outcomes.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position summary:

Senior analyst and shift anchor for the SOC pod, and L2 for PulseAI Managed Services. Owns triage quality across the rotation, handles high-severity security incidents through to handoff, and remediates PulseAI and OpenShift incidents to the Standard/Premium restoration objectives (Severity 1 in 4/2 hours), executes changes and patching within maintenance windows, and manages escalations to L3 platform engineering and to hardware vendors.

Key Roles & Responsibilities:

  • Act as shift senior: final quality gate on investigations and escalations; handle P1/P2 security alerts end-to-end to L3 handoff including scoping and evidence preservation.
  • Drive shift-level metrics: SLA adherence, false-positive rate, reopen rate.
  • Remediate PulseAI platform incidents within the authority matrix - control-plane and platform-service recovery, authentication/SSO and RBAC faults, tenancy and quota-enforcement failures, endpoint deployment and model-serving failures, observability outages - and restore PulseAI configuration state (organisation/department/project hierarchy, quota allocations, users and roles) from Gruve backups when required.
  • Remediate OpenShift cluster incidents - node NotReady, scheduling and capacity, operator degradation, cluster networking, storage/PVC faults, image registry and ingress - using oc/kubectl and cluster diagnostics; hand structured RCAs to L3.
  • Execute PulseAI patch releases and OpenShift z-stream patches in the monthly maintenance window, firmware and GPU driver updates as required, and emergency security remediation within the tier window (72 hours Premium / 5 business days Standard); enforce the pre-change gate that no cluster upgrade proceeds without the customer's written confirmation of a verified backup.
  • Manage hardware escalations opened by L1: drive OEM/neocloud/storage vendor cases to closure, keep the customer informed at the SLA cadence, and pause/resume restoration clocks correctly when waiting on the customer, a vendor or a change approval.
  • Own platform observability hygiene: Grafana dashboards, alert thresholds recorded in the SLA appendix (performance degradation, capacity), and runbooks for recurring platform faults.
  • Coach and quality-review the L1 analysts; own runbook accuracy for the SOC pod; coordinate cross-tower with the NOC anchor on ambiguous events.

Mandatory Qualifications:

  • BE/BTech (CS/IT/E&TC) or equivalent.
  • 4-6 years SOC or 24×7 platform-operations experience with demonstrable incident handling and remediation ownership.
  • Strong multi-source triage across cloud, network and identity telemetry.
  • Hands-on triage of Kubernetes/container security alerts - kube-audit events, workload anomalies and pod-level network flows (Cilium/Hubble) - ideally on GKE.
  • Hands-on Red Hat OpenShift / Kubernetes administration in production - troubleshooting pods, nodes, operators, storage and cluster networking with oc/kubectl; applying z-stream patches and operator updates within change control; reading control-plane and workload logs in a metrics/logging observability stack (Grafana).
  • Working understanding of GPU-node operations - NVIDIA GPU Operator and driver stack, DCGM-class telemetry, GPU firmware/driver update process, common GPU health failure modes - and of the monitor / remediate / escalate boundaries between platform, hardware and customer workload.
  • Audit-grade documentation; ability to run a shift independently and calmly under Severity 1 pressure and to communicate status to customer authorised contacts.
  • Mentoring aptitude - this role carries the shift's junior bench.

Preferred Qualifications:

  • Incident-handling / forensics certification-level knowledge (e.g., GCIH, CHFI or equivalent); scripting for enrichment/automation.
  • CKA/CKS or KCSA; exposure to admission controls and image/runtime security tooling.
  • Red Hat OpenShift Administration certification (EX280) or RHCSA; exposure to NVIDIA AI Enterprise (NIM microservices), model-serving endpoints and GPU workload scheduling on OpenShift.
  • Exposure to HashiCorp Vault, SAML 2.0 SSO / RBAC troubleshooting, and n8n or similar platform applications.
  • Prior GPU-cloud, neocloud, hyperscale or data-center customer exposure.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,688 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Pune
In office • 2+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Node JS
Node JS
Commander.js
Databases
PostgreSQL
AI/ML
AI Agents
LLM
DevOps
Terraform
GitHub Actions
OpenTelemetry
PagerDuty
Prometheus
CI/CD
Git
AWS
Kubernetes
Grafana
Platform Engineering
Self-Healing
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
Incident Management
SLI/SLO/SLA
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
DNS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Apply
$33k – $73k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Terraform
OpenShift
Helm
OpenTelemetry
Prometheus
AWS
Docker
Kubernetes
Grafana
Thanos
Amazon EKS
Incident Management
TCP/IP
DNS
Cybersecurity
ISO 27001
PCI DSS
SOC 2
HIPAA
Apply
$21k – $45k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Penza
Bash
Databases
PostgreSQL
DevOps
Ansible
Helm
Prometheus
GitLab CI
CI/CD
Git
Docker
Kubernetes
Ubuntu
Nginx
Grafana
Linux
Apply
Golang Developer 2 days ago
$33k – $70k per year (Estimated) • In office • Full-Time • Bengaluru
Go
SQL
DevOps
Rest API
CI/CD
Git
Docker
Kubernetes
Apply
$89k – $238k per year (Estimated) • Remote • 3+ years exp
Python
Go
Bash
AI/ML
Copilot
Cursor
Claude
AI Agents
LLM
RAG
Agentic Workflows
DevOps
Terraform
GCP
CloudFormation
Pulumi
Azure
CI/CD
Git
AWS
Docker
Kubernetes
FinOps
Incident Management
GitHub
IAM
Linux
Unix
Apply
$15k – $43k per year (Estimated) • In office • 4+ years exp • Pune
Python
AI/ML
NCCL
NVLink
DevOps
Ansible
GCP
Red Hat
OpenShift
Cilium
Kubernetes
Cloudflare
Grafana
kubectl
Google GKE
HPC
Linux
BGP
Apply
$30k – $68k per year (Estimated) • In office • 8+ years exp • Pune
DevOps
GCP
OpenShift
Cilium
Kubernetes
Grafana
Google GKE
SLI/SLO/SLA
VPN
Cybersecurity
HashiCorp Vault
SIEM
Apply
$13k – $31k per year (Estimated) • In office • 2+ years exp • Pune
AI/ML
NVLink
DevOps
GCP
Red Hat
OpenShift
Cilium
Kubernetes
kubectl
Google GKE
SLI/SLO/SLA
Linux
BGP
Management
ITSM
Apply
$25k – $55k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Mumbai
DevOps
GCP
Azure
AWS
Incident Management
TCP/IP
VPN
Cybersecurity
Palo Alto NGFW
DLP
Management
ITIL
Apply
$125k – $180k per year • Remote • 8+ years exp • Bachelor's Degree
Cybersecurity
Zero Trust
Apply
In office • Full-Time • Pune
Apply
In office • Pune
Apply
$25k – $53k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Pune
JavaScript
DevOps
GCP
Azure
AWS
Apply
Remote/Hybrid • 4+ years exp • Pune
Python
SQL
Analytics
ETL/ELT
Informatica
Master Data Management
Apply
$9k – $18k per year (Estimated) • In office • Internship • Pune
Apply
See all jobs
This is one of many
698,688 more open roles from verified company boards, updated every day.