698,699open jobs
41,391companies
99,310added this week
Browse all
Salary
$30k – $68k per year (Estimated)
Location
In office (Pune)
Seniority
Staff · 8+ years exp
Overview
Company
Impact
Profile match
Gruve delivers AI-native infrastructure & managed AI cybersecurity services built for enterprise workloads — with speed, governance, and measurable outcomes.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position summary: 

Senior technical lead for the SOC and for the PulseAI platform layer. Runs major-incident response for security and PulseAI platform events, owns the detection-content and tuning program - including Kubernetes/OpenShift security telemetry - and is L3 for PulseAI platform operations: owns root cause analysis, works platform defects with the Gruve PulseAI engineering team, leads OpenShift cluster lifecycle operations and emergency remediation, acts as vendor engineering liaison for Red Hat, the NVIDIA software stack and third-party platform vendors, partners with the Network Operations Consultant on infrastructure-side events, and deputises for the Security Operation Manager.

Key responsibilities:

  • Lead technical response on Severity 1/2 security and PulseAI platform incidents until management command engages; own the SLA status cadence (hourly / every 30 minutes on Severity 1) to customer authorised contacts; hand infrastructure-side events (network fabric, GPU hardware, storage) to the Network Operations Consultant and stay engaged on cross-domain incidents.
  • Own root cause analysis for PulseAI platform and OpenShift cluster incidents; reproduce and characterise platform defects and route them to Gruve PulseAI product engineering; own the corrective-action backlog and problem management to eliminate recurring incidents.
  • Own the PulseAI platform operations practice: severity classification standards, diagnostic runbooks for the PulseAI and OpenShift layers (control plane, operators, authentication/SSO, RBAC, tenancy and quota enforcement, model-serving endpoints, observability), Grafana observability standards, alert thresholds for the SLA appendix, and the monitor / remediate / escalate matrix as applied per customer.
  • Lead OpenShift cluster lifecycle operations for PulseAI customers - y-stream upgrades on customer approval within the agreed window of the Red Hat release, operator and platform-component changes, emergency vulnerability remediation of the cluster and platform - enforcing the pre-change backup gate, and coordinating node-pool, GPU driver/firmware and fabric changes with the Network Operations Consultant.
  • Own Kubernetes/OpenShift security operations for the engagement: onboarding kube-audit and workload telemetry into the SIEM, detection content for container attack paths (MITRE ATT&CK for Containers), Cilium/Hubble and OVN-Kubernetes flow use cases, admission-control and image/runtime security posture, and RBAC / service-account hygiene reviews.
  • Act as vendor engineering liaison for the platform stack: Red Hat for OpenShift product defects, NVIDIA for GPU Operator / driver / NVIDIA AI Enterprise (NIM) software issues, and the third-party platform variant (Rafay, vCluster, vNode, NVIDIA Run:ai, Red Hat OpenShift AI) where restoration is best-effort with committed vendor escalation.
  • Lead customer onboarding technically for the platform layer: deploy and validate OpenShift and PulseAI within the 14-day Ready-for-Install window, configure identity provider federation and initial tenancy structure, onboard the platform to monitoring via the agreed connectivity pattern (outbound collector, site-to-site VPN or jump host with just-in-time elevation), and complete the platform sections of the countersigned environment validation checklist.
  • Own the detection-content backlog and tuning program; own log-pipeline integration health with escalation into engineering.
  • Approve and execute high-risk security and platform changes; cross-train the SOC L1/L2 bench on Kubernetes/OpenShift platform operations; drive shift-quality audits and post-incident reviews for the SOC pod.
  •  

Mandatory Qualifications:

  • 8-11 years with prior senior/lead experience in SOC or 24×7 platform operations and genuine security-to-platform cross-domain fluency.
  • Incident command capability; deep SIEM content and query skills.
  • Strong Kubernetes/GKE security operations depth - onboarding kube-audit and workload telemetry, building detection content for container attack paths (MITRE ATT&CK for Containers), and tuning Cilium/Hubble-based use cases.
  • Deep Red Hat OpenShift / Kubernetes operations experience in production - multi-node cluster administration, operators, y/z-stream upgrades, RBAC, storage and networking, observability stack (Grafana, metrics, logs, alerting), backup/restore of cluster and platform state - with a track record of platform diagnostics and RCA.
  • GPU-cluster platform operations experience: NVIDIA GPU Operator/driver lifecycle, NVIDIA AI Enterprise components (NIM), DCGM-class telemetry, node health and capacity management for AI inference workloads on RTX PRO 6000 / HGX B300-class servers or equivalent.
  • Experience operating against contractual SLAs (acknowledgement, restoration, availability, service credits) and running vendor engineering escalations through to fix.
  • Working network troubleshooting sufficient to scope cluster-networking versus fabric faults jointly with the NOC; automation mindset.

Preferred Qualifications:

  • Advanced incident-handling / intrusion-analysis certification (e.g., GCIH, GCIA, GCFA or equivalent).
  • CKS or equivalent; experience running K8s posture/runtime security tooling in production.
  • Red Hat certifications (EX280 / EX380 / RHCE); exposure to Rafay, vCluster, NVIDIA Run:ai or Red Hat OpenShift AI; MLOps or AI-platform operations (model-serving endpoints, GPU scheduling, quota governance).
  • Exposure to HashiCorp Vault, SAML 2.0 SSO federation, and CSI/NFS storage for model artefacts.
  • MSSP/managed-services background; AI-SOC tooling exposure.
  • Correlating GPU-platform performance anomalies (utilisation, thermal, fabric saturation) with security events to separate abuse, crypto-mining or misconfiguration from genuine workload load.

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,699 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Pune
In office
Python
Go
Bash
Databases
ArangoDB
AI/ML
AI Agents
Edge AI
DevOps
Terraform
GCP
CircleCI
Prometheus
CI/CD
GitOps
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Linux
Apply
$35k – $84k per year (Estimated) • Remote • Full-Time • Moscow
Databases
PostgreSQL
ClickHouse
Qdrant
DevOps
Ansible
Loki
Prometheus
Yandex Cloud
CI/CD
Docker
Kubernetes
Grafana
Harbor
GitLab
Linux
VPN
Cybersecurity
Keycloak
LDAP
Apply
In office • 4+ years exp • Bachelor's Degree
Python
Go
JavaScript
TypeScript
Databases
Apache Kafka
Frontend
React.js
DevOps
OpenTelemetry
Prometheus
Kubernetes
Grafana
Platform Engineering
Amazon S3
Linux
Apply
DevOps Engineer 1 day ago
$26k – $56k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad • Bengaluru • Mumbai • Chennai • Pune
DevOps
Terraform
Ansible
OpenShift
Helm
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Configuration Management
Amazon EKS
GitHub
Apply
$28k – $66k per year (Estimated) • In office • Full-Time • 10+ years exp • Shanghai
Python
Go
SQL
Databases
GraphDB
Databricks
RabbitMQ
Apache Kafka
AI/ML
LangGraph
LangChain
DeepSpeed
vLLM
MLFlow
Vertex AI
Fine-tuning
AI Agents
SGLang
torchtune
AWS Bedrock
Kubeflow
PyTorch
LLM
RAG
Ray
OpenAI
LLMOps
SFT
Megatron-LM
Feature Store
InfiniBand
LLM Guardrails
Edge AI
Frontend
GraphQL
DevOps
Rest API
gRPC
Terraform
Ansible
GCP
OpenTelemetry
CloudFormation
Prometheus
GitLab CI
SLURM
Azure
CI/CD
Git
AWS
Kubernetes
Cloudflare
Grafana
Platform Engineering
JFrog Artifactory
CentOS Stream
FinOps
GitHub
GitLab
Amazon S3
IAM
Linux
Cybersecurity
LDAP
Management
Confluence
Jira
Agile
Scrum
Apply
$15k – $43k per year (Estimated) • In office • 4+ years exp • Pune
Python
AI/ML
NCCL
NVLink
DevOps
Ansible
GCP
Red Hat
OpenShift
Cilium
Kubernetes
Cloudflare
Grafana
kubectl
Google GKE
HPC
Linux
BGP
Apply
$13k – $31k per year (Estimated) • In office • 2+ years exp • Pune
AI/ML
NVLink
DevOps
GCP
Red Hat
OpenShift
Cilium
Kubernetes
kubectl
Google GKE
SLI/SLO/SLA
Linux
BGP
Management
ITSM
Apply
$25k – $55k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Mumbai
DevOps
GCP
Azure
AWS
Incident Management
TCP/IP
VPN
Cybersecurity
Palo Alto NGFW
DLP
Management
ITIL
Apply
$125k – $180k per year • Remote • 8+ years exp • Bachelor's Degree
Cybersecurity
Zero Trust
Apply
$25k – $52k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Pune
JavaScript
Java
TypeScript
Java
Spring Boot
Databases
Apache Kafka
AI/ML
Copilot
Claude
Claude Code
Prompt Engineering
AI Agents
Frontend
Angular
DevOps
Azure
CI/CD
Jenkins
Git
Docker
Kubernetes
Cybersecurity
IBM QRadar
LDAP
SIEM
Management
Jira
Agile
Scrum
ITSM
Apply
In office • Full-Time • Pune
Apply
In office • Pune
Apply
$25k – $53k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Pune
JavaScript
DevOps
GCP
Azure
AWS
Apply
Remote/Hybrid • 4+ years exp • Pune
Python
SQL
Analytics
ETL/ELT
Informatica
Master Data Management
Apply
$9k – $18k per year (Estimated) • In office • Internship • Pune
Apply
See all jobs
This is one of many
698,699 more open roles from verified company boards, updated every day.