401,286open jobs
13,975companies
77,769added this week
Browse all
Location
In office (Chennai)
Seniority
Junior · 2+ years exp
Overview
Company
Impact
Profile match
E2E Networks is a listed Indian cloud provider that pivoted to accelerated computing for artificial intelligence workloads. Founded in 2009 in Delhi, it operates large GPU clusters in Indian data centres for startups, research institutions and enterprises. Its TIR platform provides managed notebooks, fine-tuning and inference endpoints.

Roles & Responsibilities:

1. Cluster and GPU Management :

  • •Launch, validate, and maintain GPU-based AI/ML training clusters (8xH100, 8xH200, 32xH200, 64xH200 upto 1024 H200s).
  • •Verify all cluster nodes have InfiniBand enabled and GPUs correctly assigned (no Ethernet fallback).
  • •Ensure Slurm deployments are up within a minute and all workers are ready (sinfo shows all active).
  • •Validate DGCX Workbench runs successfully for:
    • •Llama3-8B on 8xH100 / 8xH200 cluster
    • •Llama3-70B on 32xH200 / 64xH200 clusters

Monitor GPU health using tools and validate performance benchmarks.

Maintain cluster reliability - all training and inference nodes should remain up and restartable

without failure.

2. Inference and Endpoint Operations :

  • •Launch and monitor vLLM inference endpoints (e.g., Llama 70B) ensuring:
    • •First startup within 10 minutes
    • •Restart within 1 minute
    • •Autoscale brings new workers up within 3 minutes
    • •Inference endpoints remain continuously reachable and 100% ready
    • •Troubleshoot and stabilize stateful workloads, notebooks, and AI services.

3. Customer-Facing Technical Support :

    • •Engage directly with customers via calls, video meetings (Google Meet/Hangout), and screen-sharing sessions.
    • •Understand the customer’s problem in real time and guide them through the solution.
    • •Diagnose complex GPU, Slurm, or inference issues and resolve them collaboratively on the call.
    • •Provide clear updates and ensure timely resolution of support tickets.
    • •Document RCA and contribute to permanent fixes or product improvements.
    • •Communicate professionally and technically with data scientists, developers, and enterprise users.Automation and Reliability
    • •Automate cluster provisioning and monitoring using Terraform, Ansible, and Python.
    • •Create scripts for routine cluster health checks, GPU utilization, and job queue validation.
    • •Collaborate with the platform and DevOps teams to implement improvements for speed and reliability.

Key Skills & Qualifications:

  • •2-4 years of experience in GPU-based cloud operations, MLOps, or infrastructure engineering.
  • •Prior exposure to customer-facing roles or live technical troubleshooting calls.
  • •Experience working with AI model training pipelines, inference endpoints, or Slurm-managed clusters.
  • •Familiarity with LLM workloads such as Llama, Mistral, or Falcon models.
  • •Linux (Ubuntu/CentOS), system performance tuning
  • •Networking: InfiniBand, VLAN, VPN, ALB, DNS, NAT
  • •Containers & orchestration: Docker, Kubernetes, Helm
  • •GPU operations: CUDA, GPU drivers, nvidia-smi, MIG configuration
  • •Distributed training: Slurm, DDP (Distributed Data Parallel) concepts
  • •AI Inference: vLLM, TensorRT, ONNX Runtime, Hugging Face models
  • •Infrastructure as Code: Terraform, Ansible
  • •Tools: ssh, curl, tcpdump, Prometheus, Grafana, ELK

Apply Now

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
401,286 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Chennai
SRE Engineer 2 hours ago
Remote/Hybrid • 4+ years exp
DevOps
Ansible
AWS
CI/CD
Datadog
Docker
Grafana
Kubernetes
OpenTelemetry
Prometheus
SLI/SLO/SLA
Terraform
Apply
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
C++
C
C++
LLVM
PyTorch C++
TensorFlow C++
C
MPI
AI/ML
CUDA
CUDA Toolkit
CUTLASS
DeepSpeed
JAX
MLIR
NCCL
PyTorch
ROCm
TensorFlow
TensorRT
Triton
vLLM
DevOps
AWS
Azure
GCP
HPC
Apply
$84k – $100k per year • Remote • 8+ years exp • Bachelor's Degree
Python
SQL
Databases
Oracle
DevOps
Ansible
AWS
Azure
CI/CD
Apply
Rust developer 1 hour ago
$21k – $54k per year (Estimated) • Remote • Full-Time • Moscow
C#
C++
Rust
Databases
Apache Kafka
ClickHouse
Kafka
Redis
Tarantool
AI/ML
Flink
Frontend
WebAssembly
DevOps
Docker
eBPF
Git
GitHub
GitLab
Grafana
gRPC
OpenTelemetry
Prometheus
Splunk
Cybersecurity
IBM QRadar
MITRE ATT&CK
Apply
In office • 4+ years exp • Master's Degree
Python
AI/ML
Image Segmentation
PyTorch
TensorFlow
DevOps
AWS
Docker
GCP
Apply
Payroll Lead 30 days ago
In office • Full-Time • 7+ years exp • Master's Degree • Delhi
Apply
HR Business Partner 30 days ago
In office • Full-Time • Bachelor's Degree • Delhi
Apply
In office • Full-Time
Apply
Server Engineer 6 months ago
$14k – $40k per year (Estimated) • In office • 4+ years exp • Noida
DevOps
Windows Server
Apply
$31k – $75k per year (Estimated) • In office • Delhi
AI/ML
CUDA
CUDA Toolkit
Kubeflow
ONNX
PyTorch
Ray
TensorFlow
TensorRT
cuDNN
CVAT
InfiniBand
ONNX Runtime
DevOps
CentOS Stream
Debian
Docker
Kubernetes
KVM
SLURM
Ubuntu
VMWare
Apply
$37k – $82k per year (Estimated) • Equity • In office • 8+ years exp • Bachelor's Degree • Chennai
Python
AI/ML
AI Agents
LangChain
DevOps
AWS
Management
ServiceNow
Marketing
Instagram
LinkedIn
Salesforce
YouTube
Apply
Salesforce Analyst 1 day ago
$14k – $31k per year (Estimated) • Equity • In office • 2+ years exp • Bachelor's Degree • Chennai
Marketing
Salesforce
Instagram
LinkedIn
YouTube
Apply
Legal Ops Lead 1 day ago
$4.3k – $13k per year • In office • Full-Time • 11+ years exp • Chennai
AI/ML
LLM
Apply
$32k – $72k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Chennai
Apply
Remote/Hybrid • Full-Time • 2+ years exp • Chennai
Apply
See all jobs
This is one of many
401,286 more open roles from verified company boards, updated every day.