368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$32k – $62k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Eka Care provides a personal health record and clinic management platform. Patients store prescriptions and reports while doctors manage consultations digitally. The company integrates with national digital health infrastructure.

Kernel Optimisation & Inference Engineer

Bengaluru · Full-time · 2-5 yrs

Somewhere between the model and the silicon, 10-20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.

About EkaCare and the mission

EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.

The role:

You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.

What you'll do

  • Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
  • Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
  • Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
  • Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
  • Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) - with eval-parity verification, not just perplexity.
  • Make on-device inference real for the hardware Indian clinics actually have.

What we look for

  • 2-5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
  • Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).
  • Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
  • Measurement discipline: you profile before optimising and verify correctness after.

Bonus

  • fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
  • Quantisation research or on-device/mobile inference experience.
  • Open-source kernels or serving contributions.

Why this is a rare gig

  • Open source, with your name on it: weights and technical reports ship publicly.
  • India-scale mission: models for a billion people in their own languages.
  • Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
  • Small senior team: you work with the people who own the recipe.
  • A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.

Full-Time Employee Benefits

  • Medical Insurance & Accidental Insurance
  • Maternity & Paternity Benefits
  • PF, Gratuity, & Leave Encashment
  • Salary Advance Policy
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
LLM Model Developer 4 hours ago
$28k – $77k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad
AI/ML
AI Agents
Fine-tuning
LLM
Apply
$135k – $147k per year • In office • Full-Time • 7+ years exp • PhD • New York
Python
Databases
Databricks
Delta Lake
AI/ML
A2A
AI Agents
Computer Vision
Gemini
Google ADK
LangChain
LangGraph
LLM
Model Context Protocol
Prompt Engineering
RAG
Vertex AI
DevOps
GCP
Apply
$20k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
AWS Bedrock
Computer Vision
Embeddings
LLM
Multimodal AI
Prompt Engineering
RAG
Time Series Forecasting
DevOps
Amazon S3
AWS
AWS Lambda
CI/CD
Git
Vector
Analytics
Power BI
Tableau
Apply
$149k – $365k per year • Remote • Full-Time • 5+ years exp • PhD • United States
Apex
Apex
Salesforce Data Cloud
AI/ML
Agentforce
AI Agents
LLM
LLM Guardrails
Prompt Engineering
Marketing
Salesforce
Apply
$141k – $306k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Belfast
Python
TypeScript
AI/ML
A2A
AI Agents
Anthropic
Claude
Claude Code
LangChain
LangGraph
LLM
Model Context Protocol
OpenAI
OpenAI Agents SDK
OpenAI Codex
DevOps
AWS
OpenTelemetry
Apply
$29k – $122k per year (Estimated) • In office • Full-Time • Bengaluru
Python
AI/ML
LLM
NLP
LLM Evaluation
Post-training
Apply
$27k – $57k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Apply
$27k – $74k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru
Python
AI/ML
LLM
TRL
vLLM
Transformers
DPO
GRPO
NVIDIA NeMo
Post-training
PPO
SFT
AI Agents
Apply
$21k – $62k per year (Estimated) • In office • Full-Time • 2+ years exp • Bengaluru
AI/ML
Megatron-LM
NVIDIA NeMo
Pre-training
Mixture of Experts
Apply
ML Engineer 22 days ago
$29k – $122k per year (Estimated) • In office • Full-Time • Bengaluru
Python
AI/ML
LLM
NLP
LLM Evaluation
Post-training
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.