368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$31k – $78k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Forbes is a leading global media and branding company that focuses on business, investing, technology, entrepreneurship, leadership, and luxury lifestyle. It is widely recognized for its high-profile journalism, in-depth financial analysis, and influential lists, such as the Forbes 400 and the 30 Under 30, which track wealth and achievement worldwide. Through its digital platform and print publications, the organization provides news and insights that empower business leaders, investors, and entrepreneurs to succeed in a changing global economy.

The candidate will have responsibilities across the following functions:

Self-Hosted LLM Infrastructure:

  • Deploy, fine-tune, and operate open-source models (Llama, Qwen, MedGemma, and successors) as our primary inference stack.
  • Work with vLLM / SGLang / TensorRT-LLM for serving at scale, with disciplined attention to throughput, tail latency, batching, KV-cache, and GPU economics.
  • Own fine-tuning workflows end-to-end (SFT, LoRA, QLoRA, DPO) on clinical notes, claims, and payer-rule data.
  • Optimise GPU usage, latency, batching, and cost; make build-vs-buy and hosted-vs-self-hosted trade-offs explicit and measured.

Knowledge Graphs and Embedding-Based Retrieval:

  • Design and maintain the knowledge graph encoding ICD-10-CM, CPT, HCPCS, modifiers, HCC, NCCI edits, LCD/NCD policies, and payer-specific rules and the relationships between them.
  • Build embedding-based retrieval over clinical notes, historical claims, denial reasons, and payer-policy corpora, including chunking, embedding model selection, hybrid search, and reranking.
  • Combine graph traversal and dense retrieval so every coded line, scrubbed edit, and appeal response is grounded in auditable evidence.
  • Own ingestion, versioning, and quality of underlying knowledge sources (CMS, AHA, AMA, NCCI, payer bulletins).

Evaluation and Monitoring:

  • Build continuous evaluation pipelines that gate every model, prompt, retrieval, and graph change before production.
  • Run offline eval suites grounded in coder- and biller-validated labels; use LLM-as-judge where appropriate, calibrated against human ground truth.
  • Monitor drift, hallucinations, regressions, and output quality in production; operate shadow-mode rollouts and per-cohort accuracy tracking (speciality, payer, chart type).
  • Track business metrics: chart-level and opportunity-level coding accuracy, denial rate impact, clean-claim rate, cost per chart, and end-to-end latency.

LLM Systems and Prompt Engineering:

  • Design prompts and context pipelines for coding (CPT, ICD, HCC, E/M), claim edits, denial classification, and appeal drafting.
  • Implement structured outputs (JSON, function calling, constrained decoding) on top of the self-hosted stack.
  • Apply RAG over medical coding standards (CMS, ICD-10 AHA, NCCI) and payer policies, grounded in the knowledge graph and embedding stores.
  • Treat prompts as a thin, well-versioned, well-evaluated layer, never the load-bearing piece.

Agentic Workflows and Tooling MCP:

  • Build MCP servers for internal tools: code lookup, NCCI / rule checks, payer logic, eligibility, denial classification.
  • Design multi-step agent workflows with audit trails and human-in-the-loop checkpoints for coder, biller, and AR-analyst review.
  • Define deterministic vs. LLM-based tool boundaries for reliability; reliability comes from knowing which is which.

Requirements:

  • 5+ years in ML/AI engineering, including 6+ months in production LLM systems.
  • Hands-on experience deploying and operating self-hosted LLMs (vLLM, SGLang, TensorRT-LLM, or equivalent).
  • Strong experience designing embedding-based retrieval and/or knowledge graphs for grounded LLM applications.
  • Demonstrated ownership of evaluation infrastructure, offline benchmarks, online monitoring, and drift and regression detection.
  • Strong Python + PyTorch + Hugging Face experience.
  • Production experience with monitoring, incidents, and system ownership.

Strongly Preferred:

  • Fine-tuning experience (SFT, LoRA, QLoRA, DPO) on domain-specific corpora.
  • Experience with graph databases (Neo4j, ArangoDB, or equivalent) and graph-aware retrieval.
  • Experience with vector databases and hybrid search (BM25 + dense, rerankers).
  • Familiarity with LLM observability tools (Langfuse, LangSmith, Arize, Braintrust, or in-house equivalents).
  • Exposure to healthcare, RCM, claims, or other regulated domains.
  • Experience with MCP or similar tool-orchestration frameworks.
  • Strong prompt-engineering and LLM-evaluation instincts.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
NLP / LLM Engineer 10 hours ago
$18k – $24k per year (net) • In office • Full-Time • 3+ years exp • Tashkent
Python
Python
FastAPI
Databases
ElasticSearch
Milvus
Pinecone
Qdrant
Weaviate
AI/ML
AI Agents
ChatGPT
DeepEval
Embeddings
Gemini
Hybrid Search
LangChain
Langfuse
LangGraph
LlamaIndex
LLM
LoRA
NLP
PEFT
Prompt Engineering
PyTorch
QLoRA
RAG
Reranking
Semantic Search
Synthetic Data
Tokenization
Triton
vLLM
Transformers
Anthropic
DPO
GraphRAG
Hugging Face
OCR
OpenAI
Semantic Search
SFT
Structured Outputs
Function Calling
TGI
DevOps
CI/CD
Docker
Git
GitHub
Analytics
A/B Testing
Apply
$180k – $230k per year • Equity 0.8–3% • In office • Full-Time • 3+ years exp • Bachelor's Degree • New York
AI/ML
AI Agents
Fine-tuning
RLHF
Apply
$140k – $180k per year • Equity 0.8–3% • In office • Full-Time • 1+ year exp • Bachelor's Degree • New York
AI/ML
AI Agents
Fine-tuning
RLHF
Apply
Applied - AI Engineer 10 hours ago
$25k – $69k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru • Pune
Java
Python
AI/ML
AI Agents
AWS Bedrock
Fine-tuning
Google ADK
LangGraph
LLM
LoRA
RAG
Semantic Search
LangChain
PEFT
A2A
Amazon SageMaker
AWS Strands Agents
NIST AI RMF
Semantic Search
Model Context Protocol
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Git
GitOps
gRPC
Kubernetes
OpenTelemetry
Rest API
Terraform
Vector
Amazon S3
IAM
GitLab
Apply
Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Istanbul
JavaScript
TypeScript
AI/ML
Fine-tuning
Frontend
Angular
RxJS
Sass
DevOps
Azure
Bitbucket
Git
Apply
Data Engineer 10 days ago
$28k – $54k per year (Estimated) • Remote • 6+ years exp
Python
SQL
Databases
Google BigQuery
Marketing
GA4
Meta Ads
Apply
$22k – $60k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Bengaluru
Go
Node JS
Python
SQL
JavaScript
Databases
Apache Kafka
DevOps
AWS
Azure
Docker
GCP
Apply
Data Engineer L3 18 days ago
$16k – $61k per year (Estimated) • Remote
Python
SQL
Databases
Google BigQuery
AI/ML
dbt
DevOps
GCP
Analytics
ETL/ELT
Marketing
GA4
Google Ads
Meta Ads
Apply
Data Scientist 19 days ago
$21k – $53k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Bengaluru
Python
AI/ML
NLP
PyTorch
TensorFlow
DevOps
Docker
Kubernetes
Apply
$115k – $251k per year (Estimated) • Remote/Hybrid • Contractor • 7+ years exp • Bachelor's Degree • Jersey City
DevOps
CI/CD
Cybersecurity
HIPAA
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 1 hour ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru • Hyderabad • Chennai • Noida
Databases
Db2
IMS
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.