368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$129k – $282k per year (Estimated)
Location
Remote (United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
We design custom-built solutions to help you transform, scale, and grow your business along with a team that cares about you. Salvo software is a global firm with near-shoring capabilities headquartered in Vancouver, WA. That provides cost-effect...

About Salvo Software

Salvo Software is a global firm that provides cost-effective software solutions to guide enterprises and startups through digital transformation. With distributed teams across the US, LATAM, and India, we partner with clients to build high-performance, scalable systems that solve complex technical challenges. Our culture values innovation, ownership, and engineering excellence.

Role Overview

We are seeking a highly skilled AI Developer with a strong backend and machine learning engineering background to design, train, optimize, and deploy LLM models in on-prem and offline environments. This role is deeply technical and hands-on.

You will work closely with our engineering and product teams to build end-to-end LLM pipelines - including data preprocessing, supervised fine-tuning, model quantization, evaluation, RAG pipeline design, and deployment using local or air-gapped infrastructure. If you enjoy working with cutting-edge open-source LLMs, building context-aware AI systems, and designing reliable backend pipelines, this role is for you.

Key Responsibilities

Core LLM Development

  • Train and fine-tune LLMs using supervised fine-tuning (SFT).
  • Work with open-source models such as LLaMA, Mistral, Qwen, and similar architectures.
  • Build LoRA / Q-LoRA pipelines for efficient fine-tuning.
  • Implement and optimize data preprocessing workflows, including tokenization and long-context handling.
  • Use and extend Hugging Face Transformers & Datasets for training and inference.
  • Parse and process structured and semi-structured data, including XML/XSD files.
  • Implement document parsing solutions for Office formats (python-docx, OpenXML).

RAG & Context-Aware Systems

  • Design and implement end-to-end Retrieval-Augmented Generation (RAG) pipelines for document-grounded question answering and knowledge retrieval.
  • Build and maintain vector stores and embedding pipelines using tools such as FAISS, Chroma, Weaviate, or pgvector.
  • Optimize retrieval strategies including hybrid search, re-ranking, and chunking approaches tailored for domain-specific corpora.
  • Develop and maintain MCP (Model Context Protocol) server integrations to enable LLMs to interact dynamically with tools, APIs, and external data sources.
  • Design agentic workflows that leverage MCP to give models structured access to internal systems and context in a controlled, auditable manner.

Offline / On-Prem Model Expertise

  • Deploy, run, and maintain models fully offline and in air-gapped environments.
  • Perform model optimization and quantization (GGUF, GPTQ, AWQ, bitsandbytes).
  • Build and maintain inference systems using frameworks like vLLM, TGI, and Ollama.
  • Optimize GPU usage (CUDA, cuDNN, VRAM-aware batching).
  • Maintain local CI/CD pipelines for ML models without cloud dependencies.
  • Manage local model registries, versioning, and artifacts.
  • Ensure RAG and MCP components are fully operational in offline and restricted network environments.

Backend & DevOps

  • Build backend services in Python for ML training and inference workflows.
  • Work with relational databases (Postgres/MySQL) and vector databases for RAG storage layers.
  • Use Docker and Git for reliable development and deployment pipelines.
  • Use Azure DevOps for CI/CD, including local runners when applicable.

Requirements

Technical Skills

  • Strong experience in Python for backend and machine learning development.
  • Expertise with ML frameworks such as PyTorch or TensorFlow, along with scikit-learn and pandas.
  • Solid knowledge of Postgres or MySQL for data storage.
  • Experience with Docker and Git.
  • Hands-on experience with LLM training, fine-tuning, and optimization.
  • Experience with Hugging Face Transformers & Datasets.
  • Familiarity with XML/XSD and Office document parsing tools.
  • Experience deploying models with vLLM, TGI, or Ollama.
  • Understanding of quantization techniques such as GGUF, GPTQ, or AWQ.
  • Experience with GPU optimization and the CUDA stack.
  • Experience building solutions for offline, on-prem, and air-gapped environments.
  • Hands-on experience designing and implementing RAG pipelines, including embedding models, vector stores, and retrieval optimization strategies.
  • Experience building or integrating MCP (Model Context Protocol) servers to connect LLMs with external tools, APIs, and structured data sources.
  • Experience with advanced RAG techniques such as HyDE or multi-hop retrieval.
  • Experience building agentic systems using MCP in production or near-production environments.

Nice to Have

  • Experience managing ML model registries in offline environments.
  • Familiarity with AWS for hybrid deployments.
  • Experience with secure environments, restricted networks, or enterprise compliance requirements.

Soft Skills

  • Experience discussing complex technical topics with both technical and non-technical stakeholders.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • Full-Time • 7+ years exp • Bachelor's Degree • Ho Chi Minh City
JavaScript
Python
AI/ML
Copilot
Cursor
DevOps
Azure DevOps
Jenkins
Azure
GitLab
QA
JMeter
k6
Playwright
Postman
Rest-Assured
Selenium
TestRail
Apply
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$15k – $39k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Bengaluru
COBOL
Java
TypeScript
JavaScript
Java
Hibernate
Databases
Apache Solr
AI/ML
AI Agents
Edge AI
Frontend
Angular
DevOps
Azure
Apply
$13k – $30k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bengaluru
Databases
Databricks
Oracle
AI/ML
Spark
AI Agents
Edge AI
DevOps
Azure
Analytics
ETL/ELT
Apply
AI Developer 4 months ago
$29k – $119k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bengaluru
Python
Databases
Chroma
FAISS
MySQL
pgvector
PostgreSQL
Weaviate
AI/ML
AWQ
Bitsandbytes
CUDA Toolkit
Fine-tuning
GGUF
GPTQ
Hybrid Search
LLM
LoRA
Mistral
Model Context Protocol
Ollama
PyTorch
Quantization
Qwen
RAG
Scikit-learn
TensorFlow
Tokenization
Transformers
vLLM
PEFT
CUDA
cuDNN
Hugging Face
SFT
AI Agents
TGI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
Git
Vector
Apply
AI Developer 10 months ago
Remote • Full-Time • 3+ years exp • Bachelor's Degree
Bash
C++
Go
Kotlin
PowerShell
Python
Rust
Databases
Apache Kafka
AI/ML
CUDA
CUDA Toolkit
Embeddings
GGUF
llama.cpp
LLM
Model Context Protocol
Ollama
Quantization
RAG
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
Git
GitHub Actions
GitLab CI
Grafana
kubectl
Kubernetes
Prometheus
Terraform
Vector
GitHub
GitLab
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.