937,957open jobs
57,076companies
155,728added this week
Browse all
Salary
≈ $97k – $223k per year (Estimated)
Location
In office (United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 24, 2026.

Overview
Company
Impact
Profile match
At Sterling, we’re more than an IT solutions provider—we’re a team of passionate innovators and committed problem solvers.

Title: Backend ML Engineer

Reports to: Senior Software Architect

Location: North Sioux City, SD

Job Description: Sterling Computers is a technology company that provides IT solutions to a variety of clients, including the federal government, state and local governments, education, and commercial entities. Sterling's Strategic Technologies Group is responsible for learning and becoming subject matter experts in new and emerging technologies. Our team uses this expertise to broaden the portfolio of products and solutions that the company sells, delivers, and manages. Our engineers work on a range of AI-integrated systems, from production RAG platforms and LLM orchestration layers to digital human solutions and intelligent automation pipelines. We are looking for a Backend ML Engineer who is interested in taking AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, integrating large language models, and operating ML infrastructure at scale. If you thrive in a collaborative, client-focused environment and enjoy shipping AI features that real users depend on, we'd love to have you on our team.

Required Technical Skills:

  • 0-3 years of experience in backend or ML engineering
  • Strong working knowledge of Python, including FastAPI
  • Experience with python libraries such as sentence-transformers, OpenCV, pillow, and other NLP or CV libraries
  • Strong understanding of scalability and latency in ML systems
  • Experience with multi step agentic systems and tooling structures
  • Hands-on experience integrating LLMs (OpenAI, Anthropic, Gemini, or open-source models) into production systems
  • Familiarity with vector databases such as Weaviate, Pinecone, or similar
  • Experience with different retrieval-augmented generation (RAG) architectures
  • Self-motivated with a positive and professional attitude
  • Ability to adapt to other languages across the stack as needed.

Required Education/Experience:

  • Bachelor’s degree in Computer Science, Machine Learning, or a related field (minimum requirement), or equivalent practical experience
  • Graduate-level coursework or specialization in ML/AI is a plus
  • Relevant cloud certifications are a plus
  • Demonstrated experience shipping ML systems to production is a plus
  • US DoD Clearance preferred or willingness to obtain such

Qualifications:

  • Strong experience building backend services with Python (FastAPI/Flask); comfort working with async APIs and request/response patterns for ML inference workloads.
  • Hands-on experience integrating LLMs and embedding models into production applications, including prompt engineering, context management, and handling rate limits, retries, and streaming responses.
  • Familiarity with RAG architectures: chunking strategies, embedding pipelines, vector search, reranking, and evaluation metrics (Recall@k, MRR, faithfulness, answer relevance).
  • Experience with vector databases (Weaviate, pgvector, Pinecone, Qdrant, or similar) and traditional databases (PostgreSQL, MariaDB) for hybrid retrieval and metadata filtering.
  • Cloud experience (AWS/GCP/Azure) for deploying ML services - including managed inference endpoints, GPU instances, or serverless model hosting.
  • Strong understanding of API authentication, secure handling of model inputs/outputs, and PII/PHI-aware design where applicable.
  • Experience with ML observability: tracking latency, token usage, cost-per-query, retrieval quality, and model drift in production.
  • Background in data pipelines, document ingestion/parsing, or evaluation frameworks (Ragas, TruLens, Docling, custom harnesses) is needed.
  • Familiarity with fine-tuning, LoRA/PEFT, or model distillation is appreciated.
  • Experience with MLOps tooling (MLflow, Weights & Biases, Kubeflow) or LLM orchestration frameworks (LangChain, LlamaIndex, Haystack, or custom orchestrators) is a plus.

Responsibilities:

  • Build, test, and maintain production ML services - inference APIs, retrieval pipelines, orchestration layers, and guardrail/evaluation components.
  • Design scalable RESTful and streaming APIs that serve ML model outputs reliably under real-world load.
  • Integrate and tune LLMs, embedding models, and rerankers; evaluate trade-offs across hosted (Anthropic, OpenAI, Vertex) and self-hosted (HF, vLLM) options on cost, latency, and quality.
  • Build ingestion and chunking pipelines for unstructured data (PDFs, HTML, transcripts) and maintain vector store schemas for multi-tenant or multi-domain retrieval.
  • Implement evaluation harnesses to measure retrieval quality, generation faithfulness, and end-to-end answer correctness; close the loop from evals back into pipeline improvements.
  • Containerize and deploy ML workloads with Docker and Kubernetes; manage GPU/CPU resource allocation and model versioning.
  • Optimize database queries, vector search performance, and caching strategies (including LLM prompt caching) to reduce latency and cost.
  • Implement CI/CD pipelines for ML services and instrument monitoring for both system metrics (latency, error rate) and ML-specific metrics (retrieval quality, hallucination rate, drift)
  • Collaborate with frontend engineers, ML researchers, and product analysts to translate model capabilities into shipped features.
  • Document backend and ML infrastructure, including model cards, evaluation results, and architectural decisions
  • Travel - must be willing to travel 25% and periodically up to 50%.

Sterling Computers Corporation (“Sterling”) is an Equal Opportunity Employer. Qualified applicants will receive consideration for employment without regard to age, race, color, creed, religion, disability, medical condition, economic status or status with regard to public assistance, citizenship status, national or social or ethnic origin, past or present membership in the uniformed services, protected veteran status, sex, pregnancy, marital or civil union or domestic partnership status, family or parental status, sexual orientation, gender expression or identity, family medical history or genetic information, HIV status, political belief, or any other status or characteristic protected by applicable law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
937,957 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
United States
GenAI Engineer 14 days ago
≈ $71k – $134k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Paris
Python
JavaScript
TypeScript
Node JS
Python
Flask
pre-commit
Databases
PostgreSQL
Redis
pgvector
ElasticSearch
Meilisearch
AI/ML
Model Context Protocol
Prompt Engineering
AI Agents
LLM
RAG
OpenAI
GPT-4
GPT-5
LLM Guardrails
EU AI Act
Agentic Workflows
DevOps
GCP
Helm
Azure DevOps
GitHub Actions
Datadog
Azure
CI/CD
AWS
Docker
Kubernetes
Azure AKS
Cybersecurity
Okta
Snyk
SonarQube
GDPR
Management
Confluence
Jira
n8n
SharePoint
Apply
$170k – $235k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • New York
Python
TypeScript
AI/ML
Embeddings
AI Agents
LLM
Reranking
Machine Learning
Apply
$145k – $217k per year • In office • Full-Time • 12+ years exp • Master's Degree • Cambridge
AI/ML
Multimodal AI
TensorFlow
PyTorch
Machine Learning
Apply
≈ $126k – $270k per year (Estimated) • Equity • Remote (Poland) • Full-Time • Bachelor's Degree
Python
Go
JavaScript
TypeScript
C#
AI/ML
Copilot
AutoGen
LangChain
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
Gemini
LLM
RAG
OpenAI
Anthropic
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
Tool Use
Copilot Studio
DevOps
Rest API
Terraform
GCP
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Shift-Left
Bicep
GitHub
Cybersecurity
Burp Suite
Snyk
OWASP ZAP
SonarQube
Trivy
HashiCorp Vault
Checkmarx
Semgrep
ISO 27001
Checkov
CIS Benchmarks
OWASP Top 10
PCI DSS
SOC 2
HIPAA
Least Privilege
Shift-Left Security
Threat Modeling
SBOM
SLSA
Veracode
Sigstore
SIEM
OWASP
Cryptography
Vault
Management
Agile
Apply
$39k – $71k per year • In office • Tokyo
Apply
≈ $109k – $240k per year (Estimated) • Remote (likely United States, Canada) • 7+ years exp
Python
AI/ML
Airflow
OpenCV
MLFlow
Triton Inference Server
Quantization
Knowledge Distillation
Diffusion Models
Computer Vision
TensorRT
Kubeflow
PyTorch
Pillow
scikit-image
TorchServe
Edge AI
ONNX Runtime
Model Distillation
DevOps
Rest API
GCP
Prometheus
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Cybersecurity
SOC 2
Apply
In office • 5+ years exp • Bachelor's Degree
Python
JavaScript
TypeScript
SQL
Node JS
AI/ML
TensorFlow
PyTorch
Edge AI
Frontend
React.js
DevOps
Rest API
gRPC
Terraform
Ansible
CloudFormation
Datadog
Prometheus
WebSockets
CI/CD
AWS
Docker
Kubernetes
Grafana
Amazon EKS
AWS Lambda
Amazon EC2
Incident Management
Amazon S3
Amazon CloudWatch
Cybersecurity
SOC 2
Apply
$75k – $90k per year • Hybrid • Full-Time • 3+ years exp • Kraków
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
Trino
BigQuery
AI/ML
Dagster
DevOps
AWS
Amazon S3
Cybersecurity
GDPR
Apply
$68k – $88k per year • Remote (likely Poland) • Warsaw
Python
SQL
Databases
PostgreSQL
PostGIS
Google BigQuery
BigQuery
DevOps
Terraform
GCP
CI/CD
Kubernetes
Apply
≈ $48k – $83k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Kraków
Python
Java
C++
DevOps
CI/CD
Jenkins
Git
Gerrit
Linux
Apply
≈ $126k – $262k per year (Estimated) • In office • TS/SCI • Full-Time • 10+ years exp • United States
Cybersecurity
ISO 27001
NIST 800-171
Management
ITSM
Microsoft Office
Apply
Graphic Designer 5 days ago
≈ $63k – $127k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
PHP
PHP
WordPress
Apply
Logistics Manager 5 days ago
≈ $89k – $175k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • United States
Apply
≈ $64k – $134k per year (Estimated) • In office • Full-Time • 3+ years exp • Associate's Degree • United States
Analytics
Microsoft Excel
Management
Outlook
Apply
≈ $75k – $189k per year (Estimated) • In office • Full-Time • Bachelor's Degree • United States
Analytics
Microsoft Excel
Management
Outlook
Marketing
Salesforce
Apply
$36k – $54k per year • Equity • Remote (United States) • Part-Time • 5+ years exp • Master's Degree • United States
Apply
UX Designer 4 hours ago
≈ $84k – $196k per year (Estimated) • In office • Full-Time • United States
Design
Figma
Apply
ROI Account Manager 4 hours ago
$36k – $40k per year • In office • High School Diploma • United States
Cybersecurity
HIPAA
Apply
Lead UX Designer 4 hours ago
≈ $133k – $244k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • United States
Design
Figma
Apply
≈ $133k – $244k per year (Estimated) • In office • 10+ years exp • United States
AI/ML
Human-in-the-Loop
Design
Figma
Apply
See all jobs
This is one of many
937,957 more open roles from verified company boards, updated every day.