{"id":1290038,"url":"https://alion.io/job/sterling-backend-ml-engineer","title":"Backend ML Engineer","company":{"id":1927540,"name":"Sterling","domain":"sterling.com","url":"https://alion.io/company/sterling-com","size_band":"1-10","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"ADP","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"junior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":97000,"max_usd":223000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":238},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Anthropic","optional":false},{"name":"FastAPI","optional":false},{"name":"Gemini","optional":false},{"name":"LLM","optional":false},{"name":"Machine Learning","optional":false},{"name":"NLP","optional":false},{"name":"OpenAI","optional":false},{"name":"OpenCV","optional":false},{"name":"Pinecone","optional":false},{"name":"Python","optional":false},{"name":"RAG","optional":false},{"name":"Sentence-Transformers","optional":false},{"name":"Weaviate","optional":false},{"name":"AWS","optional":true},{"name":"Azure","optional":true},{"name":"CI/CD","optional":true},{"name":"Docker","optional":true},{"name":"Docling","optional":true},{"name":"Fine-tuning","optional":true},{"name":"Flask","optional":true},{"name":"GCP","optional":true},{"name":"Hallucination","optional":true},{"name":"Haystack","optional":true},{"name":"Kubeflow","optional":true},{"name":"Kubernetes","optional":true},{"name":"LangChain","optional":true},{"name":"LlamaIndex","optional":true},{"name":"LLM Evaluation","optional":true},{"name":"LLM Guardrails","optional":true},{"name":"LoRA","optional":true},{"name":"MariaDB","optional":true},{"name":"MLFlow","optional":true},{"name":"Model Distillation","optional":true},{"name":"PEFT","optional":true},{"name":"pgvector","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Prompt Caching","optional":true},{"name":"Prompt Engineering","optional":true},{"name":"Qdrant","optional":true},{"name":"Ragas","optional":true},{"name":"Reranking","optional":true},{"name":"Transformers","optional":true},{"name":"TruLens","optional":true},{"name":"vLLM","optional":true},{"name":"Weights & Biases","optional":true}],"status":"live","first_seen_at":"2026-09-24T20:01:00Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-29T14:18:36Z","board_verified":true,"closed_at":null,"days_open":6,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":6},"description":"Title: Backend ML Engineer\nReports to: Senior Software Architect\nLocation: North Sioux City, SD\nJob Description: Sterling Computers is a technology company that provides IT solutions to a variety of clients, including the federal government, state and local governments, education, and commercial entities. Sterling's Strategic Technologies Group is responsible for learning and becoming subject matter experts in new and emerging technologies. Our team uses this expertise to broaden the portfolio of products and solutions that the company sells, delivers, and manages. Our engineers work on a range of AI-integrated systems, from production RAG platforms and LLM orchestration layers to digital human solutions and intelligent automation pipelines. We are looking for a Backend ML Engineer who is interested in taking AI/ML systems from prototype to production, designing inference APIs, building retrieval and orchestration pipelines, integrating large language models, and operating ML infrastructure at scale. If you thrive in a collaborative, client-focused environment and enjoy shipping AI features that real users depend on, we'd love to have you on our team.\nRequired Technical Skills:\n0-3 years of experience in backend or ML engineering\nStrong working knowledge of Python, including FastAPI\nExperience with python libraries such as sentence-transformers, OpenCV, pillow, and other NLP or CV libraries \nStrong understanding of scalability and latency in ML systems \nExperience with multi step agentic systems and tooling structures\nHands-on experience integrating LLMs (OpenAI, Anthropic, Gemini, or open-source models) into production systems\nFamiliarity with vector databases such as Weaviate, Pinecone, or similar\nExperience with different retrieval-augmented generation (RAG) architectures\nSelf-motivated with a positive and professional attitude\nAbility to adapt to other languages across the stack as needed.\nRequired Education/Experience:\nBachelor’s degree in Computer Science, Machine Learning, or a related field (minimum requirement), or equivalent practical experience\nGraduate-level coursework or specialization in ML/AI is a plus\nRelevant cloud certifications are a plus\nDemonstrated experience shipping ML systems to production is a plus\nUS DoD Clearance preferred or willingness to obtain such\nQualifications:\nStrong experience building backend services with Python (FastAPI/Flask); comfort working with async APIs and request/response patterns for ML inference workloads.\nHands-on experience integrating LLMs and embedding models into production applications, including prompt engineering, context management, and handling rate limits, retries, and streaming responses.\nFamiliarity with RAG architectures: chunking strategies, embedding pipelines, vector search, reranking, and evaluation metrics (Recall@k, MRR, faithfulness, answer relevance).\nExperience with vector databases (Weaviate, pgvector, Pinecone, Qdrant, or similar) and traditional databases (PostgreSQL, MariaDB) for hybrid retrieval and metadata filtering.\nCloud experience (AWS/GCP/Azure) for deploying ML services - including managed inference endpoints, GPU instances, or serverless model hosting.\nStrong understanding of API authentication, secure handling of model inputs/outputs, and PII/PHI-aware design where applicable.\nExperience with ML observability: tracking latency, token usage, cost-per-query, retrieval quality, and model drift in production.\nBackground in data pipelines, document ingestion/parsing, or evaluation frameworks (Ragas, TruLens, Docling, custom harnesses) is needed.\nFamiliarity with fine-tuning, LoRA/PEFT, or model distillation is appreciated.\nExperience with MLOps tooling (MLflow, Weights & Biases, Kubeflow) or LLM orchestration frameworks (LangChain, LlamaIndex, Haystack, or custom orchestrators) is a plus.\nResponsibilities:\nBuild, test, and maintain production ML services - inference APIs, retrieval pipelines, orchestration layers, and guardrail/evaluation components.\nDesign scalable RESTful and streaming APIs that serve ML model outputs reliably under real-world load.\nIntegrate and tune LLMs, embedding models, and rerankers; evaluate trade-offs across hosted (Anthropic, OpenAI, Vertex) and self-hosted (HF, vLLM) options on cost, latency, and quality.\nBuild ingestion and chunking pipelines for unstructured data (PDFs, HTML, transcripts) and maintain vector store schemas for multi-tenant or multi-domain retrieval.\nImplement evaluation harnesses to measure retrieval quality, generation faithfulness, and end-to-end answer correctness; close the loop from evals back into pipeline improvements.\nContainerize and deploy ML workloads with Docker and Kubernetes; manage GPU/CPU resource allocation and model versioning.\nOptimize database queries, vector search performance, and caching strategies (including LLM prompt caching) to reduce latency and cost.\nImplement CI/CD pipelines for ML services and instrument monitoring for both system metrics (latency, error rate) and ML-specific metrics (retrieval quality, hallucination rate, drift)\nCollaborate with frontend engineers, ML researchers, and product analysts to translate model capabilities into shipped features.\nDocument backend and ML infrastructure, including model cards, evaluation results, and architectural decisions\nTravel - must be willing to travel 25% and periodically up to 50%.\nSterling Computers Corporation (“Sterling”) is an Equal Opportunity Employer. Qualified applicants will receive consideration for employment without regard to age, race, color, creed, religion, disability, medical condition, economic status or status with regard to public assistance, citizenship status, national or social or ethnic origin, past or present membership in the uniformed services, protected veteran status, sex, pregnancy, marital or civil union or domestic partnership status, family or parental status, sexual orientation, gender expression or identity, family medical history or genetic information, HIV status, political belief, or any other status or characteristic protected by applicable law.","description_format":"text","description_chars":6103,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity","Government","Information Technology"],"lifecycle":[{"event":"open","at":"2026-09-26T07:43:19Z"}],"liveness":{"score":84,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.844,"p_room":1,"age_days":6,"expected_fill_days":23,"reasons":["conf:39","velocity","win:early","comp:junior"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/sterling-backend-ml-engineer","json_url":"https://alion.io/job/sterling-backend-ml-engineer.json","meta":{"generated_at":"2026-10-01T09:32:15Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":169,"day_limit":5000,"remaining_today":4831,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}