824,589open jobs
53,146companies
135,057added this week
Browse all
Location
In office
Seniority
Senior · 7+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Jun 5, 2026.

Overview
Company
Impact
Profile match
Conquer IT complexity with proven expertise and premium solutions from BlueAlly.

We are hiring a Senior AI Engineer to design, build, and operate enterprise AI systems across our client portfolio. You will work end-to-end across the AI stack - from inference engines and platform infrastructure (vLLM, KV cache, Dynamo-style serving, GPU-accelerated AI Factory platforms) up through application-level engineering (RAG pipelines, agent workflows, prompt engineering, evaluation methodology).

This role is for an engineer who can lead workstreams independently, mentor more junior engineers, and serve as the technical authority that clients trust to deliver production AI outcomes. You'll engage directly with client architects, data scientists, application teams, and executives - and you'll leave each engagement having raised both the client's capability and BlueAlly's practice.

Key Responsibilities

  • Lead end-to-end design, build, and operation of AI systems on AI Factory platforms (HPE PCAI, Dell AI Factory, Nutanix Enterprise AI, and adjacent ecosystem layers) across multiple client engagements.
  • Engineer and tune LLM inference serving stacks - primary depth in vLLM with breadth across the inference ecosystem - for client latency, throughput, and cost targets.
  • Tune inference performance through KV cache management, paged attention, batching strategies, and Dynamo-based disaggregated serving.
  • Architect and operate MLOps pipelines covering model lifecycle, registries, deployment, rollback, and observability.
  • Design and engineer RAG applications on top of vector databases - chunking strategies, retrieval tuning, reranking, citation handling, and context-window management.
  • Build and tune prompt-engineering patterns at production scale - system prompts, structured output, tool and function calling.
  • Design and maintain LLM evaluation harnesses - golden sets, regression suites, and online quality metrics.
  • Engineer high-performance storage and networking for AI workloads - parallel filesystems, object storage tiers, and high-throughput, low-latency RDMA fabrics.
  • Operate Kubernetes clusters underpinning AI workloads - namespaces, RBAC, resource quotas, network policies, storage classes, and ingress.
  • Build and maintain container images, registries, and CI/CD pipelines for AI/ML services.
  • Implement monitoring, alerting, logging, and capacity planning across the AI stack.
  • Harden environments to meet client security and compliance requirements.
  • Lead troubleshooting across bare metal, BIOS/firmware, OS, containers, GPUs, frameworks, and models.
  • Engage directly with client stakeholders - technical and executive - to communicate status, root cause, options, and recommendations.
  • Mentor and code-review work from less senior engineers; raise the technical bar of every engagement you join.
  • Author runbooks, reference architectures, and knowledge base content; lead client knowledge transfer and enablement sessions.
  • Participate in on-call rotation and incident response for production AI workloads.
  • Contribute reusable patterns, tooling, and reference designs back to the practice.

Required Qualifications

  • Experience: 7+ years of software, data, or infrastructure engineering, with 3+ years specifically working with modern AI / LLM systems.
  • Software engineering: Production-quality Python at engineering level - testing, code review, version control fluency, and shipping code that other engineers depend on.
  • Linux engineering: Deep production Linux experience, including system internals, performance tuning, and troubleshooting.
  • Containers: Deep proficiency with Docker - image build, registry management, runtime tuning, and container security.
  • Hardware fundamentals: Strong server-platform skills including CPU/GPU topologies, PCIe, BMC management, BIOS/firmware lifecycle, and physical-to-logical troubleshooting.
  • AI Factory platforms: Hands-on experience deploying and operating one or more of HPE PCAI, Dell AI Factory, or Nutanix Enterprise AI.
  • Inference stack - vLLM: Production experience deploying, tuning, and operating vLLM.
  • Inference stack breadth: Working knowledge of multiple inference and model-serving frameworks beyond vLLM, with the ability to choose and tune the right tool for each workload.
  • High-performance storage and networking: Hands-on experience with high-throughput, low-latency storage and network fabrics for AI workloads - including RDMA-class interconnects, parallel/object storage tiers, KV cache management, and Dynamo-style disaggregated serving.
  • MLOps: Practical experience operating MLOps tooling and patterns - model registries, deployment pipelines, GitOps, lineage, and rollback.
  • Vector databases and RAG: Hands-on experience deploying, tuning, and integrating vector databases and RAG pipelines, including the application-level engineering that sits on top of them.
  • Prompt engineering and tool use: Production experience designing system prompts, structured output, function calling, and tool-using LLM patterns.
  • Evaluation methodology: Demonstrated experience designing LLM evaluation harnesses - golden sets, regression suites, and quality/cost metrics.
  • Client-facing skills: Demonstrated ability to engage directly with client stakeholders - running working sessions, presenting recommendations, and translating technical detail for non-technical audiences.
  • Communication: Strong written and verbal communication - clear reference architectures, runbooks, and incident reports.
  • Mentorship: Track record of mentoring more junior engineers and raising team technical quality through code review and pairing.
  • Networking fundamentals: TCP/IP, DNS, load balancing, VLANs, and firewall administration.
  • Multi-client delivery: Comfort working across multiple concurrent client environments and managing competing priorities under SLA.

Preferred Qualifications

  • GPU operations: Experience with GPU drivers, CUDA toolchains, GPU partitioning (MIG/vGPU), and GPU-level monitoring.
  • NVIDIA AI Enterprise: Deployment and operations experience with the NVAIE software stack.
  • Ray: Familiarity with Ray for distributed training and inference scaling.
  • Kubernetes: Working knowledge of Kubernetes administration - Helm, ingress, RBAC, storage classes.
  • Identity and access: Integrating SSO and enterprise identity (LDAP, AD, OIDC/SAML), secrets management, tenant isolation.
  • Fine-tuning: Familiarity with LoRA/QLoRA/PEFT and supervised fine-tuning workflows.
  • Token economics: Experience optimizing inference cost - caching, prompt caching, model routing, and distillation.
  • MSP / multi-tenant operations: Service-provider experience including chargeback/showback and tenant isolation patterns.
  • Compliance frameworks: SOC 2, HIPAA, FedRAMP, FISMA, or CMMC environments.
  • Public cloud and hybrid: Working experience with one or more public clouds and hybrid architectures.
  • Infrastructure as Code: Terraform, Ansible, Helm, or similar.

Certifications (Preferred)

  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
  • Cloud certifications - AWS, Azure, or Google Cloud.
  • Linux certifications - RHCE, RHCSA, or LFCS.
  • NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO) or higher NVIDIA certifications.
  • HPE, Dell Technologies, or Nutanix platform certifications.

What Sets You Apart

  • Genuine curiosity about how AI systems work end-to-end - from kernel and GPU up through frameworks and models.
  • Track record of restoring production AI services under pressure.
  • Ability to translate complex technical concepts into clear, client-facing communication.
  • Comfort with ambiguity and rapid change in the AI/LLM ecosystem.
  • Service-oriented mindset: you treat each client environment as if it were your own.
  • Bias toward leaving the practice better than you found it - patterns, tooling, and reference designs.

About BlueAlly

BlueAlly is a leading provider of IT services and solutions, helping organizations conquer IT complexity across cloud, cybersecurity, infrastructure, data, and application modernization. Headquartered in Cary, North Carolina, with delivery teams across the United States and globally, BlueAlly serves clients ranging from mid-market enterprises to large public-sector and commercial organizations.

Founded in 2011, BlueAlly delivers across the full technology lifecycle - from strategy and design through implementation, managed services, and continuous optimization. The company is recognized on CRN's Tech Elite 150 and MSP 500 lists and partners deeply with leading technology vendors. As enterprise AI moves from pilot to production, BlueAlly is investing in the people, platforms, and practices required to deliver AI Factory outcomes for our clients - and this role is at the center of that investment.

Equal Employment Opportunity

BlueAlly is an Equal Opportunity Employer. We are committed to building a diverse and inclusive workforce and to making employment decisions based on merit, qualifications, and business need. BlueAlly does not discriminate in employment on the basis of race, color, religion, sex (including pregnancy), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, marital status, veteran status, or any other protected characteristic under applicable federal, state, or local law.

BlueAlly provides reasonable accommodations to qualified applicants and employees with disabilities. If you require an accommodation to participate in the application or interview process, please contact our People team.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
≈ $140k – $245k per year (Estimated) • Remote (United States) • Public Trust • Full-Time • Bachelor's Degree • Salt Lake City
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Microsoft Fabric
AI/ML
Spark
MLFlow
XGBoost
Scikit-learn
PyTorch
RAG
Machine Learning
DevOps
Terraform
Azure
CI/CD
Git
Cybersecurity
HIPAA
FedRAMP
Microsoft Entra ID
Analytics
Power BI
ETL/ELT
SSIS
Azure Data Factory
SSAS
Apply
$120k – $170k per year • In office • United States
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
≈ $145k – $263k per year (Estimated) • In office • Austin
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
$120k – $170k per year • In office • Chicago
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
$120k – $170k per year • In office • Los Angeles
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
≈ $109k – $239k per year (Estimated) • In office • Internship • New York
Python
Go
JavaScript
TypeScript
SQL
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
Google BigQuery
BigQuery
AI/ML
Prompt Engineering
AI Agents
LangSmith
LLM
OpenAI
Anthropic
LLM Evaluation
World Models
DevOps
GCP
Apply
$102k – $120k per year • In office • Full-Time • 6+ years exp • Waltham
Python
C#
DevOps
GitLab CI
CI/CD
Jenkins
Linux
Robotics
Teleoperation
Management
Jira
Apply
≈ $65k – $127k per year (Estimated) • In office • Full-Time • 3+ years exp • Waltham
Python
C#
AI/ML
VLM
LLM
DevOps
WebSockets
GitLab CI
CI/CD
Jenkins
Linux
Game Dev
Unity
Robotics
Teleoperation
Management
Jira
Apply
Founding Engineer 1 day ago
≈ $147k – $302k per year (Estimated) • Equity 0.5–3% • In office • Full-Time • San Francisco
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Apply
$150k – $250k per year • Equity 0.2–1.5% • In office • Full-Time • 3+ years exp • New York
Go
TypeScript
SQL
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
AI/ML
Function Calling
AI Agents
LangSmith
LLM
OpenAI
Tool Use
World Models
Frontend
Tailwind CSS
Next.js
React.js
DevOps
GCP
Vercel
Datadog
QA
Playwright
Jest
Sentry
Vitest
Apply
$80k – $95k per year • Hybrid • Full-Time • Bachelor's Degree
Apply
Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Cary
Apply
$55k – $65k per year • In office • Full-Time • 1+ year exp • High School Diploma
DevOps
SLI/SLO/SLA
Windows
DNS
DHCP
VPN
Wi-Fi
Cybersecurity
Active Directory
Management
ITIL
ITSM
Service Desk
Apply
$120k – $140k per year • Remote (United States, CT hours) • Full-Time
PowerShell
AI/ML
Copilot
DevOps
VMWare
Azure
Hyper-V
Windows
VPN
Cybersecurity
Microsoft Sentinel
Microsoft Defender
SOC 2
HIPAA
Microsoft Defender for Cloud
Microsoft Entra ID
Management
Microsoft Teams
OneDrive
SharePoint
Apply
$140k – $165k per year • In office • Full-Time • 6+ years exp • Washington
DevOps
DNS
DHCP
VPN
BGP
Cybersecurity
Wireshark
Apply
See all jobs
This is one of many
824,589 more open roles from verified company boards, updated every day.