847,566open jobs
53,927companies
143,960added this week
Browse all
Salary
$144k – $273k per year
Location
Remote (United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Sep 16, 2026. Hewlett Packard Enterprise scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Hewlett Packard Enterprise is a global technology company specializing in edge-to-cloud solutions, enterprise IT infrastructure, and intelligent software services. Formed in 2015 following the division of Hewlett-Packard Company, the organization offers servers, storage, high-performance computing, and networking capabilities through flexible consumption models like its GreenLake platform. Headquartered in Spring, Texas, the enterprise operates internationally to help commercial and public-sector clients secure, manage, and modernize their digital infrastructure.
Senior Software Engineer, InferenceThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:

HPE's Private Cloud AI organization is seeking a Senior Software Engineerto build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime - engine integration, batching, KV cache management, and distributed execution - together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.

Responsibilities

·Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution

·Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency

·Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage

·Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption

·Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling

·Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence

·Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Knowledge and Skills

Required

·Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals

·Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding

·Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them

·Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling

·Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight

·Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents

·Excellent analytical, debugging, and problem-solving abilities

Preferred

·Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe

·Disaggregated prefill/decode serving, or KV cache offload and reuse at scale

·RDMA, GPUDirect Storage, InfiniBand, or RoCE

·MIG, fractional GPU allocation, and multi-tenant GPU isolation

·On-premises, air-gapped, or regulated enterprise software delivery

Experience and Education

·Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving

·Degree in Computer Science or related field

Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here.

Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#unitedstates

Job:

Engineering

Job Level:

TCP_04"The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.

- United States of America: Annual Salary USD 144,000 - 273,000 in Colorado // 137,000 - 315,000 in North Carolina & Texas

The listed salary range reflects base salary. Variable incentives may also be offered."

Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html

The estimated job application period closure is December 30 2027; this timeline is provided for transparency and internal planning purposes.

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBTemployer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

Recruitment Fraud Alert

We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.

All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual’s own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
847,566 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Durham
≈ $98k – $191k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Atlanta
Python
C#
Databases
MS SQL
DevOps
Azure
CI/CD
Management
Agile
Apply
$110k – $140k per year • Remote (United States) • Full-Time • 10+ years exp • Bachelor's Degree • New York
JavaScript
TypeScript
C#
C#
ASP.NET Core
Blazor
WPF
Databases
Oracle
MS SQL
Apache Kafka
AI/ML
LangChain
Claude
Fine-tuning
Scikit-learn
ML.NET
Prompt Engineering
AI Agents
Semantic Kernel
TensorFlow
PyTorch
RAG
OpenAI
Anthropic
Edge AI
Machine Learning
Frontend
Angular
Mobile
MAUI
DevOps
Rest API
Azure
CI/CD
Git
Windows
Cybersecurity
Microsoft Entra ID
Management
Agile
Scrum
Apply
≈ $115k – $212k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Columbus
Java
SQL
Java
Maven
Gradle
DevOps
Rest API
Terraform
GCP
Helm
GitHub Actions
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Management
Agile
Apply
≈ $112k – $207k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Columbus
JavaScript
Java
TypeScript
SQL
Java
Maven
Gradle
Frontend
Angular
DevOps
Rest API
Terraform
GCP
Helm
GitHub Actions
OpenTelemetry
Azure
CI/CD
AWS
Kubernetes
Grafana
Management
Agile
Apply
$150k – $262k per year • Equity • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • San Diego
SQL
Databases
PostgreSQL
Redis
NATS
RabbitMQ
Apache Kafka
Amazon Aurora
AI/ML
AI Agents
DevOps
Terraform
GCP
Crossplane
etcd
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Amazon EKS
Google GKE
Azure AKS
IAM
Management
ServiceNow
Apply
≈ $45k – $111k per year (Estimated) • In office • Ho Chi Minh City
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Triton Inference Server
Quantization
SGLang
TensorRT
TensorRT-LLM
LLM
CUDA
Triton
NCCL
cuDNN
CUTLASS
Speculative Decoding
KV Cache
DevOps
OpenTelemetry
Prometheus
Grafana
Platform Engineering
Linux
Apply
≈ $44k – $110k per year (Estimated) • In office • Ho Chi Minh City
Python
Go
Bash
Databases
ClickHouse
ElasticSearch
Apache Kafka
OpenSearch
AI/ML
llama.cpp
Ray Serve
vLLM
Quantization
AI Agents
SGLang
TensorRT
TensorRT-LLM
LLM
RAG
Ray
KServe
Triton
LLMOps
KV Cache
DevOps
Terraform
GCP
Helm
Loki
OpenTelemetry
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
SLI/SLO/SLA
Linux
Apply
≈ $43k – $108k per year (Estimated) • In office • 4+ years exp • Ho Chi Minh City
Python
SQL
AI/ML
Spark
XGBoost
Scikit-learn
TensorFlow
PyTorch
Machine Learning
DevOps
CI/CD
Apply
In office • Ho Chi Minh City
Python
Bash
Databases
MySQL
Redis
Apache Kafka
AI/ML
Cursor
LangGraph
LangChain
Claude
Claude Code
Langfuse
LiteLLM
MiniMax
OpenAI
Agentic Workflows
DevOps
Terraform
Ansible
GCP
OpenTelemetry
Prometheus
GitLab CI
CI/CD
AWS
Docker
Kubernetes
Nginx
Grafana
Linux
TCP/IP
DNS
Cybersecurity
ISO 27001
Apply
≈ $44k – $118k per year (Estimated) • Hybrid • Ho Chi Minh City
Python
AI/ML
Fine-tuning
Embeddings
Multimodal AI
AI Agents
LLM
RAG
OCR
Human-in-the-Loop
LLM Guardrails
Recommender Systems
Machine Learning
Analytics
A/B Testing
Apply
$127k – $241k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Durham • Fort Collins • Colorado Springs • Houston
Python
C++
Apply
$137k – $277k per year • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Roseville
Python
DevOps
Linux
Apply
$172k – $349k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Cupertino
Python
DevOps
Debian
CI/CD
Linux
Apply
$121k – $243k per year • In office • Full-Time • 5+ years exp • PhD • Sunnyvale
Python
C++
AI/ML
Copilot
DevOps
Linux
Unix
Apply
$137k – $277k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • San Jose
Python
C++
DevOps
Linux
TCP/IP
VPN
BGP
OSPF
Apply
≈ $43k – $101k per year (Estimated) • In office • Part-Time • Durham
Apply
$32k – $34k per year • In office • Part-Time • Durham
Apply
Registrar 1 day ago
≈ $28k – $81k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Durham
Apply
≈ $22k – $51k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Durham
Management
Microsoft Office
Apply
≈ $41k – $72k per year (Estimated) • In office • 1+ year exp • Associate's Degree • Durham
Apply
See all jobs
This is one of many
847,566 more open roles from verified company boards, updated every day.