1,385,600open jobs
80,497companies
205,159added this week
Browse all
Salary
$143k – $275k per year
Location
In office (Washington, Mountain View, Hillsboro)
Seniority
Principal · 12+ years exp

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Sep 16, 2026. Microsoft scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Microsoft is an American multinational technology corporation founded in 1975 by Bill Gates and Paul Allen and headquartered in Redmond, Washington. It built the personal computing era around the Windows operating system and the Office productivity suite, and now generates the largest share of its revenue from Azure cloud infrastructure and commercial subscriptions. The company also owns GitHub, LinkedIn and the Xbox gaming business, and has invested heavily in artificial intelligence through its partnership with OpenAI and the Copilot assistants embedded across its products.
Overview

Do you want to be at the forefront of innovating the latest hardware and systems designs to propel Microsoft’s cloud growth? Are you seeking a unique career opportunity that combines technical capabilities, cross team collaboration, with business insight and strategy?

The SPARC organization is responsible for strategy, planning and architecture pathfinding, and manages Azure’s hardware roadmap from architecture concept, through production for Microsoft’s current and future offerings. The CSA team within SPARC is at the forefront of systems architecture and technology pathfinding spanning compute, memory, storage, and system interconnects. Drawing on deep insights of workloads, emerging technology trends, and focused industry engagements, CSA team’s charter is to define and evaluate novel systems architecture innovations through hardware/software co-design and advance them through technical readiness for productization.

Join our Compute System Architecture (CSA) team within the System Planning and Architecture (SPARC) organization in Azure Hardware Systems & Infrastructure (AHSI). AHSI is the team behind Microsoft’s expanding cloud business, responsible for delivering the hardware systems and infrastructure for cloud computing across Microsoft Azure, Bing, MSN, Office 365, OneDrive, Skype, Teams and Xbox Live.

The CSA team is seeking a Principal AI Software Engineer !

Responsibilities

  • Lead full system software prototyping to develop capable proof-of-concepts to evaluate hardware/software co-designed capabilities for memory TCO reduction such as through memory tiering/pooling and overcommit solutions for Azure usages and deployment scenarios.
  • Develop deep insights through workload characterization and correlation to identify systems optimization opportunities.
  • Collaborate with diverse workload experts across Microsoft and partner ISVs to engineer TCO-optimized solutions for Azure general-purpose and specialized compute fleet.
  • Influence and shapehardware architecture and industry alignment, targeting three-to-six-year timeframe, with data-driven analysis, insights and recommendations.
  • Lead characterization and optimization of Large Language Model (LLM) inference workloads with focus on KV Cache capacity, placement, migration, and utilization across GPU HBM, host DRAM, CXL memory expansion/pooling, SSD, and emerging memory tiers.
  • Develop proof-of-concepts and evaluation frameworks to assess memory-tiering architectures for AI inference, including CXL pooled memory, memory expansion solutions, context-memory platforms, SSD-backed cache tiers, and hardware/software co-designed approaches for reducing inference TCO.
  • Design and execute workload characterization studies for agentic, multi-turn, coding, reasoning, and long-context AI workloads to quantify memory consumption, latency, throughput, token efficiency, and system utilization.
  • Analyze end-to-end data movement across GPU, CPU, storage, and networking subsystems, identifying optimization opportunities within GPU Direct Storage (GDS), GPU Direct RDMA (GDR), peer-to-peer memory transfers, and distributed inference pipelines.
  • Develop software prototypes, framework extensions, and instrumentation to evaluate KV Cache offload, prefetching, migration, compression, deduplication, and memory-overcommit techniques.
  • Build performance models and simulation frameworks to predict the impact of memory hierarchy innovations on large-scale inference deployments.

Qualifications

Required/minimum qualifications

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Other Qualifications:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

  • 12+years of experiencein systems software (OS kernel, memory management, I/O stacks, Virtualization) with demonstrated track record ofsuccess, guiding architecture and software enabling.
  • 10+ years of experience leading significant hardware/software co-design projects involving CPU and/or systems architecture and influencing technical direction.
  • Deep expertise in Linux kernel internals, memory management, I/O subsystems, NUMA, DMA, and GPU/CPU/storage/network data paths.
  • Hands-on experience with NVIDIA GPU software stacks including CUDA, NCCL, GPUDirect Storage (GDS), and GPUDirect RDMA (GDR).
  • Understanding of AI inference infrastructure, large GPU clusters, inference serving architectures, and workload performance optimization.
  • Experience characterizing and optimizing KV Cache intensive workloads including long-context, agentic, multi-turn, coding, and reasoning workloads.
  • Experience with inference frameworks such as vLLM, SGLang, and TensorRT-LLM.
  • Familiarity with KV Cache technologies including LMCache, SGLang HiCache, cache offload, cache sharing, and memory tiering approaches.
  • Experience designing or extending inference runtimes, scheduling systems, memory management components, or KV Cache subsystems.
  • Understanding of disaggregated prefill/decode architectures, distributed inference serving, and multi-node cache-sharing topologies.
  • Experience with CXL memory expansion, memory pooling, and memory tiering solutions in large-scale deployments.
  • Software development skills in C/C++, Python, CUDA, and distributed systems.
  • Skilled in partnering and influencing architects, hardware engineers, and software leads
  • Ability to manage through ambiguity, bringing clarity and results orientation to engage and energize collaborators and stakeholders
  • Collaboration skills, teamwork, and sense of presumed responsibility
  • Verbal and written communication skills, and ability to articulate and engage with both technical and non-technical stakeholders at all levels.
  • Experience leading and driving complex projectswith respect and integrity, including thosewith multiple workstreams spanning different business and technical disciplines.
  • Intellectual curiosity and passion about learning and deploying new technologies.
  • Problem-solvingskills, analytical capabilities, and attention to detail

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:

https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,385,600 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Washington
≈ $182k – $371k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
AI/ML
Fine-tuning
Multimodal AI
Computer Vision
Gemini
Machine Learning
Apply
≈ $171k – $349k per year (Estimated) • In office • TS/SCI • Full-Time • 10+ years exp • Bachelor's Degree • Annapolis Junction
Python
Java
AI/ML
TensorFlow
PyTorch
Machine Learning
Apply
$296k – $431k per year • Equity • Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • San Jose
AI/ML
Google AI Studio
Apply
Chief Data/AI Engineer 2 months ago
$97k – $181k per year • In office • Secret • Full-Time • Bachelor's Degree • Arlington
Management
Microsoft Office
Apply
$156k – $198k per year • Equity • Hybrid • Full-Time • 5+ years exp • Chicago • Austin
AI/ML
AI Agents
Apply
≈ $21k – $55k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Java
TypeScript
Node JS
Python
FastAPI
Java
Spring Boot
Node JS
Express
Databases
PostgreSQL
Redis
Weaviate
Milvus
Pinecone
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
Model Context Protocol
Fine-tuning
Embeddings
Prompt Engineering
Multimodal AI
Function Calling
AI Agents
Semantic Kernel
Llama
Mistral
CrewAI
Gemini
LLM
RAG
Google ADK
Synthetic Data
OpenAI
Anthropic
Structured Outputs
Knowledge Graph
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Frontend
GraphQL
Tailwind CSS
Next.js
Angular
React.js
DevOps
Rest API
Terraform
GCP
GitHub Actions
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
≈ $49k – $119k per year (Estimated) • In office • Full-Time • 10+ years exp • Bengaluru
Python
JavaScript
AI/ML
Copilot
Cursor
LangGraph
AutoGen
LangChain
Claude Code
Model Context Protocol
Prompt Engineering
AI Agents
CrewAI
RAG
Frontend
Vue.js
Next.js
React.js
Apply
$165k – $220k per year • Remote (United States) • Top Secret • Full-Time • 3+ years exp • United States
Python
JavaScript
TypeScript
SQL
Bash
AI/ML
Multimodal AI
AI Agents
DevOps
OpenShift
Kubernetes
Linux
Apply
$127k – $175k per year • Equity • In office • Full-Time • 4+ years exp • Cambridge
Python
Databases
Databricks
AI/ML
Model Context Protocol
MLFlow
AI Agents
AWS Bedrock
RAG
Amazon SageMaker
Agentic Workflows
DevOps
Terraform
CI/CD
AWS
Kubernetes
Platform Engineering
Amazon S3
Cybersecurity
Least Privilege
Apply
Data Scientist 1 day ago
$86k – $148k per year • In office • 4+ years exp • Bachelor's Degree • Raleigh
Python
SQL
Databases
Snowflake
Databricks
AI/ML
MLFlow
XGBoost
Vertex AI
dbt
Prefect
Scikit-learn
TensorFlow
NumPy
PyTorch
LLM
RAG
Semantic Search
Anomaly Detection
Time Series Forecasting
Amazon SageMaker
Structured Outputs
Semantic Search
Interpretability
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Robotics
Digital Twin
Apply
$120k – $235k per year • In office • 4+ years exp • Bachelor's Degree • Cambridge
AI/ML
PyTorch
Post-training
Model Distillation
Machine Learning
Apply
$166k – $296k per year • In office • 9+ years exp • Bachelor's Degree • Mountain View
AI/ML
LLM
Mixture of Experts
KV Cache
Apply
$86k – $170k per year • In office • 5+ years exp • Bachelor's Degree • Washington • Mountain View
AI/ML
Copilot
FastAI
PyTorch
DevOps
Azure
GitHub
Management
Slack
Microsoft Teams
Agile
Apply
$120k – $235k per year • In office • 8+ years exp • Bachelor's Degree • Washington • Reston
Python
Go
TypeScript
C#
AI/ML
Reinforcement Learning
Synthetic Data
DevOps
Terraform
GCP
CloudFormation
Azure
CI/CD
AWS
Kubernetes
Bicep
Linux
Windows
DNS
Cybersecurity
Okta
Microsoft Entra ID
Active Directory
Apply
$120k – $235k per year • In office • 8+ years exp • Bachelor's Degree • Mountain View • Boston • Washington
Python
C++
AI/ML
CUDA Toolkit
Mixture of Experts
CUDA
Triton
ROCm
Speculative Decoding
KV Cache
DevOps
Azure
Management
OneDrive
Apply
$58k – $86k per year • In office • Full-Time • 8+ years exp • High School Diploma • Alpharetta • Washington
Management
Outlook
SharePoint
Microsoft Office
Apply
$100k – $110k per year • In office • Full-Time • 5+ years exp • Washington
Management
Microsoft Office
Marketing
Salesforce
Apply
≈ $36k – $65k per year (Estimated) • In office • Public Trust • Washington
Analytics
Microsoft Excel
Management
ServiceNow
Microsoft Office
Apply
≈ $43k – $71k per year (Estimated) • In office • Internship • Bachelor's Degree • Washington
Apply
See all jobs
This is one of many
1,385,600 more open roles from verified company boards, updated every day.