368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$97k – $306k per year
Location
In office (Nashville)
Overview
Company
Impact
Profile match
Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas. Founded in 1977 by Larry Ellison, Bob Miner, and Ed Oates, Oracle is one of the world's largest enterprise software and cloud computing infrastructure providers.

Oracle Hardware Platform Development Engineering is seeking a highly driven AI Systems Engineer to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions.

The engineer will identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound, while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results.

A key part of the role is comparative architecture analysis across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure.

Position Overview:

This position is ideal for someone who loves deep systems engineering, debugging complex hardware-software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.

Required Qualifications

  • Solid knowledge of AI / GPU or/and AI/CPU platform architecture and their capabilities.
  • Experience with the architecture, design, and implementation of modern server platforms consisting of multiple architectures and vendors, including x86 and ARM server architectures.
  • Strong communications skills and ability to clearly communicate complex technical issue across engineering disciplines as well as clearly and succinctly articulate issues for executives.
  • Experience and understanding of the latest high-speed busses and interconnect used in modern Compute and AI platforms. Familiarity with their startup connectivity and operational robustness as well as performance metrics.

  • Debugging & Reliability: Troubleshoot complex hardware-software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.

  • Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.

Preferred Qualifications

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.

  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).

  • Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.

  • Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.

  • Experience in benchmarking AI workloads across different architectures.

  • Background in working with the large scale clusters

  • Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $96,800 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

1. Medical, dental, and vision insurance, including expert medical opinion

2. Short term disability and long term disability

3. Life insurance and AD&D

4. Supplemental life insurance (Employee/Spouse/Child)

5. Health care and dependent care Flexible Spending Accounts

6. Pre-tax commuter and parking benefits

7. 401(k) Savings and Investment Plan with company match

8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

9. 11 paid holidays

10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

11. Paid parental leave

12. Adoption assistance

13. Employee Stock Purchase Plan

14. Financial planning and group legal

15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Nashville
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
MLOps Engineer 9 hours ago
$20k – $57k per year (Estimated) • In office • 2+ years exp • Ahmedabad
C++
Python
Databases
FAISS
Pinecone
Weaviate
AI/ML
Airflow
Computer Vision
Kubeflow
MLFlow
NLP
Ray
Ray Serve
TensorRT
Triton
Triton Inference Server
vLLM
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Grafana
Jenkins
Kubernetes
Prometheus
Terraform
Vector
Apply
NLP / LLM Engineer 11 hours ago
$18k – $24k per year (net) • In office • Full-Time • 3+ years exp • Tashkent
Python
Python
FastAPI
Databases
ElasticSearch
Milvus
Pinecone
Qdrant
Weaviate
AI/ML
AI Agents
ChatGPT
DeepEval
Embeddings
Gemini
Hybrid Search
LangChain
Langfuse
LangGraph
LlamaIndex
LLM
LoRA
NLP
PEFT
Prompt Engineering
PyTorch
QLoRA
RAG
Reranking
Semantic Search
Synthetic Data
Tokenization
Triton
vLLM
Transformers
Anthropic
DPO
GraphRAG
Hugging Face
OCR
OpenAI
Semantic Search
SFT
Structured Outputs
Function Calling
TGI
DevOps
CI/CD
Docker
Git
GitHub
Analytics
A/B Testing
Apply
$28k – $59k per year (Estimated) • Remote • 10+ years exp
C#
Go
Java
Node JS
Python
JavaScript
AI/ML
AI Agents
DPO
LLM
LoRA
Ray
Triton
vLLM
PEFT
DevOps
AWS
Azure
GCP
Grafana
Kubernetes
Prometheus
Vector
Apply
$15k – $43k per year (Estimated) • Equity • In office • Full-Time • Astana
PHP
Python
Python
Asyncio
FastAPI
Pydantic
SQLAlchemy
Databases
ElasticSearch
pgvector
PostgreSQL
Qdrant
AI/ML
Claude
Claude Code
CrewAI
Fine-tuning
Function Calling
Google ADK
LangChain
Langfuse
LangGraph
LLM
LoRA
Model Context Protocol
RAG
Scikit-learn
vLLM
NLP
Prompt Engineering
PEFT
Anthropic
OpenAI
OpenAI Codex
DevOps
Docker
Git
QA
Pytest
Apply
$158k – $355k per year • Equity • In office • Bachelor's Degree
C++
Go
Java
Python
Rust
SQL
Databases
OpenSearch
Oracle
AI/ML
AI Agents
Embeddings
Fine-tuning
Function Calling
Hallucination
Human-in-the-Loop
Hybrid Search
Knowledge Graph
LLM Guardrails
Model Context Protocol
Multimodal AI
RAG
Red Teaming
Reranking
DevOps
IAM
Kubernetes
Vector
Apply
$81k – $187k per year • Equity • In office • Bachelor's Degree
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Docker
GCP
Incident Management
Jenkins
Terraform
Apply
In office • PhD
DevOps
Incident Management
SLI/SLO/SLA
Apply
$170k – $355k per year • Equity • In office • PhD • Nashville
DevOps
CI/CD
GitOps
SLI/SLO/SLA
Cybersecurity
Zero Trust
Apply
In office
AI/ML
AI Agents
Apply
Full Stack Engineer 1 hour ago
$89k – $121k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Louisville • Nashville • Dallas
C#
C#
.NET
Databases
Oracle
Cybersecurity
HIPAA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.