368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office
Seniority
Principal
Overview
Company
Impact
Profile match
Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas. Founded in 1977 by Larry Ellison, Bob Miner, and Ed Oates, Oracle is one of the world's largest enterprise software and cloud computing infrastructure providers.

As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will play a critical role in designing, implementing, and maintaining the infrastructure that supports our customers AI and machine learning initiatives.

You will work closely with customer technical teams (data scientists, software engineers, and Infra/IT professionals) and internal cloud services team to ensure their AI/ML solutions are deployed efficiently, securely, and at scale.

Your expertise will be crucial in optimizing our infrastructure for performance, reliability, and cost-effectiveness. In this hands-on technical role you’ll troubleshoot issues for proof of concept (POC) and production deployments.

  • Lead the design and architecture of AI/ML infrastructure, considering performance, scalability, and security.
  • Implement and maintain AI/ML infrastructure, ensuring efficient and reliable operations.
  • Collaborate with customer technical teams to understand their AI/ML requirements and provide tailored solutions.
  • Work closely with our cloud services team to integrate AI/ML solutions into our cloud platform.
  • Ensure the secure deployment of AI/ML models and data, adhering to industry best practices.
  • Monitor and optimize AI/ML infrastructure performance, identifying and resolving bottlenecks.
  • Provide technical guidance and mentorship to junior engineers, fostering a culture of knowledge sharing.
  • Stay updated with the latest AI/ML technologies and trends, driving innovation within the team.
  • Document and communicate infrastructure designs, ensuring clear and concise documentation.
  • Engage with customers and stakeholders to gather feedback and ensure their satisfaction with our AI/ML offerings.

Qualifications:

  • Experience in scripting and automation using tools like Ansible, Terraform, Python and/or Kubernetes.

  • Experience with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (like Slurm, PBS, etc.) for managing distributed systems.

  • Solid understanding of networking concepts, security principles, and best practices.

  • Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.

  • Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders.

  • Strong documentation skills with experience documenting infrastructure designs, configurations, procedures, and troubleshooting steps to facilitate knowledge sharing, ensure maintainability, and enhance team collaboration.

  • Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.

Preferred Qualifications

  • Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala

  • Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads.

  • Understanding machine learning frameworks and libraries such as TensorFlow, PyTorch, or sci-kit-learn and their deployment in production environments is a plus.

  • Familiarity with DevOps practices and tools for continuous integration, deployment, and monitoring (e.g., Jenkins, GitLab CI/CD, Prometheus).

  • Strong experience with High-Performance Computing/GPU systems

ADDITIONAL INFORMATION:

Core Responsibilities

Strategic Thought Leadership:

- You should also have a demonstrated ability to think strategically about business, products, and technical challenges.

Business Experience:

- Understand the challenges in working with large customers.

Planning & Execution:

- Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.

Collaboration & Partnership:

- Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.

Problem Solving:

- Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem solving strategies.

Continuous Learning:

- Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.

Continuous Improvement:

- Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.

Performance and Development:

- Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.

Career Level - IC4

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Lead AI Engineer 9 hours ago
$30k – $73k per year (Estimated) • In office • Full-Time • Pune
Python
AI/ML
Fine-tuning
LLM
Reinforcement Learning
LLM Guardrails
AI Agents
DevOps
CI/CD
Docker
GitOps
Helm
Kubernetes
OpenShift
Platform Engineering
Vector
Apply
$22k – $58k per year (Estimated) • In office • Full-Time • Pune
AI/ML
Claude
Claude Code
Copilot
AI Agents
DevOps
ArgoCD
AWS
Azure
Backstage
CI/CD
Docker
GCP
GitHub Actions
Jenkins
Kubernetes
Platform Engineering
GitHub
Apply
$22k – $41k per year (Estimated) • Remote • 3+ years exp • Moscow
DevOps
Ansible
CI/CD
Grafana
Jenkins
Kubernetes
KVM
OpenShift
Prometheus
Terraform
GitLab
Cybersecurity
SonarQube
Apply
$20k – $84k per year (Estimated) • In office • Full-Time • Pune
JavaScript
SQL
Java
Kotlin
Java
Gradle
Maven
Spring Boot
Spring MVC
Spring Security
Kotlin
Mockito
DevOps
Bitbucket
CI/CD
Docker
Jenkins
Kibana
Kubernetes
OpenShift
GitLab
Apply
Senior PostgreSQL SRE 10 hours ago
$23k – $58k per year (Estimated) • In office • Full-Time • 10+ years exp • Pune
PowerShell
Python
SQL
Databases
MS SQL
Oracle
PostgreSQL
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Configuration Management
GCP
Grafana
Helm
Kubernetes
OpenShift
Prometheus
Apply
In office • PhD
DevOps
Incident Management
SLI/SLO/SLA
Apply
$170k – $355k per year • Equity • In office • PhD • Nashville
DevOps
CI/CD
GitOps
SLI/SLO/SLA
Cybersecurity
Zero Trust
Apply
In office
AI/ML
AI Agents
Apply
$115k – $235k per year • Equity • In office • Bachelor's Degree • Nashville
Python
Apply
In office
DevOps
Incident Management
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.