368,611open jobs
9,439companies
50,719added this week
Browse all
Location
Remote/Hybrid (Dublin, Ireland)
Employment
Full-Time
Overview
Company
Impact
Profile match

F5

F5 is an award-winning creative agency based in Shanghai that specializes in brand strategy, digital marketing, and tech-driven advertising solutions. Taking its name from the computer refresh key, the agency focuses on refreshing traditional ideas by integrating modern technology, AI, and storytelling into global brand campaigns. It serves major international and domestic enterprise clients, helping businesses drive innovation and connect with global audiences through creative storytelling.

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

TheAI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses onoptimizing Large Language Models (LLMs) for inference, serving diverse environments-from GPU-rich data centers to resource-constrained edge devices-with a strong emphasis on maximizing throughput, minimizing latency, andmaintainingmodel accuracy.

This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability byleveraginghardware acceleration, designing scalable infrastructure, andmonitoringsystem performance.

Key Responsibilities

High-Performance AI Serving

  • Build andmaintainrobust inference engines using tools likevLLM, TGI (Text Generation Inference), andNVIDIA Triton, ensuring high performance at scale.

  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.

Hardware Acceleration and Optimization

  • Profile andoptimizemodels for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators likeTPUs andLPUs.

  • Collaborate with hardware teams to maximizeutilizationand performance across various computational environments.

Inference Orchestration and Scalability

  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines,leveragingKubernetes for inference routing and orchestration.

  • Ensure software solutions areoptimizedfor peak performance during traffic spikes,maintainingreliability and scalability.

Performance Monitoring and Observability

  • Establish robust observability frameworks tomonitor Time to First Token (TTFT), tokens per second, and memory bandwidthutilizationagainst service-level agreements (SLAs).

  • Build and executeperformance and load testing suites toidentifybottlenecks and ensure consistent reliability at scale.

Technical Requirements

Required Skills:

  • Programming Languages: Proficiencyin programming languages such asPython,C++,Rust, orGolang specifically for high-performance AI workflows.

  • Inference Tools: Proven hands-on experience with tools likevLLM,TensorRT,Llama.cpp, andOllama for inference development and optimization.

  • Infrastructure Expertise: Strong familiarity with infrastructure technologies, includingDocker,Kubernetes, and cloud platforms such asAWS,GCP, andAzure.

  • Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling andoptimizingperformance for accelerators like NVIDIA GPUs and TPUs.

Preferred Experience:

  • Prior experience deploying Large Language Models (LLMs) with advanced techniques likeSpeculative Decoding orPagedAttention.

  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).

  • Background inMLOps orSRE roles focused onhigh-performance AI endpoints and reliability during demand surges.

  • Proficiencyin designing scalable solutions for high-throughput inference environmentsoptimizedfor traffic bursts.

Success Metrics (KPIs):

  • Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) andmaximumtokens per second.

  • Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resourceutilizationand hardware optimization.

  • Scalability: Maintainsystem stability and reliability duringtraffic spikes, ensuring performance consistency across environments.

  • Throughput Maximization: Deploy modelsoptimizedfor peak hardware usage and maximized process throughput.

Why Join F5?

F5 empowers you to push boundaries inAI optimization andhigh-performance engineering. Joining our team means:

  • Collaborating withcutting-edgetechnologies and hardware solutions to support real-time AI applications.

  • Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving.

  • Driving transformative projects that deliver real-time AI reliability to global customers whilemaintainingcost and efficiency standards.

  • Working on advancedMLOpssolutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment.

What Success Looks Like:

As anAI Inference Engineer at F5, success is measured by your ability to:

  • Combine technicalexpertiseand problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems.

  • Collaborate efficiently across cross-functional teams,participatingin knowledge sharing and system refinement.

  • Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals.

  • Translatecomplex AI and inference workflows into practical solutions that align with F5's strategicobjectives.

The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.

Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or@myworkday.com).

Equal Employment Opportunity

It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Dublin
$93k – $126k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
Node JS
SQL
TypeScript
JavaScript
Databases
MS SQL
Oracle
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
GCP
Git
GitLab
GitLab CI
Jenkins
Management
Jira
QA
JMeter
Playwright
Postman
Rest-Assured
TestNG
Apply
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
$68k – $142k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Marseille
AI/ML
Knowledge Graph
DevOps
AWS
Azure
GCP
Management
ServiceNow
Apply
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$150k per year • In office • Full-Time • New York
C++
Python
Apply
$311k – $467k per year • Equity • Remote/Hybrid • Full-Time • 12+ years exp • PhD • San Jose • Seattle
Design
Figma
InVision
Apply
$213k – $320k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • San Jose • Seattle
DevOps
Kubernetes
API Gateway
Apply
$177k – $265k per year • Equity • In office • Full-Time • PhD • San Jose • Seattle
C++
Go
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Llama
llama.cpp
Ollama
TensorRT
TGI
Triton
vLLM
TPU
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$49k – $127k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Guadalajara
Python
SQL
Databases
Databricks
Delta Lake
Snowflake
AI/ML
dbt
LLM
Spark
MLFlow
Streamlit
DevOps
Azure
Azure DevOps
CI/CD
GitHub Actions
Platform Engineering
GitHub
GitLab
Analytics
ETL/ELT
Apply
$35k – $118k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Guadalajara
SQL
DevOps
Rest API
Marketing
Marketo
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • Dublin
JavaScript
Ruby
Scala
Go
Apply
$49k – $174k per year (Estimated) • In office • Internship • Bachelor's Degree • Dublin
Go
JavaScript
Ruby
Scala
Apply
$59k – $103k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Dublin
Apply
$88k – $181k per year (Estimated) • In office • Full-Time • Dublin
Apply
$111k – $201k per year (Estimated) • In office • 10+ years exp • Dublin
Ruby
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.