1,454,406open jobs
86,438companies
225,454added this week
Browse all
Salary
≈ $184k – $376k per year (Estimated)
Location
In office (New York)
Seniority
Staff · 5+ years exp
Visa
H-1B filings in 12 months: 19 · for this role: 9 · green card filings: 2
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 9, 2026.

Overview
Company
Impact
Profile match
Groq is an American semiconductor company founded in 2016 by former Google engineers who had worked on the Tensor Processing Unit. It designs the Language Processing Unit, a deterministic single-core architecture with on-chip memory that removes the scheduling unpredictability of GPUs and delivers unusually low latency for large language model inference. Rather than selling chips alone the company runs GroqCloud, a hosted inference service where developers call open models through an API, and it has signed large capacity agreements to build data centres in the United States and the Middle East.

About Groq

Inference is the engine that powers AI, and Groq was built from the silicon up to deliver the world's fastest inference at scale. We pioneered the LPU-the first processor designed specifically for AI inference-and are transforming that innovation into a global cloud platform powering production AI workloads.

With the capital, infrastructure, and team to execute, we're uniquely positioned to define the next era of AI infrastructure. The opportunity is massive, and it's still wide open. Now let's go build it!

Mission:

GroqCore is building a high-performance software stack that converts bare metal compute into a token-producing engine This role is for a seasoned engineer with practical experience making LLMs fast, efficient, and reliable in production. This individual will work on an inference serving software stack from a trained model to a high-throughput delivery of output tokens. This will require experience on understanding/working with model internals, runtime software, and the supporting GPU & accelerator hardware underpinning it all.

Location: This role will be based in one of our three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired for this role must be based in one of these three areas. You’ll have the flexibility to work remotely while we establish our local Groq office, with the expectation that this role will transition to onsite once the office opens.

Responsibilities & opportunities in this role:

  • Build and optimize serving systems that handle high request volumes with low, predictable latency
  • Apply model compilation and graph optimization techniques, and benchmark carefully to use them only where they deliver real gains
  • Implement and tune batching, caching, scheduling, and memory management strategies for large models
  • Apply reduced-precision and quantization methods
  • Distribute models across multiple accelerators and nodes, balancing throughput & latency
  • Profile end-to-end performance, identify bottlenecks, and fix them-ranging from low-level kernels to request routing-leveraging both lab/development environments and live telemetry from production systems.
  • Partner with infrastructure and FDE teams to bring new models into production quickly
  • Partner with vendors along the inference serving path in support of optimizing the stack

Ideal candidates have/are:

  • BS / MS / PhD in CS, CE, EE, or equivalent depth from industry
  • 5+ years shipping performance-critical products, with some portion of that serving models at scale
  • A strong applied understanding of transformer architectures
  • Fluency in at least one systems language and one high-level language
  • A rigorous, measurement-driven approach to performance work
  • Clear communication about tradeoffs to both technical and business stakeholders

Ways to stand out:

  • Experience writing or tuning custom kernels.
  • Familiarity with non-GPU or specialized inference hardware.
  • Contributions to open-source serving or compiler projects.

Why Join Us:

  • Purposeful Hiring: You’re not here by accident, and neither is anyone else. Every teammate is handpicked with intention because who we build with matters.
  • Builders Wanted: You’re not just riding the rocket ship, you’re building it. Your work directly shapes the trajectory of our company.
  • Mission-Driven Work: We’re here to make a real impact. Our mission fuels everything we do.
  • Tackling Hard Problems: If easy isn’t your thing, you’re in the right place. We solve some of the most complex and exciting challenges in our space.
  • Excellence Is The Standard: High performance isn’t just encouraged, it’s the baseline. And it’s contagious.

If this sounds like you, we’d love to hear from you!

Compensation

Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is $341,400 - $401,600, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.

US Job Posting

This position may require access to technology and/or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet certain citizenship or residency criteria. Specifically, they must qualify as U.S. Persons for export control purposes (i.e., U.S. citizen, U.S. lawful permanent resident (Green Card holder), or a protected individual under 8 U.S.C. § 1324b(a)(3) such as a refugee or asylee), or otherwise be eligible for an applicable export license.

Non-US Job Postings

This position may require access to technology and/or information subject to U.S. export control laws and regulations, as well as applicable local laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet all relevant export control eligibility criteria.

Groq is an Equal Opportunity Employer. We are committed to creating an inclusive environment for all employees and applicants. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex (including gender identity, sexual orientation, and pregnancy), age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law.

Groq complies with all applicable federal, state, and local laws governing nondiscrimination in employment. We do not tolerate discrimination or harassment based on any protected characteristic.

Groq iscommitted to working with and providing reasonable accommodations to qualified individuals with physical or mental disabilities. If you require a reasonable accommodation to complete an application or to participate in the hiring process, please contact us at [email protected]. This contact is for accommodation requests only, which will be considered on a case-by-case basis.

All offers of employment are contingent upon verification of the applicant’s identity and employment authorization in accordance with federal law.

Groq encourages people with criminal record histories to apply for employment, and values diverse experiences, including prior contact with the criminal legal system. To that end, Groq welcomes such applicants in accordance with the California Fair Chance Act, Los Angeles City Fair Chance Act Ordinance, Los Angeles County Fair Chance Act Ordinance, and San Francisco Fair Chance Act Ordinance. Philadelphia applicants can review information pertaining to Philadelphia’s Fair Criminal Record Screening Standards Ordinance here: https://www.phila.gov/documents/fair-chance-hiring-law-poster.

As part of our hiring process, Groq may use artificial intelligence (“AI”) tools or automated systems to assist with activities such as reviewing applications, evaluating qualifications, scheduling interviews, analyzing assessment responses, or supporting recruiting operations. These tools are designed to assist-not replace-human decision-making, and hiring decisions are subject to human review. We may process information you provide during the application process, including resumes, application materials, interview responses, assessments, and, where applicable, audio, video, or transcript data. If legally required, we will request consent before using technologies that analyze biometric or video interview data. Candidates may request reasonable accommodations, an alternative evaluation process, additional information regarding the use of AI in the hiring process, or review of certain automated decisions by contacting [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,454,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
New York
$254k – $407k per year • In office • Full-Time • Bachelor's Degree • San Francisco • Boston • New York • Atlanta • Miami
AI/ML
AI Agents
LLM
Machine Learning
DevOps
Kubernetes
Platform Engineering
HPC
Apply
≈ $140k – $265k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Houston
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
PyTorch
RAG
CUDA
LLMOps
LLM Evaluation
Machine Learning
DevOps
AWS
Kubernetes
Management
Agile
Apply
≈ $133k – $253k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Houston • Tlaquepaque
Python
Java
C++
Python
FastAPI
C++
TensorFlow C++
PyTorch C++
AI/ML
Spark
Model Context Protocol
Scikit-learn
AWS Bedrock
TensorFlow
PyTorch
Amazon SageMaker
Edge AI
Machine Learning
DevOps
Terraform
Azure
CI/CD
AWS
Kubernetes
Management
Agile
Apply
$129k – $216k per year • Equity • In office • Master's Degree • Calabasas
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Reinforcement Learning
Scikit-learn
TensorFlow
PyTorch
Machine Learning
Chips/EDA
Cadence Virtuoso
Management
Agile
Apply
≈ $132k – $251k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Colorado Springs
Python
C#
C++
AI/ML
LangGraph
AutoGen
LangChain
Prompt Engineering
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Tool Use
DevOps
CI/CD
Git
Bitbucket
Management
Confluence
Jira
Agile
Apply
≈ $36k – $90k per year (Estimated) • In office • 5+ years exp • Lahore
Python
AI/ML
LoRA
Fine-tuning
Embeddings
Quantization
Multimodal AI
AI Agents
NLP
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Hugging Face
Machine Learning
Apply
≈ $64k – $161k per year (Estimated) • In office • 5+ years exp • Abu Dhabi
Python
AI/ML
LoRA
Fine-tuning
Embeddings
Quantization
Multimodal AI
AI Agents
NLP
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Hugging Face
Machine Learning
Apply
≈ $132k – $254k per year (Estimated) • Remote (United States) • TS/SCI • Full-Time • 5+ years exp
AI/ML
Quantization
Knowledge Distillation
Agentic Workflows
Model Distillation
Machine Learning
DevOps
Helm
Docker
Kubernetes
Cybersecurity
SIEM
Apply
AI Engineer 1 day ago
Remote (United States) • Full-Time
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
Databases
Weaviate
Pinecone
AI/ML
Claude
OpenCV
DeepSpeed
MLFlow
Triton Inference Server
Reinforcement Learning
Quantization
Prompt Engineering
Computer Vision
NLP
ONNX
Transformers
TensorFlow
PyTorch
RAG
Ray
Triton
Hugging Face
ONNX Runtime
DevOps
GCP
Azure
AWS
Apply
≈ $152k – $318k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
llama.cpp
Qwen
LoRA
vLLM
Gemma
Fine-tuning
Quantization
Knowledge Distillation
AWQ
GGUF
GPTQ
ONNX
PEFT
QLoRA
TGI
MLX ML
Transformers
PyTorch
LLM
Tokenization
Hugging Face
Edge AI
ExecuTorch
ONNX Runtime
Model Distillation
DevOps
GitHub
Management
Discord
Apply
≈ $157k – $322k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Dallas
AI/ML
Groq
Quantization
Apply
≈ $203k – $415k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
Groq
Quantization
Apply
≈ $120k – $251k per year (Estimated) • In office • Full-Time • Dallas
AI/ML
Groq
DevOps
Kubernetes
Apply
≈ $142k – $283k per year (Estimated) • In office • Full-Time • 8+ years exp • San Francisco
Python
AI/ML
Groq
Time Series Forecasting
InfiniBand
DevOps
Ansible
GCP
Azure
CI/CD
Git
AWS
HPC
Linux
TCP/IP
BGP
Apply
≈ $128k – $256k per year (Estimated) • In office • Full-Time • 8+ years exp • New York
Python
AI/ML
Groq
Time Series Forecasting
InfiniBand
DevOps
Ansible
GCP
Azure
CI/CD
Git
AWS
HPC
Linux
TCP/IP
BGP
Apply
$210k – $275k per year • In office • Full-Time • New York
Apply
$40k – $42k per year • In office • Full-Time • High School Diploma • New York
Apply
Diesel Mechanic 1 day ago
$60k – $80k per year • Equity • In office • Full-Time • 5+ years exp • New York
Apply
$120k – $145k per year • In office • New York
Apply
up to $72k per year • In office • Full-Time • 2+ years exp • New York
Management
Microsoft Office
Apply
See all jobs
This is one of many
1,454,406 more open roles from verified company boards, updated every day.