368,941open jobs
9,452companies
47,951added this week
Browse all
Salary
$86k – $197k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Groq develops processors and computing systems designed to run artificial intelligence models. Its hardware and cloud services focus on low-latency inference for large language models and other machine learning workloads.

About Groq

Inference is the engine that powers AI, and Groq was built from the silicon up to deliver the world's fastest inference at scale. We pioneered the LPU-the first processor designed specifically for AI inference-and are transforming that innovation into a global cloud platform powering production AI workloads.

With the capital, infrastructure, and team to execute, we're uniquely positioned to define the next era of AI infrastructure. The opportunity is massive, and it's still wide open. Now let's go build it!

Mission:

Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. Our infrastructure teams design and operate the systems that provide the compute, storage, networking, and platform capabilities behind Groq's rapidly growing AI infrastructure.

We are looking for a Senior Storage Design Engineer to architect, design, validate, and evolve the storage infrastructure supporting Groq's large-scale AI and compute environments.

As a Senior Storage Design Engineer, you will own storage architecture and design across high-performance data center and AI infrastructure environments. You will translate workload requirements-including capacity, throughput, latency, durability, availability, and cost-into scalable storage architectures that can be deployed and operated consistently.

This role requires deep expertise in distributed storage, high-performance storage systems, data center infrastructure, Linux, and automation. You will work closely with compute, network, platform, data center, security, and infrastructure operations teams to take storage designs from requirements and benchmarking through production deployment.

You will also help establish the engineering standards, reference architectures, validation methodologies, and technical direction that allow Groq's storage infrastructure to scale.

Responsibilities & opportunities in this role:

  • Architect highly available, high-performance storage platforms supporting large-scale AI and compute infrastructure.
  • Work closely with software teams building inference and training stacks to align storage topology with stack architecture.
  • Develop storage architectures, high-level designs (HLDs), low-level designs (LLDs), implementation standards, and reference architectures.
  • Design storage solutions optimized for high-throughput, low-latency, and highly parallel AI workloads.
  • Architect file, object, block, and local storage solutions based on workload and application requirements.
  • Evaluate distributed storage technologies and platforms such as Ceph, Lustre, Weka, VAST Data, Pure Storage, NetApp, Dell, IBM, or equivalent technologies.
  • Design scalable storage systems using technologies such as NVMe, NVMe-oF, SSD, high-capacity HDD, object storage, and distributed file systems.
  • Develop storage strategies for AI models, datasets, inference workloads, application data, logs, backups, and other infrastructure requirements.
  • Analyze application I/O patterns and translate workload characteristics into storage performance and capacity requirements.
  • Perform capacity planning and develop forecasting models for storage growth, utilization, performance, and lifecycle management.
  • Define storage availability, durability, replication, erasure coding, data protection, backup, and disaster recovery strategies.
  • Design storage architectures spanning multiple data centers or infrastructure environments where appropriate.
  • Partner closely with network engineering to optimize storage traffic, including high-bandwidth east-west connectivity and technologies such as RDMA and RoCEv2.
  • Evaluate storage servers, controllers, drives, network interfaces, and other hardware components for performance, reliability, density, and cost.
  • Build lab environments, proofs of concept, benchmarks, and design-validation frameworks before introducing new technologies into production.
  • Develop representative workload tests to measure throughput, IOPS, latency, metadata performance, scalability, and failure behavior.
  • Define and test failure scenarios involving drives, storage nodes, network connectivity, controllers, and entire failure domains.
  • Drive storage automation and infrastructure-as-code practices using Python, Ansible, APIs, Git, and CI/CD pipelines.
  • Establish storage observability requirements, including telemetry, performance metrics, health monitoring, capacity visibility, and alerting.
  • Troubleshoot complex performance and reliability issues spanning storage, network, compute, operating systems, and applications.
  • Lead technical design reviews and provide guidance on complex storage architecture decisions.
  • Partner with operations teams to ensure storage designs are maintainable, observable, upgradeable, and safe to operate at scale.
  • Mentor engineers and raise the technical bar for storage architecture, benchmarking, automation, documentation, and engineering practices.

Ideal candidates have/are:

  • 8+ years of experience in storage engineering, systems engineering, infrastructure engineering, or architecture, with significant experience designing large-scale storage environments.
  • Deep understanding of distributed storage architecture and high-performance storage systems.
  • Strong knowledge of file, object, and block storage technologies and their respective design tradeoffs.
  • Experience designing storage for data-intensive, highly parallel, or high-performance computing environments.
  • Strong understanding of NVMe, NVMe over Fabrics, SSD technologies, storage networking, and modern storage server architectures.
  • Experience with distributed file systems, object storage platforms, or software-defined storage technologies.
  • Strong understanding of storage performance characteristics, including IOPS, throughput, latency, queue depth, caching, metadata performance, and read/write patterns.
  • Experience with storage resiliency concepts including replication, erasure coding, failure domains, snapshots, backup, and disaster recovery.
  • Strong Linux systems knowledge, including filesystems, storage devices, multipathing, kernel I/O, and performance analysis.
  • Experience with storage benchmarking and performance tools and methodologies.
  • Experience automating infrastructure using Python, APIs, configuration management, or infrastructure-as-code tooling.
  • Familiarity with Git-based workflows, automated testing, and CI/CD engineering practices.
  • Experience with monitoring, telemetry, capacity management, and storage performance analysis.
  • Demonstrated ability to develop architecture documents, design standards, implementation plans, and technical specifications.
  • Ability to make architectural decisions involving performance, reliability, scalability, operational complexity, and cost.
  • Strong communication skills and the ability to influence technical decisions across engineering organizations.
  • Experience designing storage for AI/ML infrastructure, HPC environments, distributed computing systems, or large accelerator clusters.
  • Experience supporting storage environments with extremely high aggregate throughput and highly concurrent access patterns.
  • Knowledge of RDMA, RoCEv2, GPUDirect Storage, NVMe-oF, and high-performance Ethernet storage fabrics.
  • Experience with parallel filesystems such as Lustre, IBM Storage Scale/GPFS, WekaFS, or equivalent technologies.

Compensation

Groq is committed to providing competitive compensation through our Total Cash philosophy, which incorporates potential bonus value directly into base pay. The total cash salary range for this position, which is inclusive of the potential bonus value, is TBD, with individual placement determined by your geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates will vary based on local market dynamics. Beyond cash compensation, Groq also offers a Long-Term Incentive (LTI) Program and a robust suite of employee benefits.

US Job Posting

This position may require access to technology and/or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). To comply with these requirements, candidates for this role must meet certain citizenship or residency criteria. Specifically, they must qualify as U.S. Persons for export control purposes (i.e., U.S. citizen, U.S. lawful permanent resident (Green Card holder), or a protected individual under 8 U.S.C. § 1324b(a)(3) such as a refugee or asylee), or otherwise be eligible for an applicable export license.

Groq is an Equal Opportunity Employer. We are committed to creating an inclusive environment for all employees and applicants. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex (including gender identity, sexual orientation, and pregnancy), age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable law.

Groq complies with all applicable federal, state, and local laws governing nondiscrimination in employment. We do not tolerate discrimination or harassment based on any protected characteristic.

Groq iscommitted to working with and providing reasonable accommodations to qualified individuals with physical or mental disabilities. If you require a reasonable accommodation to complete an application or to participate in the hiring process, please contact us at [email protected]. This contact is for accommodation requests only, which will be considered on a case-by-case basis.

All offers of employment are contingent upon verification of the applicant’s identity and employment authorization in accordance with federal law.

Groq encourages people with criminal record histories to apply for employment, and values diverse experiences, including prior contact with the criminal legal system. To that end, Groq welcomes such applicants in accordance with the California Fair Chance Act, Los Angeles City Fair Chance Act Ordinance, Los Angeles County Fair Chance Act Ordinance, and San Francisco Fair Chance Act Ordinance. Philadelphia applicants can review information pertaining to Philadelphia’s Fair Criminal Record Screening Standards Ordinance here: https://www.phila.gov/documents/fair-chance-hiring-law-poster.

As part of our hiring process, Groq may use artificial intelligence (“AI”) tools or automated systems to assist with activities such as reviewing applications, evaluating qualifications, scheduling interviews, analyzing assessment responses, or supporting recruiting operations. These tools are designed to assist-not replace-human decision-making, and hiring decisions are subject to human review. We may process information you provide during the application process, including resumes, application materials, interview responses, assessments, and, where applicable, audio, video, or transcript data. If legally required, we will request consent before using technologies that analyze biometric or video interview data. Candidates may request reasonable accommodations, an alternative evaluation process, additional information regarding the use of AI in the hiring process, or review of certain automated decisions by contacting [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Full Stack Engineer 2 hours ago
$127k – $264k per year (Estimated) • In office • Full-Time • 9+ years exp • Sydney • Canberra • Melbourne
PowerShell
SQL
C#
TypeScript
JavaScript
C#
.NET
Databases
MS SQL
Frontend
Angular
DevOps
AWS
Azure
Azure DevOps
CI/CD
GCP
Git
Apply
$131k – $237k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Chantilly • Columbia
Bash
Python
TypeScript
JavaScript
Databases
PostgreSQL
Frontend
Angular
DevOps
Ansible
ArgoCD
AWS
Azure
buildah
CI/CD
Docker
Docker Swarm
FluxCD
GitLab
GitLab CI
Grafana
Helm
Kubernetes
Loki
Prometheus
Terraform
Twelve-Factor App
Podman
Apply
Remote • Full-Time • Bachelor's Degree • Taiwan
AI/ML
Agentforce
AI Agents
DevOps
CI/CD
Chips/EDA
PoC Library
Marketing
Salesforce
Apply
$47k – $165k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
DevOps
CI/CD
Apply
$215k – $240k per year • In office • Full-Time • 5+ years exp • Seattle
Python
Rust
AI/ML
Flyte
OpenAI
DevOps
ArgoCD
AWS
Azure
Buildkite
CI/CD
Docker
GCP
Helm
Kubernetes
Terraform
Apply
See all jobs
This is one of many
368,941 more open roles from verified company boards, updated every day.