598,796open jobs
30,416companies
86,197added this week
Browse all
Salary
$147k – $275k per year (Estimated)
Location
In office (Santa Clara)
Seniority
Architect
Overview
Company
Impact
Profile match
FuriosaAI designs high-performance, power-efficient AI accelerators (NPUs) used in data centers for computer vision, GenAI, LLMs, and demanding workloads.

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. 

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence  for every enterprise.

About the Role

FuriosaAI is looking for a Solutions Architect to bring the full potential of our powerful RNGD chips/servers to our customers by acting as the primary technical authority in AI/LLM model deployments. From running POCs to benchmarking and debugging, you will translate RNGD’s powerful system to real-world deployments of customers’ models, empowering customers with FuriosaAI’s powerful solutions.

If you are interested in providing the technical expertise in challenging the current status-quo of AI infrastructure in real-world environments, join us in our path to a sustainable future of AI.

Key Responsibilities

  • Own end-to-end technical enablement for US customers deploying AI models on FuriosaAI's RNGD NPU using the Furiosa SDK

  • Develop POCs, benchmarking studies, and live debugging sessions directly in customer environments

  • Act as the technical authority to the US BD/Sales team during pre-sales and enterprise evaluations; translate deep technical capability into business value for engineering and C-suite audiences

  • Develop deep, current expertise in FuriosaAI's hardware and software stack and demonstrate it at US technical forums, AI conferences, and customer workshops

  • Onboard and train customers on integration patterns, optimization workflows, and best practices post-purchase

  • Serve as a technical feedback loop from US customers back to Seoul HQ product and engineering teams

Minimum Qualifications

  • 2-5 years in a US customer-facing technical role: Solutions Architect, Sales Engineer, Forward Deployed Engineer, or equivalent at an AI infra, cloud, or semiconductor company

  • Actively current on the AI/LLM landscape - tracking model releases, inference frameworks, and serving stack evolution in real time

  • Hands-on experience with modern inference stacks: vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar

  • Hands-on experience with agent and orchestration frameworks: LangChain, LlamaIndex, LangGraph, AutoGen, or MCP-based tooling

  • Proficiency in Python; comfortable with DNN frameworks (PyTorch, TensorFlow)

  • Strong written and verbal communication - able to engage credibly with ML engineers at frontier labs and VP/C-suite executives

  • Authorized to work in the US; able to travel to customer sites and to Seoul HQ periodically

Preferred Qualifications

  • Prior experience at a US AI chip company, cloud silicon team, or AI infrastructure startup

  • Familiarity with NPU/GPU accelerator ecosystems, PCIe integration, and data center hardware deployment

  • Experience with inference optimization: quantization, kernel tuning, batching strategies, memory bandwidth optimization

  • Proficiency in C, C++, or Rust

  • Experience working with distributed or cross-timezone engineering teams

Why Join FuriosaAI

The defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.

With our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.

At Furiosa, you will:

Solve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.

Pioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.

Ship Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.

Partner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.

Do Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. 

Contact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
598,796 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$136k – $281k per year (Estimated) • Remote/Hybrid • Full-Time • Charlotte
Python
TypeScript
Databases
PostgreSQL
pgvector
ElasticSearch
OpenSearch
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
AI Agents
Human-in-the-Loop
LLM Guardrails
DevOps
GCP
Azure
AWS
Apply
$117k – $253k per year (Estimated) • Remote/Hybrid • Full-Time • Charlotte
Python
TypeScript
Databases
PostgreSQL
pgvector
ElasticSearch
OpenSearch
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
AI Agents
Human-in-the-Loop
LLM Guardrails
DevOps
GCP
Azure
AWS
Apply
$30k – $62k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
Python
Java
SQL
Scala
Databases
ClickHouse
Apache Iceberg
Presto
Apache Kafka
Trino
AI/ML
Copilot
Cursor
Spark
Airflow
Flink
OpenAI Codex
Apply
$168k – $245k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Jose • Reston
Python
DevOps
Terraform
Ansible
Helm
GitOps
AWS
Kubernetes
Amazon EKS
Amazon S3
Apply
$33k – $74k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bengaluru • Noida • Chennai • Gurgaon • Hyderabad
Python
Go
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Google GKE
Azure AKS
Apply
$36k – $93k per year (Estimated) • In office • Seoul
Rust
C++
Apply
$35k – $102k per year (Estimated) • Remote/Hybrid • Hwaseong
Rust
C++
Apply
$35k – $102k per year (Estimated) • In office • Seoul
C
C++
C
Embedded C
DevOps
RTOS
HPC
Apply
$35k – $84k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
Python
Rust
AI/ML
CUDA Toolkit
Quantization
PyTorch
LLM
CUDA
Triton
Hugging Face
KV Cache
DevOps
Git
Apply
$33k – $95k per year (Estimated) • In office • Bachelor's Degree • Seoul
Python
Rust
DevOps
Kubernetes
Apply
$140k – $311k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • PhD • Santa Clara
JavaScript
Java
AI/ML
AI Agents
Frontend
Svelte
Web3
Layer 2
Robotics
Digital Twin
Apply
$145k – $314k per year (Estimated) • Remote/Hybrid • Contractor • Master's Degree • Santa Clara
Python
AI/ML
llama.cpp
LoRA
vLLM
Fine-tuning
Quantization
Multimodal AI
Knowledge Distillation
AI Agents
VLM
SGLang
GGUF
TensorRT
PEFT
TensorRT-LLM
Transformers
PyTorch
LLM
Mixture of Experts
DPO
Post-training
Edge AI
Speculative Decoding
KV Cache
Multi-Agent Systems
Model Distillation
Analytics
ETL/ELT
Management
Freshdesk
Apply
$201k – $352k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
TypeScript
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
AI Agents
LLM
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
Function Calling
AI Agents
RAG
OpenAI Codex
Human-in-the-Loop
LLM Guardrails
Multi-Agent Systems
Tool Use
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 15+ years exp • Bachelor's Degree • Santa Clara
Python
Go
Java
AI/ML
Cursor
Windsurf
Claude Code
AI Agents
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
See all jobs
This is one of many
598,796 more open roles from verified company boards, updated every day.