857,935open jobs
54,174companies
144,275added this week
Browse all
Salary
≈ $195k – $328k per year (Estimated)
Location
Hybrid (Toronto, Canada)
Seniority
Staff · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 28, 2026. First seen by Alion on Sep 23, 2026. Cerebras Systems scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Cerebras Systems is an American computer hardware company founded in 2016 and headquartered in Sunnyvale, California that builds accelerators for artificial intelligence at wafer scale. Instead of assembling clusters from many small chips, it manufactures a single processor the size of an entire silicon wafer, the Wafer Scale Engine, which removes most of the communication overhead in large model training and inference. The company sells CS-series systems to research laboratories and enterprises, operates its own inference cloud known for very high token throughput, and has built large supercomputers with partners including the Gulf technology group G42.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About the Role

Cerebras is building a new generation of disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine.

We are hiring a Software Engineer to build and evolve the ML API layer that makes this heterogeneous serving system accessible, reliable, and easy to use. You will work across our inference APIs, model integration layer, request-routing services and Cerebras inference platform to deliver a consistent experience across models and accelerator backends.

This role sits at the intersection of machine learning systems, API design, model serving, and distributed systems. You will enable new model architectures and inference capabilities, define stable user-facing behavior, and ensure that features such as streaming, sampling, tool use, structured outputs, multimodal inputs, and model configuration behave correctly and consistently in production.

You will work closely with model enablement, compiler, runtime, cloud infrastructure, product, customer-facing teams, and customers directly. This is a hands-on software engineering role for someone who enjoys turning rapidly evolving ML capabilities into durable, production-quality APIs.

Responsibilities

  • Build production ML inference APIs. Design, implement, and maintain APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and other emerging inference capabilities.

  • Deliver a unified serving experience. Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.

  • Enable new models and capabilities. Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features into the serving platform.

  • Own API compatibility and evolution. Maintain compatibility with widely adopted inference interfaces while designing Cerebras-specific extensions. Establish clear versioning, deprecation, validation, and backward compatibility practices.

  • Integrate with model-serving runtimes. Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, the AMD ROCm stack, and Cerebras runtime components.

  • Support disaggregated inference. Build the control and data paths required to coordinate GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management.

  • Improve serving performance. Optimize streaming behavior, time to first token, request latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components.

  • Ensure functional and numerical correctness. Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and compatibility across serving backends.

  • Strengthen reliability and observability. Define end-to-end service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling for production inference traffic.

  • Develop testing and qualification infrastructure. Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates.

  • Improve developer experience. Build intuitive configuration, SDKs, documentation, examples, debugging tools, and self-service workflows for internal developers, customers, and partners.

  • Collaborate across the stack. Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate model and customer requirements into scalable serving capabilities.

Minimum Qualifications

  • 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems.

  • Strong programming ability in Python and Go plus experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language.

  • Experience building stable APIs with clear validation, error handling, observability, compatibility, and versioning practices.

  • Experience integrating software across service, framework, runtime, and infrastructure boundaries.

  • Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs.

  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production.

  • Ability to diagnose correctness, reliability, and performance issues across multiple components of a distributed serving system.

  • Strong communication and cross-functional execution skills, with the ability to turn ambiguous model or product requirements into production-quality software.

  • Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Preferred Qualifications

  • Experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project.

  • Experience creating API conformance, model-quality, numerical-comparison, determinism, or performance-regression test systems

  • Experience building SDKs, developer tools, model registries, configuration systems, or self-service ML platforms.

  • Experience with multi-model or multi-tenant inference platforms, including routing, admission control, fairness, quotas, rate limiting, and capacity-aware scheduling

  • Understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding.

  • Experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling.

  • Experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications.

  • Familiarity with reduced-precision inference and quantization formats such as BF16, FP8, FP4, INT8, or INT4.

Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  • Build a breakthrough AI platform beyond the constraints of the GPU.

  • Publish and open source their cutting-edge AI research.

  • Work on one of the fastest AI supercomputers in the world.

  • Enjoy job stability with startup vitality.

  • Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
857,935 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Toronto
$88k – $138k per year • In office • Full-Time • Bachelor's Degree • Calgary
Python
JavaScript
TypeScript
SQL
C#
C++
Node JS
C#
ASP.NET Core
Entity Framework Core
Databases
MS SQL
AI/ML
ChatGPT
Prompt Engineering
AI Agents
OpenAI
Multi-Agent Systems
Frontend
Next.js
Angular
Bootstrap
React.js
DevOps
Azure DevOps
Azure
CI/CD
Git
GitHub
Management
Agile
Scrum
Apply
≈ $94k – $200k per year (Estimated) • In office • Full-Time • Halifax
JavaScript
TypeScript
SQL
Databases
Snowflake
Databricks
Frontend
Angular
React.js
DevOps
Azure
Platform Engineering
Management
Agile
Apply
$79k – $107k per year • Remote (Canada) • Full-Time • 6+ years exp • Bachelor's Degree • Ontario
Python
TypeScript
Databases
PostgreSQL
Snowflake
Databricks
DevOps
AWS
Incident Management
Apply
$130k – $150k per year • Hybrid • 8+ years exp • Toronto
Python
Go
JavaScript
Java
TypeScript
C#
Python
FastAPI
Java
Spring Boot
C#
.NET
AI/ML
LangGraph
LangChain
Claude
LlamaIndex
Vertex AI
AI Agents
LLM
RAG
Google ADK
OpenAI
Frontend
React.js
DevOps
Terraform
GCP
Kustomize
CI/CD
Kubernetes
Google GKE
Management
Agile
Scrum
Kanban
Apply
≈ $130k – $242k per year (Estimated) • Remote (Canada) • Full-Time • 10+ years exp
Python
JavaScript
Rust
C#
C++
C#
.NET
DevOps
Rest API
Apply
$80k – $95k per year • Remote (location not specified)
Python
JavaScript
TypeScript
SQL
Node JS
Databases
Google BigQuery
BigQuery
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
OpenAI
Anthropic
Structured Outputs
Tool Use
Frontend
Next.js
React.js
DevOps
Rest API
Vercel
Git
Analytics
A/B Testing
Design
Figma
Management
n8n
Zapier
Apply
In office • 7+ years exp • Bachelor's Degree • Bengaluru
Python
Perl
DevOps
RTOS
CI/CD
Linux
TCP/IP
Wi-Fi
IoT
MQTT
Zigbee
Management
Agile
Apply
≈ $90k – $200k per year (Estimated) • Remote (United States, Canada, ET hours) • 8+ years exp
Python
Bash
Databases
MySQL
Cassandra
Apache Kafka
AI/ML
Cursor
Claude
DevOps
Terraform
Ansible
GCP
New Relic
OpenTelemetry
Datadog
Prometheus
Azure
CI/CD
AWS
Docker
Linux
IoT
MQTT
Apply
$168k – $282k per year • Equity • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Python
JavaScript
TypeScript
AI/ML
LLM Guardrails
Agentic Workflows
DevOps
CI/CD
Management
ServiceNow
Apply
≈ $124k – $270k per year (Estimated) • Remote (Singapore) • Singapore
Python
JavaScript
AI/ML
AI Agents
DevOps
GitHub Actions
Azure
CI/CD
GitHub
IAM
Cybersecurity
Microsoft Sentinel
Least Privilege
Apply
≈ $128k – $272k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Toronto
Python
AI/ML
Model Context Protocol
AI Agents
Cerebras
OpenAI
Edge AI
DevOps
gRPC
etcd
Prometheus
Kubernetes
Grafana
eBPF
HPC
Linux
Apply
≈ $92k – $209k per year (Estimated) • Hybrid • Full-Time • Toronto
C++
AI/ML
AI Agents
Cerebras
OpenAI
Edge AI
DevOps
Linux
Cybersecurity
Wireshark
Apply
≈ $179k – $368k per year (Estimated) • In office • Full-Time
Python
C++
AI/ML
AI Agents
Cerebras
OpenAI
Edge AI
Apply
≈ $143k – $327k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Master's Degree • Sunnyvale • Toronto
AI/ML
AI Agents
Cerebras
OpenAI
Edge AI
DevOps
HPC
Cybersecurity
Wireshark
Apply
≈ $200k – $391k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree
AI/ML
AI Agents
Cerebras
OpenAI
Edge AI
DevOps
HPC
Linux
Apply
$81k – $136k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Toronto
AI/ML
JAX
Multimodal AI
PyTorch
TPU
AWS Trainium
Machine Learning
DevOps
AWS
Apply
$107k – $178k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
AI/ML
LLM
LLM Guardrails
Apply
Remote (United States, Canada) • Toronto
Apply
$200k – $280k per year • In office • 2+ years exp • Toronto
SQL
Analytics
Tableau
Power BI
Looker
Microsoft Excel
Management
Google Sheets
Apply
$280k – $800k per year • In office • Toronto
Apply
See all jobs
This is one of many
857,935 more open roles from verified company boards, updated every day.