428,624open jobs
14,649companies
61,350added this week
Browse all
Salary
$75k – $164k per year (Estimated)
Location
Remote/Hybrid (Paris, France)
Employment
Full-Time
Overview
Company
Impact
Profile match
Sequential generation is the bottleneck. Kog couples a low-latency engine with parallel architecture to deliver 30x faster LLM inference. Request API access.

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

The hot path is a monokernel implemented with handwritten CUDA (with PTX inline assembly) on NVIDIA, and HIP (with CDNA ISA inline assembly) on AMD.

We optimize at the low level with engine/kernel/model co-design, using reverse engineering to understand and exploit the details of how the GPU hardware works at the micro level.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will perform experiments to understand GPU internals, find creative solutions to accelerate critical computational sections used in LLM inference, and write optimized GPU kernels accordingly. Then test, profile, and optimize again.

  • Contribute to our monokernel pipeline, the single persistent GPU program that covers the full decode pass from QKV projection to LM head sampling, across AMD and NVIDIA architectures.

  • Work on low-level GPU optimization, including impossibly-fast grid synchronizations and inter-GPU collectives, and optimized GEMV, GEMM, and attention kernels across batch sizes and context lengths, with the memory-bandwidth-bound batch-1 GEMV regime as the primary target.

  • Build profiling infrastructure inside a monokernel, including custom instrumentation, device-timestamp frameworks, and per-stage analysis to translate machine behavior into concrete engineering decisions.

  • Scale the stack to third-party MoE models such as DeepSeek v4 and Qwen 3 to push generation speed on the models that matter in production today.

  • Contribute to building AI agents that will perform GPU Engineering research and kernel optimization autonomously, calibrated to hardware target and workload, starting from the inference foundations we are building now.

What we look for

  • You have written GPU kernels where performance was the central constraint. Showing the code is a requirement to move forward in the process.

  • PyTorch custom ops are an acceptable starting point if the kernels show a genuine understanding of the hardware below the framework level.

  • Stronger signals include inline PTX or CDNA ISA in public repositories, experience with latency-sensitive execution paths, understanding of why MBU matters more than MFU at batch size 1, and a background in inference engine components.

  • A top engineering school or a PhD with concrete GPU work counts, even without industry experience.

What we offer

  • Direct access to AMD and NVIDIA datacenter GPUs from day one

  • A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions

  • Problems that sit on the critical path of model execution speed and that directly influence what the system can become

  • A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company.

  • Compensation aligned with top technical profiles in the Paris GPU Engineering market, including equity

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,624 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Paris
$156k – $234k per year • Remote/Hybrid • Full-Time • 10+ years exp • Irving • Jacksonville
Python
Python
Flask
FastAPI
Databases
Chroma
Milvus
Pinecone
AI/ML
LangChain
Vertex AI
Gemma
Fine-tuning
Prompt Engineering
AI Agents
NeMo Guardrails
Llama
Mistral
Pandas
NumPy
PyTorch
LLM
RAG
Google ADK
Hallucination
Hugging Face
NVIDIA NeMo
LLM Guardrails
Edge AI
Agentic Workflows
DevOps
OpenShift
CI/CD
Docker
Kubernetes
Vector
Apply
$191k – $334k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
Databases
Apache Kafka
AI/ML
Fine-tuning
Quantization
Prompt Engineering
AI Agents
TensorFlow
PyTorch
RAG
Anomaly Detection
Feature Store
LLM Guardrails
Cybersecurity
Zero Trust
Management
ServiceNow
Apply
$270k – $345k per year • Remote/Hybrid • Bachelor's Degree • San Francisco
SQL
Apex
Apex
Lightning Web Components
MuleSoft
AI/ML
Copilot
Cursor
Claude
Multimodal AI
AI Agents
Anthropic
Interpretability
Management
Stripe
Apply
$200k – $265k per year • Remote/Hybrid • 6+ years exp • Bachelor's Degree • San Francisco
AI/ML
Claude
Claude Code
Multimodal AI
AI Agents
Anthropic
Interpretability
Cybersecurity
HIPAA
Web3
Rollup
Apply
$120k – $130k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
AI/ML
AI Agents
Apply
Research Engineer 3 months ago
$74k – $163k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
AI/ML
Qwen
DeepSeek
Fine-tuning
Quantization
AI Agents
Transformers
LLM
Mixture of Experts
Post-training
Speculative Decoding
Apply
$90k – $136k per year • Equity • In office • Full-Time • 8+ years exp • Paris • London
AI/ML
AI Agents
Agentic Workflows
Apply
Apply
$56k – $114k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Paris
Databases
Databricks
AI/ML
AI Agents
Edge AI
Vision-Language-Action
Analytics
Tableau
Power BI
Management
Power Apps
Apply
$52k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Paris
Apply
$44k – $119k per year (Estimated) • Remote/Hybrid • Full-Time • Paris
Apply
See all jobs
This is one of many
428,624 more open roles from verified company boards, updated every day.