428,564open jobs
14,652companies
61,154added this week
Browse all
Salary
$74k – $163k per year (Estimated)
Location
Remote/Hybrid (Paris, France)
Employment
Full-Time
Overview
Company
Impact
Profile match
Sequential generation is the bottleneck. Kog couples a low-latency engine with parallel architecture to deliver 30x faster LLM inference. Request API access.

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it.

We pre-trained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will imagine, design, and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

  • Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first-order design input.

  • Extend the Laneformer thesis by exploring inference-aware architectural variants such as DTP, Ladder Residual, and PT-Transformer, and finding what compounds at scale.

  • Own the post-training pipeline across fine-tuning, evaluation methodology, and adaptation of existing open-weight models toward architecture variants optimized for inference speed.

  • Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time.

  • Write up findings as research papers, submit them to top venues, and present them at conferences.

  • Contribute to building AI agents that will perform architecture research and training experiments autonomously, starting from the research foundations we are building now.

What we look for

  • You have designed or changed model architecture, where the structure itself was the object of the work. Showing that work, a paper, a repository, or a thesis, is a requirement to move forward.

  • You reason about model design and hardware together, tracing how communication structure and layer dependencies shape inference behavior, with fluency in Transformers and MoE deep enough to weigh trade-offs.

  • Stronger signals include inference-aware architectural variants such as DTP, Ladder Residual, or PT-Transformer, and post-training methods such as fine-tuning, preference optimization, or quantization, including at research scale.

  • A top engineering school or a PhD with concrete architecture work counts, even without industry experience.

What we offer

  • Direct access to AMD and NVIDIA datacenter GPUs from day one

  • A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions

  • Problems that sit on the critical path of model execution speed and that directly influence what the system can become

  • A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company.

  • Compensation aligned with top AI research profiles, including equity

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,564 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Paris
$66k – $137k per year (Estimated) • Equity • Remote/Hybrid • 8+ years exp • Tokyo
AI/ML
AI Agents
DevOps
SLI/SLO/SLA
Robotics
Digital Twin
Apply
Copilot AI Engineer 2 days ago
Remote/Hybrid • 4+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
Senior AI Engineer 2 days ago
Remote/Hybrid • 9+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
$126k – $265k per year (Estimated) • Remote • 8+ years exp
Python
SQL
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
Remote/Hybrid • 8+ years exp
Python
SQL
C#
C#
.NET
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
GPU Engineer 3 months ago
$75k – $164k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
AI/ML
Qwen
DeepSeek
CUDA Toolkit
AI Agents
PyTorch
LLM
Mixture of Experts
CUDA
Speculative Decoding
Apply
$90k – $136k per year • Equity • In office • Full-Time • 8+ years exp • Paris • London
AI/ML
AI Agents
Agentic Workflows
Apply
Apply
$56k – $114k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Paris
Databases
Databricks
AI/ML
AI Agents
Edge AI
Vision-Language-Action
Analytics
Tableau
Power BI
Management
Power Apps
Apply
$52k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Paris
Apply
$44k – $119k per year (Estimated) • Remote/Hybrid • Full-Time • Paris
Apply
See all jobs
This is one of many
428,564 more open roles from verified company boards, updated every day.