974,315open jobs
58,640companies
159,915added this week
Browse all
Salary
≈ $62k – $148k per year (Estimated)
Location
Hybrid (Paris, France)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 30, 2026. First seen by Alion on Sep 29, 2026.

Overview
Company
Impact
Profile match
Sequential generation is the bottleneck. Kog couples a low-latency engine with parallel architecture to deliver 30x faster LLM inference. Request API access.

ABOUT KOG

Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels.

On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming.

On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs.

Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.

Our next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference.

The team has 10 people, including 9 engineers and researchers and 4 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

WHAT YOU WILL WORK ON

You will work directly on AGCO.

The goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization.

You will contribute to areas such as:

  • Compiler and IR design for representing and transforming LLM computations.

  • Optimization passes, lowering, and code generation.

  • Search methods for exploring different implementations and execution strategies.

  • Verification and correctness checks for generated changes.

  • GPU execution, profiling, and performance optimization.

  • LLM inference across operators, memory, parallelism, and communication.

  • Optimization loops that connect generated changes to measurements on real GPUs.

One direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer. We plan to start with focused problems, build working prototypes, and extend the system from what we learn.

Your main area will depend on your experience, skills, and interests. You may focus more on compilers, GPU systems, or LLM inference while working closely with people across the full stack.

WHAT WE LOOK FOR

We look for engineers with deep technical expertise and original work in at least one area relevant to AGCO.

Relevant experience includes:

  • Compiler engineering, including optimization passes, IRs, lowering, code generation, LLVM, or MLIR.

  • GPU programming with CUDA, HIP, Metal, Vulkan, or similar technologies.

  • GPU performance work involving kernels, memory, synchronization, profiling, or hardware behavior.

  • LLM inference engines and performance optimization.

  • Attention, MoE, parallelism, communication, or other systems-level parts of LLM execution.

  • Formal verification, equivalence checking, SAT/SMT, or related methods.

  • Systems that generate, search, test, benchmark, or optimize code automatically.

We care about what you personally built and the technical decisions behind it. Strong candidates can explain the problem, their approach, the alternatives they explored, and how they measured the result.

We review technical work during the process. This can be public code, an upstream contribution, a paper, a thesis, a technical project, or a detailed write-up based on work you can share.

WHAT WE OFFER

You will join a small team building AGCO as a core part of Kog's technology.

  • Work at the intersection of compilers, GPU systems, and LLM inference.

  • Direct access to engineers working across the full inference stack.

  • A fast loop from an optimization idea to compilation, execution, verification, and measurement on real GPUs.

  • The opportunity to go deep in your strongest technical area while expanding into the other parts of the stack.

  • High ownership over technical decisions and systems that will shape how Kog optimizes LLM inference.

This role is based in Paris, and we are looking for candidates who can relocate to Paris and work closely with the team.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
974,315 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Paris
≈ $88k – $209k per year (Estimated) • Equity • Hybrid • 4+ years exp • Bachelor's Degree • Paris
Python
SQL
AI/ML
Cursor
Claude Code
Prompt Engineering
AI Agents
OpenAI Codex
Context Engineering
Robotics
Digital Twin
Apply
≈ $53k – $127k per year (Estimated) • In office • Full-Time • Courbevoie
Python
SQL
AI/ML
Machine Learning
DevOps
CI/CD
Docker
Management
Jira
Apply
≈ $52k – $99k per year (Estimated) • In office • Full-Time • 5+ years exp • Nantes
Python
AI/ML
LangChain
Computer Vision
NLP
Mistral
TensorFlow
PyTorch
LLM
OpenAI
Hugging Face
Apply
≈ $79k – $226k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Paris
Python
TypeScript
Python
FastAPI
Celery
Databases
PostgreSQL
Redis
pgvector
AI/ML
LangGraph
LangChain
Docling
Langfuse
RAG
OCR
DevOps
Azure DevOps
Azure
CI/CD
Docker
Kubernetes
GitHub
Apply
≈ $47k – $135k per year (Estimated) • In office • Full-Time • 4+ years exp • Toulouse
Python
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
Mistral
LLM
RAG
Reranking
LLMOps
Tool Use
Machine Learning
DevOps
GCP
Azure DevOps
GitLab CI
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Management
n8n
Apply
≈ $109k – $214k per year (Estimated) • In office • 6+ years exp • Master's Degree • Chicago
Python
Databases
Databricks
AI/ML
Spark
LLM
Machine Learning
DevOps
Azure
AWS
Analytics
A/B Testing
Apply
$137k – $167k per year • In office • Secret • Full-Time • 3+ years exp • Bachelor's Degree • Redwood City
Python
JavaScript
AI/ML
LLM
Apply
$148k – $179k per year • In office • Top Secret • Full-Time • Bachelor's Degree • Tysons
Python
JavaScript
AI/ML
LLM
Apply
$150k – $190k per year • In office • TS/SCI • Full-Time • Tysons
Python
AI/ML
Reinforcement Learning
Computer Vision
LLM
Machine Learning
DevOps
GCP
Azure
AWS
Management
Agile
Apply
$120k – $150k per year • In office • Secret • Full-Time • Tysons
Python
AI/ML
Reinforcement Learning
Computer Vision
LLM
Machine Learning
DevOps
GCP
Azure
AWS
Management
Agile
Apply
Research Engineer 3 months ago
≈ $63k – $151k per year (Estimated) • Hybrid • Full-Time • Paris
AI/ML
Quantization
AI Agents
Transformers
Mixture of Experts
Post-training
Speculative Decoding
Apply
≈ $19k – $48k per year (Estimated) • Remote (Europe, Mexico) • Full-Time • 2+ years exp • Bachelor's Degree • Sofia • Budapest • Paris • Belgrade • Buenos Aires
Analytics
Microsoft Excel
Management
Outlook
Apply
In office • Paris
Apply
≈ $108k – $222k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Paris
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
JAX
AI Agents
TensorFlow
PyTorch
TPU
Edge AI
MLIR
Machine Learning
Apply
In office • Internship • Paris
Apply
In office • Internship • Paris
Apply
See all jobs
This is one of many
974,315 more open roles from verified company boards, updated every day.