Overview
News
Technologies
Salaries
Products
People
Growth
Offices
Jobs
Financials
Overview
Sequential generation is the bottleneck. Kog couples a low-latency engine with parallel architecture to deliver 30x faster LLM inference. Request API access.
News
Kog Laneformer 2B: The Latency-First Model Behind Kog Inference Engine
Today Kog is releasing the weights and model code of Laneformer 2B on Hugging Face Hub, the 2.3B-parameter instruction-tuned coding model designed for high-speed decoding.Most LLM research optimizes for benchmark quality first, and inference metrics like
Read more
Report
Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)
Today, Kog AI launches a tech preview of the Kog Inference Engine (KIE): 3,000 output tokens/s per request on 8 AMD MI300X GPUs and 2,100 on 8 NVIDIA H200 (FP16, no speculative decoding). This preview runs a 2B model, with support for large third-party Mo
Read more
Report
Delayed Tensor Parallelism for Faster Transformer Inference
DTP is a new Transformer architecture that hides communication overhead behind computation and weight streaming, enabling significantly faster batch-size-one inference on AMD and NVIDIA GPUs.
Read more
Report
Kog Reaches 3.5 Breakthrough Inference Speed on AMD Instinct MI300X
Kog AI breaks conventions with a radical approach to inference, free from legacy constraints, to address real market pain points
Read more
Report
Technologies
Tech DNA
CUDA
PyTorch
DeepSeek
LLM
Qwen
AI/ML
AI Agents
DeepSeek
LLM
Mixture of Experts
Qwen
Speculative Decoding
Fine-tuning
Post-training
Quantization
Transformers
CUDA
CUDA Toolkit
PyTorch
Stack modernity
84/100
How modern this stack is, based on technology relevance, AI adoption and the share of legacy tools.
In-demand technologies
Qwen
2 jobs
Required
DeepSeek
2 jobs
Required
AI Agents
2 jobs
Required
LLM
2 jobs
Required
Mixture of Experts
2 jobs
Required
Speculative Decoding
2 jobs
Required
Fine-tuning
1 job
Required
Quantization
1 job
Required
Salary medians are calculated from this company's open jobs and compared with the market.
Industry adoption
AI Agents
8%
LLM
7%
Fine-tuning
3%
PyTorch
3%
CUDA Toolkit
2%
Quantization
1%
Share of companies in the same industry that use each technology.
Stack changes
PyTorch
Sep 2026
CUDA Toolkit
Sep 2026
CUDA
Sep 2026
Transformers
Sep 2026
Speculative Decoding
Sep 2026
Qwen
Sep 2026
Quantization
Sep 2026
Post-training
Sep 2026
Technologies recently added to or removed from this company's stack - a signal of tech migrations and new initiatives.
Growth
Hiring Momentum
58/100
Growing
Open positions
2
0 opened / 0 closed in 30 days
ATS activity
Every ~1 hours
09/06/2026
Hiring Dynamics
+100%
Hiring Focus
The percentage next to each role is its share of the company's job openings over the last 90 days; the arrow shows the shift versus the previous period.
AI/ML
100% ▼
Activity Timeline
Added PyTorch to stack
Sep 2026
Added CUDA Toolkit to stack
Sep 2026
Added CUDA to stack
Sep 2026
Added Transformers to stack
Sep 2026
Added Speculative Decoding to stack
Sep 2026
Added Qwen to stack
Sep 2026
Added Quantization to stack
Sep 2026
Added Post-training to stack
Sep 2026
Added Mixture of Experts to stack
Sep 2026
Added LLM to stack
Sep 2026
Added Fine-tuning to stack
Sep 2026
Added DeepSeek to stack
Sep 2026
Offices
Where the company hires and what each office is for
Cities
1
Hiring now
1
Countries
1
Busiest office
Paris
Paris
France
Engineering / R&D
2 open roles 2 posted in 90 days 100% engineering
hires for AI/ML
Jobs
Research Engineer
3 months ago≈ $74k – $163k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
Qwen
DeepSeek
Fine-tuning
Quantization
AI Agents
Transformers
LLM
Mixture of Experts
Post-training
Speculative Decoding
Apply
GPU Engineer
3 months ago≈ $75k – $164k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
Qwen
DeepSeek
CUDA Toolkit
AI Agents
PyTorch
LLM
Mixture of Experts
CUDA
Speculative Decoding
Apply

NVIDIA
Cerebras Systems
Baseten
SambaNova Systems
Lila Sciences
Databricks