Overview
News
Technologies
Salaries
Products
People
Growth
Offices
Jobs
Financials

Overview

Sequential generation is the bottleneck. Kog couples a low-latency engine with parallel architecture to deliver 30x faster LLM inference. Request API access.

News

News Kog Labs 3 months ago
Kog Laneformer 2B: The Latency-First Model Behind Kog Inference Engine
Today Kog is releasing the weights and model code of Laneformer 2B on Hugging Face Hub, the 2.3B-parameter instruction-tuned coding model designed for high-speed decoding.Most LLM research optimizes for benchmark quality first, and inference metrics like
Read more
Report
News Kog Labs 4 months ago
Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)
Today, Kog AI launches a tech preview of the Kog Inference Engine (KIE): 3,000 output tokens/s per request on 8 AMD MI300X GPUs and 2,100 on 8 NVIDIA H200 (FP16, no speculative decoding). This preview runs a 2B model, with support for large third-party Mo
Read more
Report
News Kog Labs 4 months ago
Delayed Tensor Parallelism for Faster Transformer Inference
DTP is a new Transformer architecture that hides communication overhead behind computation and weight streaming, enabling significantly faster batch-size-one inference on AMD and NVIDIA GPUs.
Read more
Report
News AMD 1 year ago
Kog Reaches 3.5 Breakthrough Inference Speed on AMD Instinct MI300X
Kog AI breaks conventions with a radical approach to inference, free from legacy constraints, to address real market pain points
Read more
Report

Technologies

Sign in to see how your skills match this stack
Sign In
Tech DNA
CUDA
PyTorch
DeepSeek
LLM
Qwen
AI/ML
AI Agents
DeepSeek
LLM
Mixture of Experts
Qwen
Speculative Decoding
Fine-tuning
Post-training
Quantization
Transformers
CUDA
CUDA Toolkit
PyTorch
Stack modernity Cutting-edge AI adopter
84/100
How modern this stack is, based on technology relevance, AI adoption and the share of legacy tools.
In-demand technologies
Qwen 2 jobs Required
DeepSeek 2 jobs Required
AI Agents 2 jobs Required
LLM 2 jobs Required
Mixture of Experts 2 jobs Required
Speculative Decoding 2 jobs Required
Fine-tuning 1 job Required
Quantization 1 job Required
Salary medians are calculated from this company's open jobs and compared with the market.
Industry adoption
AI Agents
8%
LLM
7%
Fine-tuning
3%
PyTorch
3%
CUDA Toolkit
2%
Quantization
1%
Share of companies in the same industry that use each technology.
Stack changes
PyTorch Sep 2026
CUDA Toolkit Sep 2026
CUDA Sep 2026
Transformers Sep 2026
Speculative Decoding Sep 2026
Qwen Sep 2026
Quantization Sep 2026
Post-training Sep 2026
Technologies recently added to or removed from this company's stack - a signal of tech migrations and new initiatives.

Growth

Hiring Momentum
58/100
Growing
Open positions
2
0 opened / 0 closed in 30 days
ATS activity
Every ~1 hours
09/06/2026
Hiring Dynamics +100%
Hiring Focus
The percentage next to each role is its share of the company's job openings over the last 90 days; the arrow shows the shift versus the previous period.
AI/ML
100%
Activity Timeline
Added PyTorch to stack Sep 2026
Added CUDA Toolkit to stack Sep 2026
Added CUDA to stack Sep 2026
Added Transformers to stack Sep 2026
Added Speculative Decoding to stack Sep 2026
Added Qwen to stack Sep 2026
Added Quantization to stack Sep 2026
Added Post-training to stack Sep 2026
Added Mixture of Experts to stack Sep 2026
Added LLM to stack Sep 2026
Added Fine-tuning to stack Sep 2026
Added DeepSeek to stack Sep 2026

Offices

Where the company hires and what each office is for
Cities
1
Hiring now
1
Countries
1
Busiest office
Paris
Paris France
Engineering / R&D
2 open roles 2 posted in 90 days 100% engineering
hires for AI/ML
seen since Jun 2026 job postings

Jobs

Research Engineer

3 months ago
$74k – $163k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
Qwen
DeepSeek
Fine-tuning
Quantization
AI Agents
Transformers
LLM
Mixture of Experts
Post-training
Speculative Decoding
Apply

GPU Engineer

3 months ago
$75k – $164k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Paris
Qwen
DeepSeek
CUDA Toolkit
AI Agents
PyTorch
LLM
Mixture of Experts
CUDA
Speculative Decoding
Apply