{"id":1468377,"url":"https://alion.io/job/kog-labs-agentic-compiler-engineer","title":"Agentic Compiler Engineer","company":{"id":185339,"name":"Kog Labs","domain":"kog.ai","url":"https://alion.io/company/kog-labs","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Paris, France"],"countries":["FR"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":62000,"max_usd":148000,"period":"year","method":"role_country_seniority_unknown","sample_n":41},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"LLM","optional":false},{"name":"Mixture of Experts","optional":false},{"name":"MLIR","optional":false},{"name":"Quantization","optional":false},{"name":"Speculative Decoding","optional":false},{"name":"Transformers","optional":false}],"status":"live","first_seen_at":"2026-09-29T13:32:29Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-01T06:48:11Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"ABOUT KOG\nKog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels.\nOn the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming.\nOn the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs.\nKog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.\nOur next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference.\nThe team has 10 people, including 9 engineers and researchers and 4 PhDs.\nTest it at playground.kog.ai. Read the technical details on the Kog Labs blog.\nWHAT YOU WILL WORK ON\nYou will work directly on AGCO.\nThe goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization.\nYou will contribute to areas such as:\nCompiler and IR design for representing and transforming LLM computations.\n\nOptimization passes, lowering, and code generation.\n\nSearch methods for exploring different implementations and execution strategies.\n\nVerification and correctness checks for generated changes.\n\nGPU execution, profiling, and performance optimization.\n\nLLM inference across operators, memory, parallelism, and communication.\n\nOptimization loops that connect generated changes to measurements on real GPUs.\n\nOne direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer. We plan to start with focused problems, build working prototypes, and extend the system from what we learn.\nYour main area will depend on your experience, skills, and interests. You may focus more on compilers, GPU systems, or LLM inference while working closely with people across the full stack.\nWHAT WE LOOK FOR\nWe look for engineers with deep technical expertise and original work in at least one area relevant to AGCO.\nRelevant experience includes:\nCompiler engineering, including optimization passes, IRs, lowering, code generation, LLVM, or MLIR.\n\nGPU programming with CUDA, HIP, Metal, Vulkan, or similar technologies.\n\nGPU performance work involving kernels, memory, synchronization, profiling, or hardware behavior.\n\nLLM inference engines and performance optimization.\n\nAttention, MoE, parallelism, communication, or other systems-level parts of LLM execution.\n\nFormal verification, equivalence checking, SAT/SMT, or related methods.\n\nSystems that generate, search, test, benchmark, or optimize code automatically.\n\nWe care about what you personally built and the technical decisions behind it. Strong candidates can explain the problem, their approach, the alternatives they explored, and how they measured the result.\nWe review technical work during the process. This can be public code, an upstream contribution, a paper, a thesis, a technical project, or a detailed write-up based on work you can share.\nWHAT WE OFFER\nYou will join a small team building AGCO as a core part of Kog's technology.\nWork at the intersection of compilers, GPU systems, and LLM inference.\n\nDirect access to engineers working across the full inference stack.\n\nA fast loop from an optimization idea to compilation, execution, verification, and measurement on real GPUs.\n\nThe opportunity to go deep in your strongest technical area while expanding into the other parts of the stack.\n\nHigh ownership over technical decisions and systems that will shape how Kog optimizes LLM inference.\n\nThis role is based in Paris, and we are looking for candidates who can relocate to Paris and work closely with the team.","description_format":"text","description_chars":3989,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"France","iso":"FR","kind":"country"}],"hiring_excludes":[],"relocation_offered":true,"industries":["Artificial Intelligence","LLM & Generative AI"],"lifecycle":[{"event":"open","at":"2026-09-29T15:31:15Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":21,"reasons":["conf:11","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/kog-labs-agentic-compiler-engineer","json_url":"https://alion.io/job/kog-labs-agentic-compiler-engineer.json","meta":{"generated_at":"2026-10-01T10:21:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1466,"day_limit":5000,"remaining_today":3534,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}