{"id":1025912,"url":"https://alion.io/job/furiosa-software-software-engineer-compiler-kernel-optimization","title":"Software - Software Engineer, Compiler (Kernel Optimization)","company":{"id":31683,"name":"Furiosa","domain":"furiosa.ai","url":"https://alion.io/company/furiosa","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":null},"role":"Backend","role_family":"Backend","seniority":null,"employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Seoul, South Korea"],"countries":["KR"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"LLM","optional":false},{"name":"TPU","optional":true},{"name":"Triton","optional":true}],"status":"live","first_seen_at":"2026-09-18T10:26:08Z","employer_posted_date":"2026-09-18","last_verified_at":"2026-09-28T22:37:23Z","board_verified":true,"closed_at":null,"days_open":10,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":10},"description":"About FuriosaAI\nFuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon. \nOur vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.\nAbout the Job\nThe compiler plays a central role in FuriosaAI's mission to build high-performance, energy-efficient AI systems. Modern deep learning models are evolving rapidly and becoming increasingly diverse, making compilation a challenging problem. Transforming these models into efficient executable programs requires careful reasoning about complex transformations while preserving program meaning and structure.\nIn this role, you will own the performance of critical AI kernels integrated into Furiosa-LLM, FuriosaAI's software stack for large language model serving. You will analyze end-to-end serving workloads, identify performance-critical bottlenecks, and implement highly optimized kernels using TCL (Tensor Contraction Language), FuriosaAI's programming language for kernel optimization. Because real-world serving workloads are inherently dynamic, you will develop scheduling, specialization, and algorithmic techniques that map them efficiently onto compiler abstractions and execution environments optimized for static workloads.\nResponsibilities\nAnalyze end-to-end LLM serving workloads and identify performance bottlenecks that can be addressed through kernel-level optimization.\nDesign, implement, and optimize high-performance kernels in TCL for critical operations in Furiosa-LLM.\nDevelop algorithmic techniques for efficiently supporting dynamic serving workloads.\nIntegrate, benchmark, and validate optimized kernels across representative models, input shapes, and serving scenarios.\nCollaborate with compiler and serving teams to improve compiler capabilities and ensure optimized kernels work effectively within the production software stack.\nMinimum Qualifications\nBS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.\nExperience in developing low-level or performance-critical software.\nExperience in analyzing performance bottlenecks using profiling, benchmarking, and hardware performance characteristics.\nUnderstanding of parallel computation, memory hierarchies, and data movement on modern architectures.\nPreferred Qualifications\nMS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.\nExperience in optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU) for AI products.\nHands-on knowledge of kernel optimization techniques such as tiling, computation scheduling, operator fusion, memory layout transformation, vectorization, pipelining, and data movement optimization.\nExperience reasoning about trade-offs among parallelism, memory bandwidth, on-chip memory capacity, compute utilization, and synchronization overhead.\nExperience with accelerator programming or domain-specific kernel languages such as Triton, cuTile, or Pallas.\nWhy Join FuriosaAI\nThe defining bottleneck of the AI era is building the right hardware and software stack to run it at global scale. Furiosa is solving this challenge holistically from the ground up.\nWith our flagship chip, RNGD, in mass production today and our next-generation platform in development with Broadcom, we are proving that full-stack, tensor-native compute is the future of AI infrastructure. This is a pivotal moment to join our team, right as we accelerate our global expansion.\nAt Furiosa, you will:\nSolve AI’s Most Urgent Challenge. Help build the high-performance, energy-efficient inference hardware and software required to fulfill the promise of advanced AI.\nPioneer Full-Stack Co-Design. Work with teams that are architecting solutions from silicon up through the compiler (featuring innovations like Tensor Contraction Language and Virtual ISA) and serving frameworks.\nShip Real-World Silicon, Software, and Solutions. Turn breakthrough technology into commercial deployment. RNGD is in mass production with TSMC and running live enterprise workloads for global leaders like LG AI Research and Samsung SDS.\nPartner With the Industry's Best. Collaborate across an elite global ecosystem that includes TSMC, Broadcom, SK Hynix, and GUC.\nDo Your Life’s Best Work. Join a brilliant, low-ego, mission-driven team in a high-trust environment that values autonomy, intellectual curiosity, and shared ambition. \nContact\n ","description_format":"text","description_chars":4873,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips","AI Chips & Accelerators"],"lifecycle":[{"event":"open","at":"2026-09-18T11:10:58Z"}],"liveness":{"score":79,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.838,"p_room":0.945,"age_days":9,"expected_fill_days":19,"reasons":["conf:3","velocity","win:mid"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/furiosa-software-software-engineer-compiler-kernel-optimization","json_url":"https://alion.io/job/furiosa-software-software-engineer-compiler-kernel-optimization.json","meta":{"generated_at":"2026-09-29T01:54:09Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1578,"day_limit":5000,"remaining_today":3422,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}