{"id":1233203,"url":"https://alion.io/job/pointfive-ai-research-lead","title":"Head of LLM research","company":{"id":1920967,"name":"PointFive","domain":"pointfive.co","url":"https://alion.io/company/pointfive-co","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"Leadership","role_family":"Leadership","seniority":"head","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Tel Aviv, Israel"],"countries":["IL"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":143000,"max_usd":312000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":2568},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Fine-tuning","optional":false},{"name":"FinOps","optional":false},{"name":"Function Calling","optional":false},{"name":"Hallucination","optional":false},{"name":"Knowledge Distillation","optional":false},{"name":"KV Cache","optional":false},{"name":"Linux","optional":false},{"name":"Llama","optional":false},{"name":"llama.cpp","optional":false},{"name":"LLM","optional":false},{"name":"LoRA","optional":false},{"name":"MLX ML","optional":false},{"name":"Model Distillation","optional":false},{"name":"Ollama","optional":false},{"name":"ONNX Runtime","optional":false},{"name":"Python","optional":false},{"name":"Quantization","optional":false},{"name":"SGLang","optional":false},{"name":"Speculative Decoding","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"Tool Use","optional":false},{"name":"vLLM","optional":false},{"name":"Windows","optional":false},{"name":"Agentic Workflows","optional":true},{"name":"Anthropic","optional":true},{"name":"AWQ","optional":true},{"name":"AWS","optional":true},{"name":"Cloudflare","optional":true},{"name":"Computer Use","optional":true},{"name":"CUDA","optional":true},{"name":"CUDA Toolkit","optional":true},{"name":"DeepSeek","optional":true},{"name":"DPO","optional":true},{"name":"Gemma","optional":true},{"name":"GGUF","optional":true},{"name":"Go","optional":true},{"name":"GPTQ","optional":true},{"name":"GRPO","optional":true},{"name":"Hugging Face","optional":true},{"name":"Mistral","optional":true},{"name":"OpenAI","optional":true},{"name":"PEFT","optional":true},{"name":"Post-training","optional":true},{"name":"PyTorch","optional":true},{"name":"Qwen","optional":true},{"name":"Reinforcement Learning","optional":true},{"name":"ROCm","optional":true},{"name":"Snowflake","optional":true},{"name":"Synthetic Data","optional":true},{"name":"TensorRT","optional":true}],"status":"live","first_seen_at":"2026-09-17T12:29:05Z","employer_posted_date":"2026-09-17","last_verified_at":"2026-10-01T07:18:33Z","board_verified":true,"closed_at":null,"days_open":13,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":13},"description":"About PointFive\nPointFive is the AI Efficiency OS. From the cloud to the coding agent, we're the only platform that manages AI spend everywhere it happens.\nEngineering and FinOps teams use PointFive to make their organizations more efficient and their cloud and AI more effective. We don't just show what you spend. We show what you're wasting, and we fix it autonomously. NuBank saw ROI in 10 days. Customers average 1,200%+ ROI and a 4.9 rating on G2.\nFounded by the team behind IntSights (acquired by Rapid7), PointFive recently closed a $60M Series B led by Accel, with participation from Entrée Capital and Salesforce Ventures.\nAbout the Role\nPointFive is building infrastructure for the next generation of AI-powered engineering.\nAs AI agents become embedded into developer workflows, the underlying model layer is changing rapidly. Organizations are no longer choosing between a handful of hosted APIs. They are increasingly operating across frontier models, open-weight models, locally deployed models, specialized models, and dynamically routed combinations of them.\nWe’re looking for an AI researcher to take an integral part in PointFive’s research into how these models behave, how they should be evaluated, and how they can be deployed and used more efficiently across real engineering workloads.\nThis is a deeply technical research role with direct product impact.\nYou’ll study frontier and open-source models, develop evaluation frameworks, investigate inference and model optimization techniques, explore local deployment strategies, and help answer questions such as:\nWhich model should handle a particular task?\nHow much capability do we lose when we quantize it?\nCan a smaller local model replace a frontier API for a specific workload?\nHow should models be routed across latency, cost, privacy, and quality constraints?\nHow do we measure whether one model is actually better than another for agentic software engineering?\nYour work will directly shape PointFive’s model strategy and the intelligence behind our AI infrastructure products.\nWhat You'll Do\nDefine and lead PointFive’s LLM research agenda across model evaluation, inference optimization, local models, routing, and model adaptation.\n\nContinuously evaluate frontier and open-weight models across real-world software engineering and agentic workloads.\n\nDesign rigorous evaluation frameworks for model quality, reasoning, tool use, code generation, command execution, summarization, and agent behavior.\n\nBuild benchmarks that reflect actual developer workflows rather than generic academic tasks.\n\nResearch model efficiency techniques including quantization, distillation, speculative decoding, prefix caching, KV-cache optimization, batching, and context management.\n\nInvestigate when smaller or locally deployed models can replace expensive frontier models without materially degrading task quality.\n\nEvaluate model architectures, parameter sizes, quantization formats, and runtime configurations across heterogeneous hardware.\n\nResearch and benchmark local inference stacks including llama.cpp, MLX, vLLM, SGLang, ONNX Runtime, Ollama, TensorRT-LLM, and emerging runtimes.\n\nStudy inference performance across Apple Silicon, NVIDIA GPUs, AMD GPUs, CPUs, Windows workstations, Linux machines, and other endpoint configurations.\n\nDevelop model-routing strategies that optimize across quality, latency, cost, privacy, context size, and hardware availability.\n\nExplore intelligent cascades where smaller models handle common tasks and more capable models are invoked only when necessary.\n\nResearch model specialization through fine-tuning, LoRA, adapters, distillation, prompt optimization, and other model adaptation techniques.\n\nInvestigate model behavior under constrained environments, including offline execution, limited memory, limited compute, and local-only inference.\n\nEvaluate agent-specific model behavior, including planning, tool selection, shell interaction, code editing, error recovery, and long-running task execution.\n\nAnalyze failure modes such as hallucination, tool misuse, context degradation, reasoning collapse, excessive token consumption, and unstable agent loops.\n\nDesign experiments that quantify the tradeoffs between model capability, inference cost, token consumption, latency, and resource utilization.\n\nBuild internal research infrastructure for reproducible model experiments, benchmarking, dataset management, and evaluation.\n\nTrack frontier model releases and emerging research, rapidly determining which developments are meaningful for PointFive’s products.\n\nCollaborate closely with engineering and product teams to translate research results into production capabilities.\n\nBuild and lead a small, exceptional LLM research team over time.\n\nWhat We're Looking For\nMust-have\nDeep understanding of modern large language models and transformer-based architectures.\n\nStrong hands-on experience evaluating and experimenting with both frontier and open-weight models.\n\nStrong understanding of inference behavior, including prefill, decoding, KV caches, context windows, batching, memory usage, and token generation performance.\n\nExperience with model optimization techniques such as quantization, distillation, LoRA, fine-tuning, or model compression.\n\nStrong experimental mindset - able to formulate hypotheses, design controlled experiments, and draw meaningful conclusions from noisy results.\n\nExperience building evaluation frameworks for LLM quality and behavior.\n\nStrong Python proficiency and familiarity with the modern ML ecosystem.\n\nAbility to read, understand, and reproduce ideas from current ML research papers.\n\nComfortable working with ambiguous research problems where there may not yet be an established best practice.\n\nStrong ability to bridge research and production - understanding not only whether something works, but whether it is practical to deploy.\n\nNice to have\nExperience with open-weight models such as Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, or similar model families.\n\nExperience with frontier model APIs including OpenAI, Anthropic, Google, and other leading providers.\n\nExperience with inference frameworks such as vLLM, SGLang, llama.cpp, MLX, TensorRT-LLM, Ollama, or ONNX Runtime.\n\nDeep understanding of quantization techniques including FP8, INT8, INT4, AWQ, GPTQ, GGUF, and related approaches.\n\nExperience with GPU performance optimization, CUDA, Metal, ROCm, or DirectML.\n\nFamiliarity with distributed inference and multi-GPU serving.\n\nExperience with model routing, mixture-of-model systems, cascades, or adaptive inference.\n\nExperience with reinforcement learning, preference optimization, DPO, GRPO, or related post-training techniques.\n\nFamiliarity with agentic systems, coding agents, tool-use models, and computer-use models.\n\nExperience building or evaluating coding benchmarks and software-engineering agents.\n\nExperience with synthetic data generation, dataset curation, and automatic evaluation.\n\nResearch publications or meaningful contributions to open-source ML projects.\n\nExperience leading a small applied research or ML research team.\n\nResearch Areas\nSome of the problems we expect this team to work on include:\nModel Routing Choosing the optimal model dynamically based on task difficulty, latency requirements, cost, privacy constraints, and hardware availability.\n\nLocal vs. Cloud Inference Determining which workloads can reliably move from cloud models to models running directly on developer endpoints.\n\nModel Compression Understanding how far models can be quantized, distilled, or otherwise optimized before meaningful capability is lost.\n\nAgentic Model Evaluation Building evaluation methods for agents that operate over codebases, shells, developer tools, and long-running workflows.\n\nInference Efficiency Improving throughput, latency, memory consumption, and token efficiency across different model architectures and runtimes.\n\nModel Specialization Investigating whether smaller specialized models can outperform general-purpose frontier models for narrow engineering tasks.\n\nLong-Context Behavior Understanding how models behave as context grows, what information gets lost, and how context can be compressed or structured more intelligently.\n\nHardware-Aware AI Matching models and inference strategies to available GPUs, CPUs, NPUs, Apple Silicon, and other endpoint hardware.\n\nCost vs. Intelligence Tradeoffs Quantifying when additional model capability actually produces better outcomes - and when it simply produces more expensive tokens.\n\nOur Tech Stack\nPython, Go, PyTorch, Hugging Face, vLLM, SGLang, llama.cpp, MLX, CUDA, Metal, AWS, Cloudflare, Snowflake, and a rapidly evolving ecosystem of frontier and open-weight models.\nWhy PointFive\nA rare research surface\n\nWe operate where LLM research meets real enterprise infrastructure. The questions we work on have immediate implications for how thousands of engineers use AI every day.\nAccess to real workloads\n\nInstead of optimizing against abstract benchmarks, you'll be able to study how models perform across real software engineering and agentic workflows.\nResearch that ships\n\nThis isn't a research lab disconnected from product. Successful ideas can move rapidly from experiment to production.\nModels are becoming infrastructure\n\nEnterprises will increasingly operate fleets of models across cloud APIs, private infrastructure, and developer endpoints. Deciding how those models are selected, optimized, deployed, and governed is becoming a fundamental infrastructure problem.\nFrontier moves fast\n\nNew models, architectures, inference techniques, and agent systems appear constantly. Your job is to understand which developments matter - and turn them into an advantage for PointFive.\nFounders with a track record\n\nBuilt and sold IntSights to Rapid7. Backed by top-tier investors with deep conviction in the category.\nEarly-stage leverage\n\nYou'll define PointFive's LLM research strategy, research methodology, and eventually the team itself.\nEqual Opportunity Statement\nPointFive is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We welcome candidates from all backgrounds, experiences, and perspectives to apply.","description_format":"text","description_chars":10253,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Israel","iso":"IL","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Cloud Cost & Multi-Cloud Management"],"lifecycle":[{"event":"open","at":"2026-09-25T15:33:06Z"}],"liveness":{"score":75,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.836,"p_room":0.9,"age_days":13,"expected_fill_days":30,"reasons":["conf:8","velocity","win:mid"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/pointfive-ai-research-lead","json_url":"https://alion.io/job/pointfive-ai-research-lead.json","meta":{"generated_at":"2026-10-01T09:59:48Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":955,"day_limit":5000,"remaining_today":4045,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}