908,794open jobs
55,875companies
151,787added this week
Browse all
Salary
≈ $143k – $312k per year (Estimated)
Location
Hybrid (Tel Aviv, Israel)
Seniority
Architect
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 17, 2026.

Overview
Company
Impact
Profile match
PointFive makes enterprises more efficient across cloud, data, AI, and agents.

About PointFive

PointFive is the AI Efficiency OS. From the cloud to the coding agent, we're the only platform that manages AI spend everywhere it happens.

Engineering and FinOps teams use PointFive to make their organizations more efficient and their cloud and AI more effective. We don't just show what you spend. We show what you're wasting, and we fix it autonomously. NuBank saw ROI in 10 days. Customers average 1,200%+ ROI and a 4.9 rating on G2.

Founded by the team behind IntSights (acquired by Rapid7), PointFive recently closed a $60M Series B led by Accel, with participation from Entrée Capital and Salesforce Ventures.

About the Role

PointFive is building infrastructure for the next generation of AI-powered engineering.

As AI agents become embedded into developer workflows, the underlying model layer is changing rapidly. Organizations are no longer choosing between a handful of hosted APIs. They are increasingly operating across frontier models, open-weight models, locally deployed models, specialized models, and dynamically routed combinations of them.

We’re looking for an AI researcher to take an integral part in PointFive’s research into how these models behave, how they should be evaluated, and how they can be deployed and used more efficiently across real engineering workloads.

This is a deeply technical research role with direct product impact.

You’ll study frontier and open-source models, develop evaluation frameworks, investigate inference and model optimization techniques, explore local deployment strategies, and help answer questions such as:

Which model should handle a particular task?

How much capability do we lose when we quantize it?

Can a smaller local model replace a frontier API for a specific workload?

How should models be routed across latency, cost, privacy, and quality constraints?

How do we measure whether one model is actually better than another for agentic software engineering?

Your work will directly shape PointFive’s model strategy and the intelligence behind our AI infrastructure products.

What You'll Do

  • Define and lead PointFive’s LLM research agenda across model evaluation, inference optimization, local models, routing, and model adaptation.

  • Continuously evaluate frontier and open-weight models across real-world software engineering and agentic workloads.

  • Design rigorous evaluation frameworks for model quality, reasoning, tool use, code generation, command execution, summarization, and agent behavior.

  • Build benchmarks that reflect actual developer workflows rather than generic academic tasks.

  • Research model efficiency techniques including quantization, distillation, speculative decoding, prefix caching, KV-cache optimization, batching, and context management.

  • Investigate when smaller or locally deployed models can replace expensive frontier models without materially degrading task quality.

  • Evaluate model architectures, parameter sizes, quantization formats, and runtime configurations across heterogeneous hardware.

  • Research and benchmark local inference stacks including llama.cpp, MLX, vLLM, SGLang, ONNX Runtime, Ollama, TensorRT-LLM, and emerging runtimes.

  • Study inference performance across Apple Silicon, NVIDIA GPUs, AMD GPUs, CPUs, Windows workstations, Linux machines, and other endpoint configurations.

  • Develop model-routing strategies that optimize across quality, latency, cost, privacy, context size, and hardware availability.

  • Explore intelligent cascades where smaller models handle common tasks and more capable models are invoked only when necessary.

  • Research model specialization through fine-tuning, LoRA, adapters, distillation, prompt optimization, and other model adaptation techniques.

  • Investigate model behavior under constrained environments, including offline execution, limited memory, limited compute, and local-only inference.

  • Evaluate agent-specific model behavior, including planning, tool selection, shell interaction, code editing, error recovery, and long-running task execution.

  • Analyze failure modes such as hallucination, tool misuse, context degradation, reasoning collapse, excessive token consumption, and unstable agent loops.

  • Design experiments that quantify the tradeoffs between model capability, inference cost, token consumption, latency, and resource utilization.

  • Build internal research infrastructure for reproducible model experiments, benchmarking, dataset management, and evaluation.

  • Track frontier model releases and emerging research, rapidly determining which developments are meaningful for PointFive’s products.

  • Collaborate closely with engineering and product teams to translate research results into production capabilities.

  • Build and lead a small, exceptional LLM research team over time.

What We're Looking For

Must-have

  • Deep understanding of modern large language models and transformer-based architectures.

  • Strong hands-on experience evaluating and experimenting with both frontier and open-weight models.

  • Strong understanding of inference behavior, including prefill, decoding, KV caches, context windows, batching, memory usage, and token generation performance.

  • Experience with model optimization techniques such as quantization, distillation, LoRA, fine-tuning, or model compression.

  • Strong experimental mindset - able to formulate hypotheses, design controlled experiments, and draw meaningful conclusions from noisy results.

  • Experience building evaluation frameworks for LLM quality and behavior.

  • Strong Python proficiency and familiarity with the modern ML ecosystem.

  • Ability to read, understand, and reproduce ideas from current ML research papers.

  • Comfortable working with ambiguous research problems where there may not yet be an established best practice.

  • Strong ability to bridge research and production - understanding not only whether something works, but whether it is practical to deploy.

Nice to have

  • Experience with open-weight models such as Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, or similar model families.

  • Experience with frontier model APIs including OpenAI, Anthropic, Google, and other leading providers.

  • Experience with inference frameworks such as vLLM, SGLang, llama.cpp, MLX, TensorRT-LLM, Ollama, or ONNX Runtime.

  • Deep understanding of quantization techniques including FP8, INT8, INT4, AWQ, GPTQ, GGUF, and related approaches.

  • Experience with GPU performance optimization, CUDA, Metal, ROCm, or DirectML.

  • Familiarity with distributed inference and multi-GPU serving.

  • Experience with model routing, mixture-of-model systems, cascades, or adaptive inference.

  • Experience with reinforcement learning, preference optimization, DPO, GRPO, or related post-training techniques.

  • Familiarity with agentic systems, coding agents, tool-use models, and computer-use models.

  • Experience building or evaluating coding benchmarks and software-engineering agents.

  • Experience with synthetic data generation, dataset curation, and automatic evaluation.

  • Research publications or meaningful contributions to open-source ML projects.

  • Experience leading a small applied research or ML research team.

Research Areas

Some of the problems we expect this team to work on include:

  • Model Routing Choosing the optimal model dynamically based on task difficulty, latency requirements, cost, privacy constraints, and hardware availability.

  • Local vs. Cloud Inference Determining which workloads can reliably move from cloud models to models running directly on developer endpoints.

  • Model Compression Understanding how far models can be quantized, distilled, or otherwise optimized before meaningful capability is lost.

  • Agentic Model Evaluation Building evaluation methods for agents that operate over codebases, shells, developer tools, and long-running workflows.

  • Inference Efficiency Improving throughput, latency, memory consumption, and token efficiency across different model architectures and runtimes.

  • Model Specialization Investigating whether smaller specialized models can outperform general-purpose frontier models for narrow engineering tasks.

  • Long-Context Behavior Understanding how models behave as context grows, what information gets lost, and how context can be compressed or structured more intelligently.

  • Hardware-Aware AI Matching models and inference strategies to available GPUs, CPUs, NPUs, Apple Silicon, and other endpoint hardware.

  • Cost vs. Intelligence Tradeoffs Quantifying when additional model capability actually produces better outcomes - and when it simply produces more expensive tokens.

Our Tech Stack

Python, Go, PyTorch, Hugging Face, vLLM, SGLang, llama.cpp, MLX, CUDA, Metal, AWS, Cloudflare, Snowflake, and a rapidly evolving ecosystem of frontier and open-weight models.

Why PointFive

  • A rare research surface

We operate where LLM research meets real enterprise infrastructure. The questions we work on have immediate implications for how thousands of engineers use AI every day.

  • Access to real workloads

Instead of optimizing against abstract benchmarks, you'll be able to study how models perform across real software engineering and agentic workflows.

  • Research that ships

This isn't a research lab disconnected from product. Successful ideas can move rapidly from experiment to production.

  • Models are becoming infrastructure

Enterprises will increasingly operate fleets of models across cloud APIs, private infrastructure, and developer endpoints. Deciding how those models are selected, optimized, deployed, and governed is becoming a fundamental infrastructure problem.

  • Frontier moves fast

New models, architectures, inference techniques, and agent systems appear constantly. Your job is to understand which developments matter - and turn them into an advantage for PointFive.

  • Founders with a track record

Built and sold IntSights to Rapid7. Backed by top-tier investors with deep conviction in the category.

  • Early-stage leverage

You'll define PointFive's LLM research strategy, research methodology, and eventually the team itself.

Equal Opportunity Statement

PointFive is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We welcome candidates from all backgrounds, experiences, and perspectives to apply.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
908,794 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Leadership
Similar stack
Same company
Tel Aviv
≈ $61k – $158k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Netanya
AI/ML
AI Agents
Edge AI
DevOps
AWS
Management
Agile
ITSM
Apply
$102k – $127k per year • In office • Full-Time • Bachelor's Degree • United States
AI/ML
Multimodal AI
Management
Outlook
Microsoft Office
Apply
$4k – $14k per year • Equity • In office • 7+ years exp • Bachelor's Degree • New York
Apply
$146k – $190k per year • Equity • Remote (United Kingdom) • Full-Time • 8+ years exp • London
AI/ML
Multimodal AI
Function Calling
Human-in-the-Loop
Tool Use
Apply
≈ $157k – $310k per year (Estimated) • Remote (likely United States) • Full-Time • 10+ years exp • Bachelor's Degree
Python
SQL
AI/ML
Model Context Protocol
AI Agents
LLM
NIST AI RMF
Tool Use
Machine Learning
Cybersecurity
OWASP Top 10
Apply
≈ $22k – $55k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
Java
AI/ML
Fine-tuning
Embeddings
LLM
RAG
LLMOps
LLM Evaluation
LLM Guardrails
Machine Learning
Apply
≈ $30k – $83k per year (Estimated) • Remote (Argentina) • Full-Time • Buenos Aires
Python
SQL
DevOps
Rest API
Datadog
Management
n8n
Zapier
Apply
≈ $25k – $62k per year (Estimated) • Equity • In office • 8+ years exp • Bachelor's Degree • Munich
Python
Apply
≈ $41k – $93k per year (Estimated) • Hybrid • 3+ years exp • Thessaloniki
Python
PowerShell
C#
AI/ML
Red Teaming
Cybersecurity
Burp Suite
Metasploit
Nmap
Nessus
Cobalt Strike
Impacket
BloodHound
MITRE ATT&CK
OWASP
Analytics
Microsoft Excel
Management
Agile
Apply
≈ $82k – $181k per year (Estimated) • Hybrid • Full-Time • Stockholm
Python
AI/ML
Vertex AI
DevOps
GCP
Git
Management
Agile
Apply
≈ $118k – $281k per year (Estimated) • Hybrid • Full-Time • Tel Aviv
Go
TypeScript
Databases
Snowflake
AI/ML
Copilot
Cursor
llama.cpp
Claude Code
Model Context Protocol
vLLM
CUDA Toolkit
Quantization
AI Agents
LocalAI
Ollama
MLX ML
CUDA
OpenAI Codex
ROCm
ONNX Runtime
KV Cache
DevOps
Cloudflare
Hyper-V
eBPF
FinOps
Linux
Windows
Apply
≈ $134k – $275k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Tel Aviv
Python
Go
Java
Databases
PostgreSQL
Snowflake
AI/ML
Claude Code
AI Agents
LLM
DevOps
AWS
Docker
Kubernetes
FinOps
Apply
≈ $82k – $225k per year (Estimated) • Hybrid • Full-Time • Tel Aviv
Go
TypeScript
Databases
Snowflake
AI/ML
Copilot
Cursor
Claude Code
AI Agents
OpenAI Codex
DevOps
Ansible
Cloudflare
FinOps
Linux
Windows
Unix
Apply
Controller 21 days ago
≈ $80k – $176k per year (Estimated) • Equity • Hybrid • Full-Time • 5+ years exp • Tel Aviv
AI/ML
AI Agents
DevOps
FinOps
Apply
$180k – $200k per year • Remote (United States) • Full-Time • 7+ years exp • New York • Boston
Databases
Snowflake
Databricks
AI/ML
AI Agents
LLM
Agentic Workflows
DevOps
GCP
Azure
AWS
FinOps
Apply
≈ $139k – $345k per year (Estimated) • In office • Tel Aviv
AI/ML
AI Agents
Apply
≈ $105k – $228k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Tel Aviv
SQL
AI/ML
Model Context Protocol
Computer Vision
AI Agents
LLM
LLM Guardrails
Multi-Agent Systems
DevOps
CI/CD
Apply
≈ $136k – $322k per year (Estimated) • Hybrid • Full-Time • Tel Aviv
AI/ML
AI Agents
Management
Monday.com
Apply
≈ $92k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
AI/ML
Chain-of-Thought
AI Agents
LLM
Tool Use
Machine Learning
DevOps
AWS
Apply
≈ $82k – $191k per year (Estimated) • In office • Full-Time • 3+ years exp • Tel Aviv
DevOps
Linux
Apply
See all jobs
This is one of many
908,794 more open roles from verified company boards, updated every day.