1,389,029open jobs
80,602companies
206,165added this week
Browse all
Salary
≈ $136k – $241k per year (Estimated)
Location
Remote (United States)
Seniority
Middle · 4+ years exp

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Oct 7, 2026.

Overview
Company
Impact
Profile match
ScaleOps is a cloud infrastructure automation company headquartered in Tel Aviv, Israel, and founded in 2022. The company provides an autonomous platform for Kubernetes resource management that continuously optimizes CPU, memory, and GPU allocations for containerized workloads in real-time. It serves global enterprises and AI-driven organizations by reducing cloud expenditures and improving application performance through automated pod rightsizing and node consolidation.

ScaleOps is redefining autonomous cloud and AI infrastructure. We're on a mission to free DevOps and platform engineers from manual resource management so they can focus on innovation, not tuning resources. The results: maximized performance and a reduction of cloud costs by up to 80%.

As the category leader in Autonomous Cloud and AI Infrastructure Resource Management, we're trusted by leading enterprises including Adobe, Wiz, Epic Games, Northwestern Mutual, Coinbase, DocuSign, and Fortune 100 companies to autonomously manage their most critical production environments. 

Backed by Insight Partners, Lightspeed Venture Partners, and other leading VCs with over $210M in funding, ScaleOps is the leading player in a massive and growing market. We are building the autonomous infrastructure management platform that will power the next decade of enterprise compute.

What You'll Be Doing

We're forming a new AI Group and looking for a Senior AI Engineer to help shape it from an early stage - a greenfield, long-term effort to evolve how decisions are made across the ScaleOps platform using AI-driven systems. You won't just integrate APIs or build demos; you'll build the AI brain that works alongside (and increasingly drives) our core automation engine, with real production impact from day one and room to grow into technical leadership as the group scales.

  • Agentic AI Architecture: Design and build autonomous AI agents that analyze infrastructure in real time and make intelligent decisions. Work with modern agentic frameworks (LangGraph, PydanticAI) and conversational AI to create multi-agent systems - including troubleshooting, optimization, FinOps, and how-to agents. Leverage core LLM capabilities (tool-use, memory, retrieval) to operate safely in production.
  • Platform Integration & Intelligent Decision Systems: Develop MCPs to expose ScaleOps capabilities to AI agents that reason over infrastructure environments, metrics, configurations, and cost signals. Build integrations with tools like Slack, Jira, and AI-powered IDEs (Cursor, Windsurf) to deliver context-aware insights, from "why is this pod not scheduling?" to "how can we reduce costs by 30% safely?"
  • AI Model Development & MLOps: Build and deploy machine learning models that learn from infrastructure patterns - detecting the right resource policies for workloads, predicting optimal scaling triggers, and recommending GPU configurations. Own the complete ML pipeline from training to production, ensuring models are reliable, monitored, and continuously improving.
  • R&D AI Tools Development & Adoption: Build and embed internal AI tools to accelerate engineering, development, research, and support.
  • AI Tools for Business Impact: Develop AI-powered tools that help Sales and Support teams demonstrate value instantly - agents that analyze customer infrastructure, generate cost optimization reports automatically, and turn technical data into clear business recommendations.
  • End-to-End Ownership: Own AI systems from concept to production, ensuring they're fast (sub-2-second responses), reliable, safe, and cost-effective. Build evaluation frameworks to measure quality, implement security controls, and balance performance tradeoffs in production.
  • Technical Leadership: Define AI architecture and best practices as a founding member of the AI team. Make key technical decisions - choosing frameworks, designing multi-agent systems, establishing data governance - and shape how ScaleOps evolves from AI-enhanced internal tools to customer-facing AI products.

What You'll Bring

  • Core Engineering: Significant software engineering experience (typically 4+ years) with strong Python skills and solid backend engineering fundamentals.
  • Production Experience: Experience building and operating production systems in cloud environments.
  • Real-World GenAI Experience: Practical experience bringing LLM-based systems into production, including handling latency, cost control, and failure modes. Familiarity with additional agentic frameworks (e.g., LangChain, MetaGPT) and evaluation frameworks.
  • Builder Mentality: Strong ownership and the ability to operate independently while collaborating closely across teams, with the motivation to grow into technical leadership as the group expands.
  • (Advantage) Data & RAG: Experience enabling LLMs to consume structured or operational data (configurations, logs, metrics) and experience with retrieval systems (RAG) or vector databases.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,389,029 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
$235k – $260k per year • In office • Full-Time • 3+ years exp • San Francisco • Toronto
Python
Go
Rust
AI/ML
AI Agents
LLM
RAG
DevOps
GitHub
Apply
Remote (India) • Contractor • 6+ years exp
Python
Go
TypeScript
Python
FastAPI
AI/ML
Claude
NLP
RAG
Hallucination
OpenAI
Human-in-the-Loop
LLM Guardrails
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Platform Engineering
Argo Workflows
Analytics
A/B Testing
Apply
≈ $41k – $101k per year (Estimated) • Remote (Pakistan) • Contractor • 6+ years exp
Python
Go
TypeScript
Python
FastAPI
AI/ML
Claude
NLP
RAG
Hallucination
OpenAI
Human-in-the-Loop
LLM Guardrails
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Platform Engineering
Argo Workflows
Analytics
A/B Testing
Apply
$173k – $225k per year • Hybrid • Full-Time • 5+ years exp • San Carlos
Python
TypeScript
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
DynamoDB
OpenSearch
Amazon Aurora
AI/ML
LangChain
Embeddings
Multimodal AI
Function Calling
TensorRT
TensorRT-LLM
AWS Bedrock
LLM
RAG
Hallucination
Time Series Forecasting
Triton
OpenAI
Anthropic
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
CI/CD
AWS
Amazon S3
Analytics
A/B Testing
Apply
$224k – $260k per year • Hybrid • Full-Time • 8+ years exp • San Carlos
Python
TypeScript
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
DynamoDB
OpenSearch
Amazon Aurora
AI/ML
LangChain
Embeddings
Multimodal AI
Function Calling
TensorRT
TensorRT-LLM
AWS Bedrock
LLM
RAG
Hallucination
Time Series Forecasting
Triton
OpenAI
Anthropic
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
CI/CD
AWS
Amazon S3
Analytics
A/B Testing
Apply
≈ $20k – $50k per year (Estimated) • Hybrid • 3+ years exp • India
Python
JavaScript
TypeScript
SQL
Node JS
Python
Flask
FastAPI
Databases
PostgreSQL
AI/ML
LangGraph
LangChain
Claude
LlamaIndex
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
Gemini
LLM
RAG
RAGFlow
OpenAI
OpenAI Agents SDK
Structured Outputs
LLM Guardrails
Tool Use
DevOps
Rest API
GCP
Azure DevOps
Azure
CI/CD
Git
AWS
Docker
GitHub
Management
Confluence
Jira
SharePoint
Agile
Scrum
Apply
≈ $68k – $125k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
OpenAI
Anthropic
Machine Learning
DevOps
Azure
AWS
Analytics
Power BI
ETL/ELT
Master Data Management
Apply
$160k – $250k per year • In office • 4+ years exp • PhD • United States
Python
AI/ML
Scikit-learn
JAX
TensorFlow
PyTorch
Machine Learning
Quantum
Qiskit
Cirq
PennyLane
TKET
Apply
≈ $68k – $162k per year (Estimated) • In office • London
Python
SQL
Apply
≈ $74k – $170k per year (Estimated) • In office • London
Python
SQL
Apply
Researcher 1 month ago
≈ $105k – $262k per year (Estimated) • In office • 7+ years exp • Tel Aviv
Python
AI/ML
Time Series Forecasting
Machine Learning
DevOps
Kubernetes
FinOps
Cybersecurity
Wiz
Apply
AI Engineer 25 days ago
≈ $102k – $255k per year (Estimated) • In office • 4+ years exp • Tel Aviv
Python
AI/ML
Cursor
LangGraph
Windsurf
LangChain
AI Agents
Pydantic AI
LLM
RAG
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
FinOps
Cybersecurity
Wiz
Management
Slack
Jira
Apply
≈ $136k – $238k per year (Estimated) • Remote (United States) • 5+ years exp
Go
DevOps
GCP
Helm
Azure
AWS
Kubernetes
Platform Engineering
IAM
Cybersecurity
Wiz
Apply
Backend Engineer 3 days ago
≈ $110k – $206k per year (Estimated) • Remote (United States) • 4+ years exp • Bachelor's Degree
DevOps
Helm
Istio
Prometheus
Kubernetes
Cybersecurity
Wiz
Apply
≈ $87k – $212k per year (Estimated) • In office • Tel Aviv
DevOps
Kubernetes
Cybersecurity
Wiz
Design
Figma
Apply
See all jobs
This is one of many
1,389,029 more open roles from verified company boards, updated every day.