1,456,758open jobs
88,633companies
228,746added this week
Browse all
Location
In office (Bengaluru)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Jul 14, 2026.

Overview
Company
Impact
Profile match
Aion is an enterprise artificial intelligence company headquartered in San Francisco, California, and founded in 2025 by Jayden Watson and Christian Angermayer. The company offers a platform for building, fine tuning, and deploying AI systems and agents in production, sold together with forward deployed engineers who implement it inside the customer organization. It runs teams across the United States, the United Kingdom, and India, targeting enterprises that want working AI systems rather than pilots.

About aion

Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.

We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.

We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.

Who You Are

You are a visionary infrastructure architect passionate about democratizing AI compute at global scale. You thrive on solving complex technical challenges that create elegant, accessible systems from intricate infrastructure. With deep expertise in secure multi-tenancy environments, you understand how to design and implement comprehensive isolation guarantees across hardware, network, and storage layers for both VM and container workloads.

You're excited to join an ambitious AI infrastructure startup at the ground floor, where your work will directly unlock siloed compute resources and remove barriers limiting AI advancement. You have the technical depth to architect platform systems that seamlessly connect compute providers with AI engineers while maintaining robust security foundations that scale to serve diverse client requirements and compliance needs.

You're motivated by the opportunity to build something transformative-creating the infrastructure that will make high-performance compute more accessible, affordable, and user-friendly for the next generation of AI innovation.

What You'll Do

  • Observability Systems: Build and deploy comprehensive monitoring for GPU infrastructure using DCGM, NVML, and custom exporters; design metrics collection pipelines that track GPU health, utilization, thermal management, and performance across heterogeneous providers
  • LGTM Stack Architecture & Operations: Deploy and manage production-scale Loki, Grafana, Tempo, and Mimir alongside Prometheus and Thanos/VictoriaMetrics; design retention strategies, aggregation rules, and query patterns that scale to thousands of GPUs
  • Custom Exporter Development: Write Prometheus exporters in Go or Python for GPU metrics, platform services, and infrastructure components; implement proper metric naming, labeling strategies, and follow OpenMetrics standards
  • Kubernetes Controller Development: Build custom controllers and operators for GPU workload management, scheduling, and resource allocation; instrument controllers with comprehensive metrics and tracing for observability
  • Training & Inference Monitoring: Design and implement specialized observability for AI training workloads (GPU efficiency, distributed training performance, resource utilization) and inference services (latency percentiles, throughput, cost analytics)
  • SLURM Integration & Monitoring: Deploy and manage SLURM clusters for HPC workloads, build observability for batch jobs, create bridges between SLURM and Kubernetes, and design unified monitoring across orchestrators
  • Systemd Service Development: Write and deploy systemd services for bare-metal GPU nodes including monitoring agents, metric collectors, and platform daemons; implement proper logging and error handling
  • GitOps Platform Engineering: Manage infrastructure using ArgoCD and GitOps workflows, create observable platform abstractions that teams consume declaratively, build self-service capabilities with integrated monitoring
  • Alerting & SRE: Design intelligent, actionable alerting systems for GPU failures, thermal throttling, performance degradation, and workload anomalies; define platform SLOs and implement comprehensive monitoring to track reliability
  • Multi-Tenant Monitoring Isolation: Build secure observability isolation ensuring customers access only their metrics and logs while maintaining platform-wide visibility for operations; implement query-time filtering and RBAC
  • Metrics Pipeline Engineering: Implement automated collection of GPU and platform telemetry; integrate with OpenTelemetry for unified observability; manage cardinality and storage costs through intelligent aggregation
  • Cost & Utilization Analytics: Build monitoring pipelines that track GPU utilization, idle time, efficiency metrics, and cost allocation per tenant; create dashboards for platform economics and provider payout calculations
  • Dynamic Workload Management: Design systems that use observability data to detect hardware failures, trigger workload migration, and handle graceful degradation; ensure monitoring survives infrastructure changes
  • Self-Service Observability Platforms: Build intuitive dashboards, APIs, and alerting interfaces enabling providers to monitor hardware contributions and customers to track workload performance in real-time
  • Incident Response & Debugging: Use observability systems to rapidly identify root causes during production issues - GPU hardware failures, network bottlenecks, or workload problems; build tooling that reduces MTTD and MTTR
  • Performance Optimization: Leverage observability data to identify bottlenecks in distributed training, inference latency issues, networking inefficiencies, and opportunities for GPU utilization improvement

Requirements

Technical Skills & Experience

  • 6-10 years of experience in infrastructure engineering with strong focus on observability, monitoring systems, and production service development (exceptional candidates with different experience profiles will be considered)
  • GPU Observability expertise with production experience monitoring NVIDIA GPUs using DCGM, NVML, and nvidia-smi; understanding GPU-specific metrics (utilization, memory, temperature, power, ECC errors) and failure modes
  • LGTM Stack proficiency deploying and operating Loki, Grafana, Tempo, and Mimir alongside Prometheus in production; experience with at least one long-term storage solution (Thanos, VictoriaMetrics, or Mimir)
  • Systems Programming in Go or Rust for building production services including custom Prometheus exporters, monitoring agents, systemd daemons, and instrumented controllers
  • Kubernetes controller development building custom controllers and operators using controller-runtime or client-go; implementing proper instrumentation and observability for custom resources
  • Active & Passive Monitoring designing SLO/SLI frameworks, implementing intelligent alerting strategies, building health check systems, and creating anomaly detection for distributed workloads
  • HPC Systems experience deploying and managing SLURM clusters, understanding job schedulers, and monitoring batch workloads with observability requirements different from typical web services
  • Advanced Kubernetes expertise including custom resource definitions, admission controllers, scheduling extensions, and cluster-wide monitoring architectures
  • AI Workload Observability monitoring ML training jobs (GPU efficiency, distributed training metrics, NCCL performance) and inference workloads (latency, throughput, batch processing, cost per inference)
  • Metrics Architecture design including cardinality management, aggregation strategies, recording rules, retention policies, and balancing observability costs with data granularity at scale
  • GitOps & ArgoCD experience managing observable infrastructure declaratively, building platform abstractions, and creating self-service systems with integrated monitoring
  • Systems expertise writing unit files, managing service dependencies, deploying monitoring agents as systemd services, and debugging service failures on bare-metal hosts
  • Multi-tenant observability implementing secure metrics and log isolation, RBAC policies for monitoring data, and ensuring proper isolation while maintaining platform-wide visibility
  • Networking proficiency with CNI plugins, understanding network observability, and monitoring distributed training communication patterns
  • Infrastructure as Code using Terraform or similar tools to deploy observable infrastructure, building monitoring into provisioning workflows

Benefits

Preferred Attributes:

  • Founder-level ownership and bias for action.
  • Strong strategic thinking and ability to connect technical decisions to business impact.
  • Excellent communication and mentoring skills.
  • Thrives in ambiguity, fast-paced environments, and early-stage startup culture.

Why Join aion?

  • Work directly with high-pedigree founders shaping technical and product strategy.
  • Build infrastructure powering the future of AI compute globally.
  • Significant ownership and impact with equity reflective of your contributions.
  • Competitive compensation, flexible work options, and wellness benefits
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,456,758 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Bengaluru
≈ $16k – $37k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Pune
PowerShell
Databases
Azure SQL Database
DevOps
Terraform
Ansible
Docker Compose
Azure DevOps
Prometheus
Azure
CI/CD
Git
Docker
Kubernetes
Grafana
Bicep
GitHub
Cybersecurity
Microsoft Entra ID
Active Directory
QA
Selenium
JMeter
Pytest
Apply
≈ $22k – $61k per year (Estimated) • In office • 10+ years exp • Chandigarh
JavaScript
Java
Java
Spring Boot
AI/ML
AI Agents
DevOps
Rest API
CI/CD
AWS
IAM
Apply
≈ $17k – $37k per year (Estimated) • In office • 5+ years exp • Chandigarh
JavaScript
Java
Java
Spring Boot
AI/ML
AI Agents
DevOps
Rest API
CI/CD
AWS
Apply
In office • Chandigarh
Databases
MySQL
Oracle
MS SQL
DevOps
Splunk
Datadog
Dynatrace
Windows Server
AWS
Grafana
Windows
Apply
≈ $18k – $40k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Blue-Green Deployment
Platform Engineering
Incident Management
GitHub
Management
Agile
Apply
≈ $27k – $59k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Osijek
Python
MATLAB
Apply
In office • Osijek
Python
Go
Java
C
C++
C
GNU Make
DevOps
Linux
Wi-Fi
Apply
≈ $41k – $96k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Osijek
Python
Apply
Web QA Engineer 5 hours ago
In office • Bachelor's Degree • Osijek
Python
JavaScript
TypeScript
Frontend
Lighthouse
DevOps
Rest API
GitHub Actions
GitLab CI
CI/CD
Jenkins
Git
Management
Confluence
ClickUp
Jira
Agile
Scrum
QA
Selenium
JMeter
Cypress
Playwright
Postman
BrowserStack
k6
Apply
≈ $40k – $95k per year (Estimated) • In office • Bachelor's Degree • Osijek
Python
C++
DevOps
CI/CD
Linux
Apply
≈ $19k – $39k per year (Estimated) • In office • Full-Time • Bengaluru
AI/ML
Fine-tuning
Design
Figma
ProtoPie
Adobe After Effects
Rive
Apply
≈ $132k – $240k per year (Estimated) • In office • Full-Time • San Francisco
JavaScript
TypeScript
AI/ML
Cursor
Claude Code
Fine-tuning
OpenAI Codex
Frontend
Three.JS
React.js
WebGPU
React Three Fiber
Game Dev
GLSL
WGSL
Design
Blender
Adobe After Effects
Cinema 4D
Rive
Apply
In office • Full-Time • Master's Degree • Bengaluru
Go
Databases
MySQL
PostgreSQL
Redis
NATS
RabbitMQ
Apache Kafka
AI/ML
Jupyter Notebook
Fine-tuning
DevOps
gRPC
Terraform
GCP
Prometheus
SLURM
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Amazon S3
HPC
Analytics
ETL/ELT
Apply
Engineering Lead 3 months ago
≈ $17k – $38k per year (Estimated) • In office • Full-Time • Master's Degree • Bengaluru
Python
Go
Java
Rust
C++
AI/ML
Fine-tuning
LLM
DevOps
Terraform
GCP
OpenTelemetry
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Apply
AI Engineer 3 months ago
In office • Full-Time • Bengaluru
Python
Databases
PostgreSQL
Weaviate
Chroma
Milvus
pgvector
Pinecone
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
AI Agents
TensorRT
TensorRT-LLM
TGI
Llama
Mistral
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
Hallucination
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Apply
Advisor TPM 7 hours ago
≈ $24k – $49k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
Management
Agile
Apply
In office • Full-Time • 7+ years exp • Bachelor's Degree • Bengaluru
SQL
Databases
Snowflake
Apache Kafka
DevOps
GCP
Azure
AWS
Incident Management
Management
Agile
Apply
In office • Bengaluru
Apply
In office • Bengaluru
Apply
In office • Bengaluru
Apply
See all jobs
This is one of many
1,456,758 more open roles from verified company boards, updated every day.