Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs
Web site created using create-react-app.

About aion

Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.

We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.

We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.

Who You Are

You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads.

You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components.

You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops.

What You'll Do

Platform Development & Implementation

  • Build and maintain platform services across aion's Compute and Inference platforms, working closely with senior engineers and platform leads
  • Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems
  • Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage
  • Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions
  • Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs

Backend Systems & Distributed Infrastructure

  • Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving
  • Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics
  • Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components
  • Build event-driven architectures for asynchronous processing, job scheduling, and platform automation
  • Implement monitoring, logging, and alerting for platform services to ensure production reliability

Code Quality & Engineering Excellence

  • Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability
  • Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives
  • Refactor existing code to improve maintainability, performance, and scalability
  • Document design decisions, API specifications, and operational runbooks for platform services
  • Debug production issues and contribute to incident response and post-mortems

Requirements

Technical Skills & Experience

  • 2-4 years of experience in backend engineering, platform development, or distributed systems
  • Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing
  • Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles
  • Hands-on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions
  • Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases
  • Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures
  • Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms
  • Experience with containerization (Docker) and basic Kubernetes concepts
  • Knowledge of testing frameworks and practices (unit tests, integration tests, mocking)
  • Familiarity with Git, CI/CD pipelines, and modern development workflows
  • Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus
  • Experience with infrastructure-as-code (Terraform) or observability tools (Prometheus, Grafana) is beneficial

Bonus/ Good to Have

  • HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration
  • Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks
  • Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents)
  • ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems
  • Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity

Benefits

Preferred Attributes:

  • High ownership, self driven and biased for action.
  • Strong strategic thinking and ability to connect technical decisions to business impact.
  • Excellent communication and mentoring skills.
  • Thrives in ambiguity, fast-paced environments, and early-stage startup culture.

Why Join aion?

  • Work directly with high-pedigree founders shaping technical and product strategy.
  • Build infrastructure powering the future of AI computers globally.
  • Significant ownership and impact with equity reflective of your contributions.
  • Competitive compensation, flexible work options, and wellness benefits

Recommended for you based on this role

Similar stack
Same company
In your city
Network Engineer 1 month ago
In office • Full-Time • Master's Degree • Bengaluru
Bash
Python
AI/ML
Fine-tuning
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
GitOps
Grafana
Istio
Kubernetes
Linkerd
OpenTelemetry
Platform Engineering
Prometheus
Terraform
Cybersecurity
Tcpdump
Wireshark
Zero Trust
Apply
Hardware Engineer 1 month ago
In office • Full-Time • Master's Degree • Bengaluru
Bash
Python
AI/ML
Fine-tuning
DevOps
Kubernetes
Platform Engineering
Apply
AI Engineer 1 month ago
In office • Full-Time • Bengaluru
Python
Databases
Chroma
Milvus
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
AutoGen
Claude
CrewAI
Fine-tuning
Gemini
Hallucination
LangChain
LangGraph
Llama
LlamaIndex
LLM
Mistral
Multimodal AI
Prompt Engineering
PyTorch
RAG
TensorFlow
TensorRT
TensorRT-LLM
vLLM
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Rest API
Vector
Apply
In office • Contractor • Master's Degree • London
AI/ML
AI Agents
Fine-tuning
RAG
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Platform Engineering
Apply
In office • Full-Time • Master's Degree • London
C++
Go
Python
Rust
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
AI/ML
Fine-tuning
Jupyter Notebook
LLM
Quantization
TensorRT
TensorRT-LLM
vLLM
DevOps
Docker
Grafana
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
SLURM
Apply
In office • Full-Time • Master's Degree • London
JavaScript
Python
TypeScript
AI/ML
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
LoRA
Multimodal AI
Prompt Engineering
QLoRA
Quantization
RAG
RLHF
Synthetic Data
TensorRT
TensorRT-LLM
vLLM
PEFT
Frontend
GraphQL
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
WebRTC
WebSockets
Apply
In office • Full-Time • Master's Degree • Bengaluru
Python
Rust
AI/ML
Anomaly Detection
Fine-tuning
DevOps
ArgoCD
GitOps
Grafana
Kubernetes
Loki
Mimir
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
SLURM
Terraform
Thanos
VictoriaMetrics
Apply
In office • Full-Time • Master's Degree • London
JavaScript
Python
TypeScript
AI/ML
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
LoRA
Multimodal AI
Prompt Engineering
QLoRA
Quantization
RAG
RLHF
Synthetic Data
TensorRT
TensorRT-LLM
vLLM
PEFT
Frontend
GraphQL
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
WebRTC
WebSockets
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
London
Remote work
In office
Employment
Full-Time

Compensation

Benefits
Flexible schedule
Equity
Equity stake in a tech company