368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$21k – $51k per year (Estimated)
Location
In office (Almaty)
Employment
Full-Time
Overview
Company
Impact
Profile match
The True Engineer is an engineering leadership and software development publication platform hosted on Substack. Authored by Adlet Balzhanov, a senior software engineering lead with background in Big Tech and high-growth enterprise scale-ups, the publication delivers monthly long-form analytical essays, career frameworks, and strategic insights for software developers, engineering managers, and technology executives. Operating under a creator-driven digital subscription model, the newsletter focuses on technical leadership tactics, staff-level engineering practices, organizational dynamics within large tech companies, and navigating the operational impacts of AI on software development.

Why work at Higgsfield AI?

Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands.

We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like.

About the role

You will own the GPU platform behind our generative video/image products - both the inference fleet that serves production traffic and the training clusters where our models are built. The estate spans multiple providers orchestrated with Kubernetes on Talos Linux managed by Sidero Omni. The fleet target is >95% GPU utilization; every idle GPU-hour is money burned. You'll be the person who keeps training fast and inference cheap and boring.

What you'll do

  • Optimize the training clusters: distributed training at scale - NCCL tuning, InfiniBand/RoCE fabric health, topology-aware scheduling and gang placement, GPU/network throughput, fast checkpointing, job preemption and recovery. Make every training run use the hardware it paid for.

  • Own Talos / Sidero Omni cluster lifecycleacross the GPU fleet: node bootstrap and upgrades, GPU drivers / NVIDIA GPU Operator / DCGM on an immutable OS, zero-downtime rollouts.

  • Operate the multi-provider GPU fleet: capacity planning across Nebius regions and bare-metal RTX Pro pools, hardware incident escalation to providers, node lifecycle (NotReady triage, XID errors, driver upgrades).

  • Own inference autoscaling: KEDA-driven, Kafka-queue-based scaling of GPU consumers; GPU-aware scheduling; warm pools and cold-start reduction; supply/demand tuning of our in-house autoscaler (higgscaler).

  • GitOps everything: ArgoCD multi-cluster (10+ clusters from one repo), Helm, Terraform (HCP). No hand-labeled nodes, no console drift - if it's not in git, it doesn't exist.

  • Observability & SLOs: VictoriaMetrics/Logs/Traces, Prometheus, DCGM exporters; honest dashboards for utilization, training throughput, queue latency, cost per generation.

  • GPU efficiency as a discipline: hunt idle allocations, capacity/demand mismatches, starved queues - our AIOps platform (Mycelium) files these findings automatically; you close the loop with real fixes.

  • Partner with ML engineers on training runs and model-serving rollouts (runtimes, batching, memory sizing) and with the core team on AWS EKS (Karpenter, Istio, Bottlerocket, gVisor sandboxes).

    Our stack

    Talos Linux, EKS, NVIDIA GPU Operator, DCGM, CUDA, NCCL, InfiniBand/RoCE · ArgoCD, Helm, Terraform Cloud · KEDA, Karpenter, Kafka · VictoriaMetrics/Logs/Traces, Prometheus, Grafana · Istio, Cloudflare · Python/Go.

    You have

  • 3+ years running production Kubernetes as SRE/Platform/MLOps, includingGPU workloads.

  • Hands-on distributed training operations: NCCL, high-speed interconnects (InfiniBand/RoCE), multi-node job scheduling, checkpointing strategies - and the habit of measuring throughput before and after every change.

  • Bare-metal Kubernetes experience: Talos or similar immutable-OS setups; node lifecycle without a cloud safety net.

  • The NVIDIA stack: drivers, container toolkit, GPU Operator, DCGM metrics; you can debug "GPU visible but not allocatable" at 3am.

  • GitOps fluency (ArgoCD/Flux + Helm) and Terraform; strong Linux and networking (multi-cluster, VPN/TGW topologies).

  • Queue-based autoscaling (KEDA/HPA) and enough Kafka to reason about consumer lag.

  • Python or Go for automation; you write things down.

Nice to have

  • C++ and CUDA programming: custom kernels, memory/occupancy tuning, profiling with Nsight Systems/Compute.

  • Sidero Omni in production; multi-provider GPU clouds (Nebius, CoreWeave, Lambda).

  • Inference runtimes (Triton, vLLM, TensorRT) and batching economics.

  • FinOps for GPU fleets: cost per generation, commitment planning.

  • Experience building internal platforms or AIOps tooling.

Success in 6 months

  • Training throughput measurably up (tokens/steps per GPU-hour), failed-run recovery in minutes, not hours.

  • Fleet utilization sustainably >95%; idle-allocation findings trend to zero.

  • GPU node MTTR cut in half; every capacity change traceable through git.

  • New GPU capacity - provider to serving traffic - lands in days, fully through GitOps.

What we offer:

  • Competitive base salary in USD

  • Equity: participation in the company’s stock option program, giving you the opportunity to share in the company’s long-term growth.

  • On-site role in our Almaty office (we will relocate you from anywhere).

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Almaty
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
In office • 8+ years exp • Bachelor's Degree • Hyderabad
Python
Databases
Amazon Redshift
Apache Kafka
AI/ML
Spark
DevOps
AWS
AWS Lambda
CloudFormation
Amazon EventBridge
Amazon S3
Apply
$100k – $252k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Herzliya
Java
Databases
Apache Kafka
MySQL
PostgreSQL
RabbitMQ
AI/ML
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$24k – $49k per year (Estimated) • Equity • In office • Full-Time • 5+ years exp • Almaty
AI/ML
EU AI Act
Cybersecurity
GDPR
Apply
$22k – $54k per year (Estimated) • Equity • In office • Full-Time • Almaty
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
DeepSpeed
Multimodal AI
PyTorch
Reinforcement Learning
Triton
FSDP
InfiniBand
Megatron-LM
NCCL
NVLink
Mixture of Experts
Apply
$140k – $180k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
Apply
$165k – $230k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
AI/ML
Fine-tuning
Multimodal AI
Prompt Engineering
Reinforcement Learning
Post-training
Recommender Systems
SFT
Apply
$23k – $48k per year (Estimated) • In office • Full-Time • 6+ years exp • Almaty
Apply
$21k – $24k per year • In office • Almaty
Management
Confluence
Jira
Apply
$8.3k – $14k per year (net) • Remote • Part-Time • Almaty
Design
Figma
Apply
$13k – $27k per year (net) • In office • Full-Time • 2+ years exp • Almaty
Dart
Node JS
JavaScript
Dart
Bloc
Riverpod
AI/ML
LiveKit
Mobile
Agora
Firebase
Flutter
HealthKit
ML Kit
DevOps
Rest API
WebRTC
Apply
$17k – $58k per year (Estimated) • Remote • Full-Time • Almaty
Java
Scala
Scala
Akka
Databases
Apache Kafka
AI/ML
Flink
DevOps
Amazon Kinesis
AWS
Azure
Azure DevOps
Apply
$4.5k – $24k per year (Estimated) • In office • Full-Time • 2+ years exp • Almaty
SQL
Databases
MS SQL
PostgreSQL
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.