745,032open jobs
44,747companies
107,631added this week
Browse all
Salary
≈ $16k – $36k per year (Estimated)
Location
In office (Sofia)
Seniority
Junior · 2+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. DeepInfra scores C on the Alion truth index.

Overview
Company
Impact
Profile match
DeepInfra is a company founded in 2022 that runs a low-cost inference cloud for open machine learning models. Developers call hosted language, embedding, speech and image models through a simple interface and pay per token or per second without managing graphics hardware. Its focus on cheap serving of open-weight models made it popular with startups building on Llama, Mistral and Qwen.

About DeepInfra

DeepInfra is building the infrastructure layer for the next generation of AI. We believe open-source models are the future, and companies should have full control over their AI stack without being locked into proprietary providers.

Our inference platform serves trillions of tokens every week across hundreds of production workloads. We build everything from GPU infrastructure to the API layer because every millisecond matters.

We're looking for a Customer Success Engineer to own the technical relationship with the developers and companies running production workloads on our platform.

Why this role matters

When a customer's inference breaks, slows down, or costs more than they expected, you're the person who finds out why and fixes it. You have the technical depth to solve it and the access to do it yourself, instead of routing the ticket to someone else.

You'll work directly with customers ranging from solo developers to enterprises serving millions of requests a day. You'll debug real inference problems, tune their usage, and feed what you learn back into the product. Most investigations run through the terminal: you reproduce the issue, read logs and metrics, find the root cause, and explain it in writing.

This role suits someone who enjoys owning hard technical problems end to end, often with a real customer waiting on a deadline. You're comfortable working on your own and want a say in how our platform supports its customers. We hire for reasoning structured debugging, terminal fluency, and clear writing.

What You'll Do

  • Own customer issues end-to-end. This includes to triage, reproduce, root-cause, and resolve technical problems across LLM, embedding, image, TTS, and ASR workloads.
  • Debug production inference issues using logs, metrics, and dashboards: latency and TTFT regressions, rate limiting and 429s, error spikes, model behavior changes, and API integration bugs.
  • Handle rate limit and capacity requests with real data - analyze a customer's traffic pattern, concurrency, and token throughput, and recommend what actually fixes their problem.
  • Resolve billing, usage, and quota questions by tracing per-request ground truth, not guesswork.
  • Help customers use the platform well: batching, streaming, prompt caching, structured outputs, tool calling, retries and backoff, and choosing the right model and deployment type for their workload.
  • Advise on capacity planning and dedicated deployments - GPU sizing, throughput targets, and cost per million tokens.
  • File, prioritize, and drive well-documented bugs and feature requests with Engineering, and follow them through to a customer-facing answer.
  • Turn recurring issues into permanent fixes: documentation, internal runbooks, and product changes that stop the ticket from coming back.
  • Be the voice of the customer internally - surface patterns to Product and Engineering before they become churn.

What You Bring

  • 2+ years in Customer Success Engineering, Support Engineering, Solutions Engineering, Technical Account Management, or a similaole - or software engineering experience and a genuine pull toward customer work.
  • A degree or equivalent hands-on background in Computer Science or Engineering.
  • Strong technical fundamentals: you're comfortable with HTTP and REST APIs, status codes, headers, streaming responses, authentication, and reading someone else's integration code to spot what's wrong.
  • Working proficiency in Python or a similar language - enough to reproduce a customer's issue, write a script, and hand back a working snippet.
  • Comfort on the command line and with logs, metrics, and dashboards (Grafana, Prometheus, Loki, or equivalents).
  • The ability to take a vague report - "it got slow," "the output looks worse," "I was charged too much" - and turn it into a specific, evidence-backed root cause.
  • Excellent written communication: clear, concise, technically precise, and respectful under pressure.
  • Strong judgment about when to solve something yourself and when to escalate, and the follow-through to close the loop either way.
  • Ability to work independently and prioritize a queue in a fast-moving startup environment.

Bonus

  • Exposure to AI/ML infrastructure, such as inference serving, GPUs, or model deployment.
  • Experience supporting a developer-facing API or a usage-based billing product.
  • Experience being an early or first Customer Success Engineer and building the function's processes and tooling.

Why DeepInfra

  • Own the technical customer experience for a platform serving trillions of tokens a week.
  • Work on real inference problems at scale, with the access and context to solve them properly.
  • Join a small, high-performing team where your work ships fast and customers feel it immediately.
  • Help shape how companies build on some of the world's leading open-source AI models.

How we work

Three traits define the people who thrive here, and this role leans on all three.

Initiative. We take ownership and step in where we can add value. Whether it’s starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.

Drive. We’re energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving - because solving meaningful challenges is what motivates us.

Grit. Things don’t always work on the first try - and that’s expected. We stay persistent, adapt quickly, and learn as we go. We take setbacks seriously, but not personally, and use them to get better.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
745,032 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Customer Success
Similar stack
Same company
Sofia
≈ $19k – $48k per year (Estimated) • Remote (Bulgaria, ET hours) • 4+ years exp • Bachelor's Degree • Sofia
Management
Google Sheets
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
≈ $16k – $42k per year (Estimated) • In office • Full-Time • Sofia
Apply
≈ $85k – $164k per year (Estimated) • In office • Full-Time • 10+ years exp • London • Birmingham • Glasgow • Northampton
PHP
PHP
WooCommerce
DevOps
Rest API
Apply
≈ $30k – $81k per year (Estimated) • In office • Bengaluru
Python
SQL
Apply
≈ $31k – $83k per year (Estimated) • In office • Bengaluru
Python
SQL
AI/ML
Machine Learning
Apply
≈ $39k – $67k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • Toulouse
Python
SAS
Apply
$92k – $139k per year • Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Cincinnati
Python
SQL
Scala
AI/ML
Machine Learning
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Sofia
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
Multimodal AI
Diffusers
Transformers
TensorFlow
PyTorch
CUDA
NCCL
Edge AI
DevOps
Git
Apply
In office • Full-Time • Bachelor's Degree • Sofia
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
SciPy
Multimodal AI
Diffusers
Transformers
TensorFlow
NumPy
PyTorch
CUDA
NCCL
Edge AI
DevOps
Git
Management
Agile
Apply
Remote (Bulgaria) • Full-Time • 3+ years exp • Bachelor's Degree
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
Multimodal AI
Diffusers
Transformers
TensorFlow
PyTorch
CUDA
NCCL
Edge AI
DevOps
Git
Apply
$41k – $71k per year • Remote (Bulgaria) • Full-Time • Bachelor's Degree
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
SciPy
Multimodal AI
Diffusers
Transformers
TensorFlow
NumPy
PyTorch
CUDA
NCCL
Edge AI
DevOps
Git
Management
Agile
Apply
Remote (Bulgaria) • Internship • Bachelor's Degree
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
SciPy
Multimodal AI
Diffusers
Transformers
TensorFlow
NumPy
PyTorch
CUDA
OpenAI
NCCL
Edge AI
Machine Learning
DevOps
Git
Management
Agile
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Management
Service Desk
Apply
≈ $16k – $42k per year (Estimated) • In office • Full-Time • Sofia
Apply
≈ $17k – $46k per year (Estimated) • Remote (Bulgaria) • Full-Time • Sofia
Apply
See all jobs
This is one of many
745,032 more open roles from verified company boards, updated every day.