671,804open jobs
39,060companies
102,874added this week
Browse all
Salary
$54k – $144k per year (Estimated)
Location
Remote (United States, Argentina, Colombia, Mexico, Chile, Dominican Republic)
Employment
Full-Time
Overview
Company
Impact
Profile match
Azumo is a leading software development company that specializes in nearshore services. The company offers a range of solutions, including software development, dedicated teams, staff augmentation, and virtual CTO services. Azumo is particularly known for its expertise in artificial intelligence, mobile app development, data engineering, and cloud services.

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring an AI Engineer to own what those systems do once they are live: retrieval pipelines, tool-using agents, evaluation harnesses, and the guardrails that keep them dependable in front of real users. The role is fully remote across Latin America, aligned to your client's working day.

You will not be building demos. Azumo has shipped production AI since 2016, and the work here starts where the prototype ends, making a system reliable, measurable, and affordable enough to put in front of customers.

Where this role sits

Azumo's engineering organization is built around four lanes. The Data Engineer lane owns pipelines, storage, and the retrieval layer. The Data Scientist lane owns the question and the method. The Software Engineer lane owns AI-augmented product delivery. This role is the AI Engineer lane, and it owns production behavior.

One question places the boundary: when the output is wrong, whose problem is it? "The method was inappropriate" is a Data Scientist question. "The system did the wrong thing with an appropriate method" is yours.

What you will build

  • Retrieval systems. Chunking and embedding pipelines, hybrid search, reranking, and evaluation of retrieval quality, built on pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search.
  • Agentic workflows. Stateful multi-step execution, tool calling, MCP servers, structured output enforcement, context-window management, deterministic fallbacks, and human-in-the-loop gates for the decisions that need one.
  • Evaluation. Test sets that reflect the decision the system is actually making, model-as-judge scoring, regression tracking across prompt and model changes, and honest error analysis. If a change made the system better, you should be able to prove it.
  • Reliability and safety. Prompt-injection defense, output validation, guardrails, PII handling, and graceful degradation when a model or tool call fails.
  • Production operation. Containerized deployment on Azure or AWS, CI/CD, observability, and explicit latency, cost, and token budgets that you own rather than discover after the invoice.
  • Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.

How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Apply at https://azumo.com/join-our-team or write to us at [email protected].

Requirements

Basic qualifications

  • 4+ years building and shipping production software, with a modern backend language (e.g., Python) as your primary focus, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, API design, and async programming.
  • Demonstrable production experience with LLM-based systems: retrieval-augmented generation, function and tool calling, structured output enforcement, and prompt design as an engineering discipline rather than trial and error.
  • Hands-on work with vector and retrieval infrastructure such as pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search, including the retrieval-quality problems that come with it.
  • Experience with agent frameworks and tooling: LangGraph, LangChain, CrewAI, the Model Context Protocol (MCP), or native Python execution loops. We care that you have shipped a stateful, tool-using system, not which framework you used.
  • You have built an evaluation suite for an LLM system. Test-set design, model-as-judge or equivalent scoring, and regression tracking with Langfuse, Ragas, LangSmith, or something you wrote yourself. This is the requirement we screen hardest on.
  • Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).
  • Working discipline around latency, token cost, and throughput. You can explain what a feature costs to run and what you did about it.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.
  • Clear written and spoken English, B2/C1 or above, and the confidence to explain a technical trade-off directly to a client.
  • Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.

Preferred qualifications

  • Fine-tuning and adaptation of open-weight models: LoRA, QLoRA, PEFT, and a clear view of when fine-tuning is the wrong answer.
  • Self-hosted or open-weight inference and serving, and the cost and latency trade-offs against hosted APIs.
  • Multimodal systems covering vision, speech, or document understanding alongside text.
  • Security work specific to LLM systems: prompt-injection testing, red-teaming, and output sanitization.
  • Delivery under a compliance regime such as SOC 2 or HIPAA.
  • Streaming, real-time, or high-throughput inference workloads.
  • Contributions to open-source AI libraries, published technical writing, or active participation in the AI engineering community.

Benefits

  • Paid time off (PTO)·
  • U.S. Holidays·
  • AI Training and certifications·
  • Mentored career development·
  • Profit sharing·
  • $US remuneration.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
671,804 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Software Developer 2 hours ago
$70k – $80k per year • Remote • Full-Time • Bachelor's Degree • Fort Collins
Java
SQL
C#
C#
ASP.NET Core
WPF
Databases
MS SQL
AI/ML
Copilot
Cursor
ChatGPT
Claude Code
Prompt Engineering
AI Agents
OpenAI Codex
Agentic Workflows
Mobile
MVVM
DevOps
Azure DevOps
Azure
CI/CD
Git
GitHub
Cybersecurity
Keycloak
Management
Agile
Scrum
Apply
Remote/Hybrid • Full-Time • 14+ years exp • Bengaluru
Databases
Amazon Redshift
AI/ML
Copilot
Cursor
Claude Code
AWS Bedrock
LLM
RAG
Amazon SageMaker
Agentic Workflows
DevOps
AWS
FinOps
AIOps
Amazon S3
Apply
$73k – $161k per year (Estimated) • Remote • 2+ years exp
SQL
AI/ML
Copilot
Cursor
Claude
ChatGPT
Analytics
A/B Testing
Management
Slack
Agile
Apply
$58k – $117k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Dubai
Python
SQL
Analytics
Power BI
Apply
$74k – $115k per year • Remote/Hybrid • Full-Time • 5+ years exp • Vancouver
Python
JavaScript
SQL
Apex
Perl
Apex
MuleSoft
Databases
MySQL
PostgreSQL
Oracle
ActiveMQ
MS SQL
AI/ML
Model Context Protocol
Fine-tuning
Function Calling
AI Agents
RAG
Supervision
A2A
Tool Use
DevOps
CI/CD
Analytics
ETL/ELT
Talend
Pentaho
Management
ServiceNow
UiPath
Agile
Scrum
QA
Swagger
Apply
$67k – $152k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Python
Databases
Pinecone
LanceDB
AI/ML
Copilot
Cursor
LangGraph
LangChain
Claude
Claude Code
Model Context Protocol
Function Calling
AI Agents
NLP
Langfuse
Ragas
CrewAI
Hallucination
LLMOps
Human-in-the-Loop
Structured Outputs
LLM Guardrails
DevOps
Azure
CI/CD
Git
AWS
Docker
GitHub
Apply
$99k – $171k per year (Estimated) • Remote • Full-Time • 4+ years exp • San Francisco
Java
Java
Spring Boot
Hibernate
Apply
$49k – $118k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Buenos Aires
SQL
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
Amazon Redshift
Azure SQL Database
DevOps
Azure
Apply
$47k – $114k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Buenos Aires
Python
SQL
Databases
Databricks
AI/ML
Airflow
DevOps
AWS
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
671,804 more open roles from verified company boards, updated every day.