706,097open jobs
41,961companies
98,408added this week
Browse all
Salary
$42k – $115k per year (Estimated)
Location
Remote (Argentina, Brazil, Colombia, Mexico, Chile, Peru)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Azumo is a leading software development company that specializes in nearshore services. The company offers a range of solutions, including software development, dedicated teams, staff augmentation, and virtual CTO services. Azumo is particularly known for its expertise in artificial intelligence, mobile app development, data engineering, and cloud services.

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring a Data Engineer to own the layer everything else depends on: ingestion and transformation pipelines, storage and warehouse design, and the retrieval infrastructure that AI systems query. The role is fully remote across Latin America, aligned to your client's working day.

You will not be building pipelines that work until the schema changes. Azumo has shipped production data systems since 2016, and the work here is judged downstream: whether a model can be trained on what you deliver, whether a retrieval query returns the right passage, whether a number survives being questioned by the client.

Where this role sits

Azumo's engineering organization is built around four lanes. The Data Scientist lane owns the question and the method. The AI Engineer lane owns production behavior. The Software Engineer lane owns AI-augmented product delivery. This role is the Data Engineer lane, and it owns pipelines, storage, and the retrieval layer.

One question places the boundary: when the output is wrong, whose problem is it? "The data was missing, stale, or wrong by the time it arrived" is yours."The system did the wrong thing with data that was correct" is the AI Engineer's.

Not quite your profile? Check our other openings:

- If you decide what to measure and which method answers it - Data Scientist

- If you own how an AI system behaves in production - AI Engineer

- If you ship product software with agents in your toolchain -AI-Augmented Software Engineer

- If you've done all of the above and answered to the client directly - Forward Deployed Engineer

What you will build

- Ingestion and transformation pipelines. Batch and streaming ingestion on Spark, Kafka, dbt and Airflow, with idempotency, backfills, schema evolution and late-arriving data handled by design rather than by hand.

- Storage and modeling. Warehouse and lakehouse design on Snowflake, BigQuery, Redshift or Databricks, with partitioning, file layout and query cost treated as engineering decisions.

- The retrieval layer. The chunking, embedding and indexing pipelines that feed RAG systems on pgvector, Pinecone, Qdrant or Azure AI Search, and the freshness, deduplication and permission problems that come with them.

- Data quality as a contract. Tests, expectations, lineage and alerting. If a pipeline is wrong, the people downstream should hear it from you and not from the client.

- Sensitive data by default. PII classification, masking, row and column level access, retention and deletion, and an audit trail that holds up when a client asks who read what.

- Production operation. Containerized deployment on Azure or AWS, CI/CD, orchestration, observability, and explicit cost and runtime budgets that you own rather than discover after the invoice.

- Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.

How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Apply at https://azumo.com/join-our-team or write to us at [email protected].

Requirements

Basic qualifications

- 5+ years building and operating production data pipelines, with Python and SQL as your primary languages, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, and orchestration.

- Deep expertise in designing and building data warehouses or lakehouses, including dimensional modeling, incremental processing, and the cost and performance trade-offs behind each choice.

- Distributed processing at production scale with Spark, Kafka, Flink or equivalent, including the failure modes that only appear under load.

- Orchestration as an engineering discipline rather than a cron replacement: Airflow, Dagster or Prefect, with retries, idempotency and backfill strategy you can defend.

- Transformation under version control, with tests and lineage: dbt or something you built yourself.

- Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).

- Working discipline around pipeline cost, runtime and throughput. You can explain what a pipeline costs to run and what you did about it.

- Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.

- Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.

- Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.

Preferred qualifications

- Vector and retrieval infrastructure: pgvector, Pinecone, Qdrant, FAISS or Azure AI Search, and the retrieval-quality problems that come with it.

- Streaming, real-time or high-throughput workloads.

- Experience with cloud-based managed services like Airflow, Glue, Elastic stack, Amazon Redshift, Snowflake, BigQuery, Azure SQL Db, EMR, Databricks.

- Prior experience with notebooks using Jupyter, Google Collab, or similar.

- Delivery under a compliance regime such as SOC 2 or HIPAA.

- Contributions to open-source data libraries, published technical writing, or active participation in the data engineering community.

Benefits

  • Paid time off (PTO)
  • U.S. Holidays
  • AI Training
  • Mentored career development
  • Profit sharing
  • $US remuneration
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
706,097 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Data Scientist 2 hours ago
$43k – $116k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
AI/ML
Copilot
Cursor
Weights & Biases
Claude Code
LoRA
MLFlow
Vertex AI
Fine-tuning
Scikit-learn
Multimodal AI
PEFT
QLoRA
Kubeflow
Transformers
NumPy
PyTorch
LLM
OpenAI
Anthropic
Amazon SageMaker
OpenAI Codex
Machine Learning
DevOps
GCP
Azure
Git
AWS
Cybersecurity
SOC 2
HIPAA
Apply
$90k – $100k per year • In office • Full-Time • 2+ years exp • Boston
Python
SQL
Apply
$94k – $194k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Toronto
Python
JavaScript
Java
TypeScript
SQL
Node JS
Databases
Databricks
Firestore
Google BigQuery
BigQuery
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Vertex AI
Prompt Engineering
CrewAI
Gemini
LLM
RAG
Google ADK
Multi-Agent Systems
Machine Learning
Frontend
Angular
React.js
Mobile
Firebase
DevOps
Rest API
GCP
Azure DevOps
Dynatrace
Azure
CI/CD
Kubernetes
Google Cloud Run
Cybersecurity
SIEM
Management
Agile
Apply
$11k – $12k per year • In office • 1+ year exp • Zelenograd
Python
JavaScript
SQL
Python
Flask
SQLAlchemy
Databases
PostgreSQL
Frontend
Bootstrap
JQuery
DevOps
Docker
Linux
Apply
$72k – $102k per year • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • United States
Java
SQL
Databases
Amazon Redshift
Mobile
JUnit
DevOps
CI/CD
Jenkins
Git
AWS
AWS Lambda
Amazon S3
Amazon CloudWatch
Analytics
ETL/ELT
QA
TestNG
Selenium
Apply
$46k – $122k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
AI/ML
Claude Code
OpenAI
Anthropic
OpenAI Codex
DevOps
Terraform
GitHub Actions
Azure
CI/CD
AWS
Docker
Bicep
Cybersecurity
SOC 2
HIPAA
Apply
Data Scientist 2 hours ago
$43k – $116k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
AI/ML
Copilot
Cursor
Weights & Biases
Claude Code
LoRA
MLFlow
Vertex AI
Fine-tuning
Scikit-learn
Multimodal AI
PEFT
QLoRA
Kubeflow
Transformers
NumPy
PyTorch
LLM
OpenAI
Anthropic
Amazon SageMaker
OpenAI Codex
Machine Learning
DevOps
GCP
Azure
Git
AWS
Cybersecurity
SOC 2
HIPAA
Apply
$53k – $141k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
JavaScript
AI/ML
Claude Code
Model Context Protocol
OpenAI
Anthropic
OpenAI Codex
Frontend
Vue.js
React.js
DevOps
Rest API
Terraform
GitHub Actions
Azure
CI/CD
Git
Docker
Bicep
Cybersecurity
SOC 2
HIPAA
Apply
$54k – $144k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
Databases
PostgreSQL
pgvector
Pinecone
FAISS
Qdrant
AI/ML
Copilot
Cursor
LangGraph
LangChain
Claude Code
LoRA
Model Context Protocol
Fine-tuning
Multimodal AI
Function Calling
Langfuse
LangSmith
PEFT
QLoRA
Ragas
Transformers
CrewAI
LLM
RAG
Reranking
Hybrid Search
OpenAI
Anthropic
OpenAI Codex
Human-in-the-Loop
Structured Outputs
LLM Evaluation
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
Terraform
GitHub Actions
Azure
CI/CD
Git
AWS
Docker
Bicep
Cybersecurity
SOC 2
HIPAA
Apply
$64k – $162k per year (Estimated) • Remote • Full-Time
AI/ML
Copilot
Claude
ChatGPT
Prompt Engineering
Machine Learning
DevOps
CI/CD
Git
Management
Notion
Jira
Agile
Scrum
Kanban
Apply
See all jobs
This is one of many
706,097 more open roles from verified company boards, updated every day.