655,248open jobs
38,129companies
91,967added this week
Browse all
Salary
$76k – $172k per year (Estimated)
Location
Remote/Hybrid (Madrid, Spain)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Dlocal is a leading provider of payments solutions for emerging markets. We help businesses expand their reach into new markets by providing them with the tools they need to accept local payments. We offer a wide range of payment methods, including credit cards, debit cards, bank transfers, and mobile wallets.

What's the Opportunity?

You will join the AI Lab, a team whose mission is to validate high-value emerging AI and automation technologies and de-risk their adoption across dLocal. This is a rare opportunity to work at the frontier of applied AI in fintech: running rigorous experiments on the latest models and tools, and turning results into decisions that shape how a global payments company adopts AI.

This is a senior individual-contributor role. It does not require direct people management, but it carries significant technical influence within the Lab and across the teams that consume its work.

As a Staff AI Engineer in the AI Lab, you own technology scouting, prototyping, and evaluation for dLocal. You will run instrumented spikes and benchmarking on emerging AI technologies, produce clear recommendations for stakeholders across engineering, business, legal and IT based on your prototypes, and coordinate hand-offs to the teams that take validated technologies into production.

For promising technologies, you will help define the patterns, guardrails, and technical requirements needed for adoption and coordinate the hand-off to the engineering teams responsible for productionizing them.

What will I be doing?

    Technology Scouting & Evaluation

  • Run short, instrumented spikes and benchmarking on new models, tools and frameworks: LLMs, agentic systems, vector databases, orchestration frameworks, copilots, assistants and more.

  • Compare vendor and open-source options, documenting trade-offs across quality, cost, latency, security and integration complexity.

  • Deliver concise decision memos with clear recommendations: adopt, watch, or avoid.

  • Evaluation Harnesses & Sandboxes

  • Design and maintain evaluation environments (e.g. datasets, prompts, scenarios, telemetry) to test models under realistic constraints.

  • Build automation and tooling to measure quality, robustness, latency and cost, including regression tracking over time.

  • Ensure every evaluated technology has benchmark coverage and a documented risk and limitations view.

  • Prototyping & Technical Validation

  • Build enough of a system to understand how a technology behaves under realistic conditions, not just in vendor demos or isolated examples.

  • Explore architecture, integration patterns, operational constraints, security boundaries, and failure modes through working prototypes.

  • Determine what must be true for a proof of concept to become a viable production capability.

  • Prefer focused prototypes that answer specific technical questions over prematurely building production systems.

  • Recommendations, Readiness & Hand-offs

  • Translate technical findings into clear decision memos for both technical and non-technical stakeholders.

  • For validated technologies, produce readiness guidance covering recommended patterns, guardrails, known limitations, operational considerations, and integration requirements.

  • Coordinate hand-offs to the engineering teams responsible for productionization.

  • Support those teams during the transition when deep context from the evaluation is required, without becoming the permanent owner of the resulting system.

  • Track what happens after Lab recommendations and use those outcomes to improve future evaluation methods.

  • Governance, Risk & Standards

  • Work with Security, Legal, Compliance and other AI teams to document risk assessments, mitigations and governance recommendations for each evaluated technology.

  • Maintain checklists, decision templates and lightweight standards reusable across evaluations and by partner teams.

  • Incorporate learnings from third-party AI tooling already in use, such as external copilots and the AWS AI suite, into adoption guidelines.

  • Collaboration, Mentoring & Community

  • Partner with other AI teams and domain teams to ensure clear boundaries and smooth collaboration.

  • Participate in hiring as a technical evaluator and culture champion.

  • Mentor engineers in the Lab and adjacent teams on evaluation methods, benchmarking and experimental design.

  • Share knowledge through internal write-ups, tech talks and occasional external meetups and conferences.

What skills do I need?

    Technical depth

  • 8+ years of software engineering experience, including significant experience operating at senior or Staff-level scope.

  • Deep hands-on experience building and evaluating systems based on LLMs and modern AI tooling.

  • Strong software engineering fundamentals and the ability to rapidly build high-quality experimental systems.

  • Experience building agentic or multi-step AI systems involving tool use, orchestration, state, retrieval, or external integrations.

  • Strong knowledge of cloud infrastructure, preferably AWS, and the ability to run experimental workloads securely and cost-consciously.

  • Experience with observability, telemetry, testing, and benchmarking of complex systems.

  • Ability to reason about system architecture, reliability, scalability, asynchronous workflows, and distributed components where relevant.

  • Track record of designing experiments or benchmarks that influenced meaningful technical decisions.

  • Benchmarking & evaluation

  • Track record designing and running benchmarks that compare AI models and tools under real constraints.

  • Experience constructing evaluation datasets: task selection, labelling, holdout discipline, and keeping a set useful as models improve.

  • Working knowledge of LLM-as-judge methods and their failure modes, alongside human evaluation, inter-annotator agreement, and a view on when each is appropriate.

  • Able to reason about statistical significance on small samples, and to state confidence honestly rather than over-reading a result.

  • Familiarity with regression tracking, telemetry and versioning, so that a result stays reproducible months later.

  • Decision-making

  • Able to turn ambiguous "we should try this new thing" ideas into well-scoped evaluation plans with clear hypotheses and metrics.

  • Comfortable making trade-off calls across quality, latency, cost and vendor lock-in, and documenting them clearly.

  • Experience writing short, opinionated decision memos that help others move fast.

  • Collaboration & communication

  • Can explain technical results to non-specialists in concrete, concise terms.

  • Experience working with platform, product and operations teams to align evaluations with real use cases.

  • Able to influence without authority, aligning teams around shared standards and guardrails.

  • Mindset

  • Curious and biased toward experimentation, combined with disciplined measurement and risk awareness.

  • Comfortable in a small, high-leverage team without embedded PMs. You structure your own work and keep stakeholders informed.

  • Builder attitude: you prefer reusable tools, templates and playbooks over one-off work.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
655,248 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Madrid
Equity • Remote • Full-Time • 8+ years exp
Python
JavaScript
TypeScript
Bash
AI/ML
AI Agents
Frontend
React.js
DevOps
Rest API
GCP
Azure
AWS
Apply
$69k – $147k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Santiago
AI/ML
AI Agents
Marketing
Salesforce
Apply
$94k – $177k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • United States
AI/ML
AI Agents
Analytics
Tableau
Looker
Management
Asana
Monday.com
Confluence
Jira
Apply
Sr. Data Engineer 1 day ago
$120k – $160k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
Python
Python
pySpark
AI/ML
Unstructured.io
Spark
Embeddings
Scikit-learn
NLP
TensorFlow
NumPy
LLM
RAG
Anomaly Detection
Context Engineering
LLM Evaluation
LLM Guardrails
Analytics
Matplotlib
ETL/ELT
Apply
$29k – $67k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Chihuahua
JavaScript
TypeScript
SQL
C#
Node JS
C#
ASP.NET Core
xUnit
Databases
MS SQL
AI/ML
Llama
Frontend
Angular
DevOps
Rest API
Azure DevOps
Azure
CI/CD
Git
AWS
AWS Lambda
Amazon EC2
Amazon S3
IAM
Amazon CloudWatch
API Gateway
Cybersecurity
SonarQube
Analytics
SSIS
Management
Agile
Scrum
Apply
$76k – $197k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Montevideo
Cybersecurity
GDPR
Management
Outlook
Apply
Remote/Hybrid • Full-Time • Buenos Aires
DevOps
AWS
Analytics
Power BI
Apply
Account Manager 22 days ago
$19k – $53k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Buenos Aires
Apply
Technical Lead - Java 27 days ago
$55k – $139k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Barcelona
Java
Java
Hibernate
AI/ML
Anomaly Detection
DevOps
GCP
Datadog
Dynatrace
Prometheus
AWS
Grafana
Amazon CloudWatch
Apply
$45k – $110k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Madrid
Python
SQL
Databases
Databricks
AI/ML
AI Agents
Apply
$16k – $44k per year (Estimated) • In office • Part-Time • 2+ years exp • Madrid
Apply
$16k – $44k per year (Estimated) • In office • Full-Time • 2+ years exp • Madrid
Apply
$138k – $208k per year • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Lisbon • Munich • Amsterdam • Brussels • Madrid
Apply
$26k – $50k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Madrid
Management
Google Docs
Marketing
LinkedIn
Instagram
Apply
$50k – $94k per year (Estimated) • In office • 5+ years exp • Madrid
Management
Jira
Agile
Apply
See all jobs
This is one of many
655,248 more open roles from verified company boards, updated every day.