397,589open jobs
13,889companies
77,448added this week
Browse all
Location
In office (Buenos Aires)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Software Mind is a Polish technology company with roots in the Ailleron group that builds and operates software for clients in telecommunications, financial services, healthcare, media and travel. Its engineers work as dedicated teams on custom product development, cloud and platform engineering, data and artificial intelligence, and modernisation of legacy systems, delivered from centres in Poland, Romania, Moldova, Chile and the United States. Headquartered in Krakow, the company has grown through acquisitions of engineering firms in Europe and Latin America and now employs several thousand specialists.

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!

Project - the aim you'll have

Our client builds an AI copilot for process engineers in oil refineries and chemical plants: a natural-language interface where engineers ask questions about live plant data - equipment, sensor tags, process trends - and get grounded, chart-backed answers. The users are experienced engineers who are rightly skeptical of AI: in this domain, a fabricated number or a silent wrong assumption has real cost. The product wins or loses on whether the agent can be trusted.

This role owns the trust layer of that agent inside a large, active Python codebase. It is not feature work with an LLM endpoint bolted on. The work is the mechanics of agent reliability: making the agent say "I don't know" instead of inventing, surfacing every assumption it makes so the user can correct it, grounding every claim in actual data, holding output quality through model migrations, and keeping latency acceptable while doing all of the above.

To make the day-to-day concrete, this is what the engineer currently in this seat shipped in the last four months (all of it flag-gated, in small PRs, reviewed async daily by a team spread across the US and Australia):

- An assumption auditor: detects the silent assumptions the agent makes when answering (which equipment, which time window), validates them via multi-draw consensus, and surfaces them in the UI as correctable chips - the engineer can fix an assumption and rerun the analysis.

- A grounding auditor that catches reports fabricated from empty data feeds before they reach the user.

- An adversarial reviewer sidecar that critiques generated charts for correctness before display.

- Successive frontier-model evaluations (loop behavior, directive adherence, regression on a replay harness) that decided when to flip the product's default model - including, twice, deciding NOT to flip.

- Hardening of a plant-exploration tool against hallucinating structure that the data does not support.

- A latency fix: a narration side-loop was inflating query response times; capped it and made it best-effort.

- A concurrency fix making a shared data-reset path atomic, eliminating intermittent production read errors.

If reading that list is more interesting to you than building another CRUD feature, this role is for you.

Expectations - the experience you need

  • Strong Python in large, shared, evolving backend codebases: you will work daily in code you didn't write, alongside people committing to it every day.
  • You have shipped an LLM-based feature to production AND built an evaluation that changed a real decision (a model choice, a prompt rollback, a killed feature).
  • Production debugging from symptom to confirmed root cause: latency spikes, concurrency errors, failures that produce no log line.
  • Prompt work treated as engineering: measured adherence, regression testing against a fixed case set - not vibes.
  • Comfort with feature-flag discipline and staged rollouts (default-off, soak, flip), small PRs, and mostly-async collaboration across US and Australia time zones.
  • High autonomy: problems arrive ambiguous ("the agent feels slow", "the engineers don't trust the numbers") and you turn them into scoped, verifiable fixes without waiting for a spec.
  • Direct, precise written English.

Nice to have

  • GCP (Vertex AI in particular); AWS/Azure acceptable.
  • Observability tooling (tracing, structured logging, latency percentiles you actually watched).
  • Experience with charting/plotting pipelines (matplotlib or similar) feeding a UI.
  • Industrial, process, or time-series data domain experience.
  • Heavy AI-tooling development workflow (Claude Code or similar) - the team works this way.

What you will do

  • Investigate agent misbehavior reported from real customer plants and turn each case into a diagnosis, a fix, and a regression test.
  • Build and extend the evaluation harnesses that gate prompt changes and model migrations.
  • Add reliability mechanisms to the agent: assumption surfacing, grounding checks, output-quality reviewers.
  • Diagnose and resolve cross-cutting performance and concurrency issues.
  • Raise code quality in the areas you touch, within the team's review conventions.

Our Benefits

  • Educational resources
  • Flexible schedule and Work From Anywhere
  • Referral Program
  • Supportive and chill atmosphere

We are accepting applications from LATAM countries

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
397,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Buenos Aires
$130k – $180k per year • Remote • 10+ years exp • Master's Degree
Python
AI/ML
AI Agents
DeepSpeed
DPO
Fine-tuning
FSDP
Hallucination
Knowledge Graph
LangChain
LangGraph
LlamaIndex
LLM
LoRA
Multimodal AI
NLP
PEFT
PPO
PyTorch
QLoRA
RAG
Ray
RLHF
SFT
Synthetic Data
Transformers
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vector
Apply
Algorithm Engineer 1 hour ago
Remote/Hybrid • Full-Time • 3+ years exp • Netanya
C++
MATLAB
Python
AI/ML
Computer Vision
Quantization
DevOps
CI/CD
HPC
Analytics
A/B Testing
Apply
$100k – $170k per year • Remote • 12+ years exp • Bachelor's Degree
Python
DevOps
IAM
Terraform
Cybersecurity
CIS Benchmarks
HIPAA
ISO 27001
PCI DSS
SOC 2
Zero Trust
Apply
$100k – $113k per year • Remote • 7+ years exp • Bachelor's Degree
Python
DevOps
IAM
Terraform
Cybersecurity
CIS Benchmarks
HIPAA
ISO 27001
PCI DSS
SOC 2
Zero Trust
Apply
$86k – $103k per year • Remote • 8+ years exp • Bachelor's Degree
Bash
Go
Python
DevOps
Ansible
ArgoCD
AWS
Azure
CI/CD
GCP
GitOps
Helm
Istio
Jenkins
Kubernetes
Linkerd
OpenShift
Red Hat
Service Mesh
Tekton
Terraform
Cybersecurity
HIPAA
PCI DSS
SOC 2
Apply
$63k – $96k per year (Estimated) • Remote • Full-Time • 5+ years exp • Warsaw
Python
SQL
Databases
Databricks
PostgreSQL
DevOps
AWS
Azure
CI/CD
Docker
Rest API
Terraform
Apply
$68k – $104k per year (Estimated) • Remote • Full-Time • Warsaw
Python
DevOps
Ansible
AWX
Configuration Management
Grafana
IAM
Loki
OpenStack
Prometheus
Ubuntu
Apply
$63k – $102k per year (Estimated) • Remote • Full-Time • Warsaw
Python
SQL
Databases
Apache Kafka
Databricks
PostgreSQL
Kafka
DevOps
Azure
Azure AKS
Bicep
CI/CD
Kubernetes
Terraform
Apply
Remote • Full-Time • 4+ years exp • San José
Kotlin
TypeScript
Java
JavaScript
Java
Gradle
Frontend
React.js
DevOps
CI/CD
Apply
Remote • Full-Time • Buenos Aires
Java
Node JS
Python
C#
JavaScript
C#
.NET
AI/ML
Feature Store
NumPy
Recommender Systems
Scikit-learn
SciPy
Frontend
GraphQL
DevOps
CI/CD
Apply
$9.6k – $16k per year • Remote • Contractor • PhD • Cape Town • Guayaquil • Santa Cruz • Quezon City • Manila
PHP
PHP
WordPress
Design
Canva
Marketing
HubSpot
Instagram
LinkedIn
YouTube
Apply
$9.6k – $16k per year • Remote • Contractor • 2+ years exp • Monterrey • Guayaquil • Santa Cruz • Cairo • Davao City
PHP
PHP
WordPress
Design
Canva
Marketing
HubSpot
Spotify
YouTube
Apply
$12k – $30k per year • Remote • Contractor • Santo Domingo • Guayaquil • Santa Cruz • Monterrey • Cairo
Design
Canva
Management
Asana
ClickUp
Google Workspace
WhatsApp
Marketing
Instagram
Apply
$12k – $30k per year • Remote • Contractor • Cape Town • Guayaquil • Santa Cruz • Monterrey • Cairo
Analytics
A/B Testing
Management
Google Workspace
Slack
Marketing
Instagram
LinkedIn
Apply
SEO Lead 2 days ago
$24k – $36k per year • Remote • Contractor • 3+ years exp • Cape Town • Guayaquil • Santa Cruz • Monterrey • Cairo
PHP
PHP
WordPress
Management
Asana
Google Sheets
Google Workspace
Slack
Marketing
Ahrefs
GA4
Apply
See all jobs
This is one of many
397,589 more open roles from verified company boards, updated every day.