428,943open jobs
14,661companies
61,923added this week
Browse all
Salary
$35k – $97k per year (Estimated)
Location
Remote (Brazil)
Seniority
Junior · 1+ year exp
Employment
Contractor
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Benchmark Engineer | Native Language Specialist based in Brazil.

This is a remote freelance opportunity for a native-speaking software or prompt engineer passionate about multilingual AI and rigorous model evaluation. You will design realistic Terminal-Bench tasks that test how effectively large language models handle software challenges in your native language. Your work will help uncover failure modes related to multilingual prompts, non-English datasets, Unicode, encoding, localization, and terminal workflows. You will build and validate high-signal benchmark environments while ensuring that language-specific assets remain authentic rather than relying on English translations. Working across task engineering, implementation, calibration, and quality assurance, you will directly contribute to improving the reliability and multilingual robustness of AI systems. This role offers flexible, project-based work within a global community of technical and language experts.

Accountabilities:

    • Design and engineer challenging, realistic benchmark tasks that evaluate coding agents and large language models in multilingual terminal environments.
    • Create authentic task environments using datasets, files, prompts, and other assets written in your native language, ensuring they accurately reflect real-world language usage.
    • Identify model failure points related to native-language prompting, translation gaps, multilingual reasoning, and language-specific software workflows.
    • Develop robust reference implementations and highly reliable, deterministic verifier scripts, using rubric-based evaluation only when strictly necessary.
    • Analyze execution logs and calibrate task difficulty from Easy to Very Hard using standardized Terminal-Bench configurations across different model tiers.
    • Participate in a rigorous multi-layer quality process covering task creation, human review, calibration review, and final audit, alongside automated LLM-based checks.
    • Validate grammatical accuracy, linguistic authenticity, technical correctness, fairness, reproducibility, and overall benchmark integrity.
    • Apply deep knowledge of multilingual text processing to identify edge cases involving Unicode normalization, encoding and decoding, locale behavior, text I/O, string operations, and toolchain interoperability.
    • Where relevant to the target language, account for bidirectional and RTL text handling, font fallbacks, rendering, and typography in software interfaces and generated artifacts.
    • Requirements

      • At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
      • Demonstrated technical experience through work at established technology organizations and/or graduation from a strong engineering university.
      • Native or near-native fluency in the target language, with a sophisticated understanding of grammar, register, phrasing, and language-specific conventions.
      • Strong English proficiency for technical communication and collaboration.
      • Strong Python skills, along with proficiency in standard shell scripting and data processing workflows.
      • Extensive experience working with Terminal/CLI-based development environments and familiarity with coding agents or AI-assisted development tools.
      • Solid understanding of multilingual text-processing challenges, including encoding and decoding, Unicode normalization, locale-dependent casing and collation, non-Gregorian dates, text I/O, and safe string manipulation.
      • For applicable languages, familiarity with bidirectional or RTL text, font fallback behavior, and rendering or typography considerations.
      • Strong attention to detail and the ability to create deterministic, reproducible, technically rigorous evaluation tasks.
      • Ability to work independently, meet deadlines, and maintain consistent quality across project-based assignments.
      • Reliable availability and commitment once tasks are accepted, with most assignments requiring at least 2 hours per day or 10 hours per week.
      • Ability to provide an updated CV in English and successfully complete a GenAI assessment as part of the onboarding process.
      • Must be eligible to work as an independent contractor in the applicable location; geographic restrictions may apply in regions subject to international embargoes or sanctions.
      • As an independent contractor, you are responsible for your own applicable tax obligations.
      • Benefits

        • Fully remote, flexible freelance work that allows you to choose when and how much you work, without fixed hours or day-to-day micromanagement.
        • Competitive project-based compensation with prompt payments and a streamlined invoicing process.
        • Opportunity to contribute to cutting-edge AI evaluation and multilingual language technology with real-world impact.
        • Access to diverse technical and language-focused projects that can broaden your portfolio and strengthen your expertise across domains.
        • Collaboration with a global community of software engineers, linguists, language specialists, and AI experts.
        • Flexible supplemental work that can fit alongside other professional commitments, with project availability varying according to demand.
        • Exposure to emerging AI technologies, coding agents, benchmark development, and multilingual model evaluation.
        • No traditional employee benefits such as health insurance, paid time off, or retirement contributions, as this is an independent contractor opportunity.
        • No guaranteed hours or fixed workload; availability depends on project demand and successful task assignments.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,943 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$40k – $108k per year (Estimated) • Remote • Full-Time • 5+ years exp
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
dbt
LLM
Anomaly Detection
DevOps
GitHub Actions
CI/CD
Git
GitHub
Analytics
ETL/ELT
Fivetran
Looker
Management
Google Sheets
Apply
$142k – $235k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
AI/ML
Copilot
Cursor
AI Agents
LLM
DevOps
Rest API
Azure DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
GitHub
GitLab
Management
Linear
Jira
Apply
Security GRC Lead 1 day ago
$350k – $425k per year • Equity • In office • Full-Time • 7+ years exp • San Francisco
Python
SQL
AI/ML
Model Context Protocol
LLM
OpenAI
Anthropic
EU AI Act
NIST AI RMF
ISO 42001
DevOps
SLI/SLO/SLA
IAM
Cybersecurity
ISO 27001
Wiz
PCI DSS
SOC 2
HIPAA
FedRAMP
Management
SharePoint
Apply
$16k – $40k per year (Estimated) • Remote • Full-Time
AI/ML
AI Agents
LLM
Apply
$35k – $87k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
Spark
Scikit-learn
AI Agents
NLP
TensorFlow
Pandas
NumPy
PyTorch
BERT
Sentiment Analysis
Hugging Face
Amazon SageMaker
DevOps
GCP
Azure
AWS
AWS Lambda
Amazon S3
Analytics
Tableau
Power BI
Matplotlib
ETL/ELT
Management
Microsoft Teams
Apply
$19k – $47k per year (Estimated) • Remote • Contractor • Bachelor's Degree
SQL
DevOps
Kali Linux
Cybersecurity
Burp Suite
Nmap
SQLmap
Apply
Remote • Contractor
ABAP
Apply
$16k – $40k per year (Estimated) • Remote • Full-Time
AI/ML
AI Agents
LLM
Apply
$35k – $87k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
Spark
Scikit-learn
AI Agents
NLP
TensorFlow
Pandas
NumPy
PyTorch
BERT
Sentiment Analysis
Hugging Face
Amazon SageMaker
DevOps
GCP
Azure
AWS
AWS Lambda
Amazon S3
Analytics
Tableau
Power BI
Matplotlib
ETL/ELT
Management
Microsoft Teams
Apply
Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
JavaScript
Node JS
Databases
Snowflake
ElasticSearch
DevOps
Rest API
GCP
Kibana
Azure
AWS
Docker
Kubernetes
Amazon EKS
AWS Fargate
AWS Lambda
Amazon ECS
Analytics
Tableau
ETL/ELT
Apply
See all jobs
This is one of many
428,943 more open roles from verified company boards, updated every day.