716,098open jobs
42,633companies
101,517added this week
Browse all
Location
Remote (Argentina, Brazil, Colombia, Chile, Peru, Uruguay)
Employment
Contractor

Confirmed on the employer's own hiring board on Sep 23, 2026. First seen by Alion on Sep 23, 2026. Gramian Consulting Group scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Gramian Consulting Group is an international management and technology consulting firm based in Los Angeles, California. The firm provides executive-level advisory services, digital transformation strategy, software engineering, and corporate restructuring for middle-market and enterprise clients across diverse industries. By combining high-level strategic planning with hands-on technology execution, the company helps client organizations modernize their operations, optimize revenue models, and scale business growth efficiently.

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About the Role

We are seeking an AI Quality & Evaluation Specialist to validate task quality, review AI agent performance, and audit grading logic. You will analyze execution traces, tool calls, reference solutions, and evaluation criteria to identify task defects, grading errors, and unjustified model failures. The ideal candidate combines technical fluency with strong analytical judgment and the ability to provide clear, evidence-based feedback.

CONTRACT: Contractor (Hourly)

COMMITMENT: 40 hours per week

LOCATION: Remote - Latin America (LATAM)

Responsibilities

  • Validate task instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and generated deliverables to assess whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, and unsupported rubric criteria.
  • Identify cases where valid alternative solutions are unfairly penalized.
  • Investigate discrepancies between model performance, task defects, grader errors, and environment or tool failures.
  • Independently assess automated QC findings rather than accepting them without verification.
  • Document concise, evidence-backed findings and actionable recommendations.
  • Clearly communicate uncertainty and distinguish confirmed issues from potential problems.
  • Verify that task revisions resolve previously identified defects.
  • Collaborate with relevant teams to improve evaluation quality and reliability.

Requirements

  • Experience in technical QA, AI evaluation, data analysis, software testing, or a related technical field.
  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Strong analytical skills, including the ability to validate calculations and reconcile conflicting evidence.
  • Experience assessing the correctness and completeness of technical or professional deliverables.
  • Strong written English and the ability to provide specific, reproducible feedback.
  • High attention to detail when reviewing task requirements, grading behavior, and model outputs.
  • Ability to independently investigate issues and make evidence-based decisions.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
716,098 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
$48k – $127k per year (Estimated) • Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
$84k – $191k per year (Estimated) • Remote/Hybrid • Full-Time • Boston
Python
PowerShell
AI/ML
Prompt Engineering
AI Agents
Machine Learning
DevOps
Azure
AWS
Linux
Windows
Cybersecurity
ISO 27001
Active Directory
Apply
$57k – $118k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru
Python
JavaScript
Ruby
Ruby
Ruby on Rails
Databases
ElasticSearch
AI/ML
AI Agents
Frontend
React.js
DevOps
Terraform
GCP
Docker
Kubernetes
Apply
Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
$48k – $127k per year (Estimated) • Remote • Contractor • PhD
Python
AI/ML
AI Agents
Apply
$81k – $194k per year (Estimated) • Remote • Contractor • Bachelor's Degree
Python
Go
JavaScript
Java
Rust
TypeScript
C++
Node JS
Apply
Remote • Contractor • Bachelor's Degree
Python
Go
JavaScript
Java
Rust
TypeScript
C++
Node JS
Apply
See all jobs
This is one of many
716,098 more open roles from verified company boards, updated every day.