{"id":1136349,"url":"https://alion.io/job/mindrift-freelance-agent-evaluation-engineer-105","title":"Freelance Agent Evaluation Engineer","company":{"id":3475,"name":"Mindrift","domain":"mindrift.ai","url":"https://alion.io/company/mindrift","size_band":"201-500","is_staffing_agency":true,"is_intermediary":true,"ats_vendor":"Workable","truth_index":{"grade":"B","score":80,"open_postings":6,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":7,"computed_at":"2026-09-23T05:45:00Z"}},"role":"Backend","role_family":"Backend","seniority":null,"employment_type":"freelance","work_mode":"remote","remote_scope":"stated_countries","hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":["PT"],"hiring_countries_total":1,"salary":{"min":null,"max":40,"currency":"USD","period":"hour","gross":null,"usd_annual":80000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Apache Kafka","optional":false},{"name":"Docker","optional":false},{"name":"FastAPI","optional":false},{"name":"JavaScript","optional":false},{"name":"Machine Learning","optional":false},{"name":"NLP","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"Python","optional":false},{"name":"React.js","optional":false},{"name":"Redis","optional":false},{"name":"TypeScript","optional":false}],"status":"live","first_seen_at":"2026-09-23T08:18:11Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-23T16:37:51Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Please submit your CV in English and indicate your level of English proficiency.\nMindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.\nWe're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.\nYou'll create challenging tasks and evaluation criteria within realistic simulated environments:\nBuild realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history\nDesign tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent\nWrite tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient\nIterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust\nWhat this is NOT:\nNot data labeling\nNot prompt engineering\nNot writing code from scratch - the agent writes most of the code; you guide and evaluate\nWhat we look for:\n5+ years in software development\nCore stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis\nExperience writing tests (functional, integration)\nEnglish proficiency - B2+\nWhy this is hard:\nFrontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.\nRequirements and benefits\n Educational qualifications\nA Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields. \nBachelor’s degree is accepted if only candidate has 5 years of experience in the field.\nAcademic and/or Professional Experience\nCandidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles\nHow it works\nApply → Pass qualification(s) → Join a project → Complete tasks → Get paid\nCompensation:\nPaid per accepted task. Your rate depends on the qualification tier you reach and how efficiently you complete tasks - up to the equivalent of $40/hr. Because payment is per task, a faster pace raises your effective hourly rate.\nWhy this freelance opportunity might be a great fit for you?\nTake part in a part-time, remote, freelance project that fits around your primary professional or academic commitments. \nWork on advanced AI projects and gain valuable experience that enhances your portfolio. - Influence how future AI models understand and communicate in your field of expertise.","description_format":"text","description_chars":3144,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"Advanced (C1)","optional":false}]},"benefits":[],"hiring_locations":[{"name":"Portugal","iso":"PT","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity"],"lifecycle":[{"event":"open","at":"2026-09-23T08:18:11Z"}],"liveness":{"score":52,"band":"ok","label":"Likely open","p_open":1,"p_active":0.516,"p_room":1,"age_days":0,"expected_fill_days":7,"reasons":["conf:0","agency","win:early","comp:brand"],"computed_at":"2026-09-23T17:30:33Z"},"pay":{"stated_usd_annual":80000,"is_top_pay":false},"html_url":"https://alion.io/job/mindrift-freelance-agent-evaluation-engineer-105","json_url":"https://alion.io/job/mindrift-freelance-agent-evaluation-engineer-105.json","meta":{"generated_at":"2026-09-23T17:30:33Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}