368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$120k – $180k per year
Location
Remote (United States)
Employment
Part-Time
Overview
Company
Impact
Profile match
At Weekday, we help companies hire engineers who are vouched by other software engineers. We are enabling engineers to earn passive income by leveraging & monetizing the unused information in their head about the best people they have worked with.

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss.

In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities

  • Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks.
  • Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing.
  • Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible.
  • Document findings with clear technical explanations, supporting evidence, and reproducible methodologies.
  • Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria.
  • Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage.

Required Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis.
  • Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field.
  • Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems.
  • Strong proficiency in Python and Git, with the ability to build custom scripts for experimentation, testing, and analysis.
  • Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies.
  • Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable.
  • Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently.
  • Outstanding written communication skills for documenting technical findings clearly and accurately.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Preferred Qualifications

  • Experience with AI safety, red teaming, adversarial machine learning, or security research.
  • Background in benchmark design, evaluation framework development, or AI quality assurance.
  • Experience creating reproducible technical experiments and documenting complex failure analyses.
  • Familiarity with frontier AI research methodologies and model capability assessments.

Why Join

  • Help shape the future of AI evaluation by identifying critical weaknesses before they reach production.
  • Work on cutting-edge AI systems alongside researchers developing next-generation language models.
  • Apply your technical expertise to improve AI reliability, reasoning, and robustness.
  • Contribute directly to benchmark development that influences the evolution of advanced AI technologies.
  • Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives.

Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details

  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Software Developer 4 days ago
$70k – $126k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Egg Harbor Township
Bash
Python
DevOps
Ansible
CI/CD
Configuration Management
Docker
Git
Harbor
KVM
Podman
Red Hat
VMWare
Apply
$87k – $157k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Egg Harbor Township
Bash
Python
DevOps
Ansible
CI/CD
Configuration Management
Docker
Git
Harbor
KVM
Podman
Red Hat
VMWare
Apply
$118k – $197k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Dallas • Chicago
Python
SQL
JavaScript
Databases
Databricks
Delta Lake
AI/ML
ChatGPT
Claude
Copilot
Cursor
Spark
Frontend
Next.js
React.js
DevOps
Azure
Azure DevOps
CI/CD
Vercel
GitHub
Analytics
Power BI
Tableau
ETL/ELT
Apply
$108k – $195k per year • In office • Full-Time • 8+ years exp • High School Diploma • Tucson
Bash
Java
PowerShell
Python
DevOps
Docker
Apply
$70k – $126k per year • In office • Full-Time • 4+ years exp • Master's Degree • Saint Louis
Node JS
JavaScript
Databases
PostgreSQL
DevOps
Amazon EC2
AWS
Git
Amazon S3
GitLab
Management
Confluence
Jira
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.