841,941open jobs
53,762companies
142,138added this week
Browse all
Salary
$120k – $180k per year
Location
Remote (United States)
Employment
Part-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Jul 29, 2026. Weekday scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Weekday is an Indian recruitment company that sources software engineers through referrals from other engineers rather than through job advertisements or agency databases. Its model pays working engineers to vouch for former colleagues they rate, turning informal knowledge about who is genuinely good into a searchable candidate pool that companies can hire from. Based in Bengaluru and backed by Y Combinator, the platform has layered AI screening and outbound sourcing on top of that referral network, and sells to startups and technology companies hiring in the Indian market.

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking researchers from computational STEM disciplines-as well as computationally intensive social sciences and humanities-to bring the rigor of real-world research into AI evaluation.

In this role, you will transform scientific methodologies such as experimental design, hypothesis testing, and data-driven analysis into sophisticated, multi-step benchmark tasks that challenge state-of-the-art AI systems. Working closely with AI researchers, you'll help uncover subtle reasoning errors and methodological flaws that only experienced researchers can identify.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities

  • Design complex, research-oriented benchmark tasks inspired by real-world scientific workflows, including study design, experimentation, hypothesis testing, and data analysis.
  • Develop comprehensive reference solutions using Python, notebooks, and computational tools with the rigor expected in professional research.
  • Define clear evaluation standards that distinguish sound scientific reasoning from plausible but incorrect conclusions.
  • Review AI-generated solutions, identifying methodological weaknesses, analytical errors, and flawed reasoning that experienced researchers would recognize immediately.
  • Collaborate with AI researchers and fellow domain experts to improve benchmark quality, consistency, and scientific rigor.
  • Contribute to the continuous refinement of evaluation methodologies for advanced AI systems.

Required Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline, computational social science, computational humanities, or another research-intensive field involving programming and data analysis.
  • Minimum 1 year of experience in an active research role within academia, industry, government laboratories, or a similar research environment.
  • Demonstrated experience performing computational research involving Python, data analysis, simulation, modeling, machine learning, or scientific computing.
  • Strong understanding of experimental design, hypothesis testing, statistical analysis, and rigorous interpretation of research findings.
  • Working knowledge of Git, integrated development environments (IDEs), and notebook platforms such as Jupyter or Google Colab.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is preferred.
  • Excellent analytical thinking, attention to detail, creativity, and the ability to solve complex, open-ended problems independently.
  • Strong written communication skills for documenting technical methodologies and research findings.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Preferred Qualifications

  • Experience designing reproducible computational experiments or research workflows.
  • Familiarity with machine learning, large language models, or AI-assisted research tools.
  • Background in benchmark design, scientific software development, or computational research infrastructure.
  • Experience mentoring researchers, reviewing scientific work, or contributing to peer-reviewed publications.

Why Join

  • Help shape how next-generation AI systems are evaluated using rigorous scientific methodologies.
  • Collaborate with leading AI researchers working on frontier models and advanced evaluation frameworks.
  • Apply your research expertise to improve AI reasoning, reliability, and scientific accuracy.
  • Contribute to impactful work that advances the quality and robustness of AI systems across multiple disciplines.
  • Enjoy the flexibility of a fully remote engagement while working on cutting-edge AI research initiatives.

Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details

  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
841,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
$80k – $120k per year • In office • Internship • PhD • Boston
Python
C++
AI/ML
Reinforcement Learning
Embodied AI
Robotics
ROS
Teleoperation
Imitation Learning
Reinforcement Learning
Apply
$80k – $120k per year • In office • Internship • PhD • Boston
Python
C++
AI/ML
Reinforcement Learning
Embodied AI
Robotics
ROS
Teleoperation
Imitation Learning
Reinforcement Learning
Apply
≈ $157k – $339k per year (Estimated) • Remote (United States) • Contractor
AI/ML
Post-training
Apply
PostDoc Researcher 24 days ago
$55k – $65k per year • In office • Full-Time • PhD • New York
AI/ML
Reinforcement Learning
Multimodal AI
AI Agents
Machine Learning
Robotics
Imitation Learning
Reinforcement Learning
Management
Gmail
Apply
≈ $137k – $291k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Chicago
Databases
Snowflake
Databricks
AI/ML
AI Agents
Machine Learning
DevOps
AWS
Cybersecurity
SOC 2
HIPAA
Management
Agile
Apply
AI Engineer 1 day ago
≈ $147k – $317k per year (Estimated) • In office • Full-Time • Chicago
Python
AI/ML
AI Agents
LLM
OCR
Human-in-the-Loop
Machine Learning
Apply
$75k – $100k per year • Hybrid • Public Trust • 2+ years exp • Bachelor's Degree • Reston
Python
JavaScript
TypeScript
C#
C#
.NET
Databases
PostgreSQL
AI/ML
AI Agents
Agentic Workflows
Frontend
React.js
DevOps
Terraform
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
GitHub
Apply
In office • 6+ years exp • Hyderabad
Python
Go
JavaScript
TypeScript
SQL
Go
Temporal
AI/ML
Model Context Protocol
Embeddings
Multimodal AI
Function Calling
AI Agents
LangSmith
LLM
RAG
OpenAI
OpenAI Agents SDK
OCR
Text-to-Speech
Agentic Workflows
Tool Use
Frontend
GraphQL
React.js
DevOps
Rest API
GCP
CI/CD
AWS
Docker
Kubernetes
Google GKE
Amazon EC2
Amazon ECS
Apply
≈ $65k – $136k per year (Estimated) • In office • TS/SCI • 1+ year exp • Bachelor's Degree • Saint Louis
Python
JavaScript
Rust
DevOps
Terraform
Ansible
Azure DevOps
GitHub Actions
AWS CDK
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Amazon S3
Linux
Management
Agile
Scrum
Apply
≈ $63k – $129k per year (Estimated) • In office • TS/SCI • 5+ years exp • Bachelor's Degree • Saint Louis
Python
Java
DevOps
Red Hat
Configuration Management
CentOS Stream
Linux
Windows
Cybersecurity
NIST 800-53
Management
Confluence
Jira
Agile
Apply
AI Engineer 1 day ago
≈ $147k – $317k per year (Estimated) • In office • Full-Time • Chicago
Python
AI/ML
AI Agents
LLM
OCR
Human-in-the-Loop
Machine Learning
Apply
ML Engineer 2 days ago
Remote (India) • Full-Time
Python
Python
FastAPI
AI/ML
Fine-tuning
AI Agents
TensorFlow
PyTorch
LLM
Machine Learning
DevOps
GCP
Kubernetes
Apply
$42k – $52k per year • Remote (India) • Full-Time
Python
Python
FastAPI
AI/ML
AI Agents
PyTorch
LLM
Multi-Agent Systems
Machine Learning
DevOps
GCP
Kubernetes
Apply
$180k – $240k per year • Remote (United States, United Kingdom, Canada) • Part-Time
AI/ML
Ray Serve
DeepSpeed
vLLM
CUDA Toolkit
JAX
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
CUDA
Triton
Megatron-LM
FSDP
TPU
XLA
KV Cache
Apply
Senior GenAI Engineer 17 days ago
≈ $45k – $113k per year (Estimated) • Remote (India) • Full-Time
Python
SQL
AI/ML
LangGraph
LangChain
LoRA
Fine-tuning
RLHF
Quantization
Scikit-learn
Prompt Engineering
Knowledge Distillation
AI Agents
NLP
PEFT
Llama
Transformers
TensorFlow
PyTorch
LLM
RAG
OpenAI
LLMOps
TPU
Multi-Agent Systems
Model Distillation
Machine Learning
DevOps
Rest API
GCP
Azure
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
GDPR
Apply
See all jobs
This is one of many
841,941 more open roles from verified company boards, updated every day.