402,911open jobs
14,044companies
78,108added this week
Browse all
Salary
$150k – $220k per year
Location
Remote (United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Weekday is an Indian recruitment company that sources software engineers through referrals from other engineers rather than through job advertisements or agency databases. Its model pays working engineers to vouch for former colleagues they rate, turning informal knowledge about who is genuinely good into a searchable candidate pool that companies can hire from. Based in Bengaluru and backed by Y Combinator, the platform has layered AI screening and outbound sourcing on top of that referral network, and sells to startups and technology companies hiring in the Indian market.

This role is for one of our clients

Compensation: $75 - $110 per hour

We are seeking an experienced Cloud / DevOps Engineer (Infra & IaC) to contribute to a cutting-edge GenAI environment focused on building and improving large-scale AI training and inference infrastructure.

The ideal candidate will bring strong, hands-on expertise in Kubernetes, AWS cloud services, Infrastructure-as-Code (IaC), and CI/CD. You will apply your real-world infrastructure engineering experience to evaluate technical workflows, create high-quality reference solutions, identify gaps in AI-generated outputs, and help establish rigorous standards for cloud and DevOps reasoning.

This is a full-time engagement requiring 40 hours per week, Monday through Friday.

Requirements

Key Responsibilities

  • Collaborate with research and engineering teams to identify knowledge gaps and improve AI model performance across cloud infrastructure, DevOps, Kubernetes, and Infrastructure-as-Code domains.
  • Design realistic and technically challenging tasks covering Kubernetes troubleshooting, AWS service integration, infrastructure automation, and production operations.
  • Develop accurate, detailed reference solutions for complex infrastructure engineering scenarios.
  • Review and evaluate AI-generated technical solutions for correctness, reliability, scalability, security, and adherence to production best practices.
  • Provide clear, structured written feedback highlighting technical gaps, incorrect assumptions, and opportunities for improvement.
  • Create detailed evaluation criteria, rubrics, and benchmarks for assessing Kubernetes troubleshooting, IaC architecture, AWS integrations, and CI/CD reasoning.
  • Develop scenarios involving cluster failures, infrastructure automation, deployment workflows, service integrations, and operational reliability.
  • Work closely with other technical subject matter experts to maintain consistency, accuracy, and quality across evaluation datasets.
  • Translate practical production experience into structured guidance that can be used to improve AI-generated infrastructure solutions.

Core Qualifications

  • 4+ years of professional experience in Cloud Infrastructure, DevOps, Site Reliability Engineering, Platform Engineering, or a closely related field.
  • Strong hands-on experience managing Kubernetes in production environments, including diagnosing, troubleshooting, and resolving cluster failures and operational issues.
  • Experience with Kubernetes beyond simply writing manifests or consuming managed Kubernetes control planes.
  • Proven production experience with Infrastructure-as-Code, particularly Terraform and/or AWS CDK.
  • Strong practical knowledge of AWS cloud services, including production integration with services such as:
    • AWS Lambda
    • API Gateway
    • DynamoDB
  • Experience designing, implementing, and maintaining CI/CD pipelines for production workloads.
  • Strong understanding of cloud architecture, infrastructure automation, deployment strategies, observability, reliability, and operational best practices.
  • Demonstrated career progression with increasing ownership and responsibility in infrastructure, DevOps, or platform engineering.
  • Ability to commit reliably to 40 hours per week during standard weekdays.
  • Excellent written and verbal communication skills, with the ability to explain complex technical concepts and engineering decisions clearly.
  • Strong analytical and troubleshooting abilities, particularly when diagnosing distributed systems and infrastructure failures.

Preferred Skills

  • Experience working with large-scale cloud infrastructure or highly distributed systems.
  • Familiarity with Kubernetes networking, security, storage, scaling, and cluster lifecycle management.
  • Experience implementing infrastructure security and reliability best practices.
  • Knowledge of AWS architecture patterns and cloud-native application design.
  • Experience with GitOps, containerization, monitoring, logging, and observability platforms.
  • Familiarity with modern DevOps and platform engineering methodologies.
  • Experience reviewing or evaluating technical documentation, engineering solutions, or AI-generated outputs.

What You’ll Contribute

In this role, your production infrastructure expertise will help establish high-quality standards for AI systems working with complex Cloud, DevOps, Kubernetes, AWS, and IaC problems.

You will play a key role in transforming practical engineering knowledge into structured tasks, reference solutions, evaluation frameworks, and high-quality technical feedback that can improve the capabilities of next-generation AI models.

Equal Opportunity

We are committed to providing equal employment opportunities to all qualified candidates. Employment decisions are made without regard to legally protected characteristics, and reasonable accommodations are available throughout the hiring and engagement process upon request.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
402,911 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$30k – $76k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
JavaScript
Node JS
Python
TypeScript
Python
FastAPI
Flask
Databases
Milvus
pgvector
Pinecone
PostgreSQL
AI/ML
AI Agents
Anthropic
Arize Phoenix
Fine-tuning
LangChain
LangGraph
LangSmith
LlamaIndex
OpenAI
Prompt Engineering
RAG
Multi-Agent Systems
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Prometheus
Vector
Apply
$25k – $59k per year (Estimated) • In office • Full-Time • India
Bash
Python
Databases
Oracle
DevOps
Ansible
AWS
Azure
CI/CD
Datadog
GCP
Grafana
Incident Management
Prometheus
Splunk
Apply
In office • Full-Time • Bachelor's Degree • India
Java
SQL
Java
Maven
AI/ML
Prompt Engineering
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
QA
JMeter
Postman
Selenium
SoapUI
Apply
Remote/Hybrid • Full-Time • 1+ year exp • Beirut
Groovy
Java
Python
Databases
Oracle
PostgreSQL
DevOps
AWS
Azure
Azure DevOps
CI/CD
GCP
Git
GitLab CI
Incident Management
Jenkins
Apply
In office • Full-Time • Denmark
AI/ML
AI Agents
Copilot
Cursor
LLM
Agentic Workflows
DevOps
CI/CD
Git
GitHub
Apply
In office • Full-Time • Bengaluru
SystemVerilog
Chips/EDA
UVM
Apply
$160k – $320k per year • Remote • Contractor
Management
Stripe
Apply
$20k – $84k per year (Estimated) • In office • Full-Time • Hyderabad
JavaScript
SQL
TypeScript
Apex
Java
Apex
MuleSoft
Java
Spring Boot
Databases
Apache Kafka
Kafka
MS SQL
Frontend
React.js
DevOps
API Gateway
Azure
CI/CD
Kubernetes
Rest API
Apply
$20k – $83k per year (Estimated) • In office • Full-Time • Pune
Java
SQL
Java
Hibernate
Spring Boot
DevOps
Docker
Kubernetes
Rest API
Apply
$33k – $71k per year (Estimated) • In office • Full-Time • Mumbai
Analytics
Power BI
Apply
See all jobs
This is one of many
402,911 more open roles from verified company boards, updated every day.