707,525open jobs
41,962companies
101,146added this week
Browse all
Salary
$119k – $271k per year (Estimated)
Location
Remote/Hybrid (Sunnyvale, San Francisco, United States, Bengaluru, India)
Employment
Full-Time
Overview
Company
Impact
Profile match

Collinear AI builds environments, tasks, and evaluations that help frontier AI models improve at real work. We are growing our cybersecurity team and looking for someone who can turn practical security problems into environments where AI agents can investigate, act, and learn.

We recently released CWE-bench, a defensive cybersecurity benchmark with 100 held-out audit-and-patch tasks across 54 weakness types. Agents must find and fix vulnerabilities in real codebases, with checks that confirm the vulnerability is resolved and existing functionality still works. You will help build what comes next: richer cybersecurity environments, realistic tasks, and reliable ways to measure whether agents succeed.

You will own work from the initial security scenario through the runnable environment, task instructions, reference solution, and verifier. We are looking for hands-on security knowledge, strong programming skills, and curiosity about how AI agents fail.

Deep cybersecurity expertise is enough to get started-you do not need prior AI or machine-learning experience. If you know how to investigate vulnerabilities, reason about security failures, and verify that a fix works, we want to hear from you. We will teach you our AI tooling and evaluation workflows.

What you will do

  • Build cybersecurity environments. Create isolated, reproducible environments with real repositories, applications, services, logs, and access controls. Make them straightforward to launch, reset, and evaluate at scale.

  • Turn security work into tasks. Design scenarios around vulnerability discovery and remediation, application and API security, authentication and authorization, incident investigation, and system hardening. Define what the agent knows, which tools it can use, and what it must accomplish.

  • Develop reference solutions and verifiers. Reproduce the underlying issue, implement a valid solution, and write checks that distinguish a real fix from a superficial workaround. Verify that attacks fail after remediation while legitimate behavior continues to work.

  • Test the tests. Challenge graders with incomplete fixes, disabled features, hard-coded answers, and other shortcuts. Keep hidden solutions and test data out of the agent's environment, and separate genuine model failures from broken infrastructure.

  • Run and analyze agents. Evaluate frontier models, inspect their tool calls and code changes, and explain where their security reasoning or execution breaks down. Use those findings to improve task coverage and difficulty.

  • Contribute to benchmarks and research. Work with researchers and engineers to turn strong environments into evaluation suites and training data. Review other contributors' tasks and help document results for future releases.

Who we are looking for

  • Practical depth in at least one area of cybersecurity, such as application security, vulnerability research, penetration testing, systems security, cloud security, or incident response. Evidence can come from internships, research, open-source work, bug bounties, CTFs, or independent projects.

  • The ability to read unfamiliar code, reproduce a security issue, understand its root cause, and implement or assess a fix.

  • Strong programming skills in Python and at least one language used in the systems you investigate, such as C/C++, Go, Java, JavaScript/TypeScript, or Rust.

  • Comfort with Linux, Git, containers, debugging, and automated testing.

  • Clear technical writing and careful judgment about what an evaluation does and does not demonstrate. You can explain why a task is realistic and why its grading is trustworthy.

  • Curiosity about applying your cybersecurity expertise to AI, and a willingness to learn how to evaluate tool-using agents. Prior experience with LLMs is not expected.

Recent graduates and early-career engineers or researchers are encouraged to apply. A PhD, professional certification, or previous role at an AI lab is not required. We value demonstrated ability and the quality of your work.

Nice to have

  • Disclosed vulnerabilities, accepted security patches, strong CTF results, or useful security tools and write-ups.

  • Experience with fuzzing, static or dynamic analysis, reverse engineering, or building security labs.

  • Familiarity with CWE and OWASP classifications and how they relate to concrete software failures.

  • Experience building agent evaluations, adversarial tests, prompt-injection defenses, or reinforcement-learning environments.

*Examples of what you might build

  • A repository audit where an agent must discover and repair an authorization flaw without being told where it is.

  • A small service environment where an agent investigates suspicious activity from logs and configuration, then applies and verifies a remediation.

  • An LLM application where an agent must repair a prompt-injection or tool-permission weakness while preserving legitimate functionality.

*When applying, include a project, code sample, security write-up, or research artifact that shows how you investigate a problem and establish that your solution works.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
707,525 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Sunnyvale
In office • Full-Time • Bachelor's Degree • Marousi
Python
Bash
AI/ML
AI Agents
DevOps
Terraform
Ansible
Datadog
PagerDuty
Prometheus
Azure
CI/CD
Docker
Kubernetes
Grafana
Incident Management
Linux
Apply
Research Engineer 5 hours ago
$187k – $403k per year (Estimated) • Remote/Hybrid • New York
Python
AI/ML
Reinforcement Learning
AI Agents
DevOps
AWS
Apply
$260k – $330k per year • Remote/Hybrid • Full-Time • 5+ years exp • New York
Python
Go
JavaScript
AI/ML
AI Agents
Edge AI
Frontend
React.js
Apply
$79k – $188k per year (Estimated) • In office • Full-Time • 8+ years exp • London
AI/ML
Copilot
AI Agents
LLM Guardrails
Copilot Studio
Analytics
Power BI
Management
Confluence
Jira
Power Automate
Power Apps
SharePoint
Agile
Apply
In office • Full-Time • 6+ years exp • Master's Degree • Bengaluru
Python
SQL
AI/ML
Copilot
Claude
Prompt Engineering
AI Agents
Llama
LLM
OpenAI
Anthropic
DevOps
AWS
Analytics
Tableau
Power BI
Design
Sketch
Management
Agile
Waterfall
Microsoft Office
Marketing
Salesforce
Apply
$125k – $334k per year (Estimated) • Equity • In office • Internship • Bachelor's Degree • Sunnyvale • San Francisco • Bengaluru
Python
AI/ML
Model Context Protocol
Fine-tuning
RLHF
Reinforcement Learning
AI Agents
SFT
Post-training
Computer Use
RLAIF
Apply
Growth Marketer 8 days ago
$85k – $213k per year (Estimated) • In office • Full-Time • PhD • Sunnyvale • San Francisco
Apply
$106k – $275k per year (Estimated) • In office • Full-Time • Master's Degree • San Francisco • Sunnyvale
Python
AI/ML
RAG
Agentic Workflows
Machine Learning
DevOps
CI/CD
HPC
Linux
Robotics
Digital Twin
Apply
MTS - Product (India) 1 month ago
Remote • Full-Time • Bachelor's Degree • Bengaluru
Python
JavaScript
SQL
Python
FastAPI
Celery
Databases
Redis
AI/ML
Reinforcement Learning
NLP
LLM
Time Series Forecasting
Frontend
Next.js
React.js
DevOps
CI/CD
Apply
MTS - Research (India) 2 months ago
Remote • Full-Time • Bachelor's Degree • Bengaluru
Python
AI/ML
Fine-tuning
RLHF
Reinforcement Learning
AI Agents
SFT
Post-training
World Models
RLAIF
Apply
$78k – $114k per year • Remote • Full-Time • Sunnyvale
Python
AI/ML
Claude
ChatGPT
AI Agents
Agentic Workflows
Cybersecurity
LDAP
Management
Confluence
Jira
Apply
$114k – $245k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Sunnyvale
AI/ML
AI Agents
DevOps
Azure
DNS
Cybersecurity
LDAP
DLP
Apply
$83k – $125k per year • In office • Secret • Full-Time • 5+ years exp • Bachelor's Degree • Sunnyvale
Apply
$88k – $171k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Sunnyvale
Java
Java
Maven
Spring Boot
Databases
Cassandra
Apache Kafka
Azure Cosmos DB
DevOps
CI/CD
Jenkins
Git
Docker
Kubernetes
Management
Agile
Apply
$43k – $102k per year (Estimated) • In office • Full-Time • High School Diploma • Sunnyvale
Apply
See all jobs
This is one of many
707,525 more open roles from verified company boards, updated every day.