585,223open jobs
25,872companies
81,805added this week
Browse all
Salary
$200k – $300k per year
Location
In office (San Jose)
Employment
Full-Time
Overview
Company
Impact
Profile match

About Tessera Labs

Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run.

Every large enterprise carries the same weight - decades of accumulated process, data, and code that no longer match the business it has become. Changing any of it is a program measured in years and hundreds of millions of dollars, staffed by armies of consultants, and it fails more often than anyone admits. Most companies have quietly accepted this as the cost of being large.

We don't. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design - SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft - and tied to none of them.

Two things make this hard, and they're the reason the research is interesting. Governance: every action is logged, traceable, and reversible, because our customers are regulated and these are the systems that close their books. And generality: the platform has to work on landscapes it has never seen, at companies whose complexity is genuinely unique to them.

We sell a product, not a service. Our people are here to make the product successful, not the other way around - which is also why research here is a durable investment rather than a line item on an engagement.

We raised a $60M Series A led by Andreessen Horowitz, with Foundation Capital, Myriad Venture Partners, and Osage University Partners participating.

About the role

We're looking for a Research Scientist to set and pursue a research agenda for reliable long-horizon agents operating inside real enterprises.

Frontier labs optimize for general capability, and the public agent benchmarks are mostly sandboxes. Very little rigorous work exists on what it takes for an agent to reason across a system with nineteen years of undocumented decisions in it, plan a change across forty coupled steps, recover when step twelve reveals the model of the world was wrong, and be right often enough that a CFO signs the go-live. Almost nobody has the landscapes, the traces, or the customers to study it. We do.

Two properties make this an unusually good research setting. First, much of the task space is verifiable - a transformation either produces a system that builds, passes regression, and behaves equivalently, or it doesn't. That's a real reward signal, not a preference model. Second, the parts that aren't verifiable are where the interesting work is: is this reconciliation correct, or merely plausible? Was retiring that capability the right call? Designing reward and evaluation across that boundary is the central research question here.

You'll invent methods rather than only apply them, work with Research Engineers who help you run at scale, and hear from a product team within weeks whether you were right.

We'd like you to publish. Not everything, and never at the expense of shipping - but the work here is novel enough to be worth writing down.

One thing worth knowing up front: we post-train open-weight models on rented clusters and buy more compute when a result justifies it. We're constrained relative to a frontier lab. If your research only works at ten thousand GPUs, this is the wrong place.

What you'll do

  • Set and pursue a research agenda on reliable long-horizon agentic behavior in real enterprise environments - you decide which questions matter, and defend the choice.

  • Invent and validate methods for post-training agents on transformation work: reward design where verification is partial, delayed, or contested; RL formulations for long-horizon planning and tool use; curriculum and data strategy. Post-training and RL are the core of this role.

  • Define how an agent remembers. Memory architecture for runs that span forty steps and days of wall-clock - what persists, how it's structured and retrieved, how it's revised when the world turns out to be different, and how a model is trained to use it rather than ignore it. This is one of the least solved problems in agentic AI and one of the most consequential for us.

  • Own the question of what to measure. Develop evaluation methodology whose scores predict customer-observed correctness, and demonstrate where cheap automated proxies quietly fail.

  • Study how multi-agent systems fail - error compounding across long trajectories, planning under partial observability of a landscape, delegation and verification between agents - and design against it.

  • Work on the verification problem directly: how an agent, or another agent, establishes that a change preserved behavior when no test covers it.

  • Solve how a system represents an enterprise to itself. Turning process, data, and code into an ontology or knowledge graph an agent can reason over reliably - and one that stays true as the underlying systems change - is a research problem we own rather than inherit.

  • Investigate what post-training compute and data quantity buy us across model scales we can actually afford, and where the returns bend.

  • Turn findings into things that ship, with Research Engineering and product.

  • Publish papers, technical reports, and open-source artifacts, and represent Tessera's research externally.

  • Raise the team's research bar: review experiment designs, mentor engineers moving into research, and be the person who asks whether the result is real.

Representative projects

  • Designing a reward formulation for transformation work that doesn't collapse into reward hacking when the agent discovers it can pass the regression suite by removing the branch the tests don't reach.

  • Developing an evaluation methodology whose scores track expert-reviewed correctness on changes no automated test can verify, and showing where the cheap proxies disagree.

  • Characterizing error compounding across a forty-step transformation plan and proposing a verification scheme that measurably arrests it.

  • Designing a memory architecture for multi-day agent runs and showing, with evidence rather than anecdote, that it beats stuffing the context window.

  • Showing that grounding an agent in a learned ontology of a customer's landscape beats retrieval over raw artifacts - or finding that it doesn't, and saving us a year.

  • Running a post-training study across three or four open-weight model sizes and publishing what it says about where domain data quality beats parameter count.

  • Releasing an enterprise agent environment and task suite as an open benchmark, and being honest in the paper about where it doesn't transfer.

You may be a good fit if you

  • Have an MS or PhD in CS, ML, statistics, math, physics, or a related field - or research experience of comparable depth without the credential.

  • Have a track record of original research in ML: publications at NeurIPS/ICML/ICLR/ACL-tier venues, influential open-source work, or results inside a lab that clearly moved a frontier system.

  • Have deep expertise in at least one of: post-training and RL for LLMs, agent memory and long-context reasoning, knowledge representation and structured reasoning, or evaluation methodology. Depth in RL tuning is the single strongest signal for this role.

  • Are hands-on. You write the code, run the experiments, and read the logs. This is not a role that directs research from a distance.

  • Have exceptional research judgment - you pick the question well and kill your own ideas quickly when the evidence says to.

  • Hold rigor and urgency at once. We need results that are true and results that arrive.

  • Write clearly, and enjoy explaining a technical argument to people who aren't researchers.

Strong candidates may also have

  • Experience owning a research direction end to end at a frontier lab or a strong academic group.

  • Published work on agents, tool use, RL for LLMs, code generation or repair, reasoning, or evaluation.

  • Experience with verifiable-reward RL, or with the failure modes of reward models where verification is incomplete.

  • Experience with knowledge representation, ontologies, or neurosymbolic approaches to structured domains.

  • Experience with training runs at meaningful scale, including the ones that failed for three weeks.

  • A history of mentoring researchers and engineers into better work.

  • Curiosity about enterprises as a research domain. The problems here are strange and specific, and they reward people who find that interesting.

Join the Future of AI at Tessera Labs

We're looking for someone who enjoys building reliable, scalable systems that help teams move faster. If you take pride in cutting-edge AI, value clear ownership, and want to have a meaningful impact in a fast-moving environment, you’ll fit right in.

No third party may recruit, solicit candidates, publish job opportunities, use Tessera Labs’ name or branding, or represent that they are acting on behalf of Tessera Labs without prior written authorization. Any such activity conducted without explicit written consent is strictly prohibited.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
585,223 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$22k – $46k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru
Python
Go
SQL
Python
FastAPI
Celery
pySpark
Pydantic
Go
Dapr
Databases
PostgreSQL
Redis
Databricks
Delta Lake
pgvector
Qdrant
AI/ML
Copilot
Cursor
LangGraph
LangChain
Spark
Prompt Engineering
Function Calling
AI Agents
CrewAI
LLM
RAG
Hallucination
Anomaly Detection
OpenAI
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
DevOps
Azure
CI/CD
Git
Docker
Kubernetes
Self-Healing
Vector
SLI/SLO/SLA
Analytics
ETL/ELT
Azure Data Factory
Management
Agile
QA
Pytest
Apply
$19k – $47k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
TypeScript
SQL
Node JS
Apex
Apex
MuleSoft
Databases
Snowflake
Oracle
DevOps
CI/CD
AWS
Management
Agile
Scrum
Apply
$37k – $85k per year (Estimated) • In office • Full-Time • PhD • Mexico City
AI/ML
AI Agents
Agentforce
Analytics
Microsoft Excel
Marketing
Salesforce
Apply
In office • Full-Time • PhD • Paris
AI/ML
AI Agents
Agentforce
Marketing
Salesforce
Apply
$69k – $193k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Munich
AI/ML
AI Agents
Agentforce
Marketing
Salesforce
Apply
$179k – $200k per year • In office • Contractor • Los Angeles
Apply
Senior QA Engineer 10 days ago
$50k – $60k per year • Remote • Full-Time
Python
DevOps
CI/CD
Kubernetes
Platform Engineering
Cybersecurity
SOC 2
QA
Cypress
Playwright
Pytest
k6
Locust
Apply
$50k – $60k per year • Remote • Contractor
DevOps
GCP
Azure
Kubernetes
Platform Engineering
Cybersecurity
Burp Suite
OWASP ZAP
Trivy
Semgrep
ISO 27001
Grype
SOC 2
Threat Modeling
Apply
$50k – $60k per year • Remote • Contractor • 8+ years exp
AI/ML
ISO 42001
Cybersecurity
NIST 800-53
FedRAMP
Apply
$154k – $233k per year • In office • Full-Time • 6+ years exp • Los Angeles
AI/ML
AI Agents
Edge AI
Management
Agile
Scrum
Apply
DMTS 4 hours ago
$89k – $249k per year (Estimated) • In office • Boise • San Jose
Apply
$112k – $215k per year • Equity • In office • Full-Time • 8+ years exp • San Francisco • San Jose
Apply
$135k – $234k per year • Equity • In office • Full-Time • 5+ years exp • San Jose • San Francisco • Seattle • Los Angeles • Lehi
Design
Figma
Canva
Apply
$70k – $90k per year • In office • 1+ year exp • Bachelor's Degree • San Jose
Analytics
Microsoft Excel
Apply
$144k – $205k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Jose
Cybersecurity
Zscaler
Zero Trust
Management
Agile
Apply
See all jobs
This is one of many
585,223 more open roles from verified company boards, updated every day.