The Role
Research at Aristotle works on the open problems in building a realtime AI tutor. The biggest one right now is evaluation. Tutoring is multi-turn and adaptive, there is no verifiable reward signal, and existing benchmarks test single turns in toy scenarios. We are building the evaluations that make tutoring quality measurable, including simulated students realistic enough to test tutors against, and using them to improve our tutor in production.
You will own this work end to end: forming hypotheses, running experiments on real production sessions, and shipping findings into the product. You will work directly with one of the co-founders leading research, and with data that few groups have: full multimodal tutoring sessions at scale, with voice, whiteboard state, and student history.
This starts as a summer contract, with the intent to continue if it goes well. Full-time in SF is preferred; we are flexible on hours and location for the right person. We support publishing where we can, and we will tell you the constraints up front.

