AI Quality Analyst - English
Remote | Contractor | 3-Month Engagement
4-40 hours per week | 4 hours of overlap with PST required
About the Role
As an AI Quality Analyst - English, you will evaluate a new personalization feature for an advanced AI assistant. You will assess how effectively the model uses information from previous conversations, email, search activity, and video activity to make responses more relevant, natural, and helpful.
This role combines creative prompt design, analytical evaluation, attention to detail, and strong written communication. You will create prompts based on your own personal experiences and context, then evaluate how accurately and naturally the AI incorporates that information into its responses.
You will assess dimensions such as Grounding, Integration, and Helpfulness, while identifying incorrect personalization, unsupported inferences, hallucinations, and unnatural use of personal information.
What You’ll Do
Conversational Prompt Design
- Design and execute multi-turn conversational prompts, typically involving 1-5 turns.
- Create prompts that require the AI to appropriately use personal information and experiences.
- Approach evaluations creatively to thoroughly test the model’s personalization capabilities.
- Develop prompts based on realistic personal context and intended outcomes.
AI Response Evaluation
- Evaluate model responses based on the intent established in the starting prompt.
- Determine whether personalization was applied appropriately and meaningfully.
- Identify incorrect personalization, poor inferences, forced connections, and other quality issues.
- Assess whether responses are natural, relevant, useful, and easy to use.
Grounding & Personalization Analysis
- Analyze responses for grounding issues.
- Verify that claims about you are supported by available evidence.
- Identify flawed inferences, unsupported claims, and hallucinations.
- Evaluate whether personal information is incorporated accurately rather than simply referenced unnecessarily.
Integration & Naturalness
- Assess how naturally personal information is integrated into responses.
- Identify robotic behavior, excessive explanation, or unnecessary "overnarrating."
- Evaluate whether personalization improves the overall quality of the response.
Side-by-Side Evaluation
- Rigorously compare two model responses side-by-side.
- Stack-rank responses based on overall helpfulness, usability, naturalness, and enjoyment.
- Identify subtle differences between responses that may affect the user experience.
- Provide clear reasoning for your rankings.
Evaluation Documentation
- Write concise, structured, and defensible rationales for model comparisons.
- Explicitly reference relevant turn numbers when explaining issues or positive aspects.
- Provide constructive feedback and detailed annotations.
- Extract and verify available debugging information to confirm that chat summaries and relevant data sources were properly utilized.
Data Hygiene
- Maintain strict data hygiene throughout evaluation activities.
- Delete evaluation conversations as required to prevent them from affecting future chat history.
Required Qualifications
- Strong English reading and writing skills, with a high degree of comprehension.
- Ability to evaluate nuanced and ambiguous AI responses.
- Strong analytical thinking and judgment.
- Experience designing creative, multi-turn prompts based on personal context.
- Ability to identify incorrect personalization, poor inferences, and forced connections.
- Exceptional attention to detail when comparing model responses.
- Ability to write clear, concise, and structured evaluation rationales.
- Ability to provide constructive feedback and detailed annotations.
- Strong communication and collaboration skills.
- Ability to work independently in a remote environment.
- Willingness to use your primary personal Google account, rather than a testing account, and enable the required personal data sources for genuine evaluation.
- Desktop or laptop with a reliable internet connection.
- Full-time availability within your local time zone is required, as the team operates across a global 24-hour schedule.
Education & Experience
- BS/BA degree or equivalent experience in a relevant field such as: Policy
- Law
- Ethics
- Linguistics
- Journalism
- Computer Science
- A related analytical field
- Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.
Core Skill
Domain-Specific Languages
Engagement Details
- Engagement Type: Contractor
- Engagement Length: 3 months
- Availability: At least 4 hours per day, up to 40 hours per week
- Schedule Requirement: 4 hours of overlap with PST
- Work Arrangement: Remote
What Success Looks Like
Success in this role means consistently producing accurate, evidence-based, and well-documented AI evaluations. You will demonstrate strong judgment when assessing personalization quality, distinguish genuine contextual relevance from forced or incorrect personalization, and clearly explain why one model response performs better than another.
Your evaluations should be precise, defensible, and grounded in specific evidence from the conversations being assessed.
Evaluation Process
- Shortlisted candidates will receive a Job Interest Form.
- Following profile review, an assessment will be shared.
- The assessment must be completed within 24 hours.
- Candidates selected based on assessment results will be contacted regarding pre-onboarding requirements.

