Deutsche Telekom Digital Labs
We are looking for an experienced SDET - AI Testing to drive quality engineering initiatives for GenAI and Agentic AI applications. The ideal candidate should have strong expertise in traditional test automation along with hands-on experience evaluating Large Language Models (LLMs), AI Agents, RAG systems, and conversational AI platforms. You will be responsible for building scalable AI testing frameworks, defining evaluation strategies, automating AI quality checks, and ensuring reliability, accuracy, safety, and performance of AI-powered products.
The candidate will have responsibilities across the following functions:
AI Testing and Evaluation:
- Design and execute testing strategies for LLM, GenAI, RAG, and Agentic AI systems.
- Build automated evaluation pipelines for AI applications using DeepEval, OpenAI Evals, and custom evaluation frameworks.
- Define and monitor AI quality metrics including: Faithfulness, Hallucination Detection, Context Relevance, Answer Relevancy, Toxicity, Bias & Safety.
Latency and Cost:
- Create golden datasets and benchmark suites for continuous model evaluation.
- Validate AI agent workflows involving planning, reasoning, memory, and tool usage.
Automation and Quality Engineering:
- Develop robust automated test frameworks for APIs, backend services, and AI applications.
- Build CI/CD-integrated testing pipelines for AI systems.
- Automate regression, integration, performance, and end-to-end testing.
- Design synthetic test data and evaluation datasets for AI workflows.
- Implement observability and monitoring mechanisms for AI applications in production.
Agentic AI Validation:
- Test multi-agent and autonomous workflows.
- Validate tool-calling, function-calling, memory management, and orchestration logic.
- Assess agent reliability, consistency, and failure recovery mechanisms.
- Evaluate prompt effectiveness and prompt regression issues.
Performance and Reliability:
- Conduct load and stress testing for AI-powered systems.
- Measure response quality, latency, throughput, and cost efficiency.
- Perform root cause analysis for AI failures and model performance degradation.
The core requirements for the job include the following:
AI Testing and GenAI:
- Hands-on experience testing LLM and Generative AI applications.
- Strong experience with: DeepEval, OpenAI Evals, RAG Evaluation Frameworks, Prompt Testing, Agent Evaluation Frameworks.
- Understanding of: LLM architectures, RAG systems.
Embeddings and Vector Databases:
- Agentic AI workflows.
- Prompt Engineering.
- Automation Testing.
- Strong expertise in: Python (Preferred), Java, API Testing, Test Automation Frameworks, PyTest / Robot Framework.
- Experience with REST APIs, microservices, and distributed systems.
- CI/CD experience using Jenkins, GitHub Actions, or GitLab CI.
Data and AI Ecosystem:
- Experience with: LangChain, LangGraph, CrewAI, AutoGen, LlamaIndex.
- Familiarity with vector databases such as Pinecone, Weaviate, ChromaDB, FAISS, Cloud & DevOps, AWS, Azure, or GCP; Docker and Kubernetes; and monitoring and observability tools.
Preferred Qualifications:
- Experience testing enterprise GenAI or Agentic AI applications.
- Knowledge of Responsible AI, AI Governance, and AI Safety.
- Experience with ML model evaluation and MLOps workflows.
- Exposure to AI red-teaming and adversarial testing.
- Experience building AI quality dashboards and automated evaluation reports.
Success Metrics:
- AI evaluation coverage across all critical workflows.
- Reduction in hallucination and response quality issues.
- Automated AI regression testing coverage.
- Production defect leakage reduction.
- Faster release cycles through automated AI validation.
Nice to Have:
- Experience with AI observability platforms such as: Langfuse, Arize AI, Weights & Biases.
- Experience evaluating multimodal AI systems.
- Contributions to open-source AI testing frameworks.
