As featured in
Follow
Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
DeepEval is the open-source LLM evaluation framework for testing and benchmarking LLM applications — 50+ plug-and-play metrics for AI agents, RAG, chatbots, and more.
News
DeepEval for TypeScript, now fully open-source DeepEval - The LLM Evaluation Framework
DeepEval's TypeScript SDK is out in beta. Every metric, model, and tracing integration you know from Python, running as a gate in your CI/CD pipeline.
Read more
Report
Vibe Coding with DeepEval DeepEval - The LLM Evaluation Framework
Although DeepEval is great as an AI quality validation suite - pytest assertions, regression gates, CI/CD failure tracking - that's only half the use case.
Read more
Report
RAG Evaluation: The Definitive Guide to Unit Testing RAG in CI/CD - Confident AI
Unit-test RAG applications in CI/CD with DeepEval: score retrieval and generation separately and block regressions on every commit before they reach production.
Read more
Report
