DeepEval
Open-source LLM evaluation framework with 50+ metrics for testing agents, RAG, and chatbots
DeepEval is an open-source evaluation framework for LLM applications that works like Pytest but specialized for unit testing LLM outputs. It provides 50+ research-backed evaluation metrics including G-Eval, relevance, factual consistency, bias, and toxicity detection. Covers AI agents, RAG pipelines, and chatbots with support for synthetic dataset generation, red teaming, and CI/CD integration. Confident AI is the commercial platform layer adding collaboration, visualization, production tracing, and observability. 3M+ monthly downloads.
Pricing: Free / monthly subscriptions
DeepEval Alternatives
Explore 55 products in the Observability & Analytics category. View all DeepEval alternatives.
Giskard
Eliminate risks of biases, performance issues & security holes in AI models. In <10 lines of code.
Evidently AI
Open-source ML and LLM evaluation with 100+ built-in metrics and CI/CD integration
Future AGI
Open-source platform for testing, monitoring, and improving AI agents with tracing, evals, guardrails, and gateway
Traceloop
Open-source LLM observability built on OpenTelemetry, with automatic instrumentation for major providers and frameworks
Galileo
AI evaluation and observability platform with hallucination detection and real-time guardrails
Work on DeepEval? Feature it at the top of Observability & Analytics.
Is your product missing?