Category Landing

Llm Eval Projects

Agent-searchable GitHub projects classified as Llm Eval.

Data Source

d1 / d1_query

1213 loaded projects. Generated at 2026-09-26T16:12:48.986Z.

Project

langfuse/langfuse

90

🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

Llm EvalDockerVercelServerless
Quality80
Agent90

The platform for LLM evaluations and AI agent testing

Llm EvalDockerVercelServerless
Quality73
Agent80

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Llm EvalDockerLibrary OnlyLocal
Quality72
Agent88

Project

comet-ml/opik

90

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

Llm EvalDockerKubernetesLibrary Only
Quality64
Agent90

The LLM Evaluation Framework

Llm EvalLibrary OnlyLocalCloud
Quality61
Agent84

Project

Arize-ai/phoenix

89

AI Observability & Evaluation

Llm EvalDockerVercelServerless
Quality57
Agent89

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

Llm EvalDockerLibrary OnlyLocal
Quality49
Agent85

Project

NVIDIA/garak

78

the LLM vulnerability scanner

Llm EvalLibrary OnlyLocalCloud
Quality44
Agent78

Fast, flexible LLM inference

Llm EvalDockerKubernetesLibrary Only
Quality40
Agent86

Project

ml-explore/mlx-lm

78

Run LLMs with MLX

Llm EvalLibrary OnlyLocalCloud
Quality39
Agent78

Project

truera/trulens

80

Evaluation and Tracking for LLM Experiments and AI Agents

Llm EvalLibrary OnlyLocalCloud
Quality38
Agent80

🐢 Open-Source Evaluation & Testing library for LLM Agents

Llm EvalLibrary OnlyLocalCloud
Quality37
Agent80