Project
langfuse/langfuse
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Category Landing
Agent-searchable GitHub projects classified as Llm Eval.
Data Source
1213 loaded projects. Generated at 2026-09-26T16:12:48.986Z.
Project
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Project
The platform for LLM evaluations and AI agent testing
Project
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Project
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Project
The LLM Evaluation Framework
Project
AI Observability & Evaluation
Project
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Project
the LLM vulnerability scanner
Project
Fast, flexible LLM inference
Project
Run LLMs with MLX
Project
Evaluation and Tracking for LLM Experiments and AI Agents
Project
🐢 Open-Source Evaluation & Testing library for LLM Agents