Skip to content

linny006/llm-eval-tracker

Repository files navigation

LLM Eval Tracker

Live index of LLM evaluation tools and benchmarks, refreshed every 15 minutes from GitHub

Stars Last Commit Items Updated

⭐ Star this repo to bookmark — fresh data every 15 minutes

English · 中文 · 日本語 · 한국어 · Español · Português


💡 What is this?

Automatically discovers and indexes new LLM evaluation frameworks, benchmarks, and harnesses as they appear on GitHub. Generates a structured, searchable catalog with metadata like stars, activity, and category tags. Designed for ML engineers who need to stay current without manually scanning repositories.

This list is auto-updated every 15 minutes by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.


📋 Current Items

⏰ Last updated: 2026-07-20 22:15 UTC

Data source: GitHub Search API

The table below is rewritten on every cron tick. Star the repo to bookmark.

# Name Lang Updated Description
1 Eval-core/evalcore 12 Rust 2026-07-20 Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
2 promptfoo/promptfoo 23446 TypeScript 2026-07-20 Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C
3 vitorwilher/copom-rag-service 0 Python 2026-07-20 Serviço de RAG sobre atas do Copom + boletim Focus, servido como API (FastAPI+Docker), com eval harness (golden set, LLM
4 Arize-ai/phoenix 10642 Python 2026-07-20 AI Observability & Evaluation
5 truera/trulens 3448 Python 2026-07-20 Evaluation and Tracking for LLM Experiments and AI Agents
6 sammyjdev/gnomon-eval 0 Python 2026-07-20 Honest RAG evaluation harness: judge metrics with confidence intervals, cost and latency first-class, offline-first.
7 verifywise-ai/verifywise 321 TypeScript 2026-07-20 Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI framewo
8 SFX-TECH/sfx-lead-intelligence 0 2026-07-20 SFX Lead Intelligence Command Center: local-LLM hub plus lead dashboard, quality lifted 61 to 99 percent via a ground-tr
9 cklxx/ckl-bench 1 HTML 2026-07-20 ckl's personal benchmark for doc writing, infra code, and paper reading — one-click evaluation of the latest models via
10 NoesisVision/nasde-toolkit 10 Python 2026-07-20 CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini ind
11 Giskard-AI/giskard-oss 5662 Python 2026-07-20 🐢 Open-Source Evaluation & Testing library for LLM Agents
12 tkarim45/multi-agent-eval 0 Python 2026-07-20 Does multi-agent beat single-agent? Benchmarks a planner→workers→critic system vs single-agent on quality/cost/latency —
13 tkarim45/agent-eval-harness 0 Python 2026-07-20 Agent eval harness — measure task success, tool-call accuracy, step efficiency, and cost for tool-using LLM agents (Clau
14 iZenDeveloper/auditai 0 Python 2026-07-20 Developer-first LLM/RAG safety audits for CI/CD — faithfulness, relevancy, prompt injection (BYOK OpenAI/xAI). pip insta
15 monospaceai/evaldata 3 Python 2026-07-20 Evaluate AI-generated SQL with pytest.
16 jeremylongshore/j-rig-skill-binary-eval 0 TypeScript 2026-07-19 Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score e
17 lftherios/session-link 0 Go 2026-07-19 A local-first CLI that turns any LLM session into a permanent URL you can inspect, share, and revisit.
18 saddled-panicattack529/idea-evaluation-pipeline 0 2026-07-19 Streamline research idea evaluation for finance and economics to reach top journal quality using an iterative, AI-assist
19 Kondwani10/Origin-Continuum 0 2026-07-19 🌐 Define and explore the Origin ↔ Continuum framework, ensuring proper attribution and continuity in dependency relation
20 Sans-cell-art/-Project-Phoenix-The-E-Waste-Supercomputer- 0 2026-07-19 ♻️ Transform e-waste into a powerful, low-cost cloud operating system, unlocking computing potential and promoting resou
21 bhavya7995/AI_governance 1 PowerShell 2026-07-19 🤖 Streamline AI-assisted development with a governance kit for rules, enforcement, and decision-making, ensuring speed a
22 RudrenduPaul/memtrust 0 Python 2026-07-19 Independent, reproducible CLI benchmark harness for agent-memory backends (MemPalace, Mem0, Zep/Graphiti, OpenViking) --
23 AshC2004/llm-eval-harness 0 Python 2026-07-18 Automated groundedness scoring, hallucination detection, and regression tracking for RAG and agent outputs
24 homemade-software-inc/completion-kit 1 Ruby 2026-07-20 Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and c
25 pdxlab/trustmodel-mcp-server 0 TypeScript 2026-07-17 TrustModel MCP Server — trust evaluation, red-team, and governance for AI agents via the Model Context Protocol. npm: @t
26 valbaudo/awf 1 Go 2026-07-16 Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every s
27 sahuno/confounded 0 R 2026-07-15 A gym for scientific judgment. AI agents render fatally flawed analyses at publication quality — we measured whether any
28 vijayarjun7/flipcheck 0 TypeScript 2026-07-14 LLM self-consistency and sycophancy diagnostic — catches contradictions and false-premise acceptance hiding behind confi
29 QuesmaOrg/BinaryAudit 96 Shell 2026-07-14 An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.
30 harnexa/nexa-gauge 40 Python 2026-07-20 An graph-eval framework for LLM's
31 multivon-ai/multivon-eval 8 Python 2026-07-13 Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI
32 attogram/ollama-multirun 16 Shell 2026-07-12 Run a prompt against all, or some, of your models running on Ollama. Creates web pages with the output, performance stat
33 nishant20/testgen-eval 0 TypeScript 2026-07-10 An agent-eval harness that gates and judges testgen's generated test suites via MCP — deterministic checks plus an LLM-j
34 ahwurm/localshift 3 Python 2026-07-09 Migrate headless Claude/AI workloads to local LLMs with a derived, per-workload quality eval — cron job in, zero-margina
35 coffee-converter/onthemoney 0 Python 2026-07-14 An AI agent that answers questions about U.S. campaign money against real FEC data. Source-cited, graded on a public eva
36 MihirBindu/dynamo-log-report-fix 0 Python 2026-07-09 A broken Terminal-Bench 2 (Harbor) task, repaired: reproducible env, solution-leak removed, gameable verifier replaced w
37 christianmacion26/judge-harness 0 Python 2026-07-08 LLM-as-judge validated vs humans (Cohen's κ), with position-bias exposed.
38 plwslpld-arch/loopward 1 TypeScript 2026-07-07 Stress-test the tool-routing decision inside an agent loop: audit confusable tools, red-team routing robustness, compare
39 ContextJet-ai/skillvitals 1 Python 2026-07-06 Vital signs for your agent skills: measure whether a skill triggers when it should and whether it actually helps, on a c
40 IonDen/mlx-quant-fidelity 1 Python 2026-07-04 Measure MLX quantization quality loss — KL divergence, perplexity, top-token agreement for KV cache and weights
41 Sushant-Dagar/agent-eval-harness 0 Python 2026-07-04 CI-integrated eval harness for LLM agents — intent accuracy, retrieval precision/recall/MRR, and hallucination detection
42 lokesh75-kank/agenteval 0 TypeScript 2026-07-03 Reliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grou
43 Ruthwik-Data/finrag-eval 0 Python 2026-07-02 Local RAG eval on real SEC 10-Ks that catches confident financial hallucinations — and surfaced a metric bug now merged
44 hydrangeas20/safetylens 0 Jupyter Notebook 2026-07-01 Empirical study of evaluation robustness in large language models. Compares benchmark style evaluation prompts with real
45 kilocommits/campaign-eval-harness 0 Python 2026-06-30 An LLM-as-judge harness that scores AI-generated campaign phone scripts against a weighted quality rubric with a real Ha
46 Merchantlee99/myrealtrip-cancel-recovery-eval 0 Python 2026-06-30 Codex eval plugin for auditing cancel-recovery replacement-product exposure policies.
47 reaatech/agent-eval-harness 0 TypeScript 2026-07-20 End-to-end agent evaluation — trajectory eval, tool-use correctness, cost-per-task, latency budgets, regression suites w
48 TheAnacondA57/BidAgent 1 Python 2026-06-29 RAG agentique sur des documents de concession télécom publique (DSP/RIP), pensé eval-first et contrôlé en CI.
49 chquandogong/mission-spec 0 TypeScript 2026-06-29 Mission Spec — AI 에이전트 워크플로를 위한 task contract layer
50 G59-Toneli/dataset-eval-skill 1 JavaScript 2026-06-25 A Claude skill for building golden sets to test AI systems — matching, RAG, LLM-as-judge — without false greens.

🔍 How it works

Every 15 minutes, a GitHub Action runs tracker.py. That script:

  1. Fetches the latest state from GitHub Search API.
  2. Diffs against data/items.json (the previous snapshot).
  3. Rewrites the table above between the <!-- TRACKER_TABLE_* --> markers.
  4. Commits feat: +N added, -M removed (timestamp) if anything changed.

No external services. No paid APIs. Just a public data source and a free GitHub Action.


🤝 Contributing

See CONTRIBUTING.md — usually you don't need to: the tracker keeps itself current. If you spot a data-source bug or want to suggest a new column for the table, open an issue.


🔗 Related live trackers

If you find this useful, you might also like these other auto-updated trackers from the same maintainer — same mechanism, different upstream:


📜 License

MIT — see LICENSE.