LLM Benchmarks & Leaderboards
Showing 96 of 192 LLM Benchmarks & Leaderboards tools on page 1 of 2.
LLM Benchmarks & Leaderboards results
ACLUE
Official github repo for ACLUE, an evaluation benchmark focused on ancient Chinese language comprehension -…
Agent-Evaluation
A generative AI-powered framework for testing virtual agents. - awslabs/agent-evaluation
Agent Evaluation Framework 2026: Metrics, Rubrics & Benchmarks
Agent Evaluation Framework 2026: Metrics, Rubrics & Benchmarks | Galileo
Agent Evaluation Readiness Checklist
A practical checklist for agent evaluation: error analysis, dataset construction, grader design, offline & online…
AhaApple
AhaApple. AI Idea Generator. one click, many useful ideas.
Ai-Ethics-Evaluation-Report-In-Healthcare
Ai-Ethics-Evaluation-Report-In-Healthcare is a llm benchmarks & leaderboards tool. An evaluation report in healthcare
AI-Infra-Guard (Tencent)
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP…
Ai-Llm-Comparison
A website where you can compare every AI Model ✨.
AI Meme Arena
Add more credibility to your site - get a premium domain today. Straight-forward shopping experience.
AI Tools Arena
Explore our comprehensive AI tools list, showcasing the best solutions for diverse industries. Improve your workflow…
Aider Polyglot Leaderboard
Aider's leaderboard ranking models on multi-language code editing tasks.
AlpacaEval
An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.
Alpha Arena
Alpha Arena is a llm benchmarks & leaderboards tool. Live trading performance benchmark for AI models in real markets.
Appworld-Leaderboard
Leaderboard Repository for "AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding…
Arena
An open platform for crowdsourced AI benchmarking, hosted by researchers at UC Berkeley SkyLab.
Arena Commerce AI
Unlock the power of Commerce AI to enhance customer experiences and support, streamline operations, and drive more…
Artificial Analysis
Artificial Analysis is a platform that provides AI model and service provider comparisons and benchmarks to help…
Athina AI
Athina Flows Develop Observe Pricing Docs Log in Sign up Ship AI to prod 10x faster Athina is a collaborative AI…
Auto-evaluator
: a lightweight evaluation tool for question-answering using Langchain !
AutoRecruiter
India's AI-native recruitment platform. Hire in 7 days with a transparent flat success fee, far less than…
awesome-ai-agent-papers
Curated 2025–2026 papers on agent engineering, memory, eval, and workflows
Bananalyzer
Open source AI Agent evaluation framework for web tasks - reworkd/bananalyzer
Baseline-Agent
Baseline-Agent is a llm benchmarks & leaderboards tool. This is a simple AI Agent used to test the Bloodrock CORE…
BeHonest
BeHonest: Benchmarking Honesty in Large Language Models More Research Abel MathPile ReAlign BenBench OlympicArena…
BenchGecko
The data layer of the AI economy. AI model benchmark leaderboard with cross-provider pricing comparison across…
BenchLM.ai
Compare 216 ranked models and 385 tracked AI models across 414 benchmarks with BenchLM scoring, pricing, context…
Benchmark Email
Ditch the clunky tools. Benchmark is the powerfully simple email marketing software that helps you send campaigns in…
Berkeley Function-Calling Leaderboard
Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM
@bmdhodl
Building a one-person AI-operated holding company. AgentGuard (pip install agentguard47), showwork, measured…
braindecode
Deep learning software to decode EEG, ECG or MEG signals, providing standardized neural network models,…
BuildArena
First physics-aligned interactive benchmark for LLM agents in engineering construction, designing…
buyer-eval-skill
B2B software vendor evaluation skill for Claude Code — domain-expert questions, vendor AI agent conversations,…
Chat-Agent-Evalution
Evaluating the LLM Chat Agent on multiple evaluation benchmarks. - khuzaimakt/Chat-Agent-Evalution
CheckMyIdea
Validate your side business idea in minutes with our AI-powered evaluation service. Maximize your chances of success…
Chinese Large Model Leaderboard
Chinese Large Model Leaderboard is a llm benchmarks & leaderboards tool. an expert-driven benchmark for Chineses LLMs.
Claw-Eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans. - claw-eval/claw-eval
Cognitive-Security-Ai-Powered-Threat-Agent-Evaluation-For-Impact-On-Assets.
Cognitive-Security-Ai-Powered-Threat-Agent-Evaluation-For-Impact-On-Assets. is an AI tool for llm benchmarks &…
CompMix
CompMix: A Benchmark for Heterogeneous Question Answering.
Confident AI
Confident AI is the AI quality platform for enterprise teams to standardize AI evals and observability across the…
Countless
Compare AI models easily! All providers in one place. Find the best LLM for your needs with our comprehensive…
CRO Benchmark
Get a full AI CRO audit across 248 CRO best practices. Built for ecommerce teams. We surface the leaks costing you…
Cross-Model-Evaluation-Judging-Ai-Ethics-And-Alignment-Responses-With-Language-Models
This study aims to evaluate the quality of previously generated responses using various large language models (LLMs)…
DeepEval
DeepEval is the open-source LLM evaluation framework for testing and benchmarking LLM applications — 50+…
Demystifying Evals for AI Agents
Demystifying evals for AI agents \ Anthropic Skip to main content Skip to footer Research Policy Commitments Learn…
Design Arena
Design Arena is the largest global crowdsourced benchmark for design. Challenge, Vote, Crown your Winner.
Designing AI-Resistant Technical Evaluations
What we learned from three iterations of a performance engineering take-home that Claude keeps beating.
Dingo
Dingo is a llm benchmarks & leaderboards tool. Dingo - A Comprehensive Data Quality Evaluation Tool
Diplomacy-Llm
Public LLM benchmark using the results of Diplomacy games played by multiple LLM agents. - lukepoo101/diplomacy-llm
DreamBench++
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
Driving the Agent Quality Flywheel from Your Coding Agent
Google's June 2026 account of automating the eval-optimize loop for coding agents: independent AutoRaters grade…
Dubesor LLM Benchmark table
Dubesor LLM Benchmark table - Small-scale manual LLM performance comparison benchmark
Edge Arena
Edge Arena puts your business decisions on trial — competing AI agents challenge assumptions, expose weaknesses, and…
Eval
Eval is a llm benchmarks & leaderboards tool. Better coding workflow with smart assistance.
Eval Awareness in Claude Opus 4.6's BrowseComp Performance
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI…
Eval-Driven Development: Build and Evaluate Reliable AI Agents
Learn how to build reliable AI agents with our 8-stage evaluation framework. We explore DeepEval, multi-turn…
FaceRate.ai
FaceRate.ai offers a face attractiveness test, facial analysis, and golden ratio face tests. Get an in-depth…
FELM
FELM: Benchmarking Factuality Evaluation of Large Language Models
Fl_Llm_Benchmark_Dataset
This code collects congressional/parliamentary dataset across US, UK and Canada
Fogworkflowsim
An Environment for Simulation and Performance Evaluation of Workflows in Fog Computing
Foundry
Produce high-quality enterprise data, evaluate reliably, and optimize performance at scale—without web drift, IP…
Giskard
Giskard is a llm benchmarks & leaderboards tool. Testing & evaluation library for LLM applications, in particular RAGs
Go/NoGo
Evaluate RFPs instantly and get AI-powered insights to make better Go/No-Go decisions.
Goodai-Ltm-Benchmark
A library for benchmarking the Long Term Memory and Continual learning capabilities of LLM based agents. With all…
gpt-oss playground
Demo platform for OpenAI's open-weight models for developers.
GradeWrite
GradeWrite.AI is the ultimate AI assistant for grading assignments. GradeWrite streamlines grading process with…
GSM Arena
GSMArena Turnstile check One quick check before you continue... Continue
HeHealth
HeHealth is a llm benchmarks & leaderboards tool. Reliable sexual health evaluation
HELM
HELM is a llm benchmarks & leaderboards tool. Holistic evaluation across 42 scenarios
Hermes 3 Llama 3.1 405B
Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic…
HumanEval
HumanEval is a llm benchmarks & leaderboards tool. Python code generation correctness
Hume AI
Real human ratings, in a single API call. The human evaluation layer for voice, speech, and conversational AI.
HypeBridge
AI-powered influencer evaluation and discovery platform. Make data-driven partnership decisions in seconds. Join…
IdolCrush.ai
IdolCrush.ai — a new interactive AI idol experience where users become producers, create and chat with your dream…
imgsys
A generative AI arena where you can test different prompts and pick the results you like the most. Check-out the…
Impact-Academy
Auto-Enhance meta-benchmark, to measure the ability of LLM agents to improve other LLM agents -…
InfiBench
a benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world…
instruct-eval
This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on…
IntelliServer
simplifies the evaluation of LLMs by providing a unified microservice to access and test multiple AI models.
Internintelligence_Ai_Ethics_And_Bias_Evaluation
This project evaluates the fairness of a machine learning model trained on the Adult Income Dataset to predict…
Interview: About deployment, evaluation, and testing of agents with Sully Omar, the CEO of Cognosys AI
We asked the founder of Cognosys, Sully Omar, about his experience with building a product for no-code users in the…
Introduction to Generative AI | SqillPlan
: introduction to Generative AI, including models such as GANs, Variational Autoencoders, Autoregressive Models, and…
JARVIS
NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT),…
Jury
Jury helps with llm benchmarks & leaderboards workflows.
Just-Eval
A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.
LangSmith
Complete AI agent and LLM observability platform with tracing and real-time monitoring. Debug agents, find failures…
LawBench
a benchmark designed to evaluate large language models in the legal domain.
Lawful-Good
Benchmark for assessing legal capabilities of LLM agents - dluo96/lawful-good
LiveCodeBench
LiveCodeBench is a holistic and contamination-free evaluation benchmark of LLMs for code that continuously collects…
LiveCodeBench
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Llama 3.2 3B Instruct
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language…
Llama 3.3 70B
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in…
Llemma
Open language model for mathematics (7B/34B) trained on Proof-Pile-2, outperforming Minerva at equal scale on MATH…
Llf-Bench
A benchmark for evaluating learning agents based on just language feedback - microsoft/LLF-Bench
Llm-Agent-Ask-For-Help
Benchmark LLM Agents' abilities to quit sequential tasks as early as possible. - dillonmsandhu/llm-agent-ask-for-help
Llm-Agent-Benchmark-List
A banchmark list for evaluation of large language models. - zhangxjohn/LLM-Agent-Benchmark-List
Llm-Benchmarks
Llm-Benchmarks is a llm benchmarks & leaderboards tool. LLM benchmark tools for LMDeploy, vLLM, and TensorRT-LLM.