LLM Benchmarks & Leaderboards
Showing 96 of 192 LLM Benchmarks & Leaderboards tools on page 2 of 2.
LLM Benchmarks & Leaderboards results
LLM Evaluation
Home | LLM Evaluation Skip to main content Link Menu Expand (external link) Document Search Copy Copied Home Papers…
Llm_Evaluation_For_Gene_Set_Interpretation
Code space for 'Evaluation of large language models for discovery of gene set function' -…
Llm-Interest
Llm-Interest is a llm benchmarks & leaderboards tool. Tool to collect LLM eval topics
LLM Labs
Compare and test language models.
Llm_Model_Evaluation
LLM Model Evaluation for tmmluplus datasets. Contribute to LiuYuWei/llm_model_evaluation development by creating an…
LLM Price Check
Explore cost-effective LLM API solutions with LLM Price Check. Instantly compare updated prices from major providers…
Llm_Scavengerhunt
A New Benchmark made for tasks for Large Language Model Agents - UC Berkeley Scavenger Hunt -…
LLM Stats
LLM Stats, the most comprehensive LLM leaderboard, benchmarks and compares API models using daily‑updated,…
Llmagentoodgym
OOD benchmark study for LLM agents based on BrowserGym and AgentLab from ServiceNow. - rowingchenn/LLMAgentOODGym
Llmdataparser
LLMDataParser is a Python library that provides a collection of parsers for various benchmark datasets used in the…
Llmscenarioeval
Scenario-based Evaluation dataset for LLM (beta). Contribute to Turing-Project/LLMScenarioEval development by…
Lmsys Arena Hard
бенчмарк, основанный на сравнении качества ответов на реальные человеческие запросы. В роли судьи, правда, выступает…
LMSYS Chatbot Arena Leaderboard
LMSYS Chatbot Arena is a crowdsourced open platform for LLM evals. Collected over 1,000,000 human pairwise…
M3CoT
Leaderboard | M 3 CoT M 3 CoT Home Download Evaluation Annotation Paper Code Citation Contact Leaderboard Explore…
MatBench
MatBench is a llm benchmarks & leaderboards tool. Materials informatics benchmark
MathEval
MathEval是一个专注于全面评估大模型数学能力的测评基准。共包含22个数学领域测评集和近30K道数学题目,旨在全面评估大模型在包含算术,小初高竞赛和部分高等数学分支在内的各阶段、难度和数学子领域的解题能力表现,既可以作为现阶段大模…
Megaminer-Tinyarena
A tiny arena for testing Megaminer AI agents. Contribute to drusepth/Megaminer-Tinyarena development by creating an…
MiniWoB++
A collection of over 100 web interaction environments, along with JavaScript and Python interfaces.
MixEval
a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a…
MixEval
A reliable click-and-go evaluation suite compatible with both open-source and proprietary models, supporting MixEval…
ML6 x AISO Agent Workshop
ML6 x AISO Agent Workshop (March 2026). Contribute to ml6team/AISO-workshop development by creating an account on…
MMedBench
Medical Multilingual Benchmark Blog --> GitHub Huggingface MMedBench A Medical Benchmark for Multilingual…
MMToM-QA
MMToM-QA Leaderboard MMToM-QA: Multimodal Theory of Mind Question Answering ACL 2024 Outstanding Paper Award…
Mobilebench
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents - XiaoMi/MobileBench
Non finito
Non finito is a llm benchmarks & leaderboards tool. Model evaluation and sharing made simple.
Ollama Grid Search
Desktop utility for systematic model evaluation. Test multiple models, prompts, and inference parameters…
OLMO-eval
OLMO-eval is a llm benchmarks & leaderboards tool. a repository for evaluating open language models.
OlympicArena
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
oobabooga benchmark
oobabooga benchmark oobabooga benchmark The list is sorted by size (on disk) for each score. Highlighted = Pareto…
Open-RAG-Eval
RAG evaluation without the need for "golden answers" - vectara/open-rag-eval
OverallGPT
OverallGPT lets you compare answers from different AI models side-by-side. Experience the future of AI…
PaperBench (OpenAI, 2025)
Benchmark evaluating AI agents' ability to replicate 20 ICML 2024 Spotlight/Oral papers from scratch, with 8,316…
Petastorm
Enables single machine or distributed training and evaluation of deep learning models.
Pharmasimtext-Os-Llms
This is a repository including the benchmark and agents included in an under review submission to JEDM 2025. -…
PPTAgent
Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)
PredictionIO
Event collection, deployment of algorithms, evaluation, querying predictive results via APIs.
Price Per Token
Compare LLM API pricing across 200+ models from OpenAI, Anthropic, Google, and more. Includes token counters, cost…
PromptBench
Microsoft's adversarial robustness evaluation suite. `opensource` `free`
ProteinGym
Large-scale benchmark suite for protein fitness prediction and design, aggregating 200+ deep mutational scanning…
@PubliusAu
@PubliusAu supports evaluating, comparing, or benchmarking AI models.
Quantifying Infrastructure Noise in Agentic Coding Evals
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI…
Qwen3-Embedding (Alibaba)
2K★. 1 on MTEB multilingual leaderboard. Sizes: 0.6B/4B/8B. 32K context, MRL support, instruction-aware. Includes…
Rawbot
Discover Rawbot, the ultimate AI comparison tool. Boost your research, development, or business with the ideal AI…
Real or Fake Text
Real or Fake Text Real or Fake Text? Play Help About Leaderboard Log In Save Progress How good are you at knowing…
RepoRanger
RepoRanger is a llm benchmarks & leaderboards tool. AI-powered Github leaderboard for ranking users
RFxAI
RFxAI is the AI-native Unified Deal Lifecycle platform for Buyers and Sellers. Automate RFP discovery, bid response,…
Rhesis
Testing infrastructure for LLM and agentic applications with collaborative evaluation.
RL Baselines3 Zoo
A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization and…
SciBench
SciBench: Evaluating Math Reasoning in Visual Contexts --> --> More Research Chameleon ScienceQA LLaMA-Adapter (V2)…
SEAL LLM Leaderboard
SEAL LLM Leaderboard is a llm benchmarks & leaderboards tool. Expert-driven LLM benchmarks and updated AI model…
Shortcutsbench
ShortcutsBench: A Large-Scale Real-World Benchmark for API-Based Agents - EachSheep/ShortcutsBench
Sim-Court
BenCourt: A Benchmark and Framework for Court Simulation using LLM-based Agents - Miracle-2001/Sim-Court
Smartplay
SmartPlay is a benchmark for Large Language Models (LLMs). Uses a variety of games to test various important LLM…
Speechllm
This repository contains the training, inference, evaluation code for SpeechLLM models and details about the model…
Sphnx
SPHNX is a modular benchmark suite designed to evaluate and enhance the privacy management capabilities of Large…
SpotRank
AI Visibility Rank helps brands track how they appear in AI-generated answers. Discover what prompts mention your…
Startup Spotlight
Handpicked list of the best micro-startups. Curated by humans & updated weekly.
Starwhale
An MLOps/LLMOps platform for model building, evaluation, and fine-tuning.
STATE-Bench
Benchmark AI Agents on Enterprise Workflows. Contribute to microsoft/STATE-Bench development by creating an account…
Strat.Chat
Strat.Chat is a llm benchmarks & leaderboards tool. Business strategy development and evaluation support.
Stream-Bench
We propose a pioneering benchmark to evaluate LLM agents' ability to improve over time in streaming scenarios -…
SuperBench
a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing…
SuperLim
SUPERLIM LEADERBOARD ABOUT RESULTS TASKS DOCS SUBMIT Tasks [ explanation ]: ABSAbank-Imm - Argumentation sentences -…
Surge AI
Surge AI is a human-in-the-loop data labeling and model evaluation platform. It provides high-quality…
SWE-bench
SWE-bench Leaderboards SWE-bench SWE-bench Leaderboards Benchmarks SWE-bench SWE-bench Verified SWE-bench…
SWE-rebench
SWE-rebench: A Continuously Evolving and Decontaminated Benchmark for Software Engineering LLMs
T2T-Polish
T2T-Polish is a llm benchmarks & leaderboards tool. Evaluation and polishing workflows for T2T genome assemblies
TAT-QA
TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
TencentDB-Agent-Memory
TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four…
**Terminal-Bench**
A benchmark for LLMs on complicated tasks in the terminal - harbor-framework/terminal-bench-1
Terminal-Bench
Terminal-Bench is a llm benchmarks & leaderboards tool. Benchmark for agents performing real tasks inside a terminal.
TextFlint
(from Fudan) - A unified multilingual robustness evaluation toolkit for NLP.
The Interview
The Interview is a llm benchmarks & leaderboards tool. Evaluation and comparison of job candidates.
Theagentcompany
An agent benchmark with tasks in a simulated software company. - TheAgentCompany/TheAgentCompany
TheFastest.ai
Benchmarks for the fastest AI models
Time Series Library (TSLib)
A Library for Advanced Deep Time Series Models for General Time Series Analysis. - thuml/Time-Series-Library
TimeWarp
TimeWarp is a llm benchmarks & leaderboards tool. A benchmark on historical versions of web UI.
TLDL
AI learning shortcuts: podcast summaries, verified LLM API pricing data, AI tools, and company directories. Get…
topin.tech
Discover top talent with AI-powered assessments & interviews. Equip students for placements. Trusted by 600+…
tune
tune is a llm benchmarks & leaderboards tool. A benchmark for comparing Transformer-based models.
Usability-Benchmarking-Framework-Project
Evaluation of Software Manuals Using LLM-Powered GUI Agents: A Usability Benchmarking Framework -…
User Evaluation
Run user research with an AI agent: recruit real participants, hold AI-moderated interviews in 40 languages, and get…
Vancit
Vancit is a llm benchmarks & leaderboards tool. AI-powered active sourcing and code evaluation for developer hiring.
VarosAI
Varos delivers the work of business analysts 10x faster, better & cheaper with AI agents built for discovery. Our…
Visualwebarena
VisualWebArena is a benchmark for multimodal agents. - web-arena-x/visualwebarena
Voice Lab
Testing and evaluation framework for voice agents - GitHub - saharmor/voice-lab: Testing and evaluation framework…
Voxprobe
A python package for automated testing and evaluation for voice ai agents
We-Math
Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
WeatherBench2
Next-generation benchmark for data-driven global weather models with standardized evaluation framework and curated…
WebCanvas
All-in-one Web Agent framework for post-training. Start building with a few clicks! - iMeanAI/WebCanvas
Weblinx
WebLINX is a benchmark for building web navigation agents with conversational capabilities - McGill-NLP/weblinx
WHOOPS!
WHOOPS! Benchmark --> Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional…
Windowsagentarena
Windows Agent Arena (WAA) is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
Xobin AI Evaluate
Xobin provides an automated artificial intelligence (AI) scoring system to rate the accurate score for interview…
Yet-Another-Applied-Llm-Benchmark
A benchmark to evaluate language models on questions I've previously asked them to solve. -…
zoomeye
ZoomEye is a freemium online tool aimed to help aid cybersecurity in the areas of reconnaissance and threat…