Category directory

LLM Benchmarks & Leaderboards

Showing 96 of 192 LLM Benchmarks & Leaderboards tools on page 2 of 2.

All tools29,3883D & Game Assets178AI Agents & Automation1,641AI Chatbots & Assistants2,670AI Consulting & Strategy17AI Detection & Humanization89AI Infrastructure & MLOps338AI Model Interpretability12AI Research Resources4AI Safety & Alignment47AI Search Engines345API Documentation & Testing7Accounting & Bookkeeping62Ad Creative & Campaigns138Agent Frameworks & Orchestration175Animation & Motion Graphics20App Builders124Audio Editing & Cleanup15Audio Transcription111Automotive Tools2Blog & Article Writing65Browser Automation Agents106Browser Extensions210Business Planning12CRM & Relationship Management12Chatbot Builders44Code Review & Quality203Coding Assistants1,790Communities & Social Platforms109Crypto & Web349Customer Support336Data Analysis & BI502Data Extraction & Scraping129Dating & Relationships76Debugging & Error Fixing13Desktop Apps36DevOps & Cloud282Developer Documentation67Digital Humans & Avatars6Directories & Discovery222Ecommerce Tools173Education & Learning662Email & Inbox Productivity51Email Writing148Entertainment & Fun156Event Management2Fashion & Beauty152Finance & Investing454Fitness & Wellness142Food & Cooking94Forms & Surveys179Gaming Tools134Government & Public Data38Graphic Design50Healthcare & Medical266Identity & Fraud Detection38Image Editing200Image Generation2,783Image Recognition15Interior & Architecture Design79Job Search & Applications89Knowledge Bases & Q&A233LLM APIs & Gateways240LLM Benchmarks & Leaderboards192Language Learning97Legal & Contracts179Logo & Brand Design142Maps & Geospatial89Marketing Analytics32Meeting Assistants205Memory & Personalization74Mind Mapping & Diagrams59Mobile Apps29Model Libraries & Repositories71Model Training & Fine-Tuning143Music Generation839News & Media Monitoring112No-Code Automation131Notes & Knowledge Management168OCR & Document Processing56OSINT & Investigation70PDF & Document Chat216Paraphrasing & Rewriting9Personal Finance71Pet Care5Photo Enhancement176Podcasting Tools83Presentation Tools111Privacy & Compliance141Productivity Assistants361Project Management199Prompt Engineering232Real Estate & Property113Recruiting & HR508Research & Literature Review579Robotics & Hardware52SEO Tools636SQL & Database Tools171Sales Enablement401Scheduling & Calendar77Science & Engineering125Screen Recording & Demos31Security Scanning311Social Media Tools790Sports Analytics3Sports Prediction Tools4Spreadsheets & Data Entry122Summarization Tools108Templates & Generators467Testing & QA Automation49Text-to-Speech & Voice147Translation & Localization174Travel & Hospitality105UI/UX Design64Video Editing93Video Generation1,316Video Translation & Dubbing23Website Builders200Workflow Documentation104Writing Assistants1,681
Page 2 of 2

LLM Benchmarks & Leaderboards results

192 total
LLLM Benchmarks & Leaderboards

LLM Evaluation

Home | LLM Evaluation Skip to main content Link Menu Expand (external link) Document Search Copy Copied Home Papers…

LLLM Benchmarks & Leaderboards

Llm_Evaluation_For_Gene_Set_Interpretation

Code space for 'Evaluation of large language models for discovery of gene set function' -…

LLLM Benchmarks & Leaderboards

Llm-Interest

Llm-Interest is a llm benchmarks & leaderboards tool. Tool to collect LLM eval topics

LLLM Benchmarks & Leaderboards

LLM Labs

Compare and test language models.

LLLM Benchmarks & Leaderboards

Llm_Model_Evaluation

LLM Model Evaluation for tmmluplus datasets. Contribute to LiuYuWei/llm_model_evaluation development by creating an…

LLLM Benchmarks & Leaderboards

LLM Price Check

Explore cost-effective LLM API solutions with LLM Price Check. Instantly compare updated prices from major providers…

LLLM Benchmarks & Leaderboards

Llm_Scavengerhunt

A New Benchmark made for tasks for Large Language Model Agents - UC Berkeley Scavenger Hunt -…

LLLM Benchmarks & Leaderboards

LLM Stats

LLM Stats, the most comprehensive LLM leaderboard, benchmarks and compares API models using daily‑updated,…

LLLM Benchmarks & Leaderboards

Llmagentoodgym

OOD benchmark study for LLM agents based on BrowserGym and AgentLab from ServiceNow. - rowingchenn/LLMAgentOODGym

LLLM Benchmarks & Leaderboards

Llmdataparser

LLMDataParser is a Python library that provides a collection of parsers for various benchmark datasets used in the…

LLLM Benchmarks & Leaderboards

Llmscenarioeval

Scenario-based Evaluation dataset for LLM (beta). Contribute to Turing-Project/LLMScenarioEval development by…

LLLM Benchmarks & Leaderboards

Lmsys Arena Hard

бенчмарк, основанный на сравнении качества ответов на реальные человеческие запросы. В роли судьи, правда, выступает…

LLLM Benchmarks & Leaderboards

LMSYS Chatbot Arena Leaderboard

LMSYS Chatbot Arena is a crowdsourced open platform for LLM evals. Collected over 1,000,000 human pairwise…

MLLM Benchmarks & Leaderboards

M3CoT

Leaderboard | M 3 CoT M 3 CoT Home Download Evaluation Annotation Paper Code Citation Contact Leaderboard Explore…

MLLM Benchmarks & Leaderboards

MatBench

MatBench is a llm benchmarks & leaderboards tool. Materials informatics benchmark

MLLM Benchmarks & Leaderboards

MathEval

MathEval是一个专注于全面评估大模型数学能力的测评基准。共包含22个数学领域测评集和近30K道数学题目,旨在全面评估大模型在包含算术,小初高竞赛和部分高等数学分支在内的各阶段、难度和数学子领域的解题能力表现,既可以作为现阶段大模…

MLLM Benchmarks & Leaderboards

Megaminer-Tinyarena

A tiny arena for testing Megaminer AI agents. Contribute to drusepth/Megaminer-Tinyarena development by creating an…

MLLM Benchmarks & Leaderboards

MiniWoB++

A collection of over 100 web interaction environments, along with JavaScript and Python interfaces.

MLLM Benchmarks & Leaderboards

MixEval

a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a…

MLLM Benchmarks & Leaderboards

MixEval

A reliable click-and-go evaluation suite compatible with both open-source and proprietary models, supporting MixEval…

MLLM Benchmarks & Leaderboards

ML6 x AISO Agent Workshop

ML6 x AISO Agent Workshop (March 2026). Contribute to ml6team/AISO-workshop development by creating an account on…

MLLM Benchmarks & Leaderboards

MMedBench

Medical Multilingual Benchmark Blog --> GitHub Huggingface MMedBench A Medical Benchmark for Multilingual…

MLLM Benchmarks & Leaderboards

MMToM-QA

MMToM-QA Leaderboard MMToM-QA: Multimodal Theory of Mind Question Answering ACL 2024 Outstanding Paper Award…

MLLM Benchmarks & Leaderboards

Mobilebench

Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents - XiaoMi/MobileBench

NLLM Benchmarks & Leaderboards

Non finito

Non finito is a llm benchmarks & leaderboards tool. Model evaluation and sharing made simple.

OLLM Benchmarks & Leaderboards

Ollama Grid Search

Desktop utility for systematic model evaluation. Test multiple models, prompts, and inference parameters…

OLLM Benchmarks & Leaderboards

OLMO-eval

OLMO-eval is a llm benchmarks & leaderboards tool. a repository for evaluating open language models.

OLLM Benchmarks & Leaderboards

OlympicArena

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

oLLM Benchmarks & Leaderboards

oobabooga benchmark

oobabooga benchmark oobabooga benchmark The list is sorted by size (on disk) for each score. Highlighted = Pareto…

OLLM Benchmarks & Leaderboards

Open-RAG-Eval

RAG evaluation without the need for "golden answers" - vectara/open-rag-eval

OLLM Benchmarks & Leaderboards

OverallGPT

OverallGPT lets you compare answers from different AI models side-by-side. Experience the future of AI…

PLLM Benchmarks & Leaderboards

PaperBench (OpenAI, 2025)

Benchmark evaluating AI agents' ability to replicate 20 ICML 2024 Spotlight/Oral papers from scratch, with 8,316…

PLLM Benchmarks & Leaderboards

Petastorm

Enables single machine or distributed training and evaluation of deep learning models.

PLLM Benchmarks & Leaderboards

Pharmasimtext-Os-Llms

This is a repository including the benchmark and agents included in an under review submission to JEDM 2025. -…

PLLM Benchmarks & Leaderboards

PPTAgent

Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)

PLLM Benchmarks & Leaderboards

PredictionIO

Event collection, deployment of algorithms, evaluation, querying predictive results via APIs.

PLLM Benchmarks & Leaderboards

Price Per Token

Compare LLM API pricing across 200+ models from OpenAI, Anthropic, Google, and more. Includes token counters, cost…

PLLM Benchmarks & Leaderboards

PromptBench

Microsoft's adversarial robustness evaluation suite. `opensource` `free`

PLLM Benchmarks & Leaderboards

ProteinGym

Large-scale benchmark suite for protein fitness prediction and design, aggregating 200+ deep mutational scanning…

@LLM Benchmarks & Leaderboards

@PubliusAu

@PubliusAu supports evaluating, comparing, or benchmarking AI models.

QLLM Benchmarks & Leaderboards

Quantifying Infrastructure Noise in Agentic Coding Evals

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI…

QLLM Benchmarks & Leaderboards

Qwen3-Embedding (Alibaba)

2K★. 1 on MTEB multilingual leaderboard. Sizes: 0.6B/4B/8B. 32K context, MRL support, instruction-aware. Includes…

RLLM Benchmarks & Leaderboards

Rawbot

Discover Rawbot, the ultimate AI comparison tool. Boost your research, development, or business with the ideal AI…

RLLM Benchmarks & Leaderboards

Real or Fake Text

Real or Fake Text Real or Fake Text? Play Help About Leaderboard Log In Save Progress How good are you at knowing…

RLLM Benchmarks & Leaderboards

RepoRanger

RepoRanger is a llm benchmarks & leaderboards tool. AI-powered Github leaderboard for ranking users

RLLM Benchmarks & Leaderboards

RFxAI

RFxAI is the AI-native Unified Deal Lifecycle platform for Buyers and Sellers. Automate RFP discovery, bid response,…

RLLM Benchmarks & Leaderboards

Rhesis

Testing infrastructure for LLM and agentic applications with collaborative evaluation.

RLLM Benchmarks & Leaderboards

RL Baselines3 Zoo

A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization and…

SLLM Benchmarks & Leaderboards

SciBench

SciBench: Evaluating Math Reasoning in Visual Contexts --> --> More Research Chameleon ScienceQA LLaMA-Adapter (V2)…

SLLM Benchmarks & Leaderboards

SEAL LLM Leaderboard

SEAL LLM Leaderboard is a llm benchmarks & leaderboards tool. Expert-driven LLM benchmarks and updated AI model…

SLLM Benchmarks & Leaderboards

Shortcutsbench

ShortcutsBench: A Large-Scale Real-World Benchmark for API-Based Agents - EachSheep/ShortcutsBench

SLLM Benchmarks & Leaderboards

Sim-Court

BenCourt: A Benchmark and Framework for Court Simulation using LLM-based Agents - Miracle-2001/Sim-Court

SLLM Benchmarks & Leaderboards

Smartplay

SmartPlay is a benchmark for Large Language Models (LLMs). Uses a variety of games to test various important LLM…

SLLM Benchmarks & Leaderboards

Speechllm

This repository contains the training, inference, evaluation code for SpeechLLM models and details about the model…

SLLM Benchmarks & Leaderboards

Sphnx

SPHNX is a modular benchmark suite designed to evaluate and enhance the privacy management capabilities of Large…

SLLM Benchmarks & Leaderboards

SpotRank

AI Visibility Rank helps brands track how they appear in AI-generated answers. Discover what prompts mention your…

SLLM Benchmarks & Leaderboards

Startup Spotlight

Handpicked list of the best micro-startups. Curated by humans & updated weekly.

SLLM Benchmarks & Leaderboards

Starwhale

An MLOps/LLMOps platform for model building, evaluation, and fine-tuning.

SLLM Benchmarks & Leaderboards

STATE-Bench

Benchmark AI Agents on Enterprise Workflows. Contribute to microsoft/STATE-Bench development by creating an account…

SLLM Benchmarks & Leaderboards

Strat.Chat

Strat.Chat is a llm benchmarks & leaderboards tool. Business strategy development and evaluation support.

SLLM Benchmarks & Leaderboards

Stream-Bench

We propose a pioneering benchmark to evaluate LLM agents' ability to improve over time in streaming scenarios -…

SLLM Benchmarks & Leaderboards

SuperBench

a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing…

SLLM Benchmarks & Leaderboards

SuperLim

SUPERLIM LEADERBOARD ABOUT RESULTS TASKS DOCS SUBMIT Tasks [ explanation ]: ABSAbank-Imm - Argumentation sentences -…

SLLM Benchmarks & Leaderboards

Surge AI

Surge AI is a human-in-the-loop data labeling and model evaluation platform. It provides high-quality…

SLLM Benchmarks & Leaderboards

SWE-bench

SWE-bench Leaderboards SWE-bench SWE-bench Leaderboards Benchmarks SWE-bench SWE-bench Verified SWE-bench…

SLLM Benchmarks & Leaderboards

SWE-rebench

SWE-rebench: A Continuously Evolving and Decontaminated Benchmark for Software Engineering LLMs

TLLM Benchmarks & Leaderboards

T2T-Polish

T2T-Polish is a llm benchmarks & leaderboards tool. Evaluation and polishing workflows for T2T genome assemblies

TLLM Benchmarks & Leaderboards

TAT-QA

TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance

TLLM Benchmarks & Leaderboards

TencentDB-Agent-Memory

TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four…

*LLM Benchmarks & Leaderboards

**Terminal-Bench**

A benchmark for LLMs on complicated tasks in the terminal - harbor-framework/terminal-bench-1

TLLM Benchmarks & Leaderboards

Terminal-Bench

Terminal-Bench is a llm benchmarks & leaderboards tool. Benchmark for agents performing real tasks inside a terminal.

TLLM Benchmarks & Leaderboards

TextFlint

(from Fudan) - A unified multilingual robustness evaluation toolkit for NLP.

TLLM Benchmarks & Leaderboards

The Interview

The Interview is a llm benchmarks & leaderboards tool. Evaluation and comparison of job candidates.

TLLM Benchmarks & Leaderboards

Theagentcompany

An agent benchmark with tasks in a simulated software company. - TheAgentCompany/TheAgentCompany

TLLM Benchmarks & Leaderboards

TheFastest.ai

Benchmarks for the fastest AI models

TLLM Benchmarks & Leaderboards

Time Series Library (TSLib)

A Library for Advanced Deep Time Series Models for General Time Series Analysis. - thuml/Time-Series-Library

TLLM Benchmarks & Leaderboards

TimeWarp

TimeWarp is a llm benchmarks & leaderboards tool. A benchmark on historical versions of web UI.

TLLM Benchmarks & Leaderboards

TLDL

AI learning shortcuts: podcast summaries, verified LLM API pricing data, AI tools, and company directories. Get…

tLLM Benchmarks & Leaderboards

topin.tech

Discover top talent with AI-powered assessments & interviews. Equip students for placements. Trusted by 600+…

tLLM Benchmarks & Leaderboards

tune

tune is a llm benchmarks & leaderboards tool. A benchmark for comparing Transformer-based models.

ULLM Benchmarks & Leaderboards

Usability-Benchmarking-Framework-Project

Evaluation of Software Manuals Using LLM-Powered GUI Agents: A Usability Benchmarking Framework -…

ULLM Benchmarks & Leaderboards

User Evaluation

Run user research with an AI agent: recruit real participants, hold AI-moderated interviews in 40 languages, and get…

VLLM Benchmarks & Leaderboards

Vancit

Vancit is a llm benchmarks & leaderboards tool. AI-powered active sourcing and code evaluation for developer hiring.

VLLM Benchmarks & Leaderboards

VarosAI

Varos delivers the work of business analysts 10x faster, better & cheaper with AI agents built for discovery. Our…

VLLM Benchmarks & Leaderboards

Visualwebarena

VisualWebArena is a benchmark for multimodal agents. - web-arena-x/visualwebarena

VLLM Benchmarks & Leaderboards

Voice Lab

Testing and evaluation framework for voice agents - GitHub - saharmor/voice-lab: Testing and evaluation framework…

VLLM Benchmarks & Leaderboards

Voxprobe

A python package for automated testing and evaluation for voice ai agents

WLLM Benchmarks & Leaderboards

We-Math

Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

WLLM Benchmarks & Leaderboards

WeatherBench2

Next-generation benchmark for data-driven global weather models with standardized evaluation framework and curated…

WLLM Benchmarks & Leaderboards

WebCanvas

All-in-one Web Agent framework for post-training. Start building with a few clicks! - iMeanAI/WebCanvas

WLLM Benchmarks & Leaderboards

Weblinx

WebLINX is a benchmark for building web navigation agents with conversational capabilities - McGill-NLP/weblinx

WLLM Benchmarks & Leaderboards

WHOOPS!

WHOOPS! Benchmark --> Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional…

WLLM Benchmarks & Leaderboards

Windowsagentarena

Windows Agent Arena (WAA) is a scalable OS platform for testing and benchmarking of multi-modal AI agents.

XLLM Benchmarks & Leaderboards

Xobin AI Evaluate

Xobin provides an automated artificial intelligence (AI) scoring system to rate the accurate score for interview…

YLLM Benchmarks & Leaderboards

Yet-Another-Applied-Llm-Benchmark

A benchmark to evaluate language models on questions I've previously asked them to solve. -…

zLLM Benchmarks & Leaderboards

zoomeye

ZoomEye is a freemium online tool aimed to help aid cybersecurity in the areas of reconnaissance and threat…