Category directory

LLM Benchmarks & Leaderboards

Showing 96 of 192 LLM Benchmarks & Leaderboards tools on page 1 of 2.

All tools29,3883D & Game Assets178AI Agents & Automation1,641AI Chatbots & Assistants2,670AI Consulting & Strategy17AI Detection & Humanization89AI Infrastructure & MLOps338AI Model Interpretability12AI Research Resources4AI Safety & Alignment47AI Search Engines345API Documentation & Testing7Accounting & Bookkeeping62Ad Creative & Campaigns138Agent Frameworks & Orchestration175Animation & Motion Graphics20App Builders124Audio Editing & Cleanup15Audio Transcription111Automotive Tools2Blog & Article Writing65Browser Automation Agents106Browser Extensions210Business Planning12CRM & Relationship Management12Chatbot Builders44Code Review & Quality203Coding Assistants1,790Communities & Social Platforms109Crypto & Web349Customer Support336Data Analysis & BI502Data Extraction & Scraping129Dating & Relationships76Debugging & Error Fixing13Desktop Apps36DevOps & Cloud282Developer Documentation67Digital Humans & Avatars6Directories & Discovery222Ecommerce Tools173Education & Learning662Email & Inbox Productivity51Email Writing148Entertainment & Fun156Event Management2Fashion & Beauty152Finance & Investing454Fitness & Wellness142Food & Cooking94Forms & Surveys179Gaming Tools134Government & Public Data38Graphic Design50Healthcare & Medical266Identity & Fraud Detection38Image Editing200Image Generation2,783Image Recognition15Interior & Architecture Design79Job Search & Applications89Knowledge Bases & Q&A233LLM APIs & Gateways240LLM Benchmarks & Leaderboards192Language Learning97Legal & Contracts179Logo & Brand Design142Maps & Geospatial89Marketing Analytics32Meeting Assistants205Memory & Personalization74Mind Mapping & Diagrams59Mobile Apps29Model Libraries & Repositories71Model Training & Fine-Tuning143Music Generation839News & Media Monitoring112No-Code Automation131Notes & Knowledge Management168OCR & Document Processing56OSINT & Investigation70PDF & Document Chat216Paraphrasing & Rewriting9Personal Finance71Pet Care5Photo Enhancement176Podcasting Tools83Presentation Tools111Privacy & Compliance141Productivity Assistants361Project Management199Prompt Engineering232Real Estate & Property113Recruiting & HR508Research & Literature Review579Robotics & Hardware52SEO Tools636SQL & Database Tools171Sales Enablement401Scheduling & Calendar77Science & Engineering125Screen Recording & Demos31Security Scanning311Social Media Tools790Sports Analytics3Sports Prediction Tools4Spreadsheets & Data Entry122Summarization Tools108Templates & Generators467Testing & QA Automation49Text-to-Speech & Voice147Translation & Localization174Travel & Hospitality105UI/UX Design64Video Editing93Video Generation1,316Video Translation & Dubbing23Website Builders200Workflow Documentation104Writing Assistants1,681
Page 1 of 2

LLM Benchmarks & Leaderboards results

192 total
ALLM Benchmarks & Leaderboards

ACLUE

Official github repo for ACLUE, an evaluation benchmark focused on ancient Chinese language comprehension -…

ALLM Benchmarks & Leaderboards

Agent-Evaluation

A generative AI-powered framework for testing virtual agents. - awslabs/agent-evaluation

ALLM Benchmarks & Leaderboards

Agent Evaluation Framework 2026: Metrics, Rubrics & Benchmarks

Agent Evaluation Framework 2026: Metrics, Rubrics & Benchmarks | Galileo

ALLM Benchmarks & Leaderboards

Agent Evaluation Readiness Checklist

A practical checklist for agent evaluation: error analysis, dataset construction, grader design, offline & online…

ALLM Benchmarks & Leaderboards

AhaApple

AhaApple. AI Idea Generator. one click, many useful ideas.

ALLM Benchmarks & Leaderboards

Ai-Ethics-Evaluation-Report-In-Healthcare

Ai-Ethics-Evaluation-Report-In-Healthcare is a llm benchmarks & leaderboards tool. An evaluation report in healthcare

ALLM Benchmarks & Leaderboards

AI-Infra-Guard (Tencent)

A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP…

ALLM Benchmarks & Leaderboards

Ai-Llm-Comparison

A website where you can compare every AI Model ✨.

ALLM Benchmarks & Leaderboards

AI Meme Arena

Add more credibility to your site - get a premium domain today. Straight-forward shopping experience.

ALLM Benchmarks & Leaderboards

AI Tools Arena

Explore our comprehensive AI tools list, showcasing the best solutions for diverse industries. Improve your workflow…

ALLM Benchmarks & Leaderboards

Aider Polyglot Leaderboard

Aider's leaderboard ranking models on multi-language code editing tasks.

ALLM Benchmarks & Leaderboards

AlpacaEval

An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.

ALLM Benchmarks & Leaderboards

Alpha Arena

Alpha Arena is a llm benchmarks & leaderboards tool. Live trading performance benchmark for AI models in real markets.

ALLM Benchmarks & Leaderboards

Appworld-Leaderboard

Leaderboard Repository for "AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding…

ALLM Benchmarks & Leaderboards

Arena

An open platform for crowdsourced AI benchmarking, hosted by researchers at UC Berkeley SkyLab.

ALLM Benchmarks & Leaderboards

Arena Commerce AI

Unlock the power of Commerce AI to enhance customer experiences and support, streamline operations, and drive more…

ALLM Benchmarks & Leaderboards

Artificial Analysis

Artificial Analysis is a platform that provides AI model and service provider comparisons and benchmarks to help…

ALLM Benchmarks & Leaderboards

Athina AI

Athina Flows Develop Observe Pricing Docs Log in Sign up Ship AI to prod 10x faster Athina is a collaborative AI…

ALLM Benchmarks & Leaderboards

Auto-evaluator

: a lightweight evaluation tool for question-answering using Langchain !

ALLM Benchmarks & Leaderboards

AutoRecruiter

India's AI-native recruitment platform. Hire in 7 days with a transparent flat success fee, far less than…

aLLM Benchmarks & Leaderboards

awesome-ai-agent-papers

Curated 2025–2026 papers on agent engineering, memory, eval, and workflows

BLLM Benchmarks & Leaderboards

Bananalyzer

Open source AI Agent evaluation framework for web tasks - reworkd/bananalyzer

BLLM Benchmarks & Leaderboards

Baseline-Agent

Baseline-Agent is a llm benchmarks & leaderboards tool. This is a simple AI Agent used to test the Bloodrock CORE…

BLLM Benchmarks & Leaderboards

BeHonest

BeHonest: Benchmarking Honesty in Large Language Models More Research Abel MathPile ReAlign BenBench OlympicArena…

BLLM Benchmarks & Leaderboards

BenchGecko

The data layer of the AI economy. AI model benchmark leaderboard with cross-provider pricing comparison across…

BLLM Benchmarks & Leaderboards

BenchLM.ai

Compare 216 ranked models and 385 tracked AI models across 414 benchmarks with BenchLM scoring, pricing, context…

BLLM Benchmarks & Leaderboards

Benchmark Email

Ditch the clunky tools. Benchmark is the powerfully simple email marketing software that helps you send campaigns in…

BLLM Benchmarks & Leaderboards

Berkeley Function-Calling Leaderboard

Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM

@LLM Benchmarks & Leaderboards

@bmdhodl

Building a one-person AI-operated holding company. AgentGuard (pip install agentguard47), showwork, measured…

bLLM Benchmarks & Leaderboards

braindecode

Deep learning software to decode EEG, ECG or MEG signals, providing standardized neural network models,…

BLLM Benchmarks & Leaderboards

BuildArena

First physics-aligned interactive benchmark for LLM agents in engineering construction, designing…

bLLM Benchmarks & Leaderboards

buyer-eval-skill

B2B software vendor evaluation skill for Claude Code — domain-expert questions, vendor AI agent conversations,…

CLLM Benchmarks & Leaderboards

Chat-Agent-Evalution

Evaluating the LLM Chat Agent on multiple evaluation benchmarks. - khuzaimakt/Chat-Agent-Evalution

CLLM Benchmarks & Leaderboards

CheckMyIdea

Validate your side business idea in minutes with our AI-powered evaluation service. Maximize your chances of success…

CLLM Benchmarks & Leaderboards

Chinese Large Model Leaderboard

Chinese Large Model Leaderboard is a llm benchmarks & leaderboards tool. an expert-driven benchmark for Chineses LLMs.

CLLM Benchmarks & Leaderboards

Claw-Eval

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans. - claw-eval/claw-eval

CLLM Benchmarks & Leaderboards

Cognitive-Security-Ai-Powered-Threat-Agent-Evaluation-For-Impact-On-Assets.

Cognitive-Security-Ai-Powered-Threat-Agent-Evaluation-For-Impact-On-Assets. is an AI tool for llm benchmarks &…

CLLM Benchmarks & Leaderboards

CompMix

CompMix: A Benchmark for Heterogeneous Question Answering.

CLLM Benchmarks & Leaderboards

Confident AI

Confident AI is the AI quality platform for enterprise teams to standardize AI evals and observability across the…

CLLM Benchmarks & Leaderboards

Countless

Compare AI models easily! All providers in one place. Find the best LLM for your needs with our comprehensive…

CLLM Benchmarks & Leaderboards

CRO Benchmark

Get a full AI CRO audit across 248 CRO best practices. Built for ecommerce teams. We surface the leaks costing you…

CLLM Benchmarks & Leaderboards

Cross-Model-Evaluation-Judging-Ai-Ethics-And-Alignment-Responses-With-Language-Models

This study aims to evaluate the quality of previously generated responses using various large language models (LLMs)…

DLLM Benchmarks & Leaderboards

DeepEval

DeepEval is the open-source LLM evaluation framework for testing and benchmarking LLM applications — 50+…

DLLM Benchmarks & Leaderboards

Demystifying Evals for AI Agents

Demystifying evals for AI agents \ Anthropic Skip to main content Skip to footer Research Policy Commitments Learn…

DLLM Benchmarks & Leaderboards

Design Arena

Design Arena is the largest global crowdsourced benchmark for design. Challenge, Vote, Crown your Winner.

DLLM Benchmarks & Leaderboards

Designing AI-Resistant Technical Evaluations

What we learned from three iterations of a performance engineering take-home that Claude keeps beating.

DLLM Benchmarks & Leaderboards

Dingo

Dingo is a llm benchmarks & leaderboards tool. Dingo - A Comprehensive Data Quality Evaluation Tool

DLLM Benchmarks & Leaderboards

Diplomacy-Llm

Public LLM benchmark using the results of Diplomacy games played by multiple LLM agents. - lukepoo101/diplomacy-llm

DLLM Benchmarks & Leaderboards

DreamBench++

DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

DLLM Benchmarks & Leaderboards

Driving the Agent Quality Flywheel from Your Coding Agent

Google's June 2026 account of automating the eval-optimize loop for coding agents: independent AutoRaters grade…

DLLM Benchmarks & Leaderboards

Dubesor LLM Benchmark table

Dubesor LLM Benchmark table - Small-scale manual LLM performance comparison benchmark

ELLM Benchmarks & Leaderboards

Edge Arena

Edge Arena puts your business decisions on trial — competing AI agents challenge assumptions, expose weaknesses, and…

ELLM Benchmarks & Leaderboards

Eval

Eval is a llm benchmarks & leaderboards tool. Better coding workflow with smart assistance.

ELLM Benchmarks & Leaderboards

Eval Awareness in Claude Opus 4.6's BrowseComp Performance

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI…

ELLM Benchmarks & Leaderboards

Eval-Driven Development: Build and Evaluate Reliable AI Agents

Learn how to build reliable AI agents with our 8-stage evaluation framework. We explore DeepEval, multi-turn…

FLLM Benchmarks & Leaderboards

FaceRate.ai

FaceRate.ai offers a face attractiveness test, facial analysis, and golden ratio face tests. Get an in-depth…

FLLM Benchmarks & Leaderboards

FELM

FELM: Benchmarking Factuality Evaluation of Large Language Models

FLLM Benchmarks & Leaderboards

Fl_Llm_Benchmark_Dataset

This code collects congressional/parliamentary dataset across US, UK and Canada

FLLM Benchmarks & Leaderboards

Fogworkflowsim

An Environment for Simulation and Performance Evaluation of Workflows in Fog Computing

FLLM Benchmarks & Leaderboards

Foundry

Produce high-quality enterprise data, evaluate reliably, and optimize performance at scale—without web drift, IP…

GLLM Benchmarks & Leaderboards

Giskard

Giskard is a llm benchmarks & leaderboards tool. Testing & evaluation library for LLM applications, in particular RAGs

GLLM Benchmarks & Leaderboards

Go/NoGo

Evaluate RFPs instantly and get AI-powered insights to make better Go/No-Go decisions.

GLLM Benchmarks & Leaderboards

Goodai-Ltm-Benchmark

A library for benchmarking the Long Term Memory and Continual learning capabilities of LLM based agents. With all…

gLLM Benchmarks & Leaderboards

gpt-oss playground

Demo platform for OpenAI's open-weight models for developers.

GLLM Benchmarks & Leaderboards

GradeWrite

GradeWrite.AI is the ultimate AI assistant for grading assignments. GradeWrite streamlines grading process with…

GLLM Benchmarks & Leaderboards

GSM Arena

GSMArena Turnstile check One quick check before you continue... Continue

HLLM Benchmarks & Leaderboards

HeHealth

HeHealth is a llm benchmarks & leaderboards tool. Reliable sexual health evaluation

HLLM Benchmarks & Leaderboards

HELM

HELM is a llm benchmarks & leaderboards tool. Holistic evaluation across 42 scenarios

HLLM Benchmarks & Leaderboards

Hermes 3 Llama 3.1 405B

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic…

HLLM Benchmarks & Leaderboards

HumanEval

HumanEval is a llm benchmarks & leaderboards tool. Python code generation correctness

HLLM Benchmarks & Leaderboards

Hume AI

Real human ratings, in a single API call. The human evaluation layer for voice, speech, and conversational AI.

HLLM Benchmarks & Leaderboards

HypeBridge

AI-powered influencer evaluation and discovery platform. Make data-driven partnership decisions in seconds. Join…

ILLM Benchmarks & Leaderboards

IdolCrush.ai

IdolCrush.ai — a new interactive AI idol experience where users become producers, create and chat with your dream…

iLLM Benchmarks & Leaderboards

imgsys

A generative AI arena where you can test different prompts and pick the results you like the most. Check-out the…

ILLM Benchmarks & Leaderboards

Impact-Academy

Auto-Enhance meta-benchmark, to measure the ability of LLM agents to improve other LLM agents -…

ILLM Benchmarks & Leaderboards

InfiBench

a benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world…

iLLM Benchmarks & Leaderboards

instruct-eval

This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on…

ILLM Benchmarks & Leaderboards

IntelliServer

simplifies the evaluation of LLMs by providing a unified microservice to access and test multiple AI models.

ILLM Benchmarks & Leaderboards

Internintelligence_Ai_Ethics_And_Bias_Evaluation

This project evaluates the fairness of a machine learning model trained on the Adult Income Dataset to predict…

ILLM Benchmarks & Leaderboards

Interview: About deployment, evaluation, and testing of agents with Sully Omar, the CEO of Cognosys AI

We asked the founder of Cognosys, Sully Omar, about his experience with building a product for no-code users in the…

ILLM Benchmarks & Leaderboards

Introduction to Generative AI | SqillPlan

: introduction to Generative AI, including models such as GANs, Variational Autoencoders, Autoregressive Models, and…

JLLM Benchmarks & Leaderboards

JARVIS

NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT),…

JLLM Benchmarks & Leaderboards

Jury

Jury helps with llm benchmarks & leaderboards workflows.

JLLM Benchmarks & Leaderboards

Just-Eval

A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.

LLLM Benchmarks & Leaderboards

LangSmith

Complete AI agent and LLM observability platform with tracing and real-time monitoring. Debug agents, find failures…

LLLM Benchmarks & Leaderboards

LawBench

a benchmark designed to evaluate large language models in the legal domain.

LLLM Benchmarks & Leaderboards

Lawful-Good

Benchmark for assessing legal capabilities of LLM agents - dluo96/lawful-good

LLLM Benchmarks & Leaderboards

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark of LLMs for code that continuously collects…

LLLM Benchmarks & Leaderboards

LiveCodeBench

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

LLLM Benchmarks & Leaderboards

Llama 3.2 3B Instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language…

LLLM Benchmarks & Leaderboards

Llama 3.3 70B

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in…

LLLM Benchmarks & Leaderboards

Llemma

Open language model for mathematics (7B/34B) trained on Proof-Pile-2, outperforming Minerva at equal scale on MATH…

LLLM Benchmarks & Leaderboards

Llf-Bench

A benchmark for evaluating learning agents based on just language feedback - microsoft/LLF-Bench

LLLM Benchmarks & Leaderboards

Llm-Agent-Ask-For-Help

Benchmark LLM Agents' abilities to quit sequential tasks as early as possible. - dillonmsandhu/llm-agent-ask-for-help

LLLM Benchmarks & Leaderboards

Llm-Agent-Benchmark-List

A banchmark list for evaluation of large language models. - zhangxjohn/LLM-Agent-Benchmark-List

LLLM Benchmarks & Leaderboards

Llm-Benchmarks

Llm-Benchmarks is a llm benchmarks & leaderboards tool. LLM benchmark tools for LMDeploy, vLLM, and TensorRT-LLM.