$0.01Starting price
0Popularity
LLM Leaderboard featured image

About LLM Leaderboard

LLM Leaderboard aggregates rankings and pricing for AI models across multiple dimensions such as coding, reasoning, math, knowledge, and instruction following. It provides separate leaderboards for different model types including chat, image generation, video generation, text-to-speech, speech-to-text, and embeddings. The platform tracks 1,255 models and 741 benchmarks, with data updated regularly. Users can view overall rankings, efficiency metrics, and provider reliability scores. Benchmark pages detail raw model orders, coverage, and measurement scope for evaluations like GPQA, MMLU-Pro, and SWE-Bench. The leaderboard separates capability scores from price and runtime performance to help users make informed model selection decisions.

Key features

  • Overall model rankings by capability
  • Pricing comparisons per million tokens
  • Runtime performance metrics by tokens per second
  • Detailed benchmark libraries (GPQA, MMLU-Pro, SWE-Bench)
  • Model directory filtered by vendor, type, modality and price
  • Provider reliability scores
  • Separate leaderboards for chat, image, video, audio and embeddings
  • Scoring and data methodology documentation

Use cases

  • Selecting the best AI model for coding tasks
  • Evaluating models for reasoning and mathematical problem-solving
  • Comparing pricing and performance for production deployment

Pros

  • Compares models across multiple capability dimensions
  • Provides separate leaderboards for different model types and use cases
  • Tracks pricing and runtime performance independently of capability scores
  • Includes detailed benchmark breakdowns and raw model orders
  • Aggregates data from 1,255 models and 741 benchmarks

Cons

  • No free tier or trial access mentioned
  • Limited to official vendor API prices; third-party offers excluded
  • Data as of 2026-09-20 may not reflect real-time changes

Frequently asked questions about LLM Leaderboard

What is LLM Leaderboard and what does it do?

LLM Leaderboard is a platform that aggregates rankings, pricing, and performance metrics for AI models across multiple dimensions such as coding, reasoning, math, and instruction following. It provides separate leaderboards for different model types including chat, image generation, video generation, text-to-speech, speech-to-text, and embeddings.

Who should use LLM Leaderboard?

The platform is designed for developers, researchers, and businesses looking to compare AI models based on capability, price, speed, and benchmark performance. It helps users select the most suitable model for specific tasks or use cases.

How does LLM Leaderboard rank models?

Models are ranked based on capability scores, pricing, runtime performance, and provider reliability. Benchmark pages provide raw model orders, coverage, and measurement scope for evaluations like GPQA, MMLU-Pro, and SWE-Bench.

Does LLM Leaderboard provide pricing information?

Yes, the platform tracks and displays official input and output pricing for models, allowing users to compare cost efficiency alongside capability and performance metrics.

Can I compare models across different modalities on LLM Leaderboard?

Yes, the platform offers separate leaderboards for various model types, including chat, image generation, video generation, text-to-speech, speech-to-text, and embeddings, enabling cross-modal comparisons.

How often is the data on LLM Leaderboard updated?

The platform updates its data regularly, with the last update timestamp visible on the leaderboard pages to ensure users have access to the most current information.

LLM Leaderboard compared

Reviews