Arena icon

Arena

Claim

Arena is a public AI model comparison and ranking platform for chatting, voting on outputs, and exploring task-specific leaderboards across text, vision, code, search, video, and agent tasks.

Arena

Overview

Arena is a public AI model comparison and ranking platform. It lets people chat with frontier models, compare responses, vote on the better answer, and help shape community leaderboards for large language models, image models, code models, and other task-specific arenas.

The site presents itself as an official leaderboard for benchmarking models across real-world evaluation settings. Dedicated leaderboard pages cover text, vision, document understanding, image generation and editing, web development, search, text-to-video, image-to-video, video editing, and agent tasks.

Features

Side-by-side model comparison

Users can chat with models and compare responses side by side, then vote for the answer they think is better. This turns everyday prompts into a public evaluation workflow.

Multi-modal leaderboards

Arena maintains leaderboard views across multiple arenas, including text, vision, document, image generation, image editing, web development, search, and video tasks. Each view shows ranked models and score spreads.

Agent performance ranking

The Agent Arena page ranks models on real-world agentic tasks using signals such as tool reliability, task completion, and steerability. That makes it useful for judging orchestration, not just conversational quality.

Chat history search

The site includes a searchable chat history area for past conversations, with tabs for Agent, Battles, Search, Code, Image, Video, and Archived chats. This supports revisiting prior evaluations and examples.

Drill-down leaderboard views

Leaderboard pages expose high-level snapshots and links to deeper views, so users can start with a quick scan and drill into a specific arena when needed.

Use Cases

  • Benchmark chat models

    Compare responses from multiple frontier models to decide which output is most useful for a specific prompt. This is the core workflow for people evaluating general-purpose chat quality.

  • Choose models by task

    Review task-specific rankings when selecting a model for vision, document understanding, search, web development, or video generation work. The dedicated arenas make it easier to compare like with like.

  • Evaluate agent workflows

    Use the Agent Arena leaderboard to inspect how models perform when they need to orchestrate tools and complete multi-step work. The ranking emphasizes reliability, completion, and steerability.

  • Revisit saved chats

    Search past conversations and battles to revisit examples, compare previous prompts, or keep track of prior evaluations across different categories.

Pros and Cons

Pros

  • Combines chat, comparison, voting, and leaderboards in one product.
  • Covers a wide range of model categories and task types, not only text chat.
  • Provides task-specific ranking pages, including a dedicated agent leaderboard.
  • Includes searchable chat history for revisiting prior interactions and battles.

Cons

  • The pricing page is not currently available, so the public site does not clearly communicate plan details.
  • The homepage includes a privacy warning stating that conversations and certain personal information may be disclosed publicly or to model providers.

FAQ

What is Arena used for?

Arena lets you start a conversation, compare model responses, and vote on which answer is better. The site also provides search across saved chats and specialized leaderboards for agent, text, image, and other model categories.

Is pricing available on the site?

The public site shows leaderboard pages and chat interfaces, but the pricing page at `/pricing` currently returns a 404. Based on the available sources, pricing details are not published there.

Are chats private?

The rendered homepage warns that inputs are processed by third-party AI and that conversations and certain personal information may be disclosed to AI providers and may otherwise be disclosed publicly. Users are told not to submit sensitive information they would not want shared publicly.

What kinds of models or tasks does Arena rank?

The leaderboard pages are organized by task type, including text, web development, vision, document, text-to-image, image edit, image-to-webdev, search, text-to-video, image-to-video, and video edit. The agent leaderboard highlights performance for tool use, task completion, reliability, and steerability.

Quick Facts

Category
AI model leaderboard
Primary use
Compare, rank, and vote on AI model outputs
Source domain
arena.ai
Task coverage
Text, image, vision, document, search, web dev, video, and agents
Pricing page
Public pricing details were not available; `/pricing` returned 404
Privacy note
User inputs may be disclosed to AI providers and may be public