What is LMSYS Chatbot Arena?
The Large Model Systems (LMSYS) Chatbot Arena is a free, crowdsourced platform where anyone can compare artificial intelligence (AI) models by voting on their responses to the same prompt without knowing which model wrote which answer.
Researchers at UC Berkeley's Sky Computing Lab built and launched their LMSYS Chatbot Arena in May 2023 to rank large language models based on human preferences rather than static test scores. The project later moved to the LMArena domain and now operates under Arena, though the original leaderboard remains publicly accessible.
The idea behind it is straightforward: Multiple-choice benchmarks tell you whether a model picked the right answer on a fixed test, but they don't tell you which response a person would actually prefer to read. Arena closes that gap by putting real conversations in front of real people and accumulating votes to provide a ranking.
What are Mercor’s APEX benchmarks?
Mercor's APEX benchmarks evaluate AI models on expert-graded tasks that mirror real professional work. Instead of asking which response people prefer, APEX measures whether a model can successfully complete the kinds of assignments professionals encounter in fields like law, medicine, finance, and software engineering.
Each benchmark is built around domain-specific tasks and evaluated by subject matter experts using defined scoring criteria. The goal is to measure accuracy, completeness, and practical usefulness in a professional context rather than general user preferences.
Because APEX focuses on real-world performance, organizations use it to compare models for professional workflows where correctness matters as much as, or more than, conversational quality.
How does Arena differ from Mercor’s APEX benchmarks?
Although both Arena and APEX evaluate AI models, they answer different questions.
Arena measures which responses people prefer in blind comparisons, while Mercor’s APEX benchmarks measure whether a model can successfully complete expert-graded professional tasks. Organizations often use both benchmarks together because each reveals a different aspect of model performance.
That distinction is apparent in almost every part of how each one operates:
- Voters versus evaluators: Arena relies on anonymous, self-selected members of the public. APEX uses domain experts in fields such as law, medicine, finance, and software engineering to grade task output against defined criteria.
- Preference versus correctness: Arena's votes reflect which answers people liked reading. APEX scores whether answers are correct, complete, and usable in a professional context.
- General versus domain-specific prompts: Arena's prompt pool is broad and largely casual. APEX tasks are built to mirror the actual work a professional would be asked to do.
- Open community versus structured evaluation: LMSYS operates Arena as an open, community-run project with public data releases. APEX is a structured evaluation pipeline built around expert judgment on defined tasks.
Neither benchmark replaces the other. Arena is valuable for understanding human preference across general conversations, while APEX evaluates performance on professional tasks where correctness matters. Together, they provide a more complete picture of AI model quality.
Readers who want to better understand evaluation dimensions beyond simple preferences can look at how AI evaluation works for a broader idea of what rigorous testing involves.
How Arena ranks models vs. Mercor
Chatbot Arena ranks AI models by collecting millions of blind, head-to-head comparisons between pairs of models. Those votes are analyzed using the Bradley-Terry statistical model to estimate each model's relative strength, and the results are displayed as Elo-style ratings on the public leaderboard.
Here's how the process works:
- Battle: A user submits or selects a prompt, which Chatbot Arena sends to two randomly selected AI models. Both responses are shown anonymously so users judge the quality of the answers rather than the brand behind the model.
- Vote: The user selects the better response, declares a tie, or marks both responses as poor. The models' identities are revealed only after the vote is submitted, helping reduce brand bias.
- Rank (Bradley-Terry scoring): Every vote becomes a data point in the Bradley-Terry model, a statistical method designed for pairwise comparisons. Rather than assigning a fixed score, it estimates how likely one model is to outperform another based on millions of voting outcomes.
- Display (Elo ratings): The Bradley-Terry estimates are presented as Elo-style ratings on the leaderboard. Similar to chess rankings, defeating a highly rated model increases a model's rating more than defeating a lower-ranked one, allowing the leaderboard to continuously update as new comparison results are collected.
Mercor's APEX benchmarks generate results in a different way. Instead of ranking models through public preference voting, APEX evaluates model performance on structured professional tasks that are designed to reflect real-world work.
To determine the APEX benchmarks, subject matter experts assess each model's output with standardized evaluation criteria. The process produces benchmark scores based on task performance as opposed to user preference.
The APEX benchmark family includes 3 evaluations:
- APEX: This benchmark measures how well frontier models complete single-turn professional tasks across domains such as law, medicine, finance, and consulting.
- APEX-Agents: Users can evaluate AI agents on long-horizon, multi-step professional work that requires planning, reasoning, and the use of real software tools.
- APEX-SWE: Decision makers can use this benchmark to measure AI performance on real-world software engineering work, including integration and observability tasks.
Rather than asking which response people prefer, Mercor's benchmarks measure whether a model can successfully perform the work organizations expect it to complete.
Together, Arena's preference-based rankings and APEX's expert-graded benchmarks provide a more complete picture of AI model performance for deployment decisions.
Learn more about Mercor's APEX benchmarks.
Limitations of LMSYS Chatbot Arena
The leaderboard measures what people prefer, not what a model can actually accomplish, and that distinction has consequences. A handful of factors can influence the results in ways that don't track raw capability:
- Style bias: Longer, more confident-sounding, or more formatted responses tend to win votes even when they aren't more accurate, a pattern that TechCrunch's coverage of Arena's shortcomings flagged directly.
- Self-selected voters: Anyone can vote, which means the pool skews toward people who seek out the site rather than a representative sample of professional use cases.
- Prompt mix: Casual, general-purpose prompts dominate voting traffic, so a model's rank can reveal more about chat performance than about specialized or high-stakes tasks.
- Commercial incentives and data sharing: Labs can tune models with leaderboard performance in mind, and LMSYS has been transparent about sharing aggregated preference data and prompts publicly, which is good for research but means the incentive structure isn't invisible.
None of these limitations make Arena ineffective. Instead, they illustrate why organizations often combine preference-based leaderboards with task-based benchmarks such as APEX when evaluating AI models.
When to use Arena vs. APEX
Arena and APEX answer different questions about AI performance. Therefore, many organizations use them at different stages of the model selection process.
Arena is most useful when:
- Comparing newly released AI models.
- Identifying strong general-purpose performers.
- Understanding how people respond to different models in open-ended conversations.
APEX is most useful when:
- Evaluating models for legal, medical, finance, or software engineering workflows.
- Measuring whether a model can complete professional tasks accurately.
- Comparing deployment candidates using expert-graded evaluations rather than public opinion.
Many organizations use both benchmarks together. Arena helps narrow a broad field of models, while APEX helps determine whether those shortlisted models can reliably perform the work required in a production environment.
Measure real-world AI productivity with APEX benchmarks
Arena helps you find popular models. APEX shows which ones actually complete expert-graded professional tasks across law, finance, consulting, healthcare, and software engineering. For teams that want a deeper evaluation:
Talk to MercorFrequently Asked Questions
Why do AI decision-makers use benchmarks to compare models?+−
Teams use benchmarks to shortlist strong performers quickly, double-check a vendor's marketing claims, and track whether a model family is actually improving with each successive release in their own workflows and use cases.
There are different ways to measure AI model performance. A popularity signal, for example, is a model response that wins consistently against comparable competitors, a signal of strong performance. Meanwhile, task-based scores include benchmarks like Mercor’s APEX. They test whether a model can do the kind of work a business needs, which preference votes were never designed to measure.
Does Arena measure coding ability?+−
Yes. Through dedicated, code-focused categories and prompts submitted by developers, it measures a preference for code-related answers rather than running the code to verify that it works. A model that writes confident, well-formatted code can win votes without that code ever being executed or tested.
Can organizations use Arena and APEX together?+−
Yes. Many organizations use the two benchmarks at different stages of the model evaluation process. Arena helps identify promising AI models based on human preference in blind comparisons, while Mercor's APEX benchmarks evaluate those finalists on expert-graded professional tasks. Using both provides a more complete view of model performance before making deployment decisions.
Did LMSYS Chatbot Arena change its name to Arena?+−
Yes. The project began as LMSYS Chatbot Arena at UC Berkeley, later operated under the LMArena domain, and now runs under Arena. The core leaderboard and voting mechanic have stayed publicly accessible throughout each transition.

