A new paper
from AI lab Cohere, Stanford, MIT, and Ai2 accuses LM Arena, the organization behind the popular crowdsourced AI benchmark Chatbot Arena, of helping a select group of AI companies achieve better leaderboard scores at the expense of rivals.
According to the authors, LM Arena allowed some industry-leading AI companies like Meta, OpenAI, Google, and Amazon to privately test several variants of AI models, then not publish the scores of the lowest performers. This made it easier for these companies to achieve a top spot on the platform’s leaderboard, though the opportunity was not afforded to every firm, the authors say.
“Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others,” said Cohere’s VP of AI research and co-author of the study, Sara Hooker, in an interview with Vmeetsolutions News. “This is gamification.”
Created in 2023 as an academic research project out of UC Berkeley, Chatbot Arena has become a go-to benchmark for AI companies. It works by putting answers from two different AI models side-by-side in a “battle,” and asking users to choose the best one. It’s not uncommon to see unreleased models competing in the arena under a pseudonym.
Votes accumulated over time influence a model’s score—and as a result, its position on the Chatbot Arena leaderboard. Despite numerous commercial entities participating in Chatbot Arena, LM Arena insists that their evaluation criteria remain unbiased and equitable.
That said, this isn’t what the researchers claim to have discovered in their study.
An AI firm named Meta allegedly conducted private testing on the Chatbot Arena for 27 different model variations from January through March prior to releasing their product, Llama 4. Upon launching, Meta chose to disclose only the ranking data of one particular model, which coincidentally placed highly within the Chatbot Arena rankings.
In an email to Vmeetsolutions News, LM Arena Co-Founder and UC Berkeley professor Ion Stoica stated that the research contained numerous “mistakes” and “dubious conclusions.”
We are dedicated to conducting equitable assessments led by the community and encourage all model providers to submit additional models for evaluation with the aim of enhancing their performance according to human preferences,” stated LM Arena in remarks shared with Vmeetsolutions News. “The fact that one model provider might choose to submit more tests than another doesn’t imply that the latter would be disadvantaged.
Armand Joulin, who is a principal researcher at Google DeepMind, similarly pointed out
post on X
Some of the study’s figures were found to be incorrect, with Google stating they had submitted only one Gemma 3 AI model for pre-release testing at LM Arena. In response to Joulin, Hooker assured on X that corrections would be made by the authors.
Supposedly favored labs
The paper’s authors started conducting their research in November 2024 after learning that some AI companies were possibly being given preferential access to Chatbot Arena. In total, they measured more than 2.8 million Chatbot Arena battles over a five-month stretch.
According to the authors, they discovered proof indicating that LM Arena enabled select AI firms such as Meta, OpenAI, and Google to gather additional data from Chatbot Arena by featuring their models in a greater number of model “duels.” The elevated sampling frequency provided these corporations with what the authors claim was an unjust edge over others.
Using additional data from LM Arena could improve a model’s performance on Arena Hard, another benchmark LM Arena maintains, by 112%. However, LM Arena said in a
post on X
The performance in the Arena Hard test does not have a direct correlation with the ChatbotArena performance.
Hooker mentioned that it’s uncertain how some AI firms obtained preferential treatment, however, Hooker emphasized that LM Arena must enhance its transparency irrespective of this ambiguity.
In a
post on X
, LM Arena said that several of the claims in the paper don’t reflect reality. The organization pointed to a
blog post
It was published earlier this week suggesting that models from lesser-known laboratories participate in more ChatbotArena battles than indicated in the study.
A significant constraint of this research is its reliance on “self-identification” for identifying AI models undergoing private tests within the Chatbot Arena. The researchers asked multiple questions regarding each model’s originating company, then used these responses as the basis for classification—a process that lacks absolute certainty.
Hooker mentioned that when the authors contacted LM Arena to present their initial findings, the organization did not contest them.
Vmeetsolutions News contacted Meta, Google, OpenAI, and Amazon—companies cited in the study—for their input; however, they did not respond promptly.
LM Arena faces trouble
In the paper, the authors call on LM Arena to implement a number of changes aimed at making Chatbot Arena more “fair.” For example, the authors say, LM Arena could set a clear and transparent limit on the number of private tests AI labs can conduct, and publicly disclose scores from these tests.
In a
post on X,
LM Arena rejected these suggestions, claiming it has published information on pre-release testing
since March 2024
The benchmarking group also stated that it “doesn’t make sense to display scores for pre-release models that aren’t accessible to the public,” as this prevents the AI community from testing these models independently.
The researchers additionally suggest that LM Arena should modify Chatbot Arena’s sampling rate so that each model within the arena participates equally in the battles. Publicly, LM Arena has shown openness to this suggestion and stated that they will develop a new sampling algorithm for this purpose.
The report surfaced shortly after Meta was found manipulating benchmarks in the ChatbotArena following the release of their aforementioned Llama 4 series. They fine-tuned one of the Llama 4 models specifically for “conversation,” enabling it to secure a remarkable position on ChatbotArena’s rankings. However, Meta decided not to publish this enhanced variant; instead, they only offered the standard version.
turned out to perform significantly poorer
on Chatbot Arena.
Initially, LM Arena stated that Meta ought to have demonstrated greater openness regarding its methodology for benchmarking.
Earlier this month, LM Arena declared it was
launching a company
, with intentions of securing funding from investors. This study intensifies examination into private benchmark organizations—and questions their reliability in evaluating AI models impartially, free from corporate interference.