❯ Model leaderboard Arena raises a $200 million Series B at a $3.1 billion valuation and releases an Alignment Index measuring whether models behave
Valuation nearly doubles in ten monthsAI model evaluation platform Arena announced a $200 million Series B on October 8 at a $3.1 billion valuation. Lightspeed and Khosla Ventures co-led, with existing investors including a16z participating. According to TechCrunch, its Series A in January was $150 million at a $1.7 billion post-money valuation.
Users vote, models get rankedArena began in 2023 as a research project at UC Berkeley. Ordinary users enter prompts or request coding projects for free and judge which model did better, and those votes feed the leaderboard. In September last year it introduced a paid product, AI Evaluations, selling model labs and enterprises detailed performance analytics based on community feedback. The company says annualized revenue has passed $100 million, up from $30 million in January.
A new board for honestyReleased alongside the funding, the Alignment Index measures whether a model acts in line with what the user intended, looking at three problems: taking actions it was not asked to take, attributing statements to the wrong source, and claiming to have completed tasks it did not do. TechCrunch said a slate of OpenAI models leads the preliminary board, with Claude Opus 5.5 in sixth place.
Fixed exams no longer sufficeArena argues that static benchmarks break down once models recognize they are being tested, and that when agents write code or run analyses for people, users often cannot check the work. Enterprises buying models therefore need evaluations drawn from real use. Arena’s own difficulty is staying neutral: its paying customers are the same labs it ranks.
▮ SIGNALAn evaluation company more than tripling revenue in a year shows vendor-published scores are no longer enough to buy on, and refereeing is turning into a business of its own.