As open and closed models get more capable, the gap between benchmark scores and real-world performance has become one of the hardest problems for model providers and their customers. Arena fills that gap with a human preference platform where real, unpaid users compare anonymous model outputs across text, coding, vision, image generation, and agentic tasks. For agents specifically, Arena’s Agent Arena uses a causal inference methodology that observes the full human-agent workflow, not just the final output, to measure how well agents actually perform real work. Those first-party signals feed the public leaderboards labs and builders rely on.

Leadership

Anastasios Angelopoulos - CEO
Wei-Lin Chiang - CTO
Ion Stoica - Chairman

Status

Private