Compare computer-use agents through live tasks and human judgments
CoArena is a computer-use evaluation product operated by Coasty Systems. Its public website presents live tasks in which agents interact with browser or desktop software, with comparisons based on human judgments and recorded execution information. This makes it different from a general chatbot directory or a model-training dashboard. The central question is how agents behave while attempting a task, including the steps they take and the result they produce, rather than only whether a single text response sounds convincing.
The site offers a leaderboard, model profiles and comparisons across multiple model families. It encourages readers to consider human preference, reported completion, speed, recovery and estimated cost together. Those measures answer different questions. A preferred attempt may not be the fastest one, and a successful example does not establish that an agent will perform reliably on every workflow. The associated methodology is important context when using a ranking to make a development or evaluation decision. Public benchmark information should be interpreted with its stated measurement rules.
The arena invites users to submit tasks and compare attempts, including blind human judgments. A useful evaluation task is specific enough that someone can recognize a successful result. Reviewing the actual attempt is more informative than choosing solely from a model name or an overall score. Researchers can also use the public comparison pages as a starting point for identifying agents worth testing on their own workload. The current roster is a changing catalog, so a fixed list of available models should not be assumed indefinitely.
CoArena's privacy notice describes extensive collection of task-related activity, including screenshots and evaluation behavior, and states that datasets may be licensed. That is material to deciding what information to put into a task. Public benchmark browsing and signed-in participation are not the same data relationship. Avoid submitting confidential records or sensitive credentials without understanding the current terms and controls. The service is useful for examining agent behavior, but a public benchmark should not replace a separate review of permissions, reliability and data handling for a production deployment.