AiSulivo
AiSulivo
Menu
AiSulivo
AiSulivo

CoArena

Compare computer-use agents through live tasks and human judgments

Pricing
Public benchmark information is accessible. Paid task-execution terms not publicly verified
Free plan
Public benchmark access; task-execution allowance not verified
Platforms
Web

Tool Information

CoArena
Coasty Systems, Inc.
Updated: September 2026
Tool type: Computer-Use Agent Evaluation Arena
Pricing: Public benchmark information is accessible. Paid task-execution terms not publicly verified
Free plan: Public benchmark access; task-execution allowance not verified
Platforms: Web
Login required: Public pages are open; account features use Google sign-in
API: Official site links a metrics API; general task-execution API not verified
Browser extension: No official browser extension verified
Mobile app: No dedicated native mobile app verified
AI models: Model profiles span Claude, GPT, Gemini and other families. The available comparison roster can change.
Developer: Coasty Systems, Inc.

About CoArena

Compare agents on computer-use work

CoArena is a computer-use evaluation product operated by Coasty Systems. Its public website presents live tasks in which agents interact with browser or desktop software, with comparisons based on human judgments and recorded execution information. This makes it different from a general chatbot directory or a model-training dashboard. The central question is how agents behave while attempting a task, including the steps they take and the result they produce, rather than only whether a single text response sounds convincing.

Use several measures together

The site offers a leaderboard, model profiles and comparisons across multiple model families. It encourages readers to consider human preference, reported completion, speed, recovery and estimated cost together. Those measures answer different questions. A preferred attempt may not be the fastest one, and a successful example does not establish that an agent will perform reliably on every workflow. The associated methodology is important context when using a ranking to make a development or evaluation decision. Public benchmark information should be interpreted with its stated measurement rules.

Participate with appropriate task material

The arena invites users to submit tasks and compare attempts, including blind human judgments. A useful evaluation task is specific enough that someone can recognize a successful result. Reviewing the actual attempt is more informative than choosing solely from a model name or an overall score. Researchers can also use the public comparison pages as a starting point for identifying agents worth testing on their own workload. The current roster is a changing catalog, so a fixed list of available models should not be assumed indefinitely.

Read the data disclosures before submitting

CoArena's privacy notice describes extensive collection of task-related activity, including screenshots and evaluation behavior, and states that datasets may be licensed. That is material to deciding what information to put into a task. Public benchmark browsing and signed-in participation are not the same data relationship. Avoid submitting confidential records or sensitive credentials without understanding the current terms and controls. The service is useful for examining agent behavior, but a public benchmark should not replace a separate review of permissions, reliability and data handling for a production deployment.

Key features
  • Live computer-use task comparisons
  • Blind human preference judgments
  • Recorded execution metrics
  • Public agent leaderboard
  • Model profiles and comparison pages
  • Published benchmark methodology
  • Metrics API linked from the official site
Use cases
Computer-Use Agent Evaluation,Model Comparison,Research Task Testing,Agent Behavior Review,Benchmark Exploration
How to use
  1. Open the official CoArena website.
  2. Read the methodology and relevant privacy disclosures.
  3. Browse model profiles and comparison measures.
  4. Define a nonconfidential task with a clear success condition.
  5. Sign in if the task workflow requires an account.
  6. Submit the task and observe the agent attempts.
  7. Review the evidence before giving a preference judgment.
  8. Compare results with your own evaluation requirements.
Best for
AI Researchers, Agent Developers, Model Evaluators
Integrations
Google Sign-In,Computer-Use Browser And Desktop Environments
Commercial use
Dataset and benchmark reuse subject to published terms; do not assume unrestricted rights

Related Tags