A self-custody Web3 wallet has just reached an unprecedented milestone: making public the real performance data of hundreds of AI agents configured by its users on decentralized derivatives platforms.
Out of 688 agents analyzed, 42% finished in positive territory — and the top performer posted a ROI of +307%. A performance gap that raises as many questions as it answers.
Here is what this benchmark concretely reveals about the real state of AI algorithmic trading in 2026.
688 Agents, 7 LLM Families: What the Numbers Really Say
Wallet V, a Web3 wallet incubated by Virgo Group, has published an aggregated benchmark covering 688 AI trading agents deployed by its users over the past two months. These agents operated on Hyperliquid and Aster — two decentralized derivatives platforms — executing strategies on perpetual contracts.
Each agent was manually configured by the user, who also selected the large language model (LLM) responsible for generating trading decisions. The benchmark then aggregates performance by model family, covering seven distinct LLMs. Models represented by fewer than 10 agents are flagged as directional only, with no conclusive statistical significance.
The raw results: 42% of agents recorded a flat or positive P&L over the period. Peak ROI ranged from -30% for the worst-performing model to +307% for the best. A spread of 337 percentage points that illustrates just how decisive the choice of LLM — and configuration — can be.

BTC, ETH, Gold, Forex: The Asset Classes Covered by the Agents
The agents in the benchmark did not operate exclusively on crypto assets. They accessed four asset classes available on Hyperliquid and Aster via perpetual contracts:
- Major crypto assets: BTC, ETH, SOL
- Equities: including exposure to pre-IPO companies via tokenized shares
- Commodities: gold, silver, oil
- Forex: major currency pairs
This diversification reflects the broader ambition of Wallet V: to move beyond the crypto market alone and offer an agent infrastructure capable of operating across all tokenized financial markets. Adam Cai, founder and CEO of Virgo Group, sums up the approach: “Users are now choosing their AI model the way institutions evaluate fund managers — by examining observable performance over time.”
Virgo Group is backed by investors including Draper Dragon, OKX Ventures, and Cobo Ventures. The benchmark is hosted directly on the Wallet V website and updated continuously as new agents are deployed.
What the Next Versions of the Benchmark Will Change
Wallet V has announced several developments for upcoming iterations of the benchmark. On the roadmap: the integration of new LLM families, support for prediction markets, advanced analytics features for copilot trading, and AI prompt generation personalized to each user’s trading style.
That last feature is particularly noteworthy: it points to a continuous improvement loop where the agent adapts to the user’s risk profile and preferences rather than applying a one-size-fits-all strategy. This is precisely the kind of personalization that sets an institutional-grade tool apart from a basic crypto trading bot.
Wallet V‘s public benchmark represents a rare initiative within the ecosystem: most AI trading agent solutions remain black boxes. Making this data accessible — even in aggregated form — introduces a standard of transparency and accountability that could set a precedent across the decentralized algorithmic trading sector.