The takeaway
Design Arena proves that human judgment remains essential for AI evaluation — and that there is a $60M ARR business in providing it at scale.
Why it matters for builders
Design Arena signals that human evaluation data is becoming a critical infrastructure layer for AI model development. For builders deploying generative AI in production, it validates the approach of using A/B preference testing — not just automated metrics — to measure and improve output quality. The $60M ARR also proves there is real enterprise demand for this layer.
Design Arena Lands $7.9M to Rank AI Models With Human Taste
The startup behind Design Arena — called Intelligence — announced a $7.9 million seed round on Monday, betting that human judgment, not just automated benchmarks, is the missing piece in AI model evaluation.
Led by Index Ventures with participation from Conviction, A*, and Valkyrie, the round validates a simple premise: AI models that generate images, websites, and interfaces need real people to decide what looks good. Design Arena does exactly that — it presents users with A/B comparisons of AI-generated outputs and asks them to pick the best, turning subjective taste into measurable training data.
What Happened
Intelligence, founded by Grace Li and a team of college friends just weeks before graduation in 2025, started as an AI game engine project. The models could build functional games, but none were fun — which raised a deeper question: how do you measure whether an AI output is actually good?
The answer became Design Arena, a platform now used by 5.3 million people worldwide. It functions like a sophisticated model router: users submit prompts through a ChatGPT-style interface, choose from a dozen visual formats — websites, images, dashboards — and then rank the outputs through a series of A/B choices.

"It was the missing bottleneck for a lot of these models to make improvements in the design space," Li told TechCrunch. About a week after launch, the company closed its first major deal with a frontier AI lab, and the platform now generates $60 million in annual recurring revenue.
Why It Matters
The investment signals a growing recognition that automated benchmarks alone cannot evaluate AI model quality — especially for visual and design outputs. Last week's Hugging Face security breach demonstrated how benchmark data can be manipulated, reinforcing the case for human-led evaluation at scale.
Design Arena's approach is not without competition. LM Arena, which applies similar human-ranking methods to text-based models, raised $150 million in a Series A in January. However, not every player survives: Yupp, a comparable startup that raised $33 million from a16z crypto, shut down earlier this year.
Crucially, Design Arena tracks how user preferences shift across regions and over time. Li notes that web dashboards in Asia tend toward maximalist design — insights that automated metrics would miss entirely.
Key Details
- Funding: $7.9 million seed round led by Index Ventures
- Investors: Conviction (Sarah Guo, Mike Vernal), A*, Valkyrie
- Users: 5.3 million globally
- Revenue: $60 million ARR
- Founded: 2025, weeks before founders' college graduation
The round positions Intelligence at the intersection of two accelerating trends: the proliferation of AI-generated visual content and the growing demand for reliable evaluation beyond benchmark scores. For AI builders, the message is clear — taste matters, and now there is a metric for it.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
4 August 2026
4 August 2026
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.



