Two million votes measure what people prefer in generated images
Datapoint is releasing 2.16 million human preferences collected across 30 image models. A separate program provides credits for multimodal research projects.
Datapoint AI is releasing the data behind its image generation model leaderboard. The dataset contains 2,161,160 human votes across 216,116 side-by-side comparisons. At the same time, the company is creating a $1 million credit pool to help research teams fund their own studies involving images, video, and audio.
The dataset published on Hugging Face compares 30 models using 500 prompts divided into ten categories, including marketing, product design, and anime. Each category contains 50 requests designed to test several challenges, such as exact object counts, composition, element ordering, negation, and text rendering.
The protocol uses a complete round-robin between the models. The 30 systems create 435 possible pairings, each tested on every prompt for which both services produced an output. This provides approximately 5,000 direct votes for each matchup.
Each model generated only one image per prompt using a fixed seed. The instructions were submitted without rewriting, selecting from multiple candidates, or rerunning requests to improve the outcome. This method limits intervention by the organizer, but it also introduces an important constraint: a single result does not always represent the overall quality of a model whose output may vary between generations.
The evaluations were conducted blind. Two images appeared side by side without model names, and their left-right positions were randomized. Annotators answered a single question: which image did they prefer? Each pair received ten validated votes.
This simple format supports data collection at scale, but it combines several criteria into one decision. One person may favor aesthetics, while another may prioritize prompt adherence, text legibility, realism, or the absence of visible defects. The leaderboard therefore measures overall preference without identifying which specific quality determined each vote.
Datapoint says the responses came from people across more than 200 countries. Each record includes the country, response time, an anonymized identifier, and the annotator’s trust score at the time of voting. The reported median response time is approximately eleven seconds.
Geographic diversity does not, however, guarantee balanced representation across cultures, languages, or social groups. The documentation does not provide the panel’s full demographic composition or the distribution of votes by country. The dataset can support research into these differences, but its results should not automatically be treated as a universal definition of taste.
The Datapoint Image Bench leaderboard is calculated from raw votes using a Bradley–Terry fit. The ten categories receive equal weight, while FLUX.1 Schnell serves as the reference model with a fixed score of 1,000.
GPT Image 2 in high mode ranks first overall, but the gaps are narrow. Fewer than ten Elo points separate the top three models, placing them within one another’s confidence intervals. Seedream 5.0 Pro leads five categories, GPT Image 2 leads three, and Nano Banana 2 leads two. The overall top position therefore reflects consistency across different uses rather than dominance in every category.
The results also show that 18.2% of comparisons ended in a tie after the ten votes were counted. This figure highlights how closely several recent models are perceived and shows how a general ranking can conceal much less decisive preferences on individual prompts.
The leaderboard is only one possible use for the dataset. The data can be used to train or evaluate reward and preference models, prepare DPO datasets, compare aggregation methods, or audit Datapoint’s published results. Researchers can, for example, retain only matchups won by seven votes to three or more to exclude the least decisive comparisons.
A training and test split is already included. Fifty prompts, five from each category, are reserved for testing. They account for 21,663 image pairs and 216,630 responses. No prompt appears in both sets, reducing the risk of evaluating a system on instructions it encountered during training.
The dataset also includes 14,952 images, a list of the models, their API identifiers, their unit costs, and the evaluation rubrics associated with the prompts. The complete release totals 34.9 GB. The number of images is slightly below the theoretical 15,000 outputs expected from 30 models multiplied by 500 requests because some services did not return a result for every case.
Datapoint describes the release as the largest open human-preference dataset dedicated to image generation. This claim is based on comparisons with several public references, including Pick-a-Pic v2, HPS v2, and ImageReward. It comes from the company and does not necessarily account for private datasets held by AI labs.
Access is not entirely unrestricted. The repository is publicly visible, but downloading the files requires users to share their contact information with Datapoint, accept the attribution requirements of the CC BY 4.0 license, and agree not to attempt to identify annotators. The license covers the votes, prompts, and metadata. Rights to the images remain subject to the terms of the providers that produced each result.
The release is accompanied by a $1 million program for researchers who need new human evaluations. This is not cash funding. Awards are issued as credits that can be used on Datapoint to organize comparisons, rankings, ratings, or written feedback studies.
Typical awards range from $5,000 to $50,000 in credits. Depending on the complexity of a study and its targeting requirements, they can fund tens of thousands to hundreds of thousands of judgments. Applications are reviewed on a rolling basis until the pool is exhausted, with Datapoint promising a response within 24 hours.
The program is intended for PhD students, postdoctoral researchers, university labs, independent researchers with prior work, and nonprofit organizations. Projects should lead to a paper, benchmark, or open dataset. Data collection