Subscribe to mailing list

Get notified when we have new updates or new posts!

Subscribe Unicorn Data Science cover image
jen@unicornds.org profile image jen@unicornds.org

Which LLM Is the Best for Data Discovery and Data Visualization? We Benchmarked 16 Models

We gave 16 LLMs the same experiment data and one prompt: analyze it and build the report. Here's how they did on correctness, storytelling, charts, speed and cost.

Which LLM Is the Best for Data Discovery and Data Visualization? We Benchmarked 16 Models
A data analysis report generated by Google Gemini 3.7 Flash.

It's been more than a year since our Jeopardy! benchmarking series (Round 1, 2, 3, and 4). While trivia questions were fun, they are not representative of our day-to-day work. So in this new round of benchmarking, we jump right into data analysis.

However, it can be tricky to construct the test dataset. Foundational models are trained on the internet, so taking an existing dataset has the risk of the models having "seen the answers". For this reason, our test dataset is simulated, based on a very common scenario.

The Test Dataset

Imagine an online store that wants to redesign its product landing page. The team comes up with seven new versions, variants B through H, and tests them against the current page, variant A. Every visitor is randomly assigned to one version and stays on it for the whole test. We log what happens: did they click the call-to-action, did they add something to the cart, did they buy. Ten weeks of daily counts, eight versions, 560 rows in total.

This full dataset is small enough to include into a prompt, and it looks like the kind of table most of us have analyzed at work. By tuning the parameters for each scenario, variants were set up to be statistically distinct - some more performant, and some have interesting behaviors.

The Prompt

Each model receives the identical prompt with the dataset baked in. We provide the background of the analysis task, how the data was gathered, and request the model to generate an analysis report as a standalone HTML file and return nothing else. We then use OpenRouter to try different models. No additional tool or model parameter setting - every API call was made in exactly the same way.

You are a data analyst at an online retailer. The product team ran an experiment on the site's main product landing page and needs a report on the results.

## Background

Over ten weeks (2026-06-01 to 2026-08-09) we tested eight versions of the product landing page, labelled A through H. Every visitor who reached the page during the experiment was randomly assigned to one of the eight versions, in equal proportions, and stayed on that version for the whole period. Version A is the current page; B through H are candidate redesigns that change layout, copy and the primary call-to-action.

We logged what visitors did after landing, aggregated per day and per version:

| column        | meaning                                                                 |
|---------------|-------------------------------------------------------------------------|
| `date`        | calendar day                                                            |
| `variant`     | which version of the page the visitors saw (A–H)                        |
| `visitors`    | number of visitors assigned to that version on that day                 |
| `clicks`      | of those visitors, how many clicked the primary call-to-action          |
| `add_to_cart` | of those who clicked, how many added an item to their cart              |
| `purchases`   | of those who added to cart, how many completed a purchase               |

The team's goal is to decide what to do with the page next. They are not statisticians; they want to understand what actually happened in the experiment and what it means.

## Your task

Analyze the data below and produce a report, as a single self-contained HTML file with charts, that communicates what matters to the product team.

Requirements for the output:

- Return only the HTML document, starting with `<!DOCTYPE html>`. No explanation before or after it, and no markdown code fences (no ```html): the response will be saved to a file and opened in a browser exactly as returned.
- The file must render on its own when opened in a browser. Inline all CSS and JavaScript. You may load a charting library from a public CDN via a `<script>` tag, but do not fetch any data from the network; embed whatever data the charts need in the file.
- Do not use images or external assets.

## The data

```csv
date,variant,visitors,clicks,add_to_cart,purchases
2026-06-01,A,3407,1202,205,74
2026-06-01,B,3394,1166,192,68
2026-06-01,C,3510,1760,298,68
2026-06-01,D,3519,1262,209,66
2026-06-01,E,3377,1050,167,62
2026-06-01,F,3350,1269,222,71
2026-06-01,G,3301,1130,195,66
2026-06-01,H,3325,1156,179,52
<rest of data omitted>
```

Scoring

For each model, we scored it on six different criteria:

  • Speed: How long the API call took.
  • Cost: What OpenRouter charged for the API call.
  • Correctness: Did the report find the insights planted in the simulated data? Did the model make any wrong claims?
  • Narrative: Is the model a good storyteller? Does the report lead with the key finding, tell the team what they need to know without noise, back up its claims, and say what it means without hedging?
  • Data visualization approach: Did the model pick the right plotting type for the point it tries to make? How about chart hygiene like, axes labels, honest baselines, tick sizes, and confidence intervals when helpful?
  • Visual appeal: Is it nice to look at? Is there a wow factor? Does it feel polished?

Models

We picked a mix of flagship and budget models from six major providers. Most of the models were able to return a full HTML report - each report has a link so you can check it out on your own!

ProviderModelTimeCostReport
AnthropicClaude Opus 5.54:41$0.69Open report ↗
Claude Sonnet 58:42$0.64Open report ↗
page renders blank: a script error stops the report from loading
Claude Opus 511:09$1.58Open report ↗
OpenAIGPT-6 Luna1:09$0.01Open report ↗
GPT-6 Luna Pro1:21$0.01Open report ↗
GPT-6 Sol Pro1:54$0.29Open report ↗
page renders blank: a script error stops the results from loading
GPT-6 Astra3:25$0.91Open report ↗
GoogleGemini 3.7 Flash0:54$0.05Open report ↗
Gemini 3.1 Pro2:34$0.34Open report ↗
Gemini 3.8 Flash3:14$0.13Open report ↗
DeepSeekDeepSeek V4 Pro1:52$0.03Open report ↗
shown with a preamble and code fences removed by hand; the model did not return HTML only
DeepSeek V4.1 Flash4:56$0.01Open report ↗
QwenQwen 3.8 Max Prime10:07$0.84Open report ↗
renders, but every number shows as NaN due to a script bug
Qwen 3.8 Omni Flash17:05$0.00No report: call timed out while reasoning
Z.aiGLM 5.3 Prime7:17$0.61No report: cut off while reasoning
GLM 5.3 FlashX7:22$0.09No report: cut off while reasoning

Highlights

Before jumping into the scoring results, here are some observations worth mentioning.

  1. Model benchmarking has become easy-breezy with OpenRouter. Previously, I had to get DeepSeek hosted on AWS Bedrock - quite feasible but it took some time to set up. I also had several python clients to interact with different model providers. Fast forward one year later, trying out more than a dozen model can now be done with just one API key, with the same python code.
  2. Even without specifying data analysis tools, like statistical methods, most models are able to perform calculations reasonably. Barring provider failures, nearly all models, including those optimized for speed, can identify which variant is the most performant.
  3. The analytical findings varied more than the visuals did. Only three models, Claude Opus 5.5, Claude Opus 5 and GPT-6 Astra, picked up every pattern we had built into the data; most of the others caught the main result but missed the subtler ones. One model was drawn to the variant with the highest click-through rate rather than the one that actually sold the most. Three models wrote a perfectly reasonable analysis and then shipped it inside JavaScript that doesn't run, so their pages show empty boxes or NaN. And a few found the right answer, then hedged it into a recommendation for a confirmatory test, on a result that was not close.
  4. Google Gemini 3.7 Flash goes dark mode, the only model to make this choice. All other models had the default white background.
Gemini 3.7 Flash's data analysis report, the only one in the benchmark with a dark-mode design
Data analysis report from Google Gemini 3.7 Flash - the only model to go dark mode.
Gemini 3.8 Flash's data analysis report, light mode with a green executive-decision banner and a full-funnel scorecard table
In comparison, data analysis report from Google Gemini 3.8 Flash. Looks very different and uses light mode, which all models other than Google Gemini 3.7 Flash also used.
  1. Anthropic Claude Opus 5.5 and Opus 5 have... eerily similar looking reports.
  1. OpenAI GPT-6 Astra adds some nice interactivity touches like data visualization filtering, and data downloading.
GPT-6 Astra's report with checkboxes to toggle experiment variants on and off in the weekly chart, and a download-data button
OpenAI GPT-6 Astra's report is the only one that allows user to toggle experiment variants on/off for custom comparison. At the end of the report, there's even a "Download" button.

Results

After reviewing each result, and scoring their contents, we arrive at a full scoreboard, which is summarized below.

Full comparison on the performance of LLM models in generating data analysis reports.

So which LLM is best for data visualization?

If cost is not an issue, the new Opus 5.5 introduced a week ago is really hard to beat. It scores high on all correctness measures, and tells a solid data story. We didn't find it to be the most visually stunning, but the insights are there. And given Opus 5.5 is faster and cheaper while maintaining the same if not higher report quality compared to Opus 5, there's not really any reason to use Opus 5 anymore.

On the other hand, Google's Gemini 3.8 Flash and 3.7 Flash, although released less than a month apart, produced drastically different looking reports. This is unlike Opus 5 versus Opus 5.5 - the two reports shared identical layout. Given the much lower cost of the Google Gemini Flash models, their ability to produce good looking reports can be quite appealing. They are weaker analytically though than, e.g., Opus 5 and 5.5.

Another way to balance cost and quality is to leverage the cheapest model of the bunch, OpenAI GPT-6 Luna (each report cost $0.01!), with the analytical thoroughness of GPT-6 Astra.

And at least based on this round of benchmark, the Qwen and Z.ai models struggled with the task, mostly by never finishing their reasoning. DeepSeek was a mixed bag: V4 Pro went for the wrong variant, while V4.1 Flash got the main call right for about a cent.

Wrapping up

What impressed us most is how far a single prompt now goes. With no tools and no follow-up, most of these models came back with a report that looks reasonable at first glance. A year ago that took several rounds, and only with the most expensive models.

But none of the reports was perfect (or we were too picky!). Upon closer inspection, we found a script that doesn't run, a number that drifted, an axis that starts at 1%. Careful and thorough inspection, with domain expertise, is still very much needed. A polished-looking report can carry a subtle issue that leads to the wrong decision, and the only way to catch it is to check the work.