Slide creation benchmark
The frontier of AI slide creation
Ten frontier models completed the same 150 slide creation tasks. We compared output quality, speed, cost, token use, and performance as task complexity increased.
- Frontier models
- 10
- Shared successful tasks
- 139
- Selected AI ratings
- 10,600
Arena scores for every model tested
Pooled three-judge Bradley–Terry scores with Gemini 3.7 Flash anchored at 1000, alongside the median cost and time of one successful slide creation task.
| Rank | Model | Arena score | Median cost per task (USD) | Median time per task (seconds) |
|---|---|---|---|---|
| 1 | Opus 5.5 | 1175.2, best | $0.76 | 170s |
| 2 | Sonnet 5.5 | 1133.0 | $0.45 | 144s, best |
| 3 | GPT-6 Astra | 1123.6 | $1.06 | 226s |
| 4 | Fable 5.1 | 1050.2 | $2.38 | 222s |
| 5 | GPT-6.1 Sol | 1042.0 | $0.35 | 220s |
| 6 | Opus 5 | 1034.8 | $1.52 | 293s |
| 7 | GPT-6 Sol | 1023.8 | $0.33 | 168s |
| 8 | Gemini 3.8 Flash | 1002.6 | $0.31 | 207s |
| 9 | Gemini 3.7 Flash | 1000.0 | $0.24, best | 148s |
| 10 | GPT-5.6 Sol | 894.6 | $0.79 | 170s |
Cell shading ranks each column from best (darkest) to worst (lightest).
Slidely slide creation benchmark · 139 shared successful tasks · 10,600 blind ratings · September 2026
Update (September 2026): We added Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Sol, and GPT-6.1 Sol to the benchmark, all at the high reasoning setting and judged by the same three reviewers, and all GPT-6 Astra results now come from a high-reasoning run. Opus 5.5 is the quality leader with an Arena score of 1175 at a $0.76 median task cost. Sonnet 5.5 (1133) and Astra (1124) are effectively tied, with Sonnet costing $0.45 per median task against Astra's $1.06 and finishing in 144 seconds against 226. GPT-6 Sol scores 1024, a 129-point jump over GPT-5.6 Sol, while cutting the median task cost from $0.79 to $0.33. GPT-6.1 Sol adds another 18 points at 1042 for $0.35 per median task, but takes 220 seconds against GPT-6 Sol's 168. The leaderboard above and every chart below include all ten models.
TL;DR: Claude Opus 5.5 is the best slide creation model in the Slidely harness. Claude Sonnet 5.5 and GPT-6 Astra sit close together behind it, with Sonnet costing less than half of Astra per task. GPT-6 Sol brings OpenAI's price down to $0.33 per task at Gemini Flash quality, a quantum leap over GPT-5.6 Sol. GPT-6.1 Sol improves on it by 18 points for about three cents more per task, but runs slower.
Methodology in brief
Detailed methodology is available below.
We gave every model the same task: create a slide from a layout, content, visual reference, and ASCII plan, while following the layout even when some adaptation was required. The dataset contains 150 samples split across medium, high, and very-high-density slides. Quality results use the 139 tasks completed successfully by every model.
Judging slide creation proved to be extremely difficult. It is not only about visual appeal. A useful slide must follow the layout, group shapes correctly, keep text readable without overlaps, use space well, and create a clear visual hierarchy.
We tried several approaches and compared them with expert human choices on a smaller sample. Individual rubric scores were inconsistent, even with a detailed rubric. Judges that saw many outputs at once often missed important details. Agentic judges that inspected every input helped, but were expensive and still sensitive to the scoring scale.
We settled on blind pairwise comparisons with ties allowed. Three AI reviewers, Gemini 3.7 Flash, GPT-5.6 Sol, and Fable 5.1, compared two anonymous slides at a time. We pooled their decisions and fit a Bradley-Terry model to get a relative quality score. Gemini 3.7 Flash is fixed at 1000 as the reference point. The choice of anchor does not change the ranking or the distance between models.
The results correlated well with expert human choices on the smaller sample. Every creation model used the high reasoning setting, and caching was enabled for all models.
Results and analysis
Overall pooled Arena score
Bradley–Terry scores pooled across the three selected AI reviewers, with Gemini 3.7 Flash anchored at 1000.
Claude Opus 5.5 performs best in terms of quality, with an Arena score of 1175, ahead of Sonnet 5.5 at 1133 and GPT-6 Astra at 1124. Until this generation, OpenAI models performed relatively poorly compared with models from Google and Anthropic in our benchmark and in practical use. Astra reverses that pattern and shows extraordinary spatial reasoning, beating Fable 5.1 by 73 points.
Qualitatively, Astra is very good at adapting content without losing the layout's underlying structure. It uses larger and more consistent font sizes when space is available, which has been a persistent problem with every other model we tested. It also handles text contrast and purposeful icon use better. The results feel smart and, more importantly, usable.
As slide density rises, the difference between models becomes more pronounced. Astra and Fable win more often than the Gemini models on higher-density slides.
Pooled Arena score by slide density
Quality scores for medium, high, and very-high-density tasks, independently anchored at Gemini 3.7 = 1000.
Opus 5.5 leads on medium and high density, and Sonnet 5.5 climbs from 1090 on medium slides to 1200 on very-high-density slides, where it edges ahead. Astra is steady between 1118 and 1128 across all three groups. Gemini 3.8 moves in the opposite direction, falling from 1013.0 on medium slides to 985.8 on very-high-density slides. GPT-6.1 Sol does best on very-high-density slides at 1065, against 1042 on medium and 1027 on high.
Cost per slide creation task
Median cost per successful slide creation task
Internal billed credits converted to equivalent USD costs for the same 139 shared cases.
Gemini 3.7 Flash is still the cheapest model at $0.24 per median task, with GPT-6 Sol and GPT-6.1 Sol close behind at $0.33 and $0.35. Among the top three, Sonnet 5.5 costs $0.45 and Opus 5.5 $0.76, while Astra costs $1.06, about 55 percent less than Fable 5.1 at $2.38. Sonnet 5.5 has the best combination of quality and cost.
Every model costs more on very-high-density slides than on medium ones. Astra rises from $0.99 on medium slides to $1.20 on very-high-density slides, Fable from $2.01 to $2.79, and Sonnet 5.5 from $0.40 to $0.51. The scatter view makes the tradeoff clear: Gemini 3.7, GPT-6 Sol, and GPT-6.1 Sol are the economy choices, while Sonnet 5.5 and Opus 5.5 buy top quality for less than Astra.
Time per slide creation task
Median time per successful slide creation task
End-to-end creation time for the same 139 shared cases.
Sonnet 5.5 is the fastest model at 144 seconds per median task, just ahead of Gemini 3.7 Flash at 148. GPT-6 Sol takes 168 seconds and Opus 5.5 170. Astra needs 226 seconds, Fable 222, GPT-6.1 Sol 220, and Opus 5 293.
Density adds 85 seconds to Astra's median, from 202 seconds on medium slides to 287 on very-high-density slides. Sonnet 5.5 nearly doubles from 112 to 211 seconds, while GPT-5.6 Sol barely moves. Opus 5 and Fable slow the most on the hardest tasks.
Output tokens per slide creation task
Reasoning and visible output token mix
Median output tokens split into reported reasoning tokens and the remaining generated tokens.
6 Sol
5.6 Sol
6 Astra
6.1 Sol
Opus 5.5
Fable 5.1
Sonnet 5.5
Opus 5
3.7 Flash
3.8 Flash
GPT-6 Sol and Astra use the fewest output tokens, at 8,708 and 9,839 per median task. GPT-6.1 Sol produces 10,784, Opus 5.5 14,362, Fable 14,763, Sonnet 5.5 15,218, Opus 5 20,719, Gemini 3.7 25,809, and Gemini 3.8 47,953. More output clearly does not mean a better slide.
Astra also spends the smallest share of its output on reasoning among the top models, about half, where Sonnet 5.5 spends two-thirds. There has been speculation around terms such as neuralese and recurrent-depth transformers, but OpenAI's public materials about Astra do not confirm either architecture. The measurable result is simpler: Astra reasons in fewer tokens than the Anthropic models while producing slides of similar quality.
Other interesting results from our testing
Before finalizing the judging approach, we also benchmarked earlier generations of Google and OpenAI models. We did not test their quality with the current pairwise method, so the historical comparison focuses only on cost, total creation time, and output tokens.
How model economics changed over time
Cost, end-to-end creation time, and output-token use for earlier Google and OpenAI models, shown separately in launch order.
Google: Median cost per task
Google's Flash models repeatedly brought cost and time down after the heavier Gemini 3.1 Pro run. Gemini 3.7 Flash returned to a $0.24 median cost and 148-second median time, while producing roughly twice as many output tokens as Gemini 2.5 Pro.
OpenAI followed a different curve. Time rose from 150 seconds with o3 to 432 seconds with GPT-5.2 and remained high with GPT-5.4. Sol then brought the median back to 171 seconds and cut output to 9,264 tokens, although its median cost rose to $0.79.
Model progress is not one smooth efficiency curve. New releases often improve one operating metric while giving up ground on another. Still, the amount of usable work per dollar has increased substantially as quality has improved. We cannot put one exact number on that historical gain because the older outputs were evaluated using a different quality rubric.
Methodology details
Every model received the same task. For each case, it got the same slide layout, content, ASCII plan, and visual reference. The surrounding system, tools, cached inputs, and creation workflow stayed fixed. Only the model changed.
Scoring the outputs took several attempts to get right. We first asked individual AI judges to score each slide against a rubric covering content, layout, readability, and visual appeal. The scores often bunched near the top, and different judges rarely agreed on exact numbers.
We then tried agentic judges that inspected the slide, layout, and supporting material step by step. We also asked judges to rank several outputs at once. Both approaches helped, but they remained sensitive to the choice of judge and the scoring scale. Showing many slides together also made it easier to miss small but important visual problems.
Blind pairwise comparison worked best. A reviewer sees the reference layout and two anonymous slide images, then chooses the better one or calls a tie. One concrete choice at a time was easier to interpret than an absolute score.
Six models create 15 possible pairs per case. Across 139 cases and three reviewers, that produces 6,255 blind ratings. We combine those choices with a Bradley-Terry model. Each model has an unknown strength, and the calculation finds the strengths that best explain every observed win, loss, and tie. A tie counts as half a win for each model. We then convert the result to a chess-style scale and set Gemini 3.7 Flash to 1000.
The reviewers do not always agree, and we show that instead of hiding it.
Which creation models each judge preferred
Normalized positional scores grouped by judge, with the same model colors used across the article.
Sol ranks Astra first and Fable ranks Opus 5.5 first. Gemini ranks Fable 5.1 first and Astra in the middle of the pack. All three judges place GPT-5.6 Sol last. The model order in the middle depends on the reviewer, which is exactly why pooling matters.
Agreement between the selected AI judges
Exact agreement on shared pairwise decisions. Fable made all 12 tie decisions.
Sol and Fable
Gemini and Fable
Gemini and Sol
Gemini and Sol agree on 60.2 percent of decisions. Gemini and Fable agree on 61.3 percent, while Sol and Fable agree on 63.9 percent. Agreement around 60 percent is meaningful, but not strong enough to trust one judge as the final answer.
Do judges prefer models from their own lab?
Preference for each judge's model family in cross-family battles, compared with how the other two judges rated that same family.
Every judge preferred outputs from its own model family more often than the other two judges did. The Gemini judge preferred Google models in 51.6 percent of cross-family matches, compared with 45.7 percent for the other judges. Sol preferred OpenAI models 50.4 percent of the time, compared with 44.0 percent. Fable showed the largest difference, preferring Anthropic models 63.4 percent of the time, compared with 52.6 percent.
This is evidence of a possible family preference, not proof of bias. Each family contains models with different underlying quality. Pooling judges from three labs reduces the chance that one judge's family preference determines the final ranking.
The approach also scales well. When a new model arrives, we do not need to repeat every old comparison. We run the new model on the same cases, compare it with the existing models, and refit the Bradley-Terry scores with the additional results.
We also plan to make the blind arena public so that human reviewers can participate. That will give us a stronger independent signal and make it possible to compare AI and human preferences at a much larger scale.
Appendix
Cache reuse between turns
Cache hit rate after the first turn
Share of input tokens served from cache on repeat model calls across the 139 shared tasks.
The chart measures cached input tokens as a share of all input tokens after the first model call in each task. Sonnet 5.5 has the highest repeat-turn cache hit rate at 92.2 percent, followed by GPT-6.1 Sol at 92.1, Opus 5.5 at 91.6, Astra and GPT-6 Sol at 90.7, Fable at 88.5, GPT-5.6 Sol at 88.0, Opus 5 at 86.9, Gemini 3.8 at 71.4, and Gemini 3.7 at 60.4.
The long tail behind the medians
The main article uses medians because a few unusually difficult tasks can distort the mean. The p99 value is the second-highest observation in this 139-case sample. The p100 value is the maximum.
| Model | Median time | Mean time | p99 time | p100 time |
|---|---|---|---|---|
| Gemini 3.7 Flash | 148s | 171s | 458s | 471s |
| Gemini 3.8 Flash | 207s | 217s | 384s | 414s |
| GPT-5.6 Sol | 170s | 178s | 353s | 416s |
| Opus 5 | 293s | 306s | 685s | 826s |
| Fable 5.1 | 222s | 241s | 659s | 850s |
| GPT-6 Astra | 226s | 236s | 448s | 451s |
| GPT-6 Sol | 168s | 178s | 356s | 410s |
| Opus 5.5 | 170s | 183s | 446s | 513s |
| Sonnet 5.5 | 144s | 169s | 596s | 719s |
| GPT-6.1 Sol | 220s | 228s | 469s | 571s |
| Model | Median cost | Mean cost | p99 cost | p100 cost |
|---|---|---|---|---|
| Gemini 3.7 Flash | $0.236 | $0.250 | $0.545 | $0.636 |
| Gemini 3.8 Flash | $0.309 | $0.329 | $0.609 | $0.655 |
| GPT-5.6 Sol | $0.788 | $0.809 | $1.318 | $1.348 |
| Opus 5 | $1.518 | $1.574 | $3.200 | $3.664 |
| Fable 5.1 | $2.382 | $2.478 | $5.645 | $6.382 |
| GPT-6 Astra | $1.064 | $1.074 | $1.727 | $1.727 |
| GPT-6 Sol | $0.327 | $0.340 | $0.564 | $0.609 |
| Opus 5.5 | $0.759 | $0.789 | $1.409 | $1.682 |
| Sonnet 5.5 | $0.455 | $0.489 | $1.182 | $1.400 |
| GPT-6.1 Sol | $0.355 | $0.361 | $0.600 | $0.627 |
| Model | Median output tokens | Mean | p99 | p100 |
|---|---|---|---|---|
| Gemini 3.7 Flash | 25,809 | 26,710 | 48,559 | 60,782 |
| Gemini 3.8 Flash | 47,953 | 48,255 | 68,527 | 68,622 |
| GPT-5.6 Sol | 9,028 | 9,539 | 18,080 | 19,195 |
| Opus 5 | 20,719 | 21,698 | 46,550 | 57,584 |
| Fable 5.1 | 14,763 | 16,309 | 47,151 | 57,947 |
| GPT-6 Astra | 9,839 | 10,155 | 18,251 | 21,710 |
| GPT-6 Sol | 8,708 | 8,980 | 14,085 | 18,249 |
| Opus 5.5 | 14,362 | 14,853 | 34,751 | 42,610 |
| Sonnet 5.5 | 15,218 | 17,738 | 68,362 | 69,692 |
| GPT-6.1 Sol | 10,784 | 11,097 | 20,284 | 21,698 |
GPT-6 Sol and GPT-6.1 Sol have the tightest cost tails in the group, never exceeding $0.61 and $0.63. Fable's maximum task took 850 seconds and cost $6.38, while Astra's maximum took 451 seconds and cost $1.73.
Renders produced
Mean renders produced per final slide
Each render is one visible attempt produced before the final slide was selected.
Most models averaged roughly four visible render attempts per final slide, with GPT-5.6 Sol lowest at 3.7 and Sonnet 5.5 highest at 5.0. Most p99 values sit between six and eight renders, while Fable had one task that reached eleven and Sonnet 5.5 one that reached thirteen.
Pairwise match results
Pairwise wins across all three judges
Total pooled pairwise wins for each model across the three judges on the shared tasks.
| Model | Wins | Losses | Ties |
|---|---|---|---|
| Opus 5.5 | 765 | 380 | 0 |
| Sonnet 5.5 | 836 | 645 | 0 |
| GPT-6 Astra | 802 | 657 | 0 |
| Fable 5.1 | 1,172 | 910 | 3 |
| GPT-6.1 Sol | 793 | 976 | 1 |
| Opus 5 | 1,119 | 964 | 2 |
| GPT-6 Sol | 507 | 363 | 0 |
| Gemini 3.8 Flash | 1,004 | 1,074 | 7 |
| Gemini 3.7 Flash | 994 | 1,082 | 9 |
| GPT-5.6 Sol | 643 | 1,439 | 3 |
A model in the full six-way round-robin has 2,085 pooled outcomes: five matchups per case across 139 cases and three reviewers.
We need even better benchmarks, and Slidely is working on it
This benchmark is already difficult, but the scores of the top models show why the next version must be harder. Arena scores are relative to this task and this slide creation system. They should not be compared directly with unrelated leaderboards.
Quality is measured only on the 139 cases where every model succeeded. Each model produced one final output per case, so we do not yet measure variation across repeated runs. The density groups contain different cases, which means density trends are observed patterns rather than a controlled experiment. The three judges are also AI models with imperfect agreement, and all three model families appear as both contestants and reviewers.
We are working on harder slide creation cases, more human reviewers, stronger independent judges, and repeated generations that measure consistency. We are also building longer presentation creation tasks that test whether an agent can plan and maintain quality across many slides, much longer runtimes, and higher costs.
Slide creation is only half of presentation work. Our upcoming edit-hard benchmark focuses on difficult slide editing tasks where a model must change an existing slide without breaking the layout, content, or surrounding design. We will publish those results soon.