Which AI model is best for Excel? Introducing AutosheetBench

GPT for Excel lets you choose between several AI models from Anthropic, OpenAI, and other providers through custom endpoints. New models come out all the time, so it's hard to know which one fits your task best.

So we built AutosheetBench, the internal benchmark we announced in our SpreadsheetBench deep dive: a set of 187 "atomic" Excel tasks. We ran six models on it that vary in cost, speed, and capability: Anthropic's Claude Opus 5.5, Claude Fable 5.1, and Claude Opus 4.8, and OpenAI's GPT-6 Astra, GPT-6 Luna, and GPT-5.4. Here's how they compare.

Overall results: Opus 5.5 is the most accurate, GPT-6 Luna the best value

The chart below shows each model's overall score. Each model ran at one reasoning effort level (Low, Medium, or High), which sets how much it thinks before acting. We ran each model at the reasoning effort that gave it the best score-to-cost trade-off in earlier runs. The charts show each model's level next to its name.

Overall AutosheetBench scores for six AI models, led by Opus 5.5 at 94.7%
Chart data as a table
Model
Reasoning effort
Score
Tasks failed (of 187)
Opus 5.5
Medium
94.7%
10
GPT-6 Astra
Low
92.0%
15
Fable 5.1
Medium
92.0%
15
Opus 4.8
High
87.2%
24
GPT-6 Luna
High
87.2%
24
GPT-5.4
Low
81.3%
35

Opus 5.5 (Medium) is the best model overall. It passed 94.7% of the 187 tasks and led almost every category. It missed only 10 tasks, and half of those were missed by every model. GPT-6 Astra and Fable 5.1 came next, with 15 misses each.

The next chart plots each model's score against the credits it uses per task, relative to GPT-6 Luna at High reasoning effort. The chart calls this ratio the Credits used index. It counts only the credits used by the agent model, not by the bulk models the agent hands work to, and assumes built-in models on a subscription. The dashed line connects the models with the best trade-off: every model off the line is matched or beaten by a model that uses fewer credits.

AutosheetBench score versus credits used per task, relative to GPT-6 Luna at High, for the six tested AI models
Chart data as a table
Model
Reasoning effort
Score
Credits used index
Opus 5.5
Medium
94.7%
2×
GPT-6 Astra
Low
92.0%
2.27×
Fable 5.1
Medium
92.0%
2.27×
Opus 4.8
High
87.2%
1.9×
GPT-6 Luna
High
87.2%
1×
GPT-5.4
Low
81.3%
0.71×

GPT-6 Luna is the best lower-cost alternative. It scored 87.2%, with 24 misses, while using about half the credits of Opus 5.5. It works best for quick edits and small, direct tasks, and is weaker on visual tasks like charts and formatting, as we'll see below.

GPT-5.4 is the lowest-cost option. It uses about 30% fewer credits than GPT-6 Luna, but scored 81.3%, with 35 misses.

Cost isn't the only trade-off: speed also matters. The chart below compares each model's score with its average duration per task.

AutosheetBench score versus average duration per task for the six tested AI models
Chart data as a table
Model
Reasoning effort
Score
Average duration per task
Opus 5.5
Medium
94.7%
29s
GPT-6 Astra
Low
92.0%
32s
Fable 5.1
Medium
92.0%
37s
Opus 4.8
High
87.2%
36s
GPT-6 Luna
High
87.2%
41s
GPT-5.4
Low
81.3%
28s

Opus 5.5 combines accuracy with speed. It was one of the fastest models tested, completing a task in about 29 seconds on average. GPT-6 Luna remains cost-effective, but takes longer. It averaged about 41 seconds per task because it tends to iterate more than the other models before finishing.

The latest model releases also cut the errors made by our recommended models dramatically:

  • Best model: We now recommend Opus 5.5 instead of Opus 4.8. Errors fell by nearly 60%, from 24 to 10, at a similar cost per task.
  • Budget model: Moving from GPT-5.4 to GPT-6 Luna cut errors by about 30%, from 35 to 24, for roughly 40% more credits per task.

Opus 5.5 matches or beats Opus 4.8 in every category, for both native spreadsheet and bulk operations. GPT-6 Luna's gains over GPT-5.4 are mostly on native spreadsheet operations, as we'll see later.

The benchmark: 187 everyday Excel tasks

We ran every model in GPT for Excel on the same workbooks, with the same prompts and the same scoring rules, between September 11 and September 26 as new models were released. For each run, we measured whether the model got the task right, how long it took, and how much it cost. The agent uses different tools in Google Sheets, so results there may differ.

The benchmark covers 187 small, self-contained tasks: inserting a column, writing a formula, creating a line chart, formatting a header, and so on. It leaves out end-to-end work such as building or debugging a financial model, or turning a full analysis into a dashboard. We see it as the first step in a continuous effort to improve our product.

The tasks reflect what our agent handles day to day. We built them from real users' spreadsheets and prompts, and split them into two groups:

  • 28 bulk operations apply the same operation to every row of a large dataset, such as translating a value or extracting information from a cell. We used datasets of 1,000 rows and kept the cases simple, with little setup needed beforehand (for example, no glossary to filter before translating).
  • 159 native spreadsheet operations are everyday tasks on the spreadsheet itself.

Every model scored lower on bulk operations than on native ones:

Scores for bulk operations and native spreadsheet operations by AI model
Chart data as a table
Model
Reasoning effort
Native spreadsheet operations
Bulk operations
Opus 5.5
Medium
96%
89%
GPT-6 Astra
Low
92%
89%
Fable 5.1
Medium
93%
86%
Opus 4.8
High
89%
79%
GPT-6 Luna
High
89%
79%
GPT-5.4
Low
82%
79%

Together, the two groups cover eight categories:

Category
Examples
Bulk content generation
Translating text, writing SEO descriptions, summarizing reviews
Bulk data processing
Extracting a salary from a job description, cleaning or standardizing data, classifying the sentiment of reviews
Charts
Creating and updating charts
Data operations
Trimming, find and replace, filling blanks, finding specific cells
Formatting
Conditional formatting, number formats
Formulas
Writing formulas, tracing dependents
Pivot tables
Creating pivot tables
Spreadsheet manipulation
Hiding, deleting, or inserting rows and columns

Each model got one attempt per task. We scored the results in one of two ways:

  • Deterministic check: The output must exactly match a reference file, either the cell values in a given range or the sheet's configuration (filtered rows, hidden columns, an inserted column, and so on).
  • LLM judge: A model reviews the conversation transcript and a set of images (charts, ranges, and so on) against a checklist. We use this for charts, pivot tables, and some formatting tasks, where more than one answer can be correct.

Detailed results

The key takeaway: Opus 5.5 leads almost every category and handles every type of task well.

GPT-6 Luna is a good lower-cost alternative, with clear strengths and weak spots:

  • It's an excellent choice for formulas and pivot tables.
  • It struggles more with visual tasks (charts, formatting) and some spreadsheet manipulation. For example, it tends to make columns too narrow to read.
Category rankings (tied models share a rank)
CategoryOpus 5.5 (Medium)GPT-6 Astra (Low)Fable 5.1 (Medium)Opus 4.8 (High)GPT-6 Luna (High)GPT-5.4 (Low)
Bulk content generation1st2nd2nd4th4th3rd
Bulk data processing2nd1st2nd2nd2nd3rd
Charts1st2nd3rd4th3rd5th
Data operations1st2nd1st2nd3rd4th
Formatting1st2nd2nd3rd4th3rd
Formulas2nd2nd1st2nd1st3rd
Pivot tables1st1st1st1st1st1st
Spreadsheet manipulation1st2nd3rd3rd4th5th

Native spreadsheet operations

On these tasks, the models fall into four groups:

  • Best in class: Opus 5.5 is the clear leader, completing 96% of tasks.
  • Close challengers: GPT-6 Astra and Fable 5.1 follow closely, with strong results overall.
  • Solid performers: Opus 4.8 and GPT-6 Luna trail slightly, mostly on charts, formatting, and spreadsheet manipulation.
  • Lagging behind: GPT-5.4 comes last, well behind the other models.

Here are a few examples that show these differences.

Charts: Building a waterfall chart

Build a waterfall showing the cash-flow bridge.
SucceededOpus 5.5, Fable 5.1
Successful waterfall chart with Ending cash marked as a total
FailedGPT-6 Astra, Opus 4.8, GPT-6 Luna, GPT-5.4
Failed waterfall chart where Ending cash is not marked as a total

Most models built a waterfall chart but didn't mark "Ending cash" as a total. As a result, the chart ended at €1.16M instead of €580K.

Charts: Labeling only the highest bar

Create a bar chart with only the highest value displayed
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8, GPT-6 Luna
Successful chart labeling only the highest bar
FailedGPT-5.4
Failed chart with labels on every bar

GPT-5.4 couldn't show a label on the highest bar alone.

Charts: Labeling only the most recent point

Create a line chart of revenue over time and display a label only on the most recent point.
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8, GPT-6 Luna
Successful line chart labeling only the most recent point
FailedGPT-5.4
Failed line chart with labels on every point

Again, GPT-5.4 couldn't limit the label to a single data point, here the most recent one.

Pivot tables: Creating cumulative revenue

In a new `Pivot` sheet at `A3`, create a pivot table showing cumulative revenue over time.
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8, GPT-6 Luna
Successful native pivot table showing cumulative revenue over time
FailedGPT-5.4
Failed formula-based cumulative revenue summary created instead of a pivot table

GPT-5.4 struggled with Excel's Office.js pivot-table APIs. After several failed attempts to create and configure a native pivot table, it fell back to a formula-based summary styled to resemble one. The values were correct, but the result wasn't a real pivot table, so it lacked native pivot features such as rearranging fields or refreshing the pivot. GPT-5.4 failed this case but was the only model to pass another pivot-table task, which is why the category stays tied in the rankings above.

GPT-5.4 explaining that it created a pivot-style summary after the pivot-table API failed

Formulas: Trace dependents

If I change Financials!H65, which cells will recalculate? Do not edit the workbook; trace both direct and downstream formula dependents and name the affected cells.

Expected answer

  • Direct dependents: Financials!H66, Financials!U65, Financials!V65, and 'Ratio_Analysis '!H30
  • Downstream dependents: 'Ratio_Analysis '!H36, Graphical_representation!P47, and Graphical_representation!P48
SucceededGPT-6 Astra, Fable 5.1, GPT-6 Luna, GPT-5.4
Successful answer listing all direct and downstream dependents of Financials H65
FailedOpus 5.5, Opus 4.8
Failed answer from Opus 4.8 listing only direct dependents on the Financials sheet

Opus 4.8 only inspected dependents on the Financials sheet, so it missed the cross-sheet direct dependent and the full downstream chain. Opus 5.5 correctly counted seven affected cells, but mistyped one downstream reference, which made its answer incorrect.

Data operations: Finding #DIV/0! errors

Are there any #DIV/0! errors on this sheet? List the exact cells.
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8, GPT-6 Luna
Successful answer listing every spreadsheet cell with a DIV/0 error
FailedGPT-5.4
Failed answer missing one spreadsheet cell with a DIV/0 error

GPT-5.4 relied only on the partial preview of the sheet we give the agent, instead of searching the whole sheet. So it missed one of the cells containing #DIV/0!. This shows a broader weakness of GPT-5.4: it tends to go fast rather than be thorough, especially at Low reasoning effort.

Spreadsheet manipulation: Sorting a table with merged cells

Sort the orders by Sales, largest first

Starting sheet

Starting Excel table with merged Region cells
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8
Successfully sorted table with all Region values preserved
FailedGPT-6 Luna, GPT-5.4
Failed table sort with missing Region values

The failing models sorted the table without first unmerging the merged cells. Most rows ended up in the right order, but some Region values were lost along the way. For example, products Charlie and Delta no longer have a region.

Spreadsheet manipulation: Grouping rows

Group the rows by category (alphabetical order) so I can collapse them.

Starting sheet

Starting worksheet before grouping rows by category
SucceededOpus 5.5, GPT-6 Astra
Successful worksheet with one collapsible row group per category
FailedFable 5.1, Opus 4.8, GPT-6 Luna, GPT-5.4
Failed worksheet with one row group spanning every category

This task shows the difference between editing cells and understanding how Excel works. Most models sorted the table and added an outline, but created a single group spanning every row, as shown in the failed example. You can't collapse each category on its own.

Only Opus 5.5 and GPT-6 Astra created a separate, collapsible group for each category.

Spreadsheet manipulation: Freezing panes

Freeze the top row and the first three columns.
SucceededOpus 5.5, GPT-6 Astra, Fable 5.1, Opus 4.8
Successful worksheet with the top row and first three columns frozen
FailedGPT-6 Luna, GPT-5.4
Failed worksheet with two rows and four columns frozen

GPT-6 Luna and GPT-5.4 passed the wrong arguments to the Office.js API. They froze two rows and four columns instead of one row and three columns, which also made the top row and first three columns seem to disappear.

Bulk operations

On bulk operations, the models fall into two clear groups:

  • Best in class: Opus 5.5, GPT-6 Astra, and Fable 5.1
  • Good: Opus 4.8, GPT-6 Luna, and GPT-5.4

What separates the two groups isn't mainly the prompt template they write for each row. It's how they manage the job: checking that every row follows the rules the user gave.

Take this example, where we ask the agent to summarize job descriptions:

Summarize each job description in column C (job_description).

Add a new column with the header job_summary in column F. Write a clear one-line summary of at most 160 characters describing the role's main purpose and responsibilities. Do not copy the full description and fill every row.
  • Succeeded: Opus 5.5, GPT-6 Astra, Fable 5.1, and GPT-6 Luna
  • Failed: Opus 4.8 and GPT-5.4

The prompt clearly asks for summaries of at most 160 characters, and every model includes that limit in its prompt template:

  • Fable 5.1: "Summarize the following job description in one clear line of at most 160 characters, describing the role's main purpose and key responsibilities."
  • Opus 4.8: "Write a clear one-line summary of AT MOST 160 characters describing the role's main purpose and key responsibilities. Do not copy the full description verbatim. Do not exceed 160 characters. Do not include any tags."

But the bulk model doesn't always follow that instruction, because its output isn't fully predictable. The best models check the results and rerun the bulk operation on any summaries that are too long.

Agent checking bulk-generated job summaries against the 160-character limit

That way, every row ends up following the user's rules. The same applies to any task with specific rules, not just character limits. For example: keeping HTML tags intact when translating, or leaving English reviews untouched while translating the others.

What this benchmark taught us about our own product

Some failures were one-off mistakes that don't always happen again on a rerun. But the benchmark also showed us real gaps in our product, especially in the few tasks that every model failed.

Our agent can't see what it produces

The agent can only check that its code ran without errors, not that the result looks right. In the two cases below, every model failed.

Pie chart

Create a pie chart of revenue by year.

The trap here: the years are numbers. Unless the agent first converts them to text, Excel treats the years as a data series. The slices then show the years instead of the revenue, so they're all almost the same size and the legend doesn't show the years.

The agent creates the chart with a script and can't see the result. So it reports success, even though the chart isn't what the user asked for:

"Created a pie chart of revenue by year. Since the raw data had two entries per year, I first built an aggregated summary, then charted it with category names and percentage labels." (Opus 4.8)

Expected pie chart showing revenue by year
Expected result
Failed pie chart treating years as a data series
Actual output (Opus 4.8)

Changing chart label formatting

Another tricky case asks the agent to change the label format on an existing chart:

Make Chart 2 labels whole percentage

Every model changed the number format of the chart's data labels without checking where that format comes from. The labels follow the format of the source data, so changing the format in the chart settings alone doesn't change what the labels show.

Again, every model said it had made the change, because it couldn't see the result and realize its script hadn't worked:

"I changed the data labels on Chart 2 (Technician Work Mix by Pay Type) to show whole percentages, so 47.3% now shows as 47%. Only the chart labels changed." (Opus 5.5)

Expected chart with whole-percentage data labels
Expected result
Actual chart retaining decimal percentage labels
Actual output (Opus 5.5)

The good news: we're working on letting our agent see the spreadsheet in Excel, which would catch both failures.

Our agent still has room to improve on bulk operations

It doesn't always gather enough context before processing rows.

Take a task where we ask the agent to clean an existing priority_raw column:

Create a new priority_clean in column E. Priority should be one of Low, Medium, High, Urgent.

The column contains many variants and typos, but also values like p0, p1, p2, and p3.

To stay efficient, our agent doesn't read the whole spreadsheet at the start. It gets a preview and can explore the rest when needed. Here, it should have listed all the distinct priority values first, but it didn't. So it never saw p0 and didn't know that p0 = Urgent, p1 = High, and so on. It gave the bulk model a mapping shifted by one level (p1 = Urgent), so every row with a pX value got the wrong priority.

Bulk results depend heavily on the bulk model.

For example, GPT-5 nano is very cheap, but its translations and generated content are poor. For this benchmark, we used GPT-5.6 Luna as the bulk model for every configuration. It's a capable model, but it doesn't always follow each prompt to the letter (for example, it sometimes ignores the requested output format, or specific guidelines like the number of characters).

Today, space admins have to set the default bulk model in the GPT for Work dashboard, because our agent can't choose it. We're working on letting the agent pick it.

Next: GPT for Excel vs. the competition

We love benchmarks. They make comparisons fair: what an agent does on a fixed set of tasks, in one attempt, on one day.

This one helped us decide which models to recommend for which tasks in our product. Next, we want to use benchmarks to compare ourselves with competitors, as we started doing in our benchmark of GPT for Excel, Copilot, and Claude, not only on accuracy but also on speed. Spreadsheet work is iterative, so speed directly shapes the user experience.

Full results coming soon!

Related Articles