Skip to main content

Recommended models

Find the recommended built-in models for GPT for Sheets and GPT for Excel, and choose the one that best fits your use case and budget. For more information about the full range of supported models, see AI providers and models supported by GPT for Work.

Agent models​

BestBalancedCheapest

Fable 5.1 (Medium)

GPT-5.6 Luna (High)

GPT-5.4 (Low)

For the benchmark results behind these recommendations, see Agent model benchmark.

Examples​

  • For simple or straightforward tasks such as creating formulas or fixing errors, use GPT-5.6 Luna at High reasoning effort.

  • For high-level tasks such as financial modeling, overall formatting, or visual tasks, use Fable 5.1 at Medium reasoning effort.

Bulk models​

The same bulk models are recommended whether the Agent delegates a bulk task, or you run bulk tools or GPT functions yourself.

Use caseBestBalancedCheapest

Content generation, translation, categorization, extraction, scoring

gpt-5.6-sol

gpt-5.4-mini

gpt-5.4-nano

Enrichment from the web

gpt-5-search-api

perplexity-low

perplexity-fast

Agent model benchmark​

To compare Agent models, we run an internal benchmark of roughly 200 everyday spreadsheet tasks, such as writing a formula, inserting a column, creating a chart, or formatting a header. Each model runs every task once with the same spreadsheets and prompts, and each result is scored as a pass or a fail. Newly released models are always evaluated in real-world use before we recommend them, even when they score well in the benchmark.

The latest run is from September 2026, in which six models competed in GPT for Excel:

  • Fable 5.1 at Medium reasoning effort achieved the best overall performance and ranked first or tied for first in half of the task categories.

  • GPT-5.6 Luna at High reasoning effort offered the best balance of score and cost. It scored close to Fable 5.1 while consuming about half the credits per task.

note

The benchmark ran in GPT for Excel only. The Agent uses different tools in Google Sheets, so results in Sheets may differ.

Overall scores​

The following figure shows the share of tasks each model completed successfully.

70%80%90%100%Fable 5.1 (Medium): 90.9%90.9%Fable 5.1 (Medium)GPT-6 Astra (Low): 89.8%89.8%GPT-6 Astra (Low)GPT-5.6 Luna (High): 86.6%86.6%GPT-5.6 Luna (High)GPT-5.6 Sol (Medium): 86.6%86.6%GPT-5.6 Sol (Medium)Opus 4.8 (High): 85.6%85.6%Opus 4.8 (High)GPT-5.4 (Low): 79.7%79.7%GPT-5.4 (Low)

Scores vs. credit consumption​

The following figure shows each model's score against its credit consumption per task, relative to GPT-5.6 Luna at High reasoning effort. Credit consumption covers the Agent model only, not the bulk models the Agent delegates to, and assumes built-in models on a subscription. For example, Fable 5.1 consumes about twice as many credits per task as GPT-5.6 Luna. The dashed line connects the models that offer the best trade-off: Every model off the line is matched or outscored by a model that consumes fewer credits.

70%80%90%100%0×0.5×1×1.5×2×CREDIT CONSUMPTION (× GPT-5.6 LUNA (HIGH))SCORE (%)Fable 5.1 (Medium): 91%, 2.14×Fable 5.1 (Medium)GPT-6 Astra (Low): 90%, 2.15×GPT-6 Astra (Low)GPT-5.6 Luna (High): 87%, 1×GPT-5.6 Luna (High)GPT-5.6 Sol (Medium): 87%, 1.85×GPT-5.6 Sol (Medium)Opus 4.8 (High): 86%, 1.79×Opus 4.8 (High)GPT-5.4 (Low): 80%, 0.67×GPT-5.4 (Low)

Scores vs. speed​

The following figure shows each model's score against the average time it took to complete a task. Shorter times are better. The dashed line connects the models that offer the best trade-off: Every model off the line is matched or outscored by a faster model.

70%80%90%100%35s40s45s50sTIME PER TASK (SECONDS)SCORE (%)Fable 5.1 (Medium): 91%, 50sFable 5.1 (Medium)GPT-6 Astra (Low): 90%, 47sGPT-6 Astra (Low)GPT-5.6 Luna (High): 87%, 43sGPT-5.6 Luna (High)GPT-5.6 Sol (Medium): 87%, 46sGPT-5.6 Sol (Medium)Opus 4.8 (High): 86%, 44sOpus 4.8 (High)GPT-5.4 (Low): 80%, 35sGPT-5.4 (Low)

Rankings by task category​

The following table shows how the models rank in each task category. Models with the same score in a category share a rank.

Task categoryFable 5.1MediumGPT-6 AstraLowGPT-5.6 LunaHighGPT-5.6 SolMediumOpus 4.8HighGPT-5.4Low
Bulk content generation1st2nd1st2nd4th3rd
Bulk data processing2nd2nd1st2nd2nd3rd
Charts2nd1st4th4th3rd5th
Data operations1st2nd3rd2nd2nd4th
Formatting1st1st3rd3rd2nd2nd
Formulas1st2nd1st1st2nd3rd
Pivot tables2nd2nd1st1st2nd2nd
Spreadsheet manipulation2nd1st3rd2nd2nd4th

Learn more​