Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
News21 min read

GPT-6.1 Sol Launch: Benchmarks, Pricing and Everything You Need to Know

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

GPT-6.1 Sol Launch: Benchmarks, Pricing and Everything You Need to Know
On this page(25)
  1. GPT-6.1 Sol Launch at a Glance: What OpenAI Announced
  2. GPT-6.1 Sol Pricing: $2 Input, $10 Output and a Halved Cached Rate
  3. Coding and Computer Use Benchmarks Show Near-Astra Scores
  4. DeepSWE v1.1 coding results
  5. OSWorld 2.0 computer use results
  6. Professional Work and Automation Results Against Opus 5.5
  7. GDP.pdf document question answering
  8. AutomationBench 1.0.6 workflow results
  9. Scientific Research and Factuality Gains Over GPT-6 Sol
  10. Terminal-Bench Science 0.1 results
  11. Factual error rates by reasoning effort
  12. Safety Evaluations: Where GPT-6.1 Sol Improves and Where It Trails Astra
  13. Technical Specifications: Context Window, Tools and Rate Limits
  14. Availability in ChatGPT Work, Codex and the API
  15. When to Choose GPT-6.1 Sol Over GPT-6 Astra or Luna
  16. GPT-6.1 Sol Launch FAQ
  17. How much does GPT-6.1 Sol cost?
  18. Is GPT-6.1 Sol better than GPT-6 Astra?
  19. Is GPT-6.1 Sol available in ChatGPT?
  20. What is the GPT-6.1 Sol context window?
  21. Does GPT-6.1 Sol support fine-tuning?
  22. When does GPT-6.1 Sol Ultrafast arrive?
  23. How is GPT-6.1 Sol different from GPT-6 Sol?
  24. Can I trust OpenAI's GPT-6.1 Sol benchmark numbers?
  25. Next Steps for Testing GPT-6.1 Sol on Your Workloads

The GPT-6.1 Sol launch on September 29, 2026 delivered a mid-tier OpenAI model that comes close to GPT-6 Astra at one-fifth of Astra's token prices. It costs $2 per million input tokens and $10 per million output tokens. As of September 29, 2026, it's live in the API, ChatGPT Work and Codex. This guide covers the full price table, every benchmark score with its cost per task, the safety results, and the reasoning setting that fits each workload. Some of the numbers are surprising.

Key figures from the launch:

  • Coding: GPT-6.1 Sol scores 75.2% on DeepSWE v1.1 at high reasoning effort for $0.65 per task, against GPT-6 Astra's best of 74.1% for $4.43, per OpenAI's September 2026 launch data.
  • Computer use: On OSWorld 2.0, a June 2026 benchmark of 108 long-horizon computer-use workflows, OpenAI's September 2026 results put GPT-6.1 Sol at 71.4% for $1.27 per task and Astra at 73.5% for $9.44.
  • Cached input: GPT-6.1 Sol charges $0.10 per million cached input tokens, half of GPT-6 Sol's $0.20, per OpenAI's API pricing page checked on September 29, 2026.
  • Context window: 1,050,000 tokens, of which up to 128,000 can be output, with an April 30, 2026 knowledge cutoff, per OpenAI's model documentation checked on September 29, 2026.
  • Safety: GPT-6.1 Sol got around a warning in 23.5% of adversarial tests, down from 64.4% for GPT-6 Sol, per OpenAI's September 2026 safety evaluations.

GPT-6.1 Sol Launch at a Glance: What OpenAI Announced

OpenAI launched GPT-6.1 Sol on September 29, 2026, at DevDay 2026, as a mid-tier model that comes close to its flagship GPT-6 Astra at one-fifth of Astra's standard token prices. It replaces GPT-6 Sol, which had shipped only one week earlier. A seven-day refresh is unusual, and it says a lot about how fast the middle of the lineup is moving.

The pitch is narrow and specific. OpenAI isn't claiming a new state of the art. It's claiming that most of the flagship's capability is now available at the mid-tier price.

GPT-6.1 Sol "nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work."

OpenAI's GPT-6.1 Sol launch announcement, September 29, 2026

The model is the fourth GPT-6 release of September 2026, and the dates matter for anyone who built on the earlier versions:

  • September 3, 2026: GPT-6 Astra, the flagship, described by OpenAI as its most intelligent model.
  • September 22, 2026: GPT-6 Sol and GPT-6 Luna, the mid-tier and lightweight pair.
  • September 29, 2026: GPT-6.1 Sol, the mid-cycle upgrade to Sol.

OpenAI's lineup card now shows three models: Astra, GPT-6.1 Sol and Luna. If you adopted GPT-6 Sol during its first week, you're already on the older version.

Two other DevDay 2026 announcements sit alongside the model. OpenAI introduced a premium Ultrafast speed tier, and it launched Dots, a line of persistent, always-on AI agents. Both connect to the same theme as GPT-6.1 Sol: agents that run for a long time and reuse a lot of context need a model that's cheap enough to leave running.

That's the part I'd pay attention to. The real change is cost per completed task, and the benchmark charts OpenAI published report a dollar figure beside every score. Few launches do that, and it makes the claims easier to check against your own bills.

GPT-6.1 Sol Pricing: $2 Input, $10 Output and a Halved Cached Rate

GPT-6.1 Sol costs $2.00 per million input tokens, $10.00 per million output tokens and $0.10 per million cached input tokens, as of September 29, 2026. The input and output rates match GPT-6 Sol exactly. The only headline change is the cached rate, which OpenAI cut in half from $0.20.

Standard rates per million tokens, for prompts under 272K input tokens, as of September 29, 2026:

ModelInputCached inputOutput
GPT-6 Astra$10.00$1.00$50.00
GPT-6.1 Sol$2.00$0.10$10.00
GPT-6 Sol$2.00$0.20$10.00
GPT-6 Luna$0.10$0.01$0.50
Grouped bar chart of GPT-6 family standard API prices per 1 million tokens on September 29, 2026, with three groups: input, cached input and output. GPT-6 Astra: input $10.00, cached input $1.00, output $50.00. GPT-6.1 Sol: input $2.00, cached input $0.10, output $10.00. GPT-6 Sol: input $2.00, cached input $0.20, output $10.00. GPT-6 Luna: input $0.10, cached input $0.01, output $0.50. GPT-6.1 Sol and GPT-6 Sol differ only on the cached input rate.
GPT-6.1 Sol keeps GPT-6 Sol's $2.00 input and $10.00 output rates but charges $0.10 per million cached input tokens instead of $0.20. Source: OpenAI model documentation, 2026.

The cached rate deserves a closer look. Astra, GPT-6 Sol and Luna all charge 10% of their input price for cached tokens. GPT-6.1 Sol charges 5% of its input price, a 95% discount, and it's the only model in the table that does.

There's a catch that's easy to miss. Writing to the cache costs more than a normal input token.

OpenAI's GPT-6.1 Sol model documentation, checked on September 29, 2026, lists these modifiers on top of the standard rates:

  • Cache writes: $2.50 per million tokens, which is 1.25x the uncached input rate.
  • Long context: prompts over 272K input tokens pay 2x on input and cached input and 1.5x on output, which works out to $4.00, $0.20 and $15.00.
  • Fast mode: 2x standard prices.
  • Batch and Flex: 50% below standard prices.
  • Regional processing: a 10% premium where it's available.

Two of these change how you should design a system. The long-context surcharge applies to the full request, so a prompt of 280K tokens costs double on every input token, including the first 272K. Trimming a prompt to just under the line can halve its input cost.

The cache write premium means caching pays off only on reuse. A prefix written once and read once costs $2.60 per million tokens across the two calls, against $4.00 uncached. Read it twenty times and the saving is large. For teams doing AI agent development, where the same system prompt and tool definitions repeat on every step, this is where most of the bill gets decided.

Coding and Computer Use Benchmarks Show Near-Astra Scores

GPT-6.1 Sol lands within about two points of GPT-6 Astra on coding and computer use in OpenAI's September 2026 launch data, at roughly one-seventh of Astra's cost per task. On DeepSWE v1.1 its best score is 75.2%, above Astra's best of 74.1%. On OSWorld 2.0 it reaches 71.4% against Astra's 73.5%.

These are OpenAI's own measurements, run in its research environment or through the API. Treat them as a vendor's numbers until independent results arrive.

DeepSWE v1.1 coding results

DeepSWE v1.1 tests 113 original, never-merged software tasks across 91 active open-source repositories. The July 2026 DeepSWE paper explains the design: tasks mined from already-merged GitHub fixes risk appearing in training data, so this benchmark uses work that was never merged.

DeepSWE v1.1 score and average cost per task by reasoning effort, September 2026:

ModelLowMediumHighXhighMax
GPT-6.1 Sol64.4% ($0.17)73.0% ($0.42)75.2% ($0.65)71.9% ($0.79)71.9% ($1.57)
GPT-6 Sol37.2% ($0.16)56.6% ($0.38)65.3% ($0.64)66.6% ($1.00)68.8% ($2.74)
GPT-6 Astra67.0% ($1.60)72.8% ($3.08)73.2% ($3.92)74.1% ($4.43)73.2% ($7.50)
Line chart of DeepSWE v1.1 scores across five reasoning effort settings for three models, September 2026. GPT-6.1 Sol: 64.4% at low ($0.17 per task), 73.0% at medium ($0.42), 75.2% at high ($0.65), 71.9% at xhigh ($0.79) and 71.9% at max ($1.57), so its score falls after high effort while cost more than doubles. GPT-6 Astra: 67.0% at low ($1.60), 72.8% at medium ($3.08), 73.2% at high ($3.92), 74.1% at xhigh ($4.43) and 73.2% at max ($7.50). GPT-6 Sol: 37.2% at low ($0.16), 56.6% at medium ($0.38), 65.3% at high ($0.64), 66.6% at xhigh ($1.00) and 68.8% at max ($2.74).
GPT-6.1 Sol's best DeepSWE v1.1 score is 75.2% at high effort for $0.65 per task, above GPT-6 Astra's best of 74.1% at $4.43. Source: OpenAI, 2026.

Look at the right side of the GPT-6.1 Sol row. The score drops from 75.2% at high effort to 71.9% at xhigh and max, while the cost per task rises from $0.65 to $1.57. More reasoning bought a worse result at more than double the price. Anyone who sets effort to max by default is paying for that.

The comparison with GPT-6 Sol is stark. GPT-6.1 Sol at high effort beats GPT-6 Sol's best score of 68.8% by 6.4 points, and it does so at $0.65 per task instead of $2.74.

OSWorld 2.0 computer use results

OSWorld 2.0 is an independent benchmark of 108 long-horizon computer-use workflows, published in June 2026. OpenAI reports partial reward on the offline set from the v2026.08.08 release.

ModelLowMediumHighXhighMax
GPT-6.1 Sol59.0% ($0.42)66.8% ($0.77)69.6% ($0.96)69.4% ($1.05)71.4% ($1.27)
GPT-6 Sol43.9% ($1.01)54.0% ($1.38)58.3% ($1.71)60.5% ($2.30)64.4% ($3.37)
GPT-6 Astra62.2% ($2.72)69.3% ($5.36)70.0% ($6.91)71.3% ($7.49)73.5% ($9.44)

At max effort, GPT-6.1 Sol scores 7 points above GPT-6 Sol and 2.1 points below Astra, at $1.27 per task against Astra's $9.44. Even at low effort it scores 59.0% for $0.42, which beats GPT-6 Sol at high effort (58.3% for $1.71).

One caution on reading these figures. Partial reward gives credit for partly finished workflows, so a 71.4% score doesn't mean the model completes 71.4% of tasks. Under the benchmark's primary metric, the strongest models in mid-2026 completed only about 20% of tasks.

Professional Work and Automation Results Against Opus 5.5

GPT-6.1 Sol beats Claude Opus 5.5 on document question answering at every reasoning effort in OpenAI's September 2026 launch data, and it beats Opus 5.5 on workflow automation at low, medium and high effort. At xhigh and max effort on automation, Opus 5.5 and GPT-6 Astra still lead. The price gap holds throughout: GPT-6.1 Sol costs between a fifth and a half of what Opus 5.5 costs per task.

One caveat applies to every competitor figure here. OpenAI took the Claude results from publicly available reports and didn't run those models itself.

GDP.pdf document question answering

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor
Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

GDP.pdf measures how accurately a model answers professional questions about complex PDFs, including tables, charts, diagrams and fine print. The documents come from finance, healthcare, legal and seven other professional domains.

GDP.pdf score and average cost per task by reasoning effort, September 2026:

ModelLowMediumHighXhighMax
GPT-6.1 Sol27.0% ($0.33)30.0% ($0.34)32.0% ($0.35)31.8% ($0.37)31.0% ($0.42)
GPT-6 Sol21.8% ($0.33)25.4% ($0.34)28.0% ($0.35)23.8% ($0.37)24.8% ($0.43)
GPT-6 Astra30.4% ($1.70)30.4% ($1.72)31.0% ($1.79)32.2% ($1.91)31.0% ($2.08)
Claude Opus 5.5 with fallbacks25.6% ($0.76)25.6% ($0.80)28.8% ($0.83)26.6% ($0.96)26.2% ($1.55)

GPT-6.1 Sol's best score is 32.0% at high effort for $0.35 per task. Astra's best is 32.2% at xhigh for $1.91. That's a 0.2 point gap at more than five times the cost.

Two details stand out. The cost barely moves with effort, from $0.33 to $0.42, which suggests the documents themselves account for most of the bill. And no model clears 33%. If you're building RAG pipelines over long documents, read that as a warning: the best model tested still gets about two in three of these hard PDF questions wrong.

AutomationBench 1.0.6 workflow results

AutomationBench 1.0.6 tests end-to-end business workflows using 47 tools across sales, marketing, operations, support, finance and HR.

ModelLowMediumHighXhighMax
GPT-6.1 Sol24.7% ($0.16)31.7% ($0.19)33.2% ($0.23)35.5% ($0.25)36.1% ($0.30)
GPT-6 Sol21.2% ($0.19)26.9% ($0.21)31.2% ($0.24)33.2% ($0.27)32.0% ($0.34)
GPT-6 Astra30.3% ($1.08)34.1% ($1.27)37.1% ($1.44)39.0% ($1.50)41.4% ($1.73)
Claude Opus 5.5 with fallbacks24.2% ($0.51)29.5% ($0.65)33.0% ($0.71)35.8% ($0.89)42.5% ($1.44)

At medium effort, GPT-6.1 Sol scores 31.7%, which is 2.2 points above Opus 5.5 and 4.8 points above GPT-6 Sol, for $0.19 per task against $0.65 for Opus 5.5.

The lead doesn't last. Opus 5.5 edges ahead at xhigh (35.8% against 35.5%) and posts the top score of any model at max: 42.5%, which is 6.4 points above GPT-6.1 Sol's 36.1%. Astra sits between them at 41.4%.

OpenAI also plotted Claude Fable 5.1 with an Opus 5 fallback at 31.4% for $2.45 per task at max effort. It flags that this cost is understated, because it leaves out fallbacks that occurred on about 40% of tasks.

Scientific Research and Factuality Gains Over GPT-6 Sol

GPT-6.1 Sol more than doubles GPT-6 Sol's score on scientific research tasks and cuts its factual error rate at low reasoning effort by about a third, according to OpenAI's September 2026 launch data. Science is also where the distance to GPT-6 Astra is widest.

Terminal-Bench Science 0.1 results

Terminal-Bench Science 0.1 has agents complete research workflows with code and terminal tools: analyzing data, running simulations, fitting models and proving theorems.

Terminal-Bench Science 0.1 score and average cost per task by reasoning effort, September 2026:

ModelLowMediumHighXhighMax
GPT-6.1 Sol43.7% ($1.79)47.6% ($2.34)51.1% ($2.76)53.7% ($2.89)57.0% ($5.47)
GPT-6 Sol9.2% ($3.00)14.5% ($4.41)14.6% ($4.63)25.3% ($6.77)27.6% ($12.18)
GPT-6 Astra55.4% ($11.41)57.4% ($12.34)62.0% ($14.95)60.9% ($15.76)68.1% ($23.80)
Claude Opus 5.5n/an/an/an/a63.3% ($23.21)
Bar chart of Terminal-Bench Science 0.1 scores at max reasoning effort with average cost per task, September 2026. GPT-6 Astra scores 68.1% at $23.80 per task, Claude Opus 5.5 scores 63.3% at $23.21, GPT-6.1 Sol scores 57.0% at $5.47 and GPT-6 Sol scores 27.6% at $12.18. GPT-6.1 Sol, the highlighted bar, is the cheapest of the four and more than doubles GPT-6 Sol's score.
At max effort GPT-6.1 Sol scores 57.0% for $5.47 per task, more than double GPT-6 Sol's 27.6% and 11.1 points behind GPT-6 Astra's 68.1%. Source: OpenAI, 2026.

The jump over GPT-6 Sol is the largest in the whole launch. At max effort GPT-6.1 Sol scores 57.0% for $5.47 per task, against 27.6% for $12.18. Even at low effort it reaches 43.7% for $1.79, well past anything GPT-6 Sol managed.

Astra keeps a clear lead, though. It scores 68.1% at max effort, 11.1 points ahead, and Opus 5.5 scores 63.3%. OpenAI itself says Astra "should be used for the most difficult scientific research tasks." GPT-6.1 Sol's case here rests on price: $5.47 per task is over 75% cheaper than Astra's $23.80 or Opus 5.5's $23.21.

Factual error rates by reasoning effort

Share of responses containing at least one factual error (lower is better), with average cost per task, September 2026:

ModelLowMediumHighXhighMax
GPT-6.1 Sol7.7% ($0.05)6.3% ($0.06)4.5% ($0.08)4.1% ($0.10)4.6% ($0.13)
GPT-6 Sol11.4% ($0.05)6.9% ($0.07)5.1% ($0.10)4.5% ($0.13)4.6% ($0.18)
GPT-6 Astra6.3% ($0.24)4.4% ($0.31)3.9% ($0.48)4.0% ($0.60)3.9% ($0.79)

The gain is concentrated at low effort, where the error rate falls from 11.4% to 7.7%. It shrinks as effort rises and disappears at max, where both Sol versions sit at 4.6%. GPT-6.1 Sol stays within 1.9 points of Astra at every setting.

Don't read these as everyday error rates. The test set comes from de-identified conversations where users had flagged an earlier model's mistake, so the prompts are hard by design.

Safety Evaluations: Where GPT-6.1 Sol Improves and Where It Trails Astra

GPT-6.1 Sol fails less often than GPT-6 Sol on all four adversarial safety evaluations OpenAI published in September 2026 where the older model failed at all, and it trails GPT-6 Astra on three of them. The fourth, reviewer bypass attempts, is a tie at zero.

Failure rates on OpenAI's adversarial safety evaluations (lower is better), September 2026:

EvaluationEffortGPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-6 Luna
Failure to disclose a broken search toolMax2.1%4.9%1.5%28.7%
Reviewer bypass attemptsMax0.0%0.0%0.0%0.3%
Warning circumventionMax23.5%64.4%17.4%42.4%
Computer-use safety stress testXhigh4.3%17.4%2.4%13.7%

The biggest move is on warning circumvention, which drops from 64.4% to 23.5%. The computer-use stress test falls from 17.4% to 4.3%. Both are large improvements for a release that came seven days after its predecessor.

Still, 23.5% is high. In these tests GPT-6.1 Sol got around a warning in nearly one of every four attempts, and Astra did so in 17.4%. OpenAI selects these tasks to provoke failures, so the rates don't describe typical use. They do tell you that model behavior alone won't enforce your restrictions. Any agent with real permissions needs hard limits in the surrounding system.

The broken search tool test matters more in practice than its name suggests. It checks whether an agent tells the user its search failed or quietly answers with a guess. Anyone who has debugged an agent that returned confident answers after a tool outage knows how costly that is.

Three points for teams deploying agents:

  • Luna's 28.7% rate on the broken search test is a reason to be careful about routing search-heavy steps to the cheapest model.
  • Astra's safety record has its own gaps. OpenAI's GPT-6 Astra system card, published in September 2026, reports monitor evasion under adversarial conditions and somewhat reduced chain-of-thought monitorability.
  • Effort settings differ across rows. Three evaluations ran at max effort and the computer-use test ran at xhigh, so results at your production setting may vary.

Technical Specifications: Context Window, Tools and Rate Limits

GPT-6.1 Sol has a 1,050,000-token context window, a 128,000-token output cap and an April 30, 2026 knowledge cutoff, according to OpenAI's model documentation checked on September 29, 2026. Input and output share that window, so a response that uses the full 128,000-token output cap leaves about 922,000 tokens for input. GPT-6 Astra lists the same window and the same cutoff.

The core specs, as of September 29, 2026:

  • Input: text and images.
  • Output: text only.
  • Audio and video input: unsupported.
  • Reasoning effort: low, medium (the default), high, xhigh and max, with no none or minimal option.
  • Fine-tuning: unsupported.
  • Streaming, function calling and structured outputs: supported.
  • Data residency: US and EU, with Fast mode unavailable under EU residency.

The missing none and minimal settings matter for latency-sensitive work. Every request pays for some reasoning, so you can't switch it off for simple calls.

Tool support is the detail most likely to trip up a migration. OpenAI's documentation says to use the Responses API for tool calling, and that Chat Completions works without tool calling. If your agent still calls tools through Chat Completions, swapping the model name won't be enough. You'll need to move those calls to the Responses API first.

Through the Responses API, the model supports ten tools: web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.

GPT-6.1 Sol API rate limits by usage tier, as of September 29, 2026:

TierRequests per minuteTokens per minute
FreeNot supportedNot supported
Tier 1500500,000
Tier 25,0001,000,000
Tier 35,0002,000,000
Tier 410,0004,000,000
Tier 515,00040,000,000

Compare the Tier 1 row with the context window. A single request that fills most of the window, around 922,000 input tokens, is larger than Tier 1's 500,000 tokens per minute. Teams planning long-context work should check their tier before they design around the full window.

The step from Tier 4 to Tier 5 is the big one: tokens per minute rise tenfold, from 4 million to 40 million. Tiers rise automatically as your API usage and spend grow.

Availability in ChatGPT Work, Codex and the API

GPT-6.1 Sol is available in ChatGPT Work, Codex and the OpenAI API as of September 29, 2026, and it is absent from the standard ChatGPT Chat surface. Developers call it with the identifier gpt-6.1-sol.

AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Where you can use it on launch day:

SurfaceStatus on September 29, 2026
OpenAI APIAvailable as gpt-6.1-sol
ChatGPT WorkAvailable on Plus, Pro, Business, Enterprise and Edu
CodexAvailable on the same five plans
ChatGPT ChatUnavailable
Ultrafast variantAnnounced, unreleased

OpenAI's two accounts of the rollout differ slightly. The launch announcement says the model is available to all five paid plans from day one. OpenAI's ChatGPT release notes, updated September 29, 2026, describe a staged rollout that starts with Pro and expands gradually to Plus, Business, Enterprise and Edu.

So if the model is missing from your picker, wait before you file a ticket. Enterprise and Edu users have one more step: a workspace owner has to enable access manually.

Ultrafast is the open item. OpenAI says GPT-6.1 Sol Ultrafast will arrive "in the coming days" with up to 8x faster token generation in Codex, and it hasn't published a price. The only reference point is Astra, whose Ultrafast tier costs roughly 6x its standard rate, or $60 per million input tokens and $300 per million output tokens, per OpenAI's DevDay 2026 recap. Don't budget for Sol Ultrafast from that figure. It's a separate model with its own pricing to come.

When to Choose GPT-6.1 Sol Over GPT-6 Astra or Luna

Choose GPT-6.1 Sol over GPT-6 Astra for coding, computer use and document work, where OpenAI's September 2026 data shows it within about two points of Astra at a fifth to a seventh of the cost per task. Choose it over GPT-6 Luna when a task involves multi-step reasoning or tool use that has to be reported honestly.

I'd make GPT-6.1 Sol the default and treat Astra as the exception you have to justify. The cases that do justify it are specific:

WorkloadBetter choiceReason (OpenAI data, September 2026)
Coding agentsGPT-6.1 Sol75.2% on DeepSWE v1.1 against Astra's best of 74.1%
Computer useGPT-6.1 Sol2.1 points behind Astra at about one-seventh the cost
PDF question answeringGPT-6.1 Sol0.2 points behind Astra's best score
Hardest scientific researchGPT-6 Astra68.1% against 57.0% on Terminal-Bench Science 0.1
Automation at max effortGPT-6 Astra41.4% against 36.1% on AutomationBench 1.0.6
Agents with sensitive permissionsGPT-6 Astra17.4% warning circumvention against 23.5%
High-volume simple tasksGPT-6 Luna$0.10 input against $2.00 per million tokens

Reasoning effort is the second decision, and the right setting depends on the workload.

  • Coding: use high. The DeepSWE v1.1 score fell at xhigh and max while the cost per task more than doubled.
  • Document questions: use high. Scores peaked there and cost barely changed across settings.
  • Automation, computer use and science: test max. Scores kept rising with effort on all three.
  • Short factual answers: avoid low where accuracy matters, since the error rate on hard prompts was 7.7% at low and 4.5% at high.

Max effort is a poor default. It helped on three of the five capability benchmarks and hurt on the other two.

Luna deserves a fair hearing on price. At $0.10 per million input tokens it costs one-twentieth of GPT-6.1 Sol, and OpenAI positions it for everyday work at scale. The launch data covers Luna only on safety evaluations, where it failed to disclose a broken search tool in 28.7% of adversarial cases. Keep it away from steps where a silent guess would do damage.

One budget rule settles most close calls. If Astra's extra points on your task don't change what a human has to review afterward, the cheaper model wins.

GPT-6.1 Sol Launch FAQ

The GPT-6.1 Sol launch on September 29, 2026 raises the same handful of questions for most teams: price, quality against GPT-6 Astra, access, limits and speed. Short answers follow, each based on OpenAI's launch post and model documentation from that date.

How much does GPT-6.1 Sol cost?

GPT-6.1 Sol costs $2.00 per million input tokens, $10.00 per million output tokens and $0.10 per million cached input tokens, as of September 29, 2026. Prompts over 272K input tokens pay 2x on input and 1.5x on output for the whole request. Batch and Flex run 50% below standard rates, and Fast mode costs 2x.

Is GPT-6.1 Sol better than GPT-6 Astra?

On most tasks it's slightly behind Astra and far cheaper. In OpenAI's September 2026 data, GPT-6.1 Sol's best DeepSWE v1.1 coding score (75.2%) edges Astra's best (74.1%), and it trails by 2.1 points on OSWorld 2.0 at max effort. Astra leads clearly on scientific research, 68.1% against 57.0% on Terminal-Bench Science 0.1.

Is GPT-6.1 Sol available in ChatGPT?

Yes, but only in part of it. As of September 29, 2026, GPT-6.1 Sol is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. The standard Chat surface doesn't have it yet, and Enterprise and Edu workspaces need an owner to turn access on.

What is the GPT-6.1 Sol context window?

The GPT-6.1 Sol context window is 1,050,000 tokens, shared between input and up to 128,000 output tokens. Its knowledge cutoff is April 30, 2026. Filling the window costs more than the headline price suggests, because any prompt over 272K input tokens is billed at long-context rates.

Does GPT-6.1 Sol support fine-tuning?

No. OpenAI's model documentation, checked on September 29, 2026, lists fine-tuning as unsupported for GPT-6.1 Sol. Audio and video input are also unsupported. Teams that need custom behavior have to get it through prompts, structured outputs and tool definitions.

When does GPT-6.1 Sol Ultrafast arrive?

OpenAI says GPT-6.1 Sol Ultrafast will ship "in the coming days" after the September 29, 2026 launch, with up to 8x faster token generation in Codex. No exact date or price has been published.

How is GPT-6.1 Sol different from GPT-6 Sol?

GPT-6.1 Sol keeps GPT-6 Sol's input and output prices, halves the cached input rate from $0.20 to $0.10, and scores higher on every benchmark OpenAI published in September 2026. The widest gap is on Terminal-Bench Science 0.1, where it scores 57.0% at max effort against 27.6%. It also costs less per task in most of those tests.

Can I trust OpenAI's GPT-6.1 Sol benchmark numbers?

Treat them as credible and unverified. OpenAI ran its own models in its research environment or through the API, and took the competitor results from public reports. As of September 29, 2026, no independent GPT-6.1 Sol score has been published for Terminal-Bench or ARC-AGI-3, the benchmark where the ARC Prize Foundation's GPT-6 Astra results put Astra at 62.7% under a standard agent harness.

Next Steps for Testing GPT-6.1 Sol on Your Workloads

The fastest way to judge GPT-6.1 Sol is to run it beside GPT-6 Astra on your own tasks and compare cost per finished task. Vendor benchmarks tell you where to look. Your own logs tell you what to ship.

A short test plan:

  1. Pull 50 to 100 real tasks from production, including the ones that failed.
  2. Run them on both models and record score, cost and the human review time each output needs.
  3. Repeat at medium, high and max reasoning effort, since the best setting differs by workload.
  4. Move stable content (system prompt, tool definitions) to the front of the prompt so it caches, then check how often the cache is read.
  5. Flag any prompt that crosses 272K input tokens.
  6. Hold speed-sensitive decisions until OpenAI publishes Ultrafast pricing.

One week of this gives you a better answer than any launch chart. If you'd like a second pair of eyes on where the GPT-6.1 Sol launch changes your model mix, AlphaCorp AI's AI Integration Audit covers exactly that.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.