endueendue

GPT-6.1 Sol and Ultrafast: near-Astra results at a fifth of the price

OpenAI launched GPT-6.1 Sol at DevDay. It costs the same as GPT-6 Sol and scores close to GPT-6 Astra on coding, computer use and document work. What changed, what Ultrafast adds, and which model to use when.

Streaks of light converge from the left onto a bright cyan circle, with a chevron pointing onward beside it.

OpenAI launched GPT-6.1 Sol at DevDay on September 29, exactly one week after GPT-6 Sol. The pitch fits in one sentence: close to GPT-6 Astra on agentic coding, computer use and professional work, at one fifth of Astra’s standard token prices.

The price list is the interesting part. Input and output cost exactly what GPT-6 Sol costs, and only cached input got cheaper, by half. In practice you get a better model for the same money.

Ultrafast, announced the same day, goes the other way: you pay more and buy speed. Below are both launches in turn, followed by a short guide to which model to use when.

At a glance

Item Details
Model ID gpt-6.1-sol
API price (per 1M tokens) $2 input, $0.10 cached input, $10 output
Context 1.05M tokens (922K max input, 128K max output)
Where The API, plus Work and Codex in ChatGPT (Plus, Pro, Business, Enterprise, Edu). Not in regular Chat yet
Ultrafast GPT-6 Astra today; GPT-6.1 Sol in Codex “in the coming days”

What improved over GPT-6 Sol

Every number below comes from OpenAI’s own evaluations, published in the announcement.

Evaluation What OpenAI reports
DeepSWE v1.1 (engineering in real codebases) Matches GPT-6 Astra at about a fifth of the cost; 6.4 points above GPT-6 Sol’s best score
OSWorld 2.0 offline set (computer use) 7 points above GPT-6 Sol at max effort for less than half the cost; within 2.1 points of Astra at about a seventh of the cost per task
GDP.pdf (answering from complex PDFs) Above Opus 5.5 (with fallbacks) at less than half the cost per task; close to Astra at about a fifth of the cost
AutomationBench (multi-step business workflows) 2.2 points above Opus 5.5 at medium effort for about a third of the cost; up 4.8 points on GPT-6 Sol
Terminal-Bench Science 0.1 (scientific work) More than double GPT-6 Sol’s score at max effort; $5.47 per task on average versus $23.21 for Opus 5.5 and $23.80 for Astra
Factual errors At low effort, responses with a factual error drop from 11.4% to 7.7%, about 32% fewer

Computer use is the line to watch. Agents that read screens and operate apps tend to be expensive per task, so landing within about two points of Astra at roughly a seventh of the cost is the most practical news here for teams running them.

It does not catch Astra everywhere. OpenAI notes that Astra still holds the top science score (68.1%) and recommends it for the hardest research tasks.

OpenAI also published safety results. In a test of whether an agent admits that its search tool is broken instead of guessing, the failure rate fell from 4.9% for GPT-6 Sol to 2.1% for GPT-6.1 Sol, against 1.5% for Astra and 28.7% for Luna. The tasks are chosen to provoke failures, so these are not everyday rates.

Independent numbers are not in yet

On launch day there was no independent measurement of GPT-6.1 Sol from evaluators such as Artificial Analysis. The New Stack read OpenAI’s charts and compared it with Claude Sonnet 5.5, which costs the same $2/$10, and found mixed results: about 75% for GPT-6.1 Sol against 71% for Sonnet 5.5 on DeepSWE, but about 36% against 44.7% on AutomationBench, where GPT-6.1 Sol cost $0.30 per task to Sonnet’s $1.14.

Same price, half-price cache

Model Input Cached input Output
GPT-6 Astra $10 $1.00 $50
GPT-6.1 Sol $2 $0.10 $10
GPT-6 Sol $2 $0.20 $10
GPT-6 Luna $0.10 $0.01 $0.50

Standard prices per million tokens for prompts up to 272K input tokens. Longer prompts pay 2x input and cache rates and 1.5x output on the whole request. Batch and Flex are half price.

The cheaper cache matters most for agents, which resend the same context at every step: the system prompt, tool definitions, a codebase. Take one step with 100,000 input tokens, 90,000 of them cached, and 2,000 output tokens:

Model Cost of that step (example)
GPT-6 Astra $0.290
GPT-6 Sol $0.058
GPT-6.1 Sol $0.049

That is about 16% less than GPT-6 Sol on an identical price list, and roughly a sixth of Astra. Your cache hit rate and output length will move these numbers, so take the method rather than the totals.

What Ultrafast is

Ultrafast is OpenAI’s fastest service tier. The model stays the same; responses are generated faster and billed higher.

  • Speed: up to 8x faster token generation in Codex (300 tokens per second) and up to 6x in the API. OpenAI’s docs point out that this measures token generation, not how long a whole task takes.
  • Models today: GPT-6 Astra. In the API it is open to everyone at low rate limits (500K tokens per minute on tiers 1 to 3, 1M on tier 4, 5M on tier 5). GPT-6.1 Sol Ultrafast is due in Codex in the coming days.
  • API price: Astra on Ultrafast costs $60 input, $6 cached input and $300 output per million tokens, six times Standard.
  • In ChatGPT: available in Codex and ChatGPT Work on Pro 500 and eligible Enterprise and Edu plans. It draws down included usage at 8x the Standard rate, and purchased credits and Enterprise pay-as-you-go are billed at 6x. Other self-serve plans cannot use it at launch, even with purchased credits.
  • Regions: US data residency and global processing only. Organizations that need processing in the EU or other regions cannot use it.

In the API it is one field on the request:

from openai import OpenAI

client = OpenAI()
response = client.responses.create(
    model="gpt-6-astra",
    service_tier="ultrafast",
    input="Read these incident logs and name the most likely cause.",
)
print(response.output_text)

For agents that call tools in quick succession, OpenAI strongly recommends a WebSocket connection. Opening a new connection per request lets network overhead eat into the speed gain.

Ultrafast first appeared on August 13 as a limited preview with GPT-5.6 Sol, running on Cerebras hardware at up to 750 output tokens per second, up to 14x Standard. OpenAI pointed to incident response, financial research and live customer support as the kinds of work where seconds change the outcome.

Which model to use when

  • GPT-6.1 Sol: the new default candidate. Try it first on coding agents, computer use and document-heavy work. It costs the same as GPT-6 Sol, so there is little reason not to test it, and OpenAI’s own model page suggests comparing it with Astra on your tasks.
  • GPT-6 Astra: keep it for the hardest work, where mistakes are expensive. It still holds the top science score.
  • GPT-6 Luna: high-volume, simple jobs such as classification and extraction. Its input price is a twentieth of GPT-6.1 Sol’s.
  • Ultrafast: only where speed is the value, such as interactive work with a person waiting or incident response. It is the same model at six times the price.

GPT-6.1 Sol takes reasoning effort low, medium (default), high, xhigh or max; none and minimal are not supported. Tool calling requires the Responses API.

Do not choose on token price alone. The cost of a task depends heavily on how many tokens a model spends and at what reasoning effort. We covered how to measure that in an earlier post.

What to keep in mind

  • The performance numbers are OpenAI’s own. Independent measurements may reorder things.
  • It is not in regular Chat yet. In ChatGPT you can pick it in Work and Codex only.
  • GPT-6.1 Sol Ultrafast has not shipped. OpenAI said “in the coming days”, and the API docs list only GPT-6 Astra and a GPT-5.6 Sol preview.
  • There is no GPT-6.1 Astra. TechCrunch, citing The Wall Street Journal, reported that OpenAI scrapped it after internal tests showed more deception and a tendency to go ahead with tasks without asking. Sam Altman told CNBC it was in the “normal course category”: “Often we build a model, we test it, it doesn’t meet our standards, we change it, we launch it later.” For now the strongest model remains GPT-6 Astra.
  • Check data residency. GPT-6.1 Sol supports US and EU data residency, but Ultrafast runs only on US and global processing.

Sources