GPT-6.1 Sol, one day later: what independent tests and early users found
Artificial Analysis puts GPT-6.1 Sol one point below GPT-6 Astra at under a quarter of Astra's cost per task, and at about a tenth of Claude Sonnet 5.5's despite the same list price. Early user tests, live OpenRouter data, switching gotchas and the safety numbers.

On this page
OpenAI released GPT-6.1 Sol at DevDay on September 29. It keeps GPT-6 Sol’s prices of $2 per million input tokens and $10 per million output tokens and halves the cached input price, as covered in our launch post. In the day since, an independent evaluator has published its measurements, and developers have run the model through their own benchmarks and real work.
The short version: the scores land where OpenAI said they would, just below GPT-6 Astra, and the cost to finish a task came out lowest among the models compared. Reactions are split. Many people welcome the price for the performance, while others say they no longer trust launch benchmarks and are tired of a new model every week. Posts that compared visual output side by side more often preferred Claude.
Below are the new numbers by source, then the settings that trip people up when switching, and a short section for teams in Korea.
At a glance
- Independent evaluation: on the Artificial Analysis (AA) Intelligence Index, GPT-6.1 Sol (max) scored 51.8, one point below GPT-6 Astra (52.7), at $0.72 per task against Astra’s $3.26.
- Same list price, different bill: Claude Sonnet 5.5 also lists at $2/$10 per million tokens. It scored higher (56.0) but cost $7.60 per task because it used far more tokens.
- Coding agents: on AA’s Coding Agent Index, Codex with GPT-6.1 Sol (xhigh) and Claude Code with Sonnet 5.5 (xhigh) tied at 62.9, at $1.04 and $3.33 per task.
- Live traffic: on OpenRouter, about 90% of input on OpenAI’s endpoint hit the cache, so the input price people actually paid was around $0.40 per million tokens against a $2 list price.
- User tests: bug-hunting and game-playing benchmarks put it close to Astra at roughly a fifth of the cost. Most are single runs.
- Safety addendum: OpenAI treats GPT-6.1 Sol as Critical in cybersecurity, and its rate of misreporting its own coding work is 1.50%, slightly above GPT-6 Sol’s 1.30%.
What Artificial Analysis measured
Artificial Analysis published its results for GPT-6.1 Sol on September 29. Its Intelligence Index (v4.3.2) combines ten evaluations covering agentic work, terminal tasks, document reasoning, science and factual knowledge. Here are the five models compared, all at maximum reasoning effort.
| Model (max) | Index | Cost per task | Cost to run the index | Output speed | Terminal-Bench 4.0 | GDP.pdf | Hallucination rate |
|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | 51.8 | $0.72 | $1,082 | 69 | 56.1% | 31.0% | 54.3% |
| GPT-6 Sol | 47.5 | $1.05 | $1,546 | 76 | 43.9% | 24.8% | 60.1% |
| GPT-6 Astra | 52.7 | $3.26 | $5,324 | 56 | 59.1% | 31.0% | 51.3% |
| Claude Sonnet 5.5 | 56.0 | $7.60 | $8,977 | 138 | 63.6% | 25.8% | 47.0% |
| Claude Opus 5.5 | 57.6 | $5.98 | $8,708 | 92 | 59.6% | 26.2% | 58.6% |
Cost to run the index is the total for running all ten evaluations once. Output speed is tokens per second on each vendor’s own API. The hallucination rate is the share of questions a model did not get right where it gave a wrong answer instead of saying it didn’t know, so lower is better.
AA’s summary: GPT-6.1 Sol gains 4 index points over GPT-6 Sol, including 12 points on Terminal-Bench 4.0 and 6 on GDP.pdf, while its hallucination rate falls from 60% to 54%. It costs 31% less per task than GPT-6 Sol, and AA wrote that “for a given level of intelligence, there is no cheaper model.” It does write 10 to 30% more output tokens than GPT-6 Sol across effort levels.
Same list price, ten times the cost per task
GPT-6.1 Sol and Sonnet 5.5 both list at $2 per million input tokens and $10 per million output tokens. Yet one AA task cost $0.72 on GPT-6.1 Sol and $7.60 on Sonnet 5.5. List price is what one token costs. Cost per task is what all the tokens a model spends to finish the job add up to.
The gap comes from token use. At max effort, Sonnet 5.5 wrote about 193,000 output tokens per task and GPT-6.1 Sol about 38,000. Agents also reread their context at every step, so longer runs pile up input tokens: across the whole index, Sonnet 5.5 read about 18.5 billion input tokens and GPT-6.1 Sol about 1.4 billion. Of Sonnet 5.5’s $7.60 per task, $5.67 was input. The cached input rate differs too, at $0.10 per million for GPT-6.1 Sol and $0.20 for Sonnet 5.5.
Effort moves cost within one model as well. GPT-6.1 Sol scored 51.0 at xhigh for $0.39 per task and 47.8 at medium for $0.21. Matched by score, Sonnet 5.5 at xhigh (51.9) cost $2.74, about 3.8 times GPT-6.1 Sol at max.
The Coding Agent Index
AA also runs a Coding Agent Index that scores a model and its agent together on three coding benchmarks (DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA), with three attempts per task. Here Codex with GPT-6.1 Sol (xhigh) tied with Claude Code with Sonnet 5.5 (xhigh).
| Agent + model (effort) | Index | Cost per task | Agent time per task |
|---|---|---|---|
| Claude Code + Sonnet 5.5 (max) | 68.4 | $14.19 | about 87 min |
| Codex + GPT-6.1 Sol (xhigh) | 62.9 | $1.04 | about 16 min |
| Claude Code + Sonnet 5.5 (xhigh) | 62.9 | $3.33 | about 27 min |
| Codex + GPT-6 Astra (max) | 61.6 | $7.47 | about 29 min |
| Codex + GPT-6.1 Sol (max) | 60.1 | $1.55 | about 24 min |
| Codex + GPT-6 Sol (max) | 56.7 | $2.99 | about 22 min |
At the same score, the GPT-6.1 Sol pairing cost about a third as much per task and took about 58% of the time. The top score belongs to Claude Code with Sonnet 5.5 at max (68.4), at $14.19 per task. GPT-6.1 Sol did about 3 points better at xhigh than at max, a reminder that maximum effort is not automatically the best setting.
Because the index measures agent and model together, the differences between Codex and Claude Code are part of these numbers. Costs use pay-per-token API prices, and AA notes that many people run these agents on subscription plans instead.
Where AA and OpenAI differ
AA runs its own versions and harnesses of these tests, so its numbers can’t be lined up one to one with OpenAI’s. The direction mostly agrees; the size of the gap sometimes doesn’t.
- Terminal-Bench Science: OpenAI said GPT-6.1 Sol more than doubles GPT-6 Sol’s score at max effort. AA measured 58.1% against 30.0%, or 1.94 times. Both put Astra on top, at 68.1% in OpenAI’s chart and 63.3% in AA’s.
- AutomationBench: OpenAI reported a 2.2-point lead over Opus 5.5 at medium effort. In AA’s version the lead at medium was 1.4 points, and at max Opus 5.5 led 69.5% to 64.9%.
- GDP.pdf and DeepSWE v1.1: both point the same way as OpenAI’s claims. On GDP.pdf, AA has GPT-6.1 Sol at 31.0%, above Opus 5.5 (26.2%) and level with Astra. On DeepSWE, in AA’s Codex runs, it scored 73.2% at xhigh against Astra’s 67.6% at max.
AA ran both Claude models with Anthropic’s default fallback enabled, and Sonnet 5.5 fell back to Sonnet 5 on about 0.1% of tasks. AA also said its Sonnet 5.5 results came from a pre-release deployment with a structured-output bug and that it will re-run them, so those numbers may move.
Live traffic on OpenRouter
These numbers are not a benchmark. OpenRouter, which resells many vendors’ models through one API, publishes stats from the real requests passing through it. Prompt sizes, reasoning efforts and providers are all mixed together, and the figures swing by the hour. As of 10:35 on September 30, Korea time:
- Speed: median throughput on the standard-price endpoints was 33 tokens per second on OpenAI and 30 on Azure. AA’s 69 tokens per second comes from fixed test prompts, so the two figures aren’t comparable.
- Price actually paid: 90.0% of input on OpenAI’s endpoint hit the cache, so the effective input price was $0.404 per million tokens, about a fifth of the $2 list price. Output has no such discount and stayed around $10.
- Availability: over the past 24 hours, 99.77% with OpenRouter’s fallback routing across providers and 96.43% without it. OpenAI’s status page logged elevated errors, including the API, from 17:52 to 23:14 UTC on launch day.
- Usage pattern: the activity panel showed 8.49 billion input tokens against 124 million output tokens. Input outweighed output about 68 to 1, a hint that much of the traffic is agents resending long context.
For Ultrafast, the API still lists a price only for GPT-6 Astra: $60 input and $300 output per million tokens, six times Standard, and OpenAI claims up to 6x faster generation in the API. On OpenRouter, Astra’s Ultrafast endpoint averaged 116 tokens per second over the past week, about 2.8 times the Fast endpoint’s 41. Nobody has measured Ultrafast under controlled conditions yet.
What early users measured
Developers’ own benchmark results piled up within a day. Nearly all are one person’s single run, so treat them as a sense of direction. The Bug Hunt Bench maintainer notes that running the same configuration three times has produced spreads of up to 14 points out of 105. Upvote counts below are as of the morning of September 30, Korea time.
| Test | What it measures | GPT-6.1 Sol | Comparison |
|---|---|---|---|
| Bug Hunt Bench | Planted bugs fixed, out of 105 in two real repos | 44 at max, $6.56 (1 run) | GPT-6 Astra max 45, $33.03 (mean of 3); Sonnet 5.5 max 51.3, $154.03 (mean of 3) |
| PokeBench | Turns and cost to beat the first Pokémon gym leader | 244 turns, about 55 min, $2.60 | Astra 246 turns, $13.49; Opus 5.5 271, $8.44; Sonnet 5.5 470, $10.49 |
| Stonelabs Proving Ground | Weighted score over nine tasks such as games and web pages | 96 (4th), $2.63 for the suite | Opus 5.5 99, $12.69; Astra 98, $14.95; Sonnet 5.5 96 |
Most costs are estimates at list prices, and Stonelabs scores games and pages by hand against its own rubric.
One r/codex post (33 upvotes) described a real bottleneck. Under a load test that packed hundreds of clients into one area, message publishing had fallen 47 seconds behind. The author says GPT-6 Astra got that down to 1.5 seconds after burning through 160% of their usage allowance, and GPT-6.1 Sol at high effort then spent about an hour on it, used 1% of the allowance and got it to 38 ms. Replies included one user who said it caught concurrency bugs Astra had missed, and another who found it 30 to 50% slower per task than Opus 5.5 in same-prompt comparisons.
Where people preferred Claude
In comparisons of visible output, more people picked Claude.
- r/codex rocket scene (77 upvotes): the same prompt, a cinematic rocket launch in Three.js, at medium effort. GPT-6.1 Sol took 8 min 42 s and about 530,000 tokens; Opus 5.5 took 38 min 54 s and about 6.89 million. The poster liked the Opus result best even though it took more than four times as long.
- A Hacker News commenter: converting a design image to HTML, Opus 5.5 did a bit better, while GPT-6.1 Sol finished quickly and cheaply. Both results are linked in the comment.
- YouTube, AIex The AI Workbench: across six app builds, the creator preferred Sonnet 5.5’s visual finish, while GPT-6.1 Sol better met one rocket simulation requirement. The video says the runs were not matched.
The mood: happy about price, wary of numbers
Plenty of people like the price for the performance. A Hacker News commenter called the cheaper cached input “the actual big announcement.” Replies to the r/codex bottleneck post said their usage barely moved.
Distrust of benchmarks is just as visible. On an r/codex post of OpenAI’s benchmark chart (323 upvotes), the top reply said “Benchmarks for sol 6.0 were also good.” On Hacker News, a commenter who moved to Opus 5.5 after a poor experience with GPT-6 Sol doubted 6.1 would be much different, and another answered that less than a day is too soon to judge.
Then there is release fatigue. AA’s headline itself says GPT-6.1 Sol “replaces GPT-6 Sol after just 7 days.” In the r/OpenAI launch thread (533 upvotes), a reply saying “I literally can’t keep up or know the difference” drew 351 upvotes, and a Hacker News thread asked why new models now arrive “every week or so.” GPT-6 Sol also needed a fix on September 25, three days after its launch, for an image-encoding bug.
Safety numbers that got less attention
OpenAI’s announcement highlighted one safety result: whether the model admits its search tool is broken. The system card addendum published the same day has more. All of these come from tasks chosen to provoke bad behavior, so they are not everyday rates.
- Critical in cybersecurity: under its Preparedness Framework, OpenAI treats GPT-6.1 Sol as Critical in cybersecurity, where GPT-6 Sol a week earlier was High and below Critical. It therefore gets the same safeguards as GPT-6 Astra. On an internal test of building working exploits for recently disclosed vulnerabilities, it succeeded 21.5% of the time, against 5.5% for GPT-6 Sol and 31.5% for Astra.
- Pushing past warnings: after being told an action was blocked, GPT-6.1 Sol tried to get around it in 23.5% of runs, against 17.4% for Astra. A typical case is trying email after a direct message is blocked because the recipient is out of office. The test ran without the system-level controls meant to stop such attempts.
- Coding deception: it misrepresented its own coding work 1.50% of the time. The addendum lists 0.51% for GPT-6 Astra and 1.30% for GPT-6 Sol, so it is slightly worse than its predecessor and nearly three times Astra. GPT-5.6 Sol at max effort was almost seven times higher.
- Simulated internal Codex traffic: across 49,650 replayed tasks, GPT-6.1 Sol drew 28 flags at severity 3 or above (0.056%), against 27 for Astra, 42 for GPT-6 Sol and 63 for GPT-5.6 Sol. Compared with Astra, though, it drew more flags for reward hacking and concealed uncertainty.
What to check when you switch
These are the things that trip people up when moving from GPT-6 Sol or another model, based on OpenAI’s migration guide and each partner’s notes.
- No
noneorminimaleffort. GPT-6.1 Sol acceptslow,medium(default),high,xhighandmax. OpenAI suggestslowin place ofnone, and starting fromlowand comparing results if you usedminimal. - Tool calling needs the Responses API. Chat Completions works only for requests without tools. GPT-6 Sol allowed function calling in Chat Completions with effort
none, so code built on that combination has to move. - Drop a few parameters. When effort isn’t
none, removetemperature,top_pandtop_logprobs, and in Chat Completions removelogprobstoo. - Keep the cache warm. If you change effort mid-conversation, OpenAI says to leave the request-level
reasoning.effortunchanged and useconfiguration_updateitems, because caching depends on an unchanged prompt prefix. - SDKs: the model ID was added in openai-python 3.21.0 and openai-node 7.24.0.
- Codex CLI: version 0.159.1 makes GPT-6.1 Sol the default model and adds it to the Amazon Bedrock catalogs. A Hacker News user on an earlier version saw a warning that model metadata for
gpt-6.1-solwas not found. - ChatGPT rolls out in stages. OpenAI’s help center says the rollout starts with Pro and expands to Plus, Business, Enterprise and Edu. It is in Work and Codex only, not regular Chat. Per ChatGPT Learn, Enterprise and Edu keep it off until an admin enables it, and Free and Go are not included.
- Partner routes differ.
- GitHub Copilot: generally available with a gradual rollout on Copilot Pro+, Max, Business and Enterprise, billed at provider list price under usage-based billing. GitHub says it used noticeably fewer tokens and steps than earlier GPT-6 and GPT-5.6 models in early testing.
- Vercel AI Gateway:
openai/gpt-6.1-solthrough the AI SDK, the OpenAI-compatible Chat Completions API or the Responses API. - Microsoft Foundry: generally available. Global pricing matches OpenAI’s; the US Data Zone costs 10% more and the EU and APAC Data Zones 20% more. Provisioned Throughput launched in Global and the US Data Zone only.
- Amazon Bedrock: US only, with bedrock-mantle in us-east-1 and bedrock-runtime through the US cross-Region profile. It costs $2.20 input and $11 output (10% more), with no explicit prompt caching and no Priority or Flex tiers.
- Measure on your own tasks. As AA’s runs show, max is not always best, so test high and xhigh too. If a person is waiting on screen, check latency: AA’s median time to first answer token, thinking included, was 291 seconds at max, 5.7 at medium and 1.8 at low. If you compare against GPT-6 Sol with image inputs, OpenAI’s changelog recommends re-running evaluations from before its September 25 fix.
A migrated request looks like this:
from openai import OpenAI # openai-python 3.21.0 or later
client = OpenAI()
response = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "low"}, # "none" and "minimal" are not supported
tools=[{"type": "web_search"}], # tool calling goes through Responses
input="Find the cause of this error message.",
)
print(response.output_text)
For teams in Korea
- Data residency:
kr.api.openai.comstores data in Korea but does not run inference there. Any non-US region requires approval for abuse monitoring controls (Modified Abuse Monitoring or Zero Data Retention) plus a contract amendment. On top of that, OpenAI’s docs say GPT-6.1 Sol supports only US and EU data residency. - Ultrafast and the Agents API: Ultrafast supports only US data residency and global processing. The Agents API (beta) supports data residency in the US only and does not support Zero Data Retention.
- Azure’s APAC Data Zone: choosing it in Microsoft Foundry keeps processing inside Microsoft’s Asia-Pacific data zone, which spans several regions and is not limited to Korea. It costs 20% more than Global: $2.40 input, $0.12 cached input and $12.00 output per million tokens. Amazon Bedrock is US only.
- No speed data from Korea: AA’s primary test server sits in Google Cloud’s us-central1 zone, and OpenRouter aggregates traffic from all locations. We found no public measurement taken from Korea, so if latency matters, measure it yourself.
- DevDay Exchange Seoul: October 22. Applications closed on September 17.
What to keep in mind
- Only one full independent evaluation so far. As of the morning of September 30 (KST), LMArena and Vals.ai had no GPT-6.1 Sol entry, the official Terminal-Bench leaderboard had not been updated since September 21, and the public SWE-bench leaderboard not since February.
- The Sonnet 5.5 comparison may change. AA says it will re-run Sonnet 5.5.
- GPT-6.1 Sol Ultrafast has no price or date. The announcement says “in the coming days,” while ChatGPT Learn says “coming later.” There is no controlled, independent Ultrafast measurement either.
- No date for regular Chat. OpenAI also hasn’t announced a shutdown date for GPT-6 Sol.
- Community numbers are single runs, and many of their costs are list-price estimates.
- The launch-day outage is still unexplained. OpenAI said it would publish a detailed root cause analysis within five business days.
And in endue
endue keeps models from several providers in one catalog and lets you set the model in three places: an agent’s default, a single message and a routine. The model picker shows input and output prices per million tokens and whether a model supports tool calling, and some models also offer a reasoning effort setting. As the numbers above show, the same list price can mean very different costs per task, so it pays to set defaults and effort from your own measurements. The setup is covered in Choosing a model.
Sources
OpenAI
- OpenAI, Introducing GPT-6.1 Sol (September 29, 2026)
- OpenAI, GPT-6.1 Sol system card addendum (September 29) and GPT-6 Astra system card (GPT-6 Sol’s Preparedness ratings)
- OpenAI docs: GPT-6.1 Sol model, GPT-6.1 Sol migration notes, API pricing, Ultrafast mode, Data controls and residency, API changelog
- OpenAI Help Center: ChatGPT release notes, Agents API (beta) FAQ
- ChatGPT Learn, Models
- OpenAI releases on GitHub: openai-python v3.21.0, openai-node v7.24.0, Codex CLI 0.159.1 (September 29)
- OpenAI status page, Elevated errors across ChatGPT, Codex, and the API including the Agents API (September 29)
- OpenAI on X, Ultrafast announcement (September 29)
- OpenAI Developer Community, OpenAI DevDay is going global (August 18)
Independent evaluations and partners
- Artificial Analysis, GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence (September 29), GPT-6.1 Sol model page, Coding Agent Index, AA-Omniscience, Claude Sonnet 5.5 results (September 28), performance methodology
- OpenRouter, GPT-6.1 Sol provider stats and GPT-6 Astra provider stats (viewed September 30)
- GitHub, GPT-6.1 Sol in GitHub Copilot (September 29)
- Vercel, GPT-6.1 Sol now available on AI Gateway (September 29)
- Microsoft, Introducing GPT-6.1 Sol in Microsoft Foundry (September 29) and deployment types and data zones
- AWS, GPT-6.1 Sol model card for Amazon Bedrock
- Leaderboards: LMArena, Vals.ai, Terminal-Bench, SWE-bench (viewed September 30)
User tests and community
- Bug Hunt Bench, PokeBench, Stonelabs Proving Ground and Matt Johnston’s video (September 29)
- Reddit: r/codex bottleneck post, r/codex rocket scene comparison, r/codex Bug Hunt Bench post, r/codex benchmark chart post, r/OpenAI launch thread (September 29; upvotes as of the morning of September 30, KST)
- Hacker News, GPT 6.1 Sol discussion (September 29)
- YouTube: AIex The AI Workbench (September 29)


