Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 012 Latest feature

Wednesday, 30 September 2026

plexdata.online

010

Published
Author
Dorian Sotpyrc
Reading
6 minutes · Opinion
Status
Personal account · benchmark figures vendor-reported

Opinion · Models

GPT-6.1 Sol Has to Earn Back My Trust

Hours of unsuccessful revisions with GPT-6 Sol, an Astra escalation that exhausted my usage, and an Opus 5.5 handoff that finally worked. OpenAI’s new release arrives with a better value proposition, and some trust to recover.

References

Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.

Fig. 1 · AutomationBench 1.0.6Vendor-reported

Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.

Max effort · Opus includes default fallbacks · checked 30 Sep 2026
In this article · 5 sections
  1. 01The code changed. The problem remained.
  2. 02Other users report similar frustration
  3. 03Completed work is the better price comparison
  4. 04What I would like from the new Sol
  5. 05The release earns a retest
My account

Opus finished the job

GPT-6 Sol did not fix the integration issue, Astra used up my allowance, and Claude Opus 5.5 delivered the functionality with roughly 80% less code. That is my account of one job, not a controlled test.

Vendor claim

The new Sol is cheaper per task

OpenAI’s chart has GPT-6.1 Sol at 36.1% correct completion and $0.30 per task on AutomationBench. Opus 5.5 with fallbacks completes more, at $1.44.

Open question

Does it fix that job?

I have not used GPT-6.1 Sol on this incident. The release earns a retest, and only finished work earns the trust.

Last week, I spent hours trying to fix an issue in a large codebase connecting several systems. GPT-6 Sol kept circling the problem, giving me iterations of the code that achieved nothing. The root cause remained unresolved, and the functionality I needed still did not work.

I eventually handed the job to GPT-6 Astra. My available usage was exhausted almost immediately, before it had completed the work. I then moved it to Claude Opus 5.5.

Opus reduced and simplified the bloated codebase by roughly 80%, delivered the desired functionality, and integrated it correctly. After the previous attempts, that was a substantial difference.

The experience left me disappointed with OpenAI. It also changed what I wanted to hear about its next model. A claim about greater intelligence matters much more when it translates into a finished job.

GPT-6.1 Sol now arrives with reported improvements in capability and cost efficiency. I hope those improvements reach the kind of work that went badly last week. That remains an open question: I have not used it on this incident.

§ 01The code changed. The problem remained.

The frustrating part was the repeated lack of progress. Another explanation, another revision, another attempt, and the same underlying issue.

There is a cost to that beyond the model’s usage allowance. You have to follow the changes, keep track of what has been attempted, and decide whether the next explanation deserves more of your time. A model that produces a plausible patch can keep a session moving long after it has stopped being useful.

Opus’s smaller solution made an impression because it also worked. The approximate 80% reduction is my account of that job, rather than an independently measured benchmark. The models received sequential handoffs, so this was not a controlled comparison from identical starting conditions.

That limits what I can claim about the models in general. It does not change which one completed the work I needed.

§ 02Other users report similar frustration

Public complaints after the move from GPT-5.6 Sol to GPT-6 Sol describe problems that sound familiar. One Reddit user reported instruction-following failures and a costly return to Astra. Another described a maintenance check in which GPT-6 Sol relied on stale package information while its predecessor checked more thoroughly.12

These are self-selected accounts. They show that similar frustration exists, but they do not establish how common it is or whether the same cause explains my experience.

A more detailed report in OpenAI’s Codex issue tracker compared three pairs of tasks at maximum reasoning effort. Its author recorded context and completion omissions, but also a serious database defect found by GPT-6 Sol and missed by GPT-5.6 Sol. One pair had a contamination caveat.3

The picture is mixed. There is evidence worth taking seriously, including evidence against a blanket claim that GPT-6 Sol is worse. My disappointment is specific: on a difficult integration job, repeated attempts consumed time without delivering the required result.

§ 03Completed work is the more useful price comparison

The metric closest to that concern is cost per successful task: total spending across attempts divided by the number that finish correctly. Arize and Fireworks have published a benchmark using that approach, although its model set does not cover this three-way comparison.6

AutomationBench provides a useful comparison here. It tests business workflows across applications and grades the final system state. A task passes only when every required assertion holds. An agent saying it has finished does not count as success.5

OpenAI’s GPT-6.1 Sol release chart reports the following AutomationBench 1.0.6 results. All three rows use the named maximum effort setting; that does not imply identical compute budgets across providers. The Opus configuration includes default fallback models.45

AutomationBench 1.0.6, maximum effort, vendor-reported
Model / configurationCorrectly completed tasksReported cost per task, USDApprox. cost per correct completion, USD*
GPT-6.1 Sol · Max36.10%$0.2989$0.83
Claude Opus 5.5 · default fallbacks · Max42.47%$1.4400$3.39
GPT-6 Sol · Max31.96%$0.3406$1.07

*The final column is an editorial calculation from the published figures: reported cost per task ÷ success rate expressed as a fraction. For example, $0.2989 ÷ 0.361 = approximately $0.83. It treats the reported cost as the average across attempted tasks. It is not a separately measured result, a guarantee that retries will solve a particular failure, or a forecast of subscription usage.

In these results, Opus completes the largest share of tasks. The new Sol has the lowest reported cost per task and the lowest calculated spend per correct completion. Those are different advantages, and both matter.

AutomationBench concerns business workflow execution, rather than debugging my codebase. OpenAI’s published DeepSWE coding comparison does not include Opus 5.5, so it cannot supply the requested three-way coding chart.4

The fallback qualification matters too. Anthropic separately reports an Opus AutomationBench result without fallbacks; that is a different configuration and should not be mixed into this table.7

These are published evaluation results, not testing performed by PLEXData. API costs also cannot tell us how quickly my subscription allowance would have been consumed. The time spent supervising unsuccessful work is another cost this table does not capture.

§ 04What I would like to see from the new Sol

I would like to see GPT-6.1 Sol hold together the details of a connected system well enough to recognise where the failure begins. In last week’s job, another version of the code was of little value while the underlying problem remained.

I would like fewer confident explanations that lead back to the same fault. When the model has not established a cause, I would rather that uncertainty be clear before more time and capacity disappear into another round of changes.

The Opus handoff also raised my expectations about simplicity. It showed that, in this case, the desired functionality could be delivered with substantially less code. I would like the next Sol to recognise when complexity is getting in the way, while preserving the behaviour the system actually needs.

Above all, I would like enough usable capacity to finish difficult work. A model’s capability and the amount of it available to a customer belong in the same conversation. The Astra escalation was a reminder that access to a more capable model has limited practical value when the allowance is gone before the result arrives.

§ 05The release earns a retest; the work earns the trust

The published numbers give GPT-6.1 Sol a stronger case than its predecessor. They do not erase last week’s experience.

Opus 5.5 completed that job. GPT-6 Sol did not, and the move to Astra exhausted my available usage before completion. That is one personal experience, but it is the experience against which I will read the new claims.

I want OpenAI to have addressed it. A model that finishes more of the work, wastes less of the user’s time, and leaves a simpler working system would be welcome.

The release earns a retest. The work earns the trust.

§References

Sources checked 30 September 2026, Australia/Brisbane. Public accounts and published evaluations were reviewed; no model runs were performed for this article.

  1. 1Reddit: GPT-6 Sol instruction-following complaints and Astra usage — unverified personal account.reddit.com
  2. 2Reddit: GPT-6 Sol compared with GPT-5.6 Sol, including a maintenance-check example — original post and author’s follow-up; unverified.reddit.com
  3. 3openai/codex issue #47787: three paired audits — mixed findings and stated limitations; not an OpenAI finding of a general regression.github.com
  4. 4OpenAI: Introducing GPT-6.1 Sol — AutomationBench scores and costs extracted from the embedded chart data; vendor-reported. Opus result also corroborated by Zapier’s leaderboard.openai.com
  5. 5Zapier: AutomationBench — version 1.0.6, strict completion metric, evaluation method and default-fallback qualification. Benchmark repository.zapier.com
  6. 6Arize / Fireworks: Cost per successful task — metric definition and a separate study; its scores are not used in this article’s comparison.github.com
  7. 7Anthropic: Claude Opus 5.5 — AutomationBench reporting without fallback models; explains why configurations must remain distinct.anthropic.com