Asia/Shanghai
April 22, 2026

Sisyphus and GPT Finally Clicked

Sisyphus 和 GPT 终于契合了

Mingjian Shao
Sisyphus and GPT Finally Clicked
The old story was simple: Sisyphus belonged with Opus-like models, and GPT belonged with Hephaestus. That story no longer matches reality.
For a long time, the implied doctrine was straightforward:
  • Sisyphus × Opus = the natural orchestrator pairing
  • Hephaestus × GPT = the natural deep specialist pairing
That pairing made sense when GPT looked too rigid for long orchestration prompts and Opus looked more natural in multi-step agent loops.But that is not the story I see now.In the Opus 4.7 / GPT-5.4 era, my view is much sharper than the cautious public timeline:
  • Sisyphus and GPT now click
  • GPT + Hephaestus (Deep Agent) is genuinely strong
  • the bigger question is no longer "can GPT do it?"
  • the bigger question is why I increasingly do not trust Opus's uncertainty in the same way I used to
So the important question is no longer:
"When did GPT become barely viable for Sisyphus?"
The better question is:
"When did the old Sisyphus = Opus narrative stop matching what power users were actually experiencing?"
Here is the version I would defend, combining the public record with the actual shift in user experience:
  • Late Dec 2025 to Feb 2026: users were already getting good enough results with GPT that the old "obvious mismatch" story was visibly breaking
  • Feb 2026: users openly said the official warning was lagging reality, and some were already reporting GPT + Sisyphus felt better in practice than Opus-based setups
  • Late Feb to Mar 2026: maintainers started retreating from hard enforcement, but the framework still carried the old doctrine in hooks and routing behavior
  • Apr 21, 2026: the docs finally acknowledged the shift in product form: GPT-5.4 now has a dedicated Sisyphus prompt path
So yes, there was a timeline shift. But the more interesting thing is the narrative reversal:
  • first, users discovered GPT had stopped feeling like the wrong brain for Sisyphus
  • then, official support caught up
  • and now, for some of us, the more uncomfortable truth is the reverse one: Opus is no longer the unquestioned safe default
DateSourceWhat it saysWhat it implies
2025-12-21Issue #158 — Claude Max Plan banned - alternative Sisyphus model options?A user asks for an alternative Sisyphus model after Opus usage problems and says they had already tried GPT-5.2 (extra high) and wanted to know if that was the recommended replacement.By late Dec 2025, users were already treating GPT as a plausible Sisyphus substitute, not an absurd mismatch.
2026-01-27 / 2026-01-28Issue #1168 — Is Opus 4.5 "required" for prometheus -> sisyphus handover?The maintainer response says: "No, Opus 4.5 is not strictly required" and "the framework itself is model-agnostic - it's prompts and harnesses, so any capable model should work."Publicly, this is the first strong statement that GPT use is technically legitimate in principle. But it still does not prove parity in quality.
2026-02-19Issue #1969 — UX: Cannot use GPT 5.3 with SisyphusOne user says: "I still can use gpt-5.3-codex with Sisyphus... it works well." Another says: "Contrary to the official warnings, I feel like I've consistently achieved better results using UltraWorker (Sisyphus) with Codex 5.3 xhigh compared to Opus 4.6."This is the clearest public split between official policy and user sentiment. The framework still treated GPT as the wrong pairing, while users were already saying the real-world results were good.
2026-02-22 / 2026-02-24Issue #2054 — Don't force model usage with Hephaestus or other models...A collaborator says: "Forcing GPT models is overkill. A recommendation to use GPT models would suffice." The owner replies that the enforcement hook is "overly aggressive" and says the fix is to convert hard enforcement to "recommendation/warning instead of forced override."This is the clearest policy inflection point. The project had not embraced full model freedom, but it had started backing away from hard routing ideology.
2026-03-03Issue #2266 — Selecting Sisyphus + GPT-5.2 auto-switches the conversation to HephaestusThe explanation is explicit: no-sisyphus-gpt automatically switches Sisyphus → Hephaestus when a GPT model is detected. It also states that Sisyphus is optimized for Claude/GLM models and Hephaestus for GPT.As of early Mar 2026, the framework still actively enforced the old pairing. So even if users felt GPT viability had improved, the official product stance had not fully caught up.
2026-04-21agent-model-matching.mdThe docs now say: "GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation for the orchestrator."This is the strongest evidence of a real architecture-level shift. GPT moved from discouraged mismatch to explicit support — but still second-choice versus the reference pairing.
The conversation was never unanimous. There were roughly three camps.This group rejected the product’s pessimistic messaging outright.They were not saying, “GPT might become good later.” They were saying, “I’m already using it and getting solid results.”That matters because it tells us the story was never simply:
  • framework says no
  • framework updates
  • users discover yes
In reality, a chunk of users had already moved ahead of the docs.Another group was less ideological. They were not claiming GPT was definitely superior. They were saying:
  • stop auto-switching my agent
  • stop overriding my config
  • warn me if you want, but don’t decide for me
This is a different kind of signal.It suggests that even before consensus on quality, there was already a consensus forming around user agency. Once enough power users report good outcomes, hard-coded pairing starts to feel paternalistic.This is the position I think is becoming more real in the Opus 4.7 / GPT-5.4 moment.The updated guide still does not say:
GPT is now the default recommendation for Sisyphus.
What it says is more limited:
GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation.
That means the official stance is still conservative.But in practice, many advanced users are no longer debating mere compatibility. They are debating which side feels more trustworthy now.The public evidence points to three different changes, not one.By late 2025 and early 2026, users were already reporting that GPT could handle more of the Sisyphus workflow than the old model-matching assumptions allowed for.At first, that only meant "this works better than the docs say."But in the Opus 4.7 / GPT-5.4 era, the shift feels stronger than that. It is no longer just a compatibility story. It is a fit story.Sisyphus × GPT now feels coherent enough that the old warning language reads historically interesting rather than operationally useful.This is the part the public issue tracker cannot fully prove for me, but it is the real interpretive frame behind the piece.The old narrative assumed Opus was the trustworthy orchestrator brain and GPT was the specialist deep-work brain.What changed is not merely that GPT improved.What changed is that, for some of us, Opus became harder to trust in the way that mattered most: not tone, not elegance, but operational certainty inside agent loops.That is why the story feels reversed.It is not:
  • GPT rose a bit
  • Opus stayed the same
It is closer to:
  • GPT improved enough to fit both the deep-worker pattern and more of the orchestrator pattern
  • while Opus started feeling more uncertain than its reputation suggested
The April doc update matters because it shows the team stopped treating GPT support as a weird edge case and started treating it as a designed path.That is the difference between:
  • “it kind of works if you disable the guardrails”
  • and “we now officially ship a GPT-tuned prompt path for this agent”
That second milestone is the real product-level change.My current answer is: that question is already outdated.The more honest answer for the Opus 4.7 / GPT-5.4 moment is:
  • Sisyphus × GPT now feels real, not contrarian
  • GPT + Hephaestus (Deep Agent) feels extremely strong
  • and the old assumption that Opus is the more reliable default deserves to be questioned, not repeated
The public sources still only prove a narrow claim:
  • GPT became viable for many users earlier than the docs admitted
  • official hard rejection softened before official endorsement arrived
  • the eventual public stance was support, not full default replacement
But the lived conclusion I would write more directly is this:
The Sisyphus × Opus story used to be the safe story. In the GPT-5.4 / Opus 4.7 era, that safety premium no longer feels obvious. If anything, the stronger stack for me is now on the GPT side — especially once Hephaestus enters the picture.
That is not just a style difference. That is a trust difference.The interesting lesson is not really about OpenAI vs Anthropic.It’s about how agent systems evolve.At first, a team discovers a pairing that works well:
  • Sisyphus → Claude-like models
  • Hephaestus → GPT-like models
That pairing becomes doctrine.Then models improve, users experiment, and doctrine starts lagging reality.The system reacts in stages:
  • users quietly override the rules
  • user reports pile up
  • maintainers soften enforcement
  • docs add an official path
  • defaults may still stay conservative for much longer
That’s exactly what seems to have happened here.The only twist now is that the lag is no longer just about compatibility. It is about the emotional center of the pairing itself.The old "Sisyphus belongs to Opus" line increasingly feels like inherited wisdom from a previous model generation.If I had to summarize both the public record and the real interpretive shift in one sentence:GPT did not suddenly become Sisyphus-compatible on one date; users were already reporting success by Dec 2025–Feb 2026, maintainers began backing away from hard enforcement in late Feb 2026, public product support arrived on Apr 21, 2026 with a dedicated GPT-5.4 Sisyphus prompt path — and by the Opus 4.7 / GPT-5.4 era, the deeper narrative had already flipped: Sisyphus × GPT felt aligned, GPT + Hephaestus felt very strong, and Opus no longer felt like the unquestioned certainty play.That is the story I actually care about.Not "did GPT win?" Not "is Opus still smart?"But this:
  • which stack feels aligned with the agent now
  • which stack feels operationally trustworthy
  • and at what point the official story stopped matching what advanced users were already seeing
That is where the real reversal happened.
以前的默认叙事很简单:Sisyphus 天生配 Opus,GPT 更适合 Hephaestus 这种 Deep Agent。现在这套叙事已经不准了。到了 Opus 4.7 / GPT-5.4 这个阶段,我自己的体感很明确:Sisyphus 和 GPT 已经契合,而 GPT + Hephaestus 这条线更是强得很具体。真正让我不舒服的,反而是我越来越受不了 Opus 的不确定性。
过去最常见的问法是:
“GPT 是从什么时候开始,终于能像 Opus 一样跑 Sisyphus 的?”
我现在觉得这个问法已经有点落后了。更值得问的是:
“Sisyphus = Opus 这套旧叙事,是从什么时候开始不再符合真实体验的?”
因为今天的问题已经不是“GPT 能不能勉强跑”。而是:
  • Sisyphus × GPT 现在是不是一个自然搭配?
  • 为什么越来越多人开始觉得 GPT 这一侧更稳?
  • 为什么 Opus 的默认权威感,开始被消耗掉了?
如果把公开记录和真实体感一起看,我现在会更直接一点:
  • 2025 年底到 2026 年 2 月:用户已经开始用 GPT 跑 Sisyphus,而且不是“能跑就行”,而是有人已经觉得它比旧叙事说得更顺
  • 2026 年 2 月:用户开始公开反驳“GPT 不适合 Sisyphus”这套官方口径
  • 2026 年 2 月下旬到 3 月:维护者开始往后退,不再那么理直气壮地硬拦,但框架层面仍然保留旧 doctrine
  • 2026 年 4 月 21 日:文档正式承认 GPT-5.4 now has a dedicated Sisyphus prompt path
所以真正的变化不是单点爆发,而是一场叙事反转:
  • 先是用户发现 GPT 其实已经不是错配
  • 然后框架支持才补上
  • 再往后,越来越多人开始面对一个更不舒服的事实:Opus 不再像以前那样,是那个默认最稳的答案
日期来源原话 / 核心意思它说明了什么
2025-12-21Issue #158 — Claude Max Plan banned - alternative Sisyphus model options?用户在 Opus 使用受限后,已经把 GPT-5.2 (extra high) 当作 Sisyphus 的备选,并询问这是不是推荐方案。到 2025 年底,GPT 至少在部分用户心里已经不是“离谱搭配”,而是一个真实可考虑的替代。
2026-01-27 / 2026-01-28Issue #1168 — Is Opus 4.5 "required" for prometheus -> sisyphus handover?维护者明确说:“No, Opus 4.5 is not strictly required”,并补了一句:“the framework itself is model-agnostic - it's prompts and harnesses, so any capable model should work.”这说明从技术原则上,项目方已经承认 GPT 跑 Sisyphus 不是“不可能”。但这还不是“质量已经追平”的证据。
2026-02-19Issue #1969 — UX: Cannot use GPT 5.3 with Sisyphus有用户说:“I still can use gpt-5.3-codex with Sisyphus... it works well.” 还有人更直接:“Contrary to the official warnings, I feel like I've consistently achieved better results ... compared to Opus 4.6.”到这一步,最关键的分裂已经出现了:官方说不行,用户说我已经在用了而且挺好。
2026-02-22 / 2026-02-24Issue #2054 — Don't force model usage with Hephaestus or other models...协作者直接说:“Forcing GPT models is overkill. A recommendation to use GPT models would suffice.” 维护者随后承认 hook “overly aggressive”,并表示要改成 “recommendation/warning instead of forced override.”这是最清晰的政策拐点:项目开始从“我替你决定”退到“我给你建议”。
2026-03-03Issue #2266 — Selecting Sisyphus + GPT-5.2 auto-switches the conversation to Hephaestus明确解释了机制:no-sisyphus-gpt 会在检测到 GPT 时自动把 Sisyphus 切成 Hephaestus。 同时继续强调 Sisyphus 更适合 Claude/GLM到 3 月初,框架在产品层面其实仍然在坚持旧的 model-agent doctrine。
2026-04-21agent-model-matching.md文档第一次公开写出:“GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation for the orchestrator.”这是最强的一条证据:GPT 不再只是“用户硬改也许能用”,而是已经有了明确设计过的支持路径。但默认推荐仍然不是它。
公开讨论里,至少能分成三派。这批用户并不是在说“也许以后 GPT 会更适合”。他们说的是:
  • 我现在就在用
  • 我已经觉得它能打
  • 甚至我的主观体验里比 Opus 还顺
这很重要。因为它说明变化不是先从文档开始,而是先从用户实践开始。另一批人没那么想打擂台。他们不是一定要证明 GPT 比 Opus 强。他们更在意的是:
  • 别自动帮我切 agent
  • 别替我改配置
  • 你可以提示,但别强制
这类反馈本质上是“用户自治”问题,不是纯 model 排位问题。一旦有足够多的高级用户觉得“我自己知道我在干嘛”,那种硬编码的 model matching 就会开始显得居高临下。这更接近我现在的感受。4 月的文档当然还是保守的。它没说:
GPT 已经成为 Sisyphus 的默认答案。
它只是说:
GPT-5.4 现在有 dedicated Sisyphus prompt path,但 GPT 仍然不是默认推荐。
但从很多高阶用户的体感看,问题已经不只是“兼容不兼容”。而是:现在到底哪一侧更让我信?到 2025 年底 / 2026 年初,用户已经开始反馈 GPT 跑 Sisyphus 的效果,比旧 model-matching 假设好得多。最开始,这还只是“它比文档说得能打”。但到了 Opus 4.7 / GPT-5.4 这个阶段,我自己的判断已经更强了:这不再只是兼容性故事,而是 fit 的故事Sisyphus × GPT 已经不是那种“你强行关掉 guardrail 才能玩的邪道”。它开始像一条自然路线。这一点公开 issue 很难帮我完全证明,但这是我这篇文章真正想写的解释框架。旧叙事默认:
  • Opus 是更值得信赖的 orchestrator 脑子
  • GPT 是更适合 deep work 的 specialist 脑子
现在让我感到变化最大的,不只是 GPT 变强了。而是 Opus 开始没有以前那么让我信得过了不是语气,不是文风,也不是 benchmark 排名。而是在 agent loop 里,我越来越受不了它那种不确定性。所以叙事才会反过来。不是:
  • GPT 提升了一点
  • Opus 还是那个稳稳的默认答案
而更像是:
  • GPT 强到足够同时覆盖 deep-worker 模式,甚至吃进更多 orchestrator 模式
  • 而 Opus 的“默认更稳”这层光环,开始不再理所当然
4 月那条文档真正重要的地方在于,它不再把 GPT support 当成奇技淫巧,而是把它写成了正式路径。这和下面这两种状态完全不是一回事:
  • “你把 guardrail 关掉以后也许能跑”
  • “我们现在正式给这个 agent 做了一条 GPT 专用 prompt path”
后者才是真正意义上的产品层变化。我现在的回答是:这个问题本身已经有点过时了。更接近我现在体感的说法是:
  • Sisyphus × GPT 已经成立,不再只是逆风玩法
  • GPT + Hephaestus(Deep Agent)非常强
  • 而“Opus 还是那个更稳的默认答案”这句话,今天应该被质疑,而不是被复读
公开资料本身仍然只能稳稳证明三件事:
  • GPT 对很多用户来说,比文档承认得更早就已经可用了
  • 官方的硬性否定先软化,再变成支持
  • 最终的公开立场是支持,不是彻底替代默认值
但如果写我真正的结论,我会更直接一点:
在 GPT-5.4 / Opus 4.7 这个阶段,旧的 Sisyphus × Opus 安全叙事已经不够可信了。至少对我来说,更强、更顺、更让我信的那一侧,已经是 GPT —— 尤其是再加上 Hephaestus 之后。
这就不只是风格问题了,而是信任问题。这件事真正有意思的地方,其实不只是 OpenAI 和 Anthropic 谁强。更有意思的是:agent 系统的 doctrine 往往会落后于用户实践。一开始,团队发现一组搭配特别顺:
  • Sisyphus → Claude-like 模型
  • Hephaestus → GPT-like 模型
于是这个搭配被写进 prompt、写进 hook、写进文档,变成“正确姿势”。但随着模型变强、用户开始乱试、反馈不断累积,这套 doctrine 会慢慢落后于现实。通常它会按这个顺序变化:
  • 用户先偷偷绕过规则
  • 更多用户出来说“其实能用”
  • 维护者先放松 enforcement
  • 文档补上正式路径
  • 默认推荐往往还会保守很久
Sisyphus 和 GPT 的公开记录,基本就是这个过程。今天唯一新增的 twist 是:这个 lag 不再只是兼容性 lag,而是叙事中心本身的 lag。“ Sisyphus 天生属于 Opus ”这句话,现在越来越像上一代模型格局留下来的惯性认知。如果要把公开记录和我真正想表达的叙事变化压成一句话,那就是:GPT 不是在某一天突然“能跑 Sisyphus”了;用户在 2025 年底到 2026 年 2 月之间就已经开始稳定报告它可用,维护者在 2026 年 2 月下旬开始撤回过于强硬的 enforcement,文档在 2026 年 4 月 21 日补上了 dedicated GPT-5.4 Sisyphus prompt path,而到了 Opus 4.7 / GPT-5.4 这个阶段,更深的反转其实已经发生了:Sisyphus × GPT 开始显得契合,GPT + Hephaestus 显得很强,Opus 不再像以前那样是那个默认最稳的答案。这才是我真正关心的故事。不是“GPT 赢没赢”,也不是“Opus 还聪不聪明”。而是:
  • 这个 agent 现在到底和哪一侧的模型更对味
  • 哪一侧在真实工作流里更让我信得过
  • 又是从什么时候开始,官方叙事已经跟不上高阶用户的真实体验了
Share this post:
Enjoy this post? Subscribe via RSS: English | 中文