The old story was simple: Sisyphus belonged with Opus-like models, and GPT belonged with Hephaestus. That story no longer matches reality.
The Narrative Actually Reversed
- Sisyphus × Opus = the natural orchestrator pairing
- Hephaestus × GPT = the natural deep specialist pairing
- Sisyphus and GPT now click
- GPT + Hephaestus (Deep Agent) is genuinely strong
- the bigger question is no longer "can GPT do it?"
- the bigger question is why I increasingly do not trust Opus's uncertainty in the same way I used to
"When did GPT become barely viable for Sisyphus?"The better question is:
"When did the old Sisyphus = Opus narrative stop matching what power users were actually experiencing?"
The Short Answer
- Late Dec 2025 to Feb 2026: users were already getting good enough results with GPT that the old "obvious mismatch" story was visibly breaking
- Feb 2026: users openly said the official warning was lagging reality, and some were already reporting GPT + Sisyphus felt better in practice than Opus-based setups
- Late Feb to Mar 2026: maintainers started retreating from hard enforcement, but the framework still carried the old doctrine in hooks and routing behavior
- Apr 21, 2026: the docs finally acknowledged the shift in product form: GPT-5.4 now has a dedicated Sisyphus prompt path
- first, users discovered GPT had stopped feeling like the wrong brain for Sisyphus
- then, official support caught up
- and now, for some of us, the more uncomfortable truth is the reverse one: Opus is no longer the unquestioned safe default
Tight Timeline
| Date | Source | What it says | What it implies |
|---|---|---|---|
| 2025-12-21 | Issue #158 — Claude Max Plan banned - alternative Sisyphus model options? | A user asks for an alternative Sisyphus model after Opus usage problems and says they had already tried GPT-5.2 (extra high) and wanted to know if that was the recommended replacement. | By late Dec 2025, users were already treating GPT as a plausible Sisyphus substitute, not an absurd mismatch. |
| 2026-01-27 / 2026-01-28 | Issue #1168 — Is Opus 4.5 "required" for prometheus -> sisyphus handover? | The maintainer response says: "No, Opus 4.5 is not strictly required" and "the framework itself is model-agnostic - it's prompts and harnesses, so any capable model should work." | Publicly, this is the first strong statement that GPT use is technically legitimate in principle. But it still does not prove parity in quality. |
| 2026-02-19 | Issue #1969 — UX: Cannot use GPT 5.3 with Sisyphus | One user says: "I still can use gpt-5.3-codex with Sisyphus... it works well." Another says: "Contrary to the official warnings, I feel like I've consistently achieved better results using UltraWorker (Sisyphus) with Codex 5.3 xhigh compared to Opus 4.6." | This is the clearest public split between official policy and user sentiment. The framework still treated GPT as the wrong pairing, while users were already saying the real-world results were good. |
| 2026-02-22 / 2026-02-24 | Issue #2054 — Don't force model usage with Hephaestus or other models... | A collaborator says: "Forcing GPT models is overkill. A recommendation to use GPT models would suffice." The owner replies that the enforcement hook is "overly aggressive" and says the fix is to convert hard enforcement to "recommendation/warning instead of forced override." | This is the clearest policy inflection point. The project had not embraced full model freedom, but it had started backing away from hard routing ideology. |
| 2026-03-03 | Issue #2266 — Selecting Sisyphus + GPT-5.2 auto-switches the conversation to Hephaestus | The explanation is explicit: no-sisyphus-gpt automatically switches Sisyphus → Hephaestus when a GPT model is detected. It also states that Sisyphus is optimized for Claude/GLM models and Hephaestus for GPT. | As of early Mar 2026, the framework still actively enforced the old pairing. So even if users felt GPT viability had improved, the official product stance had not fully caught up. |
| 2026-04-21 | agent-model-matching.md | The docs now say: "GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation for the orchestrator." | This is the strongest evidence of a real architecture-level shift. GPT moved from discouraged mismatch to explicit support — but still second-choice versus the reference pairing. |
What Users Were Actually Saying
1. “GPT works fine for me.”
- framework says no
- framework updates
- users discover yes
2. “The hard block is the real problem.”
- stop auto-switching my agent
- stop overriding my config
- warn me if you want, but don’t decide for me
3. “The story already flipped, but the docs were late.”
GPT is now the default recommendation for Sisyphus.What it says is more limited:
GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation.That means the official stance is still conservative.But in practice, many advanced users are no longer debating mere compatibility. They are debating which side feels more trustworthy now.
What Actually Changed
Change 1: GPT stopped feeling like the wrong brain
Change 2: Opus lost some of its aura of certainty
- GPT rose a bit
- Opus stayed the same
- GPT improved enough to fit both the deep-worker pattern and more of the orchestrator pattern
- while Opus started feeling more uncertain than its reputation suggested
Change 3: the prompts and product policy finally caught up
- “it kind of works if you disable the guardrails”
- and “we now officially ship a GPT-tuned prompt path for this agent”
So Did GPT “Catch Up” to Opus?
- Sisyphus × GPT now feels real, not contrarian
- GPT + Hephaestus (Deep Agent) feels extremely strong
- and the old assumption that Opus is the more reliable default deserves to be questioned, not repeated
- GPT became viable for many users earlier than the docs admitted
- official hard rejection softened before official endorsement arrived
- the eventual public stance was support, not full default replacement
The Sisyphus × Opus story used to be the safe story. In the GPT-5.4 / Opus 4.7 era, that safety premium no longer feels obvious. If anything, the stronger stack for me is now on the GPT side — especially once Hephaestus enters the picture.That is not just a style difference. That is a trust difference.
The Better Mental Model
- Sisyphus → Claude-like models
- Hephaestus → GPT-like models
- users quietly override the rules
- user reports pile up
- maintainers soften enforcement
- docs add an official path
- defaults may still stay conservative for much longer
My Take
- which stack feels aligned with the agent now
- which stack feels operationally trustworthy
- and at what point the official story stopped matching what advanced users were already seeing
以前的默认叙事很简单:Sisyphus 天生配 Opus,GPT 更适合 Hephaestus 这种 Deep Agent。现在这套叙事已经不准了。到了 Opus 4.7 / GPT-5.4 这个阶段,我自己的体感很明确:Sisyphus 和 GPT 已经契合,而 GPT + Hephaestus 这条线更是强得很具体。真正让我不舒服的,反而是我越来越受不了 Opus 的不确定性。
大家其实把问题问浅了
“GPT 是从什么时候开始,终于能像 Opus 一样跑 Sisyphus 的?”我现在觉得这个问法已经有点落后了。更值得问的是:
“Sisyphus = Opus 这套旧叙事,是从什么时候开始不再符合真实体验的?”因为今天的问题已经不是“GPT 能不能勉强跑”。而是:
- Sisyphus × GPT 现在是不是一个自然搭配?
- 为什么越来越多人开始觉得 GPT 这一侧更稳?
- 为什么 Opus 的默认权威感,开始被消耗掉了?
先给结论
- 2025 年底到 2026 年 2 月:用户已经开始用 GPT 跑 Sisyphus,而且不是“能跑就行”,而是有人已经觉得它比旧叙事说得更顺
- 2026 年 2 月:用户开始公开反驳“GPT 不适合 Sisyphus”这套官方口径
- 2026 年 2 月下旬到 3 月:维护者开始往后退,不再那么理直气壮地硬拦,但框架层面仍然保留旧 doctrine
- 2026 年 4 月 21 日:文档正式承认 GPT-5.4 now has a dedicated Sisyphus prompt path
- 先是用户发现 GPT 其实已经不是错配
- 然后框架支持才补上
- 再往后,越来越多人开始面对一个更不舒服的事实:Opus 不再像以前那样,是那个默认最稳的答案
紧凑时间线
| 日期 | 来源 | 原话 / 核心意思 | 它说明了什么 |
|---|---|---|---|
| 2025-12-21 | Issue #158 — Claude Max Plan banned - alternative Sisyphus model options? | 用户在 Opus 使用受限后,已经把 GPT-5.2 (extra high) 当作 Sisyphus 的备选,并询问这是不是推荐方案。 | 到 2025 年底,GPT 至少在部分用户心里已经不是“离谱搭配”,而是一个真实可考虑的替代。 |
| 2026-01-27 / 2026-01-28 | Issue #1168 — Is Opus 4.5 "required" for prometheus -> sisyphus handover? | 维护者明确说:“No, Opus 4.5 is not strictly required”,并补了一句:“the framework itself is model-agnostic - it's prompts and harnesses, so any capable model should work.” | 这说明从技术原则上,项目方已经承认 GPT 跑 Sisyphus 不是“不可能”。但这还不是“质量已经追平”的证据。 |
| 2026-02-19 | Issue #1969 — UX: Cannot use GPT 5.3 with Sisyphus | 有用户说:“I still can use gpt-5.3-codex with Sisyphus... it works well.” 还有人更直接:“Contrary to the official warnings, I feel like I've consistently achieved better results ... compared to Opus 4.6.” | 到这一步,最关键的分裂已经出现了:官方说不行,用户说我已经在用了而且挺好。 |
| 2026-02-22 / 2026-02-24 | Issue #2054 — Don't force model usage with Hephaestus or other models... | 协作者直接说:“Forcing GPT models is overkill. A recommendation to use GPT models would suffice.” 维护者随后承认 hook “overly aggressive”,并表示要改成 “recommendation/warning instead of forced override.” | 这是最清晰的政策拐点:项目开始从“我替你决定”退到“我给你建议”。 |
| 2026-03-03 | Issue #2266 — Selecting Sisyphus + GPT-5.2 auto-switches the conversation to Hephaestus | 明确解释了机制:no-sisyphus-gpt 会在检测到 GPT 时自动把 Sisyphus 切成 Hephaestus。 同时继续强调 Sisyphus 更适合 Claude/GLM。 | 到 3 月初,框架在产品层面其实仍然在坚持旧的 model-agent doctrine。 |
| 2026-04-21 | agent-model-matching.md | 文档第一次公开写出:“GPT-5.4 now has a dedicated Sisyphus prompt path, but GPT is still not the default recommendation for the orchestrator.” | 这是最强的一条证据:GPT 不再只是“用户硬改也许能用”,而是已经有了明确设计过的支持路径。但默认推荐仍然不是它。 |
用户到底在说什么
1)“GPT 对我来说已经够好,甚至更好”
- 我现在就在用
- 我已经觉得它能打
- 甚至我的主观体验里比 Opus 还顺
2)“真正的问题不是质量,而是你别拦我”
- 别自动帮我切 agent
- 别替我改配置
- 你可以提示,但别强制
3)“叙事已经反了,只是文档还没完全跟上”
GPT 已经成为 Sisyphus 的默认答案。它只是说:
GPT-5.4 现在有 dedicated Sisyphus prompt path,但 GPT 仍然不是默认推荐。但从很多高阶用户的体感看,问题已经不只是“兼容不兼容”。而是:现在到底哪一侧更让我信?
真正变化的,其实不止一件事
变化 1:GPT 不再像“错的脑子”
变化 2:Opus 的确定性溢价掉了
- Opus 是更值得信赖的 orchestrator 脑子
- GPT 是更适合 deep work 的 specialist 脑子
- GPT 提升了一点
- Opus 还是那个稳稳的默认答案
- GPT 强到足够同时覆盖 deep-worker 模式,甚至吃进更多 orchestrator 模式
- 而 Opus 的“默认更稳”这层光环,开始不再理所当然
变化 3:prompt 和产品策略终于跟上了
- “你把 guardrail 关掉以后也许能跑”
- “我们现在正式给这个 agent 做了一条 GPT 专用 prompt path”
那 GPT 算不算已经“追平”了 Opus?
- Sisyphus × GPT 已经成立,不再只是逆风玩法
- GPT + Hephaestus(Deep Agent)非常强
- 而“Opus 还是那个更稳的默认答案”这句话,今天应该被质疑,而不是被复读
- GPT 对很多用户来说,比文档承认得更早就已经可用了
- 官方的硬性否定先软化,再变成支持
- 最终的公开立场是支持,不是彻底替代默认值
在 GPT-5.4 / Opus 4.7 这个阶段,旧的 Sisyphus × Opus 安全叙事已经不够可信了。至少对我来说,更强、更顺、更让我信的那一侧,已经是 GPT —— 尤其是再加上 Hephaestus 之后。这就不只是风格问题了,而是信任问题。
更好的理解方式
- Sisyphus → Claude-like 模型
- Hephaestus → GPT-like 模型
- 用户先偷偷绕过规则
- 更多用户出来说“其实能用”
- 维护者先放松 enforcement
- 文档补上正式路径
- 默认推荐往往还会保守很久
我的判断
- 这个 agent 现在到底和哪一侧的模型更对味
- 哪一侧在真实工作流里更让我信得过
- 又是从什么时候开始,官方叙事已经跟不上高阶用户的真实体验了