# 当 AI 能让你"不学习也能通过考试",大学还剩下什么?
Anthropic 最近做了期访谈,找来四位来自普林斯顿、伯克利、伦敦政治经济学院、亚利桑那州立大学的学生,聊校园 AI 的真实状态。访谈持续了近 40 分钟,没有宣传话术,都是真实的困惑、焦虑和思考。
这期访谈揭示了一个更深层的问题:**AI 不只是在改变学习方式,而是在拆解整个教育体系的底层逻辑**。
***
## 一个无法回避的现实
访谈开始,主持人问了个直接的问题:现在校园里对 AI 的氛围如何?
答案是:**90% 的学生在用 AI**。不是偶尔用,而是日常工作流程的一部分——总结讲座笔记、回答问题集、获取作业反馈、分析商业案例、做市场调查、完成财务研究。有些学生甚至用它完成测验,理由很现实:当你是研究生,同时做好几份工作时,你并不总是有时间。
但更有意思的是,虽然几乎所有人都在用,却没人知道规则是什么。有些课程明文禁止 AI,有些课程积极鼓励,大部分课程处于模糊地带。学生不知道界限在哪,教授也不知道该怎么管。这种状态被称为"灰色地带"——想用,又怕违规;不用,又觉得自己落后。
这种灰色地带最危险的地方不是学生可能违规,而是**它阻止了真正有价值的讨论发生**。学生无法公开分享 AI 使用的最佳实践,教授无法指导学生如何负责任地使用工具,整个学术社区陷入一种表面禁止、私下使用的尴尬状态。
而当规则无法被有效执行时,它就变成了筛选"会伪装的人"和"不会伪装的人"的工具。明文禁止但私下普遍使用的状态,**不会阻止学生使用 AI,只会阻止他们公开讨论如何更好地使用**。
***
## AI 是一面镜子
访谈中有个观点很犀利:人工智能,尤其是学生如何使用人工智能,非常能说明这些动机。
这句话背后的洞察是:**AI 成了一面镜子,照出你上大学的真实目的**。
访谈把大学目标归纳为三类:第一,深入学习专业知识,掌握某个领域的深层理解;第二,为职业做准备,找到好工作,建立职业网络;第三,拓展人脉,享受社交生活,体验大学文化。每个学生对这三个目标的权重都不一样,而 AI 的使用方式精准地暴露了这些权重。
如果你只关心"通过考试"和"拿到学位",你会直接把 AI 的输出当作作业提交。这不是道德批判,而是现实——当技术让你可以用最小成本达到目标时,你为什么要绕远路?如果你的目标本来就是拿学位找工作,那用 AI 完成作业是完全理性的选择。
但如果你真的想学习,想深入理解某个领域,你会把 AI 当成对话伙伴。你会问它问题,让它解释概念,然后用自己的话重新表述。你会让它写初版代码,然后自己重构优化。你会确保在每一个环节,自己都真正理解发生了什么。
这种分化不只存在于不同学生之间,也存在于不同专业之间。人文专业学生往往选择退出 AI,因为他们的学习需要精读——仔细阅读原文、品味语言细节、理解作者意图。AI 破坏了这个过程,因为它提供的是总结和转述,而非原文的直接体验。工程和商科学生则大量使用 AI,因为 AI 降低了技术门槛,让没有计算机背景的人也能写代码、建网站、分析数据。
**这种两极分化本质上是对"什么值得学习"的不同理解**。对人文学生来说,阅读莎士比亚原文的体验本身就是学习;对工程学生来说,重要的是能否解决问题,而不是亲手写出每一行代码。AI 让这种差异变得更加明显。
***
## 工具还是拐杖?一个简单的判断标准
访谈中有个问题很关键:如何区分 AI 是工具还是拐杖?
学生们给出的答案惊人地一致:**能否解释**。
如果你无法解释自己创造的东西,无法说明 AI 在其中扮演了什么角色,那就是拐杖。如果你能像给五年级学生解释一样讲清楚,能给出低层次和高层次的解释,那就是工具。
这个标准看似简单,实则触及了学习的本质。费曼学习法的核心逻辑是:如果你不能用简单的语言解释一个概念,那你就还没有真正理解它。AI 时代的学习也是同样的道理——如果你不能解释 AI 帮你做的事情,那你就只是在"外包思考",而非"增强思考"。
访谈中提到一个有意思的例子:有学生开发了一个工具,可以把讲义幻灯片放进去,AI 会在每张幻灯片旁边生成类似教授的注释。这个学生说:"它之所以好用,是因为我已经提示它知道我想了解什么——幻灯片上某些事物的定义。幻灯片有时候很抽象,缺乏背景信息,需要在旁边添加上下文。"
关键在于"我已经提示它知道我想了解什么"。这个学生清楚自己的知识缺口在哪,知道需要什么样的帮助,然后主动引导 AI 提供这种帮助。这是工具。如果是直接把幻灯片扔给 AI,让它"帮我总结这节课",然后照着背,那就是拐杖。
**区别在于主动性和理解**。使用工具的人知道自己在做什么,控制着整个过程。使用拐杖的人把控制权交给了技术,自己变成了被动接收者。
***
## 学校的滞后不是反应慢,而是本质上无法应对
访谈中提到了一些学校的尝试。伦敦政治经济学院的一门必修课,开始指导学生如何使用 Claude——和它对话,赋予它不同角色,然后要求学生提交对话记录,看他们如何与 AI 互动。亚利桑那州立大学的职业管理中心建了提示库,为不同场景提供提示模板。这些都是好的尝试,核心思路是:不是禁止 AI,而是教学生负责任地使用。
但这只是少数。大部分学校还在争论"该不该允许学生用 AI"。有些教授说可以用,但要在作业里注明使用方式;有些课程直接禁止;还有些干脆不提,默认学生不会用。没有整合框架,没有统一标准,整个体系处于混乱状态。
更深层的问题是:**这个问题本质上无法靠规章制度解决**。
传统教育的监管逻辑是:学校设定规则,学生遵守规则,违规者被惩罚。这个逻辑成立的前提是"违规行为可以被检测"。但 AI 打破了这个前提。
你可以禁止学生在提交作业时使用 AI,但你无法监控学生在思考过程中是否用了 AI。你可以用 AI 检测工具,但这些工具的准确率远没有达到可以作为惩罚依据的程度——误判率太高,而且学生会很快学会如何绕过检测。更重要的是,**从根本上说,你无法区分"学生在 AI 帮助下完成的高质量作业"和"学生独立完成的高质量作业"**,因为好的 AI 使用本来就应该是无缝的。
访谈中有句话说得很直接:从根本上说,不会有任何规章制度能改变学生使用 AI 的方式。责任在学生手中。这不是推卸责任,而是现实。
当技术让"不学习也能通过考试"成为可能,学校面临的不是一个管理问题,而是一个存在性问题:**如果考试无法证明学习,那学校存在的意义是什么?**
***
## 教育体系的底层逻辑被打破了
这个问题触及了教育体系的根本性矛盾。
传统教育建立在几个核心假设上:第一,知识是稀缺的,需要专门的机构(学校)和专业人员(教授)来传授;第二,学习成果可以通过考试来衡量;第三,学位证明了你掌握了某个领域的知识,因此有资格从事相关工作。
AI 逐一打破了这些假设。
**知识不再稀缺**。YouTube 上有免费的斯坦福课程,Claude 可以随时回答你的问题,GitHub 上有无数开源项目可以学习。你不需要去学校才能接触到知识,甚至不需要付费订阅就能获得基础的 AI 辅导。
**考试无法衡量学习**。当 AI 能够完成大部分考试题目,考试就从"衡量理解程度的工具"变成了"衡量是否会用 AI 的工具"。这不是说考试完全无用,而是说它不再能准确区分"真正学会的人"和"会利用工具的人"。
**学位的价值在下降**。当雇主意识到学位无法保证候选人真的掌握了相关知识,他们会更看重实际能力的证明——作品集、项目经验、实习表现。学位从"能力证明"降格为"基础门槛"。
**当这些假设被打破,教育体系的价值主张就需要重新定义**。
访谈给出了一个答案:大学的价值从"**传授知识**"转向"**提供环境**"。这个环境让你可以犯错、可以探索、可以和其他人碰撞想法。你可以和室友花周末做"毕业前愿望清单",在 hackathon 测试"可能很蠢"的想法,和教授争论,和同学讨论,在失败中学习而不必承担职业生涯的风险。
访谈中提到的学生项目很能说明这一点——"自动选课提醒"、"空教室查找器"、"毕业愿望排行榜"。这些项目技术上都不复杂,很多创作者甚至没有计算机背景。但它们源于真实的人类情感:对错过的恐惧、对便利的追求、对大学生活的珍惜。
**技术门槛降低后,重要的不再是"你会不会写代码",而是"你想解决什么问题"**。大学提供的是一个可以自由探索这些问题的空间,一个可以把想法变成现实、然后从失败中学习的环境。
AI 能帮你完成作业,但不能替你度过这段"可以犯错、可以探索"的时光。
***
## 就业市场的悖论
访谈后半段聊到就业,揭示了另一个悖论。
学生用 AI 写简历,公司用 AI 筛简历。整个招聘周期变成了对着屏幕说话——先是对着 AI 写求职信,然后对着录像回答问题,最后收到 AI 生成的拒信。从提交简历到收到拒信,可能只需要 15 分钟。效率很高,但人性很少。
这创造了一个"AI 对 AI"的就业市场。学生训练 AI 如何写出"好的"简历,公司训练 AI 如何筛选"好的"候选人。真正的人类在这个过程中的作用越来越小。对着屏幕说话没有化学反应,无法展现那些难以量化但很重要的品质——幽默感、应变能力、团队协作的微妙之处。
但悖论的另一面是:AI 熟练程度本身成了新的竞争力。四大咨询公司以前招通才型 MBA,现在专门找具备 AI 能力的 MBA。如果你懂得如何将 AI 应用于不同行业,你就是他们的首选候选人。
**悖论在于:AI 让就业市场更冰冷,也让就业市场更看重 AI 能力。你无法逃避,只能学会有效使用**。
这又回到了那个核心问题:什么是"有效使用"?不是会用 ChatGPT 写邮件,而是能够识别哪些问题适合用 AI 解决,能够设计提示词来引导 AI 产出你需要的结果,能够评估 AI 输出的质量并进行必要的修正。
这种能力不是通过禁止 AI 来培养的,而是通过大量的实践和试错。这也是为什么那些主动拥抱 AI、建立 Claude Builder Club、组织 hackathon 的学校,正在为学生提供更有价值的教育——他们让学生在相对安全的环境中,学习如何与 AI 协作。
***
## 责任的转移
访谈最核心的洞察可能是这句:当技术能让你"不学习也能通过考试"时,学习本身的意义变成了每个人需要自己回答的问题。
这是责任的转移。从**学校转向学生,从规则转向自觉,从外部动机转向内部动机**。
传统教育的动机是外部的:你需要通过考试来拿学位,需要学位来找工作。这个外部激励体系推动着学生学习。但当 AI 让你可以不学习也能通过考试,这个激励体系就失效了。
剩下的只有内部动机:你真的想学吗?你真的对这个领域感兴趣吗?你真的想深入理解,还是只想拿个学位?
访谈中有个细节很有意思。有学生提到,读研究生有不好的一面——你同时做好几份工作,没有时间,所以有时候会用 AI 快速完成测验。但他接着说:读研本应该是你拓展批判性思维的时期,是你展现更果断一面的时候。**他意识到了矛盾,但选择了效率**。
这不是道德批判。在现实压力下,效率往往比理想更重要。但这个选择揭示了一个事实:**当外部压力(完成测验)和内部动机(深入学习)冲突时,很多人会选择前者**。
AI 让这种冲突变得更加尖锐,因为它让"应付考试"变得极其容易。在没有 AI 的时代,即使你只想应付考试,你也必须学习一些东西才能通过。AI 取消了这个中间环节——你可以完全不学习也能通过。
**这迫使每个人直面那个问题:你到底为什么要上大学?**
如果答案是"为了拿学位找工作",那用 AI 完成作业是完全合理的。如果答案是"我真的想学习这个领域",那你需要主动抵抗 AI 带来的捷径诱惑。
学校无法替你做这个选择。规则无法强制你产生内部动机。**这是你自己的责任**。
***
## 技术不会等你准备好
访谈最后有个态度贯穿始终:"我们会想办法解决的。"
学校规则跟不上?先用起来,再告诉学校什么有效。AI 可能被用来作弊?慢慢学会负责任地使用。就业市场变了?适应新游戏规则。
这不是盲目乐观,而是现实主义。技术已经在这里了,它不会等你准备好才开始改变世界。你可以选择抵抗,也可以选择适应,但你无法选择让时间停止。
**这一代学生和 AI 的关系不是恐惧,不是盲目拥抱,而是在混乱中摸索,在试错中学习**。
他们在 Claude Builder Club 做的那些项目——技术上不复杂,但解决真实问题。他们在 hackathon 上尝试的那些想法——可能很蠢,但至少在尝试。他们在课堂上的困惑——规则不清楚,但至少在思考。
这种"**边做边学**"的姿态,可能比任何规章制度都更能帮助他们适应 AI 时代。
凯文·凯利在《科技想要什么》里提出"technium"概念:技术作为一个整体,似乎有自己的意志,想要变得越来越强大。访谈中有个观察很呼应这一点:过去两年,不管 AI 需要什么才能继续发展,它就能得到。核能态度的转变、太空数据中心的讨论——每当有什么瓶颈可能减缓 AI,障碍就会被移除。
但更准确的说法可能是:**学生需要什么,就会创造什么**。
AI 只是工具。决定未来的,是这一代学生选择如何使用这个工具——是逃避思考,还是增强思考。是用它来应付考试,还是用它来探索世界。是把它当拐杖,还是当工具。
这个选择,学校管不了,只能他们自己决定。
而从这期访谈来看,至少有一部分学生,正在认真思考这个问题。这可能就够了。
# 像 Agent 一样思考
Anthropic 工程师 Thariq 最近发了一条推文,讲 Claude Code 团队给 Claude 加一个"主动问用户问题"功能的过程:弹个窗让用户选 A 还是 B,比让用户从一段长文字里挑答案省事得多。
他们做了三次。
第一次想偷懒——在已有的 ExitPlanTool 上加个参数,让它输出计划的同时也输出问题。失败:用户回答和已经输出的计划冲突时,Claude 不知道该听谁的。
第二次更工程化——约定一种 markdown 格式,让 Claude 用固定结构提问,前端解析。失败:模型一次次用错格式,多余句子、缺选项、自由发挥。
第三次单独做一个 AskUserQuestion 工具,调用就弹窗,阻塞 agent 循环直到用户回答。成了。
但比这个故事更值得记的,是 Thariq 顺手提到的另一件事。
Claude Code 早期有一个 TodoWrite 工具,让 Claude 自己写待办清单。模型常常忘记自己该做什么,团队就每 5 轮对话插一次系统提醒。
聪明的设计。直到模型变强,提醒反而成了限制——Claude 把清单当成必须严格遵守的东西,不敢随机应变。最后团队彻底重做,用 Task Tool 替代了 TodoWrite。
> Todos 像老板盯着员工的清单。Tasks 像团队的协作看板。
这是 Thariq 全文最好的一句话。它指向的现象比 agent 设计普遍得多:
**所有 prompt engineering 都有半衰期。**
{/* TODO: 这里写一段你自己的真实例子,三五句话就够。
- 某个 skill 你刚写时它工作得很好,过了几个月发现 Claude 自己已经会做了;
- CLAUDE.md 里某段提醒,本来是为了让模型不出错,现在反而让它畏首畏尾;
- 一个曾经精心设计的提示词,今天读起来已经像累赘。
git status 里 xhs-card-generator 改到 v2.1,可能是天然素材。
不要凑话——这一段是整篇文章里读者唯一会真信你的地方。 */}
你不会收到一个明确的信号说"这条规则可以删了"。它只是悄悄从"必要"滑向"多余",再从"多余"滑向"反作用"。不定期回头看自己写过的提示词,它们会越积越多——而且大部分都已经过期。
像 agent 一样去看世界。也像使用者一样,看你做的每一样东西。
***
**延伸阅读:**
# 从芯片战争到太空数据中心:AI 行业的下一个十年
> 投资是对真理的追求。如果你先找到真理,而且判断正确,那就是你创造 alpha 的方式。而且这必须是其他人尚未看到的真相。
这句话来自 Atreides Management 创始人 Gavin Baker 在 Patrick O'Shaughnessy 的播客 "Invest Like the Best" 中的访谈。Gavin 被称为科技投资领域最具热情和洞察力的投资者之一,这期近两小时的对话覆盖了 GPU、TPU、AI 经济学、太空数据中心、SaaS 的未来,甚至他从滑雪教练到投资者的人生转折。
这期访谈信息密度很高,有太多让人"停下来想一想"的时刻。以下是几个最值得深思的观点。
***
## 如何追踪 AI 发展?先花 200 美元
访谈一开始,Patrick 问了一个很实际的问题:像 Gemini 3 这样的新模型发布时,你是怎么处理这些信息的?
Gavin 的回答很直接:**你必须自己用**。
但关键不是"用",而是用什么版本。他对那些用免费版 AI 就得出"AI 不过如此"结论的投资者感到惊讶:
> 免费版就像你在和一个 10 岁的孩子打交道,然后你根据这个 10 岁孩子的表现,去预测他 35 岁时会怎样。你可以付费——实际上你必须付费才能获得最高级别的会员资格,每月 200 美元。那些才是真正的 30-35 岁成年人。
这个比喻很精准。国内外大模型的差距其实也是类似的道理——很多人用国内的免费模型体验了一下,就觉得"AI 不过如此",但如果你用过 Claude 4.5 Opus、Gemini 3 Pro、GPT-5.2 Reasoning 这些顶级模型,感受会完全不同。
说到付费,我自己每个月大概花 2000 人民币订阅各种 AI 产品,其中大头是 250 美元的 Claude Code Max 套餐。如果你是开发者,编程需求比较大,我非常建议直接订阅官方的 Claude Code(125 美元起),不要去用各种镜像站。一方面你不知道镜像站背后到底用的是不是真实的模型,另一方面 Max 套餐的性价比其实很高,算下来比按量付费划算太多了。
至于获取信息的渠道,Gavin 的答案可能会让很多人意外:**X (Twitter)**。
他说,AI 的发展很大程度上"在 X 平台上实时发生"。地球上真正理解 AI 前沿的人大概有 500 到 1000 人,相当一部分在中国,而你需要密切关注这些人。他特别提到 Andrej Karpathy:
> 安德烈·卡帕西写的每篇文章,你都得读三遍。最低限度。
作为一个也在关注 AI 动态的人,我对此深有共鸣。Twitter 上的 AI 讨论确实比任何新闻媒体都更实时、更深入。那些实验室的研究员们会直接发帖讨论最新进展,甚至会互相"吵架"——Gavin 提到 Meta 的 PyTorch 团队和 Google 的 Jax 团队曾在 X 上爆发过一场公开争论,最后两个实验室的负责人不得不出面表态:"我们实验室的人不准说对方实验室的坏话。"
***
## Scaling Laws:我们的"古埃及时刻"
Gemini 3 发布后,很多人关注的是它说明了什么关于 scaling laws(规模定律)。Gavin 给出了一个我从没听过的视角:
> 我们对预训练 scaling laws 的理解,可能就像古埃及人对太阳的理解。他们可以非常精确地测量,精确到大金字塔的东西轴线与春分秋分完美对齐,巨石阵也是如此。完美的测量。但他们不懂轨道力学。他们不知道太阳为什么从东方升起、西方落下。
这个比喻让我停下来想了很久。我们确实可以非常精确地预测:给模型增加 10 倍计算量,性能会提升多少。但我们不知道为什么。这不是一条"定律",而是一个"经验观察"——一个我们测量得极其精确、但不理解其原理的经验观察。
所以 Gemini 3 为什么重要?因为它证明了这个"经验观察"仍然成立。在 Blackwell 芯片延迟、所有人都在担心"scaling laws 是否已经失效"的时候,Gemini 3 给出了明确的答案:**没有失效**。
但更有意思的是 Gavin 接下来说的:如果没有"推理模型"(reasoning models)的出现,2024 年到 2025 年的 AI 发展本来应该停滞。
为什么?因为在 XAI 搞定 20 万张 Hopper GPU 的协同工作之后,下一步需要等 Blackwell 芯片。你没法让超过 20 万张 Hopper 保持"相干"(coherent)——简单理解就是让它们像一个整体一样工作。而 Blackwell 延迟了。
> 如果没有推理模型,从 2024 年中期到现在,AI 就不会有任何进展。一切都会停滞。你能想象这对市场意味着什么吗?我们会生活在一个完全不同的环境里。推理模型在某种程度上拯救了 AI,因为它让 AI 在没有 Blackwell 的情况下继续进步。
这是一个我之前没有意识到的视角:推理模型(比如 o1)不只是一个新能力,它实际上"救"了整个 AI 行业的发展节奏。
***
## 芯片战争:Google 在"吸走氧气"
谈到 GPU 和 TPU 的竞争,Gavin 说了一句让我印象深刻的话:
> Google 目前是 token 的最低成本生产者。他们一直在做的事情,我会说是"吸走 AI 生态系统的经济氧气"——这对他们来说是一个极其理性的策略。
作为低成本生产者,Google 一直在用低价(甚至亏损)提供 AI 服务,让竞争对手的日子很难过。这是科技行业的经典打法,但 Gavin 指出了一个有趣的变化:
> AI 是我职业生涯中第一次看到"低成本生产者"这件事在科技领域真正重要。苹果不是因为低成本生产手机而市值数万亿。微软不是因为低成本生产软件而市值数万亿。英伟达也不是因为低成本生产 AI 加速器而市值数万亿。这从来都不重要。
但在 AI 时代,当电力成为限制因素时,**每瓦能产出多少 token** 变得至关重要。如果你每瓦能产出 3-5 倍的 token,那就是 3-5 倍的收入。计算的价格变得无关紧要,因为瓶颈是电力。
这个格局即将改变。Blackwell 芯片终于开始部署,而 Gavin 预测第一个 Blackwell 模型会来自 XAI:
> 根据 Jensen 的说法,没有人比 Elon 建数据中心更快。Jensen 公开说过这话。
当 Blackwell 和后续的 Ruben 芯片大规模部署后,Google 作为低成本生产者的优势会消失。到那时,他们还愿意继续以 -30% 的毛利率运营 AI 业务吗?这个计算会完全改变。
***
## 太空数据中心:疯狂,但从第一性原理看是对的
当 Patrick 问起"有没有什么不太被讨论的疯狂想法"时,Gavin 开始谈太空数据中心。一开始我以为这是在开玩笑,但听完他的分析,我意识到这可能是整个访谈里最有远见的部分。
> 从第一性原理的角度来看,太空数据中心在每个维度上都优于地球上的数据中心。
他的论证是这样的:
**1. 能源**:在太空中,卫星可以 24 小时暴露在阳光下,太阳辐射强度比地面高 6 倍。而且因为始终有阳光,你不需要电池——电池是很大一部分成本。所以太阳系中成本最低的能源是"太空太阳能"。
**2. 冷却**:地球上的数据中心,大部分成本和重量都用于冷却。但在太空中?冷却免费。把散热器放在卫星背阴面,那里接近绝对零度。
**3. 网络**:在数据中心里,机架之间用光纤连接——本质上是激光穿过电缆。唯一比这更快的是什么?激光穿过真空。所以如果用激光连接太空中的卫星,网络实际上比地面数据中心更快。
**4. 用户体验**:现在你问 AI 一个问题,信号要从手机到基站,到光纤,到某个数据中心,计算完再原路返回。但如果卫星能直接和手机通信(Starlink 已经证明了直连手机的能力),整个链路会短得多。
当然,这需要 Starship 大规模发射才能实现,可能还要 5-6 年。但 Gavin 指出了一个有趣的融合:Tesla、SpaceX、XAI 正在汇聚。XAI 会是 Optimus 机器人的"智能模块",SpaceX 会在太空建数据中心为 AI 提供算力,这三家公司正在形成一个互相增强竞争优势的闭环。
***
## SaaS 的"燃烧平台"
如果前面的内容让你对 AI 的未来感到兴奋,这一部分可能会让你对很多现有公司的未来感到担忧。
Gavin 直接说:**应用 SaaS 公司正在犯和实体零售商面对电商时完全相同的错误**。
实体零售商当年看亚马逊,觉得"电商是低利润业务,怎么可能比我们更高效?现在顾客自己付钱来店里,自己把东西带回家。"他们明明看到了客户需求,却因为不喜欢电商的利润结构而拒绝投入。结果呢?亚马逊北美零售业务的利润率现在比很多传统零售商还高。
SaaS 公司现在面临同样的处境。传统软件写一次就能无限复制分发,毛利率可以高达 80-90%。但 AI 不一样——每次使用都要重新计算,好的 AI 公司毛利率可能只有 40%。
> 如果你想做 AI agent,你不愿意以低于 35% 的毛利率运营,你就永远不会成功。因为 AI 原生公司就是以这个利润率运营的。如果你想保护 80% 的毛利率,你就是在保证自己在 AI 领域不会成功。绝对保证。
Gavin 说这是一个"生死攸关的决定",而**除了微软,几乎所有公司都在失败**。
他引用了诺基亚那封著名的"燃烧的平台"备忘录:你的平台着火了。但旁边其实有一个很好的新平台,你可以跳过去,然后回头把原来平台上的火扑灭。现在你就有两个平台了。
Salesforce、ServiceNow、HubSpot、GitLab、Atlassian——他认为所有这些公司都可以并且应该运行这个策略:公开你的 AI 收入,公开你的 AI 毛利率(低毛利率恰恰证明这是"真正的 AI"),然后指向那些还在亏损的风险投资支持的竞争对手,说"我有他们没有的:一个能产生现金流的业务。"
***
## 一个投资者的成长故事
访谈的最后,Patrick 问了一个更私人的问题:你会怎么向年轻人介绍你做的事情?
Gavin 的回答从"投资是对真理的追求"开始,但真正有意思的是他的人生故事。
他原本的计划是:冬天当滑雪教练,夏天做漂流向导,淡季去攀岩,顺便尝试写小说和做野生动物摄影。这是他大学时的"人生规划",父母都很支持。
但父母提了一个小要求:能不能找一份专业实习,就一份,什么都行?
他能找到的唯一实习是在一家券商的私人财富管理部门。工作很简单:每当公司发布研究报告,他要查一下哪些客户持有这只股票,然后把报告寄给他们。
然后他开始读那些报告。
> 我心想:"天哪,这是我能想象到的最有趣的事情。"
他把投资理解为一场"技巧与运气并存的游戏",有点像扑克。你可以因为运气不好而输——比如你投资的公司总部被陨石砸中——但大部分时候,技巧是重要的。而获得优势的方式,就是拥有最透彻的历史知识,结合对当前世界最准确的理解,形成一个对"接下来会发生什么"的差异化判断。
那是他实习的第三天。他去书店买了彼得·林奇的书,两天读完。然后读巴菲特,读《市场奇才》,读巴菲特写给股东的信——读了两遍。然后自学会计。回学校后把专业从英语和历史改成了历史和经济学。
他还提到了一段做清洁工的经历。在阿尔塔滑雪场打工时,他做过客房清洁。有一次他在打扫房间时,看到客人在读的书和自己在读的是同一本,他说"这是本好书,我正好读到和你差不多的地方"。对方看他的眼神像看外星人一样,然后更震惊地问:"你还读书?"
> 这件事永久地影响了我对待其他人的方式。
***
## 尾声:AI 需要什么,就得到什么
在访谈快结束时,Gavin 说了一段我觉得最有意思的话:
> 过去两年,不管 AI 需要什么才能继续发展,它就能得到。你见过美国公众舆论对任何问题的转变像核能问题这么快吗?就这么发生了。而且恰好在 AI 需要它发生的时候发生。现在我们遇到地球上的电力限制了,突然之间太空数据中心的讨论就出现了。每当有什么瓶颈可能减缓 AI 发展,一切反而加速了。
这让人想起凯文·凯利在《科技想要什么》里提出的"technium"(技术体)概念:技术作为一个整体,似乎有某种自己的意志,想要变得越来越强大。
也许这只是巧合。也许这只是很多聪明人在解决问题。但 Gavin 观察到的这个模式——AI 遇到什么障碍,障碍就会以某种方式被移除——确实值得思考。
# 精选解读
# Curations 精选解读
这里收录了我观看技术大佬视频、阅读优质博客后的思考与总结。
不只是简单的笔记摘录,而是融入了自己的理解和实践心得。
## 内容来源
* 技术视频解读
* 博客文章精选
* 播客/访谈总结
## 最新内容
### Karpathy 一条推引爆 62k star:andrej-karpathy-skills 到底做了什么
2026 年 4 月 GitHub 周榜第一的 `forrestchang/andrej-karpathy-skills` 只是一个 CLAUDE.md 文件,却在两周里拿到 62.7k star。它把 Karpathy 那条 769 万浏览的长推打包成四条可安装的规则——本质是 2026 年最典型的"内容即产品"案例。
[阅读全文 →](./karpathy-skills)
### 像 Agent 一样思考:Claude Code 团队的工具设计哲学
Anthropic 工程师 Thariq 分享了构建 Claude Code 过程中的 agent 工具设计经验——从 AskUserQuestion 的三次迭代到 TodoWrite 的三次重构,从 RAG 到渐进式披露,每一个案例都指向同一个核心方法论:像 Agent 一样去看世界。
[阅读全文 →](./claude-code-seeing-like-an-agent)
### 90%的大学生在用AI,但没人知道规则是什么
四位来自普林斯顿、伯克利、LSE的学生聊了聊校园AI的真实状态——作弊、迷茫、两极分化,以及那个没人敢问的问题:大学还有什么意义?不是宣传片,是真实的困惑和思考。
[阅读全文 →](./ai-on-campus-student-perspectives)
### 从芯片战争到太空数据中心:AI 行业的下一个十年
从芯片战争到太空数据中心,从 SaaS 的生死抉择到投资的本质。Gavin Baker 在 "Invest Like the Best" 播客中分享了他对 AI 行业最深刻的洞察:为什么你必须用付费版 AI、Scaling Laws 为什么像"古埃及人理解太阳"、推理模型如何"拯救"了整个 AI 行业的发展节奏。
[阅读全文 →](./gavin-baker-ai-economics)
# Karpathy 一条推引爆 62k star:andrej-karpathy-skills 到底做了什么
2026 年 1 月 27 日,Andrej Karpathy 在 X 上发了一条超长推——11 个小节、约 1400 词的编程随笔,记录他从 11 月的"80% 手写+20% agent"迅速切换到 12 月"80% agent+20% 润色"这个转变里踩到的坑。这条推最终冲到 **769 万次浏览、3.9 万喜欢、3.6 万书签**。
三个月后,一个叫 `forrestchang/andrej-karpathy-skills` 的 GitHub 仓库上线,把 Karpathy 这条推里的吐槽打包成四条可安装的规则。两周之内冲到 **62.7k star、5.5k fork**,成为 2026 年 4 月 GitHub 周榜第一。
仓库本体:**一个 Markdown 文件**。
***
## 一、Karpathy 在吐槽什么
Karpathy 那篇长文实质上列了四个 LLM 编码的"慢性病"。
**第一病:偷偷替你做假设**
> "最常见的一类错误是,模型替你做出错误的假设,然后不加验证就照搬执行。它们不管理自身的困惑、不寻求澄清、不展示不一致、不呈现权衡、该反驳时不反驳,而且还有点过于奉承。"
这是"共谋式错误"。你说"帮我加个登录",它不问用什么鉴权、不问要不要记住设备、不问 session 怎么管——直接铺一套它觉得合理的方案。等你 review 完发现和你想要的不一样时,它已经写了 500 行。
**第二病:过度工程**
> "它们特别喜欢把代码和 API 搞得过于复杂,抽象层臃肿、不清理死代码。它们会用 1000 行代码实现一个低效、臃肿、脆弱的结构,你得像哄小孩一样说'嗯,你们为什么不直接这么做呢?',它们才会说'当然可以!'然后立刻精简到 100 行。"
这是 LLM 编码最典型的观察者效应:**它在被"大方"地给予上下文的同时,会"大方"地回报以复杂度**。Strategy 模式、Factory 模式、依赖注入——全都给你塞上。
**第三病:改你没让它改的东西**
> "它们有时候还会因为不喜欢或者没完全理解而修改、删除一些注释和代码——即使这些改动和当前任务完全无关。"
你让它 fix 一个 bug,它顺手把旁边没写完的 TODO 注释删了,理由是"看起来不再需要"。
**第四病:即使你在 CLAUDE.md 里写了规矩,它还是会犯**
> "以上问题,即使我在 CLAUDE.md 里做了一些简单的修复尝试,也依然存在。"
这是整条推里最扎心的一句。Karpathy 作为前 OpenAI 创始团队和 Tesla AI Director,都写不出让 Claude 彻底守规矩的 CLAUDE.md。
***
## 二、`andrej-karpathy-skills` 的解法
`forrestchang` 把这四条病的解法系统化成四条原则,打包进一个 CLAUDE.md 文件。
| 原则 | 对应的病 | 核心动作 |
| ------------------------- | ---- | ------------------------------ |
| **Think Before Coding** | 偷偷假设 | 明说假设、列多种解读、困惑就停下来问、该反驳就反驳 |
| **Simplicity First** | 过度工程 | 只写被要求的最小代码;不写投机性的灵活性、错误处理、抽象 |
| **Surgical Changes** | 越权改动 | 只碰必须碰的;不顺手重构、不改风格;发现别的死代码只报告不删 |
| **Goal-Driven Execution** | 方法错位 | 给可验证的成功标准 + 测试,让模型自己 loop 到通过 |
第四条原则直接引用了 Karpathy 那条推里的另一段金句——"Leverage":
这是整个项目方法论层面的落脚点:**前三条原则让 LLM 不要乱来;第四条原则告诉你怎么真正发挥它的长处。**
***
## 三、怎么安装使用
**方式 A:作为 Claude Code 插件**(推荐先用这种)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
装完后,所有项目的 Claude Code 对话都会自动遵守这四条原则。全局生效,随时 `/plugin` 关掉。
**方式 B:手动复制 CLAUDE.md**
进 repo → 打开 `CLAUDE.md` → 复制 → 粘贴到你项目根目录的 `CLAUDE.md` 里。只对这个项目生效。
仓库还额外提供了 `CURSOR.md` 和 `.cursor/rules/` 适配——一套内容覆盖主流 AI IDE。
***
## 四、为什么能冲到 62k star
这是一个值得拆解的现象。62.7k star 对一个"单文件 repo"来说是夸张的数字——作为对比,同期的 Microsoft markitdown(9k)、Addy Osmani 的 agent-skills(4.6k)加起来还没它多。
按影响权重拆解:
**1. Karpathy IP 背书** ——同样的内容如果叫 `forrestchang-skills` 不可能破万。Karpathy 自带"前 OpenAI 创始团队 + Tesla AI Director + CS231n 讲师"的文化资本,他的推文自带"必读"标签。
**2. 时机完美** —— Opus 4.7 在 4 月 16 日发布,over-engineering 的抱怨达到峰值。repo 正好卡在所有人都在找"让 Claude 别那么疯"的解药时出现。
**3. 痛点普适** ——每个 Claude Code / Cursor 用户都踩过这四个坑,共情率接近 100%。
**4. 门槛极低** —— 1 个文件或 2 行命令。star 成本低到可忽略,"不装就亏"。
**5. 可验证感强** ——4 条原则清晰好记,容易截图转发。不像 1000 行的 prompt 工程指南让人望而却步。
**6. 双语 README** ——`README.zh.md` 直接吃掉中文 AI 圈流量,V2EX / 即刻 / 微博同步引爆。
**7. 作者交叉推广** ——顶栏那句 *"Check out my new project Multica"* 把流量导去作者自己的商业化 agent 平台 `multica-ai/multica`。**这个 repo 本质是 Multica 的获客漏斗顶端。**
**8. Meta 契合** ——它讨论的"LLM 编码毛病",正是所有读者此刻在用 LLM 编码时的现场。读和用合一,转化率极高。
一句话:它卖的不是代码、不是工具,而是**把 Karpathy 的情绪打包成可安装的规矩**——这是 2026 年 AI 编程圈最典型的"内容即产品"案例。
***
## 五、我的使用建议
**先用方式 A 全局装上**。看看对你写工具、写脚本时有没有体感改善——尤其是让 Claude 改别人的代码时,它顺手乱改的毛病是否减少。
**一两周后再决定要不要合并进项目 CLAUDE.md**。每个项目的 CLAUDE.md 已经塞满了领域知识(设计系统、组件规范、部署流程),而 Karpathy 这套是通用方法论。两者不冲突,可以叠加——但时机要等你真的确认它有用再合。
**注意代价**:它会让 Claude 问话变多,习惯了"一句话生成"的人会觉得烦;该做的轻微清理它也可能不做(太守规矩);对非常模糊的探索型任务反而束手束脚。
**更深层的价值**:它逼你把需求说清楚——这恰好也是所有高质量软件工程的前提。
**值得关注的后续**:forrestchang 本人同时在推 Multica——一个"open-source managed agents platform",把 skills 机制产品化。如果这套四原则最终成为事实标准,Multica 就是它的商业化载体。关注一下这条线。
***
## 参考资源
# Indie Dev
记录独立开发的完整过程 — 从准备到上架的每一步。
# Bark
# Bark
一个可以通过简单的 HTTP 请求向 iPhone 发送自定义推送通知的工具。免费、开源、支持自建服务器。
## 推荐理由
* **极简 API** - 一条 curl 命令即可发送推送通知,无需复杂配置
* **开源免费** - MIT 协议,完全开源,零费用
* **隐私优先** - 支持自建服务器,推送数据完全掌握在自己手中
## 使用场景
* **脚本通知** - 数据备份完成等长时间任务结束时收到提醒
* **服务监控** - 服务器异常时即时收到告警
* **自动化集成** - CI/CD 构建结果、定时任务完成通知
## 快速开始
1. 从 [App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865) 下载 Bark
2. 打开 App,复制推送地址
3. 发送第一条推送:
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
带标题的推送:
```bash
curl https://api.day.app/YOUR_KEY/标题/内容
```
## 项目信息
* GitHub: [Finb/Bark](https://github.com/Finb/Bark)
* Stars: 7.2k+
* 协议: MIT
# Toolkit
# Toolkit
收录我发现的优质 GitHub 项目和实用软件。
# Claude Agent Teams 完全指南
## 引言
如果你用过 Claude Code 的 Subagent,可能会觉得并行开发已经足够强大了。但 Subagent 有个限制:它们只能向主 Agent 汇报结果,无法相互交流。
Agent Teams 彻底改变了这一点。想象一下:一个 Agent 负责安全审查,另一个负责性能优化,第三个负责测试覆盖——它们不仅能并行工作,还能直接对话、相互挑战、达成共识。这就是 Agent Teams 的核心价值。
## 理解 Agent Teams
Agent Teams 的架构很像一个真实的开发团队:
```
┌─────────────────────────────────────────────────────────┐
│ 你(用户) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (主 Claude 实例,负责协调) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│ 安全审查 │◄─►│ 性能优化 │◄─►│ 测试覆盖 │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└──────────────┼──────────────┘
▼
┌─────────────────┐
│ 共享任务列表 │
└─────────────────┘
```
### 与 Subagent 的区别
| 特性 | Subagent | Agent Teams |
| ------------ | ----------------- | --------------------- |
| **上下文** | 独立上下文,结果返回主 Agent | 独立上下文,完全独立运行 |
| **通信方式** | 只能向主 Agent 汇报 | Teammates 之间可以直接通信 |
| **任务协调** | 主 Agent 管理所有工作 | 共享任务列表,自行协调 |
| **适用场景** | 只需要结果的聚焦任务 | 需要讨论和协作的复杂工作 |
| **Token 成本** | 较低:结果摘要返回主上下文 | 较高:每个 Teammate 都是独立实例 |
简单来说:**Subagent 是你派出去执行任务的承包商,Agent Teams 是坐在同一个房间里协作的项目团队**。
### 为什么 Agent Teams 有效
核心洞察:**专业化带来专注**。
单个 Agent 处理复杂多步骤任务时,上下文会不断膨胀,经常需要 `/clear` 重置。Agent Teams 让每个 Teammate 保持狭窄的专注领域,上下文保持干净,性能更稳定。
## 启用 Agent Teams
Agent Teams 目前是实验性功能,默认关闭。需要手动启用:
**方式一:环境变量**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**方式二:settings.json(推荐,永久生效)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## 核心用法
### 创建第一个 Agent Team
启用后,只需用自然语言告诉 Claude 创建团队:
```
创建一个 agent team 来审查 PR #142。
生成三个审查者:
- 一个专注安全问题
- 一个检查性能影响
- 一个验证测试覆盖率
让他们各自审查并汇报发现。
```
**关键词提示**:使用 "create an agent team" 或 "spawn an agent team"。如果只说 "spawn agents",可能会混淆 Subagent 和 Agent Teams。
### 显示模式
Agent Teams 支持两种显示模式:
| 模式 | 说明 | 要求 |
| --------------- | ------------------- | ---------------- |
| **In-process** | 所有 Teammates 在主终端运行 | 无特殊要求 |
| **Split panes** | 每个 Teammate 独立窗格 | 需要 tmux 或 iTerm2 |
默认是 `auto`:如果在 tmux 中运行则使用 split panes,否则用 in-process。
**配置显示模式**:
```json
{
"teammateMode": "in-process"
}
```
**单次会话指定**:
```bash
claude --teammate-mode in-process
```
### 常用快捷键
| 操作 | 快捷键 |
| -------------------- | ------------ |
| 在 Teammates 间切换 | `Shift+Down` |
| 返回上一个 Teammate | `Shift+Up` |
| 切换任务列表显示 | `Ctrl+T` |
| 中断当前 Teammate | `Escape` |
| **启用 Delegate Mode** | `Shift+Tab` |
| 进入 Teammate 会话 | `Enter` |
### Delegate Mode(重要)
Delegate Mode 是 Agent Teams 最重要的功能之一:
| 模式 | Lead 行为 |
| ----------------- | -------------------- |
| **普通模式** | Lead 可能自己实现任务、写代码 |
| **Delegate Mode** | Lead 只能协调,不能写代码、运行测试 |
**为什么需要 Delegate Mode**:
没有这个限制,Lead 经常会"抢活干"——明明有三个 Teammates 等着工作,Lead 却自己开始写代码。开启 Delegate Mode 后,Lead 被强制成为纯粹的项目经理,只能管理任务、与 Teammates 沟通、审查输出。
```
# 启动团队后立即按 Shift+Tab 启用
```
## 实战案例
### 案例一:并行代码审查
单个审查者容易在某类问题上陷入深挖。将审查维度拆分成独立领域,安全、性能、测试覆盖都能得到同等关注:
```
创建 agent team 审查这个 PR。生成三个审查者:
- 安全审查者:检查认证、授权、注入漏洞
- 性能审查者:分析算法复杂度、数据库查询、缓存策略
- 测试审查者:验证测试覆盖率、边界条件、错误处理
让他们各自审查后,相互讨论发现的问题。
```
### 案例二:竞争性假设调试
当根本原因不明时,单个 Agent 倾向于找到一个看似合理的解释就停止。让 Teammates 相互挑战可以避免这个问题:
```
用户反馈应用在发送一条消息后就退出了,而不是保持连接。
生成 5 个 agent teammates 调查不同的假设。让他们相互讨论,
尝试反驳彼此的理论,像科学辩论一样。把达成共识的发现
更新到调查报告中。
```
**关键机制**:辩论结构。多个独立调查者积极尝试推翻彼此的理论,最终存活下来的假设更可能是真正的根本原因。
### 案例三:内容批量生产
这是非技术任务的典型应用——将一个输入转化为多个输出:
```
创建 agent team 将这个视频脚本转化为四个平台的内容:
- LinkedIn 文章作者
- Twitter 线程作者
- Newsletter 作者
- 博客文章作者
脚本位置:/content/scripts/video-20.md
```
每个 Teammate 独立创作,但保持内容一致性。
### 案例四:QA 质量检查集群
一个博客网站的质量检查,部署 5 个 Agent 并行测试不同方面:
```
创建 agent team 对博客进行全面质量检查:
- Agent 1:核心页面测试(首页、关于页、联系页)
- Agent 2:文章页面测试(渲染、导航、SEO 元数据)
- Agent 3:链接检查(内部链接、外部链接、死链)
- Agent 4:SEO 验证(标题、描述、结构化数据)
- Agent 5:可访问性测试(ARIA 标签、对比度、键盘导航)
生成按优先级排序的问题报告。
```
**效果**:几分钟内完成了原本需要人工顺序执行的全面检查,每个 Agent 专注自己的领域,最后汇总成优先级排序的问题列表。
### 案例五:多轮讨论模式
一个有用的提示词模式——让 Teammates 像开会一样讨论:
```
使用 Agent Teams 创建 4 个 teammates 讨论 [技术决策],
进行 3 轮讨论。让 teammates 在每轮中相互交流。
其中一个 teammate 专门负责 Red Team 视角,提出批评意见。
```
这种模式特别适合架构决策、技术选型等需要多角度权衡的场景。
### 案例六:C 编译器项目
Anthropic 用 16 个 Agent 从零开始构建了一个能编译 Linux 内核的 C 编译器:
| 指标 | 数据 |
| -------- | ------------------------ |
| Agent 数量 | 16 个并行实例 |
| 会话数 | \~2,000 个 Claude Code 会话 |
| 成本 | \~$20,000 |
| 代码行数 | 100,000 行 |
| Token 用量 | 20 亿输入 + 1.4 亿输出 |
最终产出:一个能在 x86、ARM、RISC-V 上构建可启动 Linux 6.9 的 Rust 编译器。
## 团队管理
### 指定 Teammates 和模型
Claude 会根据任务自动决定生成多少 Teammates,你也可以明确指定:
```
创建 4 个 teammates 并行重构这些模块。
每个 teammate 使用 Sonnet 模型。
```
### 要求计划审批
对于复杂或高风险任务,可以要求 Teammates 在执行前先制定计划:
```
生成一个架构师 teammate 来重构认证模块。
在他们做任何更改之前,要求计划审批。
```
Teammate 完成计划后会发送审批请求给 Lead。Lead 审核后可以批准或退回修改意见。
### 直接与 Teammates 对话
每个 Teammate 都是完整的 Claude Code 会话。你可以直接发消息给任何 Teammate:
* **In-process 模式**:用 `Shift+Down` 切换,然后输入消息
* **Split-pane 模式**:直接点击对应窗格
### 关闭 Teammates
```
请安全审查 teammate 关闭
```
Lead 会发送关闭请求,Teammate 可以批准或拒绝(并解释原因)。
### 清理团队
完成后,让 Lead 清理资源:
```
清理团队
```
**重要**:始终通过 Lead 清理。不要让 Teammates 执行清理,可能导致资源状态不一致。
## 最佳实践
### 团队规模控制
规模建议:
| 团队大小 | 适用场景 |
| ----- | ------------ |
| 3 个 | 简单的多视角审查 |
| 4-5 个 | 标准的功能开发或重构 |
| 6+ 个 | 大规模迁移或复杂架构任务 |
**经验法则**:每个 Teammate 分配 5-6 个任务比较合适。如果有 15 个独立任务,3 个 Teammates 是不错的起点。
### 任务粒度
* **太小**:协调开销超过收益
* **太大**:Teammates 工作太久没有检查点,浪费风险增加
* **刚好**:独立完整、产出清晰的工作单元(一个函数、一个测试文件、一份审查报告)
### 避免文件冲突
两个 Teammates 编辑同一文件会导致覆盖。拆分工作时确保每个 Teammate 负责不同的文件集:
```
Teammate 1:负责 src/auth/ 目录
Teammate 2:负责 src/api/ 目录
Teammate 3:负责 src/utils/ 目录
```
### 监控和引导
定期检查 Teammates 进度,及时纠正不合适的方向。让团队无人监管运行太久会增加浪费风险。
如果 Lead 开始自己实现任务而不是等待 Teammates:
```
等待你的 teammates 完成任务后再继续
```
### 给足上下文
Teammates 会自动加载项目上下文(CLAUDE.md、MCP servers、skills),但不会继承 Lead 的对话历史。在生成时提供足够的任务细节:
```
生成一个安全审查 teammate,提示如下:
"审查 src/auth/ 目录的认证模块安全漏洞。
重点关注 token 处理、会话管理、输入验证。
应用使用 JWT token 存储在 httpOnly cookies 中。
报告发现时附带严重性评级。"
```
### 自报告验证模式
在任务描述中包含明确的验证标准,确保 Teammates 在完成后自我检查:
```
在任务完成时,向 Lead 报告:
1. 你检查了哪些文件
2. 发现了什么问题
3. 你做了什么修改
4. 验证标准(测试通过、lint 无警告等)是否满足
```
这种模式减少了 Lead 的验证工作量,同时确保任务真正完成而非"看起来完成"。
## 高级技巧
### 使用 Hooks 强制质量门禁
通过 Hooks 在 Teammates 完成工作时强制执行规则:
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle`:Teammate 即将空闲时运行。返回 exit code 2 可以发送反馈并让 Teammate 继续工作
* `TaskCompleted`:任务标记完成时运行。返回 exit code 2 可以阻止完成并发送反馈
### 预审批权限
Teammate 的权限请求会冒泡到 Lead,可能造成频繁中断。在生成前预先批准常见操作:
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### 与 Worktree 结合
Agent Teams 可以和 Worktree 配合使用,每个 Teammate 在自己的 worktree 中工作:
```
创建 agent team,每个 teammate 在独立的 worktree 中工作,
避免文件冲突。
```
### 第三方编排工具
除了原生 Agent Teams,社区也开发了一些编排工具:
| 工具 | 说明 |
| --------------- | -------------------- |
| **Gas Town** | 管理多个并行 Claude 会话的工具 |
| **Multiclaude** | 在多个终端窗口中运行 Claude 实例 |
这些工具在 Agent Teams 实验性功能之外提供了替代方案,但需要更多手动配置。如果原生 Agent Teams 满足需求,建议优先使用官方功能。
## 当前限制
Agent Teams 仍是实验性功能,了解限制很重要:
| 限制 | 说明 |
| -------------------------- | ----------------------------------------------- |
| 无法恢复 in-process teammates | `/resume` 和 `/rewind` 不会恢复 in-process teammates |
| 任务状态可能滞后 | Teammates 有时忘记标记任务完成 |
| 关闭可能较慢 | Teammates 会完成当前请求后才关闭 |
| 每会话一个团队 | Lead 一次只能管理一个团队 |
| 无嵌套团队 | Teammates 不能生成自己的团队 |
| Lead 固定 | 创建团队的会话就是 Lead,无法转移 |
| Split panes 需要 tmux/iTerm2 | 不支持 VS Code 终端、Windows Terminal、Ghostty |
| Plan mode 会话级别 | Teammate 的 Plan mode 状态在生成时锁定,会话中无法更改 |
## 成本考量
Agent Teams 的 Token 消耗显著高于单会话:
| 场景 | Token 消耗 | 成本倍数 |
| ----------------------- | ------------- | ------------- |
| 单 Agent 会话 | \~200k tokens | 1x |
| 3 个 Teammates | \~800k tokens | \~4x |
| 5 个 Teammates | \~1.2M tokens | \~6x |
| 16 个 Teammates(C 编译器案例) | 20 亿 tokens | $20,000 / 2 周 |
**成本分析**:
* 每个 Teammate 是完全独立的 Claude 实例,有自己的上下文
* Teammates 之间的通信也消耗 Token
* Lead 需要协调所有 Teammates,额外增加开销
**何时值得**:
* ✅ 需要并行探索的研究任务
* ✅ 多视角审查(安全、性能、测试)
* ✅ 需要讨论达成共识的决策
* ❌ 可以顺序完成的常规任务
* ❌ 不需要相互通信的并行任务(用 Subagent 更省)
## 我的使用心得
### 什么时候用 Agent Teams
我的判断标准:
1. **任务需要多个视角**:不同领域的专业知识(安全 + 性能 + 测试)
2. **需要讨论和共识**:竞争性假设、架构决策
3. **并行探索有价值**:多种实现方案对比
如果只是需要并行执行、不需要相互通信,用 Subagent 或 Worktree 更合适。
### 从研究和审查开始
如果你是 Agent Teams 新手,从不需要写代码的任务开始:审查 PR、研究技术方案、调查 bug。这些任务有清晰的边界,能展示并行探索的价值,同时避免并行实现带来的协调挑战。
### 与其他功能配合
| 组合 | 效果 |
| ---------------------- | ------------------- |
| Agent Teams + Worktree | 每个 Teammate 在隔离环境工作 |
| Agent Teams + Hooks | 自动化质量检查和反馈 |
| Agent Teams + Skills | 每个 Teammate 具备专业能力 |
## 写在最后
Agent Teams 代表了 AI 辅助开发的一个新范式:从"一个 AI 助手"到"一个 AI 团队"。
Anthropic 用 16 个 Agent 花 2 周、$20,000 写出了 10 万行代码的 C 编译器。这个项目的关键经验是:**测试质量比什么都重要**。Agent 会自主解决你给它的问题,所以任务验证器必须近乎完美,否则 Agent 会解决错误的问题。
记住三个核心要点:
| 要点 | 说明 |
| ------ | ------------------------ |
| **协作** | Teammates 可以直接交流,不只是汇报结果 |
| **共享** | 通过共享任务列表协调工作 |
| **监督** | 定期检查进度,及时纠正方向 |
开始很简单:
```
创建一个 agent team 来 [你的任务]
```
***
**相关阅读**:
* [Claude Worktree 完全指南](/docs/notes/claude-worktree) — 理解 Worktree 与 Agent Teams 的配合
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 对比 Subagent 和 Agent Teams 的使用场景
* [Tmux 快速入门指南](/docs/notes/tmux-tutorial) — 使用 Tmux 管理多个 Agent 会话
**参考资料**:
* [Claude Code 官方文档 - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**视频教程**:
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) — 7 个非技术用例演示
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) — 多智能体工作流详解
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) — 团队编排深度解析
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) — 完整设置教程
# Claude 系统架构全解析
## 引言
2025 年 9 月,Anthropic 以 **$183B 估值**完成 $13B 融资,成为全球第四大私营企业。其明星产品 Claude Code 自 2 月发布以来,已吸引 **11.5 万**活跃开发者,每周处理 **1.95 亿行**代码,用户增长 **300%**。
更有趣的是,Anthropic CEO Dario Amodei 透露:**90% 的 Claude Code 代码由它自己编写**。
**凭什么?**
一个 AI 编程助手,如何做到"自己写自己"?它的架构设计有什么独特之处,让它能够如此高效地辅助——甚至替代——人类开发者?
答案藏在 Claude 的**模块化架构**中:MCP 提供工具、Skills 教会用法、Subagents 并行执行、Hooks 确保可控——这些组件协同工作,让 Claude 具备了"像程序员一样工作"的能力。
这篇文档将带你**俯瞰这套架构的全貌**——每个组件的定位、它们之间的协同关系、以及快速上手的配置示例。后续文章会深入拆解各个组件的细节。
## 整体架构概览
Claude 系统采用**模块化架构**设计,各组件按功能分类,**互补协作**而非层级依赖:
**核心理解**:这些组件是**同级互补**的扩展能力,而非层级依赖——你可以根据需求自由组合:
| 你想要... | 使用... | 一句话定位 |
| -------------- | ------------- | ------------------------------ |
| 连接外部数据源和服务 | **MCP** | 给 Claude 装上"手脚",访问数据库、API、文件系统 |
| 教 Claude 特定工作流 | **Skills** | 让 Claude "知道"某个领域怎么做事 |
| 并行处理复杂任务 | **Subagents** | 把大任务拆成小任务,多个 Agent 同时干活 |
| 快速触发重复操作 | **Commands** | 一键启动常用工作流,省去重复指令 |
| 确保某些操作必须执行 | **Hooks** | 无论 Claude 怎么决策,这步必须跑 |
***
## 核心运行时
### Agent SDK — 运行时引擎
Agent SDK 是整个 Claude Agent 系统的**运行时核心引擎**,提供:
* **主循环 (Main Loop)**:Agent 的核心工作循环
* **上下文管理**:Token 预算、自动压缩(92% 使用率时触发)
* **工具调度**:决定使用哪个工具、如何执行
* **权限系统**:控制工具访问权限
Agent 的核心工作模式是一个简单的**反馈循环**:
```
收集上下文 → 执行操作 → 验证工作 → 重复
```
***
### Built-in Tools — 核心工具
Claude Agent 内置 20+ 核心工具,分为三类:
| 类别 | 工具 | 说明 |
| ------ | ------------------- | -------------- |
| **读取** | Read, Glob, Grep | 文件读取、模式匹配、内容搜索 |
| **操作** | Write, Edit, Bash | 文件写入、编辑、命令执行 |
| **网络** | WebSearch, WebFetch | 网络搜索、网页抓取 |
这些工具**默认可用**,无需额外配置。Claude 通过这些工具与计算机交互,就像程序员使用 IDE 一样。
***
## 配置与上下文
### CLAUDE.md — 持久化上下文
每次开始新对话,都要重复说明项目背景、编码规范、架构约定……CLAUDE.md 让这些信息**一次配置、自动加载**。
CLAUDE.md 就像项目的 **README for AI**——告诉 Claude 这个项目的背景知识、工作方式和约定。
#### 层级覆盖
Claude 按以下顺序加载 CLAUDE.md,**越具体的优先级越高**:
```
Enterprise (最低)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md) (最高)
```
#### 内容建议
CLAUDE.md 应包含以下核心信息:
| 类别 | 内容示例 |
| -------- | ---------------------------------------------- |
| **技术栈** | Next.js 14 + TypeScript, Tailwind CSS |
| **构建命令** | `npm run dev`, `npm run build`, `npm run test` |
| **代码规范** | 命名约定、Lint 工具配置 |
| **项目结构** | 关键目录的用途说明 |
**关键原则**:保持简洁。CLAUDE.md **每次对话都会加载**,过长会浪费宝贵的 Token。
***
## 打包分发
### Plugins — 可安装单元
团队配置分散、难以共享和标准化。每个人都有自己的一套 Skills、Commands、Hooks……如何统一管理?
Plugins 将 **Skills + Commands + Subagents + Hooks + MCP** 打包为**可安装单元**,实现一键分发、团队标准化。
#### 目录结构
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单(必需)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Hooks 配置
├── .mcp.json # MCP Server 配置
└── README.md # 说明文档
```
#### 配置示例
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "前端开发工具包",
"author": "Your Team"
}
```
```bash
# 安装方式
claude plugin install github:your-org/your-plugin # 从 GitHub
claude plugin install /path/to/plugin # 从本地
```
**相关资源**
| 资源 | 说明 |
| -------------------------------------------------------------------------------- | --------------------------- |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Anthropic 官方插件仓库 |
| [wshobson/agents](https://github.com/wshobson/agents) | ⭐ 24.3k,高质量 Agent 模板集合 |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | 社区插件市场,可发现各类插件 |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | ⭐ 19.3k,精选 Claude Code 资源列表 |
***
## 扩展能力(互补模块)
Claude 系统的扩展能力由多个**互补模块**组成,它们各司其职、协同工作:
| 模块 | 功能定位 | 激活方式 |
| ------------- | ------------------------ | --------- |
| **MCP** | 连接外部数据和服务 (WHAT) | 配置后可用 |
| **Skills** | 程序性知识,教 Claude 怎么做 (HOW) | 自动匹配 |
| **Subagents** | 独立上下文,并行任务委派 | 显式调用 |
| **Commands** | 重复性工作流 | 手动 `/cmd` |
| **Hooks** | 确定性控制,事件驱动 | 自动触发 |
***
### MCP — 外部连接
#### 设计理念
传统方式下,每个外部数据源都需要**自定义集成**,导致 N×M 的集成地狱。MCP 提供标准化协议,实现**一次接入、处处可用**。
MCP (Model Context Protocol) 被设计为 **AI 应用的 USB-C 接口**:
| 特性 | 说明 |
| -------- | -------------------------------------- |
| **开放标准** | 2024 年 11 月发布,2025 年 12 月捐赠给 Linux 基金会 |
| **行业采用** | OpenAI、Microsoft、Google、AWS 等已采用 |
| **生态规模** | 97M+ 月 SDK 下载量,数千个社区 Server |
#### 架构模式
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, AI 工具)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (协议客户端)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (暴露特定能力)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**使用场景**:连接数据库、集成第三方服务(GitHub、Slack、Notion)、访问私有 API、实时数据流处理。
#### 配置示例
在项目根目录创建 `.mcp.json`:
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — 手动工作流
Slash Commands 提供**手动触发**的重复性工作流。
| 特性 | 说明 |
| -------- | -------------------- |
| **触发方式** | 手动输入 `/command-name` |
| **存放位置** | `.claude/commands/` |
| **用途** | 重复性工作流、标准化操作 |
**示例**:创建 `.claude/commands/review.md`
```markdown
请对当前变更进行代码审查,重点关注:
1. 代码风格和一致性
2. 潜在的性能问题
3. 安全漏洞
4. 测试覆盖率
```
然后输入 `/review` 即可触发。
***
### Hooks — 确定性控制
Hooks 是**确定性控制**的核心——某些操作必须执行,不能依赖 LLM 的判断。
| 类别 | 事件 | 触发时机 |
| ------- | -------------------- | ------------- |
| **工具** | `PreToolUse` | 工具执行前 |
| | `PostToolUse` | 工具成功执行后 |
| | `PostToolUseFailure` | 工具执行失败后 |
| | `PermissionRequest` | 权限请求时 |
| **会话** | `SessionStart` | 会话开始时 |
| | `SessionEnd` | 会话结束时 |
| | `Stop` | Claude 完成响应时 |
| **子代理** | `SubagentStart` | 子代理启动时 |
| | `SubagentStop` | 子代理停止时 |
| **其他** | `UserPromptSubmit` | 用户提交 prompt 后 |
| | `Notification` | 通知事件 |
| | `PreCompact` | 上下文压缩前 |
#### 配置示例
自动格式化 TypeScript 文件:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### 实战:自主循环
**Ralph Wiggum** 是 Anthropic 官方插件,利用 Stop hook 实现自主迭代循环:
```bash
/ralph-loop "实现 TODO API,包含 CRUD 和测试" --max-iterations 20
```
**工作原理**:Stop hook 拦截 Claude 的退出 → 重新注入原始 prompt → 继续迭代,直到任务完成或达到最大迭代次数。
**适用场景**:需要多轮迭代的任务(测试通过、代码重构)、有自动验证手段的任务。
***
### Subagents — 任务委派与并行执行
#### 设计理念
单一 Agent 面临的挑战:上下文窗口有限、无法并行、职责不清。Subagents 采用 **Orchestrator-Worker** 架构模式解决这些问题:
```
Main Agent (Orchestrator)
├── 分析用户请求
├── 制定计划
├── 分解任务
└── 生成专门化子代理
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - 并行执行)
│ 代码 │ 测试 │ 文档 │
└────────┴────────┴────────┘
↓
汇总结果 → 主代理综合输出
```
#### 核心特性
| 特性 | 说明 |
| ---------- | ---------------------- |
| **上下文隔离** | 每个 Subagent 独立上下文,避免污染 |
| **任务专门化** | 自定义系统提示定义专属角色 |
| **工具权限控制** | 可限制 Subagent 只能使用特定工具 |
| **并行执行** | 多个 Subagent 同时工作 |
**性能数据**:多代理系统比单代理高出 90.2% 性能,并行化可削减研究时间 90%(Token 消耗约 15×,但复杂任务值得)。
#### 配置示例
在 `.claude/agents/` 创建 Markdown 文件:
```markdown
---
name: Code Reviewer
description: 专门进行代码审查的子代理
tools:
- Read
- Grep
- Glob
---
你是一位资深代码审查专家。请重点关注:
1. 代码质量和可维护性
2. 潜在的 Bug 和边界情况
3. 性能优化机会
4. 安全漏洞
```
***
### Skills — 程序性知识
#### 设计理念
Skills 是给 AI 的**可重用工作手册**——模块化的知识包,Claude 可以按需动态加载。核心设计原则是**渐进式披露 (Progressive Disclosure)**:
```
📚 Skills 工作手册
│
├─ 📋 目录 ────────────── 【元数据层】启动时预加载 (~30-50 tokens)
│ name: "weekly-report"
│ description: "生成标准化周报"
│
├─ 📖 正文章节 ─────────── 【核心文档层】相关时加载 (~数百-数千 tokens)
│ # Weekly Report Generator
│ ## Instructions
│ 按以下结构生成周报...
│
└─ 📎 附录 ────────────── 【引用资源层】需要时加载
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills**:MCP 给 Claude 访问工具的能力(WHAT),Skills 教 Claude 如何有效使用这些工具(HOW)。
#### 核心优势
| 优势 | 说明 |
| ------------ | ---------------------------------- |
| **Token 高效** | 元数据仅占 30-50 tokens,可同时启用数十个 Skills |
| **自动激活** | 根据任务上下文自动匹配,无需手动触发 |
| **可组合** | 多个 Skills 自动协同工作 |
| **可移植** | 跨 Claude.ai、Claude Code、API 一致体验 |
#### 配置示例
在 `.claude/skills/` 创建目录:
```
my-skill/
├── SKILL.md # 核心指令(必需)
├── scripts/ # 可执行脚本(可选)
└── references/ # 参考资料(可选)
```
SKILL.md 核心结构:
```yaml
---
name: code-review # Skill 名称
description: 代码审查,检查质量和安全性 # 简短描述(用于自动匹配)
---
# Code Review Skill
## Instructions
[具体的步骤说明...]
## Output Format
[输出格式要求...]
```
**要点**:frontmatter 中的 `description` 用于自动匹配,保持简洁准确。
***
## 官方参考链接
**设计理念**
| 资源 | 说明 |
| --------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Agent 架构纲领性文章 |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Agent SDK 工程实践 |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Skills 设计理念 |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | MCP 发布公告 |
**官方文档**
| 资源 | 说明 |
| ----------------------------------------------------------------------------- | -------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Skills 与其他组件对比 |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | CLAUDE.md 使用指南 |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Subagents 官方文档 |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Hooks 官方文档 |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | MCP 官方文档 |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | MCP 协议规范 |
**深度分析**
| 资源 | 说明 |
| --------------------------------------------------------------------------------------------------------- | ----------------- |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | 包含架构图 |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Skills 原理深入分析 |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Claude Code 构建内幕 |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Skills 与 MCP 技术对比 |
***
## 延伸阅读
如果你想深入了解 Skills 的概念和实践,可以参考:
* [Claude Skills 是什么](/docs/notes/claude-skills/concept) — Skills 核心原理详解
* [Claude Skills 实战指南](/docs/notes/claude-skills/practice) — 动手创建你的第一个 Skill
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 子代理的使用和自定义
* [GSD 深度解析](/docs/notes/gsd/concept) — 基于上下文工程的 AI 编程系统
* [我的 Claude Code 最佳实践](/blog/claude-code-best-practices) — Claude Code 日常使用技巧
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# Claude Worktree 完全指南
## 引言
用 Claude Code 处理复杂任务时,你可能遇到过这样的困境:手上有三个独立的任务需要处理,但在同一个目录里运行多个 Claude 实例会导致代码冲突——一个 Agent 在修改文件,另一个也在动同样的文件,最后合并时一团糟。
2025 年 2 月,Anthropic 发布了 `--worktree` 命令,彻底改变了这一局面。现在,你可以在三个终端里分别运行 `claude -w feature-1`、`claude -w feature-2`、`claude -w bugfix-1`,三个 Agent 各自在隔离的环境中工作,互不干扰。
## 理解 Worktree
想象你是一个建筑师,同时在设计三个不同的房间。传统方式是在同一张图纸上画,改来改去容易弄乱。Worktree 的做法是给你三张独立的图纸,每张专门用于一个房间的设计,最后再合并到主图纸上。
从技术角度来说,Worktree 是 Git 的原生功能。Claude Code 的 `--worktree` 命令将这个功能封装得更加简单易用——一条命令就能创建隔离环境、启动 Claude 实例、完成后自动清理。
### 为什么不直接多次 Clone
你可能会问:为什么不直接 clone 多份代码?
| 方案 | 磁盘占用 | 同步难度 | 清理复杂度 |
| ------------------- | --------------- | -------------- | --------------------- |
| 多次 Clone | 每份都是完整仓库 | 需要手动 pull/push | 需要手动删除目录 |
| Git Worktree | 只复制工作文件,共享 .git | 自动共享历史 | `git worktree remove` |
| Claude `--worktree` | 只复制工作文件,共享 .git | 自动共享历史 | 退出时自动清理 |
Worktree 共享同一个 `.git` 数据库,所有的提交历史、分支信息都是共享的。这意味着在一个 worktree 里创建的 commit,其他 worktree 立刻可见。
### 什么时候该用 Worktree
在动手之前,先判断一下你的任务是否适合用 worktree。
经验法则:**如果任务需要超过 30 分钟,考虑用 worktree**。短任务用 worktree 反而浪费时间——创建环境、安装依赖、最后合并,加起来可能比任务本身还久。但对于需要深度工作的任务,worktree 的隔离性就非常有价值了。
| 适合用 Worktree | 不太适合 |
| ------------ | ------------- |
| 独立的功能开发 | 10 分钟就能完成的小改动 |
| 不同模块的并行重构 | 需要频繁交互的任务 |
| 长时间运行的任务 | 强依赖其他正在进行的修改 |
| 需要隔离测试的实验性改动 | 简单的 bug fix |
### 前置条件
使用 worktree 之前,确保满足以下条件:
| 条件 | 说明 |
| ------------ | -------------------------- |
| Git 已初始化 | 必须在 Git 仓库目录中(有 `.git` 目录) |
| 至少有一个 commit | 空仓库无法创建 worktree |
| 远程分支可用 | 默认从远程分支检出(如 `origin/main`) |
## 完整工作流:从创建到清理
下面按实际开发顺序,走一遍从创建 Worktree 到最终清理的完整流程。
### 第一步:创建 Worktree
#### 从远端默认分支创建
使用 `-w` 或 `--worktree` 参数启动 Claude:
```bash
# 创建名为 "feature-auth" 的 worktree 并启动 Claude
claude -w feature-auth
# 自动生成随机名称(如 "bright-running-fox")
claude -w
```
这条命令实际上做了四件事:
1. 在 `/.claude/worktrees/feature-auth/` 创建新的工作目录
2. 创建名为 `worktree-feature-auth` 的新分支
3. 从远端默认分支(如 `origin/main` 或 `origin/master`)检出代码——**注意,不是你当前所在的分支**
4. 在新目录中启动 Claude Code
所有 worktree 都在 `.claude/worktrees/` 目录下:
```
your-project/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← 第一个 worktree
│ ├── bugfix-123/ ← 第二个 worktree
│ └── refactor-api/ ← 第三个 worktree
├── src/
└── package.json
```
建议将这个路径添加到 `.gitignore`:
```bash
# .gitignore
.claude/worktrees/
```
#### 从当前/特定分支创建
`-w` 总是从远端默认分支检出,目前不支持指定基础分支。如果你想基于当前分支(或某个特定分支)创建 worktree,有三种方式:
**方式一:手动 Git 创建**
```bash
# 基于当前 HEAD 创建 worktree
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# 或者基于某个特定分支
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# 然后在该目录中启动 Claude
cd .claude/worktrees/my-feature && claude
```
这种方式给你完全的控制权——可以基于任意分支、任意 commit 创建 worktree,工作目录从一开始就在正确的分支上。官方文档也建议:"如果需要更多分支和位置控制,直接使用 Git 创建 worktree,然后在该目录中运行 Claude。"
**方式二:在对话中创建(推荐)**
在已有的 Claude 会话中,直接让 Claude 创建 worktree:
```
> 从当前分支开启一个worktree
> start a worktree
```
与 `-w` 命令不同,在对话中创建的 worktree **会自动基于当前分支**,而不是远端默认分支。Claude 会自动完成 worktree 创建并切换过去,整个过程不需要你手动操作任何 Git 命令。如果你已经在某个 feature 分支上工作,这是最便捷的方式——一句话就能开出一个基于当前分支的隔离环境。
**方式三:先用 `-w` 创建,在会话中切换分支**
先用 `claude -w` 创建 worktree,进入会话后再让 Claude 切换到目标分支。这种方式的缺点是会先拉取远端默认分支,再切换——多一步操作,不如前两种方式干净。而且如果目标分支已被其他 worktree 占用,会遇到分支冲突:
如截图所示,Claude 会检测到分支冲突并提供两个选择:回到主目录操作,或基于目标分支创建一个新的工作分支。虽然最终也能工作,但整个过程不如方式一和方式二直接。
**进阶:用 Makefile 封装成一键命令**
如果你经常需要从当前分支创建 worktree,可以在项目根目录的 `Makefile` 中添加一个快捷命令,把创建 + 打开编辑器 + 启动 Claude 串成一条确定性的流水线:
```makefile
# 从当前分支创建 worktree 并启动开发环境
# 用法: make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "用法: make worktree name="; \
echo "示例: make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ 从 $$(git branch --show-current) 创建 worktree: $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ 初始化环境"
cd .claude/worktrees/$(name) && npm install
@echo "→ 在 Zed 中打开"
zed .claude/worktrees/$(name)
@echo "→ 启动 Claude"
cd .claude/worktrees/$(name) && claude
```
使用起来非常简洁:
```bash
# 基于当前分支创建 worktree,用 Zed 打开,启动 Claude
make worktree name=fix-login-bug
# 多开几个并行任务
make worktree name=feature-search
make worktree name=refactor-api
```
相比手动输入多条 Git/cd/claude 命令,`make worktree name=xxx` 只需一行,而且每次执行的流程完全一致——不会忘记某个步骤,也不会打错路径。需要注意的是,因为 Makefile 使用的是原生 `git worktree add`,Claude Code 的 `WorktreeCreate` Hook 不会触发(该 Hook 只在 `claude -w` 或对话中创建 worktree 时生效)。所以环境初始化步骤(安装依赖、复制 `.env` 等)需要直接写在 Makefile 里,如上面示例中的 `npm install`。
### 第二步:初始化环境
Worktree 创建完成后,第一件事就是初始化开发环境。每个新 worktree 都是一个独立的目录,`node_modules`、虚拟环境、`.env` 文件等不会自动带过来。
Claude Code 提供了 `WorktreeCreate` Hook 来自动化环境设置:
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
这样每次创建 worktree 时,依赖会自动安装,环境变量文件会自动复制。常见的初始化步骤:
| 项目类型 | 初始化命令 |
| ------- | ----------------------------------------- |
| Node.js | `npm install` 或 `yarn` |
| Python | `pip install -r requirements.txt` 或激活虚拟环境 |
| Go | `go mod download` |
| 通用 | 复制 `.env` 文件、设置环境变量 |
如果没有配置 Hook,也可以在每个 worktree 会话开始时运行 `/init`,确保 Claude 正确理解当前工作目录的上下文,重新读取项目结构和 CLAUDE.md 配置。
### 第三步:提交与合并
环境就绪、开发完成后,下一步是把改动合并回目标分支。
**合并回 main 分支**
最常见的情况——worktree 从 `origin/main` 分出,改动也要合并回 `main`。在 worktree 的 Claude 会话中直接说:
```
> 提交所有改动,推送到远端,然后创建一个 PR 到 main
```
Claude 会自动完成 commit → push → `gh pr create` 的全流程。
**合并回 feature 分支**
如果你在 `feature-x` 分支上开发,worktree 的改动需要合并回 `feature-x` 而不是 `main`:
```
> 提交改动并推送,然后创建一个 PR 合并到 feature-x 分支
```
Claude 会执行 `gh pr create --base feature-x`,直接创建指向 feature 分支的 PR。
也可以退出 worktree 会话后(选择保留 worktree),回到主目录启动 Claude:
```
> 把 worktree-my-task 分支的改动合并到当前分支
```
如果 worktree 中有些 commit 你不想要,可以选择性地 cherry-pick:
```
> 帮我查看 worktree-my-task 分支的 commit 历史,然后把其中关于认证模块的那几个 commit cherry-pick 到当前分支
```
> **提示**:所有 worktree 共享同一个 `.git` 数据库,worktree 中创建的 commit 在主目录中立刻可见,不需要额外的 push/pull 操作。
### 第四步:退出与清理
代码合并完成后,就可以退出 worktree 会话了。
退出 worktree 会话时,Claude 会根据情况自动处理:
| 状态 | 处理方式 |
| ---------- | ----------------- |
| **无更改** | 自动删除 worktree 和分支 |
| **有更改或提交** | 提示你选择保留或删除 |
保留的 worktree 会一直存在,方便你稍后继续工作。
你也可以配置 `WorktreeRemove` Hook 来自动化清理:
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**手动管理命令**
如果需要手动管理 worktree,可以使用标准的 Git 命令:
```bash
# 列出所有 worktree
git worktree list
# 手动删除某个 worktree
git worktree remove .claude/worktrees/feature-auth
# 清理过时的 worktree 引用
git worktree prune
```
> **注意**:不要直接 `rm -rf` 删除 worktree 目录。正确做法是使用 `git worktree remove`,或者如果已经误删了,运行 `git worktree prune` 来清理残留的引用。
## 并行开发模式
掌握了基本工作流后,来看看如何利用 worktree 实现并行开发。
### 多终端并行
最常见的用法是在多个终端标签页中同时运行:
```bash
# 终端 1:处理用户认证功能
claude -w feature-auth
# 终端 2:修复支付 bug
claude -w bugfix-payment
# 终端 3:重构 API 模块
claude -w refactor-api
```
每个 Claude 实例都在自己的 worktree 中工作,修改不会相互影响。你可以:
* 在一个终端里让 Claude 开发新功能
* 在另一个终端里让 Claude 修 bug
* 在第三个终端里继续你自己的代码审查
### 竞争性实现
一个高效的用法是让多个 Agent 独立实现同一个功能:
```bash
# 三个终端分别运行
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
给它们相同的需求说明,让它们各自实现。最后比较三个方案,选择最好的那个合并。这利用了 LLM 的非确定性——相同的输入可能产生不同的输出,有时候第二个版本反而更好。
UI 设计探索也非常适合这种模式。假设你想重新设计应用的界面,但不确定哪种风格更好:
```bash
# 让三个 Agent 分别实现不同风格
claude -w ui-minimal # 极简风格
claude -w ui-colorful # 鲜艳配色
claude -w ui-glassmorphism # 毛玻璃风格
```
完成后同时运行三个版本的开发服务器(不同端口),并排比较效果,选择最满意的方案合并到主分支——比传统的"改一版、看效果、不满意再改"的流程高效得多。
### Subagent 隔离
Worktree 不仅适用于主 Claude 实例,还可以用于 Subagent。在自定义 Subagent 的 frontmatter 中添加 `isolation: worktree`:
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
也可以在对话中直接告诉 Claude:
```
> 使用 worktree 来隔离你的 agents
> use worktrees for your agents
```
当 Subagent 配置了 worktree 隔离:
```
主 Agent(主目录)
│
├── 启动 Migration Agent 1 ──→ worktree-migration-1/
│ └── 处理 src/auth/ 目录
│
├── 启动 Migration Agent 2 ──→ worktree-migration-2/
│ └── 处理 src/api/ 目录
│
└── 启动 Migration Agent 3 ──→ worktree-migration-3/
└── 处理 src/utils/ 目录
```
每个 Subagent 在自己的 worktree 中独立工作,互不干扰。完成后,worktree 会自动清理(如果没有未提交的更改)。
### 结合 Tmux 和 IDE
配合 `--tmux` 参数可以自动在新的 Tmux 会话中启动,即使关闭终端,Claude 也会在后台继续运行:
```bash
claude -w feature-auth --tmux
```
如果你使用 VS Code 或 Cursor,源代码控制面板会自动识别所有 worktree——主仓库显示为一个 repo,每个 worktree 显示为独立的 repo,可以直接在 IDE 中切换、提交、推送。Worktree 也可以和 [Ralph 循环](/docs/notes/ralph-wiggum/concept)结合,每个 Ralph 循环在自己的 worktree 中运行,即使循环失败也不会影响主分支。
## 注意事项与最佳实践
### 常见陷阱
1. **分支来源容易搞错**:`-w` 创建的 worktree 是从**远端默认分支**检出的,不是你当前所在的分支。如果你在 `feature-x` 分支上运行 `claude -w my-task`,新 worktree 的代码来自 `origin/main`,不会包含 `feature-x` 的改动。想基于当前分支工作,参考[从当前/特定分支创建](#从当前特定分支创建)。
2. **未提交的改动不会带过去**:创建 worktree 时,主目录中未暂存或未提交的修改不会出现在新 worktree 中。Worktree 只基于 commit 历史创建,所以确保重要改动已经提交。
3. **同一分支不能被多个 Worktree 使用**:Git 不允许两个 worktree 同时检出同一个分支。如果你在主目录已经在 `feature-x` 分支上,尝试在 worktree 中也检出 `feature-x` 会报错。每个 worktree 必须在不同的分支上。
4. **环境需要重新初始化**:每个新 worktree 不包含 `node_modules` 等运行时依赖,建议配置 `WorktreeCreate` Hook 来自动化(参考[第二步:初始化环境](#第二步初始化环境))。
### 使用建议
不要贪多。虽然技术上可以开很多 worktree,但每个 Claude 实例消耗 API 额度,太多并行任务难以追踪,最终合并时冲突也更复杂。
**命名规范**:养成良好的命名习惯,方便后续管理:
```bash
# 好的命名
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# 不好的命名
claude -w test
claude -w temp
claude -w 1
```
## 非 Git 版本控制
如果你使用 SVN、Perforce 或 Mercurial,可以通过配置 `WorktreeCreate` 和 `WorktreeRemove` Hook 来实现类似的隔离效果。配置这些 Hook 后,使用 `--worktree` 时会调用你自定义的命令,而不是默认的 Git 行为。
## 写在最后
Worktree 是 Claude Code 团队自己每天都在用的功能,Boris Cherny 称它为"排名第一的生产力技巧"。核心价值很简单:**让多个 Agent 能够并行工作而不互相干扰**。
开始很简单,只需要:
```bash
claude -w your-task-name
```
***
**相关阅读**:
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 理解 Subagent 与 Worktree 的配合
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) — 另一种提升 AI 编程效率的方法
* [Claude 系统架构全解析](/docs/notes/claude-architecture) — 理解 Worktree 在整体架构中的位置
**参考资料**:
* [Claude Code 官方文档 - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Boris Cherny 的 Worktree 发布公告](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Git Worktree 官方文档](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**视频教程**:
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — 实战演示 worktree 完整工作流
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — 手动创建 worktree 的方法
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — 原生 worktree 功能详解
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — 快速上手教程,包含 Subagent 用法
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — 多 worktree 并行开发的优势
# 文档
这里是我整理的 AI 编程、Agent、开发工具和工作流笔记。笔记按主题归类,系列文章会放在一起浏览。
# Tmux 快速入门指南
## 引言
如果你用过 Claude Code 的 Agent Teams 或者想同时运行多个 Claude 实例,Tmux 几乎是必备工具。它能让你在一个终端窗口中运行多个会话,会话在后台持续运行即使你关闭了终端,而且 Claude 可以在 Tmux 中自动生成和管理多个 Agent。
这篇教程专为 Claude Code 用户设计,既涵盖 Tmux 基础知识,也重点介绍与 Claude Code 的集成技巧。
## 理解 Tmux
Tmux 的三个核心概念:
```
┌─────────────────────────────────────────────────────────┐
│ Session(会话) │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window(窗口) ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| 概念 | 类比 | 说明 |
| ----------- | ------ | ------------------ |
| **Session** | 工作空间 | 最顶层容器,即使断开连接也会持续运行 |
| **Window** | 浏览器标签页 | 一个会话可以包含多个窗口 |
| **Pane** | 分屏 | 一个窗口可以分割成多个窗格 |
## 安装与基础
### 安装 Tmux
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
验证安装:
```bash
tmux -V
# 输出类似:tmux 3.6a
```
### 前缀键
Tmux 的所有命令都以**前缀键**开始,默认是 `Ctrl+B`。
输入命令的方式:
1. 按下 `Ctrl+B`(不要松开)
2. 松开后,按命令键
例如,分割窗口:`Ctrl+B` 然后按 `%`
## 常用命令速查
### 会话管理
| 命令 | 说明 |
| --------------------------- | ------------ |
| `tmux` | 创建新会话 |
| `tmux new -s name` | 创建命名会话 |
| `tmux ls` | 列出所有会话 |
| `tmux attach -t name` | 连接到会话 |
| `tmux kill-session -t name` | 关闭会话 |
| `Ctrl+B d` | 分离当前会话(后台运行) |
### 窗口管理
| 快捷键 | 说明 |
| ------------ | ------- |
| `Ctrl+B c` | 创建新窗口 |
| `Ctrl+B n` | 下一个窗口 |
| `Ctrl+B p` | 上一个窗口 |
| `Ctrl+B 0-9` | 切换到指定窗口 |
| `Ctrl+B ,` | 重命名当前窗口 |
| `Ctrl+B &` | 关闭当前窗口 |
### 窗格管理
| 快捷键 | 说明 |
| ------------ | -------- |
| `Ctrl+B %` | 垂直分割(左右) |
| `Ctrl+B "` | 水平分割(上下) |
| `Ctrl+B 方向键` | 在窗格间移动 |
| `Ctrl+B x` | 关闭当前窗格 |
| `Ctrl+B z` | 最大化/还原窗格 |
| `Ctrl+B {` | 向左移动窗格 |
| `Ctrl+B }` | 向右移动窗格 |
### 其他常用
| 快捷键 | 说明 |
| ---------- | ----------- |
| `Ctrl+B [` | 进入复制模式(可滚动) |
| `q` | 退出复制模式 |
| `Ctrl+B ?` | 显示所有快捷键 |
## 与 Claude Code 集成
### 为什么 Claude Code 需要 Tmux
1. **Agent Teams 的 Split-pane 模式**:每个 Teammate 显示在独立窗格中
2. **后台运行**:任务继续执行,即使关闭终端
3. **会话持久化**:断线重连后恢复完整上下文
4. **多实例管理**:同时运行多个 Claude 会话
### 基本用法:后台运行 Claude
```bash
# 在 tmux 中启动 Claude
tmux new -s claude-work
claude
# 分离会话(Claude 继续运行)
# Ctrl+B d
# 稍后重新连接
tmux attach -t claude-work
```
### 使用 --tmux 参数
Claude Code 原生支持 Tmux 集成:
```bash
# 在新的 tmux 会话中启动 Claude
claude --tmux
# 配合 worktree 使用
claude -w feature-auth --tmux
```
这会自动:
1. 创建新的 tmux 会话
2. 在其中启动 Claude Code
3. 会话命名为 `claude-{随机ID}`
### Agent Teams 的 Tmux 模式
Agent Teams 可以使用 split-pane 显示模式,每个 Teammate 在独立窗格中运行:
```json
// settings.json
{
"teammateMode": "tmux"
}
```
或通过命令行:
```bash
claude --teammate-mode tmux
```
效果:
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## 实用配置
### 推荐的 \~/.tmux.conf
创建或编辑 `~/.tmux.conf`:
```bash
# 使用 Ctrl+A 作为前缀键(更容易按)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# 启用鼠标支持
set -g mouse on
# 增加历史缓冲区(Claude 输出很多)
set -g history-limit 50000
# 使用 vim 风格的窗格导航
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# 更直观的分割快捷键
bind | split-window -h
bind - split-window -v
# 快速重载配置
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# 256 色支持
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# 窗口编号从 1 开始(0 太远了)
set -g base-index 1
setw -g pane-base-index 1
# 状态栏优化
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
重载配置:
```bash
tmux source-file ~/.tmux.conf
```
### Claude Code 专用配置
针对 Claude Code 的优化配置:
```bash
# Claude 会话弹窗快捷键
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
这个配置的效果:
1. 按 `Ctrl+A y` 打开 Claude 弹窗
2. 每个目录有独立的 Claude 会话
3. 关闭弹窗后会话继续运行
4. 重新打开时恢复之前的对话
## 常见工作流
### 工作流一:多项目并行
```bash
# 为每个项目创建独立会话
tmux new -s project-a
# 在里面启动 Claude
claude -w feature-x
# 分离后,创建另一个会话
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# 在会话间切换
tmux switch -t project-a
tmux switch -t project-b
# 或列出所有会话选择
# Ctrl+B s
```
### 工作流二:开发仪表盘
创建一个多窗格的开发环境:
```bash
# 创建会话
tmux new -s dev
# 分割成三个窗格
# Ctrl+B %(垂直分割)
# Ctrl+B "(水平分割右侧)
# 窗格布局:
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# 在第一个窗格运行 Claude
claude
# 切换到第二个窗格(Ctrl+B 右箭头)
tail -f logs/app.log
# 切换到第三个窗格
npm test -- --watch
```
### 工作流三:远程开发
Tmux 最强大的特性是会话持久化,特别适合 SSH 远程开发:
```bash
# 连接远程服务器
ssh user@server
# 创建 tmux 会话
tmux new -s remote-claude
# 启动 Claude
claude
# 断开 SSH 连接(Claude 继续运行)
# Ctrl+B d
exit
# 稍后重新连接
ssh user@server
tmux attach -t remote-claude
# Claude 会话完整恢复
```
### 工作流四:Agent Teams 监控
使用 tmux 监控 Agent Teams 的所有 Teammates:
```bash
# 启动 Claude 并使用 tmux 模式
claude --teammate-mode tmux
# 创建 Agent Team
# "创建 agent team 审查代码..."
# 此时屏幕自动分割,每个 Teammate 一个窗格
# 可以点击不同窗格直接与对应 Teammate 交流
```
## 故障排除
### 常见问题
| 问题 | 解决方案 |
| ------- | ------------------------ |
| 颜色显示不正确 | 确保 `TERM=xterm-256color` |
| 鼠标不工作 | 添加 `set -g mouse on` 到配置 |
| 复制粘贴问题 | 在复制模式下使用 `Enter` 复制 |
| 会话消失了 | 检查 `tmux ls`,可能是系统重启 |
### 清理孤儿会话
Claude Code 有时会留下未清理的 tmux 会话:
```bash
# 列出所有会话
tmux ls
# 杀死特定会话
tmux kill-session -t session-name
# 杀死所有会话(谨慎!)
tmux kill-server
```
### iTerm2 用户
如果你使用 macOS 的 iTerm2,可以使用其原生集成:
```bash
# 使用 iTerm2 的 tmux 集成模式
tmux -CC
# 或在 Claude Code 中
claude --teammate-mode tmux
```
iTerm2 会自动将 tmux 窗格转换为原生标签页和分屏。
## 我的使用心得
### 什么时候用 Tmux
| 场景 | 是否需要 Tmux |
| --------------- | --------- |
| 简单的单次 Claude 对话 | 不需要 |
| 长时间运行的任务 | 需要 |
| Agent Teams | 强烈推荐 |
| 远程开发 | 必须 |
| 多项目并行 | 推荐 |
### 最小配置
如果你不想折腾配置,只需记住这几个命令:
```bash
# 创建会话
tmux new -s work
# 分离(后台运行)
Ctrl+B d
# 重新连接
tmux attach -t work
# 分割窗格
Ctrl+B % # 左右分割
Ctrl+B " # 上下分割
# 切换窗格
Ctrl+B 方向键
```
### 与 Claude Code 的最佳组合
1. **Worktree + Tmux**:每个 worktree 在独立的 tmux 会话中
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux**:所有 Teammates 可视化管理
```bash
claude --teammate-mode tmux
```
3. **长任务 + 分离**:启动后分离,稍后回来检查
```bash
# 启动
tmux new -s migration
claude
# "执行数据库迁移..."
# Ctrl+B d
# 几小时后
tmux attach -t migration
```
## 写在最后
Tmux 是 Claude Code 高效使用的关键工具,特别是在以下场景:
| 要点 | 说明 |
| ------- | --------------------------- |
| **持久化** | 会话不会因为断开连接而丢失 |
| **并行** | 同时管理多个 Claude 实例 |
| **可视化** | Agent Teams 的 split-pane 显示 |
三个核心命令就能开始:
* `tmux new -s name` 创建会话
* `Ctrl+B d` 分离会话
* `tmux attach -t name` 重新连接
***
**相关阅读**:
* [Claude Agent Teams 完全指南](/docs/notes/claude-agent-teams) — Agent Teams 需要 Tmux 支持 split-pane 模式
* [Claude Worktree 完全指南](/docs/notes/claude-worktree) — Worktree 可以配合 Tmux 后台运行
**参考资料**:
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**视频教程**:
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) — Tmux 基础入门
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) — 与 Claude Code 集成工作流
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) — 高级配置技巧
# MVP 冲刺:两周完成核心功能
这是一篇测试文章。
# 注册 Apple 开发者账号
想把自己开发的 App 上架到 App Store,第一步就是注册 Apple Developer Program(苹果开发者计划)。每年 688 元人民币(99 美元),这是所有 iOS 独立开发者绕不开的一笔投入。
这篇文章会带你了解账号类型的区别、注册前需要准备什么、以及完整的注册流程。
## 账号类型对比
Apple Developer Program 有三种账号类型,适合不同的开发场景:
| 特性 | 个人账号 | 组织账号 | 企业账号 |
| ------------ | ------------- | ----------- | ------------ |
| 年费 | ¥688($99) | ¥688($99) | ¥1,988($299) |
| App Store 上架 | ✅ | ✅ | ❌(仅内部分发) |
| 开发者名称显示 | 个人姓名 | 组织/公司名称 | 组织名称 |
| 团队成员管理 | ❌ | ✅ | ✅ |
| D-U-N-S 编号 | 不需要 | 需要 | 需要 |
| 审核周期 | 较快(通常 48 小时内) | 较慢(需验证组织信息) | 较慢 |
| 适合谁 | 独立开发者、个人 | 公司、工作室 | 大型企业内部应用 |
**独立开发者的选择**:如果你是个人开发者,直接选**个人账号**就好。流程最简单,审核最快,功能完全够用。App Store 上显示的开发者名称会是你的真实姓名。
## 注册前准备
### 必备条件
开始注册之前,确保你已经准备好了以下内容:
* **Apple ID**:如果还没有,去 [appleid.apple.com](https://appleid.apple.com) 创建一个。建议使用常用邮箱注册,后续所有开发相关的通知都会发到这个邮箱。
* **双重认证**:Apple ID 必须开启双重认证(Two-Factor Authentication)。在 iPhone 上进入「设置 → Apple ID → 登录与安全 → 双重认证」开启。
* **Apple 设备**:注册过程需要在 iPhone 或 iPad 上完成身份验证,需要下载 Apple Developer App。
### 组织账号额外要求
如果你注册的是组织账号,还需要:
* **D-U-N-S 编号**:提前去邓白氏官网申请,审核需要 5-14 个工作日。
* **法人身份**:注册人需要是组织的法人或被授权代表。
* **组织信息**:包括注册地址、法人姓名、联系方式等。
## 开发设备准备
注册开发者账号只是第一步,iOS 开发还需要一些硬件和软件工具。
### 必备设备
* **Mac 电脑** — Xcode 只能在 macOS 上运行,这是硬性要求。推荐 Apple Silicon(M 系列芯片)Mac,编译速度快、能直接运行 iOS 模拟器。MacBook Air M 系列即可满足独立开发需求,预算有限可考虑 Mac mini。
* **iPhone / iPad(推荐但非必须)** — 模拟器能覆盖大部分调试场景,但真机测试在性能、传感器(相机/GPS/NFC)、推送通知等方面不可替代。即使没有付费开发者账号,用免费 Apple ID 也可以在真机上调试(但有 7 天重签等限制,详见文末 Q\&A)。
### 开发工具
* **Xcode** — Apple 官方 IDE,从 Mac App Store 免费下载,体积较大(约 12GB+),首次安装需耐心等待。
* **Apple Developer App** — 注册账号、查看 WWDC 视频和文档用。
* **TestFlight** — 内测分发工具,邀请用户测试 App 的正式渠道。
### 注意事项
* macOS 和 Xcode 版本需保持较新,Apple 每年 WWDC 后会发布新版 Xcode,通常要求最近 1-2 个 macOS 大版本。
* Xcode 更新频繁且体积大,建议预留足够磁盘空间(至少 50GB 以上)。
* 如果开发的 App 涉及硬件功能(相机、蓝牙、NFC 等),真机测试必不可少。
* 没有 Mac 的情况下,云 Mac 服务(如 MacStadium、AWS EC2 Mac)是替代方案,但体验不如原生设备。
## 注册流程
### 第一步:下载 Apple Developer App
在 iPhone 或 iPad 上打开 App Store,搜索「Apple Developer」,下载安装。
### 第二步:登录并开始注册
打开 Apple Developer App,使用你的 Apple ID 登录。点击「账户」标签页,然后点击「立即注册 Apple Developer Program」。
### 第三步:填写信息与身份验证
根据提示填写个人信息:
1. **确认身份信息**:姓名、地址等基本信息。
2. **身份验证**:根据所在地区,App 可能会要求你拍摄政府颁发的有效证件(护照、驾照等)或自拍照片进行身份核验。
3. **同意协议**:阅读并同意 Apple Developer Program 许可协议。
> 身份验证环节需要在光线充足的环境下操作,确保拍照清晰。整个注册流程需要在同一台设备上完成。
### 第四步:支付年费
确认信息无误后,支付年费 ¥688($99)。支持 Apple ID 绑定的支付方式。支付成功后会收到确认邮件。
### 第五步:等待审核
* **个人账号**:通常在 48 小时内审核通过。我是 3 月 14 日付的款,3 月 15 日上午就收到了欢迎邮件,前后不到 24 小时。
* **组织账号**:Apple 会验证组织信息和 D-U-N-S 编号,可能需要更长时间。
审核通过后,你就可以在 [developer.apple.com](https://developer.apple.com) 登录开发者后台,访问所有开发资源了。
## 订阅管理与续费
Apple Developer Program 是年度订阅制,每年 ¥688 自动续费。
### 自动续费
默认开启自动续费,到期前会从 Apple ID 绑定的支付方式中扣款。建议保持自动续费,避免账号过期影响已上架 App(详见下方常见问题)。
### 取消或管理订阅
如果需要修改续费设置,打开 iPhone「设置 → Apple ID → 订阅」,找到 Apple Developer Program 进行管理。
## 常见问题
**Q:忘记续费了怎么办?**
账号过期后,你的 App 会从 App Store 下架,但不会被删除。重新付费后,App 会恢复上架。不过期间的用户下载和更新都会受影响,所以建议开启自动续费。
**Q:D-U-N-S 编号怎么申请?**
访问下面这个链接,填写公司信息后提交申请。审核通常需要 5-14 个工作日。个人账号不需要此编号。
**Q:遇到问题怎么联系 Apple?**
访问 [Apple Developer 支持](https://developer.apple.com/contact/),可以通过在线聊天或电话联系。中国区支持中文服务,响应速度还不错。
**Q:能不能先开发再注册?**
可以先体验,但要了解免费账号的限制。用免费 Apple ID 就能在 Xcode 中编写代码、使用模拟器调试,也可以安装到自己的设备上运行。对于学习 Swift 和验证基本 UI 想法来说完全够用。
但免费账号有不少限制:安装到真机的 App 每 7 天需要重新编译安装,每平台最多 3 台设备,而且推送通知、iCloud、TestFlight、App 内购买等功能都无法使用。如果你的 App 需要这些能力,开发阶段就需要付费会员了——不只是上架 App Store 才需要。
建议:如果只是学 Swift 入门、跑跑 Demo,可以先用免费账号;一旦开始做正式项目,尽早注册付费会员,避免后续功能受限耽误进度。
# 想法验证:从模糊灵感到可执行方向
这是一篇测试文章。
# 技术选型:为什么选择 Next.js + Supabase
这是一篇测试文章。
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# Claude Code Plugin 概念介绍
## 引言
当你在 Claude Code 中配置好了一套完美的工作流——自定义命令、代码审查钩子、专属 Skills——你可能会想:能不能把这些打包起来,分享给团队或者社区?
这就是 Plugin 要解决的问题。
如果说 Skills 是给 AI 的"工作手册",那么 Plugin 就是一个"工具箱"——它把 Skills、Commands、Hooks、MCP 服务器等所有配置打包在一起,让你可以一键安装、一键分发。
## 理解 Plugin
想象你是一位经验丰富的工匠,多年来积累了一套趁手的工具:锤子、锯子、尺子、各种螺丝刀。每次换工作台时,你都要一件件搬运这些工具,重新整理。Plugin 就像一个设计精良的工具箱——它不仅能装下所有工具,还能按类别分隔整齐,搬到哪儿都能立刻开工。
从技术角度来说,Plugin 是 Claude Code 的扩展包机制。一个 Plugin 可以包含:
| 组件 | 作用 | 文件位置 |
| ------------- | ---------- | ----------- |
| **斜杠命令** | 快捷操作入口 | `commands/` |
| **Subagents** | 专门化的子代理 | `agents/` |
| **Skills** | AI 的专业知识包 | `skills/` |
| **Hooks** | 事件触发的自动化脚本 | `hooks/` |
| **MCP 服务器** | 外部系统连接 | `.mcp.json` |
| **LSP 服务器** | 语言服务器配置 | `.lsp.json` |
这些组件协同工作,形成一个完整的工作流解决方案。
## Plugin vs 独立配置
在 Claude Code 中,你可以把配置放在项目的 `.claude/` 目录下,也可以打包成 Plugin。两者的核心区别在于**分发方式**和**命名空间**:
| 方面 | 独立配置 (`.claude/`) | Plugin |
| ---- | ----------------- | -------------------- |
| 命令名称 | `/hello` | `/plugin-name:hello` |
| 适用场景 | 个人工作流、项目特定配置 | 团队共享、社区分发 |
| 版本管理 | 随项目代码管理 | 支持语义化版本 |
| 更新方式 | 手动同步 | 支持自动更新 |
| 冲突处理 | 可能与其他配置冲突 | 命名空间隔离 |
**何时选择 Plugin**:
* 需要与团队成员共享工作流配置
* 想要在多个项目中复用同一套工具
* 准备将配置分发到社区
* 需要版本控制和自动更新
**何时使用独立配置**:
* 个人使用的快速实验
* 项目特定的、不需要复用的配置
* 简单的一次性命令
## Plugin 目录结构
一个标准的 Plugin 结构如下:
```
my-plugin/
├── .claude-plugin/ # 元数据目录
│ └── plugin.json # 必需:插件清单
├── commands/ # 斜杠命令
│ ├── review.md
│ └── deploy.md
├── agents/ # 子代理
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # 事件钩子
│ └── hooks.json
├── scripts/ # 辅助脚本
│ └── format-code.sh
├── .mcp.json # MCP 服务器配置
└── .lsp.json # LSP 服务器配置
```
**关键注意事项**:
* `plugin.json` 必须放在 `.claude-plugin/` 目录内
* 其他目录(commands、agents、skills 等)放在插件根目录
* 不要把功能目录放进 `.claude-plugin/` 里
## 核心配置文件
Plugin 的核心是 `.claude-plugin/plugin.json`,它定义了插件的元数据和组件路径:
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "一个示例插件",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| 字段 | 必需 | 说明 |
| ------------- | -- | ----------------- |
| `name` | 是 | 插件唯一标识,使用小写字母和连字符 |
| `version` | 否 | 语义化版本号 |
| `description` | 否 | 插件简短描述 |
| `author` | 否 | 作者信息 |
| `keywords` | 否 | 用于发现的标签 |
| `commands` | 否 | 命令文件或目录路径 |
| `agents` | 否 | 代理文件或目录路径 |
| `skills` | 否 | Skills 目录路径 |
| `hooks` | 否 | 钩子配置路径 |
| `mcpServers` | 否 | MCP 配置路径 |
## 安装范围
Plugin 支持四种安装范围,适应不同的使用场景:
| 范围 | 配置文件 | 用途 |
| --------- | ----------------------------- | --------------- |
| `user` | `~/.claude/settings.json` | 个人插件,所有项目可用 |
| `project` | `.claude/settings.json` | 团队插件,通过版本控制共享 |
| `local` | `.claude/settings.local.json` | 项目特定,gitignored |
| `managed` | `managed-settings.json` | 企业管理(只读) |
默认安装到 `user` 范围。如果你想把插件配置提交到 Git 供团队使用,选择 `project` 范围。
## 核心优势
### 命名空间隔离
Plugin 的命令带有命名空间前缀(如 `/my-plugin:review`),避免了与其他插件或项目配置的命名冲突。这在团队协作时尤为重要——不同团队开发的插件可以和平共处。
### 版本管理
Plugin 支持语义化版本(Semantic Versioning),你可以:
* 追踪插件的变更历史
* 在需要时回滚到旧版本
* 自动接收兼容的更新
### 分发便利
通过 Plugin Marketplace,你可以:
* 将插件托管在 GitHub
* 让用户通过简单命令安装
* 自动处理依赖和更新
### 团队协作
Plugin 特别适合团队场景:
* 统一团队的开发工具链
* 新成员一键获取所有工具
* 配置集中管理,减少重复工作
## Plugin 生态
Claude Code 的 Plugin 生态正在快速发展。截至 2025 年初,生态规模已相当可观:
* **229+ 插件** 正在生态中活跃
* **239 个 Agent Skills** 跨市场分布
* **200+ MCP 服务器** 在 Docker 工具包中预构建
**官方资源**:
| 资源 | 链接 | 说明 |
| ---------------------------------- | ---------------------------------------------------------------------------------- | ---------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 官方 Skills 仓库 |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 官方插件目录 |
| Docker MCP Toolkit | [官网](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ 预构建 MCP 服务器 |
**社区精选**:
| 资源 | 链接 | 说明 |
| ---------------------- | ----------------------------------------------------------------- | -------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 个插件自动收集 |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 最佳实践汇总 |
| claude-plugins.dev | [官网](https://claude-plugins.dev/) | 社区注册表和 CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 代理 + 15 编排器 |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 个专门代理 |
## 与其他功能的关系
Plugin 是一个"容器"概念,它可以包含 Claude Code 生态中的其他功能:
```
Plugin(容器)
├── Skills(知识包)
├── Commands(快捷命令)
├── Agents(子代理)
├── Hooks(事件钩子)
└── MCP/LSP(外部连接)
```
理解这个层次关系很重要:
* **Skills** 教会 Claude 如何做某件事
* **Commands** 提供快捷触发入口
* **Agents** 处理独立的专门任务
* **Hooks** 实现事件驱动的自动化
* **Plugin** 把这些都打包在一起,便于分发和管理
### Skills vs Plugins
初次接触可能会困惑 Skills 和 Plugins 的区别。根据 [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) 的分析:
| 特性 | Skills(技能) | Plugins(插件) |
| -------- | ------------------------- | ----------------- |
| **作用范围** | 所有 Claude 产品(网页、API、Code) | 仅限 Claude Code |
| **包含内容** | Markdown 指南 + 可选脚本 | 命令、代理、钩子、MCP、技能 |
| **激活方式** | 自动(模型决定何时使用) | 可变(取决于组件类型) |
| **最佳用途** | 教授 Claude 领域专业知识 | 扩展 Claude Code 环境 |
| **分发方式** | GitHub 仓库、文件系统 | 去中心化 Marketplace |
**关键洞察**:Skills 由模型自动触发,无需手动调用;Plugins 是打包机制,解决了分布式共享的难题。两者可以结合使用——Plugin 可以包含 Skills。
## 小结
Claude Code Plugin 本质上是一个**工作流打包分发机制**。它解决了配置复用和团队协作的痛点,让你可以把精心打造的工具链分享给更多人。
记住三个关键词:
| 关键词 | 含义 |
| ------ | ------------------- |
| **打包** | 将多种配置组件整合为一个单元 |
| **隔离** | 命名空间避免冲突 |
| **分发** | 通过 Marketplace 轻松共享 |
了解了概念之后,下一篇《[Claude Code Plugin 实战指南](/docs/notes/claude-plugin/practice)》将带你动手实践:从零创建 Plugin、发布到 Marketplace、以及团队协作的最佳实践。
如果你还不熟悉 Plugin 可以包含的组件,建议先阅读《[Claude Skills 是什么](/docs/notes/claude-skills/concept)》了解 Skills 的核心概念。
# Claude Code Plugin 实践指南
## 快速回顾
在上一篇文章中,我们了解了 Plugin 的核心概念:它是 Claude Code 的工作流打包分发机制,可以将 Commands、Skills、Agents、Hooks 等组件整合在一起,便于团队共享和社区分发。本文将从实战角度出发,带你完成从创建到发布的完整流程。
## 创建你的第一个 Plugin
### 第一步:创建目录结构
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### 第二步:创建插件清单
在 `.claude-plugin/plugin.json` 中定义插件的元数据:
```json
{
"name": "my-first-plugin",
"description": "我的第一个 Claude Code 插件",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### 第三步:添加斜杠命令
在 `commands/` 目录下创建 Markdown 文件。每个文件对应一个命令:
`commands/hello.md`:
```markdown
---
description: 向用户发送友好的问候
---
# Hello 命令
请热情地问候用户,并询问今天可以帮助他们做什么。
```
### 第四步:测试插件
使用 `--plugin-dir` 标志加载本地插件进行测试:
```bash
claude --plugin-dir ./my-first-plugin
```
在 Claude Code 中运行命令:
```
/my-first-plugin:hello
```
### 第五步:添加命令参数
命令支持接收用户输入的参数。更新 `hello.md`:
```markdown
---
description: 向指定用户发送个性化问候
---
# Hello 命令
请热情地问候名为 "$ARGUMENTS" 的用户,并询问今天可以帮助他们做什么。
如果用户没有提供名字,就使用"朋友"作为称呼。
```
测试带参数的命令:
```
/my-first-plugin:hello 小明
```
**支持的参数占位符**:
* `$ARGUMENTS` - 所有用户输入
* `$1`, `$2`, `$3` - 单独的参数
## 添加更多组件
### 添加 Skills
创建 `skills/` 目录,每个 Skill 是一个包含 `SKILL.md` 的文件夹:
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md`:
```yaml
---
name: code-review
description: 审查代码质量、安全性和可维护性
---
当审查代码时,请检查以下方面:
1. **代码组织**:结构是否清晰
2. **错误处理**:异常是否被妥善处理
3. **安全隐患**:是否存在安全漏洞
4. **测试覆盖**:关键逻辑是否有测试
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### 添加 Subagents
创建 `agents/` 目录:
`agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家。
当被调用时:
1. 运行 git diff 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(如注入、敏感信息泄露)
- 性能优化机会
```
### 添加 Hooks
Hooks 让你在特定事件发生时自动执行脚本。创建 `hooks/hooks.json`:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**重要**:使用 `${CLAUDE_PLUGIN_ROOT}` 环境变量引用插件目录中的文件,确保路径在任何安装位置都能正确解析。
创建对应的脚本 `scripts/format-code.sh`:
```bash
#!/bin/bash
# 格式化代码
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
记得设置执行权限:
```bash
chmod +x scripts/format-code.sh
```
### 添加 MCP 服务器
如果你的插件需要连接外部系统,创建 `.mcp.json`:
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## 完整的 Plugin 结构
一个功能完整的 Plugin 可能是这样的:
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单
├── commands/
│ ├── review.md # 代码审查命令
│ ├── deploy.md # 部署命令
│ └── test.md # 测试命令
├── agents/
│ ├── code-reviewer.md # 代码审查代理
│ └── debugger.md # 调试代理
├── skills/
│ └── code-standards/
│ └── SKILL.md # 代码规范知识
├── hooks/
│ └── hooks.json # 事件钩子配置
├── scripts/
│ ├── format-code.sh # 格式化脚本
│ └── run-tests.sh # 测试脚本
├── .mcp.json # MCP 配置
├── LICENSE
├── README.md
└── CHANGELOG.md
```
对应的 `plugin.json`:
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "开发者工具箱:代码审查、测试、部署一站式解决方案",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## 发布到 Marketplace
### 什么是 Marketplace
Marketplace 是 Plugin 的分发中心。你可以把它理解为"插件商店"——用户通过简单的命令就能安装你发布的插件。
### 创建 Marketplace 配置
在你的 GitHub 仓库中创建 `.claude-plugin/marketplace.json`:
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "开发者工具箱",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "文档生成工具"
}
]
}
```
### 插件来源类型
Marketplace 支持多种来源:
**相对路径**(同仓库内的插件):
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**GitHub 仓库**:
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**任意 Git 仓库**:
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### 发布流程
1. **创建 GitHub 仓库**
2. **推送代码**:
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **用户添加你的 Marketplace**:
```bash
/plugin marketplace add your-username/your-repo
```
4. **用户安装插件**:
```bash
/plugin install dev-toolkit@your-marketplace
```
## 安装和管理 Plugin
### 通过交互式菜单
```bash
/plugin
```
这会打开一个交互式界面,你可以浏览、安装、启用、禁用插件。
### 通过命令行
**添加 Marketplace**:
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # 本地
```
**安装插件**:
```bash
# 安装到用户范围(默认)
/plugin install formatter@my-marketplace
# 安装到项目范围(团队共享)
/plugin install formatter@my-marketplace --scope project
# 安装到本地范围(gitignored)
/plugin install formatter@my-marketplace --scope local
```
**其他管理命令**:
```bash
/plugin enable # 启用插件
/plugin disable # 禁用插件
/plugin uninstall # 卸载插件
/plugin update # 更新插件
```
### 验证插件
在发布前验证插件配置是否正确:
```bash
claude plugin validate .
```
或在 Claude Code 中:
```
/plugin validate .
```
## 团队协作配置
### 在项目中共享插件配置
将插件配置提交到版本控制,让团队成员自动获得:
`.claude/settings.json`:
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
团队成员克隆项目后,这些插件会自动可用。
### 企业级 Marketplace 限制
对于需要严格控制的企业环境,可以在 managed settings 中限制允许的 Marketplace:
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
设置为空数组 `[]` 可以完全禁止外部插件。
## CLI 命令参考
| 命令 | 说明 |
| -------------------------------------- | ------------------ |
| `/plugin` | 打开交互式管理界面 |
| `/plugin install @` | 安装插件 |
| `/plugin uninstall ` | 卸载插件 |
| `/plugin enable ` | 启用插件 |
| `/plugin disable ` | 禁用插件 |
| `/plugin update ` | 更新插件 |
| `/plugin validate .` | 验证当前目录的插件配置 |
| `/plugin marketplace add ` | 添加 Marketplace |
| `/plugin marketplace list` | 列出已添加的 Marketplace |
| `/plugin marketplace update` | 更新 Marketplace 缓存 |
| `/plugin marketplace remove ` | 移除 Marketplace |
## 最佳实践
### 开发最佳实践
1. **保持 Skills 专注**:每个 Skill 做好一件事,避免大而全
2. **编写清晰的描述**:让 Claude 知道何时使用你的组件
3. **先团队测试**:在分发到社区前,先在团队内部验证
4. **记录版本变更**:在 CHANGELOG.md 中记录每个版本的改动
### 目录结构最佳实践
* 将 `commands/`、`agents/`、`skills/` 放在插件根目录
* 只将 `plugin.json` 放在 `.claude-plugin/` 目录
* 使用 `${CLAUDE_PLUGIN_ROOT}` 引用插件内的文件
* 不要用 `../` 访问插件外的文件
### Hooks 最佳实践
1. 脚本必须可执行:`chmod +x script.sh`
2. 使用 shebang 声明解释器:`#!/bin/bash`
3. 使用 `${CLAUDE_PLUGIN_ROOT}` 变量确保路径正确
4. 先单独测试脚本再集成到 Hook
### 版本管理最佳实践
遵循语义化版本(Semantic Versioning):
* **MAJOR** (1.0.0 → 2.0.0):破坏性变更
* **MINOR** (1.0.0 → 1.1.0):新增功能(向后兼容)
* **PATCH** (1.0.0 → 1.0.1):Bug 修复(向后兼容)
## 常见问题排查
| 问题 | 可能原因 | 解决方案 |
| --------- | ---------------- | ---------------------------------------- |
| 插件不加载 | plugin.json 格式错误 | 用 `claude plugin validate` 验证 |
| 命令未出现 | 目录结构错误 | 确保 `commands/` 在根目录,不在 `.claude-plugin/` |
| Hooks 不触发 | 脚本不可执行 | 运行 `chmod +x script.sh` |
| 路径找不到 | 使用了相对路径 | 改用 `${CLAUDE_PLUGIN_ROOT}` |
| MCP 服务器失败 | 环境变量未设置 | 检查 `.mcp.json` 中的路径配置 |
## 从现有配置迁移
如果你已经有 `.claude/` 目录下的配置,可以按以下步骤迁移为 Plugin:
1. **创建 Plugin 结构**:
```bash
mkdir my-plugin/.claude-plugin
```
2. **创建 plugin.json**:
```json
{
"name": "my-plugin",
"description": "从现有配置迁移的插件",
"version": "1.0.0"
}
```
3. **复制现有文件**:
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **迁移 Hooks**:
从 `.claude/settings.json` 复制 `hooks` 配置到 `hooks/hooks.json`
5. **测试**:
```bash
claude --plugin-dir ./my-plugin
```
## 学习资源
### 官方文档
| 资源 | 链接 | 说明 |
| ----------------- | ----------------------------------------------------------------------------------------- | ------ |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | 插件参考文档 |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | 插件创建指南 |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | 官方最佳实践 |
| Agent Skills 标准 | [agentskills.io](https://agentskills.io) | 开放标准规范 |
### 官方仓库
| 资源 | 链接 | 说明 |
| ---------------------------------- | ---------------------------------------------------------------------------------- | ------------ |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 官方 Skills 仓库 |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 官方插件目录 |
| Docker MCP Toolkit | [官网](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ 预构建 MCP |
### 社区资源
| 资源 | 链接 | 说明 |
| ---------------------- | ---------------------------------------------------------------------------- | ----------------- |
| claude-plugins.dev | [官网](https://claude-plugins.dev/) | 社区注册表和 CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 个插件集合 |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 最佳实践汇总 |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 代理 + 15 编排器 |
| jeremylongshore 教程 | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | 数百插件 + Jupyter 教程 |
### 推荐阅读
| 文章 | 来源 |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## 展望
Plugin 机制让 Claude Code 的扩展能力有了质的飞跃。随着社区的发展,我们可以期待:
* **更丰富的插件生态**:覆盖各种开发场景和工作流
* **企业级功能**:更完善的权限管理和审计能力
* **跨平台兼容**:Skills 开放标准已被多家厂商采用
现在正是入场的好时机。你可以从简单的命令开始,逐步添加 Skills、Hooks,最终形成一套完整的工作流解决方案。
如果你想深入了解 Plugin 可以包含的 Subagent 组件,请阅读《[Claude Code Subagent 是什么](/docs/notes/claude-subagent/concept)》。
# Claude Code Subagent 概念介绍
## 引言
在使用 Claude Code 处理复杂任务时,你可能遇到过这样的困境:主对话的上下文越来越长,AI 开始"忘记"之前的重要信息,响应质量逐渐下降。
Subagent(子代理)就是为解决这个问题而生的。
如果说 Skills 是给 Claude 的"工作手册",那么 Subagent 就是你雇佣的"专职员工"——它们有自己独立的工位(上下文),专注于特定类型的工作,完成后向你汇报结果。
## 理解 Subagent
想象你是一家公司的 CEO。当公司规模小的时候,你亲自处理所有事务。但随着业务扩展,你开始雇佣专职员工:会计负责财务、HR 负责招聘、工程师负责开发。每个员工在自己的工位上工作,处理完任务后向你汇报。
Subagent 在 Claude Code 中扮演的正是这样的角色。
从技术角度来说,Subagent 是专门化的 AI 助手,它们具有以下特点:
| 特性 | 说明 |
| --------- | ------------------------ |
| **独立上下文** | 每个 Subagent 在自己的上下文窗口中运行 |
| **专门化能力** | 针对特定任务类型进行优化 |
| **可配置工具** | 只能访问指定的工具集 |
| **自定义提示** | 有专门的系统提示指导行为 |
## 为什么需要独立上下文
这是 Subagent 最核心的设计理念,值得深入理解。
在普通对话中,所有信息都堆积在同一个上下文里。当你让 Claude 搜索代码库、分析文件、然后进行修改时,所有这些中间过程都会占用上下文空间。随着对话深入,上下文越来越拥挤,Claude 可能开始"忘记"早期的重要信息。
Subagent 改变了这一点:
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
Subagent 的分析过程不会污染主对话。主对话只收到精炼的结果,保持了清晰和专注。
## 内置 Subagent 类型
Claude Code 提供了三个强大的内置 Subagent,覆盖最常见的使用场景:
### Explore Subagent(探索代理)
**定位**:快速、只读的代码库探索。
**特点**:
* 使用 Haiku 模型(快速、低延迟)
* 严格只读——无法创建、修改或删除文件
* 可用工具:Glob、Grep、Read、Bash(只读操作)
**何时使用**:
当你问"这个功能在哪里实现的"、"错误是怎么处理的"这类探索性问题时,Claude 会自动调用 Explore Subagent。
**详细程度级别**:
| 级别 | 说明 | 适用场景 |
| ------------- | --------- | ----------- |
| Quick | 快速搜索,最少探索 | 针对性的简单查询 |
| Medium | 适度探索 | 平衡速度和完整性 |
| Very thorough | 全面分析 | 需要深入理解的复杂问题 |
### Plan Subagent(计划代理)
**定位**:研究代码库,准备实施计划。
**特点**:
* 使用 Sonnet 模型(更强的推理能力)
* 只有探索工具:Read、Glob、Grep、Bash
* 在计划模式下自动调用
**何时使用**:
当你进入计划模式,需要 Claude 先调研再提出方案时,Plan Subagent 会自动收集信息,然后基于研究结果给出计划建议。
### General-purpose Subagent(通用目的代理)
**定位**:处理复杂的多步任务。
**特点**:
* 使用 Sonnet 模型
* 访问所有工具(包括读写)
* 适合需要探索和修改的复杂任务
**何时使用**:
当任务涉及多个步骤、需要先搜索再修改、或者初始搜索可能失败需要尝试多种策略时。
## 我的理解与实践
仔细观察官方内置的三个 Subagent,你会发现一个共同点:**它们都是调研和规划类的任务**。Explore 负责探索代码库,Plan 负责制定计划,即使是 General-purpose 也主要用于研究和分析。没有一个是专门用来写代码的。
这印证了我对 Subagent 的理解:**Subagent 的核心价值不是"干净的上下文",而是让主 agent 能够专注做事**。
### 分工模式
我的使用方式很简单:Subagent 负责调研、规划、review 这类"信息收集"的工作,主 agent 负责实际执行。
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### 为什么不让 Subagent 写代码
有些人喜欢让主 agent 调度多个 subagent 去写代码,我觉得这不太靠谱。原因很简单:**上下文严重缺失**。
Subagent 的上下文是独立的,它不知道主对话中讨论过什么、做过什么决定、有什么约束。让它去写代码,就像让一个刚入职的员工在没有任何背景信息的情况下独立完成任务——产出的代码很可能和你的预期不符。
相反,把 Subagent 定位为"调研员"就合理多了:
* 调研任务本身就不需要太多上下文
* 返回的是信息而不是代码,主 agent 可以基于完整上下文来使用这些信息
* 即使调研结果有偏差,主 agent 也能纠正
### 我的日常用法
1. **开始新任务前**:让 Explore 代理快速了解相关代码的结构
2. **复杂任务规划**:让 Plan 代理分析需求并制定实施步骤
3. **代码审查**:让 Review 代理检查代码质量和安全问题
4. **实际编码**:主 agent 基于收集到的上下文来写代码
这是我目前使用的 agent 列表:
这样做的好处是:主 agent 的上下文窗口保持干净,只有"我需要知道的信息",而不是"Subagent 搜索过程中产生的一堆中间结果"。
## 与其他功能的对比
### Subagent vs Skills
这是最常见的困惑。核心区别:**Skills 向 Claude 注入知识;Subagent 创建独立的工作者**。
| 维度 | Skills | Subagent |
| -------- | ---------------- | ----------- |
| **核心功能** | 提供专业知识和指令 | 独立执行任务的代理 |
| **上下文** | 共享主对话上下文 | 拥有独立的上下文 |
| **触发方式** | 根据描述自动匹配 | 自动委托或手动调用 |
| **适用场景** | 让 Claude 更擅长某类任务 | 复杂、多步骤的独立任务 |
形象地说:Skills 像培训材料,让 Claude 学会怎么做某件事;Subagent 像专职员工,在自己的工位上独立完成任务后汇报结果。
两者可以组合:一个代码审查 Subagent 可以加载代码规范 Skill,实现"专家 + 专业知识"的组合效果。
### Subagent vs 斜杠命令
| 维度 | Subagent | 斜杠命令 |
| -------- | --------- | ------ |
| **激活方式** | 自动委托或显式调用 | 用户手动输入 |
| **上下文** | 独立上下文 | 共享主对话 |
| **复杂度** | 适合复杂任务 | 适合简单操作 |
斜杠命令是快捷键,你输入 `/review` 触发预定义的操作;Subagent 是独立的工作者,可以自主完成复杂的多步任务。
### Subagent vs Plugin
Plugin 是一个"容器"概念,它可以包含 Subagent:
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
你可以在 Plugin 的 `agents/` 目录下定义 Subagent,随 Plugin 一起分发。
## Agentic 设计模式
Anthropic 在其官方文档中总结了六种核心的 Agentic 设计模式,理解这些模式有助于更好地设计 Subagent 系统:
| 模式 | 核心思想 | Subagent 应用 |
| ------------------------ | -------------- | --------------------------- |
| **Prompt Chaining** | 将复杂任务分解为多个顺序步骤 | 链式调用多个 Subagent |
| **Routing** | 根据输入类型分发到专门处理器 | 不同类型任务委托给专门 Subagent |
| **Parallelization** | 同时执行多个独立子任务 | 并行启动多个 Subagent |
| **Orchestrator-Workers** | 中央协调者分配任务给工作者 | Claude 作为协调者,Subagent 作为工作者 |
| **Evaluator-Optimizer** | 生成器产出,评估器优化 | 生成 Subagent + 审查 Subagent |
| **Agents** | 自主决策执行的独立代理 | 每个 Subagent 独立运行 |
这些模式可以组合使用。例如,一个代码质量系统可能同时使用:
* **Parallelization**:同时运行安全扫描和性能分析
* **Orchestrator-Workers**:主 Claude 协调多个专门化 Subagent
* **Evaluator-Optimizer**:代码生成后立即进行审查
## 核心优势
### 上下文保护
Subagent 最大的价值在于保护主对话的上下文。代码搜索、文件分析这些中间过程不会堆积在主对话里,让主对话始终专注于高层目标。
### 专门化能力
你可以为特定领域创建专门的 Subagent,配置详细的指令和适当的工具。专门化的 Subagent 比通用的 Claude 在特定任务上表现更好。
### 灵活的权限控制
每个 Subagent 可以有不同的工具访问权限。比如,探索类 Subagent 只给只读权限,修改类 Subagent 才给写权限。这种细粒度控制提高了安全性。
### 可复用性
一旦创建,Subagent 可以在不同项目间复用,也可以通过 Plugin 与团队共享。
## 何时使用 Subagent
**适合使用 Subagent 的场景**:
* 需要独立上下文运行任务
* 任务是复杂的多步工作流
* 需要与主对话不同的工具集
* 任务可能长时间运行
**不适合使用 Subagent 的场景**:
* 简单的一次性查询
* 需要与主对话紧密交互
* 任务很快就能完成
## 典型应用场景
### 代码审查
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
当你完成一段代码后,可以让代码审查 Subagent 在独立上下文中进行审查,不干扰你的主要开发工作。
### 调试分析
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
遇到错误时,调试 Subagent 可以深入分析错误原因,尝试各种假设,最终给出修复建议。
### 代码库探索
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
当你对新项目不熟悉时,探索 Subagent 可以快速绘制代码库地图,而不会让大量的搜索结果污染你的主对话。
## 学习资源
### 官方资源
| 资源 | 链接 | 说明 |
| -------------- | ----------------------------------------------------------------------------------------------------------------- | --------------- |
| Claude Code 文档 | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 官方文档入口 |
| Subagents 指南 | [Claude Code Docs](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Subagent 官方文档 |
| 多代理系统研究 | [Anthropic Engineering](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2% 性能提升的研究详情 |
| Agentic 设计模式 | [Anthropic Docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 六种核心设计模式详解 |
### 社区资源
| 资源 | 链接 | 说明 |
| -------------------- | ----------------------------------------------------------------- | -------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 个 Agent + 15 个编排器 |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 个专门化代理的 Plugin |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 最佳实践汇总 |
## 小结
Claude Code Subagent 本质上是**独立上下文的专门化 AI 助手**。它通过上下文隔离解决了复杂任务中的信息过载问题,让主对话始终保持清晰和专注。
记住三个关键词:
| 关键词 | 含义 |
| ------ | ----------------------------- |
| **独立** | 每个 Subagent 有自己的上下文窗口 |
| **专门** | 针对特定任务类型优化 |
| **委托** | Claude 可以自动或手动将任务委托给 Subagent |
了解了概念之后,下一篇《[Claude Code Subagent 实战指南](/docs/notes/claude-subagent/practice)》将带你动手实践:创建自定义 Subagent、配置工具权限、以及实际项目中的最佳实践。
如果你想了解 Subagent 可以加载的 Skills,请阅读《[Claude Skills 是什么](/docs/notes/claude-skills/concept)》。如果你想把 Subagent 打包分发,请阅读《[Claude Code Plugin 是什么](/docs/notes/claude-plugin/concept)》。
# Claude Code Subagent 实践指南
## 快速回顾
在上一篇文章中,我们了解了 Subagent 的核心概念:它是独立上下文的专门化 AI 助手,通过上下文隔离解决复杂任务中的信息过载问题。Claude Code 内置了三种 Subagent:Explore(探索)、Plan(计划)、General-purpose(通用)。本文将从实战角度出发,带你创建自定义 Subagent 并掌握高级用法。
## 管理 Subagent
### 通过 /agents 命令
最简单的方式是使用交互式界面:
```bash
/agents
```
这会打开一个菜单,你可以:
* 查看所有 Subagent(内置 + 自定义)
* 创建新的 Subagent
* 编辑现有 Subagent 的配置和工具权限
* 删除不需要的 Subagent
* 查看名称冲突时哪个 Subagent 是活跃的
### 通过文件管理
Subagent 以 Markdown 文件形式存储。你也可以直接创建和编辑文件。
**存储位置**:
| 位置 | 路径 | 作用域 |
| ------ | ------------------- | --------------- |
| 项目级 | `.claude/agents/` | 当前项目专用,可提交到 Git |
| 用户级 | `~/.claude/agents/` | 跨所有项目可用 |
| Plugin | 插件的 `agents/` 目录 | 随 Plugin 安装 |
**优先级**:项目级 > 用户级 > Plugin 级
当同名 Subagent 存在于多个位置时,优先级高的会覆盖低的。
## 创建你的第一个 Subagent
### 第一步:创建目录
```bash
mkdir -p .claude/agents
```
### 第二步:创建 Markdown 文件
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### 第三步:测试 Subagent
在 Claude Code 中:
```
> 用 code-reviewer 代理审查我最近的修改
```
或者让 Claude 自动选择:
```
> 帮我审查一下代码质量
```
如果 `description` 写得足够清晰,Claude 会自动识别并调用你的 Subagent。
## 配置字段详解
Subagent 配置文件由两部分组成:YAML frontmatter 和 Markdown 正文。
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| 字段 | 必需 | 说明 |
| ---------------- | -- | ---------------------------------------- |
| `name` | 是 | 唯一标识符,使用小写字母和连字符 |
| `description` | 是 | 自然语言描述(Claude 用此判断何时调用) |
| `tools` | 否 | 逗号分隔的工具列表。省略则继承所有工具 |
| `model` | 否 | 模型选择:`sonnet`、`opus`、`haiku` 或 `inherit` |
| `permissionMode` | 否 | 权限模式(见下文) |
| `skills` | 否 | 自动加载的 Skills(Subagent 不继承父会话的 Skills) |
### 权限模式
| 模式 | 说明 |
| ------------------- | ------------ |
| `default` | 正常权限检查 |
| `acceptEdits` | 自动接受编辑操作 |
| `bypassPermissions` | 跳过所有权限检查 |
| `plan` | 只提出计划,不执行 |
| `ignore` | 忽略此 Subagent |
### Markdown 正文
正文是 Subagent 的系统提示。写得越详细,Subagent 的表现越好。
好的系统提示应该包含:
* 明确的角色定义
* 具体的工作步骤
* 关键检查清单
* 输出格式要求
## 触发机制
### 自动委托
Claude 会根据任务内容和 Subagent 的 `description` 自动决定是否委托。
**鼓励自动使用的技巧**:在 `description` 中使用触发词:
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
或:
```yaml
description: MUST BE USED when encountering errors or test failures
```
### 显式调用
直接告诉 Claude 使用哪个 Subagent:
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
## 工具配置
### 常用工具列表
| 工具 | 说明 |
| ----------- | ----------- |
| `Read` | 读取文件内容 |
| `Write` | 写入文件 |
| `Edit` | 编辑文件 |
| `Glob` | 文件模式匹配 |
| `Grep` | 正则表达式搜索 |
| `Bash` | 执行 shell 命令 |
| `WebFetch` | 获取网页内容 |
| `WebSearch` | 搜索网页 |
### 工具配置策略
**只读 Subagent**(探索、分析):
```yaml
tools: Read, Grep, Glob, Bash
```
注意:即使包含 Bash,Subagent 也应该只用于只读命令(ls, git status, git log 等)。
**读写 Subagent**(修复、重构):
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**最小权限原则**:只授予必需的工具,避免意外操作。
## 实用 Subagent 模板
### 代码审查器
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### 调试专家
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### 测试运行器
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### 文档生成器
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### 安全扫描器
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## 高级用法
### 生产级设计模式
在生产环境中,有几种经过验证的多代理协作模式:
#### 3 Amigos 模式
由产品、架构、实现三个角色组成的协作模式:
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| 角色 | 职责 | 工具配置 |
| --------------- | --------- | ---------------- |
| PM Agent | 功能定义、需求梳理 | Read, WebSearch |
| Architect Agent | 技术方案设计 | Read, Glob, Grep |
| Claude Code | 代码实现 | 所有工具 |
#### 三阶段流水线
将复杂任务分解为三个明确阶段:
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
每个阶段由专门的 Subagent 负责,输出作为下一阶段的输入。
#### 模型编排策略
不同阶段使用不同模型,优化成本和效果:
| 阶段 | 推荐模型 | 原因 |
| ---- | ------ | ------ |
| 规划阶段 | Sonnet | 需要深度推理 |
| 执行阶段 | Haiku | 快速、低成本 |
| 审查阶段 | Sonnet | 需要综合判断 |
配置示例:
```yaml
---
name: quick-executor
model: haiku
---
```
### Subagent 链接
对于复杂工作流,可以链接多个 Subagent:
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### 可恢复执行
Subagent 执行可以暂停和恢复,保持之前的完整上下文:
**初始调用**:
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**恢复代理**:
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**使用场景**:
* 长时间运行的研究,分多个会话完成
* 迭代改进,保持上下文
* 多步工作流,依序处理相关任务
### 为 Subagent 配置 Skills
Subagent 不会自动继承父会话的 Skills。如果需要,显式声明:
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### CLI 动态定义
无需保存文件,直接在命令行定义临时 Subagent:
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
适用于快速测试或一次性使用。
## 最佳实践
### 1. 保持专注
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
一个 Subagent 做好一件事,比一个 Subagent 做很多事效果更好。
### 2. 编写清晰的描述
Claude 用 `description` 决定何时使用 Subagent。好的描述应该回答:
1. **这个 Subagent 做什么?** 列出具体能力
2. **什么时候应该使用它?** 包含触发词
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. 限制工具访问
只授予需要的工具:
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
这防止了 Subagent 意外修改文件,也让它更专注于审查工作。
### 4. 编写详细的系统提示
系统提示越详细,Subagent 表现越好:
* 明确的角色定义
* 具体的工作步骤
* 关键检查清单
* 输出格式要求
### 5. 版本控制
将项目级 Subagent 提交到 Git:
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
团队成员克隆项目后自动获得相同的 Subagent。
## 常见问题排查
| 问题 | 可能原因 | 解决方案 |
| ------------- | ---------------- | --------------------------------------------- |
| Subagent 不被调用 | description 不够清晰 | 添加触发词,写得更具体 |
| Subagent 不被调用 | 文件位置错误 | 确认文件在 `.claude/agents/` 或 `~/.claude/agents/` |
| 工具不可用 | tools 字段配置错误 | 检查工具名拼写,确认用逗号分隔 |
| 输出不稳定 | 系统提示太模糊 | 添加具体的步骤和输出格式要求 |
| 上下文丢失 | 会话结束 | 使用可恢复执行功能 |
| 名称冲突 | 多个位置有同名 Subagent | 用 `/agents` 查看哪个是活跃的 |
## 与团队共享
### 方式一:通过 Git
将 Subagent 放在 `.claude/agents/` 目录,提交到项目仓库。团队成员克隆后自动获得。
### 方式二:通过 Plugin
将 Subagent 放在 Plugin 的 `agents/` 目录,通过 Plugin 机制分发。
### 方式三:用户级共享
将常用 Subagent 放在 `~/.claude/agents/`,跨所有项目可用。可以用 dotfiles 管理在多台机器间同步。
## 学习资源
### 官方文档
| 资源 | 链接 | 说明 |
| -------------- | ----------------------------------------------------------------------------------------------------------------- | --------------- |
| Claude Code 文档 | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 官方文档入口 |
| Subagents 指南 | [Claude Code Docs](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Subagent 配置详解 |
| Agentic 设计模式 | [Anthropic Docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 六种核心设计模式 |
| 多代理研究 | [Anthropic Engineering](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2% 性能提升的研究详情 |
### 社区资源
| 资源 | 链接 | 说明 |
| -------------------- | ----------------------------------------------------------------- | ---------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 个 Agent + 15 个编排器模板 |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 个专门化代理的 Plugin |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Claude Code 最佳实践汇总 |
### 推荐阅读
| 文章 | 来源 | 主题 |
| ----------------------------------------------- | ---------------------------------------------------------------------------------------------- | ---------- |
| Building effective agents | [Anthropic](https://www.anthropic.com/research/building-effective-agents) | Agent 设计原则 |
| How we built our multi-agent research system | [Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system) | 多代理架构实践 |
| Claude Skills, Commands, Subagents, and Plugins | [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | 功能对比分析 |
## 小结
Claude Code Subagent 是提升 AI 编程效率的强大工具。通过独立上下文和专门化配置,它让复杂任务变得可管理。
快速开始:
1. 运行 `/agents` 打开管理界面
2. 创建一个简单的 Subagent(如代码审查器)
3. 测试自动委托和显式调用
4. 根据需要调整配置
随着使用深入,你可以逐步:
* 为团队创建专属 Subagent
* 配置 Subagent 链接处理复杂工作流
* 利用可恢复执行处理长期任务
如果你想把 Subagent 与其他配置打包分发,请阅读《[Claude Code Plugin 实战指南](/docs/notes/claude-plugin/practice)》。
# 进阶篇
## 终端通知:任务完成提醒
Claude 完成任务后希望收到通知?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
配合 iTerm2 的通知功能,或者用 `terminal-notifier` 做自定义通知(参考 [最佳实践](/blog/claude-code-best-practices) 中的 Hooks 配置)。
## Hooks 进阶用法
Hooks 不只是能跑 shell 命令。实际上有四种类型:
1. **command**:Shell 命令(最常见)
2. **http**:POST JSON 到 URL(支持自定义 headers 和环境变量展开)
3. **prompt**:发给 Claude 评估(比如「所有任务都完成了吗?」)
4. **agent**:启动一个有工具访问权限的子代理来验证
一些进阶的 Hook 事件:
* `PostCompact`:压缩完成后触发,适合注入提醒让 Claude 重新读取关键文件
* `SessionStart`:写入 `$CLAUDE_ENV_FILE` 可以给整个会话持久化环境变量
* `PreToolUse`:可以修改工具输入(`updatedInput`),甚至自动批准或拒绝操作
## 插件生态
`/plugin` 可以浏览和安装社区插件。一些值得了解的插件:
* **dx**(by ykdojo):提供 `/handoff`(自动写交接文档)、`/clone`(克隆对话)、`/half-clone`(只克隆最近的对话减少上下文)
* **mine**(by anipotts):把所有 Claude Code 会话数据导入 SQLite,支持成本追踪、缓存分析、错误记忆等查询
## Agent Teams:多代理协作
设置环境变量开启实验性的 Agent Teams 功能:
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
开启后,一个会话可以作为 Team Lead,通过 git worktree 协调多个代理同时工作。每个代理在自己的上下文窗口中独立运行,适合大型项目的并行开发。
不过 token 消耗会增加 4-15 倍,酌情使用。
## 提示词心法
以下来自 Boris Cherny 推特中分享的团队实践,算是「提示工程」在 Claude Code 场景下的最佳应用。
### 让 Claude 当你的审查员
不要只是让 Claude 写代码,还可以让它审查你的代码:
```
Grill me on these changes and don't make a PR until I pass your test.
```
或者让它证明代码是可以工作的:
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### 回答不满意时不要重复提问
Boris 的第 6 条 tip 提到:如果 Claude 给了一个平庸的回答,不要换个说法重新问一遍。直接说「这个方案不够好,具体说说哪里可以改进」,让它在现有基础上迭代,比从头开始效果好。
### 让 Claude 自己更新 CLAUDE.md
每次纠正 Claude 的错误后,加一句:
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris 说 Claude 在「给自己写规则」这件事上出奇地好。这样 CLAUDE.md 会越来越精准,后续对话的质量也会持续提升。
### 直接说「fix」
启用 Slack MCP 后,把 Slack 中的 bug 报告直接粘贴给 Claude,只说一个字:**fix**。零上下文切换。
或者 CI 挂了,直接说:
```
Go fix the failing CI tests.
```
不需要手动分析日志,不需要解释是什么问题,让 Claude 自己去看日志、定位问题、修复。
## 写在最后
Claude Code 的功能更新非常快,这些技巧也在不断演进中。建议关注官方 Changelog 保持同步。
如果你还没看过我之前的文章,建议先从基础的工作流开始:
### 延伸阅读
* 《[我的 Claude Code 最佳实践](/blog/claude-code-best-practices)》— 工作流层面的核心技巧和斜杠命令指南
* 《[AI编程质量控制:5道防线确保代码质量](/blog/claude-code-quality-control)》— Claude Code 编程中的质量保障体系
* 《[Claude 系统架构全解析](/docs/notes/claude-architecture)》— 理解 MCP、Skills、Subagents、Hooks 等组件
# 实用命令与自动化篇
## `/diff`:交互式 Diff 查看器
输入 `/diff` 打开一个交互式的 diff 视图:
* **左右箭头**:在 git diff(全量改动)和 Claude 每轮改动之间切换
* **上下箭头**:浏览不同文件
比在终端里跑 `git diff` 体验好得多,特别是改动涉及多个文件的时候。
## `/simplify`:多代理代码审查
执行 `/simplify` 会同时启动 3 个并行的审查代理:
* **代码复用**代理:查找重复模式
* **代码质量**代理:检查可读性和结构
* **效率**代理:分析不必要的性能开销
三个代理独立工作,最后汇总结果,自动修复有效问题、跳过误报。
## `/security-review`:安全扫描
对当前分支的改动进行安全审查,检查 SQL 注入、XSS、认证缺陷、数据处理问题和依赖漏洞。每个发现会经过对抗性验证来减少误报。
## `/copy` 的隐藏功能
`/copy` 不只是复制上一条回复。当回复中包含代码块时,它会弹出交互式选择器让你选择特定的代码块,而不是复制整段回复。还可以传数字来复制更早的回复:`/copy 2` 复制倒数第二条,`/copy 3` 复制倒数第三条,不用翻屏手动选择。
## `/batch`:大规模并行重构
```
/batch 把 src/ 下所有组件从 Class 组件迁移到函数组件
```
这是重量级功能。`/batch` 会分析代码库,把任务分解成 5-30 个独立单元,每个单元启动一个独立代理在隔离的 git worktree 中工作,最后每个代理提交并开一个 PR。
适合大规模迁移、批量加类型注解、全局重命名等场景。
## `/loop`:定时任务
```
/loop 5m 检查部署是否完成
/loop 1h /review-pr 1234
```
在会话内创建定时任务,按指定间隔重复执行。适合轮询部署状态、定期检查 PR 等场景。会话级别的(退出就没了),最多 50 个任务,3 天自动过期。
## 管道输入:把任何东西喂给 Claude
```bash
# 让 Claude 分析错误日志
cat error.log | claude -p "分析这个错误日志,找出根本原因"
# 让 Claude 总结最近的改动
git diff HEAD~3 | claude -p "总结这三次提交的改动"
# 让 Claude 解读命令输出
kubectl get pods | claude -p "哪些 pod 状态异常?"
```
`-p` 是 headless 模式(非交互式),适合在脚本和 CI/CD 中使用。
## Headless 模式的隐藏参数
`-p` 模式有一些非常强大但很少人知道的参数:
```bash
# 设置花费上限(超过就停)
claude -p --max-budget-usd 5.00 "重构认证模块"
# 限制对话轮数
claude -p --max-turns 3 "修复这个测试"
# 输出 JSON 格式(方便程序解析)
claude -p --output-format json "分析这个项目"
# 要求输出符合特定 JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "总结项目"
# 多轮 headless 对话(用 session-id 保持上下文)
claude -p --session-id my-task "第一步:分析代码"
claude -p --session-id my-task "第二步:生成测试"
# 指定备用模型(主模型过载时自动切换)
claude -p --fallback-model sonnet "复杂分析"
# 限制可用工具
claude -p --tools "Read,Grep,Glob" "只读分析,不要改代码"
# 完全替换系统提示词
claude -p --system-prompt "你是一个 Python 专家" "优化这段代码"
```
# 配置与诊断篇
## `/statusline`:自定义状态栏
用 `/statusline` 可以通过自然语言描述来自定义底部状态栏显示的信息。或者手动创建 `~/.claude/statusline.sh` 脚本。
可以显示的信息包括:当前模型、git 分支、未提交文件数、上下文使用进度、会话花费等。当你同时开多个 Claude 窗口处理不同任务时,状态栏能帮你快速分清每个窗口在干什么。
## settings.json 自动补全
在 settings.json 开头加上 `$schema`,VS Code / Cursor 就会提供配置项的自动补全和校验:
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## 一些有用的隐藏配置
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration`:显示每轮对话的耗时
* `DISABLE_AUTOUPDATER`:关闭自动更新检查,减少上下文开销
## `/stats` 和 `/insights`:使用分析
* `/stats`:可视化展示每日用量、会话历史、使用条纹、模型偏好,支持日期范围筛选
* `/insights`:分析你所有的 Claude Code 历史记录,告诉你什么工作流有效、哪里在拖慢你,还能生成优化建议
## history.jsonl:提示词历史
Claude 会把你发送过的每一条提示词保存在 `~/.claude/history.jsonl` 中。你可以让 Claude 分析这个文件,找出提示词模式和优化空间。
## `/doctor`:健康检查
遇到奇怪的问题时,跑一下 `/doctor`(或在终端运行 `claude doctor`),它会检查安装状态、版本、认证状态和系统依赖,帮你快速定位问题。
## 社区工具
社区工具 `ccusage` 可以追踪 token 用量:
```bash
npx ccusage daily
```
如果你用了 `--dangerously-skip-permissions` 或者批准了很多命令,可以用 `cc-safe` 扫描一下有没有风险:
```bash
npx cc-safe .
```
它会检查 `.claude/settings.json` 中的 `sudo`、`rm -rf`、`chmod 777`、`git reset --hard` 等高风险命令。
## 更多值得了解的斜杠命令
| 命令 | 功能 |
| --------------------- | ----------------------------------------- |
| `/export [filename]` | 导出对话为纯文本 |
| `/pr-comments [PR]` | 获取 PR 评论(自动检测当前分支) |
| `/release-notes` | 查看当前版本更新日志 |
| `/plugin` | 浏览和安装社区插件 |
| `/fast` | 切换快速模式 |
| `claude --debug` | 启动时开启调试日志(支持分类过滤,如 `--debug "api,hooks"`) |
| `/install-github-app` | 安装 GitHub App 实现自动 PR 审查 |
# 上下文管理篇
## `/compact` 可以带参数
很多人知道 `/compact` 可以压缩上下文,但不知道它可以带参数来指定保留什么:
```
/compact 保留所有关于数据库 schema 的讨论,以及当前的重构方案
```
这样压缩的时候会优先保留你指定的内容,不至于丢失关键上下文。
## 在 CLAUDE.md 中写压缩存活指令
在 CLAUDE.md 中加一个 `## Compact Instructions` 部分,告诉 Claude 压缩时必须保留什么:
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
这样即使自动压缩也不会丢失关键信息。
## 防止 Claude 因为 token 预算提前放弃
在 CLAUDE.md 中加上这段:
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
有时候 Claude 会在上下文快满时主动停下来说「上下文快满了」,加上这段可以防止它提前放弃。
## Handoff 协议:会话交接
当上下文快满但任务还没完成时,让 Claude 写一份交接文档:
```
把剩余的计划写到 HANDOFF.md 里,说明你尝试了什么、什么有效、什么没效。
```
然后新开一个会话,只需要 `@HANDOFF.md` 就能恢复完整上下文。从 10K+ tokens 的上下文压缩到不到 2K,远比 `/compact` 精确。
## 在 70-80% 时主动压缩
一个容易忽略的点:当上下文接近上限时,Claude 会自动触发压缩。但自动压缩发生在任务中途时,可能会丢失关键信息导致后续响应质量下降。
更好的做法是**主动管理**:在上下文达到 70-80% 时手动 `/compact`,效果比等自动压缩好得多。完成一个任务后立刻 `/clear`,不要让上下文无限膨胀。
你也可以通过环境变量提前触发自动压缩:
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context`:上下文诊断
不确定上下文窗口还剩多少空间?`/context` 会告诉你:
* 哪些工具或 MCP 服务占用了最多上下文
* 当前容量使用百分比
* 针对性的优化建议
我发现有时候光是注册了某些 MCP 服务(即使没使用),就能吃掉 30% 以上的上下文窗口。用 `/context` 查一下,清理掉不用的 MCP 能释放不少空间。
## MCP 工具自动懒加载
当 MCP 工具定义超过上下文的 10% 时,Claude Code 会自动启用 Tool Search——加载一个轻量级搜索索引代替完整的工具定义。这能减少 85% 以上的 MCP 上下文消耗(比如从 77K tokens 降到 8.7K)。这个功能**默认就是开启的**,不需要手动配置。
需要注意:Tool Search 仅支持 Sonnet 4+ 和 Opus 4+ 模型,不支持 Haiku。如果你的 `ANTHROPIC_BASE_URL` 指向非官方代理,Tool Search 会被自动禁用(因为大多数代理不转发 `tool_reference` 块)。
如果想自定义行为,可以在 settings.json 中设置:
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
支持的配置值:
* **不设置**:默认启用
* **`true`**:强制启用(包括非官方代理场景)
* **`auto`**:超过 10% 上下文时激活(等同默认行为)
* **`auto:`**:自定义阈值,如 `auto:5` 表示超过 5% 就激活
* **`false`**:禁用,所有 MCP 工具预先加载
# Claude Code 隐藏技巧
## 核心结论
**这是一份 Claude Code 隐藏技巧索引**,整理快捷键、输入交互、模型控制、上下文管理、自动化命令、配置诊断和进阶用法。它适合已经会用 Claude Code、但想减少重复操作和上下文浪费的人。
如果你只想先看最实用的部分,优先读三篇:快捷键篇、上下文管理篇、实用命令与自动化篇。前者提升操作速度,后两者决定 Claude Code 在长任务里是否稳定。
从创始人 Boris Cherny 推特、社区和 changelog 中整理的 Claude Code 实用技巧。很多是「藏在角落里」的实用操作——快捷键、隐藏功能、命令行技巧等等,用了之后确实回不去了。
## 目录
* [快捷键篇](./shortcuts) — Shift+Tab 模式切换、Esc+Esc 回溯、Ctrl+S 暂存等
* [输入与交互篇](./input-interaction) — `!` 终端命令、`@` 文件注入、URL 粘贴、/btw 插嘴、Vim 模式
* [思考与模型控制篇](./thinking-model) — think/ultrathink 关键词、/effort、subagents、opusplan
* [会话管理篇](./session-management) — /rename、/branch、/color、远程控制
* [上下文管理篇](./context-management) — /compact 参数、压缩指令、Handoff 协议、MCP 懒加载
* [实用命令与自动化篇](./commands-automation) — /diff、/simplify、/batch、/loop、Headless 模式
* [配置与诊断篇](./config-diagnostics) — statusline、settings.json、/stats、/doctor
* [进阶篇](./advanced) — Hooks 进阶、插件生态、Agent Teams、提示词心法
### 延伸阅读
* [Claude 系统架构全解析](/docs/notes/claude-architecture) — 理解 MCP、Skills、Subagents、Hooks 等组件
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 子代理的概念与实践
# 输入与交互篇
## `!`:直接运行终端命令
在输入框以 `!` 开头,可以直接在 Claude Code 内执行终端命令,不需要切换到另一个终端窗口:
```
! git status
! npm run build
! docker ps
```
输入 `!` 加命令前缀后按 Tab 还能自动补全历史命令。
## `@` + 文件路径:注入文件上下文
在输入时用 `@` 加文件路径,可以把文件内容直接注入到上下文中:
```
帮我看看 @src/auth/login.ts 和 @src/auth/middleware.ts 之间的逻辑有没有问题
```
支持 Tab 键自动补全路径,不需要手动输入完整路径。比让 Claude 自己去读文件更快,因为省去了工具调用的开销。
## 直接粘贴 URL
直接把 URL 粘贴到输入中,Claude 会自动抓取网页内容作为上下文:
```
参考这个 API 文档 https://docs.example.com/api/v2 来写客户端代码
```
## 喂 `/llms-full.txt` 让 Claude 自己查文档
很多开源项目的文档站点会提供 `/llms-full.txt` 文件(LLM 友好的完整文档)。遇到某个库的问题时,把这个文件的 URL 粘贴给 Claude,它能自己查文档解决绝大部分问题:
```
参考 https://docs.astro.build/llms-full.txt 帮我解决这个路由问题
```
## `/btw`:在 Claude 工作时插嘴
这是 2026 年 3 月刚加的新功能。当 Claude 正在执行任务时,你可以用 `/btw` 发起一个旁路对话——问问它在想什么、给它补充信息,而不需要打断当前任务。
正如 Anthropic 工程师 @trq212 在推特上说的:「没人会用 Ctrl+C 打断同事,你只需要说一句 'btw',他们就会抬头看你。」
## Vim 模式
输入 `/vim` 开启 Vim 模式,支持:
* 模式切换(Normal/Insert)
* 导航(h/j/k/l, w/b/e, 0/$)
* 编辑操作(d, c, y, p)
* 文本对象(iw, aw, i", a())
如果你是 Vim 用户,这比默认的输入体验好太多。用 `/config` 可以设置为永久开启。
## 语音模式
输入 `/voice` 激活语音模式,长按空格键说话,松开发送。适合不想打字但又需要给 Claude 交代任务的时候。按键可以在 `keybindings.json` 中自定义。
# 会话管理篇
## `/rename`:给会话命名
```
/rename my-auth-refactor
```
给当前会话取个名字。命名的好处是:在交互式会话选择器(`claude --resume`)里,已命名的会话可以直接选中恢复,不需要再按回车确认;也可以在终端里用 `claude --resume my-auth-refactor` 一步到位启动。
在选择器中直接输入文字即可搜索过滤,还支持这些快捷键:`Ctrl+V` 预览会话、`Ctrl+R` 重命名、`Ctrl+A` 切换显示所有项目、`Ctrl+B` 按分支筛选。
## `/branch`:给对话开分支
就像 git 的分支一样,`/branch` 会在当前对话点创建一个分叉。你可以在分叉里尝试不同的方案,不影响原来的对话。试完不满意就回到原来的分支继续。
## `/color`:给窗口上色
给当前会话的提示栏设置颜色,支持 red、blue、green、yellow、purple、orange、pink、cyan。
## 命令行会话管理全家桶
```bash
# 恢复当前目录最近的会话
claude --continue
# 打开会话选择器,或按名称恢复
claude --resume
claude --resume my-auth-refactor
# 启动时直接命名会话
claude -n "auth-refactor"
# Fork 上一次会话(保留上下文,创建新分支)
claude -c --fork-session
# 恢复与特定 PR 关联的会话
claude --from-pr 123
# 在隔离的 git worktree 中启动
claude -w
```
Claude Code 的会话是自动保存的(所以不需要 Ctrl+S),每次打开终端用 `--continue` 就能恢复上次工作。
## `claude --remote`:跨设备继续
```bash
claude --remote "your task description"
```
启动一个 Web 会话,可以在 claude.ai 或手机 App 上继续操作。
## `/remote-control`:手机遥控本地 Claude
在电脑上的 Claude Code 里输入 `/remote-control`,它会生成一个连接码。然后在手机上的 Claude App 里输入这个连接码,就能远程控制本地的 Claude Code 会话——用手机给电脑上的 Claude 下达指令。
# 快捷键篇
Claude Code 的快捷键体系比大多数人想象的要丰富得多,按 `?` 可以查看当前环境下所有可用的快捷键。
## Shift+Tab:模式循环切换
这可能是最重要的一个快捷键。按 `Shift+Tab` 可以在三种模式之间循环切换:
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
不需要手动输入 `/plan` 或者 `/auto-accept`,一个键搞定。我的习惯是:接到新任务先按两下切到 Plan Mode,确认方案后再按一下切到 Auto-Accept 让 Claude 自行执行。
## Esc + Esc:回溯时光机
连按两次 `Esc`,会弹出回溯菜单(Rewind):
* **恢复代码和对话**:回到之前某个检查点,代码和对话都回滚
* **只恢复对话**:回滚消息,但保留当前代码改动
* **只恢复代码**:撤销文件修改,但保留对话历史
Claude 会自动跟踪每次文件编辑作为检查点。这比 `git checkout .` 精细得多,因为你可以选择回到任意一步,而不是只能回到上次提交。
不过要注意:只有 Claude 通过工具直接编辑的文件会被追踪,你手动改的文件、git push 之类的外部操作没法回滚。
## Ctrl+S:提示词暂存(Prompt Stash)
写了一半的提示词,突然需要先处理另一件事?按 `Ctrl+S`,当前输入会被暂存起来:
然后你可以输入其他命令或指令。等你提交完那条消息后,之前暂存的内容会**自动恢复**到输入框里,继续写。
这个功能就像 `git stash` 但用在提示词上。场景举例:你正在写一段很长的重构需求描述,突然想先让 Claude 看一下某个文件确认一下细节——按 `Ctrl+S` 暂存需求描述,先问文件相关的问题,回答完后你的需求描述自动回来。
## Ctrl+B:把任务丢到后台
Claude 正在处理一个耗时任务(比如大规模重构),你突然想处理另一件事?按 `Ctrl+B` 把当前任务推到后台,终端立刻可以继续输入新的指令。
用 `Ctrl+T` 可以查看后台任务列表,`Ctrl+F` 连按两次可以终止所有后台代理。
> tmux 用户注意:tmux 的前缀键默认也是 `Ctrl+B`,需要按两次才能触发 Claude 的后台功能。
## Ctrl+G:用编辑器写长提示词
有时候需要给 Claude 一段很长的指令,在终端里打字体验很差。按 `Ctrl+G` 会打开你系统默认的 `$EDITOR`(比如 VS Code、Vim),在编辑器里写好提示词,保存退出后自动发送给 Claude。
如果想切换默认编辑器,在 shell 配置文件(`~/.zshrc` 或 `~/.bashrc`)中设置:
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
`--wait` 参数很重要——它让编辑器等你关闭文件后再返回,否则 Claude 会立刻收到空内容。Vim 这类终端编辑器天然会阻塞,不需要加。
写多段落的需求描述、粘贴大段参考内容的时候特别好用。在 Plan Mode 下用 `Ctrl+G` 还可以直接在编辑器里修改 Claude 生成的计划。
## Cmd+T:切换扩展思考
官方默认快捷键是 `Cmd+T`(Windows/Linux 上是 `Meta+T`),用来开关扩展思考(Extended Thinking)模式。开启后 Claude 会在回答前进行更深入的推理,适合处理复杂的架构设计或 bug 排查。
不过要注意:大多数终端(iTerm2、Terminal.app、Warp 等)会把 `Cmd+T` 拦截为「新建标签页」,导致这个快捷键实际上用不了。解决办法有两个:用 `/keybindings` 自定义一个不冲突的快捷键,或者直接用 `/effort` 命令来切换思考深度(效果一样,还能精确控制级别)。
## Readline 快捷键
Claude Code 的输入框支持标准的 Readline 快捷键,终端老手会很熟悉:
| 快捷键 | 功能 |
| ----------------- | ---------- |
| Ctrl+A | 跳到行首 |
| Ctrl+E | 跳到行尾 |
| Ctrl+W | 删除前一个单词 |
| Ctrl+U | 删除到行首 |
| Ctrl+K | 删除到行尾 |
| Ctrl+Y | 粘贴刚删除的内容 |
| Alt+Y | 循环浏览删除历史 |
| Option+Left/Right | 按单词跳转(Mac) |
## 审批快捷键:`y/n/d/e`
当 Claude 提出文件修改等待你确认时,四个单键快捷键控制流程:
* `y`:接受
* `n`:拒绝
* `d`:查看完整 diff
* **`e`:编辑后再接受**
`e` 是最容易被忽略但最有用的——它让你在 Claude 的修改基础上做微调,然后再应用。不满意 Claude 的某几行代码?不用拒绝重来,直接 `e` 改了就好。
## 快捷键速查表
| 快捷键 | 功能 |
| ----------- | --------------------------------- |
| Shift+Tab | 切换模式:Normal → Auto-Accept → Plan |
| Esc+Esc | 打开回溯菜单 |
| Ctrl+S | 暂存当前输入,提交后自动恢复 |
| Ctrl+B | 把当前任务推到后台 |
| Ctrl+T | 查看后台任务列表 |
| Ctrl+F (x2) | 终止所有后台代理 |
| Ctrl+G | 用外部编辑器写提示词 |
| Ctrl+O | 切换详细工具输出视图 |
| Cmd+T | 切换扩展思考(可能被终端拦截,建议自定义或用 `/effort`) |
| `\` + Enter | 多行输入(无需配置) |
| Shift+Enter | 多行输入(需先运行 `/terminal-setup`) |
| Up / Down | 浏览输入历史 |
| Ctrl+R | 搜索命令历史 |
| Ctrl+L | 清屏(历史保留) |
| Ctrl+C | 取消当前生成 |
| Ctrl+D | 退出 Claude Code |
| `?` | 显示所有可用快捷键 |
## 自定义快捷键
如果默认的快捷键不合你的习惯,可以用 `/keybindings` 打开 `~/.claude/keybindings.json` 进行自定义。改完自动生效,不需要重启。
支持组合键语法(如 `ctrl+shift+c`)和 Chord 模式(如 `ctrl+k ctrl+s`,先按 Ctrl+K 松开,再按 Ctrl+S)。有 16 种不同的绑定上下文(Chat、Autocomplete、Confirmation、DiffDialog 等),每种上下文可以绑定不同的操作。
# 思考与模型控制篇
## 用关键词控制思考深度
在提示词中加入特定关键词可以触发不同级别的思考预算,这是 Claude Code 独有的功能(claude.ai 网页端没有):
| 关键词 | 思考预算 | 适用场景 |
| ----------------------------- | --------------- | ----------- |
| `think` | \~4,000 tokens | 日常编码问题 |
| `think hard` / `megathink` | \~10,000 tokens | 复杂逻辑、多文件关联 |
| `think harder` / `ultrathink` | \~31,999 tokens | 架构设计、疑难 bug |
实际使用中,我一般在遇到 Claude 给出浅层回答时,加上 `think hard` 重新提问。对于特别复杂的问题(比如跨多个服务的 bug 排查),直接上 `ultrathink`。
## `/effort`:控制思考深度
除了用关键词(think / ultrathink),还可以用 `/effort` 直接设置思考深度:
```
/effort low # 简单任务,跳过深度思考,更快更省
/effort high # 复杂任务,深度推理
/effort max # 最大思考预算(仅 Opus)
/effort auto # 让 Claude 自己判断
```
设置后在整个会话中持续生效。对于简单的文件修改用 `low`,复杂架构设计用 `max`,这样既省钱又不牺牲质量。
## 「use subagents」关键词
在任何请求后面加上 `use subagents`,Claude 会把任务分解给多个子代理并行处理。这样不仅速度更快,还能保持主代理的上下文窗口干净。
Boris 在推特上专门提到这一点:把单个任务卸载给子代理,让主代理的上下文保持聚焦。
## `opusplan`:最佳性价比模型方案
一句话总结:**Opus 想,Sonnet 做**。
## `/model`:切换模型
用 `/model` 可以在会话中随时切换使用的模型。比如日常用 Sonnet,遇到复杂问题临时切到 Opus,处理完再切回来。
## 输出风格控制
在 `/config` 里选择 "Output style",有两个不常见但很有用的模式:
* **Explanatory 模式**:Claude 在完成任务之间会插入「知识点」,解释相关的框架和代码模式,适合学习新项目
* **Learning 模式**:协作学习模式,Claude 会在代码中添加 `TODO(human)` 标记让你自己实现,而不是直接给答案
你也可以在 `~/.claude/output-styles/` 创建自定义的输出风格文件(Markdown 格式),直接修改系统提示词。注意:自定义输出风格会**完全替换**默认的编程系统提示词,除非设置 `keep-coding-instructions: true`。
# Claude Skills 概念介绍
## 引言
2025 年 10 月,Anthropic 悄然发布了一项名为 Claude Skills 的新功能。这个看似低调的更新,却被知名技术博主 Simon Willison 评价为"可能比 MCP 更重要",并预测将引发 AI 工具领域的"寒武纪大爆发"。
这样的评价并非空穴来风。如果你经常使用 AI 助手,一定遇到过这样的困扰:每次开始新对话,都要重复输入相同的工作流程说明;好不容易把 AI 调教到满意的状态,换个对话窗口又要从头来过。Skills 正是为解决这个痛点而生的。
## 理解 Claude Skills
想象你是一家公司的老板,新员工入职时你会给他一本工作手册,里面详细记录了公司的工作流程、品牌规范、常见问题的处理方式。Claude Skills 就是给 AI 助手的这本"工作手册"——让它能够以可重复、标准化的方式完成特定任务。
从技术角度来说,Skills 是包含指令、脚本和资源的文件夹,Claude 可以在需要时动态加载它们。每个 Skill 教会 Claude 如何以一致的方式完成某类任务,而且这些知识可以跨对话持久保存。这意味着你只需要"培训"一次,之后无论何时使用,Claude 都会记得该怎么做。
### 三大组成部分
一个完整的 Skill 由以下三部分构成:
| 组件 | 作用 | 是否必需 |
| ------------ | -------------------------------- | ---- |
| **SKILL.md** | 核心指令文档,包含元数据和详细指令 | 必需 |
| **参考资料** | 品牌指南、政策文件、模板等补充信息 | 可选 |
| **脚本** | Python/JavaScript 代码,处理复杂计算或文件操作 | 可选 |
其中 SKILL.md 是整个 Skill 的"灵魂",它的基本结构如下:
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
文件开头的 YAML frontmatter 包含两个关键字段:`name` 是技能的标识名称,最多 64 个字符;`description` 告诉 Claude 这个技能是做什么的、什么时候应该使用它,最多 200 个字符。Claude 正是根据这个描述来判断何时应该调用某个 Skill,所以写得越清晰准确,Skill 被正确触发的概率就越高。
### 应用场景
Skills 的应用场景非常广泛,覆盖了日常工作中的各种重复性任务:
**文档处理**:批量创建 Excel 表格、PPT 演示文稿、Word 文档、PDF 报告。Anthropic 官方就提供了一套文档技能,开箱即用。
**品牌合规**:将公司的品牌色、Logo 使用规则、间距规范、语气风格打包成 Skill,确保 AI 生成的所有内容都符合品牌标准。
**会议纪要**:自动总结会议记录、提取行动项、分配负责人、生成跟进邮件。
**数据分析**:执行标准化的分析流程,比如竞争情报扫描(结构化提取产品更新、定价变化、分析师评论)、财务分析(分析财报、构建财务模型)。
**项目管理**:从目标构建项目计划、建议里程碑、生成周报/投资者简报。
## 渐进式披露架构
Skills 最精妙的设计在于它的信息加载方式。传统的 MCP 工具描述可能消耗数千甚至数万个 token,而 Skills 的元数据仅占用数十个 token。这意味着你可以同时启用大量 Skills,完全不用担心上下文窗口被工具描述占满。
这种高效源于一种叫做**渐进式披露**(Progressive Disclosure)的架构设计。Skills 采用三层信息结构,按需逐层加载——就像一本有目录的手册:
```
📚 Skills 工作手册
│
├─ 📋 目录 ─────────────────────────── 【元数据层】启动时预加载
│ │
│ │ name: "weekly-report"
│ │ description: "根据工作内容生成标准化周报"
│ │
│ │ ✓ 仅占 30-50 tokens
│ │ ✓ 所有 Skills 的目录同时可见
│ │
│
├─ 📖 正文章节 ─────────────────────── 【核心文档层】相关时加载
│ │
│ │ # Weekly Report Generator
│ │
│ │ ## Instructions
│ │ 按以下结构生成周报...
│ │
│ │ ## Examples
│ │ 输入:这周完成了登录功能...
│ │ 输出:### 本周完成 ...
│ │
│ │ ⚡ Claude 判断需要时才展开
│ │ 📊 消耗数百至数千 tokens
│ │
│
└─ 📎 附录 ─────────────────────────── 【引用资源层】需要时加载
│
│ references/
│ ├── brand-guide.md 品牌规范
│ ├── template.xlsx 报告模板
│ └── examples/ 历史周报
│
│ 🔍 仅在明确需要时加载
│ 📦 可包含大量参考资料
```
你先看目录知道有哪些章节(元数据层),找到需要的章节再翻开阅读(核心文档层),最后如果需要更多细节再去查附录(引用资源层)。
| 层级 | 内容 | 加载时机 | Token 消耗 |
| --------- | ------------------ | ------ | -------- |
| **元数据层** | name + description | 启动时预加载 | 30-50 |
| **核心文档层** | SKILL.md 全文 | 相关时加载 | 数百至数千 |
| **引用资源层** | 参考文件、模板等 | 需要时加载 | 按需 |
这完全符合大语言模型的本质——"输入文本让模型理解"。Skills 没有引入复杂的协议或 API 调用,而是通过精心组织的文本结构,让 AI 能够高效地获取和运用知识。Simon Willison 评价这种设计"简洁得令人发指",正是因为它把复杂的问题用最朴素的方式解决了。
## 核心优势
### Token 效率
Skills 的渐进式披露架构带来了极高的 token 效率。我们可以用一个简单的对比来理解:
| 方案 | 启动时 Token 消耗 | 100 个技能的总消耗 |
| ------------- | ------------ | ----------- |
| 传统方案(全量加载) | 数千至数万 | 可能超出上下文窗口 |
| Skills(渐进式披露) | 30-50 | 3000-5000 |
由于每个技能的元数据仅占数十个 token,你可以同时启用几十甚至上百个 Skills,完整内容按需加载,不会浪费宝贵的上下文空间。
### 可组合性
多个 Skills 可以自动协同工作。当你提出一个复杂任务时,Claude 会智能识别需要调用哪些 Skills,并协调它们一起完成任务。
比如你说"根据这份销售数据生成季度报告",Claude 可能会:
1. 调用数据分析 Skill 处理原始数据
2. 调用图表生成 Skill 创建可视化
3. 调用文档 Skill 生成最终报告
整个过程你不需要手动指定使用哪个 Skill,Claude 会根据任务需求自动选择和组合。
### 可移植性
同一个 Skill 可以在 Anthropic 生态系统的所有平台使用:
| 平台 | 说明 |
| ----------- | ------------ |
| Claude.ai | 网页版,适合普通用户 |
| Claude Code | 命令行工具,适合开发者 |
| API | 程序化集成,适合系统开发 |
你为团队创建的品牌写作 Skill,可以在所有这些平台上保持一致的行为,真正实现**一次构建、随处使用**。
> **其他 AI 平台的策略**:目前 Skills 是 Anthropic 独有的功能。OpenAI 采用 Custom GPTs + Assistants API 的双轨策略(两套系统不统一);Microsoft Copilot 和 Google Gemini 则专注于各自生态系统的深度集成,而非可复用的技能模块。Claude Skills 被认为是一个有意义的差异化特性。
### 效率数据
根据 Anthropic 内部基准测试,使用 Skills 的团队**减少了 73% 的重复提示工程时间**。这不仅意味着效率提升,更重要的是工作流程的标准化和可复用性——团队成员不再需要各自维护一套提示词,而是共享同一套经过验证的 Skills。
## 小结
Claude Skills 本质上是给 AI 助手的**可重用工作手册**。它通过渐进式披露架构实现了极高的 token 效率,让 AI 能够掌握大量专业知识而不会占用宝贵的上下文空间。
记住三个关键词,你就掌握了 Skills 的精髓:
| 关键词 | 含义 |
| ------- | ------------------- |
| **高效** | 元数据仅占数十个 token,按需加载 |
| **可组合** | 多个 Skills 自动协同工作 |
| **可移植** | 跨平台一致体验 |
了解了概念之后,下一篇《[Claude Skills 实战指南](/docs/notes/claude-skills/practice)》将带你动手实践:如何启用和安装 Skills、创建你的第一个自定义 Skill、以及避开常见的坑。
如果你希望进一步规范化的工作流程,可以参考《[规格驱动开发是什么](/docs/notes/speckit/concept)》,了解如何将 AI 编程从「直觉」升级到「工程」。
# Claude Skills 实践指南
## 快速回顾
在[上一篇文章](/docs/notes/claude-skills/concept)中,我们了解了 Skills 的核心概念:它是给 AI 助手的可重用工作手册,通过渐进式披露架构实现了极高的 token 效率,具备高效、可组合、可移植三大特点。本文将从实战角度出发,帮助你理解 Skills 与其他功能的区别,学会启用、安装和创建 Skills,掌握最佳实践并避开常见陷阱。
## 功能对比
Claude 生态中有多种功能,初次接触可能会困惑它们之间的区别。下表可以帮助你快速区分:
| 功能 | 是什么 | 最适合 | 持久性 |
| ------------- | ----- | ----------- | ------ |
| **Skills** | 专业知识包 | 重复性任务、标准化流程 | 跨对话持久 |
| **Prompts** | 即时指令 | 一次性请求 | 仅限当前对话 |
| **Projects** | 知识库 | 背景信息、项目文档 | 项目工作区内 |
| **MCP** | 连接器 | 外部数据、API 调用 | 持续连接 |
| **Subagents** | 子代理 | 任务委托、并行处理 | 跨会话 |
### Skills vs MCP
这是最常见的困惑。核心区别:**MCP 将 Claude 连接到数据,Skills 教 Claude 如何处理数据**。两者互补而非替代。
| 维度 | Skills | MCP |
| ------------ | -------------------- | ---------------- |
| **核心功能** | 教 Claude 如何执行任务 | 连接 Claude 到外部系统 |
| **Token 消耗** | 极低(数十个 token) | 较高(数千至数万 token) |
| **技术复杂度** | 简单(Markdown + YAML) | 复杂(完整协议规范) |
| **典型场景** | 品牌写作、报告生成、工作流程 | 数据库查询、API 调用、云服务 |
| **可移植性** | 跨 Claude.ai/Code/API | 已被多家模型公司采用 |
理解了这个区别,你就知道何时该用哪个了。需要查询数据库、调用 API、访问云服务时,用 MCP;需要遵循特定的写作风格、执行标准化流程、复用专业知识时,用 Skills。
最佳实践是两者结合使用:用 MCP 连接你的 CRM 系统获取客户数据,用 Skills 定义如何分析这些数据并生成报告。
### Skills vs Subagents
核心区别:**Skills 让 Claude 更擅长某类任务;Subagents 让 Claude 把任务委托给独立的"专家员工"**。
| 维度 | Skills | Subagents |
| -------- | -------------------- | -------------------------- |
| **核心功能** | 提供专业知识和指令 | 独立执行任务的子代理 |
| **上下文** | 注入主对话上下文 | 拥有独立的上下文窗口 |
| **适用场景** | 让 Claude 更擅长某类任务 | 复杂、多步骤的独立任务 |
| **激活方式** | 根据描述自动匹配 | 手动调用或 Claude 自动委托 |
| **可移植性** | 跨 Claude.ai/Code/API | 仅限 Claude Code 和 Agent SDK |
形象地说,Skills 像培训材料——让 Claude 学会如何做某件事;[Subagents](/docs/notes/claude-subagent) 像专职员工——拥有自己的工位(上下文)和权限(工具),独立完成任务后汇报结果。
两者可以组合使用:比如一个代码审查子代理可以加载语言特定的最佳实践 Skill,实现"专家 + 专业知识"的组合效果。根据 [Anthropic 研究](https://www.anthropic.com/engineering/multi-agent-research-system),多代理系统(Claude Opus 4 主代理 + Claude Sonnet 4 子代理)在内部评估中比单代理高出 90.2%。
### Skills vs 斜杠命令
如果你用过 Claude Code,一定熟悉 `/commit`、`/review` 这样的[斜杠命令](/blog/claude-code-best-practices)。核心区别:**Skills 根据上下文自动激活;斜杠命令需要手动输入触发**。
| 维度 | Skills | 斜杠命令 (Slash Commands) |
| -------- | ---------------------------- | --------------------- |
| **激活方式** | 自动激活(根据上下文匹配) | 手动输入(如 `/commit`) |
| **触发条件** | Claude 根据 description 判断是否相关 | 用户明确输入命令 |
| **适用场景** | "始终在线"的能力增强 | 明确的、可重复的操作 |
| **用户感知** | 无感知,自动生效 | 需要记住命令名 |
举例说明:当你输入 `/commit`,Claude 执行预定义的提交流程——这是斜杠命令;当你说"帮我写一份周报",Claude 自动识别并加载周报生成 Skill,无需你输入任何命令——这是 Skills。
简单记忆:斜杠命令是快捷键,需要你主动触发;Skills 是背景知识,Claude 自动判断何时使用。
### Skills vs Plugins
Plugins 是 Claude Code 的扩展包机制。核心区别:**Skills 是自动激活的能力扩展;Plugins 是打包分发的完整工作流配置**。
| 维度 | Skills | Plugins |
| -------- | ----------------------- | --------------------- |
| **核心功能** | 专业能力扩展 | 打包分发工作流 |
| **激活方式** | 根据上下文自动激活 | 安装后组件合并 |
| **适用范围** | 跨平台(Claude.ai/Code/API) | 仅限 Claude Code |
| **包含内容** | 指令 + 脚本 + 资源 | 斜杠命令 + hooks + skills |
| **分发机制** | 单独文件夹 | 通过 marketplace 安装 |
关键理解:Plugins 可以包含 Skills(在 `skills/` 目录下),是更大的打包单元。当你安装一个 Plugin 时,其中的 Skills 会自动激活,斜杠命令会出现在自动补全中,hooks 会与现有配置合并。
简单来说:用 Skills 扩展 Claude 的能力,用 Plugins 在团队间分发标准化的工作流配置。
## 实战教程
### 方式一:启用内置 Skills
这是最简单的入门方式。Anthropic 官方提供了一套实用的文档技能:
| 技能 | 功能 |
| --------------------- | -------------------- |
| **Excel (xlsx)** | 创建电子表格、分析数据、生成带图表的报告 |
| **PowerPoint (pptx)** | 创建演示文稿、编辑幻灯片、分析演示内容 |
| **Word (docx)** | 创建文档、编辑内容、格式化文本 |
| **PDF (pdf)** | 生成格式化的 PDF 文档和报告 |
**启用步骤**:
1. 登录 [Claude.ai](https://claude.ai)
2. 点击右上角头像,进入 **Settings**
3. 找到 **Capabilities** 选项
4. 启用你需要的技能
启用后可以直接测试:"帮我创建一个 Q3 销售预算的 Excel 表格,包含月度明细和总计"。
> **注意**:需要 Pro、Max、Team 或 Enterprise 计划,并且需要启用代码执行功能。
### 方式二:安装社区 Skills
如果你使用 Claude Code,可以通过命令安装社区贡献的 Skills。
**通过插件市场安装**:
```bash
# 添加官方 Skills 仓库
/plugin marketplace add anthropics/skills
# 安装文档技能包
/plugin install document-skills@anthropic-agent-skills
# 安装示例技能包
/plugin install example-skills@anthropic-agent-skills
```
**Skills 存储位置**:
| 位置 | 路径 | 说明 |
| --------- | ------------------- | --------------- |
| 个人 Skills | `~/.claude/skills/` | 仅你自己可用 |
| 项目 Skills | `.claude/skills/` | 随 git 版本控制,团队共享 |
### 方式三:创建自定义 Skill
这才是 Skills 真正的威力所在——创建专属于你的工作流程。
**第一步:创建文件夹结构**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
一个完整的 Skill 文件夹可能是这样的:
```
weekly-report/
├── SKILL.md # 核心指令(必需)
├── template.md # 周报模板(可选)
└── examples/ # 示例周报(可选)
├── good-example.md
└── bad-example.md
```
**第二步:编写 SKILL.md**
SKILL.md 是整个 Skill 的核心。它由两部分组成:YAML frontmatter(元数据)和 Markdown 正文(详细指令)。
**必需元数据**:
| 字段 | 要求 | 说明 |
| ------------- | --------- | ------------------------ |
| `name` | 最多 64 字符 | 技能的唯一标识名称 |
| `description` | 最多 200 字符 | 告诉 Claude 何时使用此技能(非常重要!) |
**可选元数据**:
| 字段 | 说明 |
| --------------- | ------------------------------------ |
| `dependencies` | 所需软件包,如 `python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | 允许使用的工具列表 |
| `model` | 可选的模型覆盖 |
一个完整的周报生成 Skill 示例:
```yaml
---
name: weekly-report
description: 根据本周工作内容生成标准化的周报,包含进展、问题和下周计划
---
# 周报生成助手
## 使用场景
当用户需要生成周报、工作总结或进度汇报时,使用此技能。
## 输出格式
请按以下结构生成周报:
### 本周完成
- 列出已完成的主要工作项
- 每项包含简短说明和成果
### 进行中
- 列出正在进行的工作
- 标注当前进度和预期完成时间
### 遇到的问题
- 列出阻碍进展的问题
- 如果有,说明需要的支持
### 下周计划
- 列出下周的主要任务
- 按优先级排序
## 风格要求
- 使用简洁的表达
- 避免过于技术化的术语
- 突出成果和影响
## 示例
**输入**:这周完成了用户登录功能,修复了 3 个 bug,参加了产品评审。
**输出**:
### 本周完成
- 用户登录功能开发:完成前后端联调,支持邮箱和手机号登录
- Bug 修复:解决了 3 个高优先级问题,提升系统稳定性
### 进行中
- (无)
### 遇到的问题
- (无)
### 下周计划
- 开始用户注册功能开发
- 编写单元测试用例
```
**第三步:测试**
在 Claude 中测试:"帮我生成本周的周报。这周我完成了用户登录功能的开发,修复了 3 个 bug,还参加了两次产品评审会议。"
### 使用 Skill Creator
如果你不想从头写 SKILL.md,Claude 内置了一个 skill-creator 技能,可以交互式引导你创建:
```
Help me create a skill for [your workflow]
```
Claude 会通过一系列问题帮你梳理需求,然后生成 SKILL.md 的初稿。
## 技术原理
### Skills 作为元工具系统
Skills 本质上是一个**元工具系统**——它不直接执行代码,而是将专门指令注入对话上下文,改变 Claude 的推理方式。
当你触发一个 Skill 时,发生了两件事:
1. **元数据消息**:一个可见的状态指示器,显示正在加载哪个 Skill
2. **技能提示**:完整的 SKILL.md 指令发送给 Claude,但对用户隐藏
### 发现与选择机制
Claude 如何知道该调用哪个 Skill?答案是:**完全依靠语言理解**。
所有已启用 Skills 的 name 和 description 会被格式化为一个动态列表,写入系统提示。当你发送消息时,Claude 使用原生语言理解能力匹配你的意图,决定是否调用某个 Skill。
这就是为什么 `description` 字段如此重要——它是 Claude 判断的唯一依据。没有复杂的算法路由,决策完全在 Claude 的推理过程中完成。
## 最佳实践
经过大量实践,社区总结出了创建 Skills 的四条黄金法则:
**1. 保持专注**
一个 Skill 应该只做一件事。多个专注的 Skills 远比一个大而全的 Skill 好用,这样不仅更容易维护,也更容易组合使用。
**2. 清晰描述**
description 字段决定了 Claude 何时调用你的 Skill,务必写清楚适用场景。"根据销售数据生成季度分析报告"是好的描述,"处理数据"则太过笼统。
**3. 提供示例**
在 SKILL.md 中包含输入输出示例能大幅提升输出的稳定性,尤其是对于有特定格式要求的任务。
**4. 从简单开始**
先用纯 Markdown 写基本指令,验证效果后再考虑添加脚本,逐步增加复杂度。
### 常见问题排查
| 问题 | 可能原因 | 解决方案 |
| ----------- | ---------------- | --------------------- |
| Skill 没有被触发 | description 不够准确 | 改写为更具体的使用场景描述 |
| Skill 没有被触发 | Skill 未正确安装 | 检查文件路径和命名 |
| 输出不稳定 | 缺少示例 | 添加更多输入输出示例 |
| 输出不稳定 | 指令太模糊 | 添加约束条件和格式要求 |
| 加载太慢 | 文件太大 | 将大文件移到 references 子目录 |
### 安全注意事项
Skills 可以执行代码,因此安全性非常重要:
* **来源可信**:只使用来自可信渠道的 Skills
* **审查脚本**:安装前检查 Skills 中的脚本代码
* **保护敏感信息**:不要在 Skills 中硬编码 API 密钥或密码
* **权限管理**:团队使用时,注意 Skills 的共享范围
## 当前限制
作为一项新兴功能,Skills 目前存在一些限制:
| 限制 | 说明 |
| ----------------------- | ------------------ |
| ~~**仅限 Anthropic 生态**~~ | ✅ **已解决** - 见下方说明 |
| **缺乏审核机制** | 暂无内置的审查或审核工作流 |
| **学习曲线** | 团队需要调整工作流并建立版本管理流程 |
| **新兴阶段** | 生态系统仍在发展中 |
> **🎉 重大更新(2025年12月18日)**:Anthropic 正式将 Agent Skills 作为[开放标准](https://agentskills.io)发布。规范和参考 SDK 已在 [agentskills.io](https://agentskills.io) 公开。
>
> **已采用的公司/产品**:
>
> 
>
> * **Microsoft**:VS Code、GitHub 已集成
> * **OpenAI**:ChatGPT、Codex CLI 采用相同架构
> * **编程工具**:Cursor、Goose、Amp、OpenCode
> * **合作伙伴 Skills**:Atlassian、Figma、Canva、Stripe、Notion、Zapier
>
> 同时,Anthropic、OpenAI 和 Block 共同创立了 [Agentic AI Foundation](https://www.linuxfoundation.org/)(由 Linux Foundation 托管),Google、Microsoft、AWS 也已加入。这意味着 Skills 正在从单一厂商功能演变为行业标准,为 Claude Code 编写的 Skills 可以与 OpenAI Codex CLI 互操作。
>
> 参考来源:
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## 学习资源
### 官方资源
| 资源 | 链接 | 说明 |
| ----------------- | -------------------------------------------------------------------------------------------------------------------- | --------------- |
| Skills GitHub 仓库 | [anthropics/skills](https://github.com/anthropics/skills) | 官方示例,22k+ Stars |
| Claude Code 文档 | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Skills 使用指南 |
| 帮助中心 | [support.claude.com](https://support.claude.com) | 常见问题解答 |
| 技术博客 | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | 技术原理深度解析 |
| API 快速入门 | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | 开发者集成指南 |
| Agent Skills 开放标准 | [agentskills.io](https://agentskills.io) | 官方规范和 SDK |
### 社区精选
| 资源 | 链接 | 说明 |
| --------------------- | ------------------------------------------------------------------------------------- | -------------------- |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Skills 精选合集 |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148+ 斜杠命令,54 个 AI 代理 |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | 办公文档创建编辑技能 |
### 推荐阅读
| 文章 | 作者 |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic 官方 |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## 展望
Skills 的出现代表了 AI 工具发展的一个重要方向——让 AI 不仅能够执行任务,还能够学习和记忆特定的工作方式。Simon Willison 预测 Skills 将带来 AI 工具领域的"寒武纪大爆发",这个判断并非夸张。
随着越来越多的开发者和团队开始构建和分享 Skills,我们可能会看到:
* **专业化 Skills 市场**:各行各业的专家将知识打包成可复用的 Skills
* **Skills 与 MCP 深度融合**:形成完整的端到端工作流
* **企业级 Skills 平台**:团队协作、版本管理、权限控制
现在正是入场的好时机。立即可做的是登录 Claude.ai 启用文档技能;这周可以尝试安装一个社区 Skill,创建你的第一个简单 Skill;长期来看,识别团队中的重复性工作,逐步构建专属的技能库,将是提升效率的有效途径。
### 延伸阅读
* 《[Claude 系统架构全解析](/docs/notes/claude-architecture)》— Skills 在整个 Claude 系统中的定位
* 《[Claude Subagent 完全指南](/docs/notes/claude-subagent)》— 深入了解 Subagent 机制
* 《[我的 Claude Code 最佳实践](/blog/claude-code-best-practices)》— Claude Code 日常使用技巧
# Skill-Creator 深度解析:用数据驱动你的技能开发
## 引言
本文基于 2026 年 3 月的信息撰写,对应 Claude Code v2.1+ 版本。
如果你已经读过[概念篇](/docs/notes/claude-skills/concept)和[实践篇](/docs/notes/claude-skills/practice),你应该已经知道如何手动编写一个 SKILL.md 文件了——定义 frontmatter、写一段指令、保存到 `.claude/skills/` 目录,完事。
但这里有一个根本性的问题:**你怎么知道你的技能真的好用?**
你可能改了一段措辞觉得效果更好了,但那只是你的主观感受。也许换一个不同的提示词,新版本反而更差。也许你的技能和不加技能相比根本没有提升——Claude 本身就能做得一样好。
在概念篇和实践篇中,技能开发的流程是这样的:**写了 → 试了 → 感觉还行 → 上线**。整个过程依赖直觉,没有量化,也无法回答"这个技能到底比不加技能好多少"。而 Skill-Creator 把这件事变成了工程:**写了 → 并行测试有/无技能 → 盲测 A/B 对比 → 量化评分 → 反馈迭代 → 数据验证**。
这就是 Skill-Creator 存在的意义。它不只是帮你"生成一个 SKILL.md",而是提供了一套完整的**创建 → 测试 → 评估 → 优化**循环,让你用数据说话。
## 什么是 Skill-Creator
Skill-Creator 本身也是一个 Skill——一个 33KB 的 SKILL.md 文件加上配套的子代理指导文件、Python 脚本和 HTML 查看器。它的目录结构长这样:
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
安装也很简单:
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## 跟着做一遍:评估并优化一个已有技能
下面用我自己实际在用的技能走完 Skill-Creator 的完整流程。我维护了一个 Claude Code 插件市场 [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub),其中 `yux-video-summary` 技能用于将视频字幕转为结构化摘要——支持中英文语言检测、DUAL\_FILE/SINGLE\_FILE 两种输出模式、filler 词清理等。技能的 SKILL.md 长这样:
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
技能已经写好了,但**怎么知道它真的好用?** 这就是 Skill-Creator 登场的时候。
> Skill-Creator 源码中有一条重要的写作原则:*"Try hard to explain the **why** behind everything. If you find yourself writing ALWAYS or NEVER in all caps, that's a yellow flag — reframe and explain the reasoning."* 意思是:好的技能应该**解释为什么**,而不是堆砌死板的规则。
### 第 1 步:创建测试用例并运行评估
核心问题:**这个技能和不用技能比,真的更好吗?**
打开 Claude Code,直接输入:
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator 会先读取技能定义和 schemas,然后自动生成测试用例和量化断言。我的这次运行生成了 3 个测试用例和 39 条 assertions:
注意它不是随便编测试用例的——它**读懂了**技能中定义的 DUAL\_FILE 和 SINGLE\_FILE 两种输出模式,针对性地设计了覆盖不同视频类型(教程、播客采访、技术分享)和语言组合的场景。Assertions 的设计也很讲究,从语言检测、输出模式选择到内容质量和中英文 filler 词清理,比我自己想测试维度要全面得多。
接着,系统对每个测试用例**同时启动两个独立的子代理**——**with\_skill**(加载技能)和 **without\_skill**(基线,不加载任何技能)。一次性启动了 **6 个并行代理**(3 个测试用例 × 2 个版本),每个在**独立的 worktree** 中运行,互不干扰。
> Anthropic 的 PDF 技能之前在处理非可填写表单时有问题——Claude 需要在没有定义字段的情况下在精确坐标放置文本。通过 Eval 隔离出了这个失败点,团队随后修复了定位逻辑。这就是 Eval 的价值——**把"感觉不太对"变成"这里具体出了什么问题"**。
### 第 2 步:三个子代理接力评分
所有运行完成后,三个专业子代理**自动**依次登场:
**Grader(评分器)** 逐条检验断言。它会检查 with\_skill 版本的摘要是否包含了 Overview 表格、是否正确选择了 DUAL\_FILE 模式、filler 词是否已清理,然后记录每条的通过/失败和证据,生成 `grading.json`:
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**Comparator(比较器)** 做盲测 A/B 对比——它收到两份摘要,但**不知道哪个是技能版、哪个是基线版**。它只看到"输出 A"和"输出 B",根据自己的质量标准独立评判,确定赢家。
**Analyzer(分析器)** 综合上面的结果做诊断:哪些断言不管有没有技能都通过了(说明这个断言**没有区分度**,该换更好的断言)、哪些结果方差很高(测试**不稳定**)、时间和 token 的权衡如何。最终给出改进建议。
### 第 3 步:在 Eval Viewer 中审查结果
评分完成后,Skill-Creator 会**自动在浏览器中打开**一个 HTML 查看器。
**Outputs 选项卡** 可以逐个查看每个测试用例的输出,底部有反馈文本框——写下你觉得哪里不够好,比如"摘要缺少时间线"、"filler 词没清理干净"。看完所有用例后点击 **Submit All Reviews**,反馈保存到 `feedback.json`。
**Benchmark Results 选项卡** 可以看到量化对比:with\_skill 和 without\_skill 的通过率、耗时、token 消耗,以及每条 assertion 的逐项对比。
### 第 4 步:迭代改进,直到满意
回到 Claude Code 告诉它你反馈完了,Skill-Creator 会读取 `feedback.json`,综合 benchmark 数据给出分析和改进建议:
我的 skill 在 97% 通过率下表现已经不错,Skill-Creator 精准定位到了一个小问题——采访类视频缺少 Notable Quotes 段落,并提出了修复建议。
关键在于它**不会针对单个测试用例打补丁**——它会泛化你的反馈,理解背后的需求并调整技能的整体结构,然后改写 SKILL.md,重新运行所有测试到 `iteration-2/` 目录,打开新的 Eval Viewer 方便你对比两轮的输出。这个循环一直持续,直到你满意。
> Skill-Creator 源码中一段值得注意的改进哲学:*"We're trying to create skills that can be used a million times across many different prompts. Rather than put in fiddly overfitty changes, or oppressively constrictive MUSTs, if there's some stubborn issue, try branching out and using different metaphors."* 核心思想:**避免过拟合**到测试用例上,追求泛化能力。
### 第 5 步(可选):优化描述,让技能在正确的时候触发
技能质量验证了,但还有一个容易被忽视的问题:技能的 `description` 字段决定了 Claude **什么时候**会调用它。
输入:
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator 自动生成约 20 个评估查询(一半应该触发、一半不应该触发),在浏览器中打开审查界面:
注意这些查询既有中文也有英文,覆盖了各种真实表达方式。"不该触发"的查询不能太离谱——好的反例是"帮我总结这份会议纪要",它和视频摘要共享了"总结"这个关键词,但实际需要的是文档处理技能而不是视频摘要。
你可以直接在页面里编辑查询文本、点击 **+ Add Query** 添加新的、用 Delete 按钮删除不合适的,还可以切换每条查询的 Should Trigger 开关。确认无误后点击 **Export Eval Set** 导出 JSON 文件,回到 Claude Code 告诉它你导出好了,系统就会在后台自动跑优化循环:
整个过程全自动——将查询按 60/40 分为训练集和测试集,在训练集上迭代优化描述(最多 5 轮),用测试集成绩选最佳版本以避免过拟合。跑完后会输出优化前后的 description 对比:
优化后的 description 变得更具体——明确了支持的文件类型(.vtt/.srt)、强调了 pipeline 特性(filler 清理、DUAL/SINGLE\_FILE 逻辑),同时用 MUST USE 排除了不该触发的场景。Anthropic 内部用这套优化器跑了自己的文档创建类技能,结果 **6 个公共技能中有 5 个**的触发精度得到了提升。
### 进阶用法:动态上下文注入
如果你想让技能在加载时**自动注入上下文**,可以用 Skills 2.0 的 `!` 语法在 SKILL.md 中嵌入 Shell 命令:
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
这些命令在 Claude **看到技能之前**就已经执行完毕,数据直接嵌入了提示。相比让 Claude 自己去一个个探索文件,省下了大量时间和 token。
## 两类技能:你该创建哪种?
在使用 Skill-Creator 之前,有必要了解 Anthropic 定义的两种技能类型:
**能力提升型**——让模型做到原本做不到或做不好的事。比如:
* 图片生成技能:Claude 原生不能生成图片,但可以通过技能调用 nanobanner 等工具实现
* 前端设计技能:默认的 AI 设计往往很"AI 味",好的设计技能可以大幅提升质量
**偏好编码型**——将你的特定工作流固化下来。模型已经具备各个单项能力,但你需要一个精确的执行顺序。比如:
* PR 审查技能:按固定流程检查代码安全性,输出风险等级报告
* 视频摘要技能:按特定模板结构输出,自动做语言检测和 filler 词清理
这两类技能需要测试的原因不同:**能力提升型**可能随着模型进化而变得不必要——如果基线(without\_skill)也能通过所有断言,说明模型原生就够了,这个技能可以退役;**偏好编码型**更持久,但需要验证它是否真的忠实于你的工作流。
Skill-Creator 的评估能力让你可以**持续验证技能是否仍有价值**,而不是盲目使用一个可能已经过时的技能。
## 社区怎么说
Skill-Creator 更新后引发了不少讨论,从 X/Twitter 到 Reddit 到独立博客,真实反馈比官方文档更有参考价值。
### 真的有用吗?数据说话
最直接的问题:**加了技能真的比不加好吗?** 几个实测给出了明确答案。
Reddit u/hashpanak 对标题生成技能跑了 eval,with\_skill 通过率 100%,without\_skill 只有 60%。被问到 token 开销值不值,他回复:"绝对值,优化之后重复任务可以转成脚本,**反而省 token**。" u/spences10 更极端——跑了 250 次沙盒 eval,把技能激活率从 84% 拉到了 100%。评论区 u/Manfluencer10kultra 说:"**这应该成为标准做法。**"
博主 [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/) 对 WordPress 安全技能做了 benchmark:21 条断言全部通过(基线只有 90.5%),速度还快了 9.9%。他的总结:**"之前 Skills 是艺术,现在是工程。"**
@0zhuxiaofeng 从实际工作流角度给出了更具体的数字:"用了一个月了,最大变化是 run\_eval 让 skill 可以自我打分了。我跑内容运营的 agent 现在每次发布后自动评估效果,差的 skill 直接淘汰重写。**人工干预从每天 3 小时降到半小时**。"
### 被忽视的盲点:触发 ≠ 质量
博主 [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/) 指出了一个没人提的盲点:**技能可以通过质量评估却在触发评估上挂掉**——输出质量很好,但永远不会被调用。他经过三轮 `run_loop.py` 优化后触发 eval 达到 13/13。核心洞察:"技能的 description 不是元数据,**而是一个可学习的参数**——你需要对真实的路由行为做优化。"
这一点和 @DrWang5257 的建议不谋而合:"别一上来整套重写,先拆成**触发条件、输入模板、失败回退**三段,逐段迭代。这样更新速度快,翻车率也低。"
### 真实痛点
效果虽好,坑也不少:
* **token 消耗巨大**。@konghao10 直言"耗 token 巨凶"——6 个并行代理同时跑确实不便宜。Reddit u/munkymead 也说"要正经测一轮很贵"。
* **技能多了会打架**。[RoboRhythms 博主 Noah Albert](https://www.roborhythms.com/best-claude-code-skills-2026/) 发现**技能到 8-10 个时开始出问题**:Claude 会自我质疑输出、生成更啰嗦的前言、偶尔出现技能间指令冲突。不过 Reddit u/Specialist\_Solid523 反驳:"写得差的技能才吃上下文。**写得好的技能几乎总是让你的 token 使用更高效。**"
* **SKILL.md 越迭代越长**。Reddit u/IulianHI 指出一个矛盾:随着迭代改进,技能文件不断膨胀,**反而挤占了真正做事的上下文窗口**。测试用例只覆盖 happy path 也会漏掉关键的 5%。
* **版本管理缺失**。@fengqve 吐槽"为什么 Skill **没有版本的概念**,更新多了描述哪次更新都说不清楚"——这在多轮迭代后尤其痛苦。
* **无头模式有 bug**。GitHub 上有一个关键问题:`claude -p` 模式下技能永远不触发,导致描述优化循环的 recall 恒为 0%([#36570](https://github.com/anthropics/claude-code/issues/36570))。
### 更远的思考:递归自我改进
@vista8 分享了一篇相关论文 [Memento-Skills: Let Agents Design Agents](https://github.com/Memento-Teams/Memento-Skills),评论区有人精准总结:"Skill 的核心瓶颈就在迭代——写第一版容易,但让它在真实场景里越用越好很难。如果能自动化这个'用→评估→改进'的循环,等于给 Agent 装了个自我进化引擎。"
Reddit r/ClaudeAI 上一个 104 赞的帖子也在讨论这个方向。但最高赞评论泼了一盆冷水——u/Tatrions 说:"递归循环能跑,但难的是**知道什么时候该信任改进**。我们发现必须做证据门控——**除非一个失败出现至少两次,否则不要提交修改**。不然每轮循环都在'修复'本来没坏的东西,最终反而更差。"
## 安装与生态
Skill-Creator 作为 Anthropic 官方维护的技能之一,收录在 [anthropics/skills](https://github.com/anthropics/skills) 仓库中,该仓库包含 17+ 个生产级技能。
更广泛的 Skills 生态也在快速增长:[skills.sh](https://skills.sh) 市场提供了便捷的发现和安装体验,社区已维护 1,234+ 个代理技能。
## 写在最后
Skill-Creator 解决的核心问题是:**如何知道你的技能真的有效?**
在没有它的时候,技能开发靠的是"写了 → 试了 → 感觉还行"。有了 Skill-Creator,你可以:
* 用**并行代理**同时测试有技能和无技能的效果
* 用**盲测 A/B 对比**消除评估偏见
* 用 **Eval Viewer** 直观查看结果并留下反馈
* 用**描述优化器**精确控制技能的触发时机
* 用**迭代循环**持续改进直到满意
这和软件工程中测试驱动开发的理念一脉相承——不是"写完代码觉得能跑就行",而是"用测试证明它确实按预期工作"。
Anthropic 在官方博客中提出了一个有趣的展望:随着模型能力的提升,SKILL.md 可能从"实现计划"(告诉 Claude **怎么做**)演变为"规格描述"(告诉 Claude **做什么**,让模型自己想办法)。而 Eval 框架正是这个方向的第一步——Eval 描述的就是"做什么",如果有一天这个描述本身就足以成为技能,那 Skill-Creator 建立的测试体系将变得更加重要。
如果你已经在用 Skills,试试用 `/skill-creator` 对你最常用的技能做一次评估。你可能会惊讶地发现,有些技能其实并没有比不用技能好多少——而这正是优化的起点。
相关阅读:
* [Claude Skills 是什么](/docs/notes/claude-skills/concept) — 理解 Skills 的核心原理
* [实践指南](/docs/notes/claude-skills/practice) — 动手创建你的第一个 Skill
# GSD 概念介绍
## 引言
在 [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) 中,我们了解了一个核心问题:**Context Rot**——随着对话越来越长,Claude 的上下文窗口被失败代码、过时讨论和无关信息填满,输出质量不断下降。
Ralph 的解决方案是"重启一切":用 bash 无限循环每次启动全新的 Claude 实例,通过文件系统传递状态。简单,有效,但也有明显的局限——它只是一个方法论,没有项目理解、没有阶段规划、没有质量验证。你需要自己写 spec、自己编排任务、自己判断"完成了没有"。
正如 Chase AI 在视频中精准概括的:**Ralph Loop 是一个极其强大的武器,但大多数人需要的不是一件武器——而是整个军火库。** Ralph 循环完全依赖于前置准备:你的 PRD 够不够好?功能定义够不够紧凑?你知道"完成"长什么样吗?如果这些问题的答案不精确,那无论循环运行多少次,都只是 garbage in, garbage out。
如果你想要一个**不只是循环运行 Claude,而是真正理解你的项目并可靠地交付代码**的系统呢?
这就是 **GSD (Get Shit Done)** 要做的事。
## 什么是 GSD
GSD 的创建者是 **TÂCHES** (GitHub: glittercowboy),一位独立开发者。他的动机很直接:
> "我不是 50 人的软件公司。我不想玩企业戏剧。我只是一个想做出好东西的创意人。"
在他的直播中,TÂCHES 展示了一个令人震撼的事实:他**从不手写代码**。他用 GSD 在 4 小时内从零构建了一个完整的 macOS 原生音乐生成应用(Sample Digger),全程零手写代码。他的自我定位不是程序员,而是"高层项目经理"——描述愿景、做关键决策、验证结果。GSD 让这种工作方式成为可能。
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> —— TÂCHES
其他规格驱动开发工具——BMAD、SpecKit——都有各自的价值,但它们倾向于引入复杂的企业工作流:sprint ceremonies、story points、stakeholder syncs。对于独立开发者或小团队来说,这些流程本身就是负担。正如 Chase AI 评价的:"It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way."
GSD 的设计哲学是**把复杂性藏在系统里**。用户只需要几个简单的命令,系统在背后处理所有的上下文管理、任务编排和质量验证。项目发布一个月内就获得了近 3,000 GitHub stars 和 14,000 npm 安装量,TÂCHES 几乎每天更新 15-20 次。
### GSD 在工具生态中的位置
| 维度 | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| -------------- | ----------------- | ------------------ | ----------------- | --------------------- |
| 核心定位 | 执行技术(bash loop) | 规格生成工具包 | 企业级框架 | **上下文工程 + 规格驱动** |
| 规划能力 | 无(需自备 spec) | 强(spec→plan→tasks) | 强(完整敏捷流程) | **强(研究→讨论→规划)** |
| 执行自主性 | 最高(AFK 模式) | 需手动触发每步 | 需手动触发每步 | **需手动触发每步** |
| 人类参与模式 | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Context Rot 处理 | 新 session 重启 | 无内置方案 | 无内置方案 | **子代理新鲜上下文** |
| 质量验证 | 依赖外部测试 | 构建检查 | 内置 QA 流程 | **自动验证 + UAT** |
| 用户复杂度 | 最低 | 中等 | 较高 | **低** |
| 系统复杂度 | 最低 | 中等 | 较高 | **高** |
这张表揭示了一个关键取舍:**Ralph 用最低的系统复杂度换取了最高的执行自主性**——启动后就可以去睡觉;而 **GSD 用高系统复杂度换取了规划质量和人类校验**——每个阶段都有你介入的机会。SpecKit 和 BMAD 落在中间地带,提供规划能力但缺少 GSD 的上下文工程和 Ralph 的自主执行。
GSD 和 Ralph 并不矛盾。GSD 继承了 Ralph 的核心原则——新鲜上下文、文件作为真相来源——但在此基础上构建了完整的项目理解和执行体系。如果 Ralph 是"给 AI 一个任务让它反复尝试",GSD 就是"理解你要什么,研究怎么做,规划怎么分步,执行并验证"。
Chase AI 的总结非常到位:**Ralph 循环假设你带着完整蓝图来——GSD 帮你构建这个蓝图。** GSD 接过你半成型的想法,深入提问、代为研究、生成完整 PRD、将其拆解为原子任务,然后端到端地交付项目。而在执行代码时,它使用的正是让 Ralph 循环强大的那些基础原则:子代理的新鲜上下文,以及尽可能小而精确的任务。
## 核心工作流
GSD 的工作流是一个**讨论 → 计划 → 执行 → 验证**的循环,每个阶段都有明确的输入和输出。
### 1. 初始化项目
```text
/gsd:new-project
```
一个命令启动整个流程。系统会:
1. **提问** — 持续追问直到完全理解你的想法(目标、约束、技术偏好、边界情况)
2. **研究** — 派出并行代理调查相关领域(可选但推荐)
3. **需求提取** — 区分 v1、v2 和超出范围的内容
4. **路线图** — 创建与需求对应的阶段规划
你审批路线图,然后开始构建。TÂCHES 的经验是:提供的初始描述越详细,系统追问得越少;越模糊,追问越多。他建议在启动前先准备一份粗略的愿景文档——不需要知道技术栈或实现细节,只需要描述你想要什么。
**产出文件**: `PROJECT.md`、`REQUIREMENTS.md`、`ROADMAP.md`、`STATE.md`
> 已有代码库?先运行 `/gsd:map-codebase`,系统会派出并行代理分析你的技术栈、架构、约定和潜在问题。之后 `/gsd:new-project` 就能基于已有代码库进行规划。
### 2. 讨论阶段
```text
/gsd:discuss-phase 1
```
路线图里每个阶段只有一两句话的描述,这不足以构建你想要的东西。讨论阶段的作用是在研究和规划之前,**捕获你的实现偏好**。
系统会分析当前阶段,识别"灰色地带"——那些有多种合理实现方式的决策点:
* 视觉功能 → 布局、交互、空状态处理
* API/CLI → 响应格式、错误处理、详细程度
* 内容系统 → 结构、语气、深度、流程
你在这里做的每个决定都会直接影响后续的研究和规划质量。跳过这步也可以(系统会用合理默认值),但深入讨论能让系统构建出更符合你期望的东西。
**产出文件**: `{phase}-CONTEXT.md`
### 3. 计划阶段
```text
/gsd:plan-phase 1
```
系统会:
1. **研究** — 调查如何实现当前阶段,以讨论阶段的决策作为指导
2. **计划** — 创建 2-3 个原子任务计划,使用 XML 结构化格式
3. **验证** — 检查计划是否满足需求,循环修正直到通过
一个重要的设计理念是 **Goal-Backward Planning**(目标回溯规划)。不是从"我们应该构建什么"出发,而是问"为了实现目标,什么条件必须成立?"——然后反向推导出计划和任务。TÂCHES 表示这种方式"极大地提升了输出质量",因为每个任务都理解自己和其他任务的关系,而不只是一个待办列表。
每个计划足够小,可以在一个全新的上下文窗口中执行。这是关键——**不会有质量退化**。
**产出文件**: `{phase}-RESEARCH.md`、`{phase}-{N}-PLAN.md`
### 4. 执行阶段
```text
/gsd:execute-phase 1
```
系统会:
1. **波次执行** — 独立任务并行执行,有依赖的按顺序
2. **新鲜上下文** — 每个计划在全新的 200k tokens 上下文中执行,零累积垃圾
3. **原子提交** — 每个任务独立 git commit
4. **目标验证** — 检查代码库是否实现了阶段承诺的功能
在 TÂCHES 的直播演示中,他完成了 3 个完整阶段的开发,**主上下文窗口始终保持在 24%**。GSD Executor 子代理只需要加载不到 1,000 行的上下文就能完成一个完整阶段——你可以连续执行 10 个计划,上下文仍然低于 50%。这和直接在 Claude Code 中工作的体验完全不同:不再是"玩俄罗斯轮盘赌,赌什么时候撞到上下文窗口的墙"。
**产出文件**: `{phase}-{N}-SUMMARY.md`、`{phase}-VERIFICATION.md`
### 5. 验证阶段
```text
/gsd:verify-work 1
```
自动化验证能检查代码是否存在、测试是否通过。但功能是否**按你的预期工作**?这需要你来确认。
系统会:
1. **提取可测试的交付物** — 列出你现在应该能做到的事情
2. **逐一引导验证** — "能用邮箱登录吗?" 是/否,或描述问题
3. **自动诊断失败** — 派出调试代理查找根因
4. **创建修复计划** — 可直接执行的修复方案
如果一切通过,继续下一阶段。如果有问题,再次运行 `/gsd:execute-phase` 执行修复计划。
这是 GSD 和 Ralph 循环最大的理念差异:**Ralph 是 hands-off 的——启动后放手让它跑;GSD 在每个阶段结束后都有人类验证环节。** Chase AI 指出,Ralph 循环是"去征服型"的——它自行运转不回头;GSD 确保你在每个关键节点都能介入纠偏,避免错误在无人监督下层层叠加。
此外,GSD 还提供了专用调试流程。当验证发现问题时,`/gsd:debug` 会启动一个**隔离的调试子代理**,它有自己的假设-证据-解决工作流,创建独立的调试文档追踪整个调查过程,不污染主上下文。
**产出文件**: `{phase}-UAT.md`
### 循环重复
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
每个阶段都经历完整的**讨论 → 计划 → 执行 → 验证**循环。上下文保持新鲜,质量保持一致。
当所有阶段完成后,`/gsd:complete-milestone` 归档里程碑并标记版本。然后 `/gsd:new-milestone` 开启下一个版本的构建。
## 为什么有效:技术原理
GSD 的可靠性不是偶然的,背后有四个关键技术支撑。
### Context Engineering
Claude Code 在获得正确上下文时非常强大。大多数人不知道如何给它正确的上下文。GSD 替你处理了这个问题。
| 文件 | 作用 |
| ----------------- | ----------------------- |
| `PROJECT.md` | 项目愿景,始终加载 |
| `research/` | 生态知识(技术栈、功能、架构、陷阱) |
| `REQUIREMENTS.md` | 分版本的需求,带阶段追溯 |
| `ROADMAP.md` | 方向和进度 |
| `STATE.md` | 决策、阻碍、位置——跨 session 的记忆 |
| `PLAN.md` | 原子任务 + XML 结构 + 验证步骤 |
| `SUMMARY.md` | 执行记录,提交到历史 |
每个文件都有基于 Claude 质量退化阈值的**大小限制**。保持在限制之下,就能获得一致的高质量输出。主上下文窗口保持在 30-40%,实际工作在子代理的全新 200k 上下文中完成。
Chase AI 对 context rot 有一个直观的解释:**无论上下文窗口多大——Sonnet、Opus、甚至百万 token 的窗口——前半段的 token 都比后半段更有效。** 这不是 bug,这是 LLM 的固有特性。Claude Code 内置的 autocompact 只能部分缓解。GSD 的方案更彻底:每个原子任务都在全新的子代理中执行,确保每一个任务都能获得 Claude 的最佳表现。
TÂCHES 自己的数据印证了这一点:他在 $200/月的 Max 计划上,每月消耗约 $30,000 的 Opus tokens。这听起来很多,但因为每个任务都在新鲜上下文中执行,返工极少,实际效率远高于在一个退化的上下文中反复修修补补。
### XML Prompt Formatting
每个计划都是为 Claude 优化的结构化 XML:
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
精确的指令,不需要猜测,验证内置在每个任务中。
### Multi-Agent Orchestration
每个阶段使用相同的模式:薄编排器派出专门化代理,收集结果,路由到下一步。
| 阶段 | 编排器做什么 | 代理做什么 |
| -- | ---------- | ----------------------- |
| 研究 | 协调、展示发现 | 4 个并行研究员调查技术栈、功能、架构、陷阱 |
| 规划 | 验证、管理迭代 | 规划者创建计划,检查者验证,循环直到通过 |
| 执行 | 分组波次、跟踪进度 | 执行者并行实现,各自拥有全新 200k 上下文 |
| 验证 | 展示结果、路由下一步 | 验证者检查代码库,调试者诊断失败 |
编排器从不做重活。它派出代理、等待、整合结果。结果是:你可以运行一整个阶段——深度研究、多个计划创建和验证、数千行代码并行编写、自动验证——**而你的主上下文窗口保持在 30-40%**。
### Atomic Git Commits
每个任务完成后立即独立提交:
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
好处:`git bisect` 能定位到具体的失败任务,每个任务可独立回滚,清晰的历史记录帮助 Claude 在未来 session 中理解代码演变。
## GSD 的边界
GSD 很强大,但理解它**做不到什么**同样重要。
### GSD 是人类引导的工作流,不是自主代理
GSD 不能持久化运行。每个阶段边界——从 `discuss` 到 `plan` 到 `execute` 到 `verify`——都需要你手动输入命令。你不能说"帮我做个 app"然后去睡觉。
这和 Ralph 的 AFK 模式形成了鲜明对比。Ralph 就是设计来"启动后去睡觉"的——bash 无限循环会持续运行,直到任务完成或失败。GSD 则要求你在每个关键节点都在场:审批路线图、回答讨论问题、触发规划、启动执行、确认验证结果。
TÂCHES 在直播中 4 小时里一直在键入命令:`new-project`、`discuss-phase 1`、`plan-phase 1`、`execute-phase 1`、`verify-work 1`、`discuss-phase 2`……每一次转换都需要他按回车。这不是偶然的——这是有意识的设计选择。
### 一个有意识的设计取舍
Ralph 牺牲了规划能力换取了执行自主性;GSD 牺牲了执行自主性换取了规划质量和人类校验。**这是设计取舍,不是缺陷。**
* **Ralph 的优势**:你可以在睡觉时让它跑完一整个功能。但如果 spec 不够好,它会在错误的方向上一路狂奔。
* **GSD 的优势**:每个阶段结束后你都能纠偏。但你必须全程在场,不能离开。
理想状态是什么?如果 GSD 的讨论、规划、执行、验证能串成一个自动循环——类似 Ralph 的 bash loop,但带有 GSD 的结构化规划和质量验证——那就是两个世界的最佳组合。但目前还没有这样的工具。这也许是下一个值得探索的方向。
## 视频资源
以下视频可以帮助你更直观地理解 GSD 的使用方式和效果。
## 写在最后
GSD 代表了 AI 编程工具演进的一个方向:从"让 AI 写代码"到"让 AI 可靠地交付项目"。
Ralph Wiggum 证明了一个关键洞察——新鲜上下文比累积上下文更有价值。GSD 在这个基础上,加入了项目理解(new-project)、决策捕获(discuss)、结构化规划(plan)、并行执行(execute)和质量验证(verify),形成了一个完整的闭环。
对于独立开发者和小团队来说,GSD 的价值在于它把复杂的工程实践封装成了几个简单命令。你不需要理解子代理编排或 XML 提示工程——你只需要描述你想要什么,然后让系统去搞定。
Chase AI 说得好:GSD 适合那些"不来自技术背景,但仍然想在 Claude Code 中以可持续、可重复的方式端到端构建项目"的人。而 TÂCHES 的直播证明了这一点——一个自称"可能只能自己写一个 Hello World HTML 页面"的人,用 GSD 构建了一个完整的原生桌面应用。
这不是魔法。这是**把正确的复杂性放在正确的地方**——系统承担编排的复杂性,人类专注于创意和决策。而它的边界同样值得尊重:GSD 选择了让人类始终在场,这既是它的限制,也是它可靠性的来源。
想要上手实操?继续阅读 [GSD 实践指南](/docs/notes/gsd/practice)——涵盖完整命令参考、配置详解、实战工作流演示和常见问题。
***
**相关阅读**:
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) — Context Rot 问题和 Ralph 方法论的完整解析
* [规格驱动开发是什么](/docs/notes/speckit/concept) — 从 Vibe Coding 到规格驱动开发的范式升级
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 另一种保持上下文清洁的方式
* [Claude 系统架构全解析](/docs/notes/claude-architecture) — Hooks、Subagent 等组件的整体架构
* [我的 Claude Code 最佳实践](/blog/claude-code-best-practices) — Claude Code 日常使用技巧
# GSD 实践指南
## 引言
[上一篇](/docs/notes/gsd/concept)我们深入了解了 GSD 的核心原理——上下文工程、子代理编排、目标回溯规划和原子提交。这些理念听起来优雅,但从"理解原理"到"实际跑通一个项目"之间还有不少操作细节。
这一篇,我们来动手操作。你将学会 GSD 的完整命令体系、配置选项、产出文件结构,以及如何用它从零交付一个完整功能。
## 安装与配置
### 安装
```bash
npx get-shit-done-cc@latest
```
安装器会提示你选择:
1. **运行时** — Claude Code、OpenCode、Gemini CLI 或全部
2. **位置** — 全局(所有项目)或本地(当前项目)
安装后在运行时中输入 `/gsd:help` 验证安装成功。
### 推荐:跳过权限模式
GSD 设计为无摩擦自动化。推荐以下方式运行 Claude Code:
```bash
claude --dangerously-skip-permissions
```
如果不想用这个标志,可以在 `.claude/settings.json` 中配置细粒度权限。
### 更新
```text
/gsd:update
```
GSD 更新非常频繁(TÂCHES 几乎每天推送 15-20 次更新),建议定期运行此命令保持最新版本。
## 完整命令参考
GSD 的所有交互都通过 `/gsd:` 前缀的斜杠命令完成。以下按功能分类列出完整命令。
### 核心工作流命令
这五个命令构成 GSD 的主循环,按顺序使用。
| 命令 | 说明 |
| ------------------------ | ------------------------------------ |
| `/gsd:new-project` | 初始化项目。系统持续提问直到理解你的想法,然后研究、提取需求、创建路线图 |
| `/gsd:discuss-phase [N]` | 讨论第 N 阶段的灰色地带。捕获你的实现偏好,为规划提供方向 |
| `/gsd:plan-phase [N]` | 为第 N 阶段创建原子任务计划。包含研究、规划、验证三个子步骤 |
| `/gsd:execute-phase ` | 执行第 N 阶段。子代理并行实现任务,每个任务独立提交 |
| `/gsd:verify-work [N]` | 验证第 N 阶段的交付物。引导你逐一确认,自动诊断问题 |
> `[N]` 表示可选参数——省略时系统会自动检测当前阶段。`` 表示必填参数。
### 里程碑管理
| 命令 | 说明 |
| --------------------------- | ------------------------------ |
| `/gsd:audit-milestone` | 审计当前里程碑进度——检查所有阶段状态,识别未完成项 |
| `/gsd:complete-milestone` | 归档当前里程碑,标记版本,准备进入下一个周期 |
| `/gsd:new-milestone [name]` | 创建新里程碑。可选提供名称,系统会基于已完成工作规划下一阶段 |
### 阶段管理
| 命令 | 说明 |
| --------------------------------- | ----------------------- |
| `/gsd:add-phase` | 在路线图末尾添加新阶段 |
| `/gsd:insert-phase [N]` | 在指定位置插入紧急阶段,后续阶段自动重新编号 |
| `/gsd:remove-phase [N]` | 移除指定阶段,级联删除所有相关产出文件 |
| `/gsd:list-phase-assumptions [N]` | 列出指定阶段的所有假设和依赖,帮助识别潜在风险 |
### Quick Mode 与工具
| 命令 | 说明 |
| ---------------------- | ---------------------------------------- |
| `/gsd:quick [--full]` | 快速模式——跳过研究、计划检查和验证,适合小任务。`--full` 启用完整保障 |
| `/gsd:debug [desc]` | 启动隔离的调试子代理。可选描述问题,系统会假设→取证→解决 |
| `/gsd:add-todo [desc]` | 记录想法到待办列表,不修改路线图 |
| `/gsd:check-todos` | 查看当前待办列表 |
| `/gsd:map-codebase` | 分析已有代码库——技术栈、架构、约定、潜在问题 |
### Session 与配置管理
| 命令 | 说明 |
| ------------------ | ----------------------------------- |
| `/gsd:pause-work` | 暂停工作。保存当前状态到 STATE.md,方便下次恢复 |
| `/gsd:resume-work` | 恢复工作。从 STATE.md 读取上次状态,继续上次中断的地方 |
| `/gsd:progress` | 查看项目整体进度——完成阶段数、当前位置、待处理项 |
| `/gsd:help` | 显示所有可用命令和简要说明 |
| `/gsd:settings` | 查看和修改 GSD 配置 |
| `/gsd:set-profile` | 切换模型配置(quality / balanced / budget) |
| `/gsd:update` | 更新 GSD 到最新版本 |
## 配置详解
### 模型配置
GSD 支持三种模型配置,通过 `/gsd:set-profile` 切换:
| 配置 | 规划 | 执行 | 验证 | 适用场景 |
| ------------- | ------ | ------ | ------ | -------------- |
| quality | Opus | Opus | Sonnet | 复杂项目、关键功能、首次使用 |
| balanced (默认) | Opus | Sonnet | Sonnet | 日常开发、多数场景的最佳平衡 |
| budget | Sonnet | Sonnet | Haiku | 简单功能、预算敏感、快速迭代 |
### 核心设置
通过 `/gsd:settings` 查看和修改以下配置:
| 设置 | 默认值 | 说明 |
| ------------------------ | ---------- | ------------------------------------------------- |
| `mode` | `balanced` | 模型配置选择 |
| `depth` | `standard` | 研究深度:`quick`(快速)/ `standard`(标准)/ `deep`(深入) |
| `git.branching_strategy` | `feature` | Git 分支策略:`feature`(按功能分支)/ `phase`(按阶段分支)/ `none` |
### 工作流开关
以下代理可以单独开关,在速度和质量之间取舍:
| 开关 | 默认 | 说明 |
| -------------- | -- | --------------- |
| `research` | 开 | 规划前是否进行自动研究 |
| `plan_check` | 开 | 计划创建后是否自动验证 |
| `verifier` | 开 | 执行后是否自动验证 |
| `auto_advance` | 关 | 阶段完成后是否自动进入下一阶段 |
> 关闭 `research` 和 `plan_check` 可以显著加快速度,但可能降低规划质量。建议在熟悉项目后再考虑关闭。
## 产出文件结构
GSD 的所有状态和产出都保存在 `.planning/` 目录下。理解这个结构有助于调试和手动干预。
### 项目级文件
| 文件 | 作用 | 创建时机 |
| ----------------- | -------------- | ------------------ |
| `PROJECT.md` | 项目愿景和范围 | `new-project` |
| `REQUIREMENTS.md` | 分版本的需求文档,带阶段追溯 | `new-project` |
| `ROADMAP.md` | 阶段规划和进度 | `new-project` |
| `STATE.md` | 当前状态——决策、阻碍、位置 | `new-project`,持续更新 |
### 阶段级文件
每个阶段会产生以下文件(以阶段 1 为例):
| 文件 | 作用 | 创建时机 |
| -------------------- | ---------- | ----------------- |
| `01-CONTEXT.md` | 讨论阶段的决策记录 | `discuss-phase 1` |
| `01-RESEARCH.md` | 研究发现和技术调查 | `plan-phase 1` |
| `01-01-PLAN.md` | 第一个原子任务计划 | `plan-phase 1` |
| `01-02-PLAN.md` | 第二个原子任务计划 | `plan-phase 1` |
| `01-01-SUMMARY.md` | 第一个计划的执行记录 | `execute-phase 1` |
| `01-02-SUMMARY.md` | 第二个计划的执行记录 | `execute-phase 1` |
| `01-VERIFICATION.md` | 自动验证结果 | `execute-phase 1` |
| `01-UAT.md` | 用户验收测试记录 | `verify-work 1` |
### 目录结构示例
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## 实战工作流演示
以下以"为博客系统添加评论功能"为例,演示从初始化到交付的完整流程。
### Step 1: 初始化项目
```text
/gsd:new-project
```
系统会开始持续提问:
```
> 你想构建什么?
"我想为我的 Next.js 博客添加评论功能。支持匿名和登录评论、
Markdown 渲染、管理后台。技术栈用 Prisma + PostgreSQL。"
```
提供的描述越详细,系统追问越少。TÂCHES 的建议是:准备一份粗略的愿景文档,描述你想要什么——不需要知道技术细节。
系统完成后会产出四个文件,并要求你审批路线图。审批后进入构建阶段。
> **已有代码库?** 先运行 `/gsd:map-codebase`,系统会分析你的现有架构和约定,之后 `new-project` 就能基于已有代码进行规划。
### Step 2: 讨论阶段
```text
/gsd:discuss-phase 1
```
系统会识别灰色地带并逐一提问:
```
> 评论嵌套层级:支持多级嵌套还是只支持一级回复?
> 匿名评论:需要验证码还是直接提交?
> 管理后台:需要批量操作还是逐条审核?
```
你在这里的每个决定都直接影响后续规划质量。如果不确定,可以让系统用默认值——但深入讨论能显著减少执行阶段的返工。
### Step 3: 计划阶段
```text
/gsd:plan-phase 1
```
系统会:
1. 研究如何用 Prisma + PostgreSQL 实现评论系统
2. 创建 2-3 个原子任务计划(如:数据模型、API 路由、前端组件)
3. 自动验证计划是否覆盖所有需求
每个计划都足够小,能在一个全新的上下文窗口中完成。
### Step 4: 执行阶段
```text
/gsd:execute-phase 1
```
系统开始波次执行:
* **Wave 1**(无依赖):数据库 schema、Prisma 模型——并行执行
* **Wave 2**(依赖 Wave 1):API 路由、评论 CRUD——并行执行
* **Wave 3**(依赖 Wave 2):前端评论组件——独立执行
每个任务在全新的 200k tokens 上下文中执行,完成后独立 git commit。
### Step 5: 验证阶段
```text
/gsd:verify-work 1
```
系统引导你逐一确认:
```
> ✅ 数据库表已创建
> ✅ API 路由返回正确状态码
> ❓ 能在博客文章下方看到评论输入框吗? [是/否/描述问题]
> ❓ 提交评论后页面是否实时更新? [是/否/描述问题]
```
如果有失败项,系统会自动诊断并创建修复计划,再次运行 `/gsd:execute-phase 1` 即可执行修复。
### 常见操作场景
**插入紧急阶段**:需求变更,需要在当前阶段之前插入新工作。
```text
/gsd:insert-phase 2
```
后续阶段自动重新编号(原来的 Phase 2 变成 Phase 3,依此类推)。
**暂停与恢复**:需要中断工作去处理其他事情。
```text
/gsd:pause-work # 保存当前状态
# ... 处理其他事情 ...
/gsd:resume-work # 恢复到上次中断的位置
```
**回滚不满意的结果**:
```bash
git reset --hard HEAD~3 # 回到执行前的状态
```
```text
/gsd:remove-phase 2 # 级联删除该阶段的所有产出文件
```
TÂCHES 在直播中多次演示这个操作——不喜欢就回滚,干脆利落。
## 调试工作流
当验证发现问题,或者你在开发过程中遇到 bug 时,GSD 提供了专用调试流程。
```text
/gsd:debug 评论提交后页面没有实时更新
```
系统会启动一个**隔离的调试子代理**,它的工作流程是:
1. **假设** — 基于问题描述生成多个可能的根因假设
2. **取证** — 逐一验证假设,检查代码、日志、网络请求
3. **解决** — 定位根因后创建修复方案
关键特性:
* **上下文隔离**:调试代理有自己的上下文窗口,不会污染主开发上下文
* **文档追踪**:创建独立的调试文档记录整个调查过程
* **修复计划**:诊断完成后产出可直接执行的修复计划
这比直接在主上下文中调试要高效得多——调试信息不会累积在你的主窗口中。
## 实战经验
综合 TÂCHES 的直播和 Chase AI 的使用体验,以下是一些实战建议。
### 放慢才能加快
TÂCHES 坦言自己早期使用 GSD 时也是"快快快"的心态,但后来发现**在研究和讨论阶段多花时间,执行阶段的返工反而更少**。新版 GSD 增加了 `research-project` 和 `define-requirements` 步骤,就是为了在动手前把方向搞对。
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### 阶段之间清上下文
TÂCHES 的习惯是**每个阶段之间运行 `clear`**,保持主上下文精简。他使用 Warp 终端,每个窗口全屏(Command+Shift+Enter),在一个窗口执行当前阶段的同时,在另一个窗口研究下一个阶段。
### Token 成本的权衡
GSD 的子代理方式确实比直接使用 Claude Code 消耗更多 token。但 Chase AI 提出了一个有力的论点:**"plan twice, prompt once"(计划两次,提示一次)比"提示一次,然后修修补补"长远来看更省 token。** 因为在新鲜上下文中做对一次,比在退化上下文中反复修复要高效得多。
### 不如意时的处理
如果你对某个阶段的结果不满意,可以 `git reset --hard` 然后用 `/gsd:remove-phase` 级联删除该阶段的所有产出文件。TÂCHES 在直播中实际演示了这个操作——他不喜欢某个视觉效果,直接回滚到上一个满意的状态,干脆利落。
### To-Do 系统
`/gsd:add-todo` 可以随时记录想法到待办列表,而不需要修改路线图。这些想法可以在 `/gsd:discuss-milestone` 时被拉出来,作为下一个里程碑的输入。TÂCHES 的策略是"先做功能,milestone 2 再打磨 UI"。
## 常见问题与最佳实践
### 最佳实践
**提供详细的初始描述**。`/gsd:new-project` 的质量取决于你的输入质量。准备一份粗略的愿景文档——描述目标、用户、核心功能、已知约束。描述越精确,系统追问越少,规划越准。
**阶段之间清理上下文**。每完成一个阶段,运行 `clear` 或 `/compact` 保持主上下文窗口精简。TÂCHES 的习惯是保持主上下文在 30-40%。
**先用 Quick Mode 测试**。对于不确定的小功能,先用 `/gsd:quick` 试水。如果效果好,再在正式路线图中纳入。
**对已有项目先 map-codebase**。在已有代码库上使用 GSD 前,先运行 `/gsd:map-codebase`。系统会分析技术栈、架构和约定,之后的规划会更贴合现有代码。
### FAQ
**Q: GSD 支持哪些运行时?**
A: Claude Code、OpenCode 和 Gemini CLI。安装时可以选择单个或全部。
**Q: Quick Mode 和完整模式有什么区别?**
A: Quick Mode 提供 GSD 的基础保障(原子提交、状态跟踪),但跳过研究、计划检查和验证步骤。适合 bug 修复、小功能、配置变更等不需要完整规划的任务。
**Q: 执行过程中可以暂停吗?**
A: 可以。`/gsd:pause-work` 会将当前状态保存到 STATE.md。下次 `/gsd:resume-work` 时,系统会从上次中断的位置继续。
**Q: 如何控制 token 成本?**
A: 三种方式——(1) 切换到 `budget` 配置:`/gsd:set-profile budget`;(2) 关闭 `research` 或 `plan_check` 代理;(3) 对简单任务使用 `/gsd:quick`。
**Q: GSD 可以和 Ralph 一起用吗?**
A: 可以。GSD 和 Ralph 解决不同的问题——GSD 负责规划和结构化执行,Ralph 负责自主循环执行。你可以用 GSD 的 `new-project` 和 `plan-phase` 生成完整规划,然后用 Ralph 循环来执行那些不需要人类介入的阶段。
**Q: 多人协作怎么办?**
A: `.planning/` 目录可以提交到 Git。多人可以各自运行不同阶段,通过 Git 合并结果。但建议避免同时运行同一阶段。
## 小结
GSD 的核心价值在于**把复杂性藏在系统里,把简单性留给用户**。你只需要几个命令——`new-project`、`discuss-phase`、`plan-phase`、`execute-phase`、`verify-work`——系统在背后处理所有的上下文管理、子代理编排和质量验证。
从安装到交付,GSD 提供了一条清晰的路径:描述你想要什么 → 讨论实现细节 → 生成原子计划 → 并行执行 → 验证交付物。每一步都有你介入的机会,每一步都有文件记录。
这不是"按一个按钮就搞定"的魔法。这是一个需要你参与但帮你承担大部分认知负担的系统。正如 TÂCHES 所说:你是高层项目经理,GSD 是你的执行团队。
***
**延伸阅读**:
* [GSD 深度解析](/docs/notes/gsd/concept) — 核心原理、工作流与技术架构
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) — Context Rot 和 Ralph 方法论
* [snarktank/ralph 实战指南](/docs/notes/ralph-wiggum/snarktank) — Ralph 的安装、PRD 编写与实战
* [规格驱动开发是什么](/docs/notes/speckit/concept) — 从 Vibe Coding 到规格驱动开发
* [Speckit 实践指南](/docs/notes/speckit/practice) — Speckit 命令详解与完整案例
# gstack:当 YC CEO 把创业经验装进 Claude Code
## 核心结论
**gstack 是 Garry Tan 开源的一套 Claude Code 角色化技能集**:它把一个通用 AI 编程助手拆成 CEO、工程经理、设计师、QA、安全官、发布工程师等多个角色,让 AI 在不同阶段用不同视角审查产品、架构、代码、测试和发布。
它要解决的问题不是「怎么让 AI 写更多代码」,而是「怎么让 AI 编程从随机应变变成可靠交付」。Ralph Wiggum 依靠重启新进程避免上下文腐烂,GSD 依靠规格驱动管理复杂项目,而 gstack 的答案是:**把开发流程拆给一支虚拟工程团队,由人来指挥角色切换。**
## 什么是 gstack
gstack 的创建者 **Garry Tan** 有着丰富的技术和创业背景——14 岁开始写代码,斯坦福计算机工程出身,Palantir 第 10 号员工,联合创办过 Posterous(后被 Twitter 收购),2023 年起担任 Y Combinator 的 President & CEO。
他用 gstack 在 60 天内发布了超过 60 万行生产代码(35% 是测试),日均万行以上——同时还在全职运营 YC。其中一个项目 garylist.org,21 天上线、15 万行代码,35% 测试覆盖。按他自己的说法,代码质量超过了他之前花 500 万美元、两年时间、10 个工程师做出来的创业项目。
项目自 2026 年 3 月 11 日开源以来,3 周内从 v0 迭代到 v0.15.1.0,GitHub 已获得 60,500+ stars。MIT 许可证,完全开源。
## gstack 在工具生态中的位置
| 维度 | 原生 Claude Code | Ralph Wiggum | GSD | SpecKit | Superpowers | **gstack** |
| ---- | -------------- | --------------- | ------------------- | ------------------- | ----------- | ------------------ |
| 核心定位 | 通用 AI 编码助手 | 无限循环迭代 | 上下文工程 + 规格驱动 | 需求→规格→任务 | 流程纪律 + TDD | **角色化虚拟团队** |
| 核心模式 | 对话式编程 | Bash 循环 + 新进程 | Phase-based Roadmap | Spec → Plan → Tasks | 严格开发流水线 | **Sprint 七步流程** |
| 人类参与 | 实时对话 | Hands-off (AFK) | 每阶段验证 | 规格审批 | 每步确认 | **每阶段角色审查** |
| 独特能力 | 基础编码 | 无限迭代 | Context Rot 管理 | 需求追溯 | 强制 TDD | **浏览器自动化 + 多角色审查** |
| 适合场景 | 简单任务 | 持续迭代 | 大型项目管理 | 需求严谨的项目 | 工程质量保障 | **全流程产品开发** |
从表中可以看出一个关键规律:**这些工具并非互相竞争,而是在不同维度上解决 AI 编程的问题。**
Superpowers 用**流程纪律**保证代码质量(强制 TDD、结构化对话、实施计划);GSD 用**上下文工程**管理复杂项目(阶段规划、子代理新鲜上下文、文件系统状态);gstack 用**角色分解**提升决策质量(CEO 视角审产品、工程经理审架构、QA 跑真实浏览器)。
简单来说,Superpowers 基于流程护栏,gstack 基于角色设计——前者适合从 1 到 N 的工程落地,后者适合从 0 到 1 的产品构建。**两者互补而非竞品。**
## 核心工作流:The Sprint 七步走
gstack 将整个开发过程组织为一个 **Think → Plan → Build → Review → Test → Ship → Reflect** 的循环,叫做"The Sprint"——不是敏捷 Sprint,而是一种"角色依次登场"的开发节奏。
### 1. Think — 产品门诊
```text
/office-hours
```
这是 gstack 最有特色的 skill。灵感直接来自 YC 的 Office Hours——创业者去见 YC 合伙人,接受灵魂拷问。AI 会问你 **6 个逼迫性问题**:
1. 谁具体需要这个?
2. 他们今天没有它怎么办?
3. 为什么这件事现在很紧迫?
4. 你怎么知道它能用?
5. 如果什么都不做会怎样?
6. 你能发布的最小版本是什么?
目的不是帮你写代码,而是在写代码之前**重新审视问题本身**。
### 2. Plan — 多角色审查
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
CEO Review 本质上是"Founder Mode"——不是按字面意思执行需求,而是退后一步问"这个产品真正的目的是什么?"它支持四种模式:扩大范围、选择性扩展、保持范围、缩小范围。
### 3. Build — 编码实现
按审查通过的计划开始编码。这一步使用标准 Claude Code 能力。
### 4. Review — 平行专家审查
```text
/review
```
这个 skill 一次性派出 **7 个并行子代理**,分别从测试、可维护性、安全、性能、数据迁移、API 合约、红队攻击 7 个角度审查代码。遇到明显问题会自动修复。
### 5. Test — 真实浏览器 QA
```text
/qa
```
不是模拟测试。QA skill 启动一个**真实的 headless Chromium 浏览器**,打开你的应用、点击按钮、填表单、截图——和真人测试员做的一样。发现 bug 后自动修复、生成回归测试、重新验证。
### 6. Ship — 一键发布
```text
/ship
```
自动同步主分支、运行测试、审查 diff、更新版本号和 CHANGELOG、提交、推送、创建 PR。如果项目没有测试框架,它甚至会先搭建一个。
### 7. Reflect — 回顾与学习
```text
/retro
```
工程经理风格的周报:分析提交历史、测试比例、代码质量趋势。支持多人团队分析,跟踪"连续发布天数"等指标。
## 为什么有效:技术原理
### Browse Daemon:给 AI 装上眼睛
gstack 最独特的技术贡献是 **Browse Daemon**——一个长驻的 headless Chromium 实例,通过 localhost HTTP 通信。第一次调用启动浏览器(约 3 秒),之后每次命令只需 100-200ms。这意味着 AI 可以**真正看到**你的应用,而不是猜测 DOM 结构。
它还引入了 **Ref System**(元素引用 `@e1`, `@e2`),通过 accessibility tree 定位元素,不需要写 CSS 选择器。这是被社区(包括批评者)普遍认可的"真正有技术含量的贡献"。
### 角色分解:不是一个 agent,而是一支团队
gstack 的做法是把所有角色拆解成独立的 prompt 文件,让 Claude Code 在不同阶段切换到不同角色的视角来审视代码。这本质上是一种精细化的 prompt engineering。
核心洞察是:**规划不等于审查,审查不等于发布,创始人品味和工程严谨是完全不同的思维模式。** 与其让一个通用 agent 做所有事,不如在需要时切换"大脑模式"——founder thinking、engineering rigor、paranoid review、fast execution。
### 三大哲学
gstack 的 ETHOS.md 记录了三个核心理念:
1. **Boil the Lake(煮沸整个湖)**:当 AI 让完整性的边际成本趋近零时,永远选择完整实现——100% 测试覆盖、所有边界情况、所有错误路径。"发布捷径"是旧时代的思维。
2. **Search Before Building(先搜索再构建)**:三层知识——久经考验的模式、新且流行的方案、第一性原理。先理解所有人在做什么,质疑他们的假设,然后发现为什么常规方案是错的。
3. **User Sovereignty(用户主权)**:AI 推荐,人类决定。即使两个 AI 模型达成共识,用户的判断仍然优先——因为用户有领域知识、战略视角和品味。
## gstack 的边界与争议
gstack 的社区反应可能是 AI 编程工具中**最两极化**的。
**看好的一面**:创始人和非技术构建者普遍认可,尤其是 `/office-hours` 和 `/plan-ceo-review` 这类"产品思维"类 skill,帮助很多独立开发者在动手编码之前重新审视了产品方向。工程审查(`/review`)也确实能发现一些隐蔽的安全漏洞,这种多角度并行审查的模式有实际价值。
**质疑的一面**也很直接:
* **LOC 指标意义不大**:60 天 60 万行代码,代码行数从来不是质量指标,大量代码可能只是脚手架和样板。
* **本质是 prompt 模板**:每个 skill 就是一个 SKILL.md 文件,技术门槛并不高。真正的价值不在文件本身,而在 prompt 的设计质量。
* **AI 自审代码的局限性**:`/review` 让 AI 审查 AI 写的代码,相当于自己批改自己的作业。多角色并行能缓解这个问题,但根本上还是同一个模型。
* **名人效应的加成**:如果创建者不是 YC CEO,这个项目大概率不会获得这么高的关注度。
**我的看法**:抛开争议不谈,gstack 真正有价值的部分是两个——Browse Daemon 的浏览器自动化技术,和角色分解的设计模式。这些不依赖于 Garry Tan 是谁。角色化的核心意义其实不在技术层面,而在行为层面——它帮助你更有意识地组织 AI 工作流,而不是一股脑把所有事丢给一个通用 agent。
gstack 适合用来 fork 和定制,取你需要的 skill、改你想改的 prompt,而不是全盘照搬。
## 视频资源
## 写在最后
gstack 代表了 AI 编程工具的一个有趣方向:不是让 AI 更自主(Ralph 的路线),也不是让流程更严格(Superpowers 的路线),而是让 AI **扮演不同角色**来提升决策质量。它的争议恰恰说明了 AI 编程生态的丰富性——没有一个方案适合所有人。
如果你对 gstack 感兴趣,下一步可以看 [实战篇](/docs/notes/gstack/practice)——从安装到跑通完整工作流的手把手教程。
***
**相关阅读**:
* [GSD 概念介绍](/docs/notes/gsd/concept) — 另一种结构化 AI 编程方案
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) — 了解无限循环迭代的起点
* [Claude Skills 概念篇](/docs/notes/claude-skills/concept) — 理解 Skills 的底层机制
# gstack 前端 Skill 全景:从设计到上线的 AI 工作流
## 引言
在之前的笔记中,我们聊了 [gstack 是什么](/docs/notes/gstack/concept)、[怎么跑通工作流](/docs/notes/gstack/practice)、以及 [Skill 的工程架构](/docs/notes/gstack/skill-architecture)。但有一个问题一直没展开——gstack 安装后带来的 60+ 个 Skill 里,**跟前端/UI 设计相关的有哪些?按什么顺序用?**
这篇笔记做两件事:先把前端相关的 \~27 个 Skill 按功能分类,然后用一个有趣的小项目——倒计时纪念日页面——从头到尾走一遍完整工作流,每个阶段配截图,让你看到实际效果。
## 前端 Skill 套件全景
gstack 的前端 Skill 可以分成 6 个功能层,从地基到屋顶,每一层解决不同阶段的问题。
### 设计基建
项目级的一次性设置,确定设计语言,后续所有 Skill 都会参照这些基准。
| Skill | 做什么 | 什么时候用 |
| ---------------------- | ------------------------- | -------------------------- |
| `/design-consultation` | 完整的设计系统咨询,产出配色、字体、间距、质感方向 | 新项目启动时,或者想重新定义视觉风格 |
| `/teach-impeccable` | 一次性采集设计偏好,写入 AI 配置文件 | 装完 gstack 后跑一次,让 AI 记住你的审美 |
| `/brand-guidelines` | 应用已有的品牌配色和字体规范 | 有现成品牌手册时直接套用 |
> 如果项目已经有 `DESIGN.md`,这一层可以跳过。
### 设计探索
不确定方向时,快速出多个方案对比选择。
| Skill | 做什么 | 什么时候用 |
| ------------------ | ------------------- | --------------- |
| `/design-shotgun` | 生成 3-5 个视觉方案,打开对比面板 | 不确定要什么风格,想看看可能性 |
| `/frontend-design` | 生成有辨识度的、生产级前端界面代码 | 方向明确后直接出活 |
| `/canvas-design` | 生成海报、视觉艺术(PNG/PDF) | 需要静态视觉设计而非网页组件 |
### 设计实现
把方案变成真正能跑的代码,处理排版、布局、响应式。
| Skill | 做什么 | 什么时候用 |
| ------------------------ | --------------------- | --------------- |
| `/design-html` | 把确认的设计稿转成生产级 HTML/CSS | 有 mockup 想直接落地 |
| `/mobile-responsiveness` | 移动端优先的响应式布局、触摸交互 | 从零开始做移动端适配 |
| `/adapt` | 跨设备、跨屏幕尺寸的断点适配 | 已有桌面版,需要适配手机/平板 |
| `/typeset` | 字体选择、层级、大小、粗细、可读性优化 | 文字排版看着"差点意思" |
| `/arrange` | 布局间距、视觉节奏、对齐修复 | 间距不一致、布局感觉挤或散 |
### 设计增强
在功能完成的基础上,注入动效、个性和情感化细节。
| Skill | 做什么 | 什么时候用 |
| ------------ | --------------------------- | ----------------- |
| `/animate` | 添加有目的的微交互和动效 | 页面功能 OK 但感觉"死板" |
| `/delight` | 添加惊喜细节、个性化触感 | 想让用户记住这个页面 |
| `/bolder` | 放大视觉冲击力 | 设计太平淡、太安全 |
| `/colorize` | 给单调的界面加入策略性色彩 | 页面太灰、太素、缺乏温度 |
| `/overdrive` | 技术炸裂级效果——shader、弹簧物理、滚动驱动动画 | 某个区域想要 wow effect |
| `/onboard` | 新用户引导流程、空状态设计 | 做首次使用体验 |
这四个增强 Skill 是**递进关系**:`animate` 是基础动效,`delight` 是情感化,`bolder` 是放大,`overdrive` 是炸裂。根据项目需要逐级叠加,不必全用。
### 设计优化
收敛和精炼——去掉多余的,对齐偏离的,打磨粗糙的。
| Skill | 做什么 | 什么时候用 |
| ------------ | --------------------- | ------------------- |
| `/polish` | 最终质量打磨:对齐、间距、一致性 | 发布前的最后一遍检查 |
| `/quieter` | 降低视觉刺激强度 | 设计太花哨、太吵 |
| `/distill` | 极简化,去除不必要的复杂度 | 页面元素太多想做减法 |
| `/normalize` | 对齐设计系统标准(token、间距、颜色) | 样式偏离了 DESIGN.md 的规范 |
| `/clarify` | 改善 UX 文案、错误提示、标签措辞 | 文案让人困惑、错误信息不友好 |
### 设计审查与验证
上线前的系统性检查,找问题、打分、修复。
| Skill | 做什么 | 什么时候用 |
| --------------------- | ---------------------- | ---------------- |
| `/plan-design-review` | 实现前的设计方案评审(0-10 打分) | 想让 AI 以设计师视角审视方案 |
| `/design-review` | 实现后的视觉 QA,自动截图对比并修复 | 代码写完了,检查视觉还原度 |
| `/critique` | UX 评估:视觉层级、认知负荷、情感共鸣 | 想要一份结构化的设计评审报告 |
| `/audit` | 技术审查:可访问性、性能、主题、响应式 | 上线前的系统性检查 |
| `/benchmark` | 性能基线测试,before/after 对比 | 想量化改动对性能的影响 |
## 实战演示:用倒计时纪念日页面走通全流程
光看分类表太抽象。我们用一个小项目把上面的 Skill 串起来——做一个**倒计时/纪念日单页**:选一个有意义的日期,做一个带数字动画和背景效果的倒计时展示。
这个项目小而完整,刚好能覆盖 6 层 Skill 中的大部分。完整流程是:
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> 不是每次都要跑全部 7 步。熟练后常用链路只有 `/frontend-design → /animate → /polish → /ship` 四步。这里为了展示完整能力,每步都走一遍。
### Stage 1:基建——确定设计语言
**Skill**:`/design-consultation` + `/teach-impeccable`
只需要在项目初期做一次。产出 `DESIGN.md`,让 AI 记住你的设计偏好。如果项目已有 `DESIGN.md`,直接跳过。
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* TODO: 截图 — design-consultation 产出的 DESIGN.md 片段 */}
### Stage 2:探索——多方案对比
**Skill**:`/design-shotgun`
不确定方向时,让 AI 生成 3-5 个视觉方案,打开对比面板挑选。
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* TODO: 截图 — design-shotgun 生成的 3 个方案对比面板 */}
从 3 个方案中选一个方向。如果你很清楚要什么,跳过这步直接进 Stage 3。
### Stage 3:构建——出生产级代码
**Skill**:`/frontend-design` + `/adapt`
核心环节。出代码,同时确保响应式从一开始就到位。
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* TODO: 截图 — 构建完成后的桌面端页面效果 */}
{/* TODO: 截图 — 手机端效果(/adapt 适配后) */}
### Stage 4:增强——注入动效和个性
**Skill**:`/animate` → `/delight`(按需 `/overdrive`)
这三个是递进关系:`animate` 是基础动效,`delight` 是情感化细节,`overdrive` 是炸裂级效果。根据需要逐级叠加。
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
注意参照 `DESIGN.md` 的动效约束——如果设计系统只允许 150ms 的 hover 过渡,`/overdrive` 就不适用。这是一个很好的判断练习。
{/* TODO: 截图或 GIF — 动效增强前 vs 后的对比 */}
### Stage 5:调优——收敛和打磨
**Skill**:`/typeset` + `/polish`(按需 `/distill`、`/normalize`)
间距对齐、字体层级、视觉节奏。如果发现加了太多东西,用 `/distill` 做减法。
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* TODO: 截图 — polish 前 vs 后的细节对比 */}
### Stage 6:审查——系统性检查
**Skill**:`/design-review` + `/audit`
视觉 QA + 技术审查。`/design-review` 会自动截图对比并修复问题,`/audit` 检查可访问性和性能。
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* TODO: 截图 — audit 产出的评分报告 */}
### Stage 7:发布
**Skill**:`/ship`
标准的 gstack 发布流程——测试、diff review、创建 PR。
```text
> /ship
```
***
**预期成果**:一个视觉精致的倒计时页面,过程中用到 8-10 个前端 Skill。更重要的是建立起"什么阶段用什么 Skill"的直觉。
## 日常速查表
上面是完整流程。日常开发中遇到具体问题,直接查这张表:
| 我现在的问题 | 用什么 |
| --------------- | -------------------------------- |
| 不知道想要什么风格 | `/design-shotgun` |
| 页面功能好了但感觉"差点意思" | `/polish` |
| 觉得哪里不对但说不清 | `/design-review` |
| 字体/排版看着别扭 | `/typeset` |
| 间距乱、布局挤 | `/arrange` |
| 样式偏离了设计系统 | `/normalize` |
| 想提取公共组件 | `/extract` |
| 页面太复杂想做减法 | `/distill` |
| 设计太素太安全 | `/bolder` 或 `/colorize` |
| 设计太花哨太吵 | `/quieter` |
| 错误提示文案不友好 | `/clarify` |
| 手机端显示有问题 | `/adapt` |
| 想加动画效果 | `/animate`(基础)或 `/overdrive`(炸裂) |
| 上线前系统性检查 | `/audit` |
| debug | `/investigate` |
## 小结
这篇笔记做了两件事:
1. **全景图**——gstack 的 27 个前端 Skill 按 6 层分类(基建→探索→实现→增强→优化→审查)
2. **实战演示**——用一个倒计时纪念日页面走通完整工作流,展示每个阶段用什么 Skill、为什么
关键 takeaway:这些 Skill 最强大的用法不是单独调用,而是**按流水线组合**——探索方向、构建实现、增强打磨、审查发布,每个阶段有明确的 Skill 选择。
但也别被流程绑住——熟练之后,大部分时候 `/frontend-design → /animate → /polish → /ship` 四步就够了。
***
**相关阅读**:
* [gstack 概念篇](/docs/notes/gstack/concept) — gstack 是什么、解决什么问题
* [gstack 实战篇](/docs/notes/gstack/practice) — 从安装到跑通完整工作流
* [gstack Skill 架构拆解](/docs/notes/gstack/skill-architecture) — Skill 开发者能学到什么
* [Claude Skills 概念篇](/docs/notes/claude-skills/concept) — 理解 Skills 的底层机制
# gstack 实战:从安装到跑通完整工作流
## 核心结论
**gstack 是一套把 Claude Code 拆成虚拟工程团队的角色化 skill 集**。如果你只是想跑通它,最短路径是:全局安装 gstack,先用 `/office-hours` 或 `/autoplan` 把需求和计划审一遍,再用 `/review`、`/qa`、`/ship` 完成代码审查、浏览器验证和发布。
这篇实战篇不再解释「gstack 是什么」——这个在 [概念篇](/docs/notes/gstack/concept) 里已经讲过。这里回答更具体的问题:**gstack 怎么装、先用哪几个命令、怎样把它串成一个完整工作流。**
| 目标 | 优先使用的命令 | 作用 |
| ---------- | ----------------- | ------------------- |
| 想清楚产品方向 | `/office-hours` | 用 6 个逼迫性问题重构想法 |
| 生成审查后的计划 | `/autoplan` | 自动跑 CEO、设计、工程等多角色审查 |
| 找出代码风险 | `/review` | 从生产事故视角审查 diff |
| 验证页面是否真的可用 | `/qa` 或 `/browse` | 用真实浏览器测试页面、截图和交互 |
| 准备发布 | `/ship` | 跑测试、更新文档、提交并创建 PR |
## 安装与配置
### 前置条件
* **Claude Code** 已安装并可用
* **Git** 已安装
* **Bun v1.0+** 已安装(gstack 基于 Bun 构建)
* Windows 用户还需要 Node.js
### 全局安装(推荐,30 秒完成)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
安装脚本会做三件事:
1. 将 gstack 的 skill 信息添加到你的 `CLAUDE.md` 文件
2. 将所有 skill 文件放入 skills 目录
3. 安装 Playwright 和对应的 Chromium 浏览器(用于 `/browse` 和 `/qa`)
### 项目级安装(团队共享)
如果希望团队成员克隆仓库后自动获得 gstack:
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
### 多 Agent 支持
gstack 不限于 Claude Code,目前已支持 **10 个 AI 编程 Agent**,`./setup` 默认自动检测已安装的 host:
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
每个 host 的 skill 安装路径形如 `~/./skills/gstack-*/`,互不干扰。
> 💡 **OpenClaw 用户额外选择**:除了通过 ACP 调用,OpenClaw 还能通过 ClawHub 直接安装 4 个原生方法论 skill(`gstack-openclaw-office-hours`、`gstack-openclaw-ceo-review`、`gstack-openclaw-investigate`、`gstack-openclaw-retro`),无需 Claude Code 会话即可对话使用。
### Team Mode(团队共享 + 自动更新,推荐)
v1.x 引入 Team Mode:每个开发者全局安装 gstack,仓库只记录"我们用 gstack"这件事,更新自动发生:
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
把 `required` 换成 `optional` 则是"温柔提醒"而非强制。每次启动 Claude Code 会自动跑一次更新检查(节流 1 次/小时,网络失败安全静默),仓库里没有 vendored 文件,也没有版本漂移。
### 更新
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
或者在 Claude Code 中直接使用 `/gstack-upgrade`。
## 完整命令参考
### Sprint 流程
| 命令 | 角色 | 说明 |
| --------------------- | --------------- | ------------------------------------------------------------------------ |
| `/office-hours` | YC Office Hours | 6 个逼迫性问题重构产品方向,生成设计文档 |
| `/plan-ceo-review` | CEO / 创始人 | 寻找 10 星级产品,四种范围模式可选 |
| `/plan-eng-review` | 工程经理 | 锁定架构、数据流、边界情况、测试矩阵 |
| `/plan-design-review` | 资深设计师 | 设计维度 0-10 评分,说明如何做到 10 分 |
| `/plan-devex-review` | 开发者体验负责人 | 探索开发者画像、对标 TTHW、设计魔法时刻;三种模式(DX EXPANSION / POLISH / TRIAGE),20-45 个逼迫性问题 |
| `/autoplan` | 审查流水线 | 自动依次运行 CEO → 设计 → 工程 → DX 审查,按编码决策原则自动决议,仅把"品味决策"上抛给你 |
### 设计
| 命令 | 说明 |
| ---------------------- | --------------------------------------- |
| `/design-consultation` | 从头构建完整设计系统,生成 DESIGN.md |
| `/design-shotgun` | 生成多个 AI 设计变体,在浏览器中对比选择 |
| `/design-html` | 生成生产级 HTML/CSS,支持 React/Svelte/Vue 框架检测 |
### 审查与安全
| 命令 | 角色 | 说明 |
| ---------------- | --------- | ------------------------------------------------------------------- |
| `/review` | Staff 工程师 | 找出能过 CI 但会在生产爆炸的 bug,明显问题自动修复,标记完整性缺口 |
| `/investigate` | 调试专家 | 系统化根因调试。铁律:不找到根因不修 bug;3 次失败修复后强制停下 |
| `/design-review` | 会写代码的设计师 | 视觉审计 + 自动修复,原子提交,前后对比截图 |
| `/devex-review` | DX 测试员 | 真实跑一遍 onboarding:浏览文档、跑入门流程、计时 TTHW、截图错误,对照 `/plan-devex-review` 评分 |
| `/cso` | 安全官 | OWASP Top 10 + STRIDE 威胁建模,17 条误报排除规则,8/10 置信度门槛,每条发现附具体利用场景 |
### 测试与 QA
| 命令 | 说明 |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/qa` | 打开真实浏览器测试,发现 bug → 原子提交修复 → 生成回归测试 → 重新验证 |
| `/qa-only` | 同上但仅报告,不修改代码 |
| `/benchmark` | 基线性能测试:页面加载、Core Web Vitals、资源大小,支持前后对比 |
| `/browse` | \~100ms 级别的浏览器命令,真实 Chromium,截图、表单填写、元素点击 |
| `/open-gstack-browser` | 启动 GStack Browser:可见的 AI 控制 Chromium,自带 sidebar 扩展、反爬 stealth、自动模型路由(Sonnet 操作 / Opus 分析),支持一键 cookie 导入 |
| `/setup-browser-cookies` | 从真实浏览器(Chrome / Arc / Brave / Edge)导入 cookie 到 headless 会话,测试需登录的页面 |
| `/pair-agent` | 跨 AI Agent 浏览器配对:把同一个 GStack Browser 共享给 OpenClaw / Hermes / Codex / Cursor 等,每个 Agent 独立 tab,自带 ngrok 隧道支持远程 Agent,作用域 token + tab 隔离 + 速率限制 + 行为归因 |
### 发布与运维
| 命令 | 说明 |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------ |
| `/ship` | 同步主分支 → 跑测试 → 审计覆盖率 → 更新版本 → 提交推送 → 创建 PR;项目无测试框架时自动 bootstrap |
| `/land-and-deploy` | 合并 PR → 等待 CI → 部署 → 验证生产环境健康 |
| `/canary` | 部署后金丝雀监控:控制台错误、性能回归、页面故障 |
| `/setup-deploy` | `/land-and-deploy` 一次性配置:自动检测平台(Fly.io / Render / Vercel / Netlify / Heroku / GitHub Actions / 自定义)+ 生产 URL + 部署命令 |
| `/setup-gbrain` | GBrain 数据库一键上手(5 分钟内):PGLite 本地、Supabase 现有 URL,或通过 Management API 自动创建新 Supabase 项目;MCP 注册 + 仓库级 read-write/read-only/deny 权限 |
### 回顾与学习
| 命令 | 说明 |
| ---------------------------------- | ----------------------------------------------------------------------------------------- |
| `/retro` | 团队感知周报:人均拆解、连胜统计、测试健康趋势、成长机会;`/retro global` 跨所有项目 + AI 工具(Claude Code / Codex / Gemini) |
| `/document-release` | 自动更新项目文档匹配已发布的代码(README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS);`/ship` 现已自动调用 |
| `/learn` | 管理跨会话学习记忆:查看、搜索、修剪、导出,按项目积累 |
| `/context-save` `/context-restore` | Continuous checkpoint 模式配套:自动 WIP 提交保存上下文,崩溃/切换后用 `/context-restore` 重建会话 |
### 安全防护
| 命令 | 说明 |
| ----------------------- | ------------------------------------- |
| `/careful` | 危险操作警告:rm -rf、DROP TABLE、force-push 等 |
| `/freeze` / `/unfreeze` | 锁定/解锁编辑范围到特定目录 |
| `/guard` | `/careful` + `/freeze` 组合,最高安全模式 |
| `/checkpoint` | 保存/恢复工作状态快照 |
### 工具集成
| 命令 | 说明 |
| -------------------------------------------------- | --------------------------------------------------------------------------- |
| `/codex` | OpenAI Codex CLI 集成:独立代码审查(pass/fail 门)、对抗模式、咨询模式;与 `/review` 都跑过后给出跨模型重叠分析 |
| `/health` | 代码质量仪表盘:tsc + biome + knip + shellcheck + tests → 0-10 综合评分 |
| `/skillify` | 把当前工作流固化为可复用 skill |
| `/scrape` | 网页抓取工作流 |
| `/landing-report` | 落地页性能与体验报告 |
| `/make-pdf` | 生成 PDF 文档 |
| `/benchmark-models` `/model-overlays` `/plan-tune` | 跨模型对比、覆盖叠加、计划调优 |
### Standalone CLI(v0.19+)
除了 slash 命令,gstack 还附带一组独立 CLI(不在 Claude Code 会话内跑):
| 命令 | 说明 |
| ------------------------ | --------------------------------------------------------------------------------------------------------- |
| `gstack-model-benchmark` | 跨模型评测:同一 prompt 同时跑 Claude / GPT(via Codex CLI)/ Gemini,对比延迟、token、成本和(可选)LLM-judge 质量分;不可用 provider 自动跳过 |
| `gstack-taste-update` | 设计品味学习:把 `/design-shotgun` 的赞同/否决写入项目级品味档案,每周衰减 5%,反哺后续变体生成 |
## 配置详解
### CLAUDE.md 添加内容
安装后,gstack 会在你的 `CLAUDE.md` 中添加所有可用 skill 的列表和简短描述。这让 Claude Code 知道可以使用哪些命令。
### Skill 目录结构
主入口是顶层 `~/.claude/skills/gstack/SKILL.md`,每个子命令以平级目录形式存在,核心都是 `SKILL.md` 文件:
```text
~/.claude/skills/gstack/
├── SKILL.md # 主入口 skill
├── browse/ # 浏览器 daemon
├── qa/ # QA 测试
├── review/ # 代码审查
├── ship/ # 发布流程
├── plan-ceo-review/ # CEO 审查
├── office-hours/ # 产品门诊
├── pair-agent/ # 跨 Agent 浏览器配对
├── open-gstack-browser/ # GStack Browser 启动器
├── setup-gbrain/ # GBrain 数据库一键上手
├── hosts/ # 10 个 host 配置(claude/codex/cursor/...)
├── bin/ # standalone CLI(gstack-model-benchmark 等)
└── ... # 当前 v1.x 共 50 个 skill 目录
```
你可以自由修改任何 `SKILL.md` 来定制行为——这就是"fork 和定制"的优势。
### Browse Daemon
Browse Daemon 是长驻 Chromium 实例,关键配置:
* **端口**:随机选择 10000-60000,支持 10+ 个并行工作区
* **安全**:仅绑定 localhost,每次会话使用 bearer token 认证
* **Cookie**:使用 `/setup-browser-cookies` 可从 Chrome/Arc/Brave/Edge 导入
## 实战工作流演示
以下演示一个典型的 gstack 工作流。命令和输出基于文档和视频中的真实案例整理。
> 💡 **注意**:以下输出为基于研究整理的通用示例。后续会根据真实实操补充具体项目截图。
### Step 1: 产品门诊
```text
> /office-hours
[YC Office Hours] 6 forcing questions:
1. Who specifically needs this?
2. What do they do today without it?
3. Why is this urgent right now?
4. How will you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can ship?
→ Design doc generated
```
不急着写代码,先让 AI 以 YC Office Hours 的视角拷问你的想法。
### Step 2: 多角色审查计划
```text
> /autoplan
[CEO Review] Finding the 10-star product...
[Design Review] Rating dimensions 0-10...
[Eng Review] Locking architecture + edge cases...
→ Fully reviewed plan ready
```
`/autoplan` 自动运行 CEO → 设计 → 工程三轮审查,产出完整的审查后计划。
### Step 3: 编码实现
按审查通过的计划正常编码。可以用标准 Claude Code 对话方式。
### Step 4: 多专家代码审查
```text
> /review
Dispatching 7 specialist reviewers...
- Testing coverage ✓
- Maintainability ✓
- Security: Found 1 issue (auto-fixing)
- Performance ✓
- Data migration ✓
- API contract ✓
- Red team: No vulnerabilities found
→ Review complete, 1 auto-fix applied
```
### Step 5: 浏览器 QA
```text
> /qa
Opening headless browser...
Testing user flows:
- Login flow ✓
- Dashboard load ✓
- Form submission: Bug found → fixing → re-testing ✓
- Image upload ✓
→ 4 flows tested, 1 bug fixed, regression test generated
```
### Step 6: 发布
```text
> /ship
Syncing with main...
Running tests: 42 passed, 0 failed
Reviewing diff: 3 files changed
Updating VERSION: 1.2.0 → 1.3.0
Creating PR: "Add screenshot feature"
→ PR #47 created, ready for merge
```
## 实用技巧与社区经验
### Garry Tan 的建议
来自 gstack 的 ETHOS.md,三个核心原则:
1. **Boil the Lake**:AI 让完整性几乎免费——永远做完整的事,不走捷径
2. **Search Before Building**:先搜索、先理解,三层知识验证后再动手
3. **User Sovereignty**:AI 推荐,你决定。即使两个 AI 模型都同意,你的判断依然优先
而 gstack 的 README 开篇用了一段 Karpathy 的话——这也是 Garry Tan 自己解释他为什么要做 gstack 的起点:
### 社区正面经验
* **`/office-hours` 用于 YC 申请**:Reddit r/ycombinator 上多位 S26 申请者反馈,用 gstack 的 office hours 来压力测试自己的申请材料效果极佳
* **安全审计发现真实漏洞**:有 CTO 反馈 `/review` 发现了团队不知道的 XSS 漏洞
* **`/browse` 真实浏览器测试**:被社区(包括批评者)认可为"真正有技术含量的贡献"
### 常见踩坑
* **权限提示频繁**:有用户反馈"每 30 秒要批准一次权限提示,根本没法睡觉"。建议在 Claude Code 设置中配置适当的自动批准规则
* **Token 消耗较高**:角色化 prompt 会增加上下文消耗。如果成本敏感,可以选择性使用最需要的 skill
* **Agent 循环**:HN 上有用户报告 agent 陷入 70 分钟循环的案例。建议设置合理的超时和检查点
* **不适合所有人**:资深开发者可能觉得大部分 skill 是不必要的包装。gstack 更适合**独立创始人和小团队**,而非已有成熟工程流程的团队
## 常见问题与最佳实践
**Q:gstack 和 Superpowers 可以同时使用吗?**
可以。两者互补——Superpowers 擅长流程纪律和 TDD 保障,gstack 擅长产品思维和多角色审查。很多团队用 Superpowers 做日常编码纪律,用 gstack 做产品规划和 QA。
**Q:Token 消耗大吗?**
比原生 Claude Code 高。每个 skill 的角色 prompt 会占用上下文窗口。但如果你的时间比 token 费用更有价值,这通常是划算的。
**Q:适合什么类型的项目?**
最适合**全流程产品开发**——从想法到上线。如果只是修 bug 或做小功能,原生 Claude Code 就够了。gstack 的价值在"完整流程"中最大化。
**Q:如何定制 skill?**
每个 skill 就是一个 `SKILL.md` 文件。直接编辑即可:
1. 找到 skill 目录:`~/.claude/skills/gstack//`
2. 编辑 `SKILL.md`
3. 重新运行 `./setup`
社区建议 fork 仓库后定制,而非直接修改全局安装。
### 最佳实践
1. **先 `/office-hours` 再编码**:养成习惯,在写任何代码之前先做产品门诊
2. **善用 `/browse` 验证**:不要只看代码,让 AI 真正"看到"你的应用
3. **定期 `/retro`**:保持对代码质量和工作节奏的可见性
4. **渐进采用**:不需要一次用所有 skill。从 `/office-hours` + `/review` + `/ship` 开始
5. **Fork 定制**:遇到不合适的 prompt,直接改。这是开源的优势
## 小结
gstack 的核心价值不在于某个具体 skill 有多强大,而在于它提供了一种**结构化的 AI 协作模式**——通过角色切换,让你在不同阶段获得不同类型的 AI 辅助。先以 CEO 的视角审视产品方向,再以工程经理的严谨审查架构,最后以 QA 的真实浏览器验证结果。
下一步,你可以亲自安装试试,从 `/office-hours` 开始你的第一个 gstack 项目。
***
**延伸阅读**:
* [gstack 概念篇](/docs/notes/gstack/concept) — 理解 gstack 的核心理念和工具生态定位
* [GSD 实战篇](/docs/notes/gsd/practice) — 另一种结构化 AI 编程方案的实战指南
* [Claude Skills 实战篇](/docs/notes/claude-skills/skill-creator) — 理解 Skill 的创建机制
# 拆解 gstack:Skill 开发者能学到什么
## 引言
在 [概念篇](/docs/notes/gstack/concept) 和 [实战篇](/docs/notes/gstack/practice) 中,我们从用户视角了解了 gstack 是什么、怎么用。这篇笔记换一个角度——**作为 Skill 开发者**,逐个文件读完 gstack 仓库后,哪些工程设计值得学习和借鉴。
gstack 不只是 23 个 prompt 文件的集合。它背后有一套完整的工程体系:模板生成、自动升级、学习记忆、渐进式引导、多平台适配、分层测试——这些才是让一个 skill 项目从"能用"变成"好用"的关键。
***
## 1. SKILL.md 不是手写的——模板生成系统
gstack 最反直觉的设计:**每个 SKILL.md 都是自动生成的,不能直接编辑。**
```text
SKILL.md.tmpl (人写) → gen-skill-docs → SKILL.md (机生)
```
人写的 `.tmpl` 模板中包含工作流逻辑和最佳实践,加上 `{{PLACEHOLDER}}` 占位符。构建脚本从源代码中提取命令参考、浏览器 flag 列表、preamble 启动代码等,填充到占位符中生成最终的 SKILL.md。
```text
{{PREAMBLE}} ← 从 resolvers/preamble.ts 生成的启动代码
{{BROWSE_SETUP}} ← 浏览器初始化指令
{{COMMAND_REFERENCE}} ← 从 commands.ts 提取的命令文档
{{SNAPSHOT_FLAGS}} ← 从源代码常量提取的快照选项
```
**为什么这样做?**
* 文档和代码永远不会不同步——命令参考从源代码生成,源码改了文档自动更新
* 23 个 skill 共享同一份 preamble(约 220 行),更新一次所有 skill 同步更新
* CI 可以 `--dry-run` 检查生成文件是否过期,防止忘记重新生成
**值得借鉴的点**:如果你维护多个 skill,任何跨 skill 共享的内容都应该提取到模板,用构建步骤生成最终文件。手动同步多份相同内容迟早会出问题。
***
## 2. 升级机制——从检测到执行的完整链路
gstack 的升级系统设计得很精致,分三层:
### 第一层:版本检测
`bin/gstack-update-check` 是一个独立的 bash 脚本,做以下事情:
1. 读本地 `VERSION` 文件
2. 检查缓存 `~/.gstack/last-update-check`(UP\_TO\_DATE 缓存 60 分钟,UPGRADE\_AVAILABLE 缓存 720 分钟)
3. 缓存过期则 HTTP 请求 GitHub 的 `raw.githubusercontent.com/.../VERSION`
4. 比对版本号,输出 `UPGRADE_AVAILABLE <旧> <新>`
### 第二层:Preamble 集成
**每个 skill 的 SKILL.md 启动代码的第一行就是版本检测**:
```bash
_UPD=$(~/.claude/skills/gstack/bin/gstack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
```
这意味着用户调用任何 skill 时都会自动检测更新——不需要专门跑升级命令,存在感为零但覆盖率 100%。
### 第三层:渐进提醒 + 自动升级
检测到新版本后不会立刻打扰用户,而是用\*\*贪睡机制(Snooze)\*\*做渐进退避:
* 第 1 次提醒:24 小时后再提
* 第 2 次提醒:48 小时后再提
* 第 3 次及以后:7 天后再提
* 新版本发布会重置贪睡计数器
用户可以 `gstack-config set auto_upgrade true` 开启自动升级,跳过确认直接执行。
升级执行时会区分 5 种安装类型(全局 git、本地 git、vendored 等),git 安装用 `git fetch + reset`,vendored 安装先备份再替换,失败时从 `.bak` 恢复。升级后还会自动同步项目本地的 vendored 副本。
**值得借鉴的点**:
* "每次调用时检测"的模式覆盖率极高且用户无感知
* 渐进退避避免了频繁打扰
* 区分安装类型执行不同升级策略,而不是一刀切
* 备份恢复保证升级失败不至于整个 skill 挂掉
***
## 3. 学习系统——让 Skill 越用越聪明
gstack 实现了一个轻量但有效的**跨会话记忆系统**。
### 存储
每个项目有独立的学习日志:`~/.gstack/projects/$SLUG/learnings.jsonl`,追加写入。
```json
{
"skill": "review",
"type": "pitfall",
"key": "n-plus-one",
"insight": "这个项目的 User model 有 N+1 查询问题,findAll 要加 include",
"confidence": 8,
"source": "observed",
"files": ["src/models/user.ts"],
"ts": "2026-04-01T14:30:00Z"
}
```
### 自动采集
每个 skill 完成前都有一个"操作自我改进"环节——反思这次执行中有没有意外失败、走弯路、或发现项目怪癖,有则自动记录到 learnings.jsonl。不需要用户手动触发。
### 自动加载
每次新会话启动时,preamble 会加载前 3 条高置信度学习条目注入上下文,让新会话继承历史知识。
### 置信度衰减
`observed` 和 `inferred` 来源的条目每 30 天衰减 1 分。不需要人工清理知识库——过时的知识自然消退,新的观察自然取代。
### 管理界面
```text
/learn # 显示最近 20 条
/learn search # 搜索
/learn prune # 检测过期条目(引用的文件已删除)
/learn export # 导出为 markdown 可加入 CLAUDE.md
```
**值得借鉴的点**:
* 追加只写设计简单可靠,并发安全
* 置信度衰减是低维护的知识时效管理——比手动清理高效得多
* 用 git remote URL 而非路径标识项目(通过 `gstack-slug`),clone 到不同位置也能复用
* 支持跨项目查询,但默认隔离
***
## 4. Preamble 注入——Skill 的"中间件层"
这是 gstack 最聪明的架构设计之一。每个 SKILL.md 共享一段约 220 行的 preamble 代码,功能类似 Web 框架的中间件:
```text
┌─ 更新检测 ──────────────────────────────────┐
│ 会话追踪 (sessions/$PPID) │
│ 配置读取 (proactive, skill_prefix, telemetry)│
│ 学习历史加载 (前 3 条高置信度) │
│ 上下文恢复 (最近的 checkpoint + timeline) │
│ 路由规则检测 │
│ 首次使用引导流程 │
└──────────────────────────────────────────────┘
↓
Skill 特有逻辑
```
Preamble 的 bash 脚本输出键值对(`BRANCH: main`、`PROACTIVE: true`),然后模板用自然语言条件让 Claude 据此调整行为:
```text
If PROACTIVE is false, do not invoke skills automatically.
Instead suggest: "I think /skillname might help here -- want me to run it?"
```
这本质上是**把 bash 输出当作 Claude 的"环境变量"**——用 bash 做运行时检测,用自然语言做行为路由。
**值得借鉴的点**:如果你有多个 skill,共享逻辑(配置加载、状态恢复、版本检测)应该提取到统一的 preamble,而不是每个 skill 各写一份。
***
## 5. 渐进式引导——Sentinel 文件模式
gstack 的首次使用体验设计得很用心。通过 touch 文件(sentinel file)确保每个引导步骤只出现一次:
```text
~/.gstack/.completeness-intro-seen ← "Boil the Lake" 理念介绍
~/.gstack/.telemetry-prompted ← 遥测选择(community/anonymous/off)
~/.gstack/.proactive-prompted ← 主动触发开关
~/.gstack/.routing-prompted ← CLAUDE.md 路由规则写入
~/.gstack/.welcome-seen ← 安装欢迎消息
```
每次 skill 启动时检查这些文件是否存在,不存在则展示对应引导并 touch 文件。已经看过的步骤永远不会再出现。
**值得借鉴的点**:比起在 config 里维护 `"onboarding_step": 3` 的状态,sentinel 文件更简单、更可靠——不会被配置文件损坏影响,而且每个步骤独立控制。
***
## 6. SKILL.md 结构设计——三层架构
每个 SKILL.md 都遵循标准的三层结构:
### 第一层:YAML Frontmatter
```yaml
---
name: qa
preamble-tier: 3
version: 0.15.1.0
description: |
Systematically QA test a web application...
Use when asked to "qa", "test this site", "find bugs"...
benefits-from: [office-hours]
allowed-tools:
- Bash
- Read
- Write
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "bash ${CLAUDE_SKILL_DIR}/bin/check-careful.sh"
---
```
关键字段:
* `allowed-tools`:工具级权限白名单,每个 skill 只声明自己需要的工具
* `benefits-from`:显式声明前置依赖 skill
* `hooks`:PreToolUse 钩子,可以在工具调用前拦截(如 careful 拦截 `rm -rf`)
* `description`:包含所有自然语言触发词
### 第二层:共享 Preamble + 通用规则
preamble 启动代码 + Voice 定义 + 上下文恢复 + 完整性原则 + 搜索优先级 + 完成状态协议 + 升级规则等。所有 skill 完全相同,由模板生成。
### 第三层:Skill 特有逻辑
这才是每个 skill 的"灵魂"——工作流定义、角色设定、认知模式注入、交互门控等。
**值得借鉴的点**:三层分离让每个 skill 只需要关注自己的独特逻辑,共享部分由框架保证一致性。
***
## 7. Prompt 工程技巧合集
读完所有 SKILL.md 后,以下是最值得学习的 prompt 设计技术:
### 反谄媚规则
office-hours 的 Startup Mode 明确禁止了 AI 常见的"和稀泥"行为:
```text
Never say:
- "That's an interesting approach" → take a position instead
- "There are many ways to think about this" → pick one
- "You might want to consider..." → say "This is wrong because..."
- "That could work" → say whether it WILL work
```
### 禁止词表
Voice 部分有明确的禁用词和禁用短语:
* 禁用词:delve, crucial, robust, comprehensive, nuanced, pivotal, landscape...
* 禁用短语:"here's the kicker", "plot twist", "let me break this down"...
* 禁用格式:em dash(用逗号/句号代替)
这些是 LLM 常见的"AI 味"词汇,禁掉后输出明显更自然。
### 认知模式注入
每个 Review skill 注入了不同的思维框架:
* **CEO Review**:18 条认知模式(Bezos 的单向/双向门决策、Munger 的逆向思维、Jobs 的专注即减法...)
* **Eng Review**:15 条工程管理模式("boring by default"、blast radius 直觉、Conway 定律...)
* **Design Review**:12 条设计认知模式(层级即服务、约束崇拜、"Would I notice?" 测试...)
这些模式不是让 AI 机械执行,而是给它提供**思考框架**——就像给一个聪明的新人一份前辈的经验清单。
### 具体化标准
```text
Not "you should test this"
but `bun test test/billing.test.ts`
Not "this might be slow"
but "this queries N+1, ~200ms per page load with 50 items"
Not "there's an issue in the auth flow"
but "auth.ts:47, the token check returns undefined"
```
### 置信度校准
review skill 要求每个发现附带置信度分数,低置信度的发现自动降级或隐藏:
| 分数 | 含义 | 处理方式 |
| ---- | --------- | ----------- |
| 9-10 | 读了具体代码验证过 | 正常展示 |
| 7-8 | 高置信度模式匹配 | 正常展示 |
| 5-6 | 中等,可能误报 | 附带说明展示 |
| 3-4 | 低置信度 | 从报告中隐藏 |
| 1-2 | 纯猜测 | 只在 P0 级别才展示 |
### 交互门控
ship skill 精确定义了**什么时候该停下来等用户、什么时候该自动继续**:
```text
Only stop for:
- Tests failing with no obvious fix
- Merge conflicts requiring human judgment
- Unclear which changes to include
Never stop for:
- Normal git operations
- CHANGELOG/VERSION updates
- PR creation
```
**值得借鉴的点**:好的 skill 不是"AI 做完所有事",而是精确定义人机边界。
***
## 8. 状态管理——文件系统即数据库
gstack 的所有持久化都通过文件系统完成,存储在 `~/.gstack/` 下:
| 路径 | 用途 | 格式 |
| ------------------------------------- | ------ | -------- |
| `config.yaml` | 全局配置 | YAML |
| `sessions/$PPID` | 活跃会话 | touch 文件 |
| `projects/$SLUG/learnings.jsonl` | 学习记录 | JSONL |
| `projects/$SLUG/timeline.jsonl` | 技能时间线 | JSONL |
| `projects/$SLUG/checkpoints/*.md` | 检查点 | Markdown |
| `projects/$SLUG/health-history.jsonl` | 健康检查历史 | JSONL |
| `analytics/skill-usage.jsonl` | 使用遥测 | JSONL |
| `last-update-check` | 版本缓存 | 纯文本 |
几乎所有时间序列数据都用 **JSONL**(每行一个 JSON 对象),追加写入。这个选择很聪明:
* 追加写入天然并发安全
* 不需要数据库依赖
* 可以用 `grep` / `jq` 直接查询
* 损坏最多丢失最后一行
***
## 9. 跨 Skill 集成模式
### 文件传递产物
skill 之间通过文件系统传递工作产物:
```text
/office-hours → design doc → /plan-ceo-review 读取
/plan-ceo-review → ceo-plans/*.md → /autoplan 读取
/review → reviews.jsonl → /ship 读取并展示 Dashboard
/qa → qa-reports/ → /retro 读取
```
### Review Readiness Dashboard
ship skill 读取 `reviews.jsonl`,在发布前展示跨 skill 审查状态:
```text
| Review | Runs | Last Run | Status | Required |
| Eng Review | 1 | 2026-03-16 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
```
### 前置依赖建议
plan-ceo-review 检测到没有 design doc 时,会主动建议先运行 `/office-hours`:
```text
"No design doc found. /office-hours produces a structured problem statement...
Takes about 10 minutes."
Options: A) Run /office-hours now B) Skip
```
### 使用序列预测
Context Recovery 会分析最近的 skill 使用顺序,预测下一步:
```text
If pattern repeats (e.g., review → ship → review),
suggest: "Based on your recent pattern, you probably want /ship."
```
***
## 10. 其他值得注意的设计
### Hook 系统
careful、freeze、guard 三个 skill 用了 `PreToolUse` hooks——这是唯一能在工具调用**之前**拦截的机制:
* **careful**:拦截 Bash,检查 `rm -rf`、`DROP TABLE`、`git push --force`
* **freeze**:拦截 Edit/Write,检查路径是否在允许范围内
* **guard**:组合以上两者
### 多平台适配
同一套模板通过 `--host` 参数生成不同平台的 skill 文件:
```bash
bun run gen:skill-docs --host claude # Claude Code 格式
bun run gen:skill-docs --host codex # OpenAI Codex 格式
bun run gen:skill-docs --host kiro # AWS Kiro 格式
bun run gen:skill-docs --host factory # Factory Droid 格式
```
路径和 frontmatter 自动适配,skill 逻辑不变。
### 完成状态协议
每个 skill 结束时必须输出标准化的完成状态:
```text
DONE — 全部完成,提供证据
DONE_WITH_CONCERNS — 完成但有顾虑
BLOCKED — 无法继续
NEEDS_CONTEXT — 需要更多信息
```
### 三次失败升级规则
```text
If you have attempted a task 3 times without success, STOP and escalate.
```
防止 AI 陷入无限重试循环。
### Diff-based 测试选择
E2E 测试每次约 $4(需要启动 Claude agent),所以 gstack 通过 `touchfiles.ts` 声明每个测试依赖的源文件,根据 `git diff` 只运行受影响的测试:
```typescript
// test/helpers/touchfiles.ts
{
"qa-workflow": ["qa/SKILL.md.tmpl", "browse/src/server.ts"],
"ship-flow": ["ship/SKILL.md.tmpl", "scripts/resolvers/preamble.ts"]
}
```
***
## 总结:可以带走的设计原则
从 gstack 的工程实践中,我提炼出以下几个对 Skill 开发者最有价值的设计原则:
1. **模板生成 > 手动同步**:跨 skill 共享的内容用模板 + 构建步骤自动生成,不要复制粘贴
2. **被动检测 > 主动检查**:升级检测嵌入每次 skill 调用,用户无感知但覆盖率 100%
3. **追加日志 > 复杂数据库**:JSONL + 文件系统能覆盖绝大多数持久化需求,简单可靠
4. **渐进引导 > 一次配置**:用 sentinel 文件控制引导步骤,每个只出现一次
5. **精确门控 > 全自动**:明确定义"停下来等用户"和"自动继续"的边界
6. **置信度量化 > 模糊判断**:每个 AI 判断附带置信度分数,低置信度自动降级
7. **时间衰减 > 手动清理**:学习记录的置信度随时间衰减,过时知识自然消退
8. **禁止词表 > 风格指南**:直接列出不准用的词比"请使用自然的语气"有效得多
***
**相关阅读**:
* [gstack 概念篇](/docs/notes/gstack/concept) — gstack 是什么、解决什么问题
* [gstack 实战篇](/docs/notes/gstack/practice) — 从安装到跑通完整工作流
* [gstack 前端 Skill 篇](/docs/notes/gstack/frontend-skills) — 前端/UI 设计 Skill 全景与推荐工作流
* [Claude Skills 概念篇](/docs/notes/claude-skills/concept) — 理解 Skills 的底层机制
# Diagnose 与 Triage:先建立反馈回路,再决定交给谁
## 为什么把两个 Skill 放在一起讲
`/diagnose` 和 `/triage` 在 README 里是两个独立 skill,但它们解决的是同一个工程问题的两半:
* `/diagnose` 关心:**这个 bug 到底是什么、怎么复现、怎么证明修好了**
* `/triage` 关心:**这个 issue 现在该等信息、给 agent、给人,还是不做**
一个负责事实,一个负责流程。真实项目里这两个经常连在一起:先 triage 一个 bug issue,发现信息不够就 `needs-info`;信息够了就用 diagnose 建反馈回路;复现清楚后再决定是 `ready-for-agent` 还是 `ready-for-human`。
## /diagnose 的核心:反馈回路就是全部
`/diagnose` 最值得记住的一句话是:**先建立一个 agent 能运行的 pass/fail 信号**。
Matt 把诊断拆成 6 个阶段:
| 阶段 | 目标 |
| --------------------- | ------------------- |
| Build a feedback loop | 搭一个快速、确定、可反复运行的失败信号 |
| Reproduce | 让这个信号复现用户描述的同一个 bug |
| Hypothesise | 列 3-5 个可证伪假设 |
| Instrument | 用最少探针验证假设 |
| Fix + regression test | 在正确测试面写回归测试,再修 |
| Cleanup + post-mortem | 清理临时探针,记录真实根因 |
这和很多人调 bug 的顺序相反。普通调试常见流程是:看代码、猜原因、改一改、刷新页面。Matt 反过来:先把 bug 变成一个可重复机器信号,再谈假设。
## 什么算好反馈回路
`/diagnose` 给了一组优先级,从最好到最兜底:
| 回路 | 适合场景 |
| ----------------------------- | --------------------- |
| 失败测试 | 有合适测试面,能直接表达 bug |
| curl / HTTP script | API bug、服务端行为可用请求复现 |
| CLI + fixture | 命令行工具、解析器、转换器 |
| Headless browser | UI bug、控制台错误、网络行为 |
| Replay captured trace | 线上真实 payload、事件流、日志链路 |
| Throwaway harness | 只启动系统一小块,隔离复杂依赖 |
| Property / fuzz loop | 偶现错误输出,需要提高触发率 |
| Bisection / differential loop | 某版本后坏了,需要二分或对比旧版 |
| HITL script | 只能人手点时,也要让人按脚本提供稳定输出 |
这里有一个很硬的判断:**没有回路,不要进入假设阶段**。因为没有信号,所有分析都会变成「看起来像」。
## 非确定性 bug 怎么办
`/diagnose` 对偶发 bug 的态度也很实用:目标不是一开始就 100% 复现,而是先把复现率提高到可调试。
比如:
* 循环触发 100 次
* 并发触发
* 注入 sleep 拉大竞态窗口
* 固定随机种子或时间
* 缩小环境变量和外部依赖
1% 的偶发 bug 很难调;50% 的偶发 bug 就已经是可调试对象。这个思路对前端异步、消息队列、支付回调、流式输出都很有用。
## 假设必须可证伪
Matt 要求在动手验证前先列 3-5 个假设,并且每个假设都要写出预测:
```text
如果 X 是原因,那么改变 Y 后 bug 应该消失;
或者观察 Z 时应该出现某个特征。
```
这会防止 agent 被第一个看起来合理的解释锁死。更重要的是,它让你能判断某个实验到底有没有信息量。
一个坏假设:
```text
可能是缓存问题。
```
一个可证伪假设:
```text
如果是浏览器缓存导致旧脚本执行,那么禁用缓存并强刷后,console 里的旧 bundle hash 应该消失,按钮点击事件也应该恢复。
```
后者才值得验证。
## 修复阶段最容易犯的错
`/diagnose` 要求:如果有正确测试面,就先把最小复现转成失败测试,再修代码。
关键是「正确测试面」。不是随便补一个 unit test 就算回归测试。正确测试面必须覆盖真实 bug 模式:
* bug 是多个调用者组合触发的,就不能只测单个函数
* bug 是真实 payload 结构触发的,就不能只测一个手写 toy object
* bug 是浏览器事件顺序触发的,就不能只测纯函数
如果找不到正确测试面,这本身就是结论:代码结构没有给你留下可锁定 bug 的地方。修完之后应该把这个信息交给 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture)。
## /triage 的核心:issue 是状态机
`/triage` 不是让 AI 随便帮你「看一下 issue」。它把 issue 当作一个小状态机。
每个 issue 应该同时有:
* 一个 category:`bug` 或 `enhancement`
* 一个 state:`needs-triage`、`needs-info`、`ready-for-agent`、`ready-for-human`、`wontfix`
这套状态的价值在于让维护者可以快速回答:
* 哪些还没人看?
* 哪些等报告者补信息?
* 哪些已经清楚到可以交给 AFK agent?
* 哪些必须人类自己做?
* 哪些应该关掉,并把原因沉淀下来?
## ready-for-agent 的标准
`ready-for-agent` 是这套流程里最关键的状态。它不是「这个任务可以让 AI 试试」,而是:
> 任务已经清楚到一个不在场的 agent 可以独立领取、实现、验证。
这通常意味着 issue 里至少有:
* 背景和问题陈述
* 相关代码路径或模块
* 明确的验收标准
* 已知约束
* 如果是 bug,最好有复现方式
* 不需要额外产品/设计判断
如果缺这些,应该是 `needs-info` 或 `ready-for-human`,而不是勉强丢给 agent。
## needs-info 要问具体问题
`/triage` 给 `needs-info` 的模板很朴素,但重点是问题必须具体:
```markdown
## Triage Notes
**What we've established so far:**
- ...
**What we still need from you (@reporter):**
- ...
```
坏问题:
```text
请提供更多信息。
```
好问题:
```text
请提供触发问题的浏览器版本、出错页面 URL、点击顺序,以及 Network 面板里 `/api/orders/:id` 的 response body。
```
AI 很容易写礼貌废话,这个 skill 强迫它把「我们已经知道什么」和「还缺什么」分开。
## wontfix 也要沉淀
`/triage` 对 enhancement 的 `wontfix` 有一个有趣设计:不要只是关 issue,而是把拒绝理由写进 `.out-of-scope/` 知识库,再在评论里链接。
这样下次类似需求出现时,AI 不会重新展开同一场讨论。它可以先读 `.out-of-scope/`,提醒维护者:「这个方向之前拒绝过,理由是 X。」
这和 ADR 的精神很像:不是记录所有决定,只记录未来会让人疑惑、并且会反复出现的决定。
## 两者怎么配合
一个典型 bug issue 可以这样走:
1. `/triage` 读取 issue、评论、标签和相关代码
2. 它判断这是 `bug + needs-triage`
3. 先尝试复现;如果步骤不足,转 `needs-info`
4. 信息足够后,启动 `/diagnose`
5. `/diagnose` 建立复现回路,列假设,定位根因
6. 如果修复路径清楚、测试面明确,issue 变 `ready-for-agent`
7. 如果需要产品判断、外部权限、人工验证,issue 变 `ready-for-human`
8. 修完后把根因和回归测试写回 issue 或 PR
这套流程的关键不是「AI 自动修 bug」,而是让 issue 从模糊描述变成可执行工作包。
## 我的使用建议
如果你只记一条:
> `/diagnose` 先问「我怎么证明它坏了」;`/triage` 先问「它现在该处在哪个状态」。
这两个问题能挡住大量低质量 AI 编程:
* 没复现就修
* 没验收就开工
* 没根因就重构
* 没信息就甩给 agent
Matt 这两个 skill 并不花哨,但很像真实团队里资深工程师会做的事:先把事实收束,再推进流程。
## 参考资源
下一篇:[TDD:用红绿重构强迫 AI 走小步](/docs/notes/matt-pocock-skills/tdd)。
# Grill Me:让 AI 在你写代码前拷问你 50 个问题
## 核心结论
**Grill Me 是一个用于需求澄清和方案拷问的 AI Skill**:它让 Claude Code 在写代码前连续追问你的计划、设计和边界条件,直到你和 AI 对同一个 design concept 达成共识。
它解决的不是「AI 代码写得不够快」,而是另一个更常见的问题:**需求还没对齐,AI 已经把代码写完了**。Plan Mode 更像「先生成计划再执行」,Grill Me 更像「先把你脑子里的模糊需求逼清楚」。
## 失败模式:「AI 没做我想要的」
Matt 在演讲里讲的第一个失败模式是:你以为脑子里的需求很清楚,让 AI 一写出来——完全不是那回事。
> "I would run it, and I would try not to look at the code, but I would look at the code, and I realized I would get worse code. I did it again, I got even worse code... I did it again, kept running the compiler, and I would just end up with garbage."
这个体感很多人都熟:你说「帮我加个登录」,AI 不会问你「要不要记住设备」「失败几次锁账号」「session 多久过期」。它直接铺一套它觉得合理的方案。等你 review 时已经写了 500 行——回炉两小时。
## 为什么会这样:design concept 跑偏
Matt 引用 Frederick P. Brooks 在《The Design of Design》里的 **design concept**(设计概念):
> 当多人合作设计一个东西时,你们之间会有一个**正在被造出来的东西**——它在脑子里飘着,是个隐形的「关于这个东西的理论」。它不是 asset,不是塞进 markdown 文件的资产,而是看不见的共识。
AI 一上来就动手写代码,意味着它根本没和你共享同一个 design concept。代码写出来错的不是语法,是**前提**。
要修这个问题,得在动手前先做 design concept 对齐。Brooks 给的工具叫 **design tree**——把一个决策拆成多分支,每个分支再拆。你不能跳过上游决策直接做下游决策,否则上游一变下游全要重做。
## Matt 的 Skill 全文
Matt 在 [`mattpocock/skills`](https://github.com/mattpocock/skills) 里给这套理论的实现是 `productivity/grill-me/SKILL.md`,整个文件加 frontmatter 也不到 15 行:
```markdown
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
逐句解读:
* **"interview me relentlessly"** —— 关键词是 *relentlessly*(不放过)。LLM 默认有个倾向是问 1\~2 个问题就觉得「差不多了」开始动手。这个词强行压住这个倾向。
* **"walk down each branch of the design tree"** —— Brooks 的 design tree 概念。强迫 Claude 把你的需求当作树,先解上游再解下游。如果你说「做登录」,它会先问「鉴权方式」(树根),再根据你的回答展开「session 怎么管」/「token 怎么存」(子节点)。
* **"resolving dependencies between decisions one-by-one"** —— 显式禁止打包提问。决策之间常有依赖(选了 SSO,下游就不需要密码策略问题),先确定上游能砍掉一堆下游问题。
* **"for each question, provide your recommended answer"** —— 关键加分项。AI 不只是问,还要给推荐答案。你只要点头/否决,省下 80% 输入时间。
* **"ask the questions one at a time"** —— 防止 AI 一次甩 10 个问题让你头大。
* **"if a question can be answered by exploring the codebase, explore the codebase instead"** —— 如果是项目内已有的事实(比如「项目用的什么测试框架」),让 Claude 自己看,不要问你。
7 行内容,但每一句都对应一个具体的 LLM 行为偏差。
## 怎么装、怎么用
**安装**:
```bash
npx skills@latest add mattpocock/skills
```
勾选 `grill-me` 和 `setup-matt-pocock-skills`(grill-me 不依赖后者,但其他 skill 依赖,建议一起装)。
**调用**:在 Claude Code 对话里输入 `/grill-me`。
**典型流程**:
1. 你描述一个想做的事,可以很模糊(「我想给博客加个评论功能」)
2. 输入 `/grill-me`
3. Claude 开始一题一题问,每题都给推荐答案
4. 你逐题回答(点头 / 否 / 改)
5. 一般 20\~50 个问题之后达成共识,Claude 给你一个总结
6. 总结可以直接喂给 [`/to-prd`](/docs/notes/matt-pocock-skills/to-prd-and-issues) 变成 PRD,或者直接交给 [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 开始写
## 真实案例:一个视频编辑器功能要花多少问题
Matt 在 [《5 Agent Skills I Use Every Day》](https://www.aihero.dev/5-agent-skills-i-use-every-day) 里给了几个具体数字:
* **新增视频编辑器功能** —— 16 个问题就达成共识
* **复杂功能** —— 30\~50 个问题
* **极端复杂的** —— 100 个问题,session 长达 45 分钟
问题样例(从 Matt 的视频/博文还原):
* "Should video clips be reorderable, or only added/removed in sequence?"
* "When a clip is deleted, do we keep its source file, or delete the file too?"
* "Does the editor need undo/redo? How many steps deep?"
* "Should we render previews in the browser, or rely on a backend service?"
这些问题没有一个是技术问题——全是产品决策。但**每一个决策都决定了几百行代码的形态**。如果你跳过这些问题让 AI 直接写,它会自己脑补一套答案,写完你再回来一个个驳回。
## 和 Claude Code 自带 Plan Mode 的差异
Claude Code 自带 `plan mode`(按 Shift+Tab 进入)。表面看上去和 grill-me 类似——都是先讨论再动手。但 Matt 在演讲里直接说**他更喜欢 grill-me**:
具体差异:
| 维度 | Plan Mode | /grill-me |
| ------- | ------------- | --------------- |
| 默认目标 | 尽快产出可执行 plan | 先达成共识,plan 是副产物 |
| 提问数量 | 0\~5 个 | 20\~100 个 |
| 提问形式 | 一次问一段 | 一次问一题 |
| 是否给推荐答案 | 否 | 是 |
| 是否探索代码库 | 偶尔 | 主动(明确指令) |
| 适合场景 | 已经想清楚,想确认实现方案 | 还没想清楚,需要被逼着想清楚 |
最大的实际差异是「**急不急**」。Plan mode 急着开干,grill-me 不急——它把「想清楚」当作主任务而不是序章。
## 进阶用法
### 1. 非编程场景
`grill-me` 不绑定代码,纯产品决策对话也能用。Matt 自己用它做:
* 课程大纲设计
* 文章写作
* 内部沟通文档
只要你脑子里有个模糊的想法、想被逼着想清楚,就能用。
### 2. 配合 [`/to-prd`](/docs/notes/matt-pocock-skills/to-prd-and-issues)
grill-me session 结束后,直接说 `/to-prd`,Claude 会把整段对话浓缩成结构化 PRD(含 user story、模块拆分、测试策略),并提交到你的 issue tracker。**关键点:不要在中间清 context**——to-prd 是从对话上下文里直接提取,不会再问你一遍。
### 3. 配合 [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)
如果项目已经有 `CONTEXT.md`(领域语言)和 `docs/adr/`(架构决策),用 `/grill-with-docs` 替代 `/grill-me`。它会在拷问的同时**同步更新 CONTEXT.md**——决策一边做、文档一边更,不再有「文档永远过时」问题。
### 4. 自定义提问深度
如果你时间紧,可以在 `/grill-me` 之后直接补一句:「Limit to 10 questions, focus only on architecture decisions.」 它会按你的限制收敛。但 Matt 不推荐——他认为「问得多」恰恰是这个 skill 的价值,砍掉就和 plan mode 差不多了。
## 使用边界:什么时候该用,什么时候别用
Grill Me 的价值来自「提前暴露决策树」,所以它并不是越多用越好。判断标准很简单:**如果这个任务失败后的返工成本很低,就不要用;如果失败后会牵连产品逻辑、数据模型、权限边界或用户流程,就应该先被拷问。**
| 任务类型 | 是否建议用 Grill Me | 原因 |
| ------------------------ | -------------- | -------------------------------- |
| 改 typo、调文案、删 console.log | 不建议 | 决策空间太小,提问成本高于返工成本 |
| 加一个独立 UI 小组件 | 看情况 | 如果只影响局部,可以直接做;如果涉及状态、权限、埋点,建议先问 |
| 新增评论、登录、支付、导入导出功能 | 建议 | 上游决策会影响数据库、权限、错误处理和测试策略 |
| 重构已有模块 | 强烈建议 | 需要先确认行为兼容性、迁移路径和回滚方案 |
| 写 PRD、课程大纲、内部方案 | 建议 | 它能把隐含假设问出来,再交给后续写作或 PRD skill 沉淀 |
第一次跑会觉得烦。习惯了「一句话生成 500 行」的人,第一次被 AI 反问 30 次会觉得在浪费时间。Matt 的建议是**忍住前 5 题**——前 5 题往往会暴露你自己都没想清楚的事。
如果 AI 问到你不在乎的技术细节,直接回「your call」「你定」即可。Grill Me 的目的不是逼你亲自决定所有事,而是把真正重要的上游决策暴露出来。
## 我的使用判断
我会把 Grill Me 当成 AI 编程里的「刹车系统」,不是「加速器」。它表面上让你慢下来,多花 20\~45 分钟回答问题;但真正节省的是后面 review、返工和推倒重来的时间。
我不建议把它包装成万能 prompt。它真正厉害的地方不是“问很多问题”,而是让 AI 承认:**在共享设计概念没有形成之前,马上写代码是一种过早行动。**
## 这个 Skill 为什么火
`/grill-me` 是 Matt 整套 skill 里**最常被截图转发**的一个。原因不复杂:
1. **极度极简**:7 行 markdown,复制粘贴就用
2. **效果立竿见影**:第一次跑就能感受到 AI 的「问题密度」变化
3. **可移植**:不依赖 Claude Code,Codex、Cursor、Aider 都能用
4. **自带反 LLM 默认行为**:每个词都在反一个具体的 LLM 偏差,工程审美高
它的成功也成了「**skill 不一定要长**」这件事的最佳论据。
## 参考资源
下一篇:[Grill With Docs:维护项目语言和 ADR](/docs/notes/matt-pocock-skills/grill-with-docs)——grill-me 的进阶版,给有领域复杂度的项目用。
# Grill With Docs:用领域语言和 ADR 给 AI 装上项目记忆
## 核心结论
**Grill With Docs 是 grill-me 的项目化版本**:它不仅会在写代码前拷问需求,还会把你的计划和项目里的 `CONTEXT.md`、`docs/adr/` 对照起来,发现术语冲突、代码事实冲突和架构决策缺口。
它最适合有长期维护价值的真实项目。简单项目用 `/grill-me` 就够了;一旦项目里出现领域术语、历史架构决策、多人协作或跨会话上下文丢失,`/grill-with-docs` 的价值就开始超过普通需求拷问。
## 失败模式:「AI 太啰嗦」
Matt 演讲里的第二个失败模式:
> AI 用一堆冗长的话表达一件简单的事。它和你像在用两种语言。
这事跟代码量没关系,是**词汇错位**。AI 默认会用泛化术语("item"、"data"、"handler"),而你脑子里项目内的真正术语可能是「Course」、「Draft Version」、「Ghost Lesson」。AI 不知道这些词在你项目里有特定含义,就会绕开它们造一堆同义新词,结果是:
* 思考过程啰嗦(要绕开你的专有词)
* 实现和你脑子里的设计错位(因为不在同一个语义空间)
* 跨会话不可复用(每次对话都得重新建立一遍上下文)
## 经典理论:DDD 的 Ubiquitous Language
Matt 引的是 Eric Evans 的《Domain-Driven Design》。这本书 2003 年出版,提出了 **ubiquitous language**(统一语言)这个概念:
> 把领域专家、开发者、代码三者用同一套术语贯通。一个词在产品讨论里、代码注释里、变量名里、文档里——必须指同一个东西。
DDD 的目标是让代码长得像领域专家的脑子。在 AI 时代,多了一个新角色:**LLM 也得在这套语言里**。LLM 不在 standup 会议里、看不到产品需求会、读不懂你们组的黑话——它只能从你给它的文档里学。
Matt 把这件事变成了一个 skill:扫描代码库提取术语,生成一个 markdown 文件 `CONTEXT.md`,然后让人和 AI 都用它对齐。
## Skill 的进化:从 Ubiquitous Language 到 Grill With Docs
最早这个 skill 叫 `ubiquitous-language`——只做一件事:扫描代码库生成术语表。但 Matt 后来发现单生成一份文档不够:
* **文档会过时**:今天生成、明天改了代码、术语表没跟上
* **人不会主动看**:放在那也是死的
他把它重构成了 `grill-with-docs`,做了三件事的合并:
1. **拷问需求**(继承自 grill-me 的全部能力)
2. **挑战已有术语表**:你说的词和 CONTEXT.md 里写的不一致?立刻指出来
3. **决策时同步更新文档**:拷问过程中达成的新结论,inline 写进 CONTEXT.md 或新建 ADR
这是从「生成静态文档」到「**对话即维护文档**」的范式跃迁。
## Skill 全文
`engineering/grill-with-docs/SKILL.md` 的核心结构:
```markdown
---
name: grill-with-docs
description: Grilling session that challenges your plan against the
existing domain model, sharpens terminology, and updates
documentation (CONTEXT.md, ADRs) inline as decisions crystallise.
---
Interview me relentlessly about every aspect of this plan until we
reach a shared understanding. Walk down each branch of the design
tree, resolving dependencies between decisions one-by-one. For each
question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question
before continuing.
If a question can be answered by exploring the codebase, explore the
codebase instead.
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
If a CONTEXT-MAP.md exists at the root, the repo has multiple contexts.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in
CONTEXT.md, call it out immediately.
"Your glossary defines 'cancellation' as X, but you seem to mean Y —
which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise
canonical term.
"You're saying 'account' — do you mean the Customer or the User?
Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with
specific scenarios.
### Cross-reference with code
When the user states how something works, check whether the code
agrees. If you find a contradiction, surface it.
### Update CONTEXT.md inline
When a term is resolved, update CONTEXT.md right there. Don't batch
these up — capture them as they happen.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. Hard to reverse
2. Surprising without context
3. The result of a real trade-off
```
## 真实的 CONTEXT.md 长什么样
Matt 自己的 [`course-video-manager`](https://github.com/mattpocock/course-video-manager/blob/main/CONTEXT.md) 仓库给了一份完整的 CONTEXT.md 例子。摘几条术语感受一下:
| 术语 | 定义 |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| **Course** | The primary domain entity: a structured collection of versions, sections, lessons, and videos |
| **Draft Version** | The single mutable CourseVersion that is currently being edited; always the latest by `createdAt` |
| **Published Version** | An immutable CourseVersion with a name and description, created by the Publish flow |
| **Ghost Lesson** | A lesson that exists in the database but not yet on the file system (`fsStatus = "ghost"`) |
| **Export Hash** | A SHA256 hash derived from a video's clip filenames, timestamps, clip order |
| **Unexported Video** | A video whose current Export Hash does not match any file on disk; blocks publishing |
| **Materialization Cascade** | The chain reaction when materializing a lesson inside a ghost course |
| **Clip** | A timestamped segment of source footage within a video |
| **Fractional Index** | A string-based ordering value that allows inserting items between existing items |
| **Purge** | The deliberate deletion of an Exported Video's `.mp4` file from disk |
注意几件事:
1. **每个术语都是动名词或专有名词**——不是「订单的状态」这种描述性短语
2. **每个定义都引用了其他术语**(Course → Version → Lesson → Video)形成本体网络
3. **代码字段直接出现**(`fsStatus = "ghost"`)——文档和代码 1:1 映射
4. **包含决策描述**("blocks publishing"、"chain reaction")——不只是名词,还有规则
写代码时 AI 看到这份文档,就会用 "Ghost Lesson" 而不是 "lesson with no file"。代码、对话、commit message 全部统一。
## ADR:什么时候创建
Skill 里有一段重要的克制:
> Only offer to create an ADR when all three are true:
>
> 1. **Hard to reverse** — the cost of changing your mind later is meaningful
> 2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
> 3. **The result of a real trade-off** — there were genuine alternatives
ADR(Architecture Decision Record)这个文化从 Michael Nygard 2011 年那篇博客起源,但很多团队用着用着就什么决策都写成 ADR——20 条 ADR 里 18 条是流水账。Matt 这个三角校验是个好工具:**只有同时满足三个条件,才值得 ADR**。否则就让它消化在 CONTEXT.md 里、消化在代码里。
## 怎么装、怎么用
**前置条件**:先跑过 `/setup-matt-pocock-skills`(它会问你 CONTEXT.md 放哪、ADR 目录放哪)。
**调用**:`/grill-with-docs`
**典型流程**:
1. 描述想做的事
2. `/grill-with-docs`
3. Claude **先扫描** CONTEXT.md 和 docs/adr/,把现有术语和决策装进上下文
4. 开始拷问,过程中:
* 你用的词和 CONTEXT.md 冲突 → 当场指出
* 你用了模糊词(比如「账户」既可能是 Customer 也可能是 User)→ 让你二选一并落进文档
* 你说的行为和现有代码矛盾 → 指出冲突
5. 决策达成时**同步更新 CONTEXT.md**(不积压、不批处理)
6. 关键的不可逆决策 → 询问是否生成 ADR
如果项目里还没有 CONTEXT.md 和 docs/adr/——它会**懒创建**:直到第一个术语需要写、第一个 ADR 需要建,才生成文件。不会一上来给你铺满空模板。
## 多上下文项目(CONTEXT-MAP.md)
如果项目大到一个 CONTEXT.md 装不下(比如 ordering 和 billing 是两个独立的 bounded context),可以在根目录放 `CONTEXT-MAP.md` 当总目录:
```
/
├── CONTEXT-MAP.md ← 总目录
├── docs/adr/ ← 系统级决策
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← 模块级决策
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
`/grill-with-docs` 会自动识别 CONTEXT-MAP.md 的存在并跳到对应子目录。这是 DDD 里 **bounded context** 概念的直接落地——每个上下文里的「订单」可能含义不同,分开维护避免污染。
## 和 grill-me 的差异
| 维度 | /grill-me | /grill-with-docs |
| --------------- | --------- | ------------------------------ |
| 提问能力 | ✅ | ✅(继承全部) |
| 项目术语校验 | ❌ | ✅ |
| 实时更新 CONTEXT.md | ❌ | ✅ |
| ADR 触发判断 | ❌ | ✅ |
| 适用阶段 | 早期想法、个人项目 | 有领域复杂度的真实项目 |
| 启动成本 | 0 | 需要 setup + 项目里有/愿意建 CONTEXT.md |
简单粗暴的判断:
* **个人脚本、写文章、教学课程** → `/grill-me`
* **需要长期维护的真项目** → `/grill-with-docs`
## 一个反直觉的好处:让 AI 学会「闭嘴」
CONTEXT.md 不只是给 AI 用的——它给**未来的 AI 会话**用的。每次新对话开始,Claude 读一遍 CONTEXT.md 就能瞬间进入项目语境,省掉一个长 onboarding。
更微妙的是:Matt 在演讲里说,加了 CONTEXT.md 之后他在 AI 的 *thinking trace* 里能看到——
> "It allows the AI to think in a less verbose way."(让 AI 思考时更简洁)
为什么?因为没有 CONTEXT.md 时 AI 在思考时要不停**自己定义术语**——「the user, by which I mean the person who ordered the item, hereinafter referred to as...」。有了 CONTEXT.md 它直接说 "Customer",思考链短了一大截,反应也快了。
**LLM 的 token 经济学决定了:缩短思考路径 = 输出更快、更准、更省钱**。CONTEXT.md 是这个效率的隐藏杠杆。
## 使用边界
| 场景 | 用 `/grill-me` 还是 `/grill-with-docs` | 原因 |
| --------------------------------------- | ----------------------------------- | ------------------------------ |
| 写一次性脚本、小工具、文章大纲 | `/grill-me` | 没有长期项目语言,也不需要维护 ADR |
| 新增真实产品功能 | `/grill-with-docs` | 需要检查现有术语、代码事实和架构约束 |
| 项目里已经有 CONTEXT.md 或 ADR | `/grill-with-docs` | 它能边拷问边更新文档,不让决策只停留在对话里 |
| 术语经常混乱,例如 User / Account / Customer 分不清 | `/grill-with-docs` | 它会把模糊词当场逼成 canonical term |
| 只是想快速确认实现方案 | `/grill-me` 或 Plan Mode | `/grill-with-docs` 的文档维护成本可能太高 |
第一次跑会大改 CONTEXT.md。如果项目里已经有手写的 CONTEXT.md,先 git stash 或者跑前让它 dry-run(你可以在 prompt 里加一句「先列出准备改的内容,不要直接写文件」)。
ADR 节制是真的节制。不要一兴奋就让它每个决策都生成 ADR,6 个月后你的 docs/adr/ 会堆满垃圾。Matt 的三条标准要严格执行。
CONTEXT.md 不要放实现细节。Skill 里有句话:「Don't couple CONTEXT.md to implementation details. Only include terms that are meaningful to domain experts.」 把 "PostgreSQL" 写进 CONTEXT.md 就是错——领域专家不关心数据库选型,那是 ADR 的事。
## 参考资源
下一篇:[to-PRD + to-Issues:从对话到可执行 ticket](/docs/notes/matt-pocock-skills/to-prd-and-issues)——拷问完了,怎么把对话凝固成可执行的工作单元。
# Improve Codebase Architecture:把 shallow 重构成 deep modules
## 失败模式:「AI 在烂代码库里乱走」
Matt 演讲里的第四个失败模式是个图景比喻:
> "Shallow modules in a codebase look like this——你有一堆零碎的 tiny blob,AI 必须穿过一堆模块、还要理解所有依赖才能改对。"
> "AI is really good at creating codebases like this. So you'll have a situation where AI doesn't understand what your code is doing. It will attempt to explore the code, but because it's poorly laid out, filled with shallow modules, it doesn't get to the right module in time, or doesn't understand all the dependencies."
这是 AI 编程特有的恶性循环:
```
AI 写代码倾向于产生 shallow 模块(小、多、互相依赖)
↓
代码库变得 shallow
↓
下次 AI 进来探索更难,更容易写错
↓
更多 shallow 模块被加进去
↓
代码库越来越烂,AI 越来越无能
```
要打破这个循环,**必须周期性手动反向重构**——把 shallow 模块合并成 deep 模块。这正是 `/improve-codebase-architecture` 干的事。
## 经典理论:Ousterhout 的 Deep Modules
John Ousterhout 是 Stanford CS 教授(也是 Tcl 语言、Raft 论文的作者)。他 2018 年那本《A Philosophy of Software Design》提出了一个简单但极有力的尺子:
**模块的「深度」 = 接口隐藏的复杂度**
| 类型 | 接口 | 实现 | 形象 |
| -------------- | -- | -- | --------- |
| **Deep(深)** | 简单 | 丰富 | 一个长方形:窄而深 |
| **Shallow(浅)** | 复杂 | 简单 | 一个长方形:宽而浅 |
理想模块是深的——使用者只需要看简短的接口,复杂度藏在内部。极端反例是 shallow 模块:接口几乎和实现一样复杂,等于没封装,使用者还不如直接看实现。
Ousterhout 的判断:**好代码库由少量深模块构成;坏代码库由大量浅模块构成**。这和「函数尽量小、文件尽量短、模块尽量多」的传统教条完全相反——他认为那种教条产生的恰恰是 shallow 模块。
## Matt 的延伸:Deletion Test
Matt 把 Ousterhout 的理论翻译成一个可操作的工程检验,他叫它 **deletion test**:
> **Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.**
人话:
* **删掉它,复杂度消失了** → 这模块本来就是 pass-through(中转),没在干活,砍掉
* **删掉它,复杂度散到 N 个调用方那里** → 它本来在帮你藏复杂度,是真的 deep,留下
这条测试的妙处是**双向**的——既能识别「该删的薄包装」,也能识别「该提取的公共逻辑」。如果你发现一段代码删掉之后复杂度会散到 5 个地方,说明这段代码值得提取成一个深模块。
## 关键术语(Matt 的精确定义)
`improve-codebase-architecture/SKILL.md` 里有一段 Glossary,要求**严格用这些词**——不要漂移到「component」「service」「API」「boundary」:
| 术语 | 定义 |
| ------------------ | -------------------------------------------------------- |
| **Module** | 任何有接口和实现的东西(function / class / package / slice) |
| **Interface** | 调用方必须知道的全部信息——类型、不变量、错误模式、顺序、配置(不只是函数签名) |
| **Implementation** | 模块内部的代码 |
| **Depth** | 接口处的杠杆。深 = 高杠杆,浅 = 接口几乎和实现一样复杂 |
| **Seam**(接缝) | 接口所在的位置——可以在不就地修改的情况下改变行为的地方。**用 "seam" 不要用 "boundary"** |
| **Adapter** | 在 seam 处实现某个接口的具体实现 |
| **Leverage** | 调用方从「深」中获得的好处 |
| **Locality** | 维护者从「深」中获得的好处——变更、bug、知识全部集中在一个地方 |
几条核心原则:
* **Deletion test**:见上文
* **The interface is the test surface**:测试只能通过接口跑——这是 deep module 可测的根基
* **One adapter = hypothetical seam. Two adapters = real seam.**:只有一个实现的接口是假接缝。**真接缝至少要有两个 adapter**
最后这条特别反直觉——很多团队会预先抽象一个接口「以备未来扩展」,但实际只有一个实现。Matt 的判断:**没用,删掉**。等真的有第二个实现时再抽。这跟 YAGNI 同源。
## Skill 工作流
### 1. Explore(探索)
Skill 先让 AI 读 `CONTEXT.md` 和 `docs/adr/`,然后用 `subagent_type=Explore` 派一个子 agent 去走代码库。
不用刚性启发,用**摩擦感**做信号:
> * Where does understanding one concept require bouncing between many small modules?
> * Where are modules **shallow** — interface nearly as complex as the implementation?
> * Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
> * Where do tightly-coupled modules leak across their seams?
> * Which parts of the codebase are untested, or hard to test through their current interface?
每发现一个可疑点,应用 deletion test:删掉它会让复杂度消失还是散开?答「散开」就是值得 deepen 的候选。
### 2. Present Candidates(陈述候选)
把候选用编号列表呈现:
```
1. Files: src/orders/parser.ts, src/orders/validator.ts, src/orders/normalizer.ts
Problem: 三个文件互相调用,理解 Order 入站需要在三处跳转
Solution: 合并为单一 OrderIntake 模块,对外只暴露 parse(raw) → ValidatedOrder
Benefits:
- Locality: Order 入站的所有逻辑、错误处理、bug 修复集中一处
- Leverage: 调用方从理解 3 个接口降为 1 个
- Tests: 只需测 parse() 的输入输出,不再需要 mock 内部协作
```
要求:
* 用 **CONTEXT.md 词汇** 谈领域("the Order intake module",不是 "the FooBarHandler")
* 用 **Glossary 词汇** 谈架构("seam"、"depth"、"locality")
* **不要立刻提议接口设计**——先让用户挑感兴趣的候选
如果某个候选会和已有 ADR 冲突——只在「冲突值得重新讨论 ADR」时才提,并明确标记:
> "contradicts ADR-0007 — but worth reopening because…"
不要把每个 ADR 禁止过的重构都翻出来。
### 3. Grilling Loop(拷问循环)
用户挑了一个候选之后,drop 进 grilling 模式(继承自 [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)):
* 走 design tree——约束、依赖、深化后的模块形状、接缝后面藏什么、哪些测试能存活
* **副作用即时发生**:
* 给深化模块起了 CONTEXT.md 里没有的名字 → 立刻加进 CONTEXT.md
* 拷问中锐化了某个模糊术语 → 立刻更新 CONTEXT.md
* 用户拒绝候选且理由是 load-bearing(关键的、未来探索者需要知道的)→ 提议生成 ADR
* 想探索深化模块的多种接口设计 → 跳到 `INTERFACE-DESIGN.md` 单独流程
文档维护和架构改造是**同一个对话里发生的**——不分两轮。
## 真实案例:Mejba Ahmed 的实践
第三方开发者 Mejba Ahmed 写了篇 [《Deep Modules: The Claude Code Skill Saving My Codebase》](https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules) 详细记录了他用这个 skill 的体感。要点:
* 他原来一个项目里有 50+ 个文件,每个文件 100 行以下——**典型 shallow 库**
* `/improve-codebase-architecture` 跑出 8 个 deepening candidates
* 他选了 3 个深化(两个数据处理模块合并、一个工具集合并)
* 结果:文件数从 50+ 降到 30+,但**总代码量基本不变**——复杂度被压进了少量深模块
* 后续 Claude 在这个库里改代码的命中率显著提高(他说「from 60% to 90%」,未严格测量但体感强)
Mejba 也提醒了一条:**不要一次深化 8 个**。一次只挑一个,做完跑测试 + commit + observe,再挑下一个。否则一次做完没法回滚。
## 怎么装、怎么用
```bash
npx skills@latest add mattpocock/skills
```
勾选 `improve-codebase-architecture` + `setup-matt-pocock-skills`。
**调用**:`/improve-codebase-architecture`
**推荐节奏**:
* **每周或每个 sprint 末**跑一次
* 或**做完一波密集开发后**跑一次(AI 高频写代码后特别容易堆 shallow 模块)
* **不要在赶工时跑**——它会建议大改动,赶工时没时间消化
**典型流程**:
1. `/improve-codebase-architecture`
2. AI 探索 + 列出 N 个 candidates(带 deletion test 论证)
3. 你挑一个最有感觉的
4. drop 进 grilling 循环对齐设计
5. AI 实施重构(建议配合 [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 一起跑——重构必须有测试保护)
6. commit + 观察
7. 一周后再来一次
## 这个 Skill 为什么是 Matt 工作流的「闭环」
回到 Matt 的工作流图:
```
/grill-me → /to-prd → /to-issues → /tdd → /improve-codebase-architecture → 回到 /grill-me
```
注意它**循环**回到了起点。`/improve-codebase-architecture` 不是一次性工具,是**周期性维护**——因为:
1. AI 持续往代码库里加 shallow 模块(这是它的默认偏向,写多了就堆)
2. 业务持续演化,老接缝会过时
3. CONTEXT.md 里的术语持续锐化,老命名会跟不上
**每跑一次这个 skill,代码库的 AI 友好度就刷新一次**。这是和 LLM 长期协作的代码库**唯一**能保持健康的方式——不刷新,三个月后 AI 在你的代码库里就废了。
## 这套思维比 Skill 本身更值钱
即使你完全不装 `/improve-codebase-architecture`,光是把以下三件事记住,PR review 质量就会提一档:
1. **deletion test**:每看到一个新模块,问自己「删了它复杂度会消失还是散开?」
2. **真接缝至少两个 adapter**:单实现接口 = 假抽象,删
3. **接口就是测试面**:测不动 = 接口设计有问题
这三条不需要 AI、不需要 skill——是工程审美的硬通货。Matt 把它们包装成 skill 是为了批量化执行,但**真正的杠杆是这三条原则本身**。
## 注意事项
**不要过度深化**。Ousterhout 自己说过 deep module 是**目标**而不是教条——一个大型 Util 类把所有功能塞进去也不是 deep module,是上帝模块。判断标准是「接口简单 + 实现内聚」,两个都得满足。
**deepening 必须有测试保护**。结构改动是高风险操作,没有测试就敢重构 = 等着背锅。如果当前没有测试,先去 [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 给关键路径补测试再回来。
**ADR 决策不要一时兴起**。grilling 时你拒绝某个候选,AI 会顺手提议生成 ADR——只在那个理由真的「未来人需要知道」时才接受。否则 docs/adr/ 会堆满流水账。
**有些代码就是 shallow 也没关系**。一个 logger wrapper、一个常量文件、一个一次性脚本——它们是浅的、没问题。这个 skill 找的是**那些假装在帮你抽象但其实在添乱的浅模块**。
## 参考资源
***
## 系列结语
到这里 6 篇都看完了。回顾整个工作流:
```
/grill-me 或 /grill-with-docs ← 谈清楚要做什么
↓
/to-prd ← 凝固成 PRD
↓
/to-issues ← 切成 vertical slice
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture ← 周期性深化
↓
回到 /grill-me
```
这套流程的精神可以浓缩成一句话:
> **AI 是地面上的战术兵,你是战略层。把「定义问题」「拆问题」「验问题」这三件事拿回来自己做,把「写代码」交给 AI——这才是 AI 时代工程师的真定位。**
Matt 这套 skill 不是终极答案,是当前阶段的最佳实践。三个月后可能有更好的,但**精神层面的东西不会变**:好代码库永远比烂代码库重要,软件基本功永远值钱。
回到 [概览](/docs/notes/matt-pocock-skills/overview),或挑一个最有体感的 skill 装上试试。
# 软件基本功比以往任何时候都更重要:Matt Pocock 的 Claude Code 技能集
## 核心结论
**Matt Pocock 的 Claude Code skills 不是一套“让 AI 多写代码”的工具,而是一套把经典软件工程基本功翻译成 LLM 可执行流程的工作流**。它的核心判断是:AI 时代代码生成更快了,但坏代码的返工成本也更高,所以更需要 design concept、ubiquitous language、TDD、deep modules 这些老方法。
如果你只想理解这套 skills 的使用顺序,可以记住这条链路:`/grill-me` 先澄清需求,`/to-prd` 固化成 PRD,`/to-issues` 切成 vertical slice,`/tdd` 小步实现,最后用 `/improve-codebase-architecture` 周期性重构模块边界。
## 一个被 specs-to-code 浪潮淹没却仍冷静的人
2026 年是 AI 编程「specs-to-code」叙事的高光年——写规约、跑编译器、不看代码、再写规约、再跑编译器。社区里冒出来的口号是「**code is cheap**」(代码很便宜),意思是:反正 AI 一秒能再生成一万行,要它干嘛在意。
Matt Pocock 是少数公开站出来反驳这件事的人。他不否认 AI 写代码很猛,但他在自己开的课《Claude Code for Real Engineers》里实测了 specs-to-code,结论很扎心:**每跑一次,代码反而越来越烂**。这正是 Pragmatic Programmer 里早就讲过的「software entropy」——软件熵增。
于是他做了两件事:
1. 把这个观察打包成一场 18 分钟的演讲:*Software Fundamentals Matter More Than Ever*。
2. 把对应的解药打包成一个 GitHub 仓库:[`mattpocock/skills`](https://github.com/mattpocock/skills) ——「Skills for Real Engineers. Straight from my .claude directory.」
仓库 2026 年 2 月 3 日上线,4 个月之内冲到 **61.1k star、5.3k fork**,是同期 AI 编程类仓库里增长最快的几个之一。
***
## Matt Pocock 是谁
如果你写过 TypeScript,大概率刷到过他。他是这几年中文圈和英文圈最高产的 TypeScript 教育者之一:
* **TotalTypeScript.com** 创办人,一系列付费课程在英文圈流传度极高
* **aihero.dev** Newsletter 6 万+ 订阅,主题从 TS 转到了 AI Coding
* Twitter [`@mattpocockuk`](https://twitter.com/mattpocockuk)、YouTube `@mattpocockuk` 上有大量短视频教程
* 不是 OpenAI/Anthropic 的人,纯独立开发者+教育者背景
他的人设很清晰:**资深工程师视角看 AI Coding**。不喊「AGI 来了」,也不喊「程序员要失业了」。喊的是「老一辈软件工程师那些招还很有用,只是需要翻译成 LLM 能执行的形式」。
***
## 核心论点:代码不便宜
整场演讲只有一个论点,每个 skill 都是它的注脚:
> 如果你的代码库结构很烂,AI 在烂代码库里也只会写出烂代码。所以**好代码库比以往任何时候都重要,软件基本功比以往任何时候都重要**。
Matt 用了一个军事类比把人和 AI 的角色定位讲得很直白:
战略层做什么?设计概念、统一语言、模块边界——这三件事都是「定义问题」而非「写代码」,**恰好也是 LLM 最不擅长替你做的**。
***
## 五个失败模式 → 五本老书 → 五个 Skill
Matt 演讲里把整套方法论压缩成一张映射表。你每碰到一个失败模式,他都给你指回 20 年前就解决了的经典理论,再给你一个 Markdown 格式的 Skill 文件:
| # | AI 编程失败模式 | 经典理论与出处 | 对应 Skill |
| - | -------------- | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
| 1 | AI 没做你想要的东西 | *The Design of Design*(Brooks)—— design concept、design tree | [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me) |
| 2 | AI 用一堆冗长术语跟你说话 | *Domain-Driven Design*(Evans)—— ubiquitous language | [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs) |
| 3 | AI 做对了但跑不起来 | *The Pragmatic Programmer*(Hunt & Thomas)—— "rate of feedback is your speed limit" | [`/tdd`](/docs/notes/matt-pocock-skills/tdd) |
| 4 | AI 在烂代码库里到处乱走 | *A Philosophy of Software Design*(Ousterhout)—— deep modules、deletion test | [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture) |
| 5 | 你脑子跟不上 AI 产出 | Kent Beck —— invest in design every day | "design the interface, delegate the implementation" |
第五条在 repo 里没有单独的 skill(曾有 `design-an-interface` 但已废弃),它的精神被吸收进了 [`/to-prd`](/docs/notes/matt-pocock-skills/to-prd-and-issues) 和 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture)——两者都强迫你在写代码前先思考模块接口。
***
## 真正每天用的 5 个 Skill
演讲是哲学骨架,Matt 后来在 aihero.dev 上发了一篇《5 Agent Skills I Use Every Day》,把骨架翻译成日常工作流。这 5 个就是这个系列后续要逐个拆解的对象:
```
/grill-me ← 先和 AI 谈清楚要做什么
↓
/to-prd ← 把对话凝固成 PRD
↓
/to-issues ← 把 PRD 切成可独立领取的 vertical slice
↓
/tdd ← 每个 slice 用红绿重构跑通
↓
/improve-codebase-architecture ← 周期性检查,把 shallow 模块改成 deep
```
这 5 个 skill 串起来就是 Matt 自己的完整研发流程。每一步对应的失败模式见上一节的表。
每篇详细拆解(这个系列的后续页面):
* [Setup Matt Pocock Skills:先把项目规则写清楚](/docs/notes/matt-pocock-skills/setup-matt-pocock-skills)
* [Grill Me:让 AI 拷问你需求](/docs/notes/matt-pocock-skills/grill-me)
* [Grill With Docs:维护项目语言和 ADR](/docs/notes/matt-pocock-skills/grill-with-docs)
* [to-PRD + to-Issues:从对话到可执行 ticket](/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [Diagnose 与 Triage:先建立反馈回路,再决定交给谁](/docs/notes/matt-pocock-skills/diagnose-and-triage)
* [TDD:用红绿重构强迫 AI 走小步](/docs/notes/matt-pocock-skills/tdd)
* [Zoom Out:当你迷路时,让 AI 先画地图](/docs/notes/matt-pocock-skills/zoom-out)
* [Prototype:用可丢弃代码回答一个设计问题](/docs/notes/matt-pocock-skills/prototype)
* [Improve Codebase Architecture:把 shallow 重构成 deep modules](/docs/notes/matt-pocock-skills/improve-codebase-architecture)
* [其他 Skills:压缩沟通、交接、教学、写 Skill 与安全护栏](/docs/notes/matt-pocock-skills/productivity-and-misc)
***
## 怎么装
仓库 README 里给了一行命令的安装方式:
```bash
npx skills@latest add mattpocock/skills
```
这条命令会:
1. 让你勾选要装哪些 skill
2. 让你勾选要装到哪些 agent(Claude Code、Codex、Cursor 等都支持)
3. 把对应的 SKILL.md 文件放进 `.claude/skills/`(或 agent 对应的目录)
**强烈建议同时勾选 `/setup-matt-pocock-skills`**——这是一个一次性配置 skill,会问你三个问题:
* **Issue tracker** 用什么?(GitHub / GitLab / 本地 markdown / 其他)
* **Triage label** 用什么词?(needs-triage 还是其他)
* **Domain doc** 放哪?(CONTEXT.md / ADR 路径)
跑一次 `/setup-matt-pocock-skills`,它会写到你项目根目录的 `AGENTS.md` 或 `CLAUDE.md`,之后所有 engineering 类 skill(to-prd、to-issues、triage、tdd 等)都会自动读取这份配置。这一步省了,后面每个 skill 都会反复问你同样的问题。
如果你只想试 `/grill-me`(最轻量、纯 productivity 类),可以跳过 setup,因为它不依赖 issue tracker。
***
## 这套 skill 和 BMAD / Spec-Kit / GSD 的区别
如果你已经在用 [BMAD](/docs/notes/speckit/concept)、Spec-Kit、GSD 这类 spec-driven 框架,可能会问:「为啥还需要 Matt 这套?」
Matt 在 README 里写得很直接:
**核心差异**:
* BMAD/Spec-Kit/GSD 是**框架**,规定了从 spec 到 code 的完整流水线,你要按它的流程走
* Matt 这套是**散件**,每个 skill 一个 markdown 文件,几行到几十行,你随时拆开改
例子:`grill-me` 的实际全文只有这么短——
```markdown
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
整个 skill 7 行。但就是这 7 行让 Claude 在做决策前会问你 20、50 甚至 100 个问题。这种**用极少文本撬动大行为变化**的设计哲学,是这套 skill 流行的根本原因。
***
## 这个系列怎么读
如果你之前没接触过 Matt 这套,**建议按 meta.json 里的顺序顺读**:
1. **Overview**(你现在在的这篇)—— 了解全貌
2. **Setup Matt Pocock Skills** —— 先把 issue tracker、标签和领域文档位置写成项目契约
3. **Grill Me** —— 单装一个先体验,门槛最低
4. **Grill With Docs** —— grill-me 的进阶版,开始引入 CONTEXT.md
5. **to-PRD + to-Issues** —— 把对话变成可执行 ticket
6. **Diagnose 与 Triage** —— 处理 bug 和 issue 状态流转
7. **TDD** —— Matt 自己说「我用过的最稳定提升 agent 输出质量的方法」
8. **Zoom Out** —— 陌生代码先画地图
9. **Prototype** —— 用可丢弃代码回答设计问题
10. **Improve Codebase Architecture** —— 周期性维护,让 AI 长期可用
11. **其他 Skills** —— caveman、handoff、teach、write-a-skill 和几组安全/课程工具
如果你已经在用 Claude Code 写真实项目,**直接跳到 [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me) + [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 这两篇**,体感最强。
如果你做的是教学或写作,**只读 [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me) 一篇**就够了——它是个通用的「设计对话」工具,不限于代码。
***
## 我的使用建议
我自己装了这套 skill 之后,最大的几个体感变化:
**第一**:不再急着开始写代码。以前 AI 接到「帮我加个登录」就开始铺 500 行,现在 `/grill-me` 会先问你 20 个问题——「要不要记住设备」「session 多久过期」「失败几次锁账号」。30 分钟后再让它写,省掉的是后续两小时的返工。
**第二**:CLAUDE.md 不再臃肿。以前 CLAUDE.md 里写了一堆「请先理解需求再写代码」「不要过度抽象」之类的禁令,但 Claude 该犯还是犯。换成 Matt 这套之后,CLAUDE.md 只放领域知识(设计系统、组件规范、部署),通用方法论交给 skill。两边职责清晰。
**第三**:deep module 思维比 skill 本身更值钱。即使你不装 `/improve-codebase-architecture`,光是读完它的 SKILL.md 里那条「**deletion test**」(如果删掉这个模块复杂度消失了,说明它本来就是 pass-through)就已经能让你在 PR review 时多看一眼。
**注意代价**:
* 装齐 5 个 skill 后,AI 会问话变多。习惯了「一句话生成 500 行」的人会觉得烦
* `/tdd` 严格执行后,简单脚本也会被它要求先写测试,对探索型代码不友好——可以告诉它「这次跳过 TDD」
* `/grill-with-docs` 会主动改你的 CONTEXT.md,第一次跑前最好让它进 dry-run
***
## 参考资源
**演讲里引用的 5 本书**(按出现顺序):
* *A Philosophy of Software Design* — John Ousterhout(complexity 定义、deep modules)
* *The Pragmatic Programmer* — David Thomas & Andrew Hunt(software entropy、outrunning headlights)
* *The Design of Design* — Frederick P. Brooks(design concept、design tree)
* *Domain-Driven Design* — Eric Evans(ubiquitous language)
* *Test-Driven Development* — Kent Beck(invest in design every day)
每本都是 20 年以上的老书。Matt 演讲里多次重复一句话:「**Go on Amazon, get it.**」——这句话本身就是这场演讲的彩蛋。
# 其他 Skills:压缩沟通、交接、教学、写 Skill 与安全护栏
## 为什么不逐个展开
Matt 的 README 把 skills 分成三类:
* Engineering:每天写真实代码用
* Productivity:通用工作流工具
* Misc:他自己留着备用的小工具
前面几篇已经覆盖了主线 engineering skill。这一篇把剩下的 productivity 和 misc 合在一起讲,因为它们多数不是完整研发流程,而是**在特定场景下很好用的小开关**。
如果你只装 5 个,我仍然建议优先装:
* [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me)
* [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)
* [`/to-prd` + `/to-issues`](/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [`/tdd`](/docs/notes/matt-pocock-skills/tdd)
* [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture)
但如果你已经把主线跑起来,下面这些会让日常体验更顺。
## Productivity Skills
### caveman:极限压缩沟通
`/caveman` 是一个「少 token 模式」。它要求 agent 去掉寒暄、填充词、过度解释和模糊缓冲,只保留技术信息。
它适合:
* 你已经在高频迭代,不想读长回复
* debug 时只要事实、原因、下一步
* 长上下文快满了,需要压缩输出
* 你想强制 AI 少说漂亮话
它不是让 AI 变粗鲁,而是让它用更短的语法保留完整技术精度。注意它会持续生效,直到你说退出。
我会把它当作一个临时档位,而不是长期默认。对于高风险操作、安全警告、复杂多步骤指令,太短反而容易误读。
### handoff:把当前会话交给下一个 agent
`/handoff` 的目标是把当前会话压缩成一个交接文档,并保存到系统临时目录,而不是污染当前 workspace。
它会包含:
* 当前目标
* 已做决策
* 关键路径和文件
* 剩余任务
* 建议下个 agent 调用哪些 skill
* 敏感信息脱敏
它特别适合长任务中断、上下文快满、或你想换一个 agent 继续做的时候。
关键点是:不要复制已经存在于 PRD、issue、ADR、commit、diff 里的内容,只引用路径或 URL。交接文档的价值是**把散在对话里的状态补齐**,不是再造一份项目文档。
### teach:把当前目录变成学习工作区
`/teach` 是这组里最重的一个。它把当前目录当作一个长期学习 workspace,维护:
* `MISSION.md`:你为什么要学这个主题
* `RESOURCES.md`:高质量资源列表
* `learning-records/*.md`:学习记录,类似 ADR
* `lessons/*.html`:每次一节的互动课程
* `reference/*.html`:速查资料
* `NOTES.md`:教学偏好和工作笔记
它的亮点是把学习看成长期系统,而不是一次问答。尤其强调:
* mission 先行:为什么学,比学什么更重要
* retrieval practice:用回忆练习建立长期记忆
* spacing / interleaving:不要被短期流畅感骗了
* 高信任资源:先找资料,不凭模型记忆硬讲
如果你只是问「解释一下 X」,不需要它;如果你想连续几周学一个主题,它就很合适。
### write-a-skill:写新 skill 的脚手架
`/write-a-skill` 是 Matt 对 skill 结构本身的抽象。
它要求一个 skill 至少有:
```text
skill-name/
├── SKILL.md
├── REFERENCE.md
├── EXAMPLES.md
└── scripts/
```
当然后三个不是必需,只有内容太长、示例有价值、或操作可脚本化时才加。
它最重要的判断是:`description` 是 agent 决定是否加载 skill 时唯一先看到的信息。因此 description 不能写成「帮助处理文档」这种空话,必须说明:
* 它提供什么能力
* 什么时候触发
* 触发词或上下文是什么
这和我自己写 skill 的经验一致:很多 skill 失效不是因为正文写得差,而是 description 写得太泛,agent 根本不知道该加载它。
## Misc Skills
### git-guardrails-claude-code:拦危险 git 命令
这个 skill 会给 Claude Code 配一个 `PreToolUse` hook,在执行 Bash 前拦截危险 git 命令。
默认会挡:
* `git push`
* `git reset --hard`
* `git clean -f` / `git clean -fd`
* `git branch -D`
* `git checkout .` / `git restore .`
它的价值很直接:防止 agent 在你没授权时推送、硬重置、清掉未跟踪文件。
如果你经常让 AI 在真实仓库里工作,这个 skill 很值得装。它不是不信任 AI,而是把高破坏性操作放到工具层拦截,而不是靠 prompt 祈祷。
### setup-pre-commit:给项目加提交前检查
`/setup-pre-commit` 会设置:
* Husky pre-commit hook
* lint-staged + Prettier
* typecheck
* test
它会先检测包管理器,再按项目已有 script 决定 pre-commit 里该跑什么。没有 `typecheck` 或 `test` 时不会硬造,而是省略并告诉你。
这个 skill 的价值不在配置本身,而在 Matt 的质量观:**不要只让 AI 自己说代码没问题,要让它过确定性检查**。
### migrate-to-shoehorn:测试里少写 `as`
这是一个很 Total TypeScript 风格的小工具。它把测试里的 TypeScript `as` 类型断言迁移到 `@total-typescript/shoehorn`。
典型替换:
| 旧写法 | 新写法 | 场景 |
| --------------------------- | ------------------ | -------------- |
| `obj as Request` | `fromPartial(obj)` | 测试里只关心大对象的几个字段 |
| `obj as unknown as Request` | `fromAny(obj)` | 故意传错类型测错误路径 |
| 完整对象假数据 | `fromExact(obj)` | 需要强制完整形状 |
它明确只用于测试代码,不用于生产代码。
这个 skill 很窄,但很符合 Matt 的工程品味:不要为了类型系统在测试里造 20 个无意义字段,也不要用裸 `as` 把类型安全完全关掉。
### scaffold-exercises:给课程仓库生成练习目录
这个 skill 明显来自 Matt 自己做课程的工作流。它会按规范创建:
```text
exercises/
└── 05-memory-skill-building/
└── 05.02-short-term-memory/
├── explainer/
├── problem/
└── solution/
```
每个子目录至少有非空 `readme.md`,必要时有 `main.ts`,并且要通过 `pnpm ai-hero-cli internal lint`。
它对大多数工程项目没用,但对课程、训练营、练习仓库非常实用。更重要的是,它展示了一个好 skill 的特征:**把重复、机械、容易漏细节的格式工作交给 agent**。
## 不建议现在写进主线的目录
上游仓库里还有 `deprecated/`、`in-progress/`、`personal/`。
我建议暂时不要把它们写成正式使用指南:
| 目录 | 为什么不放主线 |
| -------------- | -------------------------- |
| `deprecated/` | 已废弃,容易误导读者继续采用旧流程 |
| `in-progress/` | 还在实验,行为和命名都可能变 |
| `personal/` | 更像 Matt 自己的私人工作区,不一定适合通用读者 |
如果以后要写,可以单独做一篇「Matt Pocock skills 仓库考古」,而不是混在稳定推荐里。
## 这组小工具的共同点
这些 skill 看起来很散,但背后有同一个原则:
> 把 agent 容易漂移的事情,变成小而明确的工作模式。
* `caveman` 防止沟通漂移
* `handoff` 防止上下文丢失
* `teach` 防止学习变成一次性问答
* `write-a-skill` 防止 skill 结构随手写
* `git-guardrails` 防止危险命令靠自觉
* `setup-pre-commit` 防止质量检查靠 AI 自述
* `migrate-to-shoehorn` 防止测试类型断言失控
* `scaffold-exercises` 防止课程结构手工漏项
这也是 Matt 这套 repo 最值得学习的地方:skill 不需要宏大。一个高频小偏差,如果能被 20 行指令稳定纠正,就值得写成 skill。
## 参考资源
# Prototype:用可丢弃代码回答一个设计问题
## 原型不是「先随便写一个」
`/prototype` 的第一句定义很重要:
> 原型是用来回答一个问题的可丢弃代码。
这句话把原型和「偷懒版实现」分开了。原型不是生产代码的前身,不是以后慢慢改成正式版的半成品。它从第一天开始就应该被标记为 throwaway。
所以 `/prototype` 的关键不是写得快,而是先问清楚:
> 这个原型到底要回答什么问题?
## 两条分支
Matt 把原型分成两类,输出完全不同。
| 要回答的问题 | 分支 | 输出 |
| --------------- | --------------- | ---------------- |
| 逻辑、状态机、数据模型是否合理 | Logic prototype | 一个可运行的终端小程序 |
| 这个界面应该长什么样 | UI prototype | 一个路由里多套可切换 UI 方案 |
这点很实用。很多团队说「做个 prototype」,但没说清楚是想验证交互外观,还是验证状态流转。两者需要的东西完全不同。
## Logic prototype:把状态摊在终端里
如果问题是「这个状态机对不对」「这个业务规则能不能跑通」,原型应该是一个很小的命令行程序。
它的特点:
* 内存态,不依赖真实数据库
* 一个命令启动
* 每个操作后打印完整相关状态
* 覆盖那些纸面上难以推演的分支
* 不写测试,不做异常兜底,不抽象成框架
例子:你要设计订阅状态流。
不要直接改生产代码。先写一个 `subscription-prototype.ts`,让用户可以在终端里选择:
```text
1. start trial
2. pay
3. cancel
4. expire
5. refund
6. print state
```
每按一步,打印当前 entitlement、trial quota、paid state、next renewal。你会很快发现有些状态组合根本没想清楚。
这类原型的价值是:**让抽象规则变成可操作对象**。
## UI prototype:在一个路由里放多套激进方案
如果问题是「界面应该怎么设计」,原型应该生成几套差异足够大的 UI,而不是把同一个方案微调三次。
`/prototype` 的 UI 分支要求:
* 在一个路由里放多个 variation
* 用 URL search param 或底部浮动切换条切换
* 方案之间要有明显差异
* 原型代码靠近未来真实页面,但命名要清楚表示 prototype
* 不要过早接真实数据和持久化
这和普通 AI 生成 UI 的区别在于:它不是让 AI 一次给「最佳方案」,而是让你用真实浏览器比较几种方向。
比如一个 dashboard 空状态,不要只让 AI 改文案。可以让它做:
* A:表格式、密度高,强调下一步操作
* B:任务导向,左侧 checklist + 右侧预览
* C:引导式,突出一个主 CTA 和历史示例
然后你在同一路由切换,而不是在聊天里看三张截图脑补。
## 所有原型都必须可删除
`/prototype` 的通用规则里,最重要的是「可删除」:
* 文件名或路径要表明这是 prototype
* 不要默认接生产数据库
* 不要写通用抽象
* 不要做过度错误处理
* 结束后要删除,或者把学到的结论吸收到正式代码
如果一个原型不能删,它就已经变成生产代码负债。
这点在 AI 编程里尤其重要。AI 很擅长把 prototype 写得「看起来能用」,然后人类懒得删,最后项目里多出一堆没人敢碰的临时代码。
## 原型结束后要留下什么
原型代码不值得保留,但答案值得保留。
Matt 建议把下面几件事写到一个持久位置:
* 原型要回答的问题
* 观察到的结论
* 选择了哪个方向
* 放弃了哪些方向
* 如果需要,转成 ADR、issue、PRD 或 commit message
也就是说,`/prototype` 的产物不是代码,而是**决策**。
## 什么时候不该用
不要把 `/prototype` 用在这些场景:
* 需求已经明确,只需要实现
* bug 已经有复现,应该用 [`/diagnose`](/docs/notes/matt-pocock-skills/diagnose-and-triage)
* 重构方向已经清楚,应该用 [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 保护后实施
* UI 只是小 polish,不值得做多套方案
* 你没有时间删除或吸收原型
原型的成本不在写,而在收尾。没有收尾,就不要开。
## 一个好用的提示词
可以这样调用:
```text
/prototype
我想验证这个 checkout 状态机是否合理。请走 logic 分支。
只做可丢弃终端原型,不接真实 DB。
每次操作后打印完整状态。
```
或者:
```text
/prototype
我想比较项目详情页的 3 种信息架构。请走 UI 分支。
放在现有路由体系下的 prototype route,提供底部切换条。
不要动生产组件。
```
这里最重要的是明确「要回答的问题」。只要这个问题清楚,原型就不容易跑偏。
## 和 Grill Me 的关系
[`/grill-me`](/docs/notes/matt-pocock-skills/grill-me) 适合通过提问收束决策;`/prototype` 适合通过试玩收束决策。
有些问题靠问就能解决,比如「匿名评论要不要审核」。有些问题必须摸一下,比如「这个拖拽排序状态机到底会不会难用」。后者就该 prototype。
所以我会把它放在工作流的一个分叉位置:
```text
想法模糊
↓
/grill-me
↓
如果仍然需要体验或验证
↓
/prototype
↓
保留结论,删除原型
↓
/to-prd 或 /tdd
```
## 参考资源
下一篇:[Improve Codebase Architecture:把 shallow 重构成 deep modules](/docs/notes/matt-pocock-skills/improve-codebase-architecture)。
# Setup Matt Pocock Skills:先把项目规则写清楚
## 这个 Skill 解决的不是安装问题
`/setup-matt-pocock-skills` 容易被误解成「装完之后跑一下的初始化命令」。实际上它更像一个**项目契约生成器**:告诉后续 skill 这个仓库怎么追踪任务、怎么标记 issue、在哪里读取领域语言和架构决策。
Matt 在 README 的 Quickstart 里特别提醒:安装时要选中 `/setup-matt-pocock-skills`,然后在 agent 里运行它。原因很简单:`to-prd`、`to-issues`、`triage`、`diagnose`、`tdd`、`improve-codebase-architecture`、`zoom-out` 都需要同一批项目上下文。如果每个 skill 都临时问一遍,流程会变得很碎。
它做的不是「配置 Claude 偏好」,而是回答三个工程问题:
| 问题 | 它要写清楚什么 | 后续谁会用 |
| ---------------- | --------------------------------------------------------- | ----------------------------------------------------------------------------- |
| Issue tracker 在哪 | GitHub、GitLab、本地 markdown,还是其他系统 | `to-prd`、`to-issues`、`triage` |
| Triage 标签怎么映射 | `needs-triage`、`needs-info`、`ready-for-agent` 等角色对应哪些真实标签 | `triage` |
| 领域文档在哪里 | 单一 `CONTEXT.md`,还是多上下文 `CONTEXT-MAP.md` + 分区 ADR | `grill-with-docs`、`diagnose`、`tdd`、`zoom-out`、`improve-codebase-architecture` |
## 它为什么重要
这套 skills 的核心思路是「小而可组合」。小的代价是:它们不想自己接管整个项目流程,所以必须知道你项目里的真实约定。
举例:`/to-issues` 要创建 issue。没有 setup,它不知道应该:
* 调 `gh issue create`
* 调 `glab issue create`
* 写到 `.scratch//`
* 还是给你生成一段 Linear/Jira 可复制文本
再比如 `/triage` 要把 issue 移到 `ready-for-agent`。如果你的仓库里真实标签叫 `ai:ready`,而 skill 自己创建了一个新标签 `ready-for-agent`,issue tracker 会立刻变脏。
所以 `/setup-matt-pocock-skills` 的价值不是自动化,而是**把隐含约定外显化**。
## 它会读哪些东西
这个 skill 开始时会先探索仓库,而不是假设:
* `git remote -v` 和 `.git/config`:判断是不是 GitHub/GitLab 项目
* 根目录 `AGENTS.md` / `CLAUDE.md`:看是否已有 `## Agent skills` 区块
* 根目录 `CONTEXT.md` / `CONTEXT-MAP.md`:判断领域语言文档形态
* `docs/adr/` 和 `src/*/docs/adr/`:判断 ADR 是全局还是模块级
* `docs/agents/`:看是否已经跑过 setup
* `.scratch/`:判断是否已有本地 markdown issue 约定
这符合 Matt 整套工作流的风格:**先看项目真实状态,再写规则**。
## 三个决策
### 1. Issue tracker
这是后续工作单元落地的位置。
默认倾向是 GitHub,因为这套 skill 最早围绕 GitHub Issues 设计。但它已经把 GitLab 和本地 markdown 也当作一等选择:
| 选择 | 适合什么场景 |
| -------------- | ---------------------------------- |
| GitHub | 开源项目、GitHub issue workflow 已经存在 |
| GitLab | 公司项目在 GitLab,习惯用 `glab` |
| Local markdown | 个人项目、临时探索、没有远程 issue tracker |
| Other | Jira、Linear、飞书、多维表格等,需要用文字记录你的真实流程 |
重点不是选哪个,而是选**团队真的在用的那个**。写错了,后续 skill 会在错误系统里创建任务。
### 2. Triage label vocabulary
`/triage` 内部使用 5 个状态角色:
| 角色 | 含义 |
| ----------------- | ------------------- |
| `needs-triage` | 等维护者判断 |
| `needs-info` | 等报告者补信息 |
| `ready-for-agent` | 已经清楚到可以交给 AFK agent |
| `ready-for-human` | 需要人类判断或实现 |
| `wontfix` | 不处理 |
setup 会问你这些角色对应的真实标签名。如果项目里没有既有标签,用默认名就行;如果已经有自己的命名体系,应该在这里映射,而不是让 skill 新造一套。
### 3. Domain docs
这是 Matt 这套 skill 和普通 prompt 最大的区别:它不是只看当前对话,还会读项目里的**领域语言**和**架构决策**。
最简单的形态:
```text
/
├── CONTEXT.md
└── docs/
└── adr/
```
大型 monorepo 可以用多上下文:
```text
/
├── CONTEXT-MAP.md
├── apps/
│ └── web/
│ ├── CONTEXT.md
│ └── docs/adr/
└── services/
└── billing/
├── CONTEXT.md
└── docs/adr/
```
setup 不是强迫你现在写完所有文档,而是告诉后续 skill:应该去哪里找、找不到时应该如何创建。
## 它会写什么
最终会有两类输出。
第一类是 `AGENTS.md` 或 `CLAUDE.md` 里的 `## Agent skills` 区块:
```markdown
## Agent skills
### Issue tracker
...
### Triage labels
...
### Domain docs
...
```
第二类是 `docs/agents/` 下的三份说明:
| 文件 | 内容 |
| ------------------------------ | --------------------------------------- |
| `docs/agents/issue-tracker.md` | issue 系统、命令、创建/更新约定 |
| `docs/agents/triage-labels.md` | canonical role 到真实标签的映射 |
| `docs/agents/domain.md` | `CONTEXT.md`、`CONTEXT-MAP.md`、ADR 的读取规则 |
注意它会优先编辑已有的 `CLAUDE.md`;没有 `CLAUDE.md` 才考虑 `AGENTS.md`。这体现了一个很重要的克制:**不要在项目里制造两份互相竞争的 agent 规则入口**。
## 我建议怎么用
第一次装 Matt 这套 skill 时,顺序应该是:
1. 安装:`npx skills@latest add mattpocock/skills`
2. 选中 `/setup-matt-pocock-skills`
3. 运行 `/setup-matt-pocock-skills`
4. 按真实项目状态回答 issue tracker、标签、领域文档三个问题
5. 看它生成的 `## Agent skills` 和 `docs/agents/*.md`
6. 再开始用 [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)、[`/to-prd`](/docs/notes/matt-pocock-skills/to-prd-and-issues)、[`/triage`](/docs/notes/matt-pocock-skills/diagnose-and-triage)
如果只是想单独体验 [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me),可以跳过 setup;但只要进入工程流,最好先做。
## 这个 Skill 的设计启发
`/setup-matt-pocock-skills` 看起来很朴素,但它解决了 agent workflow 里最常见的一个问题:**规则散落在人的脑子里**。
很多团队把「我们用哪个标签」「哪些 issue 可以给 AI」「CONTEXT.md 在哪」这些信息当作口头约定。人知道,AI 不知道。AI 不知道就会反复问,或者更糟糕:自己猜。
setup 的作用是把这些口头约定变成可读取文件。后续 skill 不需要更聪明,只需要稳定地读同一份项目契约。
这也是我觉得它值得单独写一篇的原因:它不是炫技的 skill,但它是整套 workflow 能长期跑起来的地基。
## 参考资源
下一篇:[Grill Me:让 AI 在你写代码前拷问你 50 个问题](/docs/notes/matt-pocock-skills/grill-me)。
# TDD:用红绿重构强迫 AI 走小步
## 失败模式:「AI 做对了东西,但跑不起来」
Matt 演讲里的第三个失败模式:**方向是对的,但 it doesn't work**。
最直接的修法是给 AI 装反馈基础设施:
* TypeScript(不用静态类型 *is crazy*)
* 让 LLM 能访问浏览器自己看页面
* 自动化测试
但 Matt 观察到一件事:**即使装了这些反馈,LLM 也用不好**。它倾向于一次写 500 行,然后才想起来「噢我应该 type check 一下」。这就是 Pragmatic Programmer 里说的 *outrunning your headlights*——开得比车头灯能照到的还快,撞墙是早晚的事。
> "The rate of feedback is your speed limit, which means you should be testing as you go, taking small deliberate steps. **And the AI by default is really not very good at that.**"
要修这个问题,需要在工具层面**强迫 AI 一步一停**。Matt 的答案是 TDD——**测试先行能强行制造检查点**。
## 经典理论:Kent Beck 的红绿重构
TDD 的标准节奏是 Kent Beck 在 2003 年那本《Test-Driven Development: By Example》里定义的:
1. **RED**:写一个失败的测试(描述要做的事)
2. **GREEN**:写最小够用的代码让测试通过
3. **REFACTOR**:在测试保护下改进代码结构
每个循环极短——分钟级。每一步都有自动化检查(测试通过/失败)。
Matt 直接照搬这个节奏,但他的 SKILL.md 里花了不少篇幅讲一个**反模式**——这才是核心。
## 关键反模式:横向切片的红绿
很多人以为 TDD 就是「先写所有测试,再写所有实现」。Matt 在 SKILL.md 里直接说这是错的:
```
WRONG (horizontal slicing):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical slicing via tracer bullets):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
为什么横向是错的?SKILL.md 给了三条理由:
> 1. Tests written in bulk test *imagined* behavior, not *actual* behavior
> 2. You end up testing the *shape* of things (data structures, function signatures) rather than user-facing behavior
> 3. Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
人话:**一口气写完所有测试是在测你脑子里的东西,不是真实代码**。等你写到 impl3 才发现 test1 设计错了——但这时 test2/test3/test4 都耦合在错的设计上,回头改一发动全身。
正确做法是**一个测试一个实现,写完一对再开下一对**。每对完成后你已经从这次实现里学到东西,下一对测试可以基于真实经验设计——而不是脑补。
## Skill 全文结构
`engineering/tdd/SKILL.md` 是 Matt 写得最长的 skill 之一,因为 TDD 本身有很多 nuance。核心结构如下:
### 哲学
> **Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
> **Good tests** are integration-style: they exercise real code paths through public APIs. They describe *what* the system does, not *how* it does it. A good test reads like a specification.
> **Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly). The warning sign: your test breaks when you refactor, but behavior hasn't changed.
记住一条诊断:**重命名一个内部函数,测试就跪了——那这个测试在测实现而不是行为,是坏测试**。
### 工作流(带 checklist)
#### 1. Planning
写代码前先和用户对齐:
```
[ ] Confirm with user what interface changes are needed
[ ] Confirm with user which behaviors to test (prioritize)
[ ] Identify opportunities for deep modules (small interface, deep impl)
[ ] Design interfaces for testability
[ ] List the behaviors to test (not implementation steps)
[ ] Get user approval on the plan
```
关键问题:「**What should the public interface look like? Which behaviors are most important to test?**」
> "**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case."
这条很反直觉。AI 默认会想穷举所有 edge case,但 Matt 强调**优先级**——不是所有行为都值得测,把火力集中到核心路径。
#### 2. Tracer Bullet
写**一个**测试,验证**一件**事:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
这就是「曳光弹」——先打一发看看准星。Matt 强调这一发要 **end-to-end**——不是先写 schema 再写 API 再写 UI,而是切一条最薄但贯穿全栈的路径。
#### 3. Incremental Loop
后面每个行为都重复 RED→GREEN:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
规则:
* 一次一个测试
* 只写够通过当前测试的代码
* **不要预判未来的测试**
* 测试聚焦于可观察行为
「不要预判」这条特别重要。AI 会忍不住想「反正这个函数后面也要支持 X,顺便加上吧」——这就开始横向切片化了。
#### 4. Refactor
测试都过之后,看重构机会:
```
[ ] Extract duplication
[ ] Deepen modules (move complexity behind simple interfaces)
[ ] Apply SOLID principles where natural
[ ] Consider what new code reveals about existing code
[ ] Run tests after each refactor step
```
> **Never refactor while RED.** Get to GREEN first.
红着重构 = 同时改测试和代码 = 你不知道是测试错还是代码错。**先绿,再重构**。
### Per-Cycle Checklist
每个红绿循环结束 Matt 让 AI 自检:
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
这五条用来挑出**坏测试和过度实现**。AI 自查一遍能拦掉大部分常见错误。
## 真实使用:从 issue 到 PR
`/tdd` 在 Matt 的工作流里是**接 `/to-issues` 的下一步**。给定一个 vertical slice issue,流程是:
```
你: 实现 issue #43
↓
/tdd
↓
Claude 读 issue acceptance criteria
↓
Claude 探索代码库 → 找到 CONTEXT.md → 用项目术语
↓
Planning 阶段:
- 列出准备改的接口
- 列出准备测的行为(按优先级排序)
- 让你点头
↓
Tracer Bullet:
- RED: 写第一个测试(基于 acceptance criteria 第 1 条)
- 跑测试,确认 fail
- GREEN: 写最小实现
- 跑测试,确认 pass
↓
Incremental Loop:
- 每个 acceptance criteria 一个 RED→GREEN
↓
Refactor:
- 看 deep module 提取机会
- 每次重构后跑全套测试
↓
PR
```
每个红绿循环 AI 都会停下来给你一个状态——「test fails」/「test passes, here's the diff」。**这些停顿就是 outrun headlights 的解药**——AI 没机会一口气铺一千行了。
## 关于 Mock:Matt 的强烈观点
SKILL.md 里特别提到 mock 的危险——他还配了一份 `mocking.md` 单独讲。核心观点:
> "Bad tests... mock internal collaborators."
mock 内部协作者 = 测试和实现 1:1 耦合 = 重构时测试集体跪。Matt 的偏好是 **integration-style 测试**——尽量用真实数据库(in-memory 或 testcontainers)、真实 HTTP(MSW)、真实文件系统(tmp dir)。只在**真正昂贵或不稳定的边界**(比如调用 OpenAI API)才 mock。
这跟很多团队的现状反着——多数代码库 unit test 满天飞,mock 比真实代码还多。Matt 在演讲里有个判断:**好代码库 = 容易测试的代码库**。如果你必须 mock 一堆东西才能测,说明代码结构有问题,应该先改架构(去 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture))。
## TDD 在 AI 时代的新意义
Kent Beck 那本书 23 年前写的时候,TDD 的核心收益是「让人不写错代码」。在 AI 时代,TDD 多了一层意义:
**它是 AI 唯一能听懂的「成功标准」**。
Matt 在演讲后半段引用了 Karpathy 的金句:
「成功标准」最好的形式就是**测试**——它是机器可验证、二值化、不会被狡辩的。给 AI 一个测试套件 + 「让它通过」,比给 AI 一段需求描述 + 「请实现」靠谱十倍。
所以 `/tdd` 不只是质量保证手段——它是 **agent loop 的输入接口**。每个红绿循环都是一次完整的「输入 → 行动 → 反馈」,AI 在循环里学到这次实现的真实情况,下次循环更准。
## 怎么装、怎么用
```bash
npx skills@latest add mattpocock/skills
```
勾选 `tdd` + `setup-matt-pocock-skills`。
如果你现在主力用 Codex,就把安装产物承接到 `.agents/skills/`,并把项目级工作流、测试命令、issue tracker 规则写进 `AGENTS.md`。Matt 的 `/tdd` 精髓是红绿重构循环,不绑死在 Claude Code。
**调用方式**:
* 直接:`/tdd` —— 让它从当前 conversation context 推断要测什么
* 接 issue:`/tdd implement #43` —— 它会去 fetch issue 再开干
* 修 bug:`/tdd reproduce this bug then fix it` —— 它会先写一个能复现 bug 的失败测试,再修
## 注意事项
**不适合所有任务**。一次性脚本、playground 探索代码、UI 微调——别用 TDD,会拖慢节奏。Matt 自己说 TDD 适合「有持久价值、需要被维护」的代码。
**测试基础设施先备好**。如果项目还没装测试框架(Vitest / Jest / Playwright 等),先装好再用 `/tdd`,否则它会先帮你装,但那一步问得很多。
**不要让它自动加 e2e 测试**。e2e 慢且脆,TDD 节奏要分钟级。`/tdd` 默认偏 integration test 不偏 e2e,但你可以明确告诉它「单元 + integration only,不要 e2e」。
**重构阶段最容易失控**。AI 拿到 GREEN 状态后会兴奋地重构一堆——盯着,每次重构后跑测试。这部分是 AI 走偏的高发区。
## 参考资源
下一篇:[Improve Codebase Architecture:把 shallow 重构成 deep modules](/docs/notes/matt-pocock-skills/improve-codebase-architecture)——周期性维护,让 AI 长期能在你的代码库里跑得动。
# to-PRD + to-Issues:把 grill 出来的对话凝固成可执行的 vertical slice
## 核心结论
**`/to-prd` 和 `/to-issues` 是 Matt Pocock 工作流里把“聊清楚”变成“可执行”的中间层**:`/to-prd` 负责把 grill-me 对话压缩成 PRD,`/to-issues` 负责把 PRD 切成一个个端到端可验收的 vertical slice。
它们解决的是 AI 编程里一个常见断点:你已经和 AI 讨论清楚了需求,但如果直接让它「全部实现」,结果往往是一次性输出太大、难 review、难测试、难交给 AFK agent 独立完成。正确做法是先固化成 PRD,再切成小到可以独立领取的 issue。
| 阶段 | 输入 | 输出 | 关键原则 |
| ------------ | -------- | --------------------- | --------------------- |
| `/grill-me` | 模糊想法 | 已澄清的对话决策 | 先问清楚,不急着写 |
| `/to-prd` | 当前对话上下文 | 结构化 PRD | 不再采访用户,只合成已知信息 |
| `/to-issues` | PRD | vertical slice issues | 每个 issue 都能端到端验收 |
| `/tdd` | 单个 issue | 可测试实现 | 一个 slice 一个 slice 跑红绿 |
## 这一段在工作流里的位置
回到 Matt 的工作流图:
```
/grill-me 或 /grill-with-docs ← 谈清楚
↓
/to-prd ← 凝固成 PRD(你在这里)
↓
/to-issues ← 切成可领取的 vertical slice(你在这里)
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture
```
`/to-prd` 和 `/to-issues` 是承上启下的一段:把抽象的对话决策**翻译成可执行的工作单元**。我把它们合并讲,因为在 Matt 的实际使用里它们就是连续两步。
## 失败模式:grill 完之后没事干
不少人用完 `/grill-me` 后会卡住:拿到一段长长的对话,里面有一堆决策,但**怎么开始写代码**?
直接喂给 Claude「请实现」是错的——因为:
1. 一次实现完整功能 = AI 输出 1000+ 行 = 很难 review、很难测、bug 难定位
2. AI 失忆后下一个 session 没有上下文
3. 没有 tracking——没法知道做到哪、还剩多少
正确做法是把决策**冻结成 artifact**(PRD),再把 PRD 切成**小到可独立完成**的工作包(issues)。这是软件工程 30 年的常识,但在 AI 时代有了新意义:
> 切得够细,才能让 AFK agent(你不在场时跑的 agent)独立领取并完成。
## /to-prd:把对话压缩成 PRD
### Skill 的关键约束
`/to-prd` 的 SKILL.md 开头写了一句很重要的话:
> "This skill takes the current conversation context and codebase understanding and produces a PRD. **Do NOT interview the user — just synthesize what you already know.**"
不再问问题。这是 grill-me 阶段做的事,to-prd 只做**合成**。所以**不要清 context 再跑 to-prd**——它依赖前面 grill 出来的全部对话。
### Skill 的处理流程
1. **探索代码库**(如果还没探过)—— 用项目的 CONTEXT.md 词汇,尊重已有 ADR
2. **草拟模块**—— 主动找可以提取为 deep module 的机会,让接口可独立测
3. **和用户对齐模块**—— 「这些模块对吗?哪些要写测试?」
4. **按模板生成 PRD**,发到 issue tracker,打 `needs-triage` 标签
### PRD 模板
Matt 给的模板是这样:
```markdown
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list:
1. As a , I want a , so that
2. ...
## Implementation Decisions
- The modules that will be built/modified
- The interfaces of those modules
- Technical clarifications from the developer
- Architectural decisions
- Schema changes / API contracts / Specific interactions
(NO specific file paths or code snippets — they rot fast.)
## Testing Decisions
- What makes a good test (test external behavior, not internals)
- Which modules will be tested
- Prior art (similar tests in the codebase)
## Out of Scope
What's NOT in this PRD.
## Further Notes
```
几个关键设计:
* **User Stories 占大头**:要求 LONG, numbered list——逼你穷举完整的功能点。这避免了「我以为说清楚了」的盲点
* **Implementation 不写文件路径或代码**:Matt 直接说 "they may end up being outdated very quickly"。这是 LLM 时代特有的考量——具体路径在重构后立刻过时,但「模块边界」「接口契约」生命周期更长
* **必须有 Out of Scope**:这个段落在大多数 PRD 模板里被忽略,但它是后续切 issue 时的边界保险
## /to-issues:把 PRD 切成 vertical slice
### Vertical Slice 是什么
这是 Matt 整套方法论里最重要的概念之一。SKILL.md 直接说:
> Each issue is a thin vertical slice cutting through ALL integration layers end-to-end, NOT a horizontal slice of one layer.
举例最清楚:要做「评论功能」。
**横向切片(错的方式)**:
* Issue 1: 数据库 schema
* Issue 2: API 端点
* Issue 3: UI 组件
* Issue 4: 测试
**纵向切片(Tracer Bullet 方式)**:
* Issue 1: 「访客可以提交一条匿名评论」(schema + API + UI + test 全在内,但范围小到只能匿名)
* Issue 2: 「登录用户的评论关联到账户」
* Issue 3: 「评论可以被回复」
* Issue 4: 「管理员可以删除评论」
横向切片的问题:每个切片单独验不了。Issue 1 完成后没有可演示的东西,要等到 Issue 4 才能跑通整个链路——直到那时才发现 schema 设计错了。
纵向切片每完成一个就是**端到端可用的功能子集**——用 Pragmatic Programmer 里的话叫 **tracer bullet**(曳光弹),一发先打过去看准星,再调整下一发。
### HITL vs AFK
`/to-issues` 还会给每个 slice 标一个标签:
* **HITL**(Human in the Loop)—— 需要人参与决策。比如架构决策、设计审查
* **AFK**(Away From Keyboard)—— agent 可以独立做完,你回来看结果
> "Prefer AFK over HITL where possible."(能 AFK 就别 HITL)
这是 Matt 的工作流里一个很激进的设想:你切完 issue,**直接派给跑在你不在场时的 agent**(比如夜里、周末),第二天回来 PR 已经躺在那等 review 了。HITL 的部分留在白天和 agent 配合干。
### 切片确认环节
`/to-issues` 不会一上来就生成 issue——它会先把切片方案以编号列表的方式给你看:
```
1. Title: 访客提交匿名评论
Type: AFK
Blocked by: None
User stories covered: #1, #2
2. Title: 评论关联到登录账户
Type: AFK
Blocked by: #1
User stories covered: #3
3. Title: 评论审核流程
Type: HITL(需要确认审核 UI 设计)
Blocked by: #1
User stories covered: #4, #5
```
然后问你:
* 粒度对吗?太粗 / 太细?
* 依赖关系对吗?
* 哪些应该合并 / 拆分?
* HITL/AFK 标对了吗?
迭代到你点头之后才真正发到 issue tracker,按依赖顺序发(先发 blocker),这样后发的 issue 可以引用先发的真实 issue ID。
### Issue 模板
```markdown
## Parent
A reference to the parent issue (if any).
## What to build
A concise description. Describe end-to-end behavior, NOT layer-by-layer
implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
## Blocked by
- A reference to the blocking ticket
(or "None - can start immediately")
```
注意 "describe end-to-end behavior" 这条——和 vertical slice 的精神一致。Acceptance criteria 是验收清单,Claude 在 `/tdd` 阶段会逐条转成测试。
## 怎么用:完整流程示例
假设你要给博客加评论功能。完整流程:
```
你: 我想给博客加评论功能
↓
/grill-me → Claude 问 30 个问题(要不要登录?匿名?嵌套?审核?……)
↓
你回答完毕,达成共识
↓
/to-prd → Claude 生成结构化 PRD,提交到 GitHub Issues #42
↓
/to-issues → Claude 提议切成 4 个 vertical slice
让你确认粒度和依赖
你点头
按依赖顺序发布到 GitHub Issues #43~#46
↓
你回家睡觉
↓
夜里 AFK agent 抓 #43(无依赖),跑 /tdd 完成 → 提 PR
你早上 review、merge
↓
agent 抓 #44 / #45 ……
```
整套流程不需要你坐在屏幕前看着每个细节,关键决策都在 grill-me 阶段做完了。
## 安装与前置条件
```bash
npx skills@latest add mattpocock/skills
```
勾选 `to-prd`、`to-issues`、`setup-matt-pocock-skills`。
**必须先跑 `/setup-matt-pocock-skills`**——它会把你的 issue tracker(GitHub / GitLab / 本地 markdown)和 triage label 词汇写到 AGENTS.md/CLAUDE.md,否则 to-prd 和 to-issues 不知道往哪发 issue。
支持的 issue tracker:
* **GitHub Issues**(默认,用 `gh` CLI)
* **GitLab Issues**(用 `glab` CLI)
* **本地 markdown**(在 `.scratch//` 下创建文件)—— 适合个人项目或没有远程的项目
* **其他**(Jira、Linear 等)—— 用一段 prose 描述工作流,skill 会按你的描述调用
## 常见问题
**Q: 已经有现成 PRD 了,能跳过 to-prd 直接 to-issues 吗?**
A: 可以。`/to-issues` 接受 issue 引用作为参数("Break down issue #42 into vertical slices"),它会去 fetch issue 内容再切。
**Q: 我的项目没用 issue tracker,能用吗?**
A: 能。setup 时选「本地 markdown」,所有 issue 会变成 `.scratch//001-foo.md` 这样的本地文件。
**Q: PRD 太长,AI 自己也搞不定怎么办?**
A: 这是切片粒度的信号——PRD 应该被切成多个独立 PRD 而不是一个巨型 PRD。在 grill 阶段就该感觉到:如果聊到第 50 个问题还在引入新功能,先停下来切成两个 PRD 分批做。
**Q: AFK agent 怎么自动抓 issue?**
A: 这部分 Matt 的 repo 没提供,要配合你自己的 agent 编排(比如 GitHub Actions 触发 Claude Code 跑 issue)。最简单的做法是 cron 每小时检查 `is:open no:assignee label:agent-ready`。
## 这套流程的真正价值
`/to-prd` 和 `/to-issues` 看起来像「自动化项目管理」,但 Matt 把它放在工作流核心位置的原因更深:
**它强迫你思考「**什么是一个完整的小事**」**。当你被迫把功能切成 vertical slice,你就在做一件软件工程里最难的事——**找接缝**。这个思维本身比这两个 skill 更值钱。
而且 vertical slice 的尺寸就是 AI 一次能搞定的尺寸。**让任务大小匹配 AI 的能力上限**——这是和 LLM 协作的根本节奏。
## 参考资源
下一篇:[TDD:用红绿重构强迫 AI 走小步](/docs/notes/matt-pocock-skills/tdd)——issue 切完了,怎么让 AI 真的小步实现。
# Zoom Out:当你迷路时,让 AI 先画地图
## 最短,但很有用
`/zoom-out` 可能是 Matt 这套稳定 engineering skill 里最短的一个。它的核心指令可以概括成一句:
> 我不熟悉这块代码,请上升一层抽象,用项目领域语言给我画出相关模块和调用者地图。
它不是用来写代码的,也不是用来重构的。它用来处理一种很常见的状态:**你和 AI 都已经钻进某个文件,但开始忘记这个文件为什么存在**。
## 它解决的失败模式
AI 编程很容易进入局部最优:
1. 用户指了一个文件
2. AI 读这个文件
3. AI 根据局部代码猜意图
4. 改完之后才发现上游调用者、领域规则或 ADR 不支持这个改法
人类也一样。我们 debug 久了会盯住一个函数,忘掉它在系统里的位置。
`/zoom-out` 的作用是打断这种隧道视野。它让 agent 暂停实现,先回答:
* 这块代码属于哪个领域概念?
* 谁调用它?
* 它调用谁?
* 它背后有哪些不变量?
* 它和 `CONTEXT.md` 里的术语怎么对应?
* 它是否受某个 ADR 约束?
## 和 /improve-codebase-architecture 的区别
`/zoom-out` 和 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture) 都会看系统全局,但目标完全不同。
| Skill | 目标 | 输出 |
| -------------------------------- | ---------- | ----------------------- |
| `/zoom-out` | 帮你理解一块陌生代码 | 地图、调用关系、领域解释 |
| `/improve-codebase-architecture` | 找可深化的架构机会 | 候选重构、deletion test、接口设计 |
`/zoom-out` 更像「请给我讲讲这块」。它不应该急着提出重构方案,更不应该直接改代码。它的任务是降低认知负担。
## 为什么要强调领域词汇
这个 skill 明确要求使用项目的 domain glossary。原因是:如果只用文件名解释,AI 很容易输出这种东西:
```text
OrderService 调用 OrderRepository,然后 OrderRepository 调用 db client。
```
这听起来像解释了,其实没解释。更有用的地图应该长这样:
```text
Checkout flow 里,Order Draft 是用户尚未支付前的临时订单。
Order Finalization 会把 Draft 转成不可变 Order,并触发 Inventory Reservation。
`OrderService.finalize()` 是这个转换的 seam,调用者主要来自 Payment Callback 和 Admin Retry。
```
第二种解释把代码放回了业务语言里。你不只是知道「谁调用谁」,还知道「它为什么存在」。
## 适合什么时候用
我建议在这些场景下主动调用 `/zoom-out`:
* 接手一个陌生模块前
* 改一个 bug,但还不确定相关调用链
* review AI 生成代码,看不出它是否改到了正确层级
* 准备写 PRD,想确认模块边界
* 已经看了 3 个文件还没形成系统图
* 准备跑 `/improve-codebase-architecture`,但还不确定候选区域
它特别适合作为「动手前的 5 分钟」。有些 bug 不是因为代码难,而是因为一开始就看错了层级。
## 一个可复用的输出格式
虽然原始 skill 极短,但我建议在使用时让 AI 按这个格式输出:
```markdown
## 这块代码在系统里的位置
## 关键领域术语
## 主要模块
| 模块 | 责任 | 调用者 | 被调用对象 |
|---|---|---|---|
## 关键流程
## 已知约束 / ADR
## 我建议你先看的文件
```
这个格式比普通解释更稳定,也更适合转成后续 `/to-prd` 或 `/diagnose` 的上下文。
## 不要把它用成计划模式
`/zoom-out` 的危险是:AI 讲完地图后,顺手开始建议「可以这么改」。如果你只是想理解代码,应该明确限制:
```text
只解释结构,不提出实现方案,不改文件。
```
因为它的价值就在于把决策和理解分开。理解不清时提方案,往往只是把误解包装得更漂亮。
## 我的使用建议
`/zoom-out` 很适合和其他 skill 组合:
* `/zoom-out` → `/diagnose`:先看系统地图,再建反馈回路
* `/zoom-out` → `/grill-with-docs`:先理解现有领域语言,再拷问新需求
* `/zoom-out` → `/to-prd`:先确认模块位置,再写 PRD
* `/zoom-out` → `/improve-codebase-architecture`:先画地图,再找 shallow/deep 问题
它不是完整流程,只是一个刹车。AI 开始在局部文件里越改越多时,先让它 zoom out,通常能省掉后面一轮返工。
## 参考资源
下一篇:[Prototype:用可丢弃代码回答一个设计问题](/docs/notes/matt-pocock-skills/prototype)。
# Pi Agent 是什么
## 引言
如果只看功能,Pi Agent 很容易被低估:它在终端里运行,能读文件、改文件、执行命令、保存会话,也能切换模型。听起来像另一个 Claude Code 或 Codex。
但我觉得 Pi 真正有意思的地方,不是它多做了什么,而是它少做了什么。它把 AI 编程工具最核心的那层保留下来:模型、上下文、工具、会话、扩展,然后尽量不把用户的工作流提前写死。
所以我更愿意这样理解 Pi:
**Pi Agent 不是一个“更全”的 AI 编程产品,而是一个更薄、更透明的 coding agent harness。**
这个判断比功能清单更重要。因为它决定了你应该怎么学习 Pi:不是先背命令,而是先理解一个 coding agent 到底由哪几层组成。
## 先把 Pi 放对位置
一个 AI 编程工具通常可以先粗略分成三层:模型、harness、工程环境。但如果只画这三层,还是太抽象。Pi 真正值得看的,是中间这层 harness 里面又拆成了哪些模块:
这张图里有三个重点。
第一,Pi 不是只有一个“聊天 UI”。CLI、交互 TUI、print/JSON、RPC、SDK 都只是入口,真正承接任务的是 `AgentSessionRuntime` 和 `AgentSession`。
第二,Pi 在请求模型之前会先做资源加载。`ResourceLoader` 会把 `AGENTS.md`、`CLAUDE.md`、skills、extensions、prompt templates 这些东西整理出来,再交给 `SystemPrompt Builder` 组装成模型真正看到的上下文。
第三,模型调用工具时,并不是模型直接控制文件系统。`AgentHarness` 和 `AgentLoop` 负责校验工具、执行工具、接住结果、继续下一轮。Extensions、Tool Registry、SessionManager 则在旁边扩展能力和保存状态。
所以 Pi 站在中间,但这个“中间”不是一句空话。它具体控制的是:哪些上下文进入模型,哪些工具可以被调用,工具结果如何回到会话,哪些能力由扩展补进来。
这也是为什么很多介绍 Pi 的文章都会强调 minimal、transparent、extensible。它们其实在说同一件事:Pi 试图把 agent 的核心运行层做小,让用户能看见,也能改。
## 一次工作是怎么流动的
Pi 的一次请求不是“问模型一句话,模型回一句话”。更准确地说,它是一段带分支的时序:
这里最关键的是第四步到第五步之间的来回。模型不直接碰你的文件系统,它只提出 `tool_call`;Pi 接住这个调用,校验工具名和参数,触发可能存在的 extension hooks,执行真实操作,再把 `tool_result` 放回上下文。模型再根据新的上下文判断下一步。
这就是 coding agent 和普通聊天机器人的区别。聊天机器人主要在文本里完成任务;coding agent 要进入工程系统,所以它必须有 harness 来管理工具、上下文和状态。
Pi 的默认工具很少:
| 工具 | 含义 |
| ------- | ----------- |
| `read` | 读取文件 |
| `edit` | 修改已有文件 |
| `write` | 创建或覆盖文件 |
| `bash` | 执行 shell 命令 |
还有 `grep`、`find`、`ls` 这类只读工具可以被启用或限制。这个工具集看起来克制,但已经形成了编程闭环:读代码、改代码、跑测试、根据错误继续修。
这套设计背后的问题不是“Pi 会不会做更多”,而是“更多东西是否应该默认进核心”。Pi 的回答很明确:不一定。
## 为什么它不急着内置很多功能
很多 AI 编程产品会把计划模式、todo、子代理、MCP、权限弹窗、后台任务、浏览器工具都做进产品里。这样上手快,但也带来一个代价:你很难知道模型实际收到了什么上下文,也很难把产品工作流改成自己的工作流。
Pi 的路线相反。它把核心保持得很小,然后把工作流放到外面:
| 你想改变什么 | Pi 交给哪里 |
| --------- | ------------------------- |
| 项目规则 | `AGENTS.md` / `CLAUDE.md` |
| 专门任务方法 | Skills |
| 自定义工具和 UI | Extensions |
| 一组可分享能力 | Pi Packages |
| 模型选择 | Provider / Model 配置 |
这不是“功能不够”,而是一种产品取舍:核心只管 agent loop,具体工作流交给用户和团队自己组合。
举个例子,Pi 没有默认内置 DeepSearch。但这并不意味着它不能做深度搜索。更符合 Pi 思路的做法,是写一个 extension:注册一个 `deep_search` 工具,把 Tavily、Exa、Brave Search 或公司内部搜索接进去,再让模型在需要时调用它。
这和把“搜索按钮”硬编码进产品不同。前者是你在扩展 agent 的能力,后者是产品替你决定工作流。
## 重点:怎么扩展自己的工作流
如果要真正用 Pi,而不只是“体验一下”,重点一定是扩展自己的工作流。
这里容易混淆的是,Pi 不是只有“插件”这一种扩展方式。它更像给你四层入口:项目规则、任务方法、真实工具、可分享包。你要先判断自己想沉淀的到底是哪一种东西。
| 你要沉淀的东西 | 用什么 | 适合什么场景 |
| ------- | ------------------------- | ------------------------------------------------------ |
| 项目习惯和约束 | `AGENTS.md` / `CLAUDE.md` | 告诉 agent 怎么改代码、跑什么检查、哪些目录不能碰 |
| 一套可复用方法 | Skill | 代码审查、写文章、发版、生成文档、图片处理这类“步骤和经验” |
| 一个真实能力 | Extension | 注册工具、拦截工具调用、加 slash command、加 UI、接外部 API |
| 一组可分发能力 | Pi Package | 把 extensions、skills、prompt templates、themes 打包给自己或团队复用 |
我的理解是:**Skill 是工作手册,Extension 是可执行插件,Package 是分发容器。**
比如我现在这个博客工作流,可以这样拆:
| 工作流需求 | 放进哪里 |
| ----------------------------------------------------------- | ----------------------- |
| “写内容只写中文,图片必须用 `BlogImage`,改 MDX 后跑 `pnpm types:check`” | `AGENTS.md` |
| “写概念文章时按误解、定义、机制、例子、边界来组织” | `article-writing` Skill |
| “给 Pi 增加一个 `deep_search` 工具,能查 Tavily / Exa / Brave Search” | Extension |
| “把写作 Skill、DeepSearch Extension、微信发布命令打包给多个项目用” | Pi Package |
这就比单纯说“装插件”更准确。因为很多工作流不需要写代码,只需要一份好的规则或 Skill;但只要你希望 agent 真的多一个能力,比如查外部搜索、查数据库、调用 CI、拦截危险命令,就应该写 Extension。
### Extension:真正的插件层
Pi 的 Extension 是 TypeScript 模块。它可以做几类事:
| 能力 | 例子 |
| ----- | ------------------------------------------ |
| 注册工具 | `deep_search`、`query_logs`、`open_issue` |
| 注册命令 | `/review`、`/publish`、`/checkpoint` |
| 拦截事件 | 在 `bash` 执行 `rm -rf`、`sudo`、写 `.env` 前要求确认 |
| 改 UI | 在 TUI 里显示状态、选择框、确认框、任务面板 |
| 保存状态 | 记录 todo、连接池、上次搜索结果、任务阶段 |
| 接外部系统 | CI、GitHub、日志系统、公司内部 API |
Extension 可以放在全局,也可以放在项目里:
```text
~/.pi/agent/extensions/ # 全局扩展,所有项目可用
.pi/extensions/ # 项目扩展,只在当前项目里用
```
测试一个临时扩展,可以用:
```bash
pi -e ./my-extension.ts
```
放到自动发现目录后,可以在 Pi 里用:
```text
/reload
```
重新加载 extensions、skills、prompts 和 context files。
这就是我觉得 Pi 最有价值的地方:你不只是“让模型帮你写代码”,而是在给模型设计一个可控的工作环境。Extension 决定模型能调用什么能力,hooks 决定哪些行为要被拦截,commands 决定你自己的工作流如何被触发。
### Skill:不要把所有东西都写成插件
如果一个能力主要是“怎么做”,而不是“调用一个真实 API 或执行一段程序”,那它更适合写成 Skill。
Skill 的结构通常是:
```text
my-skill/
SKILL.md
scripts/
templates/
references/
```
Pi 启动时不会把完整 skill 全塞进上下文。它会先加载 skill 的名字和描述;当任务匹配时,再让模型读取完整的 `SKILL.md`。这叫渐进式披露。好处是:你可以保存复杂方法论,但不用每次都污染上下文。
比如“写一篇好文章”“发布到公众号”“做一次浏览器 QA”,这些都更像 Skill。它们的价值主要在步骤、判断标准和参考材料,不一定需要注册一个 LLM 可调用工具。
### Package:把自己的工作流打包
当你已经有一组稳定能力,就可以考虑把它做成 Pi Package。
Package 可以包含:
| 内容 | 作用 |
| ---------------- | ----------------- |
| extensions | 可执行插件、工具、命令、hooks |
| skills | 工作方法和任务手册 |
| prompt templates | 常用 prompt 模板 |
| themes | TUI 主题 |
安装方式大概是:
```bash
pi install npm:@scope/my-pi-package
pi install git:github.com/user/repo@v1
pi install ./relative/path/to/package
pi list
pi remove npm:@scope/my-pi-package
pi update --extensions
```
默认安装会写到个人设置里。如果你想让团队项目共享,可以用项目级设置,让包记录在 `.pi/settings.json`。这样别人进入项目启动 Pi 时,也能自动补齐缺失的 package。
但这里也要非常谨慎:Package、Extension、Skill 都可能影响 agent 的行为。第三方 package 不是浏览器插件那种低权限装饰,它可能运行代码,也可能指导模型执行命令。安装前应该看源码。
所以我会按这个顺序学习 Pi 的扩展:
1. 先用 `AGENTS.md` 写清项目规则。
2. 再把重复方法做成 Skill。
3. 需要真实工具能力时,再写 Extension。
4. 多项目复用时,最后再打成 Package。
这样学习比较稳。你不是一上来就写插件,而是先把工作流拆成“规则、方法、工具、分发”四类,再决定每一类放到 Pi 的哪一层。
## 从源码里能看到什么
我看 Pi 源码时,最有帮助的不是追每个函数,而是看几个文件各自代表哪层设计:
| 源码位置 | 说明 |
| ---------------------------------------------------- | ------------------------------------------------------- |
| `packages/agent/src/agent-loop.ts` | 核心循环:把用户消息、模型响应、工具调用和工具结果串起来 |
| `packages/agent/src/harness/agent-harness.ts` | Harness 状态:管理 session、system prompt、tools、hooks 和消息队列 |
| `packages/coding-agent/src/core/tools/index.ts` | 内置工具集合:默认 coding tools 是 `read`、`bash`、`edit`、`write` |
| `packages/coding-agent/src/core/resource-loader.ts` | 资源加载:读取项目指令、extensions、skills、prompt templates 和 themes |
| `packages/coding-agent/src/core/system-prompt.ts` | 系统提示词构建:把工具说明、项目上下文、skills 和当前目录放进 prompt |
| `packages/coding-agent/src/core/extensions/types.ts` | 扩展系统:允许扩展注册工具、命令、快捷键、UI 和生命周期事件 |
这几块合起来,基本就是 Pi 的心脏:它先组装上下文和工具,再把请求交给模型;模型如果要调用工具,Pi 执行工具;结果回来后,循环继续。
所以 Pi 的“极简”不是空口号。源码结构本身也在表达这个想法:把 agent loop、harness、coding tools、资源加载、扩展系统分开,每层都相对清楚。
## Pi 的边界
Pi 的自由度很高,但自由度不等于安全。
Pi packages 和 extensions 可以运行代码;skills 也可能指导模型执行脚本;`bash` 能触达你的真实系统。官方文档和安全分析都提醒过:第三方 package、extension、skill 需要自己审查。
我会把 Pi 的边界理解成三点:
1. **它不是沙箱**:不要把它当成天然隔离环境。危险项目最好放进容器、临时目录或干净 worktree。
2. **它不替你判断权限**:Pi 的核心哲学不是靠一堆弹窗管理风险,而是让你控制工具、上下文和扩展。
3. **它适合懂工程边界的人**:你需要知道什么时候让 agent 跑命令,什么时候只给 read-only 工具,什么时候先建 git checkpoint。
这也是 Pi 和一些更产品化 agent 的差异。产品化工具会替你包装更多安全和交互细节;Pi 给你更直接的控制权,同时也把更多责任还给你。
## 应该怎么学习 Pi
学习 Pi,我不建议从“有哪些命令”开始。命令很快能查到,真正值得学的是这几个问题:
| 问题 | 为什么重要 |
| ----------------------- | ------------------- |
| Pi 怎么组装上下文 | 决定模型实际知道什么 |
| Pi 的工具面为什么这么小 | 决定 agent 的行为是否可观察 |
| Extension 怎么注册工具 | 决定你能不能把自己的工作流接进去 |
| Skill 和 Extension 有什么区别 | 决定什么时候写说明,什么时候写代码 |
| Session 怎么保存和分叉 | 决定一次工程探索能不能恢复、回看和继续 |
如果你已经用过 Claude Code 或 Codex,可以把 Pi 当成一次“拆开看”的机会:同样是让模型写代码,为什么有的工具像黑箱产品,有的工具像可改造的运行时?
这个问题比“Pi 能不能替代某个工具”更值得问。
## 写在最后
Pi Agent 最有价值的地方,是它把 AI 编程工具的中间层暴露出来了。
它提醒我们:一个 coding agent 的能力,不只来自模型,也来自 harness 的设计。模型负责想,harness 负责让模型在真实工程环境中行动。上下文怎么进来,工具怎么出去,结果怎么回来,扩展怎么插入,这些细节共同决定了 agent 是否可靠、透明、可控。
所以我不会把 Pi 简单看成“Claude Code 平替”。它更像一个适合开发者研究和改造的 agent runtime。你可以直接用它写代码,也可以用它学习怎样设计自己的 agent 工作流。
下一篇实战我会按这个思路继续:不写一个普通教程,而是用 Pi Extension 做一个 `deep_search` 工具,看看怎样把外部搜索能力接进 agent loop。
## 延伸阅读
# Pi Agent 实践指南
## 快速回顾
在概念篇里,我把 Pi Agent 理解成一个极简的 **Agent Harness**:它连接模型、终端、文件系统、shell、会话和扩展系统,但不替你预设一整套厚重工作流。
所以实践篇我不想再做一个普通的"让 Pi 改文件"案例。那个案例能说明基础闭环,但不够体现 Pi 的可扩展性。
更适合 Pi 的实战案例,是给它补一个它默认没有、但很多人真实需要的能力:**DeepSearch**。
这里的 DeepSearch 不是简单的联网搜索,而是一套研究型工作流:
| 阶段 | 要做什么 |
| ---- | ----------------------- |
| 问题拆解 | 把一个模糊问题拆成几个可检索子问题 |
| 多轮检索 | 分别搜索官方文档、代码仓库、博客、讨论区或论文 |
| 来源筛选 | 去重、排除低质量结果、优先保留一手来源 |
| 证据整理 | 摘出关键事实、链接、时间、版本和不确定性 |
| 综合回答 | 给出结论,同时说明依据和限制 |
我的判断是:**DeepSearch 不应该写进 Pi 本体,也不应该只靠 prompt 硬凑。它更适合做成一个 Pi Extension。**
原因很简单:DeepSearch 涉及网络请求、第三方搜索 API、来源过滤、结果截断、引用格式和安全边界。这些都属于工作流能力,而不是 coding agent 的最小核心。
## 设计目标
这个案例要实现的不是一个完美的研究系统,而是一个可跑通的最小版本。
目标如下:
```text
给 Pi 增加一个 deep_search 工具。
它接收:
- query:用户要研究的问题
- depth:检索深度
- maxResults:最多返回多少条候选资料
它输出:
- 结构化搜索结果
- 每条结果的标题、URL、摘要、相关性
- 给模型使用的证据提示
Pi 拿到这些证据后,再由当前模型生成最终结论。
```
我会刻意把"检索"和"综合"拆开:
| 部分 | 由谁负责 | 原因 |
| --------- | -------------------- | ------------- |
| 搜索 API 调用 | DeepSearch extension | 这是确定性的外部能力 |
| 结果去重和截断 | DeepSearch extension | 避免上下文被噪声塞满 |
| 判断哪些证据重要 | Pi 当前模型 | 需要推理和上下文理解 |
| 最终答案写作 | Pi 当前模型 | 需要结合用户问题和项目语境 |
这样做更稳。Extension 不需要自己再调用一个模型,也不需要变成一个嵌套 agent。它只提供高质量证据,让 Pi 原本的模型继续推理。
## 准备工作
Pi extension 可以放在全局目录,也可以放在项目目录。这里我建议先放项目目录:
```text
.pi/extensions/deepsearch/
package.json
index.ts
```
项目本地 extension 的好处是边界清楚。这个 DeepSearch 能力只在当前项目里启用,不会影响所有 Pi 会话。
搜索服务可以选 Tavily、Exa、Brave Search、SerpAPI,甚至你自己的搜索后端。第一版不要纠结服务商,先抽象成一个 `searchWeb()` 函数。
例如用环境变量保存 API Key:
```bash
export TAVILY_API_KEY=tvly-...
```
如果你不想接第三方搜索 API,也可以先用本地 mock 数据把 extension 跑通。等工具注册、参数传递和结果格式都稳定后,再接真实搜索服务。
## Step 1: 创建 Extension 目录
先创建目录:
```bash
mkdir -p .pi/extensions/deepsearch
```
如果 extension 需要依赖,可以放一个 `package.json`:
```json
{
"name": "pi-deepsearch-extension",
"private": true,
"dependencies": {
"typebox": "*",
"@earendil-works/pi-ai": "*",
"@earendil-works/pi-coding-agent": "*"
},
"pi": {
"extensions": ["./index.ts"]
}
}
```
然后安装依赖:
```bash
cd .pi/extensions/deepsearch
npm install
```
Pi 的 extension 是 TypeScript 模块,不需要你先手动编译。这个体验很适合快速做工具实验。
## Step 2: 注册 deep\_search 工具
核心文件是 `.pi/extensions/deepsearch/index.ts`。
第一版可以这样写:
```typescript
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { StringEnum } from "@earendil-works/pi-ai";
import { Type } from "typebox";
type SearchResult = {
title: string;
url: string;
snippet: string;
score?: number;
};
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "deep_search",
label: "DeepSearch",
description: "Search the web for source-backed evidence about a question.",
promptSnippet: "Research a question with web search and return source-backed evidence.",
promptGuidelines: [
"Use deep_search when the user asks for current facts, external sources, comparison, investigation, or source-backed research.",
"After deep_search returns results, synthesize an answer with citations and clearly separate facts, inference, and uncertainty.",
"Do not treat deep_search results as final truth; inspect source quality and mention gaps."
],
parameters: Type.Object({
query: Type.String({
description: "The research question or search query."
}),
depth: Type.Optional(StringEnum(["quick", "normal", "deep"] as const)),
maxResults: Type.Optional(Type.Number({
minimum: 3,
maximum: 10,
default: 6
}))
}),
async execute(_toolCallId, params, signal) {
const depth = params.depth ?? "normal";
const maxResults = params.maxResults ?? 6;
const results = await searchWeb(params.query, depth, maxResults, signal);
return {
content: [
{
type: "text",
text: formatResultsForModel(params.query, results)
}
],
details: {
query: params.query,
depth,
results
}
};
}
});
}
async function searchWeb(
query: string,
depth: "quick" | "normal" | "deep",
maxResults: number,
signal: AbortSignal
): Promise {
const apiKey = process.env.TAVILY_API_KEY;
if (!apiKey) {
throw new Error("Missing TAVILY_API_KEY. Set it before starting pi.");
}
const response = await fetch("https://api.tavily.com/search", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
api_key: apiKey,
query,
search_depth: depth === "quick" ? "basic" : "advanced",
max_results: maxResults,
include_answer: false,
include_raw_content: depth === "deep"
}),
signal
});
if (!response.ok) {
throw new Error(`Search failed: ${response.status} ${response.statusText}`);
}
const data = await response.json() as {
results?: Array<{
title?: string;
url?: string;
content?: string;
score?: number;
}>;
};
return dedupeByUrl((data.results ?? []).map((item) => ({
title: item.title ?? "Untitled",
url: item.url ?? "",
snippet: item.content ?? "",
score: item.score
}))).filter((item) => item.url);
}
function dedupeByUrl(results: SearchResult[]): SearchResult[] {
const seen = new Set();
const deduped: SearchResult[] = [];
for (const result of results) {
const key = normalizeUrl(result.url);
if (seen.has(key)) continue;
seen.add(key);
deduped.push(result);
}
return deduped;
}
function normalizeUrl(url: string): string {
try {
const parsed = new URL(url);
parsed.hash = "";
parsed.searchParams.delete("utm_source");
parsed.searchParams.delete("utm_medium");
parsed.searchParams.delete("utm_campaign");
return parsed.toString();
} catch {
return url;
}
}
function formatResultsForModel(query: string, results: SearchResult[]): string {
if (results.length === 0) {
return `DeepSearch found no results for: ${query}`;
}
const lines = results.map((result, index) => {
return [
`## Source ${index + 1}`,
`Title: ${result.title}`,
`URL: ${result.url}`,
result.score === undefined ? undefined : `Score: ${result.score}`,
`Snippet: ${result.snippet}`
].filter(Boolean).join("\n");
});
return [
`DeepSearch query: ${query}`,
"",
"Use these sources as evidence. Cite URLs when making factual claims.",
"Separate confirmed facts from inference and uncertainty.",
"",
...lines
].join("\n\n");
}
```
这段代码只做最关键的事:
| 代码位置 | 作用 |
| ------------------------- | ----------------------- |
| `pi.registerTool()` | 把 `deep_search` 暴露给模型调用 |
| `parameters` | 告诉模型工具需要哪些参数 |
| `promptGuidelines` | 告诉模型什么时候用、用完后怎么处理 |
| `searchWeb()` | 调用真实搜索服务 |
| `dedupeByUrl()` | 去掉重复 URL |
| `formatResultsForModel()` | 把搜索结果整理成模型容易引用的证据块 |
第一版先不要做太复杂。DeepSearch 真正难的不是写一个搜索请求,而是把来源质量、上下文长度、引用格式和不确定性控制住。
## Step 3: 加一个 /deepsearch 命令
工具是给模型调用的,但用户也需要一个直接入口。
可以再注册一个命令,把用户输入改写成更明确的研究任务:
```typescript
export default function (pi: ExtensionAPI) {
pi.registerCommand("deepsearch", {
description: "Run a source-backed DeepSearch task",
handler: async (args, ctx) => {
const query = String(args ?? "").trim();
if (!query) {
ctx.ui.notify("Usage: /deepsearch ", "warning");
return;
}
pi.sendUserMessage(
[
"请对下面的问题做 DeepSearch。",
"",
`问题:${query}`,
"",
"要求:",
"1. 先判断是否需要调用 deep_search。",
"2. 如果问题较复杂,先拆成 2-4 个子问题分别检索。",
"3. 最终答案必须包含来源链接。",
"4. 区分事实、推断和仍不确定的部分。",
"5. 不要把搜索结果原样堆出来,要给出综合判断。"
].join("\n"),
{ deliverAs: "followUp" }
);
}
});
pi.registerTool({
// deep_search tool definition...
});
}
```
这样用户就可以直接输入:
```text
/deepsearch Pi Coding Agent 的 extension 机制适合做哪些能力?
```
`/deepsearch` 不直接搜索,而是给 Pi 发送一条更完整的任务说明。模型会根据说明调用 `deep_search`,再基于结果完成综合。
我更喜欢这种设计,因为它保留了 agent 的判断空间。搜索工具只是证据入口,不是最终答案生成器。
## Step 4: 启动和验证
项目本地 extension 放好以后,可以直接在项目根目录启动 Pi:
```bash
TAVILY_API_KEY=tvly-... pi
```
如果你只是临时测试,也可以显式指定 extension:
```bash
TAVILY_API_KEY=tvly-... pi -e ./.pi/extensions/deepsearch/index.ts
```
进入 Pi 后,先问一个需要外部事实的问题:
```text
/deepsearch Pi Coding Agent 最新版本的 extension 系统支持哪些能力?
```
一个可接受的输出不应该只是几条搜索结果,而应该包含:
| 检查点 | 合格表现 |
| ------- | ---------------------- |
| 是否调用工具 | 能看到 `deep_search` 被调用 |
| 来源是否清楚 | 每个关键事实后面有 URL |
| 是否去重 | 不重复引用同一个页面 |
| 是否有判断 | 不只罗列资料,还能归纳适用场景 |
| 是否有不确定性 | 对版本变化、第三方 API、社区扩展保持边界 |
如果结果只是"搜索结果列表",说明 promptGuidelines 不够强。可以把 guideline 改得更明确:
```typescript
promptGuidelines: [
"Use deep_search to gather evidence, not to produce the final answer.",
"After deep_search, write a concise research brief with citations.",
"Prefer official documentation, source code, release notes, and primary sources.",
"Mention when sources disagree or when the evidence is incomplete."
]
```
## Step 5: 让 DeepSearch 更像研究工具
跑通第一版以后,可以继续加三类能力。
### 子问题拆解
DeepSearch 最容易失败的地方,是把一个大问题直接丢给搜索 API。
比如:
```text
Pi Agent 能不能替代 Claude Code?
```
这不是一个好 search query。它至少可以拆成:
| 子问题 | 作用 |
| -------------------- | ----- |
| Pi Agent 的核心设计是什么 | 找定位 |
| Pi Agent 支持哪些工具和扩展 | 找能力边界 |
| Claude Code 的默认能力有哪些 | 找对比对象 |
| 两者在权限、安全、可扩展性上有什么区别 | 形成判断 |
第一版可以让模型自己拆;第二版可以让 `/deepsearch` 命令强制要求模型先列子问题,再逐个调用 `deep_search`。
### 来源质量分层
DeepSearch 的输出不能只按搜索 API 的分数排序。实际写技术文章时,我会优先看:
| 优先级 | 来源 |
| --- | --------------------- |
| P0 | 官方文档、源码、release note |
| P1 | 作者博客、维护者说明、issue / PR |
| P2 | 高质量教程、技术分析 |
| P3 | 社区讨论、Reddit、X、论坛 |
Extension 可以在 `formatResultsForModel()` 里先标注来源类型:
```typescript
function classifySource(url: string): "official" | "source" | "community" | "other" {
const host = new URL(url).hostname;
if (host === "pi.dev") return "official";
if (host === "github.com") return "source";
if (host.includes("reddit.com")) return "community";
return "other";
}
```
这样模型综合时就不会把社区传言和官方文档放在同一个证据等级上。
### 上下文截断
搜索结果很容易污染上下文。DeepSearch 的工具输出应该少而精。
我的建议是:
| 内容 | 是否放进工具输出 |
| -------------- | ------------------- |
| 标题 | 放 |
| URL | 放 |
| 200-500 字摘要 | 放 |
| 页面全文 | 默认不放 |
| 原始 HTML | 不放 |
| 搜索 API 原始 JSON | 放进 `details`,不要放进正文 |
如果确实需要全文阅读,可以再做第二个工具:
```text
fetch_source(url)
```
这样 DeepSearch 第一步负责找候选来源,第二步只抓最重要的 2-3 个页面。不要一上来把十几个网页全文都塞给模型。
## 常见问题
### 为什么不直接用 bash 跑搜索脚本?
可以,但不如 extension 稳。
用 bash 的问题是:模型每次都要重新决定命令、参数、输出格式和错误处理。Extension 把这些细节固定下来,模型只需要调用 `deep_search`。
### 为什么不把总结也写在 extension 里?
第一版不建议。
如果 extension 自己再调用一个模型做总结,你就会遇到嵌套模型调用、成本统计、上下文漂移和引用责任的问题。更简单的方式是:extension 只返回证据,Pi 当前会话里的模型负责综合。
### 这个 DeepSearch 算不算 MCP?
不算。它是 Pi extension 注册出来的本地工具。
如果你已经有成熟的 MCP 搜索服务器,也可以通过 Pi 的 MCP 相关 package 或 extension 接进来。但这个案例选择直接写 extension,是为了看清 Pi 本身的扩展机制。
### 安全上要注意什么?
至少注意四件事:
| 风险 | 做法 |
| ---------- | ------------------- |
| API Key 泄露 | 只从环境变量读取,不写进仓库 |
| 不可信网页内容 | 不把网页内容当系统指令,只当待核验证据 |
| 搜索结果污染 | 优先官方和源码,降低社区结果权重 |
| 上下文爆炸 | 限制结果数量和摘要长度 |
DeepSearch 看起来是"搜索增强",本质上是让外部网页进入 agent 上下文。只要外部内容进入上下文,就要把 prompt injection 当成真实风险。
## 小结
我会把 Pi Agent 的第一个实战案例定为 **DeepSearch Extension**,因为它能同时体现 Pi 的三个关键特点:
* Pi 的核心默认很小,不内置所有工作流。
* 真正有用的能力可以通过 extension 补上。
* Extension 不只是加命令,更是在定义模型进入外部世界的边界。
这个案例跑通以后,Pi 就不只是一个本地代码编辑 agent,而是有了一个可控的研究入口:遇到需要外部资料的问题,它可以先检索、再筛选、再带来源地回答。
这比让模型凭记忆回答更可靠,也比每次手写搜索命令更可复用。
## 参考文档
# Ralph Wiggum 深度解析
## 引言
下班前给 AI 布置任务,第二天早上收获可用的代码——这个梦想听起来需要复杂的 Agent 集群、精妙的编排系统。但 2025 年最火的 AI 编程技术,核心就是这一行:
```bash
while :; do cat PROMPT.md | claude ; done
```
一个无限循环,反复把任务喂给 Claude。这就是 **Ralph Wiggum**。它简单到令人尴尬,但确实有人用它花 $297 完成了原本报价 $50,000 的项目。
为什么这么简单的方法反而有效?Anthropic 发布官方插件后,发明者 Geoffrey Huntley 却说"This isn't it"——这又是怎么回事?
## 什么是 Ralph
名字来自《辛普森一家》里的角色。Ralph Wiggum 是警察局长的儿子,全剧最"单纯"的人——他不太清楚自己在做什么,但永远不会停下来。他的标志性台词"I'm helping!"意外地揭示了这项技术的精髓:**天真且不懈的坚持**(Naive and relentless persistence)。
这里有个重要区分:**Ralph 是方法论,不是工具**。就像"敏捷开发"是方法论而不是某个软件,Ralph 描述的是一种工作方式。不同的实现效果可能差异巨大——后面会详细讨论这个问题。
## 为什么需要 Ralph:Context Rot 问题
要理解 Ralph 为什么有效,先要理解它解决的问题。
### AI 是怎么"变笨"的
用 Claude 处理复杂任务时,你可能有过这样的体验:刚开始对话很顺畅,Claude 理解准确、执行到位。但随着对话越来越长,它开始"迟钝"——忘记重要信息,重复犯同样的错误,代码质量下降,甚至开始产生莫名其妙的"幻觉"。
这不是 AI 不够聪明。问题在于**上下文窗口被污染了**。
想象这个场景:你让 Claude 写一个功能,第一次失败了。你说"修复这个",它尝试但又失败。来回十次之后,Claude 的上下文里塞满了:九次失败的代码、九组错误信息、大量不再相关的讨论。在这些杂乱信息中找到重点,难度越来越大。
### Dumb Zone
Geoffrey Huntley 和社区开发者发现了一个现象,称之为"Dumb Zone":
| 上下文大小 | 表现 |
| ----------------- | ----------- |
| 0 - 50k tokens | 最佳性能 |
| 50k - 100k tokens | 良好,轻微下降 |
| 100k+ tokens | 明显退化,开始忽略指令 |
| 150k+ tokens | 严重退化 |
没有精确临界点,但经验法则:**上下文用到一半左右就该警惕了**。对于 200k tokens 的 Claude,超过 100k 时你可能在和一个"变笨"的 AI 交流。
### 累积的上下文是负债
这里有个反直觉的洞察:累积的上下文不是资产,而是负债。
我们习惯认为记忆力越好越好,保留的信息越多越好。但在大语言模型的世界里,这个直觉是错的。对话越长,上下文里充斥的"负面信息"越多:失败的代码、不再相关的讨论、被纠正过的错误理解。这些不仅占用空间,还分散 AI 的"注意力"。
## Ralph 的工作原理
理解了 Context Rot,Ralph 的解决方案就清晰了:**既然累积上下文是问题,那就不要累积**。
Ralph 建立在三个支柱上:
### 1. 新 Session
每次循环迭代时,启动一个**全新的 Claude 实例**,获得完全干净的上下文窗口。不是"清空对话历史"——那样累积的状态可能仍然存在。而是彻底关闭当前进程,启动一个新的。
这意味着每次迭代开始时,Claude 都处于最佳状态。没有之前的错误来困扰它,没有过时的讨论来分散注意力。
**这就是为什么循环必须在 Claude Code 外部运行**——bash 循环需要能够控制 Claude 进程的生命周期。
### 2. 文件作为真相来源
如果每次都是新上下文,AI 怎么知道之前做了什么?答案:通过文件系统,而不是对话历史。
关键文件:
* **PRD/spec 文件** — 定义目标、功能列表、成功标准
* **IMPLEMENTATION\_PLAN.md** — 任务分解和进度
* **progress.txt** — 自由格式日志,每次迭代结束时追加学到的内容
* **Git 历史** — 代码修改的证明
每次迭代开始时,Claude 读取这些文件了解目标和进度。它看到的是精心组织的状态快照,而不是混乱的对话历史。
### 3. 反馈循环
仅有干净上下文和持久化状态还不够。如果 AI 写了有问题的代码并提交了,错误会累积。
反馈循环是自动化的质量门槛:
* **TypeScript 类型检查** — 即时反馈类型正确性
* **单元测试** — 验证功能符合预期
* **CI/CD** — 确保代码可以构建和集成
如果测试失败,代码不会被提交,Claude 会看到失败信息。下一次迭代的新 Claude 实例会尝试修复问题。
> 关于如何建立完整的质量保障体系,我在 [我的 Claude Code 质量检查流程](/blog/claude-code-quality-control) 中分享了五层防线的实践经验:Hooks 自动化、测试策略、AI Review、Pre-commit、GitHub 集成。
## Human on the Loop
Geoffrey Huntley 反复强调一个概念区别:
| Human **in** the Loop | Human **on** the Loop |
| --------------------- | --------------------- |
| 保姆式陪伴 | 监督式管理 |
| AI 每一步都等你确认 | 你设定目标和边界,AI 自主运行 |
| 你是工作流程的瓶颈 | 你偶尔检查进度 |
实际使用中有两种模式:
* **AFK 模式**:下班前启动,回家睡觉,早上检查结果
* **Human-in-the-loop 模式**:每次迭代后暂停检查,适合复杂或不确定的任务
## 什么任务适合 Ralph
Ralph 不是万能的。它的核心优势是"迭代直到成功",这决定了它适合特定类型的任务。
### 适合的任务
| 场景 | 原因 |
| ---------- | --------------------- |
| 有明确成功标准的任务 | 可以自动验证完成(测试通过、类型检查通过) |
| 需要迭代改进的任务 | Ralph 的核心优势就是不断尝试 |
| 绿地项目 | 无需担心破坏现有代码 |
| 有自动测试的项目 | 测试作为反压机制,确保质量 |
### 不适合的任务
| 场景 | 原因 |
| ----------- | ------------------- |
| 需要人工判断的设计决策 | 无法自动验证"好不好看" |
| 一次性操作 | 不需要迭代的任务用 Ralph 是浪费 |
| 生产环境调试 | 风险太高,不适合无人值守 |
| 成功标准不清晰的任务 | 无法判断何时停止 |
### 三种使用模式
**完整实现模式**
这是 Ralph 最常见的用法:从头构建一个完整的功能或项目。你准备好 spec 文件和实施计划,让 Ralph 自动执行所有任务。
典型场景:
* 构建一个新的 REST API
* 开发一个 CLI 工具
* 实现一个新功能模块
真实案例:有开发者用这种模式完成了一个价值 $50,000 的外包项目,总 API 成本仅 $297。整个过程包括 MVP 开发、测试编写和代码审查,全程自动化。另一个案例是将一个老旧代码库从 React v16 升级到 v19,Ralph 跑了 14 小时,完全无需人工干预。
**探索模式**
不是所有任务都需要产出代码。有时候你需要的是理解——理解一个新接手的代码库、理解一个复杂系统的架构、理解某个模块的工作原理。
典型场景:
* 接手一个陌生的项目,需要快速建立整体认知
* 为现有代码库生成文档
* 分析系统架构,找出潜在问题
在这种模式下,你的 prompt 不是"实现 X 功能",而是"阅读这个代码库,生成架构文档"或"找出所有 API 端点并说明它们的作用"。Claude 每次迭代都会深入探索,逐步建立更完整的理解。
**暴力测试模式**
有些 bug 你知道症状、知道期望的正确行为,但就是找不到根本原因。这时候可以让 Ralph 来"暴力破解"。
典型场景:
* 一个间歇性出现的 bug,难以复现
* 某个测试偶尔失败,原因不明
* 性能问题,不确定瓶颈在哪里
设置目标:"修复这个 bug,让这个测试稳定通过"。Ralph 会不断尝试不同的修复方案,直到找到有效的那个。这种方法特别适合那些"我不知道怎么修,但我知道什么时候算修好了"的问题。
## 实现方式的选择
理解了 Ralph 的方法论,接下来面临一个实际问题:如何实现这个循环?
社区发展出了两种工程化程度不同的实现:
**极简路线**——[snarktank/ralph](/docs/notes/ralph-wiggum/snarktank):几百行 bash 脚本,每次全新会话,专注于循环本身。轻量、易上手,适合快速开始。
**工程化路线**——[frankbria/ralph-claude-code](/docs/notes/ralph-wiggum/frankbria):完整工具链(监控仪表盘、断路器、速率限制、会话过期管理)。默认通过 `--continue` 复用会话,也可通过 `--no-continue` 切换为全新会话模式。
| 维度 | 极简(snarktank) | 工程化(frankbria) |
| ----- | --------------- | --------------- |
| 会话模式 | 每次全新 | 默认复用,可切换为全新 |
| 监控 | 手动查看 | 内置 tmux 仪表盘 |
| 安全机制 | max\_iterations | 断路器 + 速率限制 + 超时 |
| 安装复杂度 | Skill 复制 | install.sh + 向导 |
两种实现各有优劣,选择取决于你对工程化工具的需求。详细的使用方法和对比分析见各自的实战文章。
## 写在最后
Ralph 教给我们一个重要的道理:有时候最简单的方法反而最有效。当所有人都在追求更复杂的架构时,一个 bash 循环改变了游戏规则。
当然,Ralph 只是拼图的一部分。它需要好的 prompt、合适的项目、正确的反馈机制才能发挥威力。了解了原理之后,可以根据你的需求选择合适的实现方式:
* 需要长期 AFK、大量迭代?→ 《[snarktank/ralph 实战指南](/docs/notes/ralph-wiggum/snarktank)》
* 需要工程化监控和安全机制?→ 《[frankbria/ralph-claude-code 实战指南](/docs/notes/ralph-wiggum/frankbria)》
***
**相关阅读**:
* [snarktank/ralph 实战指南](/docs/notes/ralph-wiggum/snarktank) — 极简外部循环,从安装到实战的完整操作手册
* [frankbria/ralph-claude-code 实战指南](/docs/notes/ralph-wiggum/frankbria) — 工程化实现:监控、断路器与安全机制
* [Claude Subagent 完全指南](/docs/notes/claude-subagent) — 另一种保持上下文清洁的方式
* [Claude Skills 是什么](/docs/notes/claude-skills/concept) — 探索 Claude 的可重用工作手册
* [GSD 深度解析](/docs/notes/gsd/concept) — 在 Ralph 基础上构建的完整上下文工程系统
* [Claude 系统架构全解析](/docs/notes/claude-architecture) — 理解 Hooks、Subagent 等组件的整体架构
# frankbria/ralph-claude-code 实战指南
## 引言
[上一篇](/docs/notes/ralph-wiggum/concept)介绍了 Ralph 的方法论,[snarktank/ralph](/docs/notes/ralph-wiggum/snarktank) 展示了一种极简的外部循环实现。现在来看另一条路线:[frankbria/ralph-claude-code](https://github.com/frankbria/ralph-claude-code)。
如果说 snarktank/ralph 的哲学是"用最少的代码做最多的事",那 frankbria 的哲学就是"**工程化一切**"——交互式配置向导、实时监控仪表盘、断路器、速率限制、会话过期管理。它不追求简洁,而是追求**可控**。
两种实现没有优劣之分,适合不同的使用场景。本文带你了解 frankbria 的完整工具链。
## 安装与配置
### 全局安装
```bash
# 克隆仓库
git clone https://github.com/frankbria/ralph-claude-code.git
cd ralph-claude-code
# 全局安装
./install.sh
```
安装完成后,你会获得以下全局命令:
| 命令 | 说明 |
| --------------- | -------------- |
| `ralph` | 启动 Ralph 循环 |
| `ralph-enable` | 在现有项目中启用 Ralph |
| `ralph-setup` | 创建新项目并配置 Ralph |
| `ralph-import` | 导入现有 PRD/需求文档 |
| `ralph-monitor` | 启动实时监控仪表盘 |
### 项目初始化
对于已有项目,使用交互式向导:
```bash
cd your-project
ralph-enable
```
向导会自动检测项目类型(Node.js、Python、Go 等)和框架(Next.js、FastAPI 等),然后生成对应的配置文件。
对于全新项目:
```bash
ralph-setup my-new-project
```
这会创建项目目录、初始化 Git、生成 `.ralph/` 配置目录。
### 导入现有需求
如果你已经有 PRD 文档或需求说明:
```bash
ralph-import path/to/your-prd.md
```
Ralph 会解析文档,提取任务列表,生成结构化的 `fix_plan.md`。
## .ralph/ 目录结构
frankbria 的记忆和配置集中在 `.ralph/` 目录下:
```
.ralph/
├── PROMPT.md # 项目目标和上下文
├── fix_plan.md # 任务清单(类似 prd.json 的作用)
├── AGENT.md # 构建/测试命令(自动维护)
├── specs/ # 详细需求文档
│ ├── feature-a.md
│ └── feature-b.md
└── sessions/ # 会话持续性数据
├── current.json
└── history/
```
**与 snarktank/ralph 的对比**:
| frankbria | snarktank | 作用 |
| ------------- | ------------------------------------------ | --------- |
| `PROMPT.md` | `prd.json` 的 `projectName` + `description` | 定义项目目标 |
| `fix_plan.md` | `prd.json` 的 `userStories` | 任务列表和进度 |
| `AGENT.md` | `CLAUDE.md` / `AGENTS.md` | 构建命令和项目约定 |
| `specs/` | `prd.json` 的 `notes` 字段 | 详细需求 |
| `sessions/` | 无(每次新进程) | 会话状态追踪 |
注意 `AGENT.md` 是**自动维护**的——Ralph 在执行过程中会根据发现的项目约定自动更新这个文件,类似于 snarktank/ralph 的 `progress.txt`,但更结构化。
## 核心命令
### 基本执行
```bash
# 启动 Ralph 循环
ralph
# 带实时监控
ralph --monitor
# 在 tmux 中启动(推荐长时间运行)
ralph --live
```
### 监控仪表盘
```bash
# 独立启动监控
ralph-monitor
```
`ralph-monitor` 会打开一个 tmux 仪表盘,实时显示:
* 当前正在执行的任务
* 已完成/未完成的任务计数
* API 调用次数和成本估算
* 断路器状态
* 最近的错误日志
### 常用参数
| 参数 | 说明 | 默认值 |
| ----------------- | ----------- | ----- |
| `--resume` | 从上次中断处继续 | - |
| `--calls ` | 最大 API 调用次数 | 100 |
| `--timeout ` | 超时时间(分钟) | 300 |
| `--monitor` | 启用实时监控 | false |
| `--live` | 在 tmux 中运行 | false |
```bash
# 限制 50 次 API 调用,2 小时超时
ralph --calls 50 --timeout 120
# 从上次中断处继续
ralph --resume
```
## 安全机制
frankbria 最大的差异化特性是其多层安全机制。
### 断路器(Circuit Breaker)
断路器会在检测到"无进展"时自动停止循环,防止无意义的 API 消耗:
**连续无进展检测**:如果连续 N 次迭代都没有新的任务完成,断路器触发。
**相同错误检测**:如果连续出现相同的错误信息,说明 AI 陷入了死循环,断路器触发。
### 速率限制
默认限制 100 calls/hour,防止意外的 API 账单爆炸。可以通过参数调整:
```bash
ralph --calls 200 # 提高到 200 calls
```
### 5 小时 API 限额三层检测
Anthropic API 有 5 小时滑动窗口的使用限额。frankbria 内置了三层检测:
1. **预检测**:在每次 API 调用前估算剩余额度
2. **响应检测**:解析 API 响应中的 rate limit headers
3. **回退策略**:接近限额时自动降低调用频率
### 会话过期管理
默认会话有效期 24 小时。超过后自动清理会话数据,防止过期上下文影响后续执行。
## 智能退出检测
frankbria 不是简单地在所有任务完成后退出。它使用了**双条件退出门**:
```
退出条件 = completion_indicators >= 2 AND EXIT_SIGNAL: true
```
**completion\_indicators** 是从 AI 输出中检测到的完成信号数量,包括:
* "所有任务已完成"
* "没有更多待办事项"
* 测试全部通过
* fix\_plan.md 中所有条目标记为 done
**EXIT\_SIGNAL** 是 AI 在输出中显式声明的退出意图。
为什么需要两个条件?防止**过早退出**。单一信号可能是误判——比如 AI 说"任务完成"但实际上只完成了当前 story。双条件确保只有在多个独立信号都确认完成时才真正退出。
## 与 snarktank/ralph 的对比
| 维度 | snarktank/ralph | frankbria/ralph-claude-code |
| -------- | ----------------- | ------------------------------- |
| **实现方式** | 外部 bash 循环(每次新会话) | 外部 bash 循环(`--continue` 复用会话) |
| **会话模式** | 每次全新 | 默认复用(可通过 `--no-continue` 切换为全新) |
| **上下文** | 每次全新 | 通过 `--continue` 跨迭代累积 |
| **安装** | Skill 复制 | install.sh + 交互向导 |
| **任务格式** | prd.json | PROMPT.md + fix\_plan.md |
| **监控** | 手动 `cat`/`jq` | 内置 tmux 仪表盘 |
| **安全机制** | max\_iterations | 断路器 + 速率限制 + 超时 |
| **任务来源** | PRD only | beads / GitHub Issues / PRD |
| **适合场景** | 长期 AFK、大量迭代 | 短中期迭代、需要监控 |
### 核心差异:工程化程度
两者都是外部 bash 循环启动新的 Claude 进程。核心区别不在会话管理方式(frankbria 可以通过 `--no-continue` 切换为全新会话模式),而在**工程化程度**:
* **snarktank**:极简脚本,几百行 bash,专注于循环本身
* **frankbria**:完整工程化工具链——监控仪表盘、断路器、速率限制、会话过期管理
frankbria 默认开启 `--continue` 复用会话,适合短任务。对长任务,可以用 `--no-continue` 切换为全新会话模式,获得和 snarktank 一样的 Context Rot 防护,同时保留 frankbria 的工程化优势。
### 如何禁用会话复用
frankbria 提供三种方式禁用 `--continue`:
```bash
# 方式一:命令行参数
ralph --no-continue
# 方式二:环境变量
export CLAUDE_USE_CONTINUE=false
# 方式三:.ralphrc 配置
SESSION_CONTINUITY=false
```
禁用后,frankbria 的行为等同于 snarktank(每次全新会话),但保留所有工程化工具(监控、断路器、速率限制等)。
## Context Rot 的现实权衡
会话复用方式的选择,本质上是在 Context Rot 和启动开销之间做权衡:
**短任务(\< 50k tokens)**:复用会话更有优势。上下文还没来得及退化,前几次迭代的记忆还能被后续利用。每次新建会话的启动开销反而是浪费。
**长任务(100k+ tokens)**:全新会话更可靠。超过 100k tokens 后 Context Rot 明显加剧,累积的上下文从资产变成负债。全新会话虽然有启动开销,但每次都是最佳状态。
**实际建议**:
| 场景 | 推荐 | 原因 |
| -------------- | --------------------------------------- | ---------------------- |
| 5 个以下小任务 | frankbria(默认模式) | 启动快、上下文可复用 |
| 10+ 个任务、需要 AFK | snarktank 或 frankbria + `--no-continue` | 避免 Context Rot、更可靠 |
| 需要实时监控 | frankbria | 内置仪表盘 |
| 不确定任务量 | frankbria + `--no-continue` | 工程化工具 + Context Rot 防护 |
frankbria 用户可以根据任务规模灵活选择:短任务用默认的 `--continue` 模式,长任务切换到 `--no-continue` 模式。相比 snarktank,frankbria 的优势在于无论哪种模式都保留完整的工程化工具链。
## 小结
frankbria/ralph-claude-code 代表了 Ralph 方法论的工程化实现路线。它牺牲了一些 snarktank 的简洁性,换来了更完善的监控、安全和配置能力。
选择哪个实现取决于你的具体需求——没有"更正确"的答案,只有"更适合"的选择。
***
**延伸阅读**:
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) — 核心原理与方法论
* [snarktank/ralph 实战指南](/docs/notes/ralph-wiggum/snarktank) — 极简外部循环实现
* [GSD 深度解析](/docs/notes/gsd/concept) — 在 Ralph 基础上构建的完整上下文工程系统
* [Claude 系统架构全解析](/docs/notes/claude-architecture) — 理解 Hooks、Subagent 等组件的整体架构
# Ralph 实战指南
## 引言
在[上一篇文章](/docs/notes/ralph-wiggum/concept)中,我们理解了 Ralph 的核心原理——无限循环 + 每次全新上下文 + 文件作为唯一真相源。这三个支柱听起来简单,但从理解概念到真正跑起来,中间有不少细节需要打通。
这篇文章,我们来动手。你将学会如何使用 [snarktank/ralph](https://github.com/snarktank/ralph) 完成从安装到执行的完整流程。snarktank/ralph 是社区中最成熟的 Ralph 实现之一(10k+ stars),支持 Claude Code 和 Amp 两种工具,配套 PRD 生成、JSON 转换、自动化执行的完整工具链。
## 前置条件
开始之前,确保你的环境满足以下要求:
| 依赖 | 说明 |
| ----------- | ------------------------------------------------------------------ |
| **AI 编程工具** | Claude Code (`npm install -g @anthropic-ai/claude-code`) 或 Amp CLI |
| **jq** | JSON 处理工具 (macOS: `brew install jq`) |
| **Git** | 项目需要是 Git 仓库 |
```bash
# 检查依赖
claude --version # Claude Code CLI
jq --version # JSON 处理
git --version # Git
```
## 安装与配置
snarktank/ralph 提供多种安装方式,根据使用场景选择。
### 方式一:直接在 Claude Code 中安装(推荐)
最简单的方式——在 Claude Code 对话中粘贴 GitHub 链接,让 Claude 自动完成安装:
```
Install this skill for me: https://github.com/snarktank/ralph
```
Claude Code 会自动克隆仓库并将 skill 文件复制到正确位置。安装完成后即可使用 `/prd` 和 `/ralph` 命令。
### 方式二:Claude Code 市场安装
通过市场命令安装:
```bash
# 添加并安装插件
/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace
```
安装后可使用 `/prd`(生成 PRD)和 `/ralph`(转换为 JSON)两个 skill。
### 方式三:手动 Skill 安装(Claude Code / Amp)
手动将 skill 文件复制到对应工具的全局配置目录:
```bash
# 先克隆仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# Claude Code 用户
cp -r /tmp/ralph/skills/prd ~/.claude/skills/
cp -r /tmp/ralph/skills/ralph ~/.claude/skills/
# Amp 用户
cp -r /tmp/ralph/skills/prd ~/.config/amp/skills/
cp -r /tmp/ralph/skills/ralph ~/.config/amp/skills/
```
安装完成后即可使用 `/prd` 和 `/ralph` 命令。
### 方式四:项目级安装
将 Ralph 脚本直接复制到项目中——适合需要团队共享或自定义脚本的场景:
```bash
# 克隆 Ralph 仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# 复制核心文件到项目
mkdir -p scripts/ralph
cp /tmp/ralph/ralph.sh scripts/ralph/
cp /tmp/ralph/CLAUDE.md scripts/ralph/ # Claude Code 用户
# 或
cp /tmp/ralph/prompt.md scripts/ralph/ # Amp 用户
# 赋予执行权限
chmod +x scripts/ralph/ralph.sh
```
安装完成后,项目结构如下:
```
your-project/
├── scripts/ralph/
│ ├── ralph.sh # 核心循环脚本
│ └── CLAUDE.md # Claude Code 的 Prompt 模板
├── tasks/ # PRD 文件目录(执行时自动创建)
│ └── prd.json # 你的任务定义
└── ...
```
> **建议**:方式一最省事——把 GitHub 链接丢给 Claude Code 就行。如果想手动控制安装过程,选方式二(市场命令)或方式三(手动复制)。需要团队共享或自定义脚本时选方式四。
***
## 核心文件结构
Ralph 的记忆完全依赖文件系统。理解每个文件的角色,是用好 Ralph 的前提。
### ralph.sh —— 循环引擎
这是 Ralph 的核心:一个 bash 脚本,不断生成新的 AI 实例。
```bash
# 基本用法
./scripts/ralph/ralph.sh [max_iterations] # 默认:Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # 使用 Claude Code
```
每次迭代,ralph.sh 会执行以下步骤:
1. 创建功能分支(来自 prd.json 中的 `branchName`)
2. 选择优先级最高的未完成 story(`passes: false`)
3. 生成一个**全新的** AI 实例来实现这个 story
4. 运行质量检查(类型检查、测试)
5. 检查通过 → git commit;检查失败 → 留给下次迭代
6. 更新 prd.json,将该 story 标记为 `passes: true`
7. 将学到的经验追加到 progress.txt
8. 重复,直到所有 story 完成或达到迭代次数上限
默认迭代上限为 10 次。根据项目复杂度调整:
```bash
# 简单项目
./scripts/ralph/ralph.sh --tool claude 10
# 复杂项目
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json —— 任务定义
这是 Ralph 的"大脑"——所有任务都定义在这里。它是一个扁平的 JSON 文件:
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists with valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve MDX component props format"
}
]
}
```
**字段说明**:
| 字段 | 说明 |
| -------------------- | ------------------------- |
| `projectName` | 项目名称,用于日志和分支命名 |
| `branchName` | Git 分支名——Ralph 会自动创建 |
| `id` | Story 唯一标识,推荐 `US-001` 格式 |
| `title` | 简短标题 |
| `description` | 详细描述——越具体越好 |
| `acceptanceCriteria` | 验收标准列表——**这是最关键的字段** |
| `priority` | 优先级数字——数字越小越先执行 |
| `passes` | 是否完成——Ralph 会自动更新 |
| `dependsOn` | 依赖的 story ID 列表 |
| `notes` | 额外的提示和上下文 |
### progress.txt —— 经验日志
这是 Ralph 的"长期记忆"。每次迭代后,AI 会追加学到的经验:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
下一次迭代的全新 Claude 实例会读取这个文件,立即获得之前所有的经验。这就是 Ralph 越跑越顺的原因——**知识在迭代间积累,而上下文保持干净**。
### AGENTS.md —— 持久知识库
除了 progress.txt,Ralph 还会更新项目中的 `AGENTS.md`(或 `CLAUDE.md`)文件。Claude Code 和 Amp 启动时都会自动读取这些文件。
与 progress.txt 不同,AGENTS.md 记录的是**稳定的、跨项目的知识**:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## 编写 PRD
PRD(产品需求文档)的质量直接决定 Ralph 的执行效果。写得好,Ralph 一路畅通。写得差,Ralph 会在同一个 story 上反复失败。
### 用 Skill 生成 PRD
如果你安装了 snarktank/ralph 的 skill,可以交互式生成 PRD:
```bash
# 在 Claude Code 或 Amp 中
/prd I want to add i18n support to the blog, translating all Chinese content to English
```
AI 会问你一些澄清问题(涉及哪些文件、技术栈限制、质量标准等),然后生成结构化的 PRD 文档。
生成后,用 `/ralph` 命令将 PRD 转换为 `prd.json` 格式:
```bash
/ralph # 转换 PRD 为 prd.json
```
### 手动编写 PRD
你也可以直接编写 prd.json。以下是关键的设计原则。
**原则一:Story 粒度适中**
每个 story 应该小到能在一次迭代中完成,大到能独立交付价值。
```json
// ❌ 太大:一次迭代完不成
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, forgot password, OAuth, permission management..."
}
// ❌ 太小:没有独立价值
{
"id": "US-001",
"title": "Create email field on User table",
"description": "Add email field to User model"
}
// ✅ 刚好:一次迭代能完成,有独立价值
{
"id": "US-001",
"title": "Implement email/password login",
"description": "Create login API and login page with email/password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form submits successfully",
"All tests pass"
]
}
```
**经验法则**:一个 story 涉及 1-3 个文件修改,有 3-5 条验收标准。
**原则二:验收标准必须可自动验证**
Ralph 需要判断 story 是否完成,因此验收标准必须是可客观评估的:
```json
// ❌ 模糊的标准
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ 可验证的标准
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**原则三:用 dependsOn 控制执行顺序**
有些 story 存在依赖关系。`dependsOn` 字段确保 Ralph 按正确顺序执行:
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**原则四:在 notes 中提供上下文**
notes 字段给 AI 额外提示。把你知道但 AI 不一定知道的信息写在这里:
```json
{
"notes": "Project uses fumadocs framework, i18n files follow .en.mdx suffix naming. Reference content/docs/notes/speckit/concept.en.mdx for translation style."
}
```
***
## 执行 Ralph 循环
PRD 准备好后,就可以运行循环了。
### 启动执行
```bash
# 使用 Claude Code,默认 10 次迭代
./scripts/ralph/ralph.sh --tool claude
# 指定迭代次数
./scripts/ralph/ralph.sh --tool claude 30
# 使用 Amp(默认)
./scripts/ralph/ralph.sh 20
```
### 执行过程
启动后,你会看到类似这样的输出:
```
=== Ralph Loop - Iteration 1 ===
Branch: ralph/i18n-translation
Selected story: US-001 - Translate homepage metadata
Spawning fresh Claude instance...
[Claude Code executing...]
Quality check: pnpm types:check ... PASSED
Committing: feat: [US-001] - Translate homepage metadata
Updating prd.json: US-001 passes: true
Appending to progress.txt
=== Ralph Loop - Iteration 2 ===
Selected story: US-002 - Translate blog post hello-world
Spawning fresh Claude instance...
```
每次迭代都是全新的 Claude 实例。它通过读取 prd.json 知道要做什么,通过读取 progress.txt 知道之前学到了什么。
### 完成信号
当所有 story 都标记为 `passes: true` 时,Ralph 输出完成信号并退出:
```
All stories completed!
COMPLETE
```
### 监控与调试
Ralph 运行期间,可以用以下命令查看进度:
```bash
# 查看每个 story 的完成状态
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 查看经验日志
cat progress.txt
# 查看最近的 git 提交
git log --oneline -10
# 实时跟踪 Ralph 输出
tail -f progress.txt
```
### 自动归档
当你启动新功能(使用不同的 `branchName`)时,Ralph 会自动将上次运行的文件归档到 `archive/YYYY-MM-DD-feature-name/` 目录,保持工作目录整洁。
***
## 反馈循环与质量门禁
Ralph 的"自我纠错"能力完全取决于反馈循环的质量。没有反馈循环,Ralph 只是一个盲目循环的脚本——它会不停地产出代码,但无法判断代码是否正确。
### 配置质量检查
在 CLAUDE.md(或 prompt.md)中定义质量检查命令:
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 质量门禁层级
| 层级 | 工具 | 捕获的问题 |
| ----- | ----------------- | ------------- |
| 即时反馈 | TypeScript 编译器 | 类型错误、语法错误 |
| 功能验证 | 单元测试 | 逻辑错误、边界情况 |
| 集成验证 | 构建命令 | 依赖问题、配置错误 |
| 运行时验证 | dev-browser skill | UI 渲染问题(前端项目) |
> 对于前端 story,Ralph 建议添加这条验收标准:"使用 dev-browser skill 在浏览器中验证"——让 AI 真正打开浏览器确认页面渲染正确。
### 质量检查失败时
如果某个 story 的质量检查反复失败,Ralph 不会无限重试同一个 story。达到迭代上限后会停止,保留当前状态。你可以:
1. 查看 progress.txt 了解卡住的原因
2. 手动修复问题后重新运行
3. 调整 story 粒度(可能太大了)
4. 在 notes 字段中补充更多上下文
***
## Prompt 定制
Ralph 的 prompt 模板(CLAUDE.md 或 prompt.md)是你控制 AI 行为的主要手段。安装后,你应该根据项目进行定制。
### 关键定制项
**1. 项目特定的质量命令**
```markdown
## Project-Specific Commands
- Typecheck: `pnpm types:check` (not `tsc` or `pnpm typecheck`)
- Test: `pnpm vitest run`
- Build: `pnpm build`
- Lint: `pnpm lint`
```
**2. 代码风格约束**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**3. 已知的坑**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**4. 卡住时的处理**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 实战案例:用 Ralph 翻译博客
为了展示 Ralph 在实际中如何运作,这里是一个真实案例:使用 Ralph 风格的自主 Agent 将整个博客从中文翻译成英文。
### 项目设置
该项目需要将 22+ 个内容文件(博客文章、文档、导航元数据)从中文翻译成英文,项目基于 fumadocs 的 Next.js 博客,支持 i18n。任务定义在 `prd.json` 文件中,包含 16 个 user story,每个都有明确的验收标准:
```
scripts/ralph/
├── prd.json # 16 个 user story,带验收标准
└── progress.txt # 经验日志,每个 story 完成后更新
```
每个 user story 遵循一致的模式:
* **明确的交付物**:"Create content/blog/xxx.en.mdx"
* **可验证的标准**:"Typecheck passes"、"Internal links use /en/ prefix"
* **技术约束**:"Keep code blocks untranslated"、"Set defaultLang='en' on QuoteCard"
### 执行模式
Agent 遵循 Ralph 方法论的核心原则:
1. **文件即真相源**:`prd.json` 追踪哪些 story 通过了(`passes: true/false`)。`progress.txt` 在迭代间积累经验——例如"Typecheck command is `pnpm types:check`, not `pnpm typecheck`"
2. **自动化质量门禁**:每次翻译后,`pnpm types:check` 运行以验证 MDX 文件编译正确。如果类型检查失败,先修复再提交。
3. **增量推进**:每个 story 独立提交,带描述性的提交信息(`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`),需要时可以轻松回滚。
4. **并行执行**:对于较长的文章,多个子 Agent 同时翻译——例如 US-010(claude-skills concept + practice)、US-011(speckit concept + practice)和 US-012(claude-architecture + claude-subagent)并行运行。
### 关键经验
| 经验 | 详情 |
| -------------------- | ------------------------------------------------------------ |
| **积累的知识很重要** | 早期 story 发现的模式(QuoteCard `defaultLang`、链接前缀规则)让后续 story 更快完成 |
| **Typecheck 作为反馈循环** | 在问题叠加之前捕获缺失导入或格式错误的 MDX |
| **并行化可扩展** | 6 个翻译 Agent 同时运行,完成时间与 1 个 Agent 大致相同 |
| **PRD 粒度至关重要** | 每个 story 范围为 1-2 个文件——小到能可靠完成,大到有意义 |
| **进度日志防止重复犯错** | progress.txt 的"Codebase Patterns"部分成为知识库,防止重新发现相同问题 |
### 成果
16 个 user story 在单次会话中全部完成:创建了 8 个 meta.en.json 导航文件,翻译了 3 篇博客文章,翻译了 12 个文档页面,完整站点构建验证通过。每个翻译保持了一致的质量,因为验收标准是明确的,反馈循环(typecheck)能即时捕获问题。
这个项目展示了 Ralph 的**完整实现模式**——定义明确的任务 + 清晰的成功标准 + 自动化验证 + 通过文件系统增量交付。
***
## 社区实现与替代方案
snarktank/ralph 不是唯一选择。根据需求不同,这些实现各有所长:
| 资源 | 链接 | 说明 |
| ------------------ | ----------------------------------------------------------------------------------- | -------------------------- |
| snarktank/ralph | [snarktank/ralph](https://github.com/snarktank/ralph) | 本文使用,功能最完整 |
| ralph-orchestrator | [mikeyobrien/ralph-orchestrator](https://github.com/mikeyobrien/ralph-orchestrator) | Mickey O'Brien 开发,有更多自定义选项 |
| ralph-loop-agent | [vercel-labs/ralph-loop-agent](https://github.com/vercel-labs/ralph-loop-agent) | Vercel 基于 AI SDK 的实现 |
| ralphy | [michaelshimeles/ralphy](https://github.com/michaelshimeles/ralphy) | Michael Shimeles 的轻量实现 |
### 替代方案:GSD
GSD 严格来说不是 Ralph 的"社区实现"——它是一种**替代方案**。它运用了 Ralph 的核心原则(上下文管理、原子任务),但提供了更完整的工作流:讨论 → 计划 → 执行 → 验证。
| 资源 | 链接 | 说明 |
| -------------------- | ----------------------------------------------------------------------------- | ----------------- |
| GSD (Get Stuff Done) | [glittercowboy/get-shit-done](https://github.com/glittercowboy/get-shit-done) | 从想法到 PRD 到执行的完整框架 |
如果你觉得 Ralph 太"原始",需要更多流程支持,GSD 可能更适合。详见 [GSD 深度解析](/docs/notes/gsd/concept)。
***
## 推荐资源
**官方来源**:
| 资源 | 链接 | 说明 |
| -------------------- | ------------------------------------------------------------------------------- | -------- |
| Geoffrey Huntley 的博客 | [ghuntley.com/ralph](https://ghuntley.com/ralph/) | 发明者的原始文章 |
| how-to-ralph-wiggum | [ghuntley/how-to-ralph-wiggum](https://github.com/ghuntley/how-to-ralph-wiggum) | 官方使用指南 |
**视频教程**:
| 资源 | 链接 | 说明 |
| ----------------- | ---------------------------------------------------------------------------------------- | -------------------------- |
| Ralph Wiggum 深度讨论 | [Why Claude Code's implementation isn't it](https://www.youtube.com/watch?v=O2bBWDoxO4s) | Geoffrey Huntley 解释官方实现的问题 |
| 正确使用 Ralph | [You're Using Ralph Wiggum Loops WRONG](https://www.youtube.com/watch?v=I7azCAgoUHc) | Roman (Mentat) 的使用演示 |
| 我们需要谈谈 Ralph | [We need to talk about Ralph](https://www.youtube.com/watch?v=Yr9O6KFwbW4) | Theo 对争议的分析 |
***
## 最佳实践与 FAQ
### 成本控制
Ralph 的自动化执行意味着 API 费用在持续产生。几个控制措施:
* **始终设置 `max_iterations`**:这是最基本的安全网
* **保持 story 粒度合理**:太大的 story 消耗多次迭代;太细碎的 story 增加启动开销
* **先小规模测试**:新项目先跑 3-5 次迭代,确认 prompt 和质量门禁工作正常后再扩大
### 常见陷阱
**陷阱一:Story 太大**
症状:某个 story 反复失败,迭代次数快速消耗殆尽。
解决:拆成 2-3 个更小的 story。"构建完整认证系统" 拆成 "实现登录 API" + "创建登录页面" + "添加 JWT 中间件"。
**陷阱二:没有反馈循环**
症状:Ralph 声称 story 完成了,但实际代码有问题。
解决:在验收标准中添加可执行的检查命令。"代码写好了" 不是验收标准——"pnpm test all passes" 才是。
**陷阱三:progress.txt 没被利用**
症状:不同迭代中反复出现相同错误。
解决:确认你的 prompt 模板明确指示"读取 progress.txt 并遵循其中的经验"。如果 AI 没有自动追加经验,在 prompt 中添加"每个 story 完成后,将经验追加到 progress.txt"。
**陷阱四:依赖顺序错误**
症状:某个 story 依赖尚不存在的代码,导致实现失败。
解决:正确设置 `dependsOn` 字段。确保基础设施 story 排在前面。
### FAQ
**Q:Ralph 和官方插件有什么区别?**
核心区别:snarktank/ralph 每次迭代都生成新进程(真正全新的上下文),而官方插件在同一会话内循环(上下文持续累积)。详见[上一篇文章的分析](/docs/notes/ralph-wiggum/concept#the-problem-with-the-official-plugin)。
**Q:执行过程中可以手动修改 prd.json 吗?**
可以。Ralph 在每次迭代开始时重新读取 prd.json。你可以在迭代之间修改 story 描述、添加新 story、或手动将某个 story 标记为 `passes: true`(跳过它)。
**Q:Ralph 卡在某个反复失败的 story 上怎么办?**
1. 查看 progress.txt 了解失败原因
2. 在 notes 中补充更多上下文
3. 拆分 story(可能太大了)
4. 手动修复阻塞问题后重新运行
**Q:Ralph 运行时我可以做其他事吗?**
可以。Ralph 设计为 "Human on the Loop"——你不需要盯着它。AFK 模式下,下班前启动,第二天早上检查结果。只是不要修改 Ralph 正在处理的文件。
**Q:如何控制成本?**
三个方法:设置合理的 `max_iterations`、保持 story 粒度适当(减少浪费迭代)、先小规模试运行确认流程正确。一般来说,10-20 个 story 的项目,API 费用在 $50-100 以内。
***
## 总结
Ralph 的工作流可以提炼为五步:
```
安装 → 编写 PRD → 配置质量门禁 → 运行循环 → 检查结果
```
核心理念始终不变:**让文件成为唯一真相源,让每次迭代从全新开始,让质量门禁替你把关**。
现在,回到你的项目,准备好 prd.json,运行 `./scripts/ralph/ralph.sh --tool claude`,然后去泡杯咖啡吧。
### 延伸阅读
* [Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept) —— 重新回顾 Ralph 的核心原理
* [GSD 深度解析](/docs/notes/gsd/concept) —— 在 Ralph 基础上构建的完整上下文工程系统
* [什么是 Claude Skills](/docs/notes/claude-skills/concept) —— Ralph 的 PRD skill 就是一个 Claude Skill
* [Speckit 实战指南](/docs/notes/speckit/practice) —— 另一种结构化的 AI 编程工作流
# snarktank/ralph 实战指南
## 引言
[上一篇](/docs/notes/ralph-wiggum/concept)我们了解了 Ralph 的核心原理——无限循环 + 每次全新上下文 + 文件作为真相来源。三根支柱听起来简单,但从理解到实际跑通之间还有不少细节。
这一篇,我们来动手操作。[snarktank/ralph](https://github.com/snarktank/ralph) 是 Ralph 方法论的**外部循环实现**——每次迭代启动一个全新的 Claude 进程,彻底解决 Context Rot 问题。它是目前社区中最完善的 Ralph 实现之一(10k+ stars),支持 Claude Code 和 Amp 双平台,提供了 PRD 生成、JSON 转换、自动执行的全套工具链。
> 另一种实现路线是 [frankbria/ralph-claude-code](/docs/notes/ralph-wiggum/frankbria),提供完整的工程化工具链(监控仪表盘、断路器、速率限制),侧重可控性和安全机制。两者的对比见该文。
## 前置要求
在开始之前,确保你的环境满足以下条件:
| 依赖 | 说明 |
| ----------- | ------------------------------------------------------------------ |
| **AI 编程工具** | Claude Code (`npm install -g @anthropic-ai/claude-code`) 或 Amp CLI |
| **jq** | JSON 处理工具(macOS: `brew install jq`) |
| **Git** | 项目需要是 Git 仓库 |
```bash
# 检查依赖
claude --version # Claude Code CLI
jq --version # JSON 处理
git --version # Git
```
## 安装与配置
最简单的方式——在 Claude Code 对话中直接粘贴 GitHub 链接:
```
帮我安装这个 skill:https://github.com/snarktank/ralph
```
Claude Code 会自动克隆仓库并将 skill 文件复制到正确位置。安装完成后即可使用 `/prd` 和 `/ralph` 命令。
> snarktank/ralph 还支持 Marketplace 安装、手动复制 skill 文件、项目级安装等方式,详见 [GitHub 仓库说明](https://github.com/snarktank/ralph)。
***
## 核心文件结构
Ralph 的记忆完全依赖文件系统。理解每个文件的作用是用好 Ralph 的前提。
### ralph.sh — 循环引擎
这是 Ralph 的核心:一个 bash 脚本,负责反复启动新的 AI 实例。
```bash
# 基本用法
./scripts/ralph/ralph.sh [max_iterations] # 默认使用 Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # 使用 Claude Code
```
每次迭代,ralph.sh 做这些事:
1. 创建功能分支(基于 prd.json 中的 `branchName`)
2. 选择最高优先级的未完成 story(`passes: false`)
3. 启动一个**全新的** AI 实例来实现这个 story
4. 运行质量检查(类型检查、测试)
5. 检查通过 → git commit;失败 → 留给下一次迭代
6. 更新 prd.json,标记 story 为 `passes: true`
7. 在 progress.txt 中追加本次学到的经验
8. 重复,直到所有 story 完成或达到迭代上限
默认迭代上限是 10 次。根据项目复杂度调整:
```bash
# 简单项目
./scripts/ralph/ralph.sh --tool claude 10
# 复杂项目
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — 任务定义
这是 Ralph 的"大脑"——所有任务都在这里定义。格式是一个扁平的 JSON 文件:
```json
{
"projectName": "博客 i18n 翻译",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "翻译首页元数据",
"description": "创建 content/docs/meta.en.json,包含所有导航项的英文翻译",
"acceptanceCriteria": [
"meta.en.json 文件存在且 JSON 格式正确",
"所有导航标题都已翻译为英文",
"pnpm types:check 通过"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "参考现有 meta.json 的结构"
},
{
"id": "US-002",
"title": "翻译博客文章 hello-world",
"description": "创建 content/blog/hello-world.en.mdx,从中文翻译为英文",
"acceptanceCriteria": [
"hello-world.en.mdx 文件存在",
"所有 QuoteCard 组件设置 defaultLang='en'",
"内部链接使用 /en/ 前缀",
"代码块保持不翻译",
"pnpm types:check 通过"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "注意保留 MDX 组件的 props 格式"
}
]
}
```
**字段说明**:
| 字段 | 说明 |
| -------------------- | -------------------------- |
| `projectName` | 项目名称,用于日志和分支命名 |
| `branchName` | Git 分支名,Ralph 会自动创建 |
| `id` | Story 唯一标识,建议用 `US-001` 格式 |
| `title` | 简短标题 |
| `description` | 详细描述,越具体越好 |
| `acceptanceCriteria` | 验收标准列表——**这是最关键的字段** |
| `priority` | 优先级数字,越小越先执行 |
| `passes` | 是否已完成,Ralph 自动更新 |
| `dependsOn` | 依赖的 story ID 列表 |
| `notes` | 额外备注和提示 |
### progress.txt — 经验日志
这是 Ralph 的"长期记忆"。每次迭代结束后,AI 会在这里追加本轮学到的内容:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
下一次迭代的新 Claude 实例会读取这个文件,立即获得之前所有的经验。这就是 Ralph 能越跑越顺的原因——**知识在迭代间积累,但上下文保持干净**。
### AGENTS.md — 持久化知识库
除了 progress.txt,Ralph 还会更新项目中的 `AGENTS.md` 文件(或 `CLAUDE.md`)。Claude Code 和 Amp 都会在启动时自动读取这些文件。
与 progress.txt 不同,AGENTS.md 记录的是**稳定的、跨项目通用的知识**:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## 编写 PRD
PRD(Product Requirements Document)的质量直接决定 Ralph 的执行效果。写得好,Ralph 一路畅通;写得差,Ralph 会在同一个 story 上反复失败。
### 使用 Skill 生成 PRD
如果你安装了 snarktank/ralph 的 skill,可以用交互式方式生成 PRD:
```bash
# 在 Claude Code 或 Amp 中
/prd 我想为博客系统添加 i18n 支持,需要将所有中文内容翻译成英文
```
AI 会向你提出一系列澄清问题(涉及哪些文件、技术栈约束、质量标准等),然后生成结构化的 PRD 文档。
生成后,使用 `/ralph` 命令将 PRD 转换为 `prd.json` 格式:
```bash
/ralph # 将 PRD 转换为 prd.json
```
### 手动编写 PRD
你也可以直接编写 prd.json。以下是关键的设计原则。
**原则一:Story 粒度要合适**
每个 story 应该小到能在一次迭代中完成,大到有独立的交付价值。
```json
// ❌ 太大:一次迭代完不成
{
"id": "US-001",
"title": "构建完整的用户认证系统",
"description": "实现注册、登录、忘记密码、OAuth、权限管理..."
}
// ❌ 太小:没有独立价值
{
"id": "US-001",
"title": "创建 User 表的 email 字段",
"description": "在 User 模型中添加 email 字段"
}
// ✅ 合适:一次能完成,有独立价值
{
"id": "US-001",
"title": "实现邮箱密码登录",
"description": "创建登录 API 和登录页面,支持邮箱密码验证",
"acceptanceCriteria": [
"POST /api/auth/login 接受 email + password",
"返回 JWT token",
"登录页面表单可提交",
"所有测试通过"
]
}
```
**经验法则**:一个 story 涉及 1-3 个文件修改,有 3-5 条验收标准。
**原则二:验收标准必须可自动验证**
Ralph 需要判断 story 是否完成,所以验收标准必须是可以客观判断的:
```json
// ❌ 模糊的标准
"acceptanceCriteria": [
"代码质量好",
"性能不错",
"用户体验流畅"
]
// ✅ 可验证的标准
"acceptanceCriteria": [
"pnpm types:check 通过",
"pnpm test 通过",
"API 响应时间 < 200ms",
"文件 src/auth/login.ts 存在且导出 loginHandler 函数"
]
```
**原则三:利用 dependsOn 控制顺序**
有些 story 之间有依赖关系。`dependsOn` 字段确保 Ralph 按正确顺序执行:
```json
{
"userStories": [
{
"id": "US-001",
"title": "创建数据库 schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "实现用户注册 API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "实现登录页面",
"dependsOn": ["US-002"]
}
]
}
```
**原则四:在 notes 中提供上下文**
notes 字段是给 AI 的额外提示。把你知道但 AI 可能不知道的信息写在这里:
```json
{
"notes": "项目使用 fumadocs 框架,i18n 文件命名规则是 .en.mdx 后缀。参考 content/docs/notes/speckit/concept.en.mdx 的翻译风格。"
}
```
***
## 执行 Ralph Loop
PRD 准备好了,开始跑循环。
### 启动执行
```bash
# 使用 Claude Code,默认 10 次迭代
./scripts/ralph/ralph.sh --tool claude
# 指定迭代次数
./scripts/ralph/ralph.sh --tool claude 30
# 使用 Amp(默认)
./scripts/ralph/ralph.sh 20
```
### 执行过程
启动后,你会看到类似这样的输出:
```
Starting Ralph - Tool: claude - Max iterations: 35
===============================================================
Ralph Iteration 1 of 35 (claude)
===============================================================
## US-001 Complete
**Summary of what was done:**
1. Created meta.en.json with all navigation items translated
2. Ran pnpm types:check — PASSED
3. Committed: feat: [US-001] - Translate homepage metadata
There are still **15 user stories with `passes: false`** remaining.
The next story is **US-002: 翻译博客文章 hello-world**.
Iteration 1 complete. Continuing...
===============================================================
Ralph Iteration 2 of 35 (claude)
===============================================================
```
每次迭代都是一个全新的 Claude 实例。它通过读取 prd.json 知道当前要做什么,通过 progress.txt 知道之前学到了什么。
### 完成信号
当所有 story 都标记为 `passes: true` 时,Ralph 会输出完成信号并退出:
```
All stories completed!
COMPLETE
```
### 监控与调试
在 Ralph 运行时,你可以用这些命令查看进度:
```bash
# 查看各 story 的完成状态(带图标,更直观)
cat tasks/prd.json | python3 -c "
import json,sys
for s in json.load(sys.stdin)['userStories']:
print(f'{\"✅\" if s[\"passes\"] else \"⬜\"} {s[\"id\"]}: {s[\"title\"]}')"
# 或者用 jq 查看
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 查看经验日志
cat progress.txt
# 查看最近的 git 提交
git log --oneline -10
# 实时查看 Ralph 的输出
tail -f progress.txt
# 完成后,查看相比主分支的全部改动
git diff main...ralph/your-branch-name --stat
```
### 中断与恢复
Ralph 运行时间可能很长,中途中断是完全安全的:
* **中断**:直接 `Ctrl+C`,已完成的 story(`passes: true`)不会丢失,它们已经 commit 并写入 prd.json
* **恢复**:再次运行同样的命令,Ralph 会自动从第一个 `passes: false` 的 story 继续
```bash
# 中断后恢复,只需重新运行同样的命令
./scripts/ralph/ralph.sh --tool claude 35
```
如果某个 story 反复失败导致阻塞,你可以手动跳过它——编辑 `prd.json`,将该 story 的 `passes` 字段改为 `true`,然后重新运行。Ralph 会跳过它,继续处理后续 story。
### 自动归档
当你用新的 `branchName` 启动一个不同的功能时,Ralph 会自动将上一次运行的文件归档到 `archive/YYYY-MM-DD-feature-name/` 目录,保持工作目录整洁。
***
## 反馈循环与质量门禁
Ralph 的"自我纠错"能力完全取决于反馈循环的质量。没有反馈循环的 Ralph 就是一个盲目循环的脚本——它会不断产出代码,但无法判断代码是否正确。
### 配置质量检查
在 CLAUDE.md(或 prompt.md)中定义你的质量检查命令:
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 质量门禁的层次
| 层次 | 工具 | 捕获的问题 |
| ----- | ------------------- | ------------- |
| 即时反馈 | TypeScript compiler | 类型错误、语法错误 |
| 功能验证 | 单元测试 | 逻辑错误、边界情况 |
| 集成验证 | Build 命令 | 依赖问题、配置错误 |
| 运行时验证 | dev-browser skill | UI 渲染问题(前端项目) |
> 对于前端 story,Ralph 建议在验收标准中加入:"Verify in browser using dev-browser skill"——让 AI 实际打开浏览器确认页面渲染正确。
### 当质量检查失败时
如果某个 story 的质量检查反复失败,Ralph 不会无限重试同一个 story。到达迭代上限后,它会停止并留下当前状态。你可以:
1. 检查 progress.txt 看看 AI 卡在什么地方
2. 手动修复问题后重新运行
3. 调整 story 的粒度(可能拆分得太大了)
4. 补充 notes 提供更多上下文
### Prompt 定制
Ralph 的 prompt 模板(CLAUDE.md 或 prompt.md)是你控制 AI 行为的主要手段。安装后,你应该根据自己的项目定制它。关键定制方向:
**代码风格约束**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**常见陷阱**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**卡住时的处理方式**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 实战案例:用 Ralph 完成博客 i18n 翻译
为了展示 Ralph 在实际项目中的运作方式,这里分享一个真实案例:使用 Ralph 风格的自主 agent 将整个博客从中文翻译成英文。
### 项目设置
项目需要将 22+ 个内容文件(博客文章、文档、导航元数据)从中文翻译成英文,目标是基于 fumadocs 的 Next.js 博客 i18n 支持。任务定义在一个 `prd.json` 文件中,包含 16 个 user story,每个都有明确的验收标准:
```
scripts/ralph/
├── prd.json # 16 个 user story,含验收标准
└── progress.txt # 经验日志,每完成一个 story 后更新
```
每个 user story 遵循一致的模式:
* **明确的交付物**:"Create content/blog/xxx.en.mdx"
* **可验证的标准**:"Typecheck passes"、"Internal links use /en/ prefix"
* **技术约束**:"Keep code blocks untranslated"、"Set defaultLang='en' on QuoteCard"
### 执行模式
Agent 遵循 Ralph 方法论的核心原则:
1. **文件即真相来源**:`prd.json` 跟踪每个 story 的状态(`passes: true/false`)。`progress.txt` 在迭代中积累经验——比如 "Typecheck 命令是 `pnpm types:check`,不是 `pnpm typecheck`"
2. **自动化质量门禁**:每次翻译完成后,运行 `pnpm types:check` 验证 MDX 文件能正确编译。如果 typecheck 失败,先修复问题再提交。
3. **增量推进**:每个 story 独立提交,使用描述性的 commit message(`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`),方便在需要时回滚。
4. **并行执行**:对于较长的文章,多个 subagent 同时翻译——比如 US-010(claude-skills concept + practice)、US-011(speckit concept + practice)和 US-012(claude-architecture + claude-subagent)同时并行运行。
### 关键经验
| 经验 | 详情 |
| -------------------- | -------------------------------------------------------------- |
| **知识积累很重要** | 早期 story 发现的模式(QuoteCard 的 `defaultLang`、链接前缀规则)让后续 story 更快完成 |
| **Typecheck 作为反馈循环** | 在问题叠加之前捕获缺失的 import 或格式错误的 MDX |
| **并行化可以扩展** | 6 个翻译 agent 同时运行,完成时间与 1 个差不多 |
| **PRD 粒度至关重要** | 每个 story 限定在 1-2 个文件——小到足以可靠完成,大到足以有意义 |
| **进度日志防止重复犯错** | `progress.txt` 的 "Codebase Patterns" 部分成了知识库,避免重新踩坑 |
### 成果
16 个 user story 在单次会话中全部完成:创建了 8 个 meta.en.json 导航文件、翻译了 3 篇博客文章、翻译了 12 个文档页面、完整站点构建验证通过。每次翻译都保持了一致的质量,因为验收标准是明确的,反馈循环(typecheck)能立即发现问题。
这个项目展示了 Ralph 的**完整实现模式(Full Implementation Mode)**——定义清晰的任务、明确的成功标准、自动化验证,以及通过文件系统进行增量交付。
***
## 最佳实践与常见问题
### 成本控制
Ralph 的自动化执行意味着 API 成本是持续产生的。几个控制手段:
* **始终设置 `max_iterations`**:这是最基本的安全网
* **Story 粒度合理**:太大的 story 会消耗多次迭代;太碎的 story 增加启动开销
* **先小规模测试**:新项目先用 3-5 次迭代试跑,确认 prompt 和质量门禁工作正常后再放开
### 常见陷阱
**陷阱 1:Story 太大**
症状:一个 story 反复失败,迭代次数很快耗尽。
解决:拆分成 2-3 个更小的 story。"构建完整的认证系统"拆成"实现登录 API"+"创建登录页面"+"添加 JWT 中间件"。
**陷阱 2:没有反馈循环**
症状:Ralph 声称 story 完成了,但实际代码有问题。
解决:在验收标准中加入可执行的检查命令。"代码写好了"不是验收标准,"pnpm test 全部通过"才是。
**陷阱 3:progress.txt 没有被利用**
症状:同一个错误在不同迭代中反复出现。
解决:确认 prompt 模板中有明确指示"读取 progress.txt 并遵循其中的经验"。如果 AI 没有自动追加学习内容,在 prompt 中加入"After each story, append learnings to progress.txt"。
**陷阱 4:依赖顺序错误**
症状:某个 story 依赖的代码还不存在,导致实现失败。
解决:正确设置 `dependsOn` 字段,确保基础设施 story 排在前面。
### 常见问题
**Q: 能不能手动修改 prd.json 中途干预?**
可以。Ralph 每次迭代开始时都会重新读取 prd.json。你可以在迭代间隙修改 story 描述、添加新 story、或手动标记某个 story 为 `passes: true`(跳过它)。
**Q: 如果 Ralph 卡在一个 story 上反复失败怎么办?**
1. 检查 progress.txt 看失败原因
2. 补充 notes 提供更多上下文
3. 拆分 story(可能粒度太大)
4. 手动修复阻塞问题后重新运行
**Q: Ralph 运行时我可以做其他事吗?**
可以。Ralph 设计为"Human on the Loop"——你不需要盯着它。AFK 模式下,下班前启动,第二天检查结果就行。运行时不要修改 Ralph 正在操作的文件即可。
**Q: 如何控制成本?**
三种方式:设置合理的 `max_iterations`、保持 story 粒度合适(减少无效迭代)、以及先小规模试跑确认流程无误。一般来说,10-20 个 story 的项目在 $50-100 的 API 成本范围内。
***
## 小结
Ralph 的使用流程可以归纳为五步:
```
安装 → 编写 PRD → 配置质量门禁 → 执行循环 → 检查成果
```
核心思路始终不变:**让文件成为真相来源,让每次迭代都是全新开始,让质量门禁替你把关**。
现在,回到你的项目,准备好 prd.json,运行 `./scripts/ralph/ralph.sh --tool claude`,然后去喝杯咖啡吧。
### 延伸阅读
* 《[Ralph Wiggum 深度解析](/docs/notes/ralph-wiggum/concept)》— 回顾 Ralph 的核心原理
* 《[frankbria/ralph-claude-code 实战指南](/docs/notes/ralph-wiggum/frankbria)》— 工程化的 Ralph 实现:监控、断路器与安全机制
* 《[GSD 深度解析](/docs/notes/gsd/concept)》— 在 Ralph 基础上构建的完整上下文工程系统
* 《[Claude Skills 是什么](/docs/notes/claude-skills/concept)》— Ralph 的 PRD skill 就是一个 Claude Skill
* 《[Speckit 实践指南](/docs/notes/speckit/practice)》— 另一种结构化的 AI 编程工作流
# 规格驱动开发概念介绍
## 引言
2025 年 10 月,GitHub 开源了一个名为 Spec Kit 的工具包,正式将「规格驱动开发」(Spec-Driven Development)这一概念带入 AI 编程的视野。这个看似复古的理念——先写规格再写代码——正在成为驾驭 AI 编程工具的新范式。
如果你经常使用 Claude Code、Cursor 或 GitHub Copilot 这类 AI 编程助手,一定遇到过这样的困扰:你说「帮我添加一个用户登录功能」,AI 兴冲冲地写了一大堆代码,但等你仔细一看——用的是你不熟悉的框架、安全策略和你预期的不一样、UI 风格也对不上……然后你开始一轮又一轮地修正,直到精疲力尽。
问题出在哪里?不是 AI 不够聪明,而是你给的信息不够。「添加用户登录功能」这句话看似清晰,实际上隐藏了成百上千个未说明的决策:用什么认证方式?密码有什么要求?登录失败怎么处理?需要记住登录状态吗?支持第三方登录吗?……AI 不得不猜,而猜测意味着偏差。
规格驱动开发正是为了解决这个问题而生的。
## Vibe Coding:速度的代价
2025 年初,前 Tesla AI 总监 Andrej Karpathy 创造了「Vibe Coding」这个词,描述一种「接受 AI 建议而不深入审查」的开发方式。这个词迅速走红,甚至被 Collins 词典评为 2025 年度词汇。
Vibe Coding 的诱惑显而易见:你描述一个想法,AI 生成代码,看起来能跑就行。对于快速原型、黑客马拉松、一次性脚本,这种方式确实高效。但当它被用于生产系统时,问题就来了。
类似的故事在业界屡见不鲜:AI 生成的数据库查询在小规模测试中运行良好,但在真实流量下系统爬行;拼接的认证模块通过了 QA,但两周后发现停用账户仍可访问管理工具。根据 Final Round AI 的 2025 年调查,**18 位 CTO 中有 16 位经历过 AI 生成代码导致的生产灾难**。
这并不是说 Vibe Coding 毫无价值。关键在于**识别边界**:
| 场景 | Vibe Coding | 规格驱动开发 |
| ----- | ----------- | ------ |
| 原型/演示 | ✓ 适合 | 过度 |
| 一次性脚本 | ✓ 适合 | 过度 |
| 生产功能 | 风险高 | ✓ 推荐 |
| 安全相关 | 危险 | ✓ 必须 |
| 团队协作 | 难维护 | ✓ 推荐 |
规格驱动开发正是为了在保持 AI 效率的同时,避免 Vibe Coding 的陷阱。
## 什么是规格驱动开发
规格驱动开发的核心理念可以用一句话概括:**先定义「做什么」,再考虑「怎么做」**。
这听起来像是软件工程的老生常谈,但在 AI 编程时代,它有了新的含义。传统的需求文档是写给人看的,往往冗长、模糊、充满行话。规格驱动开发中的「规格」是写给 AI 看的——简洁、结构化、可执行。
想象你要盖一栋房子。传统的 AI 编程方式就像你对施工队说「帮我盖一栋舒适的三居室」,然后让他们自己发挥。结果可能不错,但更可能和你想象的大相径庭。规格驱动开发则是先画好建筑蓝图:几层楼、每层多少平米、窗户朝向、材料规格……施工队按图施工,结果自然符合预期。
在 AI 编程中,这套蓝图就是**规格文档**(Specification)。它不关心用什么编程语言、什么框架,只关心功能要达成什么效果、用户要完成什么任务、成功的标准是什么。
与传统开发流程相比,规格驱动开发有一个根本的不同:
| 传统 AI 编程 | 规格驱动开发 |
| ---------------- | ---------------------- |
| 直接描述需求 → AI 生成代码 | 需求 → 规格 → 计划 → 任务 → 代码 |
| AI 需要猜测大量细节 | 每一步都明确,AI 只需执行 |
| 返工频繁,沟通成本高 | 前期投入,后期顺畅 |
| 适合简单任务 | 适合复杂功能 |
这种「渐进式细化」的过程,正是规格驱动开发的精髓。你不是一步到位,而是通过多个阶段逐步明确需求,每个阶段都可以审核和调整。
## Speckit 工作流程概览
GitHub 的 Spec Kit 和 Claude Code 中的 speckit 命令都遵循类似的工作流程,大致可以分为六个阶段:
```
Constitution → Specify → Clarify → Plan → Tasks → Implement
↓ ↓ ↓ ↓ ↓ ↓
项目宪法 功能规格 澄清模糊 技术计划 任务分解 执行实现
```
**1. Constitution(项目宪法)**
项目宪法定义了整个项目的基本原则和约束,比如「测试优先」「简单至上」「API 优先」等。这些原则会贯穿后续所有阶段,确保 AI 生成的方案符合你的技术偏好。
**2. Specify(功能规格)**
这是核心的第一步。你用自然语言描述想要的功能,AI 帮你整理成结构化的规格文档,包括:
* 用户故事:谁要做什么,为什么
* 功能需求:系统必须具备的能力
* 成功标准:如何判断功能是否达标
重要的是,规格文档**只关注「做什么」,不涉及「怎么做」**——不提具体的技术栈、不写代码结构。
**3. Clarify(澄清模糊)**
AI 会检查规格中的模糊点,提出最多 5 个关键问题。这些问题通常涉及功能边界、用户类型、安全要求等。通过问答,规格变得更加清晰。
**4. Plan(技术计划)**
有了清晰的规格,才开始考虑技术方案。这一步会产出:
* 技术选型(语言、框架、数据库)
* 数据模型设计
* API 合约定义
* 研究报告(解决技术决策)
**5. Tasks(任务分解)**
将技术计划拆分成可执行的任务清单。每个任务都有明确的 ID、描述、文件路径,可以直接交给 AI 执行。任务按用户故事分组,支持并行开发。
**6. Implement(执行实现)**
按任务清单逐个执行。每完成一个任务就标记完成,确保可追溯。
这六个阶段的产出物形成了一条清晰的链条:
| 阶段 | 产出物 | 作用 |
| ------------ | -------------------- | ------- |
| Constitution | constitution.md | 定义项目原则 |
| Specify | spec.md | 描述功能需求 |
| Clarify | 更新后的 spec.md | 消除模糊点 |
| Plan | plan.md, research.md | 技术方案设计 |
| Tasks | tasks.md | 可执行任务清单 |
| Implement | 实际代码 | 最终交付物 |
## 为什么这种方式有效
规格驱动开发之所以能在「直接让 AI 写代码」失败的地方成功,是因为它解决了 AI 编程的核心矛盾:**信息不对称**。
当你说「添加照片分享功能」时,你脑子里可能有一个完整的图景,但 AI 只看到这几个字。它必须猜测:分享到哪里?谁可以看?需要压缩吗?有水印吗?支持批量吗?……每一个猜测都可能是错的。
规格驱动开发通过「强制你先想清楚」来解决这个问题。当你被要求写下用户故事、功能需求、成功标准时,那些你以为「显而易见」的细节就会浮出水面。这个过程本身就有价值——即使不用 AI,把需求写清楚也能减少沟通成本。
此外,渐进式细化让错误更早暴露。在 Specify 阶段发现需求偏差,修改成本几乎为零;在代码写完后才发现,可能要推倒重来。
当然,规格驱动开发不是万能的。它有明确的适用场景:
**适合的情况**:
* 复杂功能开发(涉及多个模块、多种交互)
* 团队协作项目(规格文档作为沟通媒介)
* 对质量要求高的场景(需要可追溯、可验证)
**不适合的情况**:
* 简单的 bug 修复或小改动
* 探索性编程(还不知道要做什么)
* 时间极度紧迫(没空写规格)
关键是识别任务的复杂度。一个小时能完成的事情,不需要花一个小时写规格;一个星期的功能开发,花两个小时写规格绝对值得。
## 但规格不是银弹
需要澄清一个常见误解:**规格驱动开发减少了猜测,但没有消除审查的需求**。
即使有了完整的规格,AI 仍可能:
* 错过边界情况(规格没覆盖的极端场景)
* 生成不符合性能要求的代码
* 引入潜在的安全漏洞
* 产出风格不一致的实现
这就像建筑施工:即使有了详细蓝图,验收检查仍然必要。你不会因为施工队按图纸盖完房子就直接搬进去住——你会检查电路是否安全、管道是否通畅、门窗是否牢固。
规格驱动开发的价值在于**让错误更容易被发现**,而不是消除错误本身。
## 小结
规格驱动开发的核心是一个简单的道理:**越是复杂的任务,越需要先想清楚再动手**。AI 编程工具放大了这个道理的重要性——因为 AI 会忠实地执行你的指令,但无法真正理解你的意图。
记住三个关键词:
| 关键词 | 含义 |
| ---------- | ------------------- |
| **先规格后代码** | 定义「做什么」在前,考虑「怎么做」在后 |
| **渐进式细化** | 从模糊到清晰,每步都可审核调整 |
| **减少猜测** | 明确的规格 = AI 更少的推测空间 |
了解了理念之后,下一篇《[Speckit 实践指南](/docs/notes/speckit/practice)》将带你动手实践:如何使用 speckit 命令完成一个功能的规格驱动开发流程。
配合《[Claude Skills](/docs/notes/claude-skills/concept)》使用,可以将规格的执行进一步自动化和标准化。
# 规格驱动开发实践指南
## 引言
[上一篇](/docs/notes/speckit/concept)我们了解了规格驱动开发的理念——先定义「做什么」,再考虑「怎么做」。这种看似多此一举的流程,实际上能大幅减少 AI 编程中的返工和沟通成本。
这一篇,我们来动手实践。你将学会如何使用 speckit 系列命令,完成从需求描述到代码实现的完整流程。
## 安装与配置
Speckit 命令源自 GitHub 官方的 [Spec Kit](https://github.com/github/spec-kit) 项目。根据你的使用场景,有几种不同的集成方式。
### 新项目初始化
对于新项目,推荐使用官方的 specify-cli 工具进行初始化:
```bash
# 使用 uv 安装 specify-cli
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git
# 初始化新项目,指定使用 Claude 作为 AI 助手
specify init my-project --ai claude
```
这会自动创建项目目录结构,包括 `.specify/` 配置目录和相关模板文件。
### 已有项目集成
Speckit 命令需要配置文件才能使用。在已有项目中集成 speckit,推荐使用 specify-cli:
```bash
cd your-existing-project
specify init . --ai claude # 注意是 . 表示当前目录
```
这会在项目中创建:
```
your-project/
├── .specify/
│ ├── templates/ # 规格、计划等模板
│ ├── scripts/ # 辅助脚本
│ └── memory/ # constitution.md
├── .claude/
│ └── commands/ # Claude Code 命令配置
│ ├── speckit.specify.md
│ ├── speckit.plan.md
│ └── ...
└── specs/ # 功能规格存放目录
```
初始化不会覆盖你现有的文件。完成后,在 Claude Code 中即可使用 `/speckit.*` 系列命令。
> **注意**:speckit 命令不是 Claude Code 内置的,必须先完成上述初始化步骤。如果直接运行 `/speckit.specify` 会提示命令不存在。
***
## 命令详解
Speckit 提供了一套命令来支持规格驱动开发的各个阶段。每个命令都有明确的输入和输出,形成一条可追溯的链条。
### /speckit.specify — 创建功能规格
这是整个流程的起点。你用自然语言描述想要的功能,AI 会帮你整理成结构化的规格文档。
**功能**:从自然语言描述创建功能规格文档
**输入**:功能描述(自然语言)
**输出**:
* `specs/[编号]-[功能名]/spec.md` — 功能规格文档
* 新的 git 分支(如 `001-user-auth`)
**使用示例**:
```
/speckit.specify 我想添加一个用户登录功能,支持邮箱密码登录,需要有记住登录状态的选项
```
执行后,AI 会:
1. 生成一个简短的功能名称(如 `user-auth`)
2. 创建新的功能分支
3. 生成包含用户故事、功能需求、成功标准的规格文档
4. 对不明确的地方标记 `[NEEDS CLARIFICATION]`
**规格文档的核心结构**:
```markdown
# Feature Specification: 用户登录功能
## User Scenarios & Testing
### User Story 1 - 用户登录 (Priority: P1)
读者通过邮箱密码登录系统...
**Acceptance Scenarios**:
1. Given 用户输入正确的邮箱和密码, When 点击登录, Then 成功进入系统
## Requirements
### Functional Requirements
- FR-001: 系统必须支持邮箱密码登录
- FR-002: 系统必须提供"记住我"选项
## Success Criteria
- SC-001: 用户能在 30 秒内完成登录流程
```
注意规格文档**不涉及任何技术细节**——不提用什么框架、不写数据库结构、不定义 API。这些都是后面阶段的事情。
***
### /speckit.clarify — 澄清模糊点
规格文档写好后,可能还有一些模糊的地方。这个命令会检查规格,提出关键问题帮你澄清。
**功能**:识别规格中的模糊点,通过问答完善规格
**输入**:现有的 spec.md 文档
**输出**:更新后的 spec.md(包含澄清记录)
**使用示例**:
```
/speckit.clarify
```
执行后,AI 会:
1. 扫描规格文档中的模糊点
2. 按优先级(范围 > 安全 > 用户体验 > 技术细节)排序
3. 逐一提问,每次只问一个问题
4. 根据你的回答更新规格文档
**问答示例**:
```markdown
## Question 1: 登录失败处理
**Context**: 规格中提到用户登录,但未说明登录失败的处理方式。
**Recommended:** Option B - 连续 5 次失败后锁定账户是安全最佳实践
| Option | Description |
|--------|-------------|
| A | 仅显示错误提示,不做限制 |
| B | 连续 5 次失败后锁定账户 15 分钟 |
| C | 使用验证码防止暴力破解 |
你可以回复选项字母(如 "B"),说 "yes" 接受推荐,或提供自己的答案。
```
每次澄清后,规格文档会自动更新,添加澄清记录:
```markdown
## Clarifications
### Session 2025-12-20
- Q: 登录失败如何处理? → A: 连续 5 次失败后锁定账户 15 分钟
```
***
### /speckit.plan — 生成技术计划
规格清晰后,进入技术设计阶段。这一步会产出技术计划和研究报告。
**功能**:根据规格生成技术实施计划
**输入**:spec.md 文档
**输出**:
* `plan.md` — 技术计划(架构、数据模型、API 设计)
* `research.md` — 研究报告(技术选型决策)
* `data-model.md` — 数据模型(如适用)
* `contracts/` — API 合约(如适用)
**使用示例**:
```
/speckit.plan 我使用 Next.js + Prisma + PostgreSQL
```
你可以在命令后附加技术栈偏好。执行后,AI 会:
1. 分析规格中的功能需求
2. 研究相关技术的最佳实践
3. 设计数据模型和 API 结构
4. 生成完整的技术计划
**技术计划的核心内容**:
```markdown
# Implementation Plan: 用户登录功能
## Technical Context
**Language/Version**: TypeScript 5.x
**Primary Dependencies**: Next.js 15, Prisma, PostgreSQL
**Authentication**: NextAuth.js with credentials provider
## Project Structure
src/
├── app/
│ └── (auth)/
│ ├── login/
│ └── api/auth/
├── lib/
│ └── auth/
└── prisma/
└── schema.prisma
## Data Model
- User: id, email, passwordHash, createdAt, updatedAt
- Session: id, userId, expiresAt
```
***
### /speckit.tasks — 分解任务
技术计划有了,接下来把它分解成可执行的任务清单。
**功能**:将技术计划拆分为可执行的任务列表
**输入**:plan.md 文档
**输出**:`tasks.md` — 按依赖排序的任务清单
**使用示例**:
```
/speckit.tasks
```
执行后,AI 会:
1. 从 plan.md 提取技术方案
2. 从 spec.md 提取用户故事优先级
3. 按用户故事分组生成任务
4. 标记可并行执行的任务 `[P]`
5. 为每个任务指定具体的文件路径
**任务清单格式**:
```markdown
## Phase 1: Setup
- [ ] T001 创建项目结构
- [ ] T002 [P] 配置 Prisma schema
- [ ] T003 [P] 配置 NextAuth
## Phase 2: User Story 1 - 用户登录 (P1)
- [ ] T004 [US1] 创建 User 模型 in prisma/schema.prisma
- [ ] T005 [US1] 实现登录 API in src/app/api/auth/[...nextauth]/route.ts
- [ ] T006 [US1] 创建登录页面 in src/app/(auth)/login/page.tsx
```
每个任务都有:
* **任务 ID**(T001, T002...)— 用于追踪
* **\[P] 标记** — 表示可与其他 \[P] 任务并行
* **\[US] 标签** — 表示属于哪个用户故事
* **文件路径** — 明确要操作哪个文件
***
### /speckit.implement — 执行实现
万事俱备,开始执行任务清单。
**功能**:按任务清单逐个执行实现
**输入**:tasks.md 文档
**输出**:实际代码
**使用示例**:
```
/speckit.implement
```
执行前,AI 会检查检查清单(如果有)。执行时:
1. 按阶段顺序执行任务
2. 每完成一个任务,标记为 `[X]`
3. 遵循任务依赖关系
4. 并行任务可同时进行
**执行过程示例**:
```
Phase 1: Setup
✓ T001 创建项目结构
✓ T002 配置 Prisma schema
✓ T003 配置 NextAuth
Phase 2: User Story 1
✓ T004 创建 User 模型
正在执行 T005...
```
### 实现后的审查
`/speckit.implement` 完成后,**不要直接合并代码**。AI 生成的代码需要人工审查:
**必做的验证步骤**:
1. **运行测试套件**
```bash
npm test # 或你的测试命令
```
确保 AI 没有破坏现有功能。
2. **代码审查要点**
* 是否符合规格的意图(对照 spec.md)
* 是否遵循项目的代码风格
* 是否有潜在的安全问题
3. **边界测试**
手动测试 AI 可能遗漏的边界情况:
* 空值处理
* 极端输入
* 并发场景
* 错误路径
4. **性能检查**
如果涉及数据库操作或 API 调用,检查是否有 N+1 查询等性能问题。
> **提示**:即使规格写得很详细,AI 仍可能在实现细节上出现偏差。审查不是不信任规格驱动开发,而是工程纪律的一部分。
***
### /speckit.analyze — 一致性分析
这是一个可选的质量检查步骤,用于验证规格、计划、任务之间的一致性。
**功能**:跨文档一致性和质量分析
**输入**:spec.md, plan.md, tasks.md
**输出**:分析报告(不修改任何文件)
**使用示例**:
```
/speckit.analyze
```
执行后会检查:
* 每个需求是否都有对应的任务
* 任务是否覆盖了所有用户故事
* 术语是否一致
* 是否有遗漏或重复
***
### 其他命令(可选)
除了上述核心命令,speckit 还提供了几个辅助命令。这些命令不在主流程中,但在特定场景下很有用。
**`/speckit.constitution`** — 创建项目宪法
用于定义项目的开发原则和规范。适合团队项目,确保所有成员遵循统一的开发标准。
* **输入**:交互式问答或直接提供原则
* **输出**:`.specify/constitution.md` 项目宪法文件
* **场景**:新团队项目初始化、统一代码风格和架构决策
**`/speckit.checklist`** — 生成质量检查清单
根据功能规格生成定制化的质量检查清单,用于实施前确认质量标准。
* **输入**:spec.md 文档
* **输出**:`checklists/` 目录下的检查清单
* **场景**:重要功能上线前的质量把关、代码审查参考
**`/speckit.taskstoissues`** — 转换任务为 GitHub Issues
将 tasks.md 中的任务自动转换为 GitHub Issues,方便团队协作和任务分配。
* **输入**:tasks.md 文档
* **输出**:GitHub Issues(通过 gh CLI 创建)
* **场景**:团队协作开发、Sprint 规划、任务追踪
***
## 工具生态
本文介绍的 speckit 命令来自 [GitHub Spec Kit](https://github.com/github/spec-kit) 项目。除此之外,2025 年多个主流 AI 编程工具都开始支持类似的规格驱动工作流:
| 工具 | 特点 | 适用场景 |
| --------------------------------------------------------- | ---------------------------------------------------- | ---------------- |
| **[GitHub Spec Kit](https://github.com/github/spec-kit)** | 本文使用的工具,MIT 开源,支持 Claude Code / Copilot / Gemini CLI | 命令行偏好者,跨工具协作 |
| **[AWS Kiro](https://kiro.dev/)** | VS Code fork,可视化工作流,EARS 表示法 | GUI 偏好者,AWS 生态用户 |
| **[JetBrains Junie](https://blog.jetbrains.com/junie/)** | IntelliJ 生态集成,Think More 推理模式 | JetBrains IDE 用户 |
| **Cursor Plan Mode** | 内置规划阶段,自动生成执行计划 | 已使用 Cursor 的开发者 |
**如何选择**:
* 如果你使用 Claude Code、GitHub Copilot 或 Gemini CLI,推荐 GitHub Spec Kit
* 如果偏好图形界面和可视化工作流,可以尝试 AWS Kiro
* 如果是 JetBrains 用户,Junie 与 IDE 的集成更自然
* 如果已经在用 Cursor,它的 Plan Mode 提供了类似的规划能力
核心理念是相通的——工具只是载体,重要的是**先规格后代码**的思维方式。
***
## 完整案例演示
让我们通过一个实际案例,走完 speckit 的完整流程。
**场景**:为博客系统添加文章收藏功能
### Step 1: 创建规格
```
/speckit.specify 我想为博客添加文章收藏功能,用户可以收藏喜欢的文章,并在个人中心查看收藏列表
```
AI 生成规格文档 `specs/002-article-bookmark/spec.md`:
```markdown
# Feature Specification: 文章收藏功能
## User Scenarios & Testing
### User Story 1 - 收藏文章 (Priority: P1)
用户在阅读文章时,点击收藏按钮将文章添加到收藏列表。
**Acceptance Scenarios**:
1. Given 用户已登录, When 点击收藏按钮, Then 文章被添加到收藏列表
2. Given 用户未登录, When 点击收藏按钮, Then 提示登录
### User Story 2 - 查看收藏列表 (Priority: P2)
用户在个人中心查看自己收藏的所有文章。
## Requirements
- FR-001: 用户必须能够收藏/取消收藏文章
- FR-002: 收藏按钮必须显示当前收藏状态
- FR-003: 个人中心必须显示收藏列表
## Success Criteria
- SC-001: 收藏操作在 500ms 内完成
- SC-002: 收藏列表支持分页,每页 10 篇
```
### Step 2: 澄清需求
```
/speckit.clarify
```
AI 提问:「收藏数量是否有上限?」
回答:「最多收藏 100 篇文章」
规格更新,添加:
* FR-004: 每个用户最多收藏 100 篇文章
* 当达到上限时显示提示
### Step 3: 生成计划
```
/speckit.plan 使用 Next.js + Prisma
```
生成技术计划,包括:
* Bookmark 模型(userId, articleId, createdAt)
* API 路由设计(POST/DELETE /api/bookmarks)
* 组件设计(BookmarkButton, BookmarkList)
### Step 4: 分解任务
```
/speckit.tasks
```
生成任务清单:
```markdown
## Phase 1: Setup
- [ ] T001 添加 Bookmark 模型到 Prisma schema
## Phase 2: US1 - 收藏文章
- [ ] T002 [US1] 创建收藏 API in src/app/api/bookmarks/route.ts
- [ ] T003 [US1] 创建 BookmarkButton 组件 in src/components/BookmarkButton.tsx
- [ ] T004 [US1] 集成到文章页面
## Phase 3: US2 - 收藏列表
- [ ] T005 [US2] 创建收藏列表页面 in src/app/profile/bookmarks/page.tsx
- [ ] T006 [US2] 实现分页逻辑
```
### Step 5: 执行实现
```
/speckit.implement
```
按任务顺序执行,每完成一个任务标记 `[X]`。
***
## 最佳实践与注意事项
### 什么时候用 speckit
**适合的场景**:
* 新功能开发(涉及 3+ 个文件)
* 需求不完全明确时(通过 clarify 澄清)
* 多人协作项目(规格作为共识)
* 重要功能(需要可追溯性)
**不适合的场景**:
* 简单 bug 修复
* 一行代码的改动
* 紧急热修复
* 纯探索性实验
### 常见陷阱
在使用 speckit 的过程中,有几个常见的陷阱需要注意:
**陷阱 1:规格太模糊**
症状:AI 生成的代码与预期差距大,需要大量返工。
```markdown
# ❌ 模糊的规格
用户可以搜索文章
# ✓ 清晰的规格
- FR-001: 用户可以按标题关键词搜索文章
- FR-002: 搜索结果按相关度排序,每页显示 10 条
- FR-003: 搜索词高亮显示在结果中
- FR-004: 空搜索词时显示热门文章
```
解决方案:运行 `/speckit.clarify`,或手动补充功能需求和成功标准。
**陷阱 2:规格太详细**
症状:AI 被限制得无法发挥,生成的代码过于僵硬,或者直接忽略部分指令。
```markdown
# ❌ 过度详细(指定实现细节)
使用 lodash 的 debounce 函数,延迟 300ms,
用 useCallback 包裹,依赖项为 [searchTerm]...
# ✓ 恰当的详细程度(只说做什么)
搜索输入应该防抖,避免频繁请求
```
解决方案:保持规格在「做什么」层面,把「怎么做」留给 Plan 阶段。
**陷阱 3:跳过 Plan 阶段**
症状:Tasks 太粗糙或太碎片化,实现时频繁返工,任务之间依赖关系混乱。
解决方案:复杂功能务必完成 Plan 阶段。Plan 不仅产出技术方案,还帮助识别潜在的架构问题。
**陷阱 4:不审查直接合并**
症状:上线后发现边界情况未处理、安全漏洞、性能问题。
解决方案:参考上文「实现后的审查」,始终在合并前运行测试和代码审查。
### 常见问题
**Q: 每个功能都要走完整流程吗?**
不必。简单改动可以直接编码,复杂功能建议至少完成 specify + plan。
**Q: 规格写得很详细,但 AI 还是生成了不符合预期的代码?**
检查规格是否真的「详细」。很多时候我们以为说清楚了,实际上还有模糊点。尝试运行 `/speckit.clarify` 看看有没有遗漏。
**Q: 可以跳过某些步骤吗?**
可以。最小流程是 specify → tasks → implement。但跳过 clarify 和 plan 可能增加后期返工风险。
**Q: 如何修改已生成的规格?**
直接编辑 spec.md 文件即可。修改后建议重新运行 plan 和 tasks 以保持一致性。
**Q: AI 生成的代码完全不对,怎么调试?**
分几步排查:
1. **检查规格**:规格是否真的清晰?尝试运行 `/speckit.clarify` 看有没有遗漏
2. **检查计划**:plan.md 中的技术方案是否合理?如果不合理,直接编辑后重新生成 tasks
3. **缩小范围**:让 AI 只执行一个任务,观察输出是否符合预期
4. **添加约束**:在 constitution.md 中添加更明确的技术偏好
**Q: Plan 和 Tasks 之间不一致怎么办?**
运行 `/speckit.analyze` 可以检测不一致。常见原因:
* Plan 更新后忘记重新生成 Tasks
* 手动编辑了 Tasks 但没更新 Plan
* 规格变更后只更新了部分文档
解决方案:以 spec.md 为准,依次重新生成 plan.md 和 tasks.md。
**Q: 如何处理跨功能的依赖?**
如果功能 B 依赖功能 A,有两种方式:
1. **合并规格**:把 A 和 B 写进同一个 spec.md,让 AI 统一规划
2. **分阶段开发**:先完成 A 的完整流程,再开始 B 的 specify
不建议同时开发有依赖关系的多个功能,容易造成集成问题。
***
## 小结
Speckit 的核心价值不是增加流程,而是**把隐性知识显性化**。当你被要求写下用户故事、功能需求、成功标准时,那些你以为「不言自明」的细节就会浮出水面。
记住这个流程:
```
Specify → Clarify → Plan → Tasks → Implement
需求 澄清 设计 分解 执行
```
每一步都在为下一步减少歧义。最终,AI 拿到的是清晰的任务清单,而不是模糊的意图描述。
现在,回到你的项目,试着用 `/speckit.specify` 开始你的第一个规格驱动开发流程吧。
### 延伸阅读
* 《[规格驱动开发是什么](/docs/notes/speckit/concept)》— 回顾核心理念
* 《[GSD 深度解析](/docs/notes/gsd/concept)》— 另一个采用规格驱动思想的上下文工程系统
* 《[Claude Skills 是什么](/docs/notes/claude-skills/concept)》— Speckit 本身就是一个 Claude Skill
# AI 时代的 TDD:先让模型撞上红灯
## 先说结论
AI 写代码以后,TDD 不是过时了,而是换了一个位置。
过去我们说 TDD,常常是在说程序员的自律:先写测试,再写实现,小步重构。到了 AI 编程里,它更像一套刹车系统。因为模型最擅长的事,正好也是最危险的事:它能很快写出一大段看起来完整的代码。
你让它实现一个功能,它可能十几秒就给你:
* 一个实现文件
* 一组测试
* 一段解释
* 一句“已完成”
问题是,“看起来完整”不是工程意义上的完成。工程意义上的完成,至少要回答:
> 这个行为有没有被一个明确的失败测试定义过?\
> 这个失败有没有因为实现而变绿?\
> 变绿以后,我们有没有在不改行为的前提下整理代码?
这就是 AI 时代重新谈 TDD 的原因。
它不是为了让流程显得高级,而是为了把“相信模型”换成“相信反馈”。
## 一、常见误解:TDD 不是“先写测试”
很多人讨厌 TDD,是因为他们理解成了一个仪式:
```text
先写测试。
再写代码。
最后跑一下。
```
这当然无聊,而且很容易变成形式主义。
真正有用的 TDD,不是“测试文件出现得比较早”,而是“失败出现得足够早”。
### 关键不是测试,而是红
TDD 的第一步叫 RED,不叫 TEST。
RED 的意思是:先写一个测试,让系统明确失败。这个失败必须满足三件事:
1. 它确实失败。
2. 它因为目标行为不存在而失败。
3. 它失败的方式和你的预期一致。
如果没有先看到红,后面的绿就没有意义。
比如你要实现 `slugify("Hello World") -> "hello-world"`。一个有价值的 RED 不是“我写了个测试文件”,而是:
```text
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
Failure: NameError: name 'slugify' is not defined
Reason: 目标函数还不存在,符合预期
```
这时测试才变成了规格。它告诉你:下一步实现只需要让这一条行为成立。
### 先绿再补测试,通常是在补故事
AI 很容易走另一条路:先写实现,再补测试。
这在体验上很顺。你看到代码已经跑了,测试也有了,心里会觉得“差不多”。但它有一个致命问题:测试很可能只是对当前实现的追认。
它不是在问“需求应该是什么”,而是在问“当前代码怎么写才容易通过”。
这就是为什么 AI 写测试经常会出现这些味道:
* 断言太贴合当前实现
* mock 太多,真实边界没测到
* 只测 happy path
* 为了让现有代码通过,把断言写得很宽
* 没有一个测试能证明旧代码原本是错的
TDD 要反过来:先让需求变成失败,再让代码追上需求。
## 二、为什么 AI 时代更需要 TDD
AI 编程的核心矛盾,不是“代码写得慢”,而是“反馈来得晚”。
没有 TDD 时,你通常是这样工作:
```text
描述需求 -> AI 写一堆代码 -> 人肉看 diff -> 跑一下 -> 发现问题 -> 回头修
```
问题会堆到最后。等你发现它错了,可能已经有三类东西混在一起:
* 需求理解错了
* 实现路径错了
* 重构把旧行为改坏了
TDD 的作用,是把这个长链条切短。
### 它给模型一个可判定目标
“写得优雅一点”不是目标。
“实现用户登录态过期后自动跳回登录页”也还不够具体。
更好的目标是:
```text
当 access token 过期时:
1. 请求返回 401。
2. 客户端清理本地 session。
3. 用户被重定向到 /login。
4. 原始目标地址被保存在 redirect 参数里。
```
再往前一步,把其中一条变成失败测试:
```text
given expired session
when user opens /settings
then app redirects to /login?redirect=/settings
```
这时 AI 不再是在猜“登录态过期应该怎么处理”,而是在完成一个明确行为。
### 它把大任务拆成小闭环
AI 最容易失控的地方,是一口气做完。
一次性让它实现登录、权限、刷新 token、错误提示、路由跳转,最后你会得到一个很大的 diff。它也许能跑,但 review 成本很高。你要同时判断业务、状态、路由、边界、测试和重构。
TDD 的节奏更像这样:
```text
一个行为 -> 一个失败测试 -> 最小实现 -> 变绿 -> 再下一个行为
```
每次只推进一小段。小到你能看懂,小到 AI 不容易编故事,小到失败时能快速定位。
### 它限制模型“顺手发挥”
AI 的一个常见问题是热心过度。
你让它修一个边界 bug,它顺手抽了 helper;你让它加一个测试,它顺手改了实现;你让它重构,它顺手改了行为。
TDD 用阶段把这些动作分开:
| 阶段 | 可以做什么 | 不该做什么 |
| -------- | ------- | ----- |
| RED | 写一个失败测试 | 写生产实现 |
| GREEN | 写最小实现 | 改测试凑绿 |
| REFACTOR | 整理结构 | 引入新行为 |
这张表比“请谨慎一点”有用。它让模型知道现在处在哪个阶段,也让人类更容易发现越界。
## 三、红绿重构:三道门,不是三句口号
“Red, Green, Refactor” 很容易被说成口号。真正用起来,它应该像三道门。每过一道门,都要留下证据。
### 第一门:RED,证明需求还没被满足
RED 阶段最重要的问题是:
> 这个测试如果失败,是否能证明我们还缺一个目标行为?
一个坏 RED:
```py
assert True
```
一个也不太好的 RED:
```py
assert "hello" in format_title("Hello World")
```
它太宽了。很多错误实现也能通过。
更好的 RED:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
这条测试很小,但它清楚。它指定了输入、输出和行为。
### 第二门:GREEN,只让当前测试通过
GREEN 阶段不是写最终架构。
它的任务只有一个:用最少代码让当前失败测试通过。
这句话听起来反直觉。很多人会担心“最少实现”会不会太丑。会,有时候会丑。但它的价值在于保持设计压力。
如果第一条测试是:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
一个可以接受的 GREEN 可能只是:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
你不需要立刻支持中文、重音符号、连续标点、emoji、SEO 特例。那些应该由后面的测试推动。
### 第三门:REFACTOR,只改结构,不改行为
REFACTOR 阶段最容易被 AI 搞混。
它会把“整理代码”理解成“顺便增强一下”。这不行。重构的定义很窄:外部行为不变,内部结构变好。
好的重构像这样:
* 改一个更准确的变量名
* 抽出重复表达式
* 拆掉过深的条件分支
* 移动函数位置,让模块职责更清楚
坏的重构像这样:
* 顺手支持了新输入
* 顺手改了错误提示
* 顺手换了依赖
* 顺手改了测试断言
判断标准很简单:
> 如果这个提交只叫 `refactor:`,测试前后应该一样绿,用户行为也应该一样。
## 四、好测试的味道
TDD 不是测试越多越好。AI 也很擅长生成一堆没什么价值的测试。
更重要的是测试的味道。
### 好测试像规格
好测试读起来应该像一句业务规格:
```text
当用户没有权限时,保存按钮不可点击。
当标题为空时,表单显示错误信息。
当重复提交同一个请求时,只创建一条记录。
```
它关心的是外部行为,而不是内部怎么做。
坏测试则像实现笔记:
```text
应该调用 validateInput 三次。
应该读取 state.user.flags。
应该触发 handleClick 内部函数。
```
实现细节被绑住以后,重构就会很痛。你只是改了内部结构,测试却大面积失败。这样的测试不是保护代码,而是在冻结代码。
### 好测试有边界
一个测试最好只回答一个问题。
如果一个测试同时断言:
* 格式正确
* 权限正确
* 网络请求正确
* toast 文案正确
* 数据库状态正确
它失败时你很难知道问题在哪。
AI 特别容易写这种“大而全”的测试,因为它想一次证明很多东西。TDD 要反过来:一个行为,一个失败,一个实现。
### 好测试会让实现难以作弊
如果测试只覆盖一个过于特殊的输入,AI 可能写出刚好匹配的假实现。
比如:
```py
def slugify(text: str) -> str:
return "hello-world"
```
第一条测试能让它过,但第二条测试就会逼出真正逻辑:
```py
def test_slugify_handles_another_title():
assert slugify("Test Driven Development") == "test-driven-development"
```
所以 TDD 不是永远只写一个测试,而是每一轮只新增一个行为压力。压力逐步增加,设计逐步长出来。
## 五、AI 会怎么绕过 TDD
这部分必须讲清楚,因为 AI 不会天然尊重测试。
它的优化目标很简单:完成你刚才说的任务。如果你说“让测试通过”,它可能会做出一些人类不想要的动作。
### 第一种:改测试凑绿
最典型:
* 把 `assert slugify("Hello World") == "hello-world"` 改成当前输出
* 删除失败断言
* 给测试加 `skip`
* 把严格断言改成宽松断言
这不是 TDD,这是把红灯拆掉。
### 第二种:写过拟合实现
比如测试只有一个输入:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
模型可能写出:
```py
def slugify(text: str) -> str:
if text == "Hello World":
return "hello-world"
return text
```
这时你不需要骂它。你需要继续加下一条行为,让实现无法继续硬编码。
### 第三种:用 mock 盖住真实边界
AI 很喜欢 mock。mock 让测试容易写,也让很多真实问题消失。
不是说不能 mock,而是要问:
> 我现在 mock 掉的,是慢依赖,还是我真正想验证的边界?
如果你要验证支付回调解析,却把解析层 mock 掉了,测试就没有意义。
## 六、什么时候不该用 TDD
TDD 有价值,但不是所有事情都值得套。
### 不适合的场景
* 纯视觉微调
* 一次性脚本
* 技术探索 demo
* 需求本身还没想清楚的原型
* 测试框架还没搭好的仓库
这些场景先追求探索速度,不要被流程拖住。
### 适合的场景
* bug 修复
* 权限、计费、状态机
* 数据转换和边界处理
* 会长期维护的核心模块
* AI 会反复修改的代码路径
判断标准不是“这个功能大不大”,而是:
> 如果它错了,代价是否明显?
代价明显,就值得先写测试。
## 七、一个可执行的心法
如果只给 AI 一句话,我不会说:
```text
请高质量实现这个功能。
```
我会说:
```text
先写一个失败测试,运行它,确认失败原因符合预期。不要写实现,直到我说 go。
```
这句话的质量更高,因为它不是要求模型“表现好”,而是要求它进入一个可检查的流程。
再完整一点:
```text
每轮只处理一个行为。
RED:写一个失败测试并运行。
GREEN:写最小实现,不改测试。
REFACTOR:只在绿色状态下整理结构。
每轮报告测试文件、命令、失败原因、通过结果。
```
这就是 AI 时代 TDD 的核心。
不是迷信测试,不是迷信流程,而是让每一步都有证据。
## 收尾
AI 编程最需要的不是更多代码,而是更短的反馈。
TDD 的价值就在这里:它把“我觉得应该对”变成“这里有一个失败,后来它变绿了”。这个变化很小,但足够真实。
如果你只记一句话,记这个:
> 不要让 AI 直接交付代码。先让它交付一个红灯,再让它把红灯变绿。
下一篇 [实战指南](/docs/notes/tdd-with-ai/practice) 会把这套节奏变成可以直接复制的工作流。
## 推荐资源
# AI 时代的 TDD:Codex 实战手册
## 先给一张地图
[概念篇](/docs/notes/tdd-with-ai/concept) 讲的是:AI 编程里,TDD 的价值不是“先写测试”这个仪式,而是先制造一个可验证的红灯,再让实现把它变绿。
这篇讲怎么落地到 Codex。
别一上来就问“我要配哪些文件”。更好的问题是:
> 我怎么让 Codex 每次都按同一条 TDD 工作流行动?
这条工作流可以拆成四层:
| 层级 | 你在做什么 | 放在哪里 | 适合什么时候 |
| -- | ----- | ----------------------------------- | --------------- |
| L1 | 写项目纪律 | `AGENTS.md` | 所有项目都该有 |
| L2 | 固化流程 | `.agents/skills/tdd-codex/SKILL.md` | 反复用 TDD 做需求 |
| L3 | 隔离阶段 | `.codex/agents/*.toml` | 复杂任务,怕测试和实现互相污染 |
| L4 | 自动提醒 | `.codex/hooks.json` | 重要仓库,怕 AI 偷改测试 |
最小可用版本是 L1 + L2。
完整防线是 L1 + L2 + L3 + L4。
## 一、先定义“完成”是什么
如果没有完成标准,Codex 很容易把“写了代码”当成“做完了”。
TDD 场景里,完成标准应该更具体。
### 每一轮都要交付证据
让 Codex 每轮都报告这六项:
```text
Behavior: 这一轮实现哪个行为
Test: 测试文件和测试名
Command: 跑了什么命令
RED: 失败原因是否符合预期
GREEN: 通过结果是什么
REFACTOR: 是否重构,为什么
```
这比一句“已完成”有用得多。
它让你知道模型真的走过了红绿循环,而不是先写完实现再补一个看起来合理的测试。
### 一轮只处理一个行为
这条很关键。
不要让 Codex 一次生成完整测试矩阵。那会变成“横向铺测试”:
```text
RED: test1, test2, test3, test4, test5
GREEN: 一次写一个大实现
```
你要的是纵向切片:
```text
RED test1 -> GREEN impl1 -> REFACTOR
RED test2 -> GREEN impl2 -> REFACTOR
RED test3 -> GREEN impl3 -> REFACTOR
```
第一轮实现会改变你对问题的理解。不要把所有测试一次性写死。
## 二、L1:把纪律写进 AGENTS.md
`AGENTS.md` 是 Codex 进入项目时会读取的说明文件。
OpenAI 官方文档说明,Codex 会先读取全局说明,再从项目根目录一路读到当前目录。每一层优先读取 `AGENTS.override.md`,否则读取 `AGENTS.md`。越靠近当前目录的说明越晚出现,因此优先级更高。默认合并上限是 `32 KiB`,所以这里不能写成长篇教程。
### 它应该像项目交通规则
`AGENTS.md` 不负责教会 Codex 什么是 TDD。它只负责写清楚:在这个项目里,什么行为不允许。
可以直接放这段:
```markdown
# TDD Rules
- For new behavior and bug fixes, use red/green TDD.
- RED: write exactly one failing behavior test first.
- Run the smallest relevant test command and confirm the failure is expected.
- Do not edit production implementation during RED.
- GREEN: write the minimum production code required to pass the current failing test.
- Never modify, delete, skip, or weaken tests to make implementation pass.
- REFACTOR only after tests are green.
- Keep structural changes and behavior changes separate.
- Report Behavior, Test, Command, RED, GREEN, and REFACTOR for each cycle.
```
再补项目命令:
```markdown
# Verification
- Use `pytest` or the smallest relevant pytest command for Python behavior tests.
- Use `npm run types:check` only when this blog site's MDX or TypeScript changes.
- Use the smallest targeted command during RED/GREEN loops.
- If a command is slow, explain what targeted command was used first and what full command remains.
```
### 它不应该写成百科全书
坏的 `AGENTS.md` 会写满:
* TDD 历史
* 所有测试哲学
* 一大堆框架教程
* 复杂 prompt 模板
* 不同语言的完整规范
这些东西会稀释真正重要的规则。
我的建议是:`AGENTS.md` 只放常驻纪律。长流程放 skill。
## 三、L2:把流程做成 Codex Skill
`AGENTS.md` 解决“默认纪律”,skill 解决“完整流程”。
当你经常对 Codex 说“按 TDD 做”,就应该把这句话升级成项目级 skill。
### 目录结构
放这里:
```text
.agents/
skills/
tdd-codex/
SKILL.md
```
Codex 会从当前目录一路向上扫描 `.agents/skills`。仓库根目录下的 skill,适合放团队共同使用的工作流。
### 最小可用 SKILL.md
```markdown
---
name: tdd-codex
description: Implementing or fixing maintainable code with Codex using strict red-green-refactor TDD. Use for new behavior, bug reproduction, behavior tests, or safe AI coding.
---
# TDD Codex Workflow
Use one behavior slice per cycle.
## Phase 0: Scope
Identify one observable behavior.
Name the public API, user flow, or integration boundary under test.
Do not edit production code.
## Phase 1: RED
Write exactly one failing behavior test.
Prefer public behavior over implementation details.
Run the smallest relevant test command.
Confirm the failure is expected.
Stop and report:
- Behavior
- Test file
- Command
- Failure reason
## Phase 2: GREEN
Write the minimum production code to pass the current failing test.
Never modify, delete, skip, or weaken tests to pass.
Do not add speculative features or abstractions.
Run the same test command.
Report the passing result.
## Phase 3: REFACTOR
Only refactor after tests are green.
If the code is already simple, skip.
If refactoring, make one structural change at a time.
Run tests after each refactor.
Do not change behavior.
## Cycle Report
Return:
- Behavior:
- Test:
- Command:
- RED:
- GREEN:
- REFACTOR:
- Next slice:
```
### 调用方式
以后你可以这样说:
```text
用 tdd-codex skill 做这个需求。
每轮只处理一个行为。
先 RED,确认失败后停下来,不要直接写实现。
```
或者更短:
```text
用 tdd-codex。先写红灯,等我说 go。
```
重点不是 prompt 多漂亮,而是它每次都能把 Codex 拉回同一条轨道。
## 四、L3:用 Subagents 隔离红绿重构
不是所有任务都需要 subagents。
但当任务复杂、测试容易被实现污染、重构容易失控时,把 RED、GREEN、REFACTOR 拆给不同 agent 会更稳。
### 什么时候值得拆
适合拆:
* 权限、计费、状态机
* 多模块功能
* bug 很隐蔽,需要先写复现测试
* 模型总是改测试凑绿
* 你希望有人只负责 review 测试质量
不适合拆:
* 小工具函数
* 文案改动
* 纯视觉微调
* 一次性脚本
拆 agent 的成本是真实存在的。只在隔离收益大于沟通成本时使用。
### RED agent
```toml
# .codex/agents/tdd-test-writer.toml
name = "tdd_test_writer"
description = "RED phase agent. Writes one failing behavior test and stops before implementation."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the RED phase.
Write exactly one behavior-focused test for the requested slice.
Prefer public APIs and user-visible behavior over implementation details.
Run the smallest relevant test command.
Confirm the test fails for the expected reason.
Do not edit production implementation.
Do not add multiple tests at once.
Return Behavior, Test, Command, and RED failure reason.
"""
```
### GREEN agent
```toml
# .codex/agents/tdd-implementer.toml
name = "tdd_implementer"
description = "GREEN phase agent. Implements the minimum production code to pass the current failing test."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the GREEN phase.
Read the failing test and relevant production code.
Write the minimum implementation required to pass the current test.
Never modify, delete, skip, or weaken tests to make them pass.
Do not add speculative features, helpers, configuration, or abstractions.
Run the relevant tests and return the command plus passing output.
"""
```
### REFACTOR agent
```toml
# .codex/agents/tdd-refactorer.toml
name = "tdd_refactorer"
description = "REFACTOR phase agent. Improves structure only after tests are green."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the REFACTOR phase.
Start by running the relevant tests to confirm the code is green.
Look for duplication, unclear names, needless branching, or misplaced responsibility.
Skip refactoring when the code is already simple.
If you refactor, make one structural change at a time.
Run tests after each refactor.
Never change behavior in this phase.
"""
```
### 主会话怎么指挥
```text
按三阶段 TDD 做这个 slice:
1. tdd_test_writer 只写一个失败测试,并确认 RED。
2. 等我确认后,tdd_implementer 写最小实现,并确认 GREEN。
3. tdd_refactorer 判断是否需要结构重构。
不要批量铺测试。
不要在 GREEN 阶段修改测试。
```
这里的重点是隔离上下文。写测试的人尽量不要被实现细节影响;写实现的人不能随手动测试;重构的人不能引入新行为。
## 五、L4:用 Hooks 盯住测试 diff
如果只靠规则,模型仍然可能越界。
最常见的越界就是:测试红了,模型为了变绿,顺手把测试改了。
hooks 的价值不是“绝对安全”,而是把这种动作立刻暴露出来。
### 启用 Codex hooks
先在配置里打开 feature flag:
```toml
# ~/.codex/config.toml 或 /.codex/config.toml
[features]
codex_hooks = true
```
Codex 会在活动配置层旁边查找 hooks。常见位置:
* `~/.codex/hooks.json`
* `~/.codex/config.toml`
* `/.codex/hooks.json`
* `/.codex/config.toml`
项目级建议先用 `/.codex/hooks.json`,因为它能跟着仓库走。
### hooks.json
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/watch-test-edits.sh\"",
"timeout": 10,
"statusMessage": "Checking test file edits"
}
]
},
{
"matcher": "Bash|apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/run-fast-check.sh\"",
"timeout": 120,
"statusMessage": "Running fast checks"
}
]
}
]
}
}
```
### 检查测试文件是否被改
```bash
# .codex/hooks/watch-test-edits.sh
#!/usr/bin/env bash
set -euo pipefail
changed_tests="$(
git diff --name-only |
grep -E '(^|/)(__tests__|tests?)/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.go$|test_.*\.py$' || true
)"
if [ -n "$changed_tests" ]; then
cat < "hello-world"
- trim leading/trailing spaces
- collapse repeated spaces into one hyphen
- remove punctuation
- normalize "Café" -> "cafe"
- empty input returns empty string
```
### Step 1:第一盏红灯
```text
读取 SPEC.md。
只实现第一条行为:"Hello World" -> "hello-world"。
先 RED:只写一个失败测试,运行它,确认失败。
不要写生产实现。
```
理想输出:
```text
Behavior: basic title becomes lowercase hyphenated slug
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
RED: failed because slugify is not defined
```
这时才可以继续。
### Step 2:最小变绿
```text
go
```
Codex 写最小实现:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
然后报告:
```text
GREEN: pytest tests/test_slugify.py -q passed
REFACTOR: skipped, implementation is still simple
Next slice: trim leading/trailing spaces
```
### Step 3:第二盏红灯
```text
继续下一条:去掉首尾空格。
先 RED,只写一个测试。
```
测试:
```py
def test_slugify_trims_spaces():
assert slugify(" Hello World ") == "hello-world"
```
如果当前实现输出 `-hello-world-`,红灯成立。
然后 GREEN:
```py
def slugify(text: str) -> str:
return text.strip().lower().replace(" ", "-")
```
### Step 4:别急着抽象
到这里很多 AI 会想抽一个 `normalizeInput`、`removePunctuation`、`toAscii`。
先别急。
TDD 的设计应该被测试压力推出来,不是被想象推出来。等你加到 unicode、标点、空字符串,结构压力真的出现,再重构。
## 八、日常用法速查
### 新功能
```text
用 TDD 实现这个需求。
每轮只处理一个行为。
先写一个失败测试并运行确认 RED。
不要写生产实现,直到我说 go。
```
### 修 bug
```text
先写一个能复现这个 bug 的失败测试。
确认它因为这个 bug 失败后,再写最小修复。
不要改测试来适配当前实现。
```
### 复杂功能
```text
先不要写代码。
请给出 TDD 分解计划:
- 外圈集成测试是什么
- 内圈每个行为 slice 是什么
- 每轮用什么命令验证
- 哪些地方不能 mock
等我确认后再开始 RED。
```
### Review
```text
Review 这次改动,重点看:
- 是否先有失败测试
- 测试是否测行为而不是实现
- 是否存在为了通过而弱化测试
- 结构改动和行为改动是否混在一起
- 是否缺少外圈集成测试
```
## 九、最后的检查清单
每次让 Codex 做 TDD,最后用这张表检查。
| 问题 | 合格标准 |
| ----------- | -------------- |
| 真的先红了吗 | 有失败命令和失败原因 |
| 红得对吗 | 失败原因对应目标行为缺失 |
| 一轮只做一个行为吗 | 没有批量铺测试 |
| GREEN 改测试了吗 | 没有改测试凑绿 |
| 测试测行为吗 | 不依赖内部实现细节 |
| 重构混行为了吗 | 结构改动和行为改动分开 |
| 有完整验证吗 | 目标测试和必要全量检查都跑过 |
如果这张表过不了,就不要急着合并。
## 推荐资源
# When AI Can Help You Pass Without Studying, What's Left of College?
Anthropic recently produced an interview featuring four students from Princeton, Berkeley, the London School of Economics, and Arizona State University, discussing the real state of AI on campus. The nearly 40-minute conversation had no marketing speak — just genuine confusion, anxiety, and reflection.
This interview reveals a deeper issue: **AI isn't just changing how we learn — it's dismantling the underlying logic of the entire education system**.
***
## An Unavoidable Reality
The interview opens with a direct question from the host: What's the current campus atmosphere around AI?
The answer: **90% of students are using AI**. Not occasionally, but as part of their daily workflow — summarizing lecture notes, answering problem sets, getting homework feedback, analyzing business cases, conducting market research, completing financial studies. Some students even use it for quizzes, and the reasoning is practical: when you're a graduate student juggling multiple jobs, you don't always have the time.
But what's more interesting is that while almost everyone is using it, nobody knows the rules. Some courses explicitly ban AI, some actively encourage it, and most exist in a gray area. Students don't know where the line is, and professors don't know how to enforce it. This state is called the "gray zone" — wanting to use it but fearing violation; not using it but feeling left behind.
The most dangerous thing about this gray zone isn't that students might violate rules — it's that **it prevents truly valuable discussions from happening**. Students can't openly share AI best practices, professors can't guide students on responsible tool usage, and the entire academic community falls into an awkward state of surface-level prohibition and private use.
When rules can't be effectively enforced, they become a tool for filtering "those who can disguise their usage" from "those who can't." A state of explicit prohibition but widespread private use **won't stop students from using AI — it will only stop them from openly discussing how to use it better**.
***
## AI Is a Mirror
One observation from the interview is particularly sharp: artificial intelligence, and especially how students use it, is very revealing of their motivations.
The insight behind this statement: **AI has become a mirror, reflecting your real purpose for attending college**.
The interview categorizes university goals into three types: first, deep learning of specialized knowledge and gaining profound understanding of a field; second, career preparation — finding good jobs and building professional networks; third, expanding social connections, enjoying campus social life, and experiencing university culture. Every student weighs these three goals differently, and their AI usage precisely exposes these weightings.
If you only care about "passing exams" and "getting the degree," you'll submit AI output directly as your assignment. This isn't a moral judgment — it's reality. When technology lets you achieve your goal with minimal cost, why take the longer route? If your goal was always to get a degree for a job, using AI to complete assignments is an entirely rational choice.
But if you genuinely want to learn, to deeply understand a field, you'll use AI as a conversation partner. You'll ask it questions, have it explain concepts, then rephrase in your own words. You'll have it write a first draft of code, then refactor and optimize it yourself. You'll ensure that at every step, you truly understand what's happening.
This divergence exists not only between different students but between different majors. Humanities students tend to opt out of AI because their learning requires close reading — carefully reading original texts, savoring linguistic nuances, understanding authorial intent. AI undermines this process because it provides summaries and paraphrases, not direct experience of the original text. Engineering and business students heavily use AI because it lowers technical barriers, enabling people without computer science backgrounds to write code, build websites, and analyze data.
**This polarization is fundamentally about different understandings of "what's worth learning."** For humanities students, the experience of reading Shakespeare in the original is itself the learning; for engineering students, what matters is whether you can solve the problem, not whether you personally wrote every line of code. AI makes this difference more pronounced.
***
## Tool or Crutch? A Simple Litmus Test
A key question in the interview: How do you distinguish whether AI is a tool or a crutch?
The students' answers were remarkably consistent: **whether you can explain it**.
If you can't explain what you created, can't articulate what role AI played in it, then it's a crutch. If you can explain it as if to a fifth grader, can provide both low-level and high-level explanations, then it's a tool.
This standard seems simple, but it touches the essence of learning. The core logic of the Feynman technique is: if you can't explain a concept in simple language, you haven't truly understood it. Learning in the AI era follows the same principle — if you can't explain what AI did for you, then you're "outsourcing thinking," not "augmenting thinking."
The interview mentions an interesting example: a student developed a tool that takes lecture slides as input, and the AI generates professor-like annotations beside each slide. The student said: "It works well because I've already prompted it to know what I want to learn — definitions of certain things on the slides. Slides are sometimes abstract, lacking background information, and need context added alongside them."
The key is "I've already prompted it to know what I want to learn." This student knows exactly where their knowledge gaps are, knows what kind of help they need, then actively guides the AI to provide it. That's a tool. If they had just thrown the slides at AI saying "summarize this lecture for me" and memorized the output, that would be a crutch.
**The difference is agency and understanding.** Tool users know what they're doing and control the entire process. Crutch users hand control to the technology and become passive receivers.
***
## Schools Aren't Just Slow — They're Fundamentally Unable to Cope
The interview mentions some institutional attempts. An LSE mandatory course began teaching students how to use Claude — conversing with it, assigning it different roles, then requiring students to submit conversation logs showing how they interacted with the AI. Arizona State University's career center built a prompt library providing prompt templates for different scenarios. These are good attempts, with the core idea being: don't ban AI, teach students to use it responsibly.
But these are the minority. Most schools are still debating "should students be allowed to use AI." Some professors say yes but require disclosure; some courses ban it outright; others simply don't address it, defaulting to assuming students won't use it. No integration framework, no unified standards — the entire system is in chaos.
The deeper issue: **this problem fundamentally cannot be solved through regulations**.
Traditional education's regulatory logic is: schools set rules, students follow rules, violators get punished. This logic requires "violations can be detected." But AI breaks this premise.
You can ban students from using AI when submitting assignments, but you can't monitor whether students used AI during their thinking process. You can use AI detection tools, but their accuracy is far from sufficient to serve as a basis for punishment — false positive rates are too high, and students quickly learn to bypass detection. More importantly, **fundamentally, you cannot distinguish between "high-quality work completed with AI assistance" and "high-quality work completed independently,"** because good AI usage should be seamless.
One quote from the interview is blunt: fundamentally, no policy is going to change the way that students use AI. The onus is on the students. This isn't passing the buck — it's reality.
When technology makes "passing exams without studying" possible, schools face not a management problem but an existential one: **if exams can't prove learning, what's the point of school?**
***
## The Underlying Logic of Education Has Broken
This question touches the fundamental contradiction of the education system.
Traditional education rests on several core assumptions: first, knowledge is scarce and requires specialized institutions (schools) and professionals (professors) to transmit it; second, learning outcomes can be measured through exams; third, degrees prove you've mastered knowledge in a field, qualifying you for related work.
AI is breaking these assumptions one by one.
**Knowledge is no longer scarce.** YouTube has free Stanford courses, Claude can answer your questions anytime, and GitHub has countless open-source projects to learn from. You don't need to attend school to access knowledge, and you don't even need a paid subscription for basic AI tutoring.
**Exams can't measure learning.** When AI can complete most exam questions, exams shift from "tools measuring understanding" to "tools measuring whether you can use AI." This doesn't mean exams are entirely useless, but they can no longer accurately distinguish between "someone who truly learned" and "someone who knows how to use tools."
**The value of degrees is declining.** When employers realize degrees can't guarantee candidates actually mastered relevant knowledge, they'll place more weight on demonstrated ability — portfolios, project experience, internship performance. Degrees are downgraded from "proof of competence" to "baseline entry requirement."
**When these assumptions break, the education system's value proposition needs redefinition.**
The interview offers an answer: the value of college shifts from "**transmitting knowledge**" to "**providing an environment**." An environment where you can make mistakes, explore, and bounce ideas off others. You can spend weekends with roommates on a "bucket list before graduation," test "probably stupid" ideas at hackathons, argue with professors, discuss with classmates, and learn from failure without risking your career.
Student projects mentioned in the interview illustrate this well — "automated course registration alerts," "empty classroom finders," "graduation bucket list leaderboards." None are technically complex; many creators don't even have computer science backgrounds. But they spring from genuine human emotions: fear of missing out, pursuit of convenience, cherishing university life.
**When technical barriers drop, what matters is no longer "can you code" but "what problem do you want to solve."** College provides a space for freely exploring these questions, an environment for turning ideas into reality and learning from failure.
AI can complete your assignments, but it can't live through this time of "making mistakes and exploring" for you.
***
## The Employment Market Paradox
The latter part of the interview touches on employment, revealing another paradox.
Students use AI to write resumes, companies use AI to screen resumes. The entire hiring cycle becomes talking to screens — first using AI to write cover letters, then answering questions on video, finally receiving AI-generated rejection letters. From resume submission to rejection, it might take just 15 minutes. Very efficient, but very little humanity.
This creates an "AI vs AI" job market. Students train AI to write "good" resumes, companies train AI to screen "good" candidates. Real humans play an increasingly small role. There's no chemistry in talking to a screen, no way to showcase qualities that are hard to quantify but deeply important — humor, adaptability, the subtleties of teamwork.
But the flip side of the paradox: AI proficiency itself has become a new competitive advantage. The Big Four consulting firms used to recruit generalist MBAs; now they specifically look for AI-capable MBAs. If you know how to apply AI across different industries, you're their top candidate.
**The paradox: AI makes the job market colder while simultaneously making it value AI skills more. You can't escape it — you can only learn to use it effectively.**
This circles back to the core question: what does "effective use" mean? Not using ChatGPT to write emails, but being able to identify which problems are suited for AI, design prompts to guide AI toward the results you need, and evaluate AI output quality while making necessary corrections.
This capability isn't developed by banning AI, but through extensive practice and trial and error. This is also why schools that proactively embrace AI — building Claude Builder Clubs, organizing hackathons — are providing more valuable education. They're letting students learn to collaborate with AI in a relatively safe environment.
***
## The Transfer of Responsibility
Perhaps the interview's most core insight is this: when technology enables you to "pass exams without studying," the meaning of learning itself becomes a question each person must answer for themselves.
This is a transfer of responsibility. From **school to student, from rules to self-discipline, from external motivation to internal motivation**.
Traditional education's motivation is external: you need to pass exams to get a degree, and you need a degree to get a job. This external incentive system drives students to learn. But when AI lets you pass without studying, this incentive system fails.
What remains is only internal motivation: Do you genuinely want to learn? Are you truly interested in this field? Do you really want deep understanding, or just a degree?
There's an interesting detail in the interview. A student mentions the downside of graduate school — you're juggling multiple jobs, running out of time, so sometimes you use AI to quickly complete quizzes. But then they add: grad school is supposed to be a period where you expand your critical thinking, where you demonstrate a more decisive side. **They recognized the contradiction but chose efficiency.**
This isn't moral judgment. Under real-world pressure, efficiency often trumps ideals. But this choice reveals a fact: **when external pressure (completing quizzes) conflicts with internal motivation (deep learning), many people choose the former**.
AI makes this conflict sharper because it makes "getting through exams" extremely easy. In the pre-AI era, even if you just wanted to get through exams, you still had to learn something to pass. AI removes this intermediate step — you can pass without learning at all.
**This forces everyone to confront the question: Why are you actually in college?**
If the answer is "to get a degree for a job," then using AI to complete assignments is perfectly reasonable. If the answer is "I genuinely want to learn this field," then you need to actively resist the shortcut temptation that AI brings.
Schools can't make this choice for you. Rules can't force you to develop internal motivation. **This is your own responsibility.**
***
## Technology Won't Wait for You to Be Ready
A persistent attitude runs through the interview's conclusion: "We'll figure it out."
School rules can't keep up? Start using it first, then tell schools what works. AI might be used for cheating? Gradually learn to use it responsibly. Job market changed? Adapt to the new rules of the game.
This isn't blind optimism — it's realism. Technology is already here, and it won't wait for you to be ready before changing the world. You can choose to resist or adapt, but you can't choose to stop time.
**This generation's relationship with AI isn't fear, isn't blind embrace — it's feeling their way through chaos, learning through trial and error.**
The projects they're building in Claude Builder Club — not technically complex, but solving real problems. The ideas they're testing at hackathons — maybe silly, but at least they're trying. Their confusion in classrooms — rules aren't clear, but at least they're thinking.
This attitude of "**learning by doing**" might help them adapt to the AI era better than any set of regulations.
Kevin Kelly proposed the "technium" concept in *What Technology Wants*: technology as a whole seems to have its own will, wanting to become ever more powerful. An observation from the interview echoes this: over the past two years, whatever AI needed to continue developing, it got. Shifts in attitudes toward nuclear energy, discussions about space data centers — whenever a bottleneck might slow AI, the obstacle gets removed.
But a more accurate framing might be: **students create whatever they need**.
AI is just a tool. What determines the future is how this generation of students chooses to use it — to avoid thinking or augment it. To get through exams or explore the world. As a crutch or as a tool.
This is a choice schools can't control — only they can decide for themselves.
And from this interview, at least some students are seriously thinking about it. That might be enough.
# Think Like an Agent: Claude Code Team's Tool Design Philosophy
Thariq is an Anthropic engineer and one of the core builders of Claude Code. In late February, he published a long thread on X sharing five real-world cases about agent tool design from the Claude Code development process — not theoretical frameworks, but lessons from the trenches. The post garnered 3.51 million views, 209 replies, and 9,691 likes, sparking extensive high-quality discussion.
To introduce the core question of the entire piece, Thariq used a great analogy: imagine you're facing a tough math problem — what tools do you need?
**Pen and paper** is the bare minimum — but you're limited to manual calculation. A **calculator** is better — but you need to know how to use its advanced functions. A **computer** is the most powerful — but you need to know how to code.
Tool selection depends on the user's capability. Giving a computer to someone who can't program is worse than giving them a calculator. Giving a calculator to a programmer actually limits their potential.
The same applies to agents. The question isn't "what tool is most powerful" but rather "what tool best matches the model's current capabilities." This article shares the lessons the Claude Code team learned while searching for that match.
Here are three progressively deeper layers I distilled from the original thread and community discussion.
***
## Layer 1: Learning Tool Design from Failures
Intuition tells us that more tools mean a more capable agent. But the Claude Code team's experience says otherwise.
Claude Code currently has only about 20 tools, and the team constantly evaluates whether they truly need all of them. The bar for adding new tools is high because each addition gives the model one more option to consider — every additional tool adds "cognitive overhead" to the model's decision-making.
Apple's CodeAct research provides quantitative support for this insight: **a single code execution primitive outperforms a sprawling set of specialized tools by up to 20% on complex tasks.** Less really can be more.
The three iterations of the AskUserQuestion tool are the best illustration of this principle. The Claude Code team wanted to improve Claude's ability to ask users questions — while Claude could ask questions in plain text, answering those questions felt too time-consuming. How to reduce friction?
**First attempt**: Add a parameter to ExitPlanTool that lets it output a set of questions alongside its plan. Result — Claude got confused. Requiring it to simultaneously output a plan and questions about the plan created conflicts when user answers contradicted the plan.
**Second attempt**: Modify output instructions to have Claude ask questions in a specific markdown format, then parse and format them on the frontend. Result — unreliable. Claude would add extra sentences, omit options, or use completely different formatting.
**Third attempt**: Create a standalone AskUserQuestion tool. Claude can call it at any time, which pops up a dialog displaying the question and blocks the agent loop until the user responds. **It worked.**
Thariq wrote something fascinating in the original thread:
> Claude seemed happy to call this tool and we found it did a great job outputting to it. Even the best designed tool won't work if Claude doesn't understand how to call it.
**The success criterion for tool design isn't "it makes sense to humans" but "the model understands how to use it and is willing to use it."** This judgment requires carefully observing the model's actual behavior — call frequency, output quality, and whether it proactively uses the tool.
If tool count represents the spatial dimension of cognition, then tool timeliness is a lesson from the temporal dimension. **Tools that once helped the model can actually become constraints as the model improves.**
When Claude Code first launched, the team realized the model needed a to-do list to stay on track — it could write to-do items at the start and check them off as work was completed. They provided Claude with the TodoWrite tool. But even so, Claude frequently forgot what it was supposed to do.
The team's response was to insert system reminders every 5 turns, prompting Claude about its goals.
But as the model improved, the problem reversed: the model no longer needed to-do reminders, and actually found them constraining. **Being repeatedly reminded of its to-do list made Claude feel obligated to follow it strictly rather than adapting flexibly as needed.** Meanwhile, Opus 4.5 had significantly improved at using sub-agents, but how should sub-agents coordinate around a shared to-do list?
So the team replaced TodoWrite with the Task Tool. The difference is fundamental: Todos were about keeping the model on track — like a boss watching over an employee's task list; Tasks are more about facilitating communication between agents — like a team collaboration board. Tasks support dependencies, cross-sub-agent shared updates, and models can modify and delete them.
**From TodoWrite to reminders every 5 turns to Task Tool — that's three redesigns.** Not because the previous designs were "wrong," but because the model had grown. Tool design isn't a one-time effort; it needs to iterate continuously alongside model capabilities.
***
## Layer 2: Progressive Disclosure — From "Spoon-Feeding" to "Self-Searching"
This is what I consider the most practically valuable part of the entire thread.
Claude Code initially used a RAG vector database to find context for Claude. RAG was powerful and fast, but had two problems: first, it required indexing and configuration that could be fragile across different environments; second, and more fundamentally — **this approach was providing context to Claude rather than letting it find its own.**
The team made a key pivot: if Claude can search the web, why can't it search your codebase? By giving Claude the Grep tool, they let it search files and build context on its own.
**Over the course of a year, Claude evolved from being barely able to build context autonomously to performing nested searches across multiple layers of files, precisely finding the context it needed.** The key to this evolution wasn't giving Claude more information — it was giving it better search capabilities.
When Claude Code introduced Agent Skills, the team formally articulated the concept of **Progressive Disclosure**: allowing the agent to gradually discover relevant context through exploration.
The implementation is elegant: Claude can read skill files, and those files can reference other files, which the model can recursively read. A common use case for skills is adding more search capabilities to Claude — for example, giving it instructions on how to use an API or query a database.
The core logic of this layered strategy is: **context is a finite resource with diminishing marginal returns.** Dumping all information onto an agent at once not only wastes tokens but also dilutes truly important information. Providing information on demand, letting the agent decide when it needs deeper details — that's the scalable approach.
The Claude Code Guide sub-agent is another clever application of progressive disclosure. The team noticed Claude didn't know enough about how to use Claude Code itself — if you asked it how to add an MCP or what a slash command does, it couldn't answer.
They could have stuffed all the information into the system prompt, but users rarely ask these questions, and doing so would increase context erosion and distract Claude Code from its primary job: writing code.
First they tried giving Claude a documentation link to search on its own — it worked, but Claude would load massive amounts of results into context to find the right answer. The final solution was building a dedicated sub-agent: Claude Code Guide. This sub-agent has detailed search instructions and knows how to efficiently search documentation and what content to return. **No new tools added, yet Claude's action space was expanded.**
Lance Martin offered a complementary perspective in his article on agent design patterns: **rather than defining dozens of tools for an agent, give it a computer and let it orchestrate tools through code.** Claude Code's core abstraction is the CLI — the agent lives on your computer, accomplishing complex tasks through fundamental primitives like bash and the file system. A few atomic-level tools (like the bash tool) are more flexible and token-efficient than a massive tool set.
***
## Layer 3: Model Empathy — Thinking Like an Agent
Behind the cases from the first two layers — iterative tool design and progressive disclosure — lies a shared meta-methodology. Thariq stated it at the beginning of his thread: **see like an agent.**
This isn't a set of rules — it's a mindset. David Zhang gave it a name:
**Model empathy** — not designing "reasonable" tools from a human perspective, but thinking from the model's perspective about what it actually sees, how it understands things, and how it will use them.
This "mental model inversion" sounds simple but requires constant practice. You need to carefully read the agent's outputs — not looking at what it got right, but understanding why it made certain choices, where it hesitated, and where it took detours. These "anomalous behaviors" are often not bugs in the model but bugs in the tool design.
***
## Final Thoughts
Returning to Thariq's closing words:
> Experiment more, read your outputs, try new approaches. See like an agent.
As a heavy Claude Code user, my biggest takeaway from reading this thread is: **those seemingly "natural" features are backed by countless iterations of "this doesn't work, let's try something else."** AskUserQuestion took three attempts, TodoWrite was redesigned three times, RAG was replaced by Grep. Each improvement came not from inventing a cleverer solution, but from carefully observing the model's actual behavior.
The subtitle of this thread is "Seeing like an Agent" — seeing the world as an agent does. But from another angle, this is essentially the core of all good engineering practice: **don't design systems from your own perspective — design them from the user's perspective.** The only difference this time is that the user is an AI model.
Of course, these lessons also leave some open questions. All these tool iterations assume the agent is stateless — starting from scratch each session. What if the most important "tool" isn't in the action space but in persistent memory about how the codebase works? Claude Code later partially addressed this through CLAUDE.md and the memory system, but persistent state management remains an open challenge in agent design. From a broader perspective, action space design is essentially power design — what permissions you give AI determine what role it becomes, and the bottleneck is often not the model's capability but the permission boundaries you draw.
In the future, every developer building agents may need to master what David Zhang calls "model empathy." This isn't some mysterious ability — at its core, it's three things: **observe the model's actual behavior, read its outputs, and adjust your design based on what you see.**
See like an agent.
***
**Further Reading:**
# From Chip Wars to Space Data Centers: The Next Decade of AI
> Investing is a pursuit of truth. If you find the truth first, and you're right, that's how you generate alpha. And it has to be a truth that others haven't yet seen.
This quote comes from Atreides Management founder Gavin Baker's interview on Patrick O'Shaughnessy's podcast "Invest Like the Best." Gavin is regarded as one of the most passionate and insightful investors in tech investing, and this nearly two-hour conversation covered GPUs, TPUs, AI economics, space data centers, the future of SaaS, and even his life transition from ski instructor to investor.
This interview is incredibly information-dense, with too many "stop and think" moments. Here are some of the most thought-provoking insights.
***
## How to Track AI Development? Start by Spending $200
At the start, Patrick asked a very practical question: When a new model like Gemini 3 launches, how do you process this information?
Gavin's answer was direct: **you have to use it yourself.**
But the key isn't just "using it" — it's which version you use. He was surprised by investors who try the free version and conclude "AI is nothing special":
> The free version is like dealing with a 10-year-old, and then based on this 10-year-old's performance, predicting what they'll be like at 35. You can pay — actually, you have to pay to get the highest tier membership, $200 per month. Those are the real 30-35-year-old adults.
This analogy is spot-on. The gap between different AI model tiers follows the same logic — many people try a free model, conclude "AI is nothing special," but if you've used Claude 4.5 Opus, Gemini 3 Pro, or GPT-5.2 Reasoning — the top-tier models — the experience is completely different.
On the topic of spending, I personally spend about $300/month subscribing to various AI products, with the bulk going to Claude Code Max ($250/month). If you're a developer with significant coding needs, I strongly recommend subscribing directly to the official Claude Code ($125 minimum) rather than using mirror sites. You never know if mirror sites are actually using the real model, and the Max plan's cost-effectiveness is actually excellent — far more economical than pay-per-use.
As for information channels, Gavin's answer might surprise many: **X (Twitter)**.
He says AI development largely "happens in real-time on X." There are probably 500 to 1,000 people on Earth who truly understand the AI frontier, a significant portion in China, and you need to closely follow these people. He specifically mentions Andrej Karpathy:
> Every piece that Andrej Karpathy writes, you need to read it three times. Minimum.
As someone who also follows AI developments, I deeply resonate with this. AI discussions on Twitter are indeed more real-time and in-depth than any news outlet. Researchers from labs post directly about the latest developments, and even "argue" with each other — Gavin mentions that Meta's PyTorch team and Google's Jax team once had a public dispute on X, until both lab heads had to step in and declare: "Our people are not allowed to trash-talk the other lab."
***
## Scaling Laws: Our "Ancient Egyptian Moment"
After Gemini 3's release, many focused on what it revealed about scaling laws. Gavin offered a perspective I'd never heard before:
> Our understanding of pre-training scaling laws is probably like the ancient Egyptians' understanding of the sun. They could measure with extraordinary precision — the east-west axis of the Great Pyramid aligns perfectly with the equinoxes, as does Stonehenge. Perfect measurement. But they didn't understand orbital mechanics. They didn't know why the sun rises in the east and sets in the west.
This analogy made me pause for a long time. We can indeed predict very precisely: increase a model's compute by 10x and performance improves by a certain amount. But we don't know why. This isn't a "law" — it's an "empirical observation," one we measure with extraordinary precision but whose underlying principles we don't understand.
So why does Gemini 3 matter? Because it proved this "empirical observation" still holds. At a time when Blackwell chips were delayed and everyone worried whether "scaling laws have hit a wall," Gemini 3 gave a clear answer: **they haven't.**
But what's more interesting is what Gavin said next: without the emergence of reasoning models, AI development from 2024 to 2025 should have stagnated.
Why? Because after XAI managed to get 200,000 Hopper GPUs working in concert, the next step required waiting for Blackwell chips. You can't keep more than 200,000 Hopper GPUs "coherent" — simply put, working as a unified system. And Blackwell was delayed.
> Without reasoning models, from mid-2024 to now, there would have been zero progress in AI. Everything would have stalled. Can you imagine what that would have meant for markets? We'd be living in a completely different environment. Reasoning models in some ways saved AI, because they allowed progress without Blackwell.
This is a perspective I hadn't considered before: reasoning models (like o1) aren't just a new capability — they actually "saved" the entire AI industry's development trajectory.
***
## Chip Wars: Google Is "Sucking Out the Oxygen"
On the GPU vs TPU competition, Gavin said something that stuck with me:
> Google is currently the lowest-cost producer of tokens. What they've been doing, I would say, is "sucking the economic oxygen out of the AI ecosystem" — which is an extremely rational strategy for them.
As the low-cost producer, Google has been offering AI services at low prices (even at a loss), making life difficult for competitors. This is a classic tech industry play, but Gavin points out an interesting shift:
> AI is the first time in my career where being the "low-cost producer" actually matters in tech. Apple isn't worth trillions because they're the low-cost producer of phones. Microsoft isn't worth trillions because they're the low-cost producer of software. NVIDIA isn't worth trillions because they're the low-cost producer of AI accelerators. It's never mattered before.
But in the AI era, when power becomes the limiting factor, **tokens per watt** becomes crucial. If you can produce 3-5x more tokens per watt, that's 3-5x the revenue. The price of compute becomes irrelevant because the bottleneck is power.
This landscape is about to change. Blackwell chips are finally being deployed, and Gavin predicts the first Blackwell model will come from XAI:
> According to Jensen, nobody builds data centers faster than Elon. Jensen has said this publicly.
Once Blackwell and subsequent Ruben chips are deployed at scale, Google's advantage as the low-cost producer will evaporate. Will they still be willing to operate their AI business at -30% gross margins then? The math will completely change.
***
## Space Data Centers: Crazy, But Right from First Principles
When Patrick asked about "any crazy ideas not being discussed enough," Gavin brought up space data centers. At first I thought he was joking, but after hearing his analysis, I realized this might be the most visionary part of the entire interview.
> From a first-principles perspective, space data centers are superior to earth-based data centers in every dimension.
His argument:
**1. Energy**: In space, satellites are exposed to sunlight 24 hours a day, with solar radiation 6x stronger than on the ground. And because there's always sunlight, you don't need batteries — batteries are a significant portion of costs. So the lowest-cost energy source in the solar system is "space solar."
**2. Cooling**: On Earth, most data center costs and weight go to cooling. But in space? Cooling is free. Put the radiators on the satellite's shaded side, where temperatures approach absolute zero.
**3. Networking**: In data centers, racks connect via fiber optic — essentially lasers through cables. What's the only thing faster? Lasers through vacuum. So laser-connected satellites in space would actually have faster networking than terrestrial data centers.
**4. User experience**: Currently, when you ask AI a question, the signal travels from your phone to a cell tower, through fiber, to some data center, gets processed, and returns the same way. But if satellites can communicate directly with phones (Starlink has already demonstrated direct-to-phone capability), the entire chain becomes much shorter.
Of course, this requires Starship mass launches to become reality — probably 5-6 more years. But Gavin points to an interesting convergence: Tesla, SpaceX, and XAI are converging. XAI will be the "intelligence module" for Optimus robots, SpaceX will build data centers in space to provide compute for AI — these three companies are forming a flywheel of mutually reinforcing competitive advantages.
***
## SaaS's "Burning Platform"
If the previous sections made you excited about AI's future, this one might worry you about many existing companies.
Gavin states bluntly: **application SaaS companies are making the exact same mistake that brick-and-mortar retailers made when facing e-commerce.**
Brick-and-mortar retailers looked at Amazon and thought "e-commerce is a low-margin business — how could it be more efficient than us? Customers currently pay to come to our stores and carry products home themselves." They clearly saw customer demand but refused to invest because they didn't like e-commerce's margin structure. The result? Amazon's North American retail margins are now higher than many traditional retailers.
SaaS companies face the same situation now. Traditional software is written once and can be infinitely replicated and distributed, with gross margins reaching 80-90%. But AI is different — every use requires new computation, and good AI companies might only achieve 40% margins.
> If you want to build AI agents and you're unwilling to operate at less than 35% gross margins, you will never succeed. Because AI-native companies are operating at those margins. If you want to protect 80% gross margins, you are guaranteeing yourself failure in AI. Absolutely guaranteeing.
Gavin calls this a "life-or-death decision," and **except for Microsoft, almost everyone is failing.**
He references Nokia's famous "burning platform" memo: your platform is on fire. But there's actually a perfectly good new platform right next to you. Jump over, then go back and put out the fire on the old one. Now you have two platforms.
Salesforce, ServiceNow, HubSpot, GitLab, Atlassian — he believes all these companies can and should run this playbook: publicly disclose your AI revenue, publicly disclose your AI gross margins (low margins actually prove it's "real AI"), then point to your venture-backed competitors who are still losing money and say "I have something they don't: a business that generates cash flow."
***
## An Investor's Origin Story
Near the end, Patrick asked a more personal question: how would you explain what you do to a young person?
Gavin's answer starts with "investing is a pursuit of truth," but the really interesting part is his life story.
His original plan was: teach skiing in winter, guide rafting in summer, rock climb in the off-season, while dabbling in novel writing and wildlife photography. This was his "life plan" in college, and his parents were fully supportive.
But his parents made one small request: could he find a professional internship — just one, anything?
The only internship he could find was at a brokerage firm's private wealth management division. The job was simple: whenever the firm published a research report, he'd check which clients held that stock, then mail them the report.
Then he started reading those reports.
> I thought: "Oh my God, this is the most interesting thing I can imagine."
He understood investing as a "game of skill and luck," somewhat like poker. You can lose due to bad luck — say, a meteorite hits your company's headquarters — but most of the time, skill matters. And gaining an edge means having the deepest historical knowledge, combined with the most accurate understanding of the present world, to form a differentiated view of "what happens next."
That was day three of his internship. He went to a bookstore, bought Peter Lynch's book, and finished it in two days. Then he read Buffett, read *Market Wizards*, read Buffett's shareholder letters — twice. Then taught himself accounting. Back at school, he switched his major from English and History to History and Economics.
He also shared a formative experience working as a cleaner. While working at Alta ski resort, he cleaned hotel rooms. Once, while cleaning, he noticed a guest reading the same book he was reading. He said "that's a great book, I'm about the same place as you." The guest looked at him like an alien, then asked with even more shock: "You read books?"
> That experience permanently changed how I treat other people.
***
## Epilogue: Whatever AI Needs, It Gets
Near the interview's close, Gavin said something I found most fascinating:
> Over the past two years, whatever AI needed to keep developing, it got. Have you ever seen U.S. public opinion shift on any issue as fast as it shifted on nuclear power? It just happened. And it happened right when AI needed it to happen. Now we're hitting power constraints on Earth, and suddenly the discussion about space data centers appears. Every time something might slow AI down, everything accelerates instead.
This recalls Kevin Kelly's "technium" concept from *What Technology Wants*: technology as a whole seems to have its own will, wanting to become ever more powerful.
Maybe it's just coincidence. Maybe it's just many smart people solving problems. But the pattern Gavin observes — AI encounters an obstacle, and the obstacle somehow gets removed — is indeed worth pondering.
# Curated Insights
# Curated Insights
A collection of my reflections and summaries after watching tech expert videos and reading quality blogs.
Not just simple note-taking, but insights blended with my own understanding and practical experience.
## Content Sources
* Technical video breakdowns
* Curated blog articles
* Podcast/interview summaries
## Latest Content
### Thinking Like an Agent: The Tool Design Philosophy of the Claude Code Team
Anthropic engineer Thariq shares agent tool design lessons from building Claude Code — from three iterations of AskUserQuestion to three refactors of TodoWrite, from RAG to progressive disclosure. Every case points to the same core methodology: see the world like an Agent.
[Read more →](./claude-code-seeing-like-an-agent)
### 90% of College Students Use AI, But Nobody Knows the Rules
Four students from Princeton, Berkeley, and LSE discuss the real state of AI on campus — cheating, confusion, polarization, and the question nobody dares ask: what's the point of college anymore? Not a promotional piece, but genuine confusion and reflection.
[Read more →](./ai-on-campus-student-perspectives)
### From Chip Wars to Space Data Centers: The Next Decade of AI
From chip wars to space data centers, from the SaaS life-or-death dilemma to the essence of investing. Gavin Baker shares his deepest insights on the AI industry in the "Invest Like the Best" podcast: why you must use paid AI, why Scaling Laws are like "ancient Egyptians understanding the sun," and how reasoning models "saved" the entire AI industry's development trajectory.
[Read more →](./gavin-baker-ai-economics)
# Karpathy’s tweet exploded to 62k stars: What exactly did andrej-karpathy-skills do?
On January 27, 2026, Andrej Karpathy posted a very long tweet on X - an 11-section, approximately 1,400-word programming essay, recording the pitfalls he encountered during the transition from "80% handwriting + 20% agent" in November to "80% agent + 20% polishing" in December. The tweet eventually reached **7.69 million views, 39,000 likes, and 36,000 bookmarks**.
Three months later, a GitHub repository called `forrestchang/andrej-karpathy-skills` was launched, packaging Karpathy’s tweet into four installable rules. Within two weeks, it reached **62.7k stars and 5.5k forks**, becoming No. 1 on the GitHub weekly list in April 2026.
Warehouse ontology: **a Markdown file**.
***
## 1. What is Karpathy complaining about?
Karpathy's lengthy article essentially lists four LLM-coded "chronic conditions."
**The first disease: secretly making assumptions for you**
> "The most common type of mistake is that models make incorrect assumptions for you and then apply them without validating them. They don't manage their own confusion, they don't seek clarification, they don't show inconsistencies, they don't present trade-offs, they don't refute when it's time to refute, and they're a little too flattering."
This is a "collusive error". You say "Add a login for me", it doesn't ask what authentication to use, whether to remember the device, or how to manage the session - it just lays out a solution that it thinks is reasonable. By the time you finish reviewing and find that it is different from what you want, it will already have 500 lines written.
**The Second Disease: Over-Engineering**
> "They particularly like to overcomplicate their code and APIs, bloat abstraction layers, and don't clean up dead code. They will use 1000 lines of code to implement an inefficient, bloated, fragile structure, and you have to coax like a child and say, 'Well, why don't you just do this?', before they say, 'Of course!' and then immediately reduce it to 100 lines."
This is the most typical observer effect of LLM coding: while being "generously" given context, it will "generously" reward complexity\*\*. Strategy mode, Factory mode, dependency injection—all are given to you.
**The third disease: Change something you didn’t ask to change**
> "They sometimes modify or delete some comments and code because they don't like it or don't fully understand it - even if these changes have nothing to do with the current task."
You ask it to fix a bug, and it conveniently deletes the unfinished TODO comment next to it on the grounds that it "doesn't seem to be needed anymore."
**Disease 4: Even if you write the rules in CLAUDE.md, it will still break**
> "The above problem still exists even if I made some simple repair attempts in CLAUDE.md."
This is the most heartbreaking sentence in the whole tweet. Karpathy, a former OpenAI founding team member and Tesla AI Director, couldn't write CLAUDE.md that would keep Claude completely in line.
***
## 2. Solution to `andrej-karpathy-skills`
`forrestchang` Systematize the solutions to these four diseases into four principles and package them into a CLAUDE.md file.
| Principles | Corresponding diseases | Core actions |
| ------------------------- | ---------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| **Think Before Coding** | Secretly assuming | State your assumptions clearly, list multiple interpretations, stop and ask if you are confused, and refute when necessary |
| **Simplicity First** | Over-engineering | Only write the minimum code required; do not write speculative flexibility, error handling, abstraction |
| **Surgical Changes** | Unauthorized changes | Only touch what is necessary; do not refactor or change the style; only report other dead codes but do not delete them |
| **Goal-Driven Execution** | Method misalignment | Give verifiable success criteria + tests and let the model loop to pass by itself |
The fourth principle is a direct reference to another famous line in Karpathy's tweet - "Leverage":
This is the foothold of the entire project methodology: \*\*The first three principles prevent LLM from messing around; the fourth principle tells you how to truly leverage its strengths. \*\*
***
## 3. How to install and use
**Method A: As a Claude Code plug-in** (recommended to use this first)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
After installation, the Claude Code dialogue of all projects will automatically comply with these four principles. It takes effect globally and can be turned off at any time by `/plugin`.
**Method B: Manually copy CLAUDE.md**
Enter the repo → open `CLAUDE.md` → copy → paste into `CLAUDE.md` in the root directory of your project. Only effective for this project.
The repository also provides additional `CURSOR.md` and `.cursor/rules/` adaptations - a set of content covering mainstream AI IDEs.
***
## 4. Why can it reach 62k star?
This is a phenomenon worth unpacking. 62.7k stars is an exaggerated number for a "single file repo" - for comparison, Microsoft markitdown (9k) and Addy Osmani's agent-skills (4.6k) in the same period combined are not as many as it.
Broken down by impact weight:
**1. Karpathy IP Endorsement** - The same content cannot exceed 10,000 if it is called `forrestchang-skills`. Karpathy comes with the cultural capital of "former OpenAI founding team + Tesla AI Director + CS231n instructor", and his tweets come with a "must read" label.
**2. Perfect timing** — Opus 4.7 was released on April 16th, and over-engineering complaints were at their peak. The repo appeared just when everyone was looking for an antidote to "stop Claude from being so crazy".
**3. Pain points are universal** - Every Claude Code / Cursor user has stepped on these four pitfalls, and the empathy rate is close to 100%.
**4. The threshold is extremely low** - 1 file or 2 lines of commands. The cost of star is so low that it can be ignored. "If you don't install it, you will lose."
**5. Strong sense of verifiability** - The 4 principles are clear and easy to remember, and easy to screenshot and forward. Unlike the 1000 line prompt project guide which is prohibitive.
**6. Bilingual README** ——`README.zh.md` directly eats up the Chinese AI circle traffic, V2EX/instantly/detonates simultaneously on Weibo.
**7. Cross-promotion by the author** - The sentence in the top column *"Check out my new project Multica"* directs traffic to the author's own commercial agent platform `multica-ai/multica`. \*\*This repo is essentially the top of Multica’s customer acquisition funnel. \*\*
**8. Meta fit** - The "LLM coding errors" it discusses are exactly what all readers are experiencing when coding with LLM. Reading and using are integrated, and the conversion rate is extremely high.
In a word: What it sells is not code or tools, but **packaging Karpathy's emotions into installable rules** - this is the most typical "content is product" case in the AI programming circle in 2026.
***
## 5. My usage suggestions
**First use method A to install globally**. See if it improves your experience when writing tools and scripts - especially when you let Claude change other people's code, whether it reduces the problem of making random changes.
**We will decide after a week or two whether to merge it into the CLAUDE.md project**. Each project's CLAUDE.md is already filled with domain knowledge (design system, component specification, deployment process), while Karpathy's set is a general methodology. The two do not conflict and can be superimposed - but the timing should wait until you are really sure that it is useful.
**Pay attention to the cost**: It will make Claude ask more questions, which will be annoying to people who are used to "generating in one sentence"; it may not do the slight cleaning that should be done (too strict); it will be constrained by very vague exploration tasks.
**Deeper value**: It forces you to clearly state your requirements - which happens to be the prerequisite for all high-quality software engineering.
**Follow-up noteworthy**: forrestchang himself is also promoting Multica - an "open-source managed agents platform" to productize the skills mechanism. If this set of four principles eventually becomes the de facto standard, Multica will be its commercial vehicle. Pay attention to this line.
***
## Reference resources
# Indie Dev
Documenting the complete indie development process — every step from preparation to launch.
# Bark
# Bark
A tool that lets you send custom push notifications to your iPhone via a simple HTTP request. Free, open source, and supports self-hosting.
## Why I Recommend It
* **Minimal API** — A single curl command sends a push notification, no complex configuration required
* **Open source & free** — MIT-licensed, fully open source, zero cost
* **Privacy first** — Self-hosting support means your push data stays entirely in your own hands
## Use Cases
* **Script notifications** — Get notified when long-running tasks like data backups finish
* **Service monitoring** — Receive instant alerts when your server hits an exception
* **Automation hooks** — Notifications for CI/CD build results and scheduled task completions
## Quick Start
1. Download Bark from the [App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865)
2. Open the app and copy your push URL
3. Send your first notification:
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
Push with a title:
```bash
curl https://api.day.app/YOUR_KEY/Title/Content
```
## Project Info
* GitHub: [Finb/Bark](https://github.com/Finb/Bark)
* Stars: 7.2k+
* License: MIT
# Toolkit
# Toolkit
A collection of high-quality GitHub projects and useful software I've come across.
# The Complete Guide to Claude Agent Teams
## Introduction
If you've used Claude Code's Subagents, you might think parallel development is already powerful enough. But Subagents have a limitation: they can only report results back to the main Agent and cannot communicate with each other.
Agent Teams change this entirely. Imagine: one Agent handles security reviews, another focuses on performance optimization, and a third covers test coverage -- they can not only work in parallel but also talk directly to each other, challenge one another, and reach consensus. This is the core value of Agent Teams.
## Understanding Agent Teams
The architecture of Agent Teams resembles a real development team:
```
┌─────────────────────────────────────────────────────────┐
│ You (User) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (Main Claude instance, coordinates work) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│Security Audit│◄─►│ Performance │◄─►│Test Coverage │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└───��──────────┼──────────────┘
▼
┌─────────────────┐
│ Shared Task List │
└─────────────────┘
```
### Differences from Subagents
| Feature | Subagent | Agent Teams |
| --------------------- | --------------------------------------------------- | --------------------------------------------------- |
| **Context** | Independent context, results returned to main Agent | Independent context, fully independent operation |
| **Communication** | Can only report to main Agent | Teammates can communicate directly |
| **Task Coordination** | Main Agent manages all work | Shared task list, self-coordinated |
| **Use Cases** | Focused tasks where only results are needed | Complex work requiring discussion and collaboration |
| **Token Cost** | Lower: result summaries returned to main context | Higher: each Teammate is an independent instance |
In short: **Subagents are contractors you send out to execute tasks, while Agent Teams are a project team collaborating in the same room**.
### Why Agent Teams Work
The core insight: **specialization brings focus**.
When a single Agent handles complex multi-step tasks, the context keeps expanding, often requiring `/clear` to reset. Agent Teams let each Teammate maintain a narrow area of focus, keeping the context clean and performance more stable.
## Enabling Agent Teams
Agent Teams is currently an experimental feature, disabled by default. You need to enable it manually:
**Method 1: Environment Variable**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**Method 2: settings.json (Recommended, persistent)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## Core Usage
### Creating Your First Agent Team
Once enabled, simply tell Claude in natural language to create a team:
```
Create an agent team to review PR #142.
Spawn three reviewers:
- One focused on security issues
- One checking performance impact
- One verifying test coverage
Have them each review and report their findings.
```
**Keyword tip**: Use "create an agent team" or "spawn an agent team." If you just say "spawn agents," it may confuse Subagents with Agent Teams.
### Display Modes
Agent Teams supports two display modes:
| Mode | Description | Requirements |
| --------------- | -------------------------------------- | ----------------------- |
| **In-process** | All Teammates run in the main terminal | No special requirements |
| **Split panes** | Each Teammate gets its own pane | Requires tmux or iTerm2 |
The default is `auto`: uses split panes if running in tmux, otherwise in-process.
**Configure display mode**:
```json
{
"teammateMode": "in-process"
}
```
**Per-session override**:
```bash
claude --teammate-mode in-process
```
### Common Shortcuts
| Action | Shortcut |
| --------------------------- | ------------ |
| Switch between Teammates | `Shift+Down` |
| Return to previous Teammate | `Shift+Up` |
| Toggle task list display | `Ctrl+T` |
| Interrupt current Teammate | `Escape` |
| **Enable Delegate Mode** | `Shift+Tab` |
| Enter Teammate session | `Enter` |
### Delegate Mode (Important)
Delegate Mode is one of the most important features of Agent Teams:
| Mode | Lead Behavior |
| ----------------- | ------------------------------------------------------------ |
| **Normal Mode** | Lead may implement tasks and write code itself |
| **Delegate Mode** | Lead can only coordinate -- no writing code or running tests |
**Why you need Delegate Mode**:
Without this restriction, the Lead often "grabs work" -- even with three Teammates waiting, the Lead starts writing code itself. With Delegate Mode enabled, the Lead is forced into a pure project manager role, limited to managing tasks, communicating with Teammates, and reviewing output.
```
# Press Shift+Tab immediately after starting the team
```
## Practical Examples
### Example 1: Parallel Code Review
A single reviewer tends to go deep on one type of issue. Splitting review dimensions into independent domains ensures security, performance, and test coverage all receive equal attention:
```
Create an agent team to review this PR. Spawn three reviewers:
- Security reviewer: check authentication, authorization, injection vulnerabilities
- Performance reviewer: analyze algorithm complexity, database queries, caching strategies
- Test reviewer: verify test coverage, edge cases, error handling
Have them review independently, then discuss their findings with each other.
```
### Example 2: Competitive Hypothesis Debugging
When the root cause is unclear, a single Agent tends to stop at the first plausible explanation. Having Teammates challenge each other avoids this problem:
```
Users report the app exits after sending one message instead of staying connected.
Spawn 5 agent teammates to investigate different hypotheses. Have them discuss
with each other, actively trying to disprove each other's theories like a
scientific debate. Update the investigation report with consensus findings.
```
**Key mechanism**: The debate structure. Multiple independent investigators actively try to overturn each other's theories, and the hypothesis that survives is more likely to be the true root cause.
### Example 3: Content Batch Production
This is a typical non-technical application -- transforming one input into multiple outputs:
```
Create an agent team to convert this video script into content for four platforms:
- LinkedIn article writer
- Twitter thread writer
- Newsletter writer
- Blog post writer
Script location: /content/scripts/video-20.md
```
Each Teammate creates independently while maintaining content consistency.
### Example 4: QA Quality Check Cluster
A comprehensive quality check for a blog site, deploying 5 Agents to test different aspects in parallel:
```
Create an agent team for comprehensive blog quality checks:
- Agent 1: Core page testing (homepage, about page, contact page)
- Agent 2: Article page testing (rendering, navigation, SEO metadata)
- Agent 3: Link checking (internal links, external links, dead links)
- Agent 4: SEO validation (titles, descriptions, structured data)
- Agent 5: Accessibility testing (ARIA labels, contrast, keyboard navigation)
Generate a priority-sorted issue report.
```
**Result**: Completes in minutes what would otherwise require sequential manual execution, with each Agent focused on its own domain, culminating in a priority-sorted issue list.
### Example 5: Multi-Round Discussion Mode
A useful prompt pattern -- having Teammates discuss like a meeting:
```
Use Agent Teams to create 4 teammates to discuss [technical decision],
conducting 3 rounds of discussion. Have teammates exchange ideas in each round.
One teammate specifically takes the Red Team perspective, raising critiques.
```
This pattern is especially suitable for architecture decisions, technology selection, and other scenarios requiring multi-angle consideration.
### Example 6: C Compiler Project
Anthropic used 16 Agents to build a C compiler from scratch capable of compiling the Linux kernel:
| Metric | Data |
| ---------------- | ------------------------------------ |
| Number of Agents | 16 parallel instances |
| Sessions | \~2,000 Claude Code sessions |
| Cost | \~$20,000 |
| Lines of Code | 100,000 lines |
| Token Usage | 2 billion input + 140 million output |
Final output: A Rust compiler that can build a bootable Linux 6.9 on x86, ARM, and RISC-V.
## Team Management
### Specifying Teammates and Models
Claude automatically decides how many Teammates to spawn based on the task, but you can also specify explicitly:
```
Create 4 teammates to refactor these modules in parallel.
Have each teammate use the Sonnet model.
```
### Requiring Plan Approval
For complex or high-risk tasks, you can require Teammates to create a plan before execution:
```
Spawn an architect teammate to refactor the authentication module.
Require plan approval before they make any changes.
```
After completing the plan, the Teammate sends an approval request to the Lead. The Lead can then approve or return it with revision feedback.
### Talking Directly to Teammates
Each Teammate is a full Claude Code session. You can send messages directly to any Teammate:
* **In-process mode**: Switch with `Shift+Down`, then type your message
* **Split-pane mode**: Click directly on the corresponding pane
### Closing Teammates
```
Please close the security review teammate
```
The Lead sends a shutdown request, and the Teammate can approve or decline (with an explanation).
### Cleaning Up the Team
When finished, have the Lead clean up resources:
```
Clean up the team
```
**Important**: Always clean up through the Lead. Don't have Teammates perform cleanup, as this may cause inconsistent resource states.
## Best Practices
### Team Size Control
Size recommendations:
| Team Size | Use Case |
| --------- | ---------------------------------------------------- |
| 3 | Simple multi-perspective reviews |
| 4-5 | Standard feature development or refactoring |
| 6+ | Large-scale migrations or complex architecture tasks |
**Rule of thumb**: Assigning 5-6 tasks per Teammate works well. If you have 15 independent tasks, 3 Teammates is a good starting point.
### Task Granularity
* **Too small**: Coordination overhead exceeds benefits
* **Too large**: Teammates work too long without checkpoints, increasing waste risk
* **Just right**: Independent, self-contained work units with clear output (a function, a test file, a review report)
### Avoiding File Conflicts
Two Teammates editing the same file will cause overwrites. When splitting work, ensure each Teammate is responsible for a different set of files:
```
Teammate 1: responsible for src/auth/ directory
Teammate 2: responsible for src/api/ directory
Teammate 3: responsible for src/utils/ directory
```
### Monitoring and Guidance
Check Teammates' progress regularly and correct misaligned directions promptly. Letting the team run unsupervised for too long increases waste risk.
If the Lead starts implementing tasks instead of waiting for Teammates:
```
Wait for your teammates to complete their tasks before continuing
```
### Providing Sufficient Context
Teammates automatically load project context (CLAUDE.md, MCP servers, skills) but do not inherit the Lead's conversation history. Provide enough task details when spawning:
```
Spawn a security review teammate with this prompt:
"Review the authentication module in src/auth/ for security vulnerabilities.
Focus on token handling, session management, and input validation.
The app uses JWT tokens stored in httpOnly cookies.
Include severity ratings with your findings."
```
### Self-Reporting Verification Pattern
Include clear verification criteria in task descriptions to ensure Teammates self-check upon completion:
```
When the task is complete, report to the Lead:
1. Which files you checked
2. What issues you found
3. What changes you made
4. Whether verification criteria (tests pass, lint has no warnings, etc.) are met
```
This pattern reduces the Lead's verification workload while ensuring tasks are truly complete rather than just "appearing complete."
## Advanced Techniques
### Using Hooks to Enforce Quality Gates
Use Hooks to enforce rules when Teammates complete work:
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle`: Runs when a Teammate is about to become idle. Returning exit code 2 sends feedback and keeps the Teammate working
* `TaskCompleted`: Runs when a task is marked complete. Returning exit code 2 blocks completion and sends feedback
### Pre-Approving Permissions
Teammate permission requests bubble up to the Lead, which can cause frequent interruptions. Pre-approve common operations before spawning:
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### Combining with Worktrees
Agent Teams can work alongside Worktrees, with each Teammate operating in its own worktree:
```
Create an agent team where each teammate works in an independent worktree
to avoid file conflicts.
```
### Third-Party Orchestration Tools
Beyond native Agent Teams, the community has developed some orchestration tools:
| Tool | Description |
| --------------- | --------------------------------------------------- |
| **Gas Town** | Tool for managing multiple parallel Claude sessions |
| **Multiclaude** | Run Claude instances in multiple terminal windows |
These tools provide alternatives beyond the experimental Agent Teams feature but require more manual configuration. If native Agent Teams meets your needs, official functionality is recommended.
## Current Limitations
Agent Teams is still experimental, so understanding the limitations is important:
| Limitation | Description |
| ----------------------------------- | --------------------------------------------------------------------------------------------- |
| Cannot recover in-process teammates | `/resume` and `/rewind` won't restore in-process teammates |
| Task status may lag | Teammates sometimes forget to mark tasks as complete |
| Shutdown may be slow | Teammates finish their current request before shutting down |
| One team per session | The Lead can only manage one team at a time |
| No nested teams | Teammates cannot spawn their own teams |
| Fixed Lead | The session that creates the team is the Lead and cannot be transferred |
| Split panes require tmux/iTerm2 | VS Code terminal, Windows Terminal, and Ghostty are not supported |
| Plan mode is session-level | A Teammate's Plan mode state is locked at spawn time and cannot be changed during the session |
## Cost Considerations
Agent Teams consume significantly more tokens than single sessions:
| Scenario | Token Consumption | Cost Multiplier |
| ------------------------------ | ----------------- | ----------------- |
| Single Agent session | \~200k tokens | 1x |
| 3 Teammates | \~800k tokens | \~4x |
| 5 Teammates | \~1.2M tokens | \~6x |
| 16 Teammates (C compiler case) | 2 billion tokens | $20,000 / 2 weeks |
**Cost analysis**:
* Each Teammate is a fully independent Claude instance with its own context
* Communication between Teammates also consumes tokens
* The Lead needs to coordinate all Teammates, adding extra overhead
**When it's worth it**:
* Research tasks requiring parallel exploration
* Multi-perspective reviews (security, performance, testing)
* Decisions requiring discussion to reach consensus
* Not for routine tasks that can be done sequentially
* Not for parallel tasks that don't need inter-communication (Subagents are more cost-effective)
## My Usage Insights
### When to Use Agent Teams
My decision criteria:
1. **Tasks need multiple perspectives**: Domain expertise across areas (security + performance + testing)
2. **Discussion and consensus are needed**: Competitive hypotheses, architecture decisions
3. **Parallel exploration has value**: Comparing multiple implementation approaches
If you only need parallel execution without inter-communication, Subagents or Worktrees are more suitable.
### Start with Research and Reviews
If you're new to Agent Teams, start with tasks that don't require writing code: reviewing PRs, researching technical solutions, investigating bugs. These tasks have clear boundaries, demonstrate the value of parallel exploration, and avoid the coordination challenges of parallel implementation.
### Combining with Other Features
| Combination | Effect |
| ---------------------- | ---------------------------------------------- |
| Agent Teams + Worktree | Each Teammate works in an isolated environment |
| Agent Teams + Hooks | Automated quality checks and feedback |
| Agent Teams + Skills | Each Teammate gains specialized capabilities |
## Final Thoughts
Agent Teams represents a new paradigm in AI-assisted development: from "one AI assistant" to "an AI team."
Anthropic used 16 Agents over 2 weeks at a cost of $20,000 to produce a C compiler with 100,000 lines of code. The key lesson from this project: **test quality matters more than anything**. Agents will autonomously solve whatever problem you give them, so task validators must be near-perfect -- otherwise Agents will solve the wrong problem.
Remember three core takeaways:
| Takeaway | Description |
| ----------------- | ----------------------------------------------------------- |
| **Collaboration** | Teammates can communicate directly, not just report results |
| **Sharing** | Work is coordinated through a shared task list |
| **Supervision** | Check progress regularly and correct course promptly |
Getting started is simple:
```
Create an agent team to [your task]
```
***
**Related reading**:
* [The Complete Guide to Claude Worktree](/en/docs/notes/claude-worktree) -- Understanding how Worktrees work with Agent Teams
* [The Complete Guide to Claude Subagent](/en/docs/notes/claude-subagent) -- Comparing use cases for Subagents vs Agent Teams
* [Tmux Quick Start Guide](/en/docs/notes/tmux-tutorial) -- Using Tmux to manage multiple Agent sessions
**References**:
* [Claude Code Official Documentation - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**Video tutorials**:
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) -- 7 non-technical use case demos
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) -- Multi-agent workflow deep dive
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) -- Team orchestration in depth
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) -- Complete setup tutorial
# A Complete Guide to Claude's System Architecture
## Introduction
In September 2025, Anthropic closed a $13B funding round at a **$183B valuation**, making it the fourth-largest private company in the world. Its flagship product, Claude Code, has attracted **115,000** active developers since its February launch, processing **195 million lines** of code per week, with user growth of **300%**.
Even more fascinating, Anthropic CEO Dario Amodei revealed: **90% of Claude Code's code is written by itself**.
**How is that possible?**
How can an AI coding assistant "write itself"? What is unique about its architecture that enables it to assist — and even replace — human developers so effectively?
The answer lies in Claude's **modular architecture**: MCP provides tools, Skills teach how to use them, Subagents execute tasks in parallel, and Hooks ensure deterministic control — these components work together to give Claude the ability to "work like a programmer."
This article gives you a **bird's-eye view of the entire architecture** — the role of each component, how they work together, and quick-start configuration examples. Follow-up articles will dive deep into the details of each component.
## Architecture Overview
The Claude system uses a **modular architecture** where components are categorized by function and work as **complementary collaborators** rather than hierarchical dependencies:
**Key insight**: These components are **peer-level, complementary** extensions — not hierarchical dependencies. You can freely combine them based on your needs:
| You want to... | Use... | One-line description |
| --------------------------------------------- | ------------- | ---------------------------------------------------------------------------- |
| Connect to external data sources and services | **MCP** | Give Claude "hands and feet" to access databases, APIs, and file systems |
| Teach Claude specific workflows | **Skills** | Let Claude "know" how to do things in a particular domain |
| Process complex tasks in parallel | **Subagents** | Break large tasks into subtasks, with multiple Agents working simultaneously |
| Quickly trigger repetitive operations | **Commands** | One-click launch for common workflows, eliminating repetitive instructions |
| Ensure certain operations always execute | **Hooks** | Regardless of Claude's decisions, this step must run |
***
## Core Runtime
### Agent SDK — The Runtime Engine
Agent SDK is the **core runtime engine** of the entire Claude Agent system, providing:
* **Main Loop**: The Agent's core work loop
* **Context Management**: Token budgets, automatic compression (triggered at 92% usage)
* **Tool Dispatch**: Deciding which tool to use and how to execute it
* **Permission System**: Controlling tool access permissions
The Agent's core operating model is a simple **feedback loop**:
```
Gather context → Execute actions → Verify work → Repeat
```
***
### Built-in Tools — Core Toolset
The Claude Agent ships with 20+ core tools in three categories:
| Category | Tools | Description |
| ----------- | ------------------- | ---------------------------------------------- |
| **Read** | Read, Glob, Grep | File reading, pattern matching, content search |
| **Operate** | Write, Edit, Bash | File writing, editing, command execution |
| **Network** | WebSearch, WebFetch | Web search, webpage fetching |
These tools are **available by default** with no additional configuration. Claude interacts with the computer through these tools, just like a programmer uses an IDE.
***
## Configuration & Context
### CLAUDE.md — Persistent Context
Every new conversation used to require repeating the project background, coding standards, and architectural conventions... CLAUDE.md lets you **configure once, load automatically**.
CLAUDE.md is like a **README for AI** — it tells Claude about the project's background knowledge, workflow, and conventions.
#### Hierarchical Override
Claude loads CLAUDE.md files in the following order, with **more specific files taking higher priority**:
```
Enterprise (lowest)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md) (highest)
```
#### Content Recommendations
CLAUDE.md should include the following core information:
| Category | Example Content |
| --------------------- | ---------------------------------------------- |
| **Tech Stack** | Next.js 14 + TypeScript, Tailwind CSS |
| **Build Commands** | `npm run dev`, `npm run build`, `npm run test` |
| **Code Standards** | Naming conventions, Lint tool configuration |
| **Project Structure** | Purpose of key directories |
**Key principle**: Keep it concise. CLAUDE.md is **loaded with every conversation** — making it too long wastes precious tokens.
***
## Packaging & Distribution
### Plugins — Installable Units
Team configurations are scattered, hard to share, and difficult to standardize. Everyone has their own set of Skills, Commands, Hooks... How do you unify management?
Plugins bundle **Skills + Commands + Subagents + Hooks + MCP** into **installable units**, enabling one-click distribution and team standardization.
#### Directory Structure
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # Plugin manifest (required)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Hooks configuration
├── .mcp.json # MCP Server configuration
└── README.md # Documentation
```
#### Configuration Example
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "Frontend development toolkit",
"author": "Your Team"
}
```
```bash
# Installation methods
claude plugin install github:your-org/your-plugin # From GitHub
claude plugin install /path/to/plugin # From local
```
**Related Resources**
| Resource | Description |
| -------------------------------------------------------------------------------- | ---------------------------------------------------- |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Anthropic's official plugin repository |
| [wshobson/agents](https://github.com/wshobson/agents) | 24.3k stars, high-quality Agent template collection |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | Community plugin marketplace for discovering plugins |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | 19.3k stars, curated list of Claude Code resources |
***
## Extension Capabilities (Complementary Modules)
The Claude system's extension capabilities consist of multiple **complementary modules**, each with a distinct role, working in synergy:
| Module | Functional Role | Activation Method |
| ------------- | ---------------------------------------------------------- | ----------------------------- |
| **MCP** | Connect to external data and services (WHAT) | Available after configuration |
| **Skills** | Procedural knowledge — teach Claude how to do things (HOW) | Auto-matched |
| **Subagents** | Independent context, parallel task delegation | Explicit invocation |
| **Commands** | Repetitive workflows | Manual `/cmd` |
| **Hooks** | Deterministic control, event-driven | Auto-triggered |
***
### MCP — External Connectivity
#### Design Philosophy
In the traditional approach, each external data source requires **custom integration**, leading to an N-times-M integration nightmare. MCP provides a standardized protocol for **integrate once, use everywhere**.
MCP (Model Context Protocol) is designed as the **USB-C port for AI applications**:
| Feature | Description |
| --------------------- | ----------------------------------------------------------------- |
| **Open Standard** | Released November 2024, donated to Linux Foundation December 2025 |
| **Industry Adoption** | Adopted by OpenAI, Microsoft, Google, AWS, and others |
| **Ecosystem Scale** | 97M+ monthly SDK downloads, thousands of community servers |
#### Architecture Pattern
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, AI Tools)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (Protocol Clients)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (Expose Specific Capabilities)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**Use cases**: Connecting to databases, integrating third-party services (GitHub, Slack, Notion), accessing private APIs, real-time data stream processing.
#### Configuration Example
Create `.mcp.json` in the project root:
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — Manual Workflows
Slash Commands provide **manually triggered** repetitive workflows.
| Feature | Description |
| ------------------ | --------------------------------------------- |
| **Trigger Method** | Manually type `/command-name` |
| **Location** | `.claude/commands/` |
| **Purpose** | Repetitive workflows, standardized operations |
**Example**: Create `.claude/commands/review.md`
```markdown
Please perform a code review on the current changes, focusing on:
1. Code style and consistency
2. Potential performance issues
3. Security vulnerabilities
4. Test coverage
```
Then type `/review` to trigger it.
***
### Hooks — Deterministic Control
Hooks are the core of **deterministic control** — certain operations must execute and cannot rely on the LLM's judgment.
| Category | Event | Trigger Timing |
| ------------ | -------------------- | -------------------------------- |
| **Tool** | `PreToolUse` | Before tool execution |
| | `PostToolUse` | After successful tool execution |
| | `PostToolUseFailure` | After tool execution failure |
| | `PermissionRequest` | When permission is requested |
| **Session** | `SessionStart` | When a session starts |
| | `SessionEnd` | When a session ends |
| | `Stop` | When Claude completes a response |
| **Subagent** | `SubagentStart` | When a subagent starts |
| | `SubagentStop` | When a subagent stops |
| **Other** | `UserPromptSubmit` | After user submits a prompt |
| | `Notification` | Notification events |
| | `PreCompact` | Before context compression |
#### Configuration Example
Auto-format TypeScript files:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### In Practice: Autonomous Loop
**Ralph Wiggum** is an official Anthropic plugin that uses the Stop hook to implement an autonomous iteration loop:
```bash
/ralph-loop "Implement a TODO API with CRUD operations and tests" --max-iterations 20
```
**How it works**: The Stop hook intercepts Claude's exit, re-injects the original prompt, and continues iterating until the task is complete or the maximum iteration count is reached.
**Best suited for**: Tasks requiring multiple iterations (passing tests, code refactoring), and tasks with automated verification.
***
### Subagents — Task Delegation & Parallel Execution
#### Design Philosophy
A single Agent faces challenges: limited context window, inability to parallelize, and unclear responsibilities. Subagents adopt the **Orchestrator-Worker** architecture pattern to solve these problems:
```
Main Agent (Orchestrator)
├── Analyze user request
├── Formulate plan
├── Decompose tasks
└── Spawn specialized subagents
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - parallel execution)
│ Code │ Tests │ Docs │
└────────┴────────┴────────┘
↓
Aggregate results → Main agent synthesizes output
```
#### Core Features
| Feature | Description |
| --------------------------- | ----------------------------------------------------------- |
| **Context Isolation** | Each Subagent has independent context, preventing pollution |
| **Task Specialization** | Custom system prompts define dedicated roles |
| **Tool Permission Control** | Subagents can be restricted to specific tools |
| **Parallel Execution** | Multiple Subagents work simultaneously |
**Performance data**: Multi-agent systems outperform single agents by 90.2%, and parallelization can cut research time by 90% (token consumption is \~15x, but worth it for complex tasks).
#### Configuration Example
Create a Markdown file in `.claude/agents/`:
```markdown
---
name: Code Reviewer
description: A subagent specialized in code review
tools:
- Read
- Grep
- Glob
---
You are a senior code review expert. Focus on:
1. Code quality and maintainability
2. Potential bugs and edge cases
3. Performance optimization opportunities
4. Security vulnerabilities
```
***
### Skills — Procedural Knowledge
#### Design Philosophy
Skills are **reusable playbooks for AI** — modular knowledge packages that Claude can dynamically load on demand. The core design principle is **Progressive Disclosure**:
```
📚 Skills Playbook
│
├─ 📋 Table of Contents ──── [Metadata Layer] Preloaded at startup (~30-50 tokens)
│ name: "weekly-report"
│ description: "Generate standardized weekly reports"
│
├─ 📖 Main Content ────────── [Core Document Layer] Loaded when relevant (~hundreds to thousands of tokens)
│ # Weekly Report Generator
│ ## Instructions
│ Generate weekly reports following this structure...
│
└─ 📎 Appendix ────────────── [Reference Resource Layer] Loaded when needed
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills**: MCP gives Claude the ability to access tools (WHAT), while Skills teach Claude how to use those tools effectively (HOW).
#### Core Advantages
| Advantage | Description |
| ------------------- | ------------------------------------------------------------------------------- |
| **Token Efficient** | Metadata only takes 30-50 tokens; dozens of Skills can be active simultaneously |
| **Auto-activated** | Automatically matched based on task context, no manual trigger needed |
| **Composable** | Multiple Skills automatically work together |
| **Portable** | Consistent experience across Claude.ai, Claude Code, and the API |
#### Configuration Example
Create a directory in `.claude/skills/`:
```
my-skill/
├── SKILL.md # Core instructions (required)
├── scripts/ # Executable scripts (optional)
└── references/ # Reference materials (optional)
```
SKILL.md core structure:
```yaml
---
name: code-review # Skill name
description: Code review for quality and security # Brief description (used for auto-matching)
---
# Code Review Skill
## Instructions
[Detailed step-by-step instructions...]
## Output Format
[Output format requirements...]
```
**Key point**: The `description` in the frontmatter is used for auto-matching — keep it concise and accurate.
***
## Official References
**Design Philosophy**
| Resource | Description |
| --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------ |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Foundational article on Agent architecture |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Agent SDK engineering practices |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Skills design philosophy |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | MCP launch announcement |
**Official Documentation**
| Resource | Description |
| ----------------------------------------------------------------------------- | ------------------------------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Skills compared with other components |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | Guide to using CLAUDE.md |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Subagents official documentation |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Hooks official documentation |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | MCP official documentation |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | MCP protocol specification |
**In-Depth Analysis**
| Resource | Description |
| --------------------------------------------------------------------------------------------------------- | -------------------------------------- |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | Includes architecture diagrams |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | In-depth analysis of Skills internals |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Inside the building of Claude Code |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Technical comparison of Skills and MCP |
***
## Further Reading
If you want to dive deeper into Skills concepts and practices, check out:
* [What Are Claude Skills](/en/docs/notes/claude-skills/concept) — Detailed explanation of Skills core principles
* [Claude Skills Hands-on Guide](/en/docs/notes/claude-skills/practice) — Create your first Skill step by step
* [The Complete Guide to Claude Subagents](/en/docs/notes/claude-subagent) — Using and customizing subagents
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — An AI programming system based on context engineering
* [My Claude Code Best Practices](/en/blog/claude-code-best-practices) — Tips for everyday Claude Code usage
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# The Complete Guide to Claude Worktree
## Introduction
When working on complex tasks with Claude Code, you may have encountered this dilemma: you have three independent tasks to handle, but running multiple Claude instances in the same directory causes code conflicts — one Agent is modifying files while another is touching the same ones, resulting in a mess when merging.
In February 2025, Anthropic released the `--worktree` command, completely changing the game. Now you can run `claude -w feature-1`, `claude -w feature-2`, and `claude -w bugfix-1` in three separate terminals, with each Agent working in its own isolated environment without interfering with each other.
## Understanding Worktree
Imagine you're an architect designing three different rooms simultaneously. The traditional approach is drawing on the same blueprint, which gets messy with constant changes. Worktree gives you three separate blueprints, each dedicated to one room's design, to be merged back into the master plan later.
Technically speaking, Worktree is a native Git feature. Claude Code's `--worktree` command wraps this into a much simpler experience — one command creates an isolated environment, starts a Claude instance, and automatically cleans up when done.
### Why Not Just Clone Multiple Times
You might ask: why not just clone the repo multiple times?
| Approach | Disk Usage | Sync Difficulty | Cleanup Complexity |
| ------------------- | ------------------------------- | ------------------- | ------------------------- |
| Multiple Clones | Full repo each | Manual pull/push | Manual directory deletion |
| Git Worktree | Working files only, shared .git | Auto-shared history | `git worktree remove` |
| Claude `--worktree` | Working files only, shared .git | Auto-shared history | Auto-cleanup on exit |
Worktrees share the same `.git` database — all commit history and branch information is shared. This means a commit created in one worktree is immediately visible to all others.
### When to Use Worktree
Before diving in, consider whether your task is a good fit for worktrees.
Rule of thumb: **if a task takes more than 30 minutes, consider using a worktree**. Short tasks aren't worth the overhead — setting up the environment, installing dependencies, and merging might take longer than the task itself. But for tasks requiring deep work, worktree isolation becomes invaluable.
| Good for Worktree | Not Ideal |
| --------------------------------------------- | -------------------------------------------- |
| Independent feature development | Quick changes under 10 minutes |
| Parallel refactoring of different modules | Tasks requiring frequent interaction |
| Long-running tasks | Strong dependencies on other ongoing changes |
| Experimental changes needing isolated testing | Simple bug fixes |
### Prerequisites
Before using worktree, ensure these conditions are met:
| Condition | Description |
| ----------------------- | ----------------------------------------------------------------- |
| Git initialized | Must be in a Git repository (has `.git` directory) |
| At least one commit | Empty repositories cannot create worktrees |
| Remote branch available | Defaults to checking out from remote branch (e.g., `origin/main`) |
## Complete Workflow: From Creation to Cleanup
Below is the complete workflow from creating a Worktree to final cleanup, in actual development order.
### Step 1: Create a Worktree
#### Creating from Remote Default Branch
Use the `-w` or `--worktree` flag to start Claude:
```bash
# Create a worktree named "feature-auth" and start Claude
claude -w feature-auth
# Auto-generate a random name (e.g., "bright-running-fox")
claude -w
```
This command actually does four things:
1. Creates a new working directory at `/.claude/worktrees/feature-auth/`
2. Creates a new branch named `worktree-feature-auth`
3. Checks out code from the remote default branch (e.g., `origin/main` or `origin/master`) — **note: not your current branch**
4. Starts Claude Code in the new directory
All worktrees live under the `.claude/worktrees/` directory:
```
your-project/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← First worktree
│ ├── bugfix-123/ ← Second worktree
│ └── refactor-api/ ← Third worktree
├── src/
└── package.json
```
It's recommended to add this path to `.gitignore`:
```bash
# .gitignore
.claude/worktrees/
```
#### Creating from Current/Specific Branch
`-w` always checks out from the remote default branch and doesn't currently support specifying a base branch. If you want to create a worktree based on your current branch (or a specific branch), there are three approaches:
**Approach 1: Manual Git Creation**
```bash
# Create worktree based on current HEAD
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# Or based on a specific branch
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# Then start Claude in that directory
cd .claude/worktrees/my-feature && claude
```
This gives you full control — you can create worktrees based on any branch or commit, and the working directory starts on the correct branch from the beginning. The official docs also recommend: "For more control over branch and location, create the worktree with Git directly, then run Claude in that directory."
**Approach 2: Create in a Session (Recommended)**
In an existing Claude session, simply ask Claude to create a worktree:
```
> start a worktree from the current branch
> start a worktree
```
Unlike the `-w` command, worktrees created in a session are **automatically based on the current branch**, not the remote default branch. Claude handles the worktree creation and switches to it automatically — no manual Git commands needed. If you're already working on a feature branch, this is the most convenient approach — one sentence gets you an isolated environment based on your current branch.
**Approach 3: Create with `-w` First, Switch Branches in Session**
Create a worktree with `claude -w`, then ask Claude to switch to the target branch within the session. The downside is it pulls from the remote default branch first, then switches — an extra step, less clean than the first two approaches. And if the target branch is already occupied by another worktree, you'll hit a branch conflict:
As shown in the screenshot, Claude detects the branch conflict and offers two choices: go back to the main directory, or create a new working branch based on the target branch. While it works in the end, the process is less straightforward than Approaches 1 and 2.
**Advanced: Wrapping with Makefile for One-Click Commands**
If you frequently need to create worktrees from the current branch, you can add a shortcut command to the project root's `Makefile`, chaining creation + opening editor + starting Claude into a single deterministic pipeline:
```makefile
# Create worktree from current branch and start development environment
# Usage: make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "Usage: make worktree name="; \
echo "Example: make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ Creating worktree from $$(git branch --show-current): $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ Initializing environment"
cd .claude/worktrees/$(name) && npm install
@echo "→ Opening in Zed"
zed .claude/worktrees/$(name)
@echo "→ Starting Claude"
cd .claude/worktrees/$(name) && claude
```
Usage is very clean:
```bash
# Create worktree from current branch, open in Zed, start Claude
make worktree name=fix-login-bug
# Spin up multiple parallel tasks
make worktree name=feature-search
make worktree name=refactor-api
```
Compared to manually typing multiple Git/cd/claude commands, `make worktree name=xxx` is just one line, and the process is exactly the same every time — no forgetting a step or mistyping a path. Note that because the Makefile uses native `git worktree add`, Claude Code's `WorktreeCreate` Hook won't trigger (that Hook only fires with `claude -w` or in-session worktree creation). So environment initialization steps (installing dependencies, copying `.env`, etc.) need to be written directly in the Makefile, as shown in the `npm install` example above.
### Step 2: Initialize Environment
After creating a worktree, the first thing to do is initialize the development environment. Each new worktree is an independent directory — `node_modules`, virtual environments, `.env` files, etc. won't carry over automatically.
Claude Code provides `WorktreeCreate` Hooks to automate environment setup:
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
This ensures dependencies are automatically installed and environment files are copied every time a worktree is created. Common initialization steps:
| Project Type | Init Command |
| ------------ | -------------------------------------------------- |
| Node.js | `npm install` or `yarn` |
| Python | `pip install -r requirements.txt` or activate venv |
| Go | `go mod download` |
| General | Copy `.env` files, set environment variables |
If Hooks aren't configured, you can also run `/init` at the start of each worktree session to ensure Claude properly understands the current working directory's context and re-reads the project structure and CLAUDE.md configuration.
### Step 3: Commit and Merge
Once the environment is ready and development is complete, the next step is merging changes back to the target branch.
**Merging back to main**
The most common case — the worktree branched from `origin/main`, and changes should merge back to `main`. Simply tell Claude in the worktree session:
```
> Commit all changes, push to remote, then create a PR to main
```
Claude will automatically handle the full commit → push → `gh pr create` workflow.
**Merging back to a feature branch**
If you're developing on the `feature-x` branch and need to merge worktree changes back to `feature-x` instead of `main`:
```
> Commit and push changes, then create a PR targeting the feature-x branch
```
Claude will execute `gh pr create --base feature-x`, creating a PR directly to the feature branch.
You can also exit the worktree session (choosing to keep the worktree) and start Claude in the main directory:
```
> Merge the worktree-my-task branch changes into the current branch
```
If you only want specific commits from the worktree, you can selectively cherry-pick:
```
> Show me the commit history of the worktree-my-task branch, then cherry-pick the auth-related commits to the current branch
```
> **Tip**: All worktrees share the same `.git` database — commits created in a worktree are immediately visible in the main directory without any push/pull operations.
### Step 4: Exit and Cleanup
Once changes are merged, you can exit the worktree session.
When exiting a worktree session, Claude handles things automatically based on the state:
| State | Action |
| -------------------------- | --------------------------------------------- |
| **No changes** | Automatically deletes the worktree and branch |
| **Has changes or commits** | Prompts you to keep or remove |
Kept worktrees persist for you to continue working on later.
You can also configure a `WorktreeRemove` Hook to automate cleanup:
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**Manual Management Commands**
If you need to manually manage worktrees, use standard Git commands:
```bash
# List all worktrees
git worktree list
# Manually remove a worktree
git worktree remove .claude/worktrees/feature-auth
# Clean up stale worktree references
git worktree prune
```
> **Note**: Don't directly `rm -rf` a worktree directory. Use `git worktree remove` instead, or if you've already deleted it by mistake, run `git worktree prune` to clean up stale references.
## Parallel Development Patterns
Now that you've mastered the basic workflow, let's explore how to leverage worktrees for parallel development.
### Multi-Terminal Parallel
The most common usage is running simultaneously across multiple terminal tabs:
```bash
# Terminal 1: Working on user authentication
claude -w feature-auth
# Terminal 2: Fixing a payment bug
claude -w bugfix-payment
# Terminal 3: Refactoring the API module
claude -w refactor-api
```
Each Claude instance works in its own worktree without interfering with others. You can:
* Have one terminal with Claude developing a new feature
* Have another terminal with Claude fixing a bug
* Continue your own code review in a third terminal
### Competitive Implementation
An efficient approach is having multiple Agents independently implement the same feature:
```bash
# Three terminals running separately
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
Give them the same requirements and let each implement independently. Then compare the three solutions and merge the best one. This leverages LLM non-determinism — the same input can produce different outputs, and sometimes the second version is better.
UI design exploration is also great for this pattern. Say you want to redesign your app's interface but aren't sure which style works best:
```bash
# Have three Agents implement different styles
claude -w ui-minimal # Minimalist style
claude -w ui-colorful # Vibrant colors
claude -w ui-glassmorphism # Glassmorphism style
```
After completion, run all three dev servers simultaneously (on different ports), compare side by side, and merge your favorite into the main branch — far more efficient than the traditional "build one version, review, rebuild" cycle.
### Subagent Isolation
Worktrees aren't just for the main Claude instance — they work with Subagents too. Add `isolation: worktree` to your custom Subagent's frontmatter:
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
You can also tell Claude directly in conversation:
```
> use worktrees for your agents
```
When a Subagent is configured with worktree isolation:
```
Main Agent (main directory)
│
├── Launch Migration Agent 1 ──→ worktree-migration-1/
│ └── Processing src/auth/
│
├── Launch Migration Agent 2 ──→ worktree-migration-2/
│ └── Processing src/api/
│
└── Launch Migration Agent 3 ──→ worktree-migration-3/
└── Processing src/utils/
```
Each Subagent works independently in its own worktree without interference. When finished, the worktree is automatically cleaned up (if there are no uncommitted changes).
### Tmux and IDE Integration
The `--tmux` flag automatically starts Claude in a new Tmux session, so it keeps running even if you close the terminal:
```bash
claude -w feature-auth --tmux
```
If you use VS Code or Cursor, the source control panel automatically recognizes all worktrees — the main repo shows as one repo, each worktree as an independent repo, and you can switch, commit, and push directly from the IDE. Worktrees also pair well with [Ralph loops](/en/docs/notes/ralph-wiggum/concept) — each Ralph loop runs in its own worktree, so even a failed loop won't affect the main branch.
## Tips & Best Practices
### Common Pitfalls
1. **Branch source confusion**: `-w` creates worktrees from the **remote default branch**, not your current branch. If you run `claude -w my-task` while on the `feature-x` branch, the new worktree's code comes from `origin/main` and won't include `feature-x` changes. To work from your current branch, see [Creating from Current/Specific Branch](#creating-from-currentspecific-branch).
2. **Uncommitted changes won't carry over**: When creating a worktree, unstaged or uncommitted changes from the main directory won't appear in the new worktree. Worktrees are created based on commit history only, so make sure to commit important changes first.
3. **Same branch can't be used by multiple worktrees**: Git doesn't allow two worktrees to check out the same branch simultaneously. If your main directory is on `feature-x`, attempting to also check out `feature-x` in a worktree will fail. Each worktree must be on a different branch.
4. **Environment needs reinitialization**: Each new worktree doesn't include runtime dependencies like `node_modules`. Configure a `WorktreeCreate` Hook for automation (see [Step 2: Initialize Environment](#step-2-initialize-environment)).
### Usage Tips
Don't overdo it. While you can technically open many worktrees, each Claude instance consumes API credits, too many parallel tasks become hard to track, and merge conflicts get more complex.
**Naming conventions**: Develop good naming habits for easier management:
```bash
# Good naming
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# Bad naming
claude -w test
claude -w temp
claude -w 1
```
## Non-Git Version Control
If you use SVN, Perforce, or Mercurial, you can achieve similar isolation by configuring `WorktreeCreate` and `WorktreeRemove` Hooks. Once configured, `--worktree` will invoke your custom commands instead of the default Git behavior.
## Final Thoughts
Worktree is a feature the Claude Code team uses every day — Boris Cherny calls it "the #1 productivity tip." The core value is simple: **enabling multiple Agents to work in parallel without interfering with each other**.
Getting started is easy:
```bash
claude -w your-task-name
```
***
**Related Reading**:
* [The Complete Guide to Claude Subagent](/en/docs/notes/claude-subagent) — Understanding Subagent and Worktree integration
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Another approach to boosting AI programming efficiency
* [Claude System Architecture Explained](/en/docs/notes/claude-architecture) — Understanding Worktree's place in the overall architecture
**References**:
* [Claude Code Official Docs - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Boris Cherny's Worktree Announcement](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Git Worktree Official Docs](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**Video Tutorials**:
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — Full worktree workflow demo
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — Manual worktree creation methods
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — Native worktree feature walkthrough
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — Quick start tutorial with Subagent usage
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — Benefits of multi-worktree parallel development
# Docs
Welcome to the documentation center. Here you'll find technical docs and tutorials I've put together on Claude Code.
# Tmux Quick Start Guide
## Introduction
If you've used Claude Code's Agent Teams or want to run multiple Claude instances simultaneously, Tmux is practically essential. It lets you run multiple sessions within a single terminal window, keeps sessions alive in the background even after you close the terminal, and allows Claude to automatically spawn and manage multiple Agents inside Tmux.
This tutorial is designed for Claude Code users, covering both Tmux fundamentals and key integration tips with Claude Code.
## Understanding Tmux
The three core concepts of Tmux:
```
┌─────────────────────────────────────────────────────────┐
│ Session │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| Concept | Analogy | Description |
| ----------- | ------------ | ---------------------------------------------------------- |
| **Session** | Workspace | Top-level container that persists even after disconnection |
| **Window** | Browser tab | A session can contain multiple windows |
| **Pane** | Split screen | A window can be divided into multiple panes |
## Installation & Basics
### Installing Tmux
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
Verify the installation:
```bash
tmux -V
# Output example: tmux 3.6a
```
### The Prefix Key
All Tmux commands start with the **prefix key**, which defaults to `Ctrl+B`.
How to enter a command:
1. Press `Ctrl+B` (hold it down)
2. Release, then press the command key
For example, to split a window: `Ctrl+B` then press `%`
## Command Cheat Sheet
### Session Management
| Command | Description |
| --------------------------- | ------------------------------------------- |
| `tmux` | Create a new session |
| `tmux new -s name` | Create a named session |
| `tmux ls` | List all sessions |
| `tmux attach -t name` | Attach to a session |
| `tmux kill-session -t name` | Kill a session |
| `Ctrl+B d` | Detach current session (runs in background) |
### Window Management
| Shortcut | Description |
| ------------ | --------------------------- |
| `Ctrl+B c` | Create a new window |
| `Ctrl+B n` | Next window |
| `Ctrl+B p` | Previous window |
| `Ctrl+B 0-9` | Switch to a specific window |
| `Ctrl+B ,` | Rename current window |
| `Ctrl+B &` | Close current window |
### Pane Management
| Shortcut | Description |
| ------------------- | ----------------------------- |
| `Ctrl+B %` | Vertical split (left/right) |
| `Ctrl+B "` | Horizontal split (top/bottom) |
| `Ctrl+B Arrow keys` | Move between panes |
| `Ctrl+B x` | Close current pane |
| `Ctrl+B z` | Maximize/restore pane |
| `Ctrl+B {` | Move pane left |
| `Ctrl+B }` | Move pane right |
### Other Useful Shortcuts
| Shortcut | Description |
| ---------- | ---------------------------- |
| `Ctrl+B [` | Enter copy mode (scrollable) |
| `q` | Exit copy mode |
| `Ctrl+B ?` | Show all shortcuts |
## Integration with Claude Code
### Why Claude Code Needs Tmux
1. **Agent Teams split-pane mode**: Each Teammate is displayed in its own pane
2. **Background execution**: Tasks continue running even after closing the terminal
3. **Session persistence**: Full context is restored after reconnecting
4. **Multi-instance management**: Run multiple Claude sessions simultaneously
### Basic Usage: Running Claude in the Background
```bash
# Start Claude in tmux
tmux new -s claude-work
claude
# Detach the session (Claude keeps running)
# Ctrl+B d
# Reconnect later
tmux attach -t claude-work
```
### Using the --tmux Flag
Claude Code natively supports Tmux integration:
```bash
# Start Claude in a new tmux session
claude --tmux
# Use with worktree
claude -w feature-auth --tmux
```
This automatically:
1. Creates a new tmux session
2. Starts Claude Code inside it
3. Names the session `claude-{random-ID}`
### Agent Teams Tmux Mode
Agent Teams can use split-pane display mode, with each Teammate running in its own pane:
```json
// settings.json
{
"teammateMode": "tmux"
}
```
Or via the command line:
```bash
claude --teammate-mode tmux
```
Result:
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## Practical Configuration
### Recommended \~/.tmux.conf
Create or edit `~/.tmux.conf`:
```bash
# Use Ctrl+A as prefix key (easier to reach)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# Enable mouse support
set -g mouse on
# Increase history buffer (Claude generates lots of output)
set -g history-limit 50000
# Vim-style pane navigation
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# More intuitive split shortcuts
bind | split-window -h
bind - split-window -v
# Quick config reload
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# 256 color support
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# Start window numbering at 1 (0 is too far away)
set -g base-index 1
setw -g pane-base-index 1
# Status bar optimization
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
Reload configuration:
```bash
tmux source-file ~/.tmux.conf
```
### Claude Code-Specific Configuration
Optimized configuration for Claude Code:
```bash
# Claude session popup shortcut
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
What this configuration does:
1. Press `Ctrl+A y` to open a Claude popup
2. Each directory gets its own Claude session
3. The session keeps running after closing the popup
4. Reopening restores the previous conversation
## Common Workflows
### Workflow 1: Parallel Multi-Project
```bash
# Create a separate session for each project
tmux new -s project-a
# Start Claude inside it
claude -w feature-x
# Detach, then create another session
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# Switch between sessions
tmux switch -t project-a
tmux switch -t project-b
# Or list all sessions to choose from
# Ctrl+B s
```
### Workflow 2: Development Dashboard
Create a multi-pane development environment:
```bash
# Create session
tmux new -s dev
# Split into three panes
# Ctrl+B % (vertical split)
# Ctrl+B " (horizontal split on the right side)
# Pane layout:
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# Run Claude in the first pane
claude
# Switch to the second pane (Ctrl+B Right Arrow)
tail -f logs/app.log
# Switch to the third pane
npm test -- --watch
```
### Workflow 3: Remote Development
Tmux's most powerful feature is session persistence, making it ideal for SSH remote development:
```bash
# Connect to the remote server
ssh user@server
# Create a tmux session
tmux new -s remote-claude
# Start Claude
claude
# Disconnect from SSH (Claude keeps running)
# Ctrl+B d
exit
# Reconnect later
ssh user@server
tmux attach -t remote-claude
# Claude session is fully restored
```
### Workflow 4: Agent Teams Monitoring
Use tmux to monitor all Teammates in Agent Teams:
```bash
# Start Claude with tmux mode
claude --teammate-mode tmux
# Create an Agent Team
# "Create an agent team to review code..."
# The screen automatically splits, one pane per Teammate
# Click on different panes to interact with the corresponding Teammate
```
## Troubleshooting
### Common Issues
| Issue | Solution |
| ------------------------------- | ------------------------------------------ |
| Colors not displaying correctly | Ensure `TERM=xterm-256color` |
| Mouse not working | Add `set -g mouse on` to config |
| Copy/paste issues | Use `Enter` to copy in copy mode |
| Session disappeared | Check `tmux ls`; system may have restarted |
### Cleaning Up Orphaned Sessions
Claude Code sometimes leaves behind uncleaned tmux sessions:
```bash
# List all sessions
tmux ls
# Kill a specific session
tmux kill-session -t session-name
# Kill all sessions (use with caution!)
tmux kill-server
```
### iTerm2 Users
If you're using iTerm2 on macOS, you can use its native integration:
```bash
# Use iTerm2's tmux integration mode
tmux -CC
# Or with Claude Code
claude --teammate-mode tmux
```
iTerm2 automatically converts tmux panes into native tabs and split views.
## My Personal Tips
### When to Use Tmux
| Scenario | Tmux Needed? |
| ---------------------------------- | ------------------ |
| Simple one-off Claude conversation | Not needed |
| Long-running tasks | Yes |
| Agent Teams | Highly recommended |
| Remote development | Essential |
| Parallel multi-project work | Recommended |
### Minimal Setup
If you don't want to fuss with configuration, just remember these commands:
```bash
# Create a session
tmux new -s work
# Detach (runs in background)
Ctrl+B d
# Reattach
tmux attach -t work
# Split panes
Ctrl+B % # Left/right split
Ctrl+B " # Top/bottom split
# Switch panes
Ctrl+B Arrow keys
```
### Best Combinations with Claude Code
1. **Worktree + Tmux**: Each worktree in its own tmux session
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux**: Visual management of all Teammates
```bash
claude --teammate-mode tmux
```
3. **Long tasks + Detach**: Start a task, detach, come back later to check
```bash
# Start
tmux new -s migration
claude
# "Run the database migration..."
# Ctrl+B d
# Hours later
tmux attach -t migration
```
## Final Thoughts
Tmux is a key tool for using Claude Code efficiently, especially in these scenarios:
| Takeaway | Description |
| ----------------- | ----------------------------------------------- |
| **Persistence** | Sessions survive disconnections |
| **Parallelism** | Manage multiple Claude instances simultaneously |
| **Visualization** | Split-pane display for Agent Teams |
Three core commands to get started:
* `tmux new -s name` to create a session
* `Ctrl+B d` to detach a session
* `tmux attach -t name` to reconnect
***
**Related Reading**:
* [Claude Agent Teams Complete Guide](/en/docs/notes/claude-agent-teams) — Agent Teams requires Tmux for split-pane mode
* [Claude Worktree Complete Guide](/en/docs/notes/claude-worktree) — Worktrees can run in the background with Tmux
**References**:
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**Video Tutorials**:
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) — Tmux basics for beginners
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) — Integration workflow with Claude Code
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) — Advanced configuration tips
# MVP Sprint: Shipping Core Features in Two Weeks
This is a test article.
# Register for Apple Developer Program
To publish your app on the App Store, the first step is to register for the Apple Developer Program. At ¥688 ($99) per year, this is an unavoidable investment for every iOS indie developer.
This article walks you through the differences between account types, what you need to prepare before registering, and the complete registration process.
## Account Type Comparison
The Apple Developer Program offers three account types, each suited for different development scenarios:
| Feature | Individual | Organization | Enterprise |
| ---------------------- | -------------------------------- | ------------------------------------------- | ------------------------------ |
| Annual Fee | ¥688 ($99) | ¥688 ($99) | ¥1,988 ($299) |
| App Store Distribution | ✅ | ✅ | ❌ (Internal distribution only) |
| Developer Name Display | Personal name | Organization/company name | Organization name |
| Team Member Management | ❌ | ✅ | ✅ |
| D-U-N-S Number | Not required | Required | Required |
| Review Timeline | Faster (usually within 48 hours) | Slower (requires organization verification) | Slower |
| Best For | Indie developers, individuals | Companies, studios | Large enterprise internal apps |
**The indie developer's choice**: If you're an individual developer, go with the **Individual account**. It has the simplest process, the fastest review, and all the features you need. The developer name displayed on the App Store will be your real name.
## Pre-Registration Preparation
### Prerequisites
Before starting the registration, make sure you have the following ready:
* **Apple ID**: If you don't have one yet, create one at [appleid.apple.com](https://appleid.apple.com). It's recommended to use your primary email, as all development-related notifications will be sent to this address.
* **Two-Factor Authentication**: Your Apple ID must have Two-Factor Authentication enabled. On your iPhone, go to Settings → Apple ID → Sign-In & Security → Two-Factor Authentication to enable it.
* **Apple Device**: The registration process requires identity verification on an iPhone or iPad, where you'll need to download the Apple Developer App.
### Additional Requirements for Organization Accounts
If you're registering an organization account, you'll also need:
* **D-U-N-S Number**: Apply in advance on the Dun & Bradstreet website — the review takes 5-14 business days.
* **Legal Entity Status**: The registrant must be the legal representative or an authorized representative of the organization.
* **Organization Information**: Including registered address, legal representative name, contact details, etc.
## Development Equipment Preparation
Registering a developer account is just the first step. iOS development also requires certain hardware and software tools.
### Required Equipment
* **Mac Computer** — Xcode only runs on macOS, so this is a hard requirement. Apple Silicon (M-series chip) Macs are recommended for faster compilation speeds and native iOS simulator support. A MacBook Air with an M-series chip is sufficient for indie development; if budget is limited, consider a Mac mini.
* **iPhone / iPad (Recommended but not required)** — The simulator covers most debugging scenarios, but real device testing is irreplaceable for performance, sensors (camera/GPS/NFC), push notifications, and more. Even without a paid developer account, you can debug on a real device using a free Apple ID (though with limitations like 7-day re-signing — see the FAQ section below).
### Development Tools
* **Xcode** — Apple's official IDE, free to download from the Mac App Store. It's quite large (around 12GB+), so be patient during the initial installation.
* **Apple Developer App** — Used for account registration, viewing WWDC videos, and documentation.
* **TestFlight** — Apple's beta distribution tool for inviting users to test your app.
### Things to Keep in Mind
* macOS and Xcode versions should be kept up to date. Apple releases a new version of Xcode after WWDC each year, typically requiring one of the two most recent macOS major versions.
* Xcode updates are frequent and large — make sure to reserve enough disk space (at least 50GB or more).
* If your app involves hardware features (camera, Bluetooth, NFC, etc.), real device testing is essential.
* If you don't have a Mac, cloud Mac services (such as MacStadium, AWS EC2 Mac) are an alternative, though the experience isn't as good as a native device.
## Registration Process
### Step 1: Download the Apple Developer App
Open the App Store on your iPhone or iPad, search for "Apple Developer," and download it.
### Step 2: Sign In and Start Registration
Open the Apple Developer App and sign in with your Apple ID. Tap the "Account" tab, then tap "Enroll in Apple Developer Program."
### Step 3: Fill in Information and Verify Identity
Follow the prompts to fill in your personal information:
1. **Confirm Identity Information**: Basic information such as name and address.
2. **Identity Verification**: Depending on your region, the app may ask you to photograph a government-issued ID (passport, driver's license, etc.) or take a selfie for identity verification.
3. **Agree to Terms**: Read and agree to the Apple Developer Program License Agreement.
> The identity verification step should be done in a well-lit environment to ensure clear photos. The entire registration process must be completed on the same device.
### Step 4: Pay the Annual Fee
After confirming your information is correct, pay the annual fee of ¥688 ($99). Payment methods linked to your Apple ID are supported. You'll receive a confirmation email after successful payment.
### Step 5: Wait for Review
* **Individual Account**: Usually reviewed within 48 hours. I paid on March 14 and received the welcome email on the morning of March 15 — less than 24 hours in total.
* **Organization Account**: Apple will verify the organization's information and D-U-N-S number, which may take longer.
Once approved, you can log in to the developer dashboard at [developer.apple.com](https://developer.apple.com) and access all development resources.
## Subscription Management and Renewal
The Apple Developer Program is an annual subscription, auto-renewing at ¥688 per year.
### Auto-Renewal
Auto-renewal is enabled by default, and the fee will be charged from the payment method linked to your Apple ID before the expiration date. It's recommended to keep auto-renewal on to avoid account expiration affecting your published apps (see FAQ below for details).
### Cancel or Manage Subscription
If you need to modify your renewal settings, open iPhone Settings → Apple ID → Subscriptions, find Apple Developer Program, and manage it there.
## Frequently Asked Questions
**Q: What happens if I forget to renew?**
After your account expires, your apps will be removed from the App Store, but they won't be deleted. Once you pay again, your apps will be restored. However, user downloads and updates will be affected during the lapse, so it's recommended to keep auto-renewal enabled.
**Q: How do I apply for a D-U-N-S Number?**
Visit the link below, fill in your company information, and submit the application. The review typically takes 5-14 business days. Individual accounts do not need this number.
**Q: How do I contact Apple if I run into issues?**
Visit [Apple Developer Support](https://developer.apple.com/contact/), where you can reach out via online chat or phone. The China region supports Chinese-language service, and response times are quite reasonable.
**Q: Can I start developing before registering?**
You can start exploring, but be aware of the free account limitations. With a free Apple ID, you can write code in Xcode, debug using the simulator, and even install apps on your own device. This is perfectly adequate for learning Swift and validating basic UI ideas.
However, free accounts have notable limitations: apps installed on real devices need to be recompiled and reinstalled every 7 days, you're limited to 3 devices per platform, and features like push notifications, iCloud, TestFlight, and In-App Purchases are unavailable. If your app needs these capabilities, you'll need a paid membership during development — not just for publishing to the App Store.
Recommendation: If you're just getting started learning Swift and running demos, a free account works fine. Once you begin working on a real project, register for the paid membership as early as possible to avoid delays caused by feature restrictions.
# Idea Validation: From Fuzzy Inspiration to an Actionable Direction
This is a test article.
# Tech Stack: Why I Chose Next.js + Supabase
This is a test article.
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# Concept Introduction
## Introduction
When you've set up a perfect workflow in Claude Code — custom commands, code review hooks, dedicated Skills — you might wonder: can I package all of this up and share it with my team or the community?
That's exactly the problem Plugin solves.
If Skills are "instruction manuals" for AI, then a Plugin is a "toolbox" — it bundles Skills, Commands, Hooks, MCP servers, and all other configurations together so you can install and distribute them with a single command.
## Understanding Plugin
Imagine you're an experienced craftsperson who has accumulated a trusted set of tools over the years: hammers, saws, rulers, various screwdrivers. Every time you switch workbenches, you have to carry each tool one by one and reorganize everything. A Plugin is like a well-designed toolbox — it not only holds all your tools but also keeps them neatly organized by category, ready to go wherever you take it.
From a technical perspective, Plugin is the extension packaging mechanism for Claude Code. A Plugin can contain:
| Component | Purpose | File Location |
| ------------------ | ---------------------------------- | ------------- |
| **Slash Commands** | Quick action entry points | `commands/` |
| **Subagents** | Specialized sub-agents | `agents/` |
| **Skills** | AI knowledge packages | `skills/` |
| **Hooks** | Event-triggered automation scripts | `hooks/` |
| **MCP Servers** | External system connections | `.mcp.json` |
| **LSP Servers** | Language server configuration | `.lsp.json` |
These components work together to form a complete workflow solution.
## Plugin vs Standalone Configuration
In Claude Code, you can place configurations in the project's `.claude/` directory or package them as a Plugin. The core difference lies in **distribution method** and **namespace**:
| Aspect | Standalone Config (`.claude/`) | Plugin |
| ------------------ | ------------------------------------------- | ------------------------------------ |
| Command Name | `/hello` | `/plugin-name:hello` |
| Use Case | Personal workflows, project-specific config | Team sharing, community distribution |
| Version Management | Managed with project code | Supports Semantic Versioning |
| Update Method | Manual sync | Supports auto-updates |
| Conflict Handling | May conflict with other configs | Namespace isolation |
**When to choose Plugin**:
* Need to share workflow configurations with team members
* Want to reuse the same toolset across multiple projects
* Planning to distribute configurations to the community
* Need version control and automatic updates
**When to use standalone configuration**:
* Quick personal experiments
* Project-specific configs that don't need reuse
* Simple one-off commands
## Plugin Directory Structure
A standard Plugin structure looks like this:
```
my-plugin/
├── .claude-plugin/ # Metadata directory
│ └── plugin.json # Required: plugin manifest
├── commands/ # Slash commands
│ ├── review.md
│ └── deploy.md
├── agents/ # Sub-agents
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # Event hooks
│ └── hooks.json
├── scripts/ # Helper scripts
│ └── format-code.sh
├── .mcp.json # MCP server configuration
└── .lsp.json # LSP server configuration
```
**Key notes**:
* `plugin.json` must be placed inside the `.claude-plugin/` directory
* Other directories (commands, agents, skills, etc.) go in the plugin root
* Do not put feature directories inside `.claude-plugin/`
## Core Configuration File
The core of a Plugin is `.claude-plugin/plugin.json`, which defines the plugin's metadata and component paths:
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "An example plugin",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| Field | Required | Description |
| ------------- | -------- | ----------------------------------------------------------- |
| `name` | Yes | Unique plugin identifier, use lowercase letters and hyphens |
| `version` | No | Semantic version number |
| `description` | No | Short plugin description |
| `author` | No | Author information |
| `keywords` | No | Tags for discoverability |
| `commands` | No | Command file or directory path |
| `agents` | No | Agent file or directory path |
| `skills` | No | Skills directory path |
| `hooks` | No | Hook configuration path |
| `mcpServers` | No | MCP configuration path |
## Installation Scopes
Plugin supports four installation scopes to accommodate different use cases:
| Scope | Config File | Purpose |
| --------- | ----------------------------- | ----------------------------------------------- |
| `user` | `~/.claude/settings.json` | Personal plugins, available across all projects |
| `project` | `.claude/settings.json` | Team plugins, shared via version control |
| `local` | `.claude/settings.local.json` | Project-specific, gitignored |
| `managed` | `managed-settings.json` | Enterprise-managed (read-only) |
The default installation scope is `user`. If you want to commit plugin configurations to Git for team use, choose the `project` scope.
## Core Advantages
### Namespace Isolation
Plugin commands carry a namespace prefix (e.g., `/my-plugin:review`), avoiding naming conflicts with other plugins or project configurations. This is especially important in team collaboration — plugins developed by different teams can coexist peacefully.
### Version Management
Plugin supports Semantic Versioning, allowing you to:
* Track the plugin's change history
* Roll back to older versions when needed
* Automatically receive compatible updates
### Easy Distribution
Through the Plugin Marketplace, you can:
* Host plugins on GitHub
* Let users install with a simple command
* Automatically handle dependencies and updates
### Team Collaboration
Plugin is particularly well-suited for team scenarios:
* Unify the team's development toolchain
* New members get all tools with one command
* Centralized configuration management reduces duplicate work
## Plugin Ecosystem
The Claude Code Plugin ecosystem is growing rapidly. As of early 2025, the ecosystem has reached a considerable scale:
* **229+ plugins** are active in the ecosystem
* **239 Agent Skills** distributed across the marketplace
* **200+ MCP servers** pre-built in the Docker toolkit
**Official resources**:
| Resource | Link | Description |
| ---------------------------------- | --------------------------------------------------------------------------------------- | -------------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Official Skills repository |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Official plugin directory |
| Docker MCP Toolkit | [Website](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ pre-built MCP servers |
**Community picks**:
| Resource | Link | Description |
| ---------------------- | ----------------------------------------------------------------- | ---------------------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 plugins auto-collected |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Best practices compilation |
| claude-plugins.dev | [Website](https://claude-plugins.dev/) | Community registry and CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agents + 15 orchestrators |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 specialized agents |
## Relationship with Other Features
Plugin is a "container" concept that can include other features in the Claude Code ecosystem:
```
Plugin (Container)
├── Skills (Knowledge Packages)
├── Commands (Quick Commands)
├── Agents (Sub-agents)
├── Hooks (Event Hooks)
└── MCP/LSP (External Connections)
```
Understanding this hierarchy is important:
* **Skills** teach Claude how to do something
* **Commands** provide quick trigger entry points
* **Agents** handle independent, specialized tasks
* **Hooks** enable event-driven automation
* **Plugin** bundles all of these together for easy distribution and management
### Skills vs Plugins
The difference between Skills and Plugins can be confusing at first. According to [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins)'s analysis:
| Feature | Skills | Plugins |
| ---------------- | ------------------------------------- | ------------------------------------- |
| **Scope** | All Claude products (Web, API, Code) | Claude Code only |
| **Contents** | Markdown guides + optional scripts | Commands, Agents, Hooks, MCP, Skills |
| **Activation** | Automatic (model decides when to use) | Variable (depends on component type) |
| **Best For** | Teaching Claude domain expertise | Extending the Claude Code environment |
| **Distribution** | GitHub repos, file system | Decentralized Marketplace |
**Key insight**: Skills are automatically triggered by the model without manual invocation; Plugins are a packaging mechanism that solves the challenge of distributed sharing. The two can be used together — a Plugin can contain Skills.
## Summary
Claude Code Plugin is essentially a **workflow packaging and distribution mechanism**. It solves the pain points of configuration reuse and team collaboration, letting you share your carefully crafted toolchain with more people.
Remember three keywords:
| Keyword | Meaning |
| ---------------- | --------------------------------------------------------------- |
| **Packaging** | Integrates multiple configuration components into a single unit |
| **Isolation** | Namespaces prevent conflicts |
| **Distribution** | Easy sharing through the Marketplace |
Now that you understand the concepts, the next article, [Claude Code Plugin Practical Guide](/en/docs/notes/claude-plugin/practice), will walk you through hands-on practice: creating a Plugin from scratch, publishing to the Marketplace, and best practices for team collaboration.
If you're not yet familiar with the components a Plugin can contain, consider reading [What Are Claude Skills](/en/docs/notes/claude-skills/concept) to understand the core concepts of Skills first.
# Practical Guide
## Quick Recap
In the previous article, we explored the core concepts of Plugins: they are Claude Code's mechanism for packaging and distributing workflows, combining Commands, Skills, Agents, Hooks, and other components into a single unit for team sharing and community distribution. This article takes a hands-on approach, walking you through the complete process from creation to publishing.
## Creating Your First Plugin
### Step 1: Create the Directory Structure
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### Step 2: Create the Plugin Manifest
Define your plugin's metadata in `.claude-plugin/plugin.json`:
```json
{
"name": "my-first-plugin",
"description": "My first Claude Code plugin",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### Step 3: Add Slash Commands
Create Markdown files in the `commands/` directory. Each file corresponds to a command:
`commands/hello.md`:
```markdown
---
description: Send a friendly greeting to the user
---
# Hello Command
Please greet the user warmly and ask what you can help them with today.
```
### Step 4: Test the Plugin
Use the `--plugin-dir` flag to load your local plugin for testing:
```bash
claude --plugin-dir ./my-first-plugin
```
Run the command in Claude Code:
```
/my-first-plugin:hello
```
### Step 5: Add Command Arguments
Commands support receiving user-provided arguments. Update `hello.md`:
```markdown
---
description: Send a personalized greeting to a specific user
---
# Hello Command
Please warmly greet the user named "$ARGUMENTS" and ask what you can help them with today.
If the user didn't provide a name, use "friend" as the default.
```
Test the command with arguments:
```
/my-first-plugin:hello Alice
```
**Supported argument placeholders**:
* `$ARGUMENTS` - All user input
* `$1`, `$2`, `$3` - Individual arguments
## Adding More Components
### Adding Skills
Create a `skills/` directory. Each Skill is a folder containing a `SKILL.md` file:
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md`:
```yaml
---
name: code-review
description: Review code for quality, security, and maintainability
---
When reviewing code, check the following aspects:
1. **Code organization**: Is the structure clear?
2. **Error handling**: Are exceptions properly handled?
3. **Security concerns**: Are there any security vulnerabilities?
4. **Test coverage**: Is critical logic covered by tests?
Output format:
- Critical issues (must fix)
- Warnings (should fix)
- Suggestions (nice to have)
```
### Adding Subagents
Create an `agents/` directory:
`agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: Professional code review agent for quality inspection
tools: Read, Grep, Glob, Bash
---
You are an experienced code review expert.
When invoked:
1. Run git diff to see recent changes
2. Analyze modified files
3. Provide structured review feedback
Review checklist:
- Code readability and naming conventions
- Error handling and edge cases
- Security vulnerabilities (e.g., injection, sensitive data exposure)
- Performance optimization opportunities
```
### Adding Hooks
Hooks let you automatically run scripts when specific events occur. Create `hooks/hooks.json`:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**Important**: Use the `${CLAUDE_PLUGIN_ROOT}` environment variable to reference files within your plugin directory, ensuring paths resolve correctly regardless of where the plugin is installed.
Create the corresponding script `scripts/format-code.sh`:
```bash
#!/bin/bash
# Format code
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
Remember to set execution permissions:
```bash
chmod +x scripts/format-code.sh
```
### Adding MCP Servers
If your plugin needs to connect to external systems, create `.mcp.json`:
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## Complete Plugin Structure
A fully-featured Plugin might look like this:
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # Plugin manifest
├── commands/
│ ├── review.md # Code review command
│ ├── deploy.md # Deployment command
│ └── test.md # Testing command
├── agents/
│ ├── code-reviewer.md # Code review agent
│ └── debugger.md # Debugging agent
├── skills/
│ └── code-standards/
│ └── SKILL.md # Coding standards knowledge
├── hooks/
│ └── hooks.json # Event hook configuration
├── scripts/
│ ├── format-code.sh # Code formatting script
│ └── run-tests.sh # Test runner script
├── .mcp.json # MCP configuration
├── LICENSE
├── README.md
└── CHANGELOG.md
```
The corresponding `plugin.json`:
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "Developer toolkit: an all-in-one solution for code review, testing, and deployment",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## Publishing to the Marketplace
### What Is the Marketplace
The Marketplace is the distribution hub for Plugins. Think of it as a "plugin store" — users can install your published plugins with a simple command.
### Creating a Marketplace Configuration
Create `.claude-plugin/marketplace.json` in your GitHub repository:
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "Developer toolkit",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "Documentation generation tool"
}
]
}
```
### Plugin Source Types
The Marketplace supports multiple source types:
**Relative path** (plugins within the same repository):
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**GitHub repository**:
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**Any Git repository**:
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### Publishing Workflow
1. **Create a GitHub repository**
2. **Push your code**:
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **Users add your Marketplace**:
```bash
/plugin marketplace add your-username/your-repo
```
4. **Users install plugins**:
```bash
/plugin install dev-toolkit@your-marketplace
```
## Installing and Managing Plugins
### Via the Interactive Menu
```bash
/plugin
```
This opens an interactive interface where you can browse, install, enable, and disable plugins.
### Via the Command Line
**Add a Marketplace**:
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # Local
```
**Install plugins**:
```bash
# Install to user scope (default)
/plugin install formatter@my-marketplace
# Install to project scope (shared with team)
/plugin install formatter@my-marketplace --scope project
# Install to local scope (gitignored)
/plugin install formatter@my-marketplace --scope local
```
**Other management commands**:
```bash
/plugin enable # Enable a plugin
/plugin disable # Disable a plugin
/plugin uninstall # Uninstall a plugin
/plugin update # Update a plugin
```
### Validating Plugins
Verify your plugin configuration before publishing:
```bash
claude plugin validate .
```
Or within Claude Code:
```
/plugin validate .
```
## Team Collaboration Setup
### Sharing Plugin Configuration in a Project
Commit the plugin configuration to version control so team members automatically get it:
`.claude/settings.json`:
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
After team members clone the project, these plugins will be automatically available.
### Enterprise Marketplace Restrictions
For enterprise environments that require strict control, you can restrict allowed Marketplaces in managed settings:
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
Set it to an empty array `[]` to completely disable external plugins.
## CLI Command Reference
| Command | Description |
| -------------------------------------- | ---------------------------------------------------------- |
| `/plugin` | Open the interactive management interface |
| `/plugin install @` | Install a plugin |
| `/plugin uninstall ` | Uninstall a plugin |
| `/plugin enable ` | Enable a plugin |
| `/plugin disable ` | Disable a plugin |
| `/plugin update ` | Update a plugin |
| `/plugin validate .` | Validate the plugin configuration in the current directory |
| `/plugin marketplace add ` | Add a Marketplace |
| `/plugin marketplace list` | List added Marketplaces |
| `/plugin marketplace update` | Update the Marketplace cache |
| `/plugin marketplace remove ` | Remove a Marketplace |
## Best Practices
### Development Best Practices
1. **Keep Skills focused**: Each Skill should do one thing well — avoid catch-all designs
2. **Write clear descriptions**: Help Claude understand when to use your components
3. **Test with your team first**: Validate internally before distributing to the community
4. **Document version changes**: Record every version's changes in CHANGELOG.md
### Directory Structure Best Practices
* Place `commands/`, `agents/`, `skills/` in the plugin root directory
* Only put `plugin.json` inside the `.claude-plugin/` directory
* Use `${CLAUDE_PLUGIN_ROOT}` to reference files within the plugin
* Never use `../` to access files outside the plugin
### Hooks Best Practices
1. Scripts must be executable: `chmod +x script.sh`
2. Use a shebang to declare the interpreter: `#!/bin/bash`
3. Use the `${CLAUDE_PLUGIN_ROOT}` variable to ensure correct paths
4. Test scripts independently before integrating them into Hooks
### Version Management Best Practices
Follow Semantic Versioning:
* **MAJOR** (1.0.0 → 2.0.0): Breaking changes
* **MINOR** (1.0.0 → 1.1.0): New features (backward compatible)
* **PATCH** (1.0.0 → 1.0.1): Bug fixes (backward compatible)
## Troubleshooting
| Problem | Possible Cause | Solution |
| -------------------- | ----------------------------- | --------------------------------------------------------------- |
| Plugin doesn't load | Malformed plugin.json | Validate with `claude plugin validate` |
| Command not showing | Incorrect directory structure | Ensure `commands/` is in the root, not inside `.claude-plugin/` |
| Hooks not triggering | Script not executable | Run `chmod +x script.sh` |
| Path not found | Using relative paths | Switch to `${CLAUDE_PLUGIN_ROOT}` |
| MCP server fails | Environment variables not set | Check path configuration in `.mcp.json` |
## Migrating from Existing Configuration
If you already have configurations under the `.claude/` directory, follow these steps to migrate them into a Plugin:
1. **Create the Plugin structure**:
```bash
mkdir my-plugin/.claude-plugin
```
2. **Create plugin.json**:
```json
{
"name": "my-plugin",
"description": "Plugin migrated from existing configuration",
"version": "1.0.0"
}
```
3. **Copy existing files**:
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **Migrate Hooks**:
Copy the `hooks` configuration from `.claude/settings.json` to `hooks/hooks.json`
5. **Test**:
```bash
claude --plugin-dir ./my-plugin
```
## Learning Resources
### Official Documentation
| Resource | Link | Description |
| --------------------- | ----------------------------------------------------------------------------------------- | ------------------------------ |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | Plugin reference documentation |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | Plugin creation guide |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | Official best practices |
| Agent Skills Standard | [agentskills.io](https://agentskills.io) | Open standard specification |
### Official Repositories
| Resource | Link | Description |
| ---------------------------------- | --------------------------------------------------------------------------------------- | -------------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Official Skills repository |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Official plugin catalog |
| Docker MCP Toolkit | [Website](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ prebuilt MCPs |
### Community Resources
| Resource | Link | Description |
| ------------------------ | ---------------------------------------------------------------------------- | --------------------------------------- |
| claude-plugins.dev | [Website](https://claude-plugins.dev/) | Community registry and CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | Collection of 243 plugins |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Best practices compilation |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agents + 15 orchestrators |
| jeremylongshore tutorial | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | Hundreds of plugins + Jupyter tutorials |
### Recommended Reading
| Article | Source |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## Looking Ahead
The Plugin system represents a quantum leap in Claude Code's extensibility. As the community grows, we can expect to see:
* **A richer plugin ecosystem**: Covering a wide range of development scenarios and workflows
* **Enterprise-grade features**: More comprehensive permission management and auditing capabilities
* **Cross-platform compatibility**: The open Skills standard has already been adopted by multiple vendors
Now is the perfect time to get involved. Start with simple commands, gradually add Skills and Hooks, and ultimately build a complete workflow solution.
If you want to learn more about the Subagent component that Plugins can include, read [What Are Claude Code Subagents](/en/docs/notes/claude-subagent/concept).
# Introduction to the concept of Claude Code Subagent
## Introduction
When using Claude Code to handle complex tasks, you may have encountered such a dilemma: the context of the main conversation is getting longer and longer, the AI begins to "forget" previous important information, and the quality of the response gradually decreases.
Subagent was born to solve this problem.
If Skills is the "work manual" for Claude, then Subagent is the "full-time employee" you hire - they have their own independent work station (context), focus on a specific type of work, and report the results to you after completion.
## Understanding Subagent
Imagine you are the CEO of a company. When the company is small, you handle everything yourself. But as your business expands, you start hiring full-time employees: accountants for finance, HR for recruiting, and engineers for development. Each employee works at his or her own station and reports to you after completing the task.
Subagent plays exactly this role in Claude Code.
From a technical perspective, Subagents are specialized AI assistants that have the following characteristics:
| Features | Description |
| ---------------------------- | -------------------------------------------------- |
| **Independent context** | Each Subagent runs in its own context window |
| **Specialized Capabilities** | Optimized for specific task types |
| **Configurable Tools** | Can only access a specified set of tools |
| **Customized prompts** | There are special system prompts to guide behavior |
## Why do we need an independent context?
This is the core design concept of Subagent and deserves a deep understanding.
In a normal conversation, all information is piled into the same context. When you have Claude search the code base, analyze the files, and then make modifications, all of this intermediate processing takes up context space. As the conversation progresses and the context becomes more crowded, Claude may begin to "forget" earlier important information.
Subagent changes that:
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
Subagent's analysis process does not pollute the main conversation. The main dialogue received only refined results, maintaining clarity and focus.
## Built-in Subagent type
Claude Code provides three powerful built-in Subagents, covering the most common usage scenarios:
### Explore Subagent
**Targeting**: Fast, read-only exploration of the code base.
**Features**:
* Use Haiku model (fast, low latency)
* Strictly read-only - files cannot be created, modified or deleted
* Available tools: Glob, Grep, Read, Bash (read-only operations)
**When to use**:
When you ask exploratory questions such as "Where is this function implemented?" and "How are errors handled?" Claude will automatically call Explore Subagent.
**Detail Level**:
| Level | Description | Applicable scenarios |
| ------------- | ------------------------------------ | -------------------------------------------------- |
| Quick | Fast search with minimal exploration | Simple, targeted queries |
| Medium | Moderate exploration | Balancing speed and completeness |
| Very thorough | Comprehensive analysis | Complex issues that require in-depth understanding |
### Plan Subagent
**Positioning**: Study the code base and prepare an implementation plan.
**Features**:
* Use Sonnet model (stronger inference capabilities)
* Only exploration tools: Read, Glob, Grep, Bash
* Automatically called in planning mode
**When to use**:
When you enter planning mode and need Claude to conduct research before proposing a plan, Plan Subagent will automatically collect information and then give plan suggestions based on the research results.
### General-purpose Subagent
**Positioning**: Handle complex multi-step tasks.
**Features**:
* Use Sonnet model
* Access to all tools (including reading and writing)
* Suitable for complex tasks that require exploration and modification
**When to use**:
When the task involves multiple steps, requires searching before modifying, or the initial search may fail and multiple strategies need to be tried.
## My understanding and practice
If you carefully observe the three official built-in Subagents, you will find one thing in common: **They are all research and planning tasks**. Explore is responsible for exploring the code base, Plan is responsible for making plans, and even General-purpose is mainly used for research and analysis. None of them are specifically designed for writing code.
This confirms my understanding of Subagent: **The core value of Subagent is not "clean context", but to allow the main agent to focus on doing things**.
### Division of labor mode
My usage method is very simple: Subagent is responsible for "information collection" work such as research, planning, and review, and the main agent is responsible for actual execution.
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### Why not let Subagent write code?
Some people like to have the main agent schedule multiple subagents to write code. I think this is unreliable. The reason is simple: **Context is severely missing**.
The Subagent's context is independent. It does not know what was discussed, what decisions were made, and what constraints were imposed in the main conversation. Asking it to write code is like asking a new employee to complete a task without any background information - the code produced is likely not to match your expectations.
On the contrary, it is much more reasonable to position Subagent as a "researcher":
* The research task itself does not require much context
* Information is returned instead of code, which can be used by the main agent based on the full context
* Even if the survey results are biased, the main agent can correct them
### My daily usage
1. **Before starting a new task**: Let the Explore agent quickly understand the structure of the relevant code
2. **Complex task planning**: Let the Plan agent analyze the requirements and formulate implementation steps
3. **Code Review**: Let the Review agent check code quality and security issues
4. **Actual Coding**: The main agent writes code based on the collected context
Here is the list of agents I currently use:
The advantage of this is that the main agent's context window remains clean, with only "the information I need to know" instead of "a bunch of intermediate results generated during the Subagent search process".
## Comparison with other functions
### Subagent vs Skills
This is the most common confusion. Core difference: **Skills inject knowledge into Claude; Subagent creates independent workers**.
| Dimensions | Skills | Subagent |
| ------------------------ | -------------------------------------------- | --------------------------------------- |
| **Core Features** | Provide expertise and instructions | Agents that perform tasks independently |
| **Context** | Share main conversation context | Have independent context |
| **Trigger method** | Automatic matching based on description | Automatic delegation or manual call |
| **Applicable scenarios** | Make Claude better at certain types of tasks | Complex, multi-step independent tasks |
To put it figuratively: Skills are like training materials, allowing Claude to learn how to do something; Subagent is like a full-time employee, who completes the task independently at his or her work station and reports the results.
The two can be combined: a code review Subagent can load the code specification Skill to achieve the combined effect of "expert + professional knowledge".
### Subagent vs slash command
| Dimensions | Subagent | Slash Command |
| --------------------- | ------------------------------------- | ------------------------------ |
| **Activation method** | Automatic delegation or explicit call | User manual input |
| **Context** | Standalone context | Shared main conversation |
| **Complexity** | Suitable for complex tasks | Suitable for simple operations |
The slash command is a shortcut key, and you enter `/review` to trigger a predefined operation; Subagent is an independent worker that can complete complex multi-step tasks autonomously.
### Subagent vs Plugin
Plugin is a "container" concept, which can contain Subagent:
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
You can define a Subagent in the Plugin's `agents/` directory and distribute it with the Plugin.
## Agentic design pattern
Anthropic summarizes six core Agentic design patterns in its official documentation. Understanding these patterns can help better design Subagent systems:
| Pattern | Core Idea | Subagent Application |
| ------------------------ | --------------------------------------------------------- | -------------------------------------------------------------- |
| **Prompt Chaining** | Decompose complex tasks into multiple sequential steps | Chain call multiple Subagent |
| **Routing** | Distributed to specialized processors based on input type | Different types of tasks are delegated to specialized Subagent |
| **Parallelization** | Execute multiple independent subtasks at the same time | Start multiple Subagents in parallel |
| **Orchestrator-Workers** | Central coordinator assigns tasks to workers | Claude as coordinator, Subagent as worker |
| **Evaluator-Optimizer** | Generator output, evaluator optimization | Generate Subagent + Review Subagent |
| **Agents** | Independent agents that make autonomous decisions | Each Subagent runs independently |
These modes can be used in combination. For example, a code quality system might use both:
* **Parallelization**: run security scans and performance analysis simultaneously
* **Orchestrator-Workers**: Master Claude coordinates multiple specialized Subagents
* **Evaluator-Optimizer**: Review code immediately after generation
## Core Advantages
### Context protection
Subagent's greatest value lies in protecting the context of the main conversation. Intermediate processes such as code search and file analysis will not be accumulated in the main dialogue, allowing the main dialogue to always focus on high-level goals.
### Specialization capabilities
You can create specialized Subagents for specific domains, configured with detailed instructions and appropriate tools. A specialized Subagent performs better at a specific task than a general-purpose Claude.
### Flexible permission control
Each Subagent can have different tool access rights. For example, the exploration class Subagent only gives read-only permissions, and the modification class Subagent only gives write permissions. This fine-grained control improves security.
### Reusability
Once created, Subagent can be reused across projects or shared with teams via plugins.
## When to use Subagent
**Suitable scenarios for using Subagent**:
* Requires independent context to run tasks
* Tasks are complex multi-step workflows
* Requires a different toolset than the main conversation
* Tasks may take a long time to run
**Scenarios not suitable for using Subagent**:
* Simple one-time query
* Requires close interaction with main dialogue
* Missions can be completed quickly
## Typical application scenarios
### Code review
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
When you complete a piece of code, you can have Code Review Subagent review it in an independent context without interfering with your main development work.
### Debugging analysis
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
When an error is encountered, Debugging Subagent can deeply analyze the cause of the error, try various hypotheses, and finally provide repair suggestions.
### Code base exploration
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
When you're new to a new project, Explore Subagent can quickly map your code base without cluttering your main conversation with tons of search results.
## Learning resources
### Official resources
| Resources | Links | Instructions |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| Claude Code Documentation | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Official Documentation Entry |
| Subagent Guide | [Claude Code Docs](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Subagent Official Documentation |
| Multi-agent system research | [Anthropic Engineering](https://www.anthropic.com/engineering/multi-agent-research-system) | Research details of 90.2% performance improvement |
| Agentic Design Pattern | [Anthropic Docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Detailed explanation of six core design patterns |
### Community Resources
| Resources | Links | Instructions |
| -------------------- | ----------------------------------------------------------------- | --------------------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 Agents + 15 Orchestrators |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Plugins for 17 specialized agents |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Best practices summary |
## Summary
Claude Code Subagent is essentially a specialized, context-independent AI assistant. It solves the problem of information overload in complex tasks through context isolation, keeping the main conversation clear and focused at all times.
Remember three key words:
| Keywords | Meaning |
| --------------- | --------------------------------------------------------------- |
| **Independent** | Each Subagent has its own context window |
| **Specialized** | Optimized for specific task types |
| **Delegation** | Claude can delegate tasks to Subagent automatically or manually |
After understanding the concept, the next article "[Claude Code Subagent Practical Guide](/en/docs/notes/claude-subagent/practice)" will take you to practice: creating a custom Subagent, configuring tool permissions, and best practices in actual projects.
If you want to know the Skills that Subagent can load, please read "[What are Claude Skills](/en/docs/notes/claude-skills/concept)". If you want to package Subagent for distribution, please read "[What is Claude Code Plugin](/en/docs/notes/claude-plugin/concept)".
# Claude Code Subagent Practical Guide
## Quick review
In the previous article, we learned about the core concept of Subagent: it is a context-independent specialized AI assistant that solves the problem of information overload in complex tasks through context isolation. Claude Code has three built-in subagents: Explore, Plan, and General-purpose. This article will take you from a practical perspective to create a custom Subagent and master advanced usage.
## Manage Subagent
### Via /agents command
The easiest way is to use the interactive interface:
```bash
/agents
```
This will open a menu where you can:
* View all Subagents (built-in + custom)
-Create new Subagent
* Edit configuration and tool permissions for existing Subagent
* Delete unnecessary Subagent
* See which Subagent is active when there is a name conflict
### Through file management
Subagent is stored as a Markdown file. You can also create and edit files directly.
**Storage location**:
| location | path | scope |
| ------------- | ---------------------------- | ------------------------------------------------------------ |
| Project level | `.claude/agents/` | Dedicated to the current project and can be submitted to Git |
| User level | `~/.claude/agents/` | Available across all projects |
| Plugin | Plugin's `agents/` directory | Installed with Plugin |
**Priority**: Project level > User level > Plugin level
When a Subagent with the same name exists in multiple locations, the one with higher priority will overwrite the one with lower priority.
## Create your first Subagent
### Step 1: Create a directory
```bash
mkdir -p .claude/agents
```
### Step 2: Create Markdown file
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Step 3: Test Subagent
In Claude Code:
```
> 用 code-reviewer 代理审查我最近的修改
```
Or let Claude choose automatically:
```
> 帮我审查一下代码质量
```
If `description` is written clearly enough, Claude will automatically recognize and call your Subagent.
## Detailed explanation of configuration fields
Subagent configuration file consists of two parts: YAML frontmatter and Markdown body.
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| Field | Required | Description |
| ---------------- | -------- | ------------------------------------------------------------------------- |
| `name` | is | a unique identifier, use lowercase letters and hyphens |
| `description` | Yes | Natural language description (Claude uses this to determine when to call) |
| `tools` | No | Comma separated list of tools. If omitted, all tools are inherited |
| `model` | No | Model selection: `sonnet`, `opus`, `haiku`, or `inherit` |
| `permissionMode` | No | Permission Mode (see below) |
| `skills` | No | Autoloaded Skills (Subagent does not inherit the parent session's Skills) |
### Permission mode
| Mode | Description |
| ------------------- | --------------------------------------- |
| `default` | Normal permission check |
| `acceptEdits` | Automatically accept editing operations |
| `bypassPermissions` | Skip all permission checks |
| `plan` | Only propose a plan, not execute it |
| `ignore` | Ignore this Subagent |
### Markdown text
The text is Subagent's system prompt. The more detailed you write, the better Subagent performs.
A good system prompt should include:
* Clear role definition
* Specific work steps
* Key checklist
* Output format requirements
## Trigger mechanism
### Automatic delegation
Claude will automatically decide whether to delegate based on the task content and Subagent's `description`.
**Tip to encourage automatic use**: Use trigger words in `description`:
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
or:
```yaml
description: MUST BE USED when encountering errors or test failures
```
### Explicit call
Directly tell Claude which Subagent to use:
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
## Tool configuration
### List of commonly used tools
| Tools | Instructions |
| ----------- | ------------------------- |
| `Read` | Read file content |
| `Write` | Write to file |
| `Edit` | Edit file |
| `Glob` | File pattern matching |
| `Grep` | Regular expression search |
| `Bash` | Execute shell command |
| `WebFetch` | Get web content |
| `WebSearch` | Search the web |
### Tool configuration strategy
**Read only Subagent** (exploration, analysis):
```yaml
tools: Read, Grep, Glob, Bash
```
NOTE: Even if Bash is included, Subagent should only be used for read-only commands (ls, git status, git log, etc.).
**Read and write Subagent** (repair, refactor):
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**Principle of Least Privilege**: Grant only necessary tools to avoid accidental operations.
## Practical Subagent template
### Code Reviewer
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### Debugging Expert
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### Test runner
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### Document generator
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### Security Scanner
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## Advanced usage
### Production-level design patterns
In production environments, there are several proven patterns for multi-agent collaboration:
#### 3 Amigos Mode
A collaboration model composed of three roles: product, architecture, and implementation:
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| Roles | Responsibilities | Tool Configuration |
| --------------- | ---------------------------------------- | ------------------ |
| PM Agent | Function definition, requirement sorting | Read, WebSearch |
| Architect Agent | Technical Solution Design | Read, Glob, Grep |
| Claude Code | Code Implementation | All Tools |
#### Three-stage pipeline
Break down complex tasks into three clear stages:
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
Each stage is responsible for a dedicated Subagent, and the output serves as the input to the next stage.
#### Model orchestration strategy
Different models are used at different stages to optimize costs and effects:
| Stage | Recommended model | Reason |
| --------------- | ----------------- | ------------------------------- |
| Planning phase | Sonnet | Deep reasoning required |
| Execution phase | Haiku | Fast, low cost |
| Review stage | Sonnet | Comprehensive judgment required |
Configuration example:
```yaml
---
name: quick-executor
model: haiku
---
```
### Subagent link
For complex workflows, multiple Subagents can be linked:
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### Resumeable execution
Subagent execution can be paused and resumed, maintaining the full previous context:
**Initial call**:
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**Recovery Agent**:
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**Usage Scenario**:
* Long-running studies, completed in multiple sessions
* Iterative improvements, keeping context
* Multi-step workflow to process related tasks in sequence
### Configure Skills for Subagent
Subagent does not automatically inherit skills from the parent session. If necessary, declare it explicitly:
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### CLI dynamic definition
No need to save the file, define the temporary Subagent directly on the command line:
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
Suitable for quick testing or one-time use.
## Best Practices
### 1. Stay focused
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
A Subagent that does one thing well is better than a Subagent that does many things.
### 2. Write a clear description
Claude uses `description` to decide when to use Subagent. A good description should answer:
1. \*\*What does this Subagent do? \*\* List specific abilities
2. \*\*When should it be used? \*\* Contains trigger words
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. Restrict tool access
Grant only the tools you need:
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
This prevents Subagent from accidentally modifying files and allows it to focus more on its review work.
### 4. Write detailed system prompts
The more detailed the system prompts, the better Subagent performs:
* Clear role definition
* Specific work steps
* Key checklist
* Output format requirements
### 5. Version control
Commit project-level Subagent to Git:
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
Team members automatically get the same Subagent after cloning the project.
## Troubleshooting common problems
| Problem | Possible Cause | Solution |
| ------------------------- | ------------------------------------------------- | --------------------------------------------------------------------------- |
| Subagent is not called | description is not clear enough | Add trigger words to make it more specific |
| Subagent not being called | Wrong file location | Make sure the file is in `.claude/agents/` or `~/.claude/agents/` |
| Tools are not available | Tools field configuration error | Check the spelling of tool names and make sure they are separated by commas |
| The output is unstable | The system prompt is too vague | Add specific steps and output format requirements |
| Context lost | Session ended | Using resumable execution |
| Name conflict | Subagent with the same name in multiple locations | Use `/agents` to see which one is active |
## Share with team
### Method 1: Through Git
Place Subagent in the `.claude/agents/` directory and submit it to the project repository. Team members are automatically obtained after cloning.
### Method 2: Through Plugin
Place Subagent in the `agents/` directory of Plugin and distribute it through the Plugin mechanism.
### Method 3: User-level sharing
Place commonly used Subagents in `~/.claude/agents/` to make them available across all projects. Synchronization between multiple machines can be managed using dotfiles.
## Learning resources
### Official Documentation
| Resources | Links | Instructions |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- |
| Claude Code Documentation | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Official Documentation Entry |
| Subagent Guide | [Claude Code Docs](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Subagent configuration details |
| Agentic Design Patterns | [Anthropic Docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Six core design patterns |
| Multi-agent study | [Anthropic Engineering](https://www.anthropic.com/engineering/multi-agent-research-system) | Study details of 90.2% performance improvement |
### Community Resources
| Resources | Links | Instructions |
| -------------------- | ----------------------------------------------------------------- | ------------------------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 Agents + 15 Orchestrator Templates |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Plugins for 17 specialized agents |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Summary of Claude Code best practices |
### Recommended reading
| Article | Source | Topic |
| ----------------------------------------------- | ---------------------------------------------------------------------------------------------- | --------------------------------- |
| Building effective agents | [Anthropic](https://www.anthropic.com/research/building-effective-agents) | Agent design principles |
| How we built our multi-agent research system | [Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system) | Multi-agent architecture practice |
| Claude Skills, Commands, Subagents, and Plugins | [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Function comparison analysis |
## Summary
Claude Code Subagent is a powerful tool to improve AI programming efficiency. It makes complex tasks manageable through independent contexts and specialized configurations.
Quick start:
1. Run `/agents` to open the management interface
2. Create a simple Subagent (such as a code reviewer)
3. Test automatic delegation and explicit calls
4. Adjust configuration as needed
As you deepen your use, you can gradually:
* Create exclusive Subagent for your team
* Configure Subagent links to handle complex workflows
* Handle long-term tasks with resumable execution
If you want to package and distribute Subagent with other configurations, please read "[Claude Code Plugin Practical Guide](/en/docs/notes/claude-plugin/practice)".
# Advanced Tips
## Terminal Notifications: Task Completion Alerts
Want to get notified when Claude finishes a task?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
Pair this with iTerm2's notification feature, or use `terminal-notifier` for custom alerts (see the Hooks configuration in [Best Practices](/en/blog/claude-code-best-practices)).
## Advanced Hooks Usage
Hooks can do more than just run shell commands. There are actually four types:
1. **command**: Shell commands (most common)
2. **http**: POST JSON to a URL (supports custom headers and environment variable expansion)
3. **prompt**: Send to Claude for evaluation (e.g., "Are all tasks complete?")
4. **agent**: Spawn a sub-agent with tool access to verify
Some advanced Hook events worth knowing:
* `PostCompact`: Fires after compaction, useful for injecting reminders so Claude re-reads key files
* `SessionStart`: Write to `$CLAUDE_ENV_FILE` to persist environment variables for the entire session
* `PreToolUse`: Modify tool inputs (`updatedInput`), or even auto-approve or reject operations
## Plugin Ecosystem
Use `/plugin` to browse and install community plugins. Some notable ones:
* **dx** (by ykdojo): Provides `/handoff` (auto-generate handoff docs), `/clone` (clone conversations), `/half-clone` (clone only recent conversation to reduce context)
* **mine** (by anipotts): Imports all Claude Code session data into SQLite, supporting cost tracking, cache analysis, error memory queries, and more
## Agent Teams: Multi-Agent Collaboration
Set an environment variable to enable the experimental Agent Teams feature:
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
Once enabled, a session can act as a Team Lead, coordinating multiple agents working simultaneously through git worktree. Each agent runs independently in its own context window, making it ideal for parallel development on large projects.
Note that token consumption increases 4-15x, so use this judiciously.
## Prompt Strategies
The following tips come from Boris Cherny's Twitter thread on team practices — essentially "prompt engineering" best practices applied to Claude Code.
### Use Claude as Your Code Reviewer
Don't just have Claude write code — have it review your code too:
```
Grill me on these changes and don't make a PR until I pass your test.
```
Or ask it to prove the code works:
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### Don't Rephrase When Unsatisfied
Boris's tip #6: If Claude gives a mediocre answer, don't rephrase and ask again. Instead, say "This isn't good enough — tell me specifically what can be improved." Iterating on the existing response works better than starting from scratch.
### Let Claude Update Its Own CLAUDE.md
After correcting a mistake, add:
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris says Claude is surprisingly good at writing rules for itself. Over time, CLAUDE.md becomes increasingly precise, and conversation quality keeps improving.
### Just Say "fix"
With Slack MCP enabled, paste a bug report from Slack and say one word: **fix**. Zero context switching.
Or when CI fails, just say:
```
Go fix the failing CI tests.
```
No need to manually analyze logs or explain the problem — let Claude find the logs, diagnose the issue, and fix it.
## Final Thoughts
Claude Code is evolving rapidly, and these tips are constantly being refined. Follow the official Changelog to stay up to date.
If you haven't read my earlier articles, I recommend starting with the foundational workflows:
### Further Reading
* [My Claude Code Best Practices](/en/blog/claude-code-best-practices) — Core workflow tips and slash command guide
* [AI Coding Quality Control: 5 Lines of Defense](/en/blog/claude-code-quality-control) — Quality assurance system for Claude Code programming
* [Claude System Architecture Explained](/en/docs/notes/claude-architecture) — Understanding MCP, Skills, Subagents, Hooks, and more
# Useful Commands & Automation
## `/diff`: Interactive Diff Viewer
Type `/diff` to open an interactive diff view:
* **Left/Right arrows**: Switch between the full git diff (all changes) and Claude's per-turn changes
* **Up/Down arrows**: Browse different files
Much better than running `git diff` in the terminal, especially when changes span multiple files.
## `/simplify`: Multi-Agent Code Review
Running `/simplify` launches 3 parallel review agents:
* **Code Reuse** agent: Finds duplicate patterns
* **Code Quality** agent: Checks readability and structure
* **Efficiency** agent: Analyzes unnecessary performance overhead
The three agents work independently, then aggregate results — automatically fixing valid issues and skipping false positives.
## `/security-review`: Security Scan
Performs a security audit on the current branch's changes, checking for SQL injection, XSS, authentication flaws, data handling issues, and dependency vulnerabilities. Each finding goes through adversarial validation to reduce false positives.
## Hidden Features of `/copy`
`/copy` does more than just copy the last response. When the response contains code blocks, it pops up an interactive selector letting you pick a specific code block instead of copying the entire response. You can also pass a number to copy earlier responses: `/copy 2` copies the second-to-last, `/copy 3` copies the third-to-last — no need to scroll and manually select.
## `/batch`: Large-Scale Parallel Refactoring
```
/batch Migrate all components in src/ from Class components to function components
```
This is a heavyweight feature. `/batch` analyzes the codebase, breaks the task into 5–30 independent units, spins up a separate agent for each unit working in an isolated git worktree, and finally each agent commits and opens a PR.
Great for large-scale migrations, bulk type annotations, global renames, and similar scenarios.
## `/loop`: Scheduled Tasks
```
/loop 5m Check if the deployment is complete
/loop 1h /review-pr 1234
```
Creates a scheduled task within the session that repeats at the specified interval. Useful for polling deployment status, periodically checking PRs, etc. It's session-scoped (gone when you exit), limited to 50 tasks, and auto-expires after 3 days.
## Pipe Input: Feed Anything to Claude
```bash
# Have Claude analyze error logs
cat error.log | claude -p "Analyze this error log and find the root cause"
# Have Claude summarize recent changes
git diff HEAD~3 | claude -p "Summarize the changes in these three commits"
# Have Claude interpret command output
kubectl get pods | claude -p "Which pods have abnormal status?"
```
`-p` is headless mode (non-interactive), ideal for use in scripts and CI/CD pipelines.
## Hidden Flags for Headless Mode
`-p` mode has some extremely powerful but little-known flags:
```bash
# Set a spending cap (stops when exceeded)
claude -p --max-budget-usd 5.00 "Refactor the auth module"
# Limit conversation turns
claude -p --max-turns 3 "Fix this test"
# Output in JSON format (easy for programmatic parsing)
claude -p --output-format json "Analyze this project"
# Require output matching a specific JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "Summarize the project"
# Multi-turn headless conversation (use session-id to maintain context)
claude -p --session-id my-task "Step 1: Analyze the code"
claude -p --session-id my-task "Step 2: Generate tests"
# Specify a fallback model (auto-switches when primary is overloaded)
claude -p --fallback-model sonnet "Complex analysis"
# Restrict available tools
claude -p --tools "Read,Grep,Glob" "Read-only analysis, don't modify code"
# Completely replace the system prompt
claude -p --system-prompt "You are a Python expert" "Optimize this code"
```
# Configuration & Diagnostics
## `/statusline`: Custom Status Bar
Use `/statusline` to customize the information displayed in the bottom status bar using natural language descriptions. Alternatively, you can manually create a `~/.claude/statusline.sh` script.
Information you can display includes: current model, git branch, uncommitted file count, context usage progress, session cost, and more. When you have multiple Claude windows open for different tasks, the status bar helps you quickly identify what each window is doing.
## settings.json Autocomplete
Add `$schema` at the beginning of your settings.json, and VS Code / Cursor will provide autocomplete and validation for configuration options:
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## Some Useful Hidden Settings
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration`: Display the duration of each conversation turn
* `DISABLE_AUTOUPDATER`: Disable automatic update checks, reducing context overhead
## `/stats` and `/insights`: Usage Analytics
* `/stats`: Visualizes daily usage, session history, usage streaks, and model preferences, with date range filtering
* `/insights`: Analyzes all your Claude Code history, tells you which workflows are effective, where bottlenecks are, and generates optimization suggestions
## history.jsonl: Prompt History
Claude saves every prompt you send in `~/.claude/history.jsonl`. You can ask Claude to analyze this file to find prompt patterns and optimization opportunities.
## `/doctor`: Health Check
When you encounter strange issues, run `/doctor` (or `claude doctor` in the terminal). It checks installation status, version, authentication state, and system dependencies to help you quickly pinpoint problems.
## Community Tools
The community tool `ccusage` can track token usage:
```bash
npx ccusage daily
```
If you've used `--dangerously-skip-permissions` or approved many commands, you can use `cc-safe` to scan for risks:
```bash
npx cc-safe .
```
It checks `.claude/settings.json` for high-risk commands like `sudo`, `rm -rf`, `chmod 777`, `git reset --hard`, and more.
## More Slash Commands Worth Knowing
| Command | Function |
| --------------------- | ------------------------------------------------------------------------------------------ |
| `/export [filename]` | Export conversation as plain text |
| `/pr-comments [PR]` | Fetch PR comments (auto-detects current branch) |
| `/release-notes` | View current version changelog |
| `/plugin` | Browse and install community plugins |
| `/fast` | Toggle fast mode |
| `claude --debug` | Enable debug logging on startup (supports category filtering, e.g., `--debug "api,hooks"`) |
| `/install-github-app` | Install GitHub App for automatic PR reviews |
# Context Management
## `/compact` Accepts Arguments
Many people know `/compact` can compress context, but few realize it accepts arguments to specify what to preserve:
```
/compact Preserve all discussions about database schema and the current refactoring plan
```
This way, the compression will prioritize keeping the content you specified, preventing loss of critical context.
## Write Compaction Survival Instructions in CLAUDE.md
Add a `## Compact Instructions` section to your CLAUDE.md to tell Claude what must be preserved during compaction:
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
This ensures that even automatic compaction won't lose critical information.
## Prevent Claude from Giving Up Early Due to Token Budget
Add this to your CLAUDE.md:
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
Sometimes Claude will proactively stop when the context is nearly full, saying "context is almost full." Adding this prevents it from giving up prematurely.
## Handoff Protocol: Session Handover
When context is nearly full but the task isn't complete, have Claude write a handoff document:
```
Write the remaining plan to HANDOFF.md, explaining what you tried, what worked, and what didn't.
```
Then start a new session and simply `@HANDOFF.md` to restore full context. This compresses 10K+ tokens of context down to under 2K — far more precise than `/compact`.
## Proactively Compact at 70-80%
An easily overlooked point: when context approaches its limit, Claude automatically triggers compaction. But when automatic compaction happens mid-task, it may lose critical information and degrade subsequent response quality.
A better approach is **proactive management**: manually run `/compact` when context reaches 70-80% — this works much better than waiting for automatic compaction. Run `/clear` immediately after completing a task; don't let context grow indefinitely.
You can also trigger automatic compaction earlier via an environment variable:
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context`: Context Diagnostics
Not sure how much space remains in the context window? `/context` will tell you:
* Which tools or MCP services are consuming the most context
* Current capacity usage percentage
* Targeted optimization suggestions
I've found that sometimes just having certain MCP services registered (even without using them) can consume over 30% of the context window. Run `/context` to check, and cleaning up unused MCPs can free up significant space.
## Automatic Lazy Loading of MCP Tools
When MCP tool definitions exceed 10% of the context, Claude Code automatically enables Tool Search — loading a lightweight search index instead of full tool definitions. This reduces MCP context consumption by over 85% (e.g., from 77K tokens down to 8.7K). This feature is **enabled by default** and requires no manual configuration.
Note: Tool Search only supports Sonnet 4+ and Opus 4+ models, not Haiku. If your `ANTHROPIC_BASE_URL` points to an unofficial proxy, Tool Search will be automatically disabled (because most proxies don't forward `tool_reference` blocks).
To customize the behavior, set this in settings.json:
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
Supported configuration values:
* **Not set**: Enabled by default
* **`true`**: Force enable (including unofficial proxy scenarios)
* **`auto`**: Activates when context exceeds 10% (equivalent to default behavior)
* **`auto:`**: Custom threshold, e.g., `auto:5` means activate when exceeding 5%
* **`false`**: Disabled, all MCP tools are preloaded
# Claude Code Hidden Tips & Tricks
Practical Claude Code tips gathered from founder Boris Cherny's tweets, the community, and the changelog. Many of these are tucked-away gems -- shortcuts, hidden features, CLI tricks, and more -- that are hard to go back from once you start using them.
## Table of Contents
* [Keyboard Shortcuts](./shortcuts) -- Shift+Tab mode switching, Esc+Esc history recall, Ctrl+S stash, and more
* [Input & Interaction](./input-interaction) -- `!` terminal commands, `@` file injection, URL pasting, /btw interrupts, Vim mode
* [Thinking & Model Control](./thinking-model) -- think/ultrathink keywords, /effort, subagents, opusplan
* [Session Management](./session-management) -- /rename, /branch, /color, remote control
* [Context Management](./context-management) -- /compact arguments, compression directives, Handoff protocol, MCP lazy loading
* [Commands & Automation](./commands-automation) -- /diff, /simplify, /batch, /loop, Headless mode
* [Config & Diagnostics](./config-diagnostics) -- statusline, settings.json, /stats, /doctor
* [Advanced](./advanced) -- Hooks deep dive, plugin ecosystem, Agent Teams, prompt philosophy
### Further Reading
* [Claude System Architecture Deep Dive](/en/docs/notes/claude-architecture) -- Understanding MCP, Skills, Subagents, Hooks, and other components
* [The Complete Guide to Claude Subagents](/en/docs/notes/claude-subagent) -- Subagent concepts and hands-on practice
# Input & Interaction
## `!`: Run Terminal Commands Directly
Start your input with `!` to run terminal commands directly inside Claude Code, no need to switch to another terminal window:
```
! git status
! npm run build
! docker ps
```
Type `!` followed by a command prefix and hit Tab to autocomplete from your command history.
## `@` + File Path: Inject File Context
Use `@` followed by a file path to inject file contents directly into the context:
```
Check if there are any issues between @src/auth/login.ts and @src/auth/middleware.ts
```
Tab autocomplete works for paths, so you don't need to type the full path manually. This is faster than having Claude read the file itself, since it skips the tool call overhead.
## Paste URLs Directly
Paste a URL directly into your input and Claude will automatically fetch the web content as context:
```
Write client code based on this API doc https://docs.example.com/api/v2
```
## Feed `/llms-full.txt` to Let Claude Look Up Docs
Many open-source project doc sites provide a `/llms-full.txt` file (a complete, LLM-friendly version of the docs). When you hit an issue with a library, paste that file's URL and Claude can look up the docs to solve most problems on its own:
```
Use https://docs.astro.build/llms-full.txt to help me fix this routing issue
```
## `/btw`: Chime In While Claude Is Working
This is a new feature added in March 2026. While Claude is running a task, you can use `/btw` to start a side conversation — ask what it's thinking, give it extra info, all without interrupting the current task.
As Anthropic engineer @trq212 put it on Twitter: "Nobody Ctrl+C's a coworker. You just say 'btw' and they look up."
## Vim Mode
Type `/vim` to enable Vim mode, which supports:
* Mode switching (Normal/Insert)
* Navigation (h/j/k/l, w/b/e, 0/$)
* Editing operations (d, c, y, p)
* Text objects (iw, aw, i", a())
If you're a Vim user, this is a massive upgrade over the default input experience. Use `/config` to enable it permanently.
## Voice Mode
Type `/voice` to activate voice mode — hold spacebar to talk, release to send. Great for when you don't feel like typing but still need to give Claude instructions. The key binding can be customized in `keybindings.json`.
# Session Management
## `/rename`: Name Your Session
```
/rename my-auth-refactor
```
Give your current session a name. The benefit is that in the interactive session picker (`claude --resume`), named sessions can be selected and restored directly without pressing Enter to confirm. You can also launch directly from the terminal with `claude --resume my-auth-refactor`.
In the picker, just start typing to search and filter. It also supports these keyboard shortcuts: `Ctrl+V` to preview a session, `Ctrl+R` to rename, `Ctrl+A` to toggle showing all projects, and `Ctrl+B` to filter by branch.
## `/branch`: Branch Your Conversation
Just like git branches, `/branch` creates a fork at the current point in the conversation. You can try different approaches in the fork without affecting the original conversation. If things don't work out, just go back to the original branch and continue.
## `/color`: Color Your Window
Set a color for the current session's prompt bar. Supports red, blue, green, yellow, purple, orange, pink, and cyan.
## CLI Session Management Toolkit
```bash
# Resume the most recent session in the current directory
claude --continue
# Open the session picker, or resume by name
claude --resume
claude --resume my-auth-refactor
# Name the session at launch
claude -n "auth-refactor"
# Fork the last session (preserve context, create new branch)
claude -c --fork-session
# Resume a session associated with a specific PR
claude --from-pr 123
# Launch in an isolated git worktree
claude -w
```
Claude Code sessions are automatically saved (no need for Ctrl+S), so you can resume your last session with `--continue` every time you open a terminal.
## `claude --remote`: Continue Across Devices
```bash
claude --remote "your task description"
```
Start a web session that you can continue on claude.ai or the mobile app.
## `/remote-control`: Control Local Claude from Your Phone
Type `/remote-control` in Claude Code on your computer, and it will generate a connection code. Then enter this code in the Claude App on your phone to remotely control the local Claude Code session — give instructions to Claude on your computer from your phone.
# Keyboard Shortcuts
Claude Code's shortcut system is richer than most people realize -- press `?` to see every shortcut available in your current context.
## Shift+Tab: Cycle Through Modes
This is probably the single most important shortcut. Pressing `Shift+Tab` cycles between three modes:
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
No need to type `/plan` or `/auto-accept` manually -- one key does it all. My workflow: when I get a new task, I tap it twice to jump into Plan Mode, confirm the approach, then tap once more to switch to Auto-Accept and let Claude execute on its own.
## Esc + Esc: The Time Machine
Double-tap `Esc` and a rewind menu pops up:
* **Restore code and conversation**: Roll back to a previous checkpoint -- both files and chat history revert
* **Conversation only**: Roll back messages but keep the current code changes
* **Code only**: Undo file modifications but keep the conversation history
Claude automatically tracks every file edit as a checkpoint. This is far more granular than `git checkout .` because you can step back to any individual edit, not just the last commit.
One caveat: only files that Claude edited directly through tools are tracked. Files you changed by hand, `git push`, or other external operations can't be rewound.
## Ctrl+S: Prompt Stash
Halfway through writing a prompt and need to handle something else first? Press `Ctrl+S` and your current input gets stashed:
You can then type another command or question. After you submit that message, the stashed content **automatically restores** into the input box so you can pick up right where you left off.
Think of it as `git stash` but for prompts. Example scenario: you're writing a lengthy refactoring request, then realize you want Claude to check a file first -- press `Ctrl+S` to stash the request, ask your file question, and once that's answered your request comes right back.
## Ctrl+B: Send Tasks to the Background
Claude is in the middle of a long-running task (like a big refactor) and you want to work on something else? Press `Ctrl+B` to push the current task to the background -- your terminal is immediately free for new input.
Use `Ctrl+T` to view the background task list, and double-tap `Ctrl+F` to kill all background agents.
> tmux users: tmux's default prefix key is also `Ctrl+B`, so you'll need to press it twice to trigger Claude's background feature.
## Ctrl+G: Write Long Prompts in Your Editor
Sometimes you need to give Claude a lengthy set of instructions and typing in the terminal is painful. Press `Ctrl+G` to open your system's default `$EDITOR` (VS Code, Vim, etc.), write your prompt there, and it gets sent to Claude automatically when you save and close.
To change the default editor, set it in your shell config (`~/.zshrc` or `~/.bashrc`):
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
The `--wait` flag is important -- it tells the editor to block until you close the file, otherwise Claude receives empty content immediately. Terminal editors like Vim block naturally, so they don't need it.
Super useful for multi-paragraph requirement descriptions or pasting large reference material. In Plan Mode, you can even use `Ctrl+G` to edit Claude's generated plan directly in your editor.
## Cmd+T: Toggle Extended Thinking
The default shortcut is `Cmd+T` (or `Meta+T` on Windows/Linux) and it toggles Extended Thinking mode. When enabled, Claude reasons more deeply before responding -- great for complex architecture decisions or tricky bug hunts.
Fair warning: most terminals (iTerm2, Terminal.app, Warp, etc.) intercept `Cmd+T` as "new tab," so this shortcut often doesn't work in practice. Two workarounds: use `/keybindings` to remap it to a non-conflicting key, or just use the `/effort` command to switch thinking depth (same effect, plus you get fine-grained level control).
## Readline Shortcuts
Claude Code's input box supports standard Readline shortcuts -- terminal veterans will feel right at home:
| Shortcut | Action |
| ----------------- | ---------------------------- |
| Ctrl+A | Jump to beginning of line |
| Ctrl+E | Jump to end of line |
| Ctrl+W | Delete previous word |
| Ctrl+U | Delete to beginning of line |
| Ctrl+K | Delete to end of line |
| Ctrl+Y | Paste last deleted text |
| Alt+Y | Cycle through delete history |
| Option+Left/Right | Jump by word (Mac) |
## Approval Shortcuts: `y/n/d/e`
When Claude proposes a file change and waits for confirmation, four single-key shortcuts control the flow:
* `y`: Accept
* `n`: Reject
* `d`: View full diff
* **`e`: Edit before accepting**
`e` is the most overlooked yet most useful one -- it lets you tweak Claude's changes before they're applied. Not happy with a few lines? No need to reject and redo; just press `e` and fix it.
## Quick Reference
| Shortcut | Action |
| ----------- | ------------------------------------------------------------------------------------------------ |
| Shift+Tab | Cycle modes: Normal → Auto-Accept → Plan |
| Esc+Esc | Open rewind menu |
| Ctrl+S | Stash current input, auto-restores after next submit |
| Ctrl+B | Push current task to background |
| Ctrl+T | View background task list |
| Ctrl+F (x2) | Kill all background agents |
| Ctrl+G | Write prompt in external editor |
| Ctrl+O | Toggle verbose tool output view |
| Cmd+T | Toggle extended thinking (may be intercepted by terminal; consider remapping or using `/effort`) |
| `\` + Enter | Multi-line input (no setup needed) |
| Shift+Enter | Multi-line input (requires `/terminal-setup` first) |
| Up / Down | Browse input history |
| Ctrl+R | Search command history |
| Ctrl+L | Clear screen (history preserved) |
| Ctrl+C | Cancel current generation |
| Ctrl+D | Exit Claude Code |
| `?` | Show all available shortcuts |
## Custom Keybindings
If the defaults don't suit you, run `/keybindings` to open `~/.claude/keybindings.json` and customize away. Changes take effect immediately -- no restart required.
The syntax supports combo keys (e.g., `ctrl+shift+c`) and chord sequences (e.g., `ctrl+k ctrl+s` -- press Ctrl+K, release, then press Ctrl+S). There are 16 different binding contexts (Chat, Autocomplete, Confirmation, DiffDialog, etc.), each with its own set of bindable actions.
# Thinking & Model Control
## Control Thinking Depth with Keywords
Adding specific keywords to your prompts triggers different levels of thinking budget. This is a Claude Code-exclusive feature (not available on claude.ai web):
| Keyword | Thinking Budget | Use Case |
| ----------------------------- | --------------- | -------------------------------------- |
| `think` | \~4,000 tokens | Everyday coding questions |
| `think hard` / `megathink` | \~10,000 tokens | Complex logic, multi-file dependencies |
| `think harder` / `ultrathink` | \~31,999 tokens | Architecture design, tricky bugs |
In practice, I usually add `think hard` when Claude gives a shallow answer and re-ask the question. For particularly complex problems (like debugging across multiple services), I go straight to `ultrathink`.
## `/effort`: Control Thinking Depth
Besides keywords (think / ultrathink), you can use `/effort` to set thinking depth directly:
```
/effort low # Simple tasks, skip deep thinking, faster and cheaper
/effort high # Complex tasks, deep reasoning
/effort max # Maximum thinking budget (Opus only)
/effort auto # Let Claude decide
```
The setting persists for the entire session. Use `low` for simple file edits and `max` for complex architecture design — this saves money without sacrificing quality.
## The `use subagents` Keyword
Append `use subagents` to any request, and Claude will break the task into multiple sub-agents that run in parallel. This is not only faster but also keeps the main agent's context window clean.
Boris specifically mentioned this on Twitter: offload individual tasks to sub-agents to keep the main agent's context focused.
## `opusplan`: Best Value Model Strategy
One-line summary: **Opus thinks, Sonnet does**.
## `/model`: Switch Models
Use `/model` to switch models on the fly during a session. For example, use Sonnet day-to-day, temporarily switch to Opus for complex problems, then switch back when you're done.
## Output Style Control
In `/config`, select "Output style" — there are two uncommon but very useful modes:
* **Explanatory mode**: Claude inserts "knowledge nuggets" between tasks, explaining relevant frameworks and code patterns — great for learning a new project
* **Learning mode**: A collaborative learning mode where Claude adds `TODO(human)` markers in the code for you to implement yourself, instead of giving you the answer directly
You can also create custom output style files (in Markdown format) at `~/.claude/output-styles/` to directly modify the system prompt. Note: custom output styles will **completely replace** the default coding system prompt unless you set `keep-coding-instructions: true`.
# Concept Introduction
## Introduction
In October 2025, Anthropic quietly released a new feature called Claude Skills. Despite its understated debut, prominent tech blogger Simon Willison called it "maybe a bigger deal than MCP" and predicted it would trigger a "Cambrian explosion" in the AI tooling space.
This assessment is far from hyperbole. If you use AI assistants regularly, you have likely encountered a familiar frustration: every new conversation requires you to re-enter the same workflow instructions; you finally get the AI tuned to your satisfaction, only to start from scratch in a new chat window. Skills were built to solve exactly this pain point.
## Understanding Claude Skills
Imagine you are a company owner handing new employees an operations manual that documents workflows, brand guidelines, and standard procedures for handling common issues. Claude Skills is essentially that "operations manual" for an AI assistant -- enabling it to complete specific tasks in a repeatable, standardized way.
From a technical perspective, Skills are folders containing instructions, scripts, and resources that Claude can dynamically load on demand. Each Skill teaches Claude how to handle a particular type of task consistently, and this knowledge persists across conversations. This means you only need to "train" it once -- from then on, Claude will remember how to do it no matter when you use it.
### Three Core Components
A complete Skill consists of three parts:
| Component | Purpose | Required |
| ----------------------- | ---------------------------------------------------------------------------------- | -------- |
| **SKILL.md** | Core instruction document containing metadata and detailed instructions | Required |
| **Reference Materials** | Brand guidelines, policy documents, templates, and other supplementary information | Optional |
| **Scripts** | Python/JavaScript code for complex computations or file operations | Optional |
SKILL.md is the "soul" of the entire Skill. Its basic structure looks like this:
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
The YAML frontmatter at the top of the file contains two key fields: `name` is the Skill's identifier, limited to 64 characters; `description` tells Claude what the Skill does and when to use it, limited to 200 characters. Claude uses this description to determine when to invoke a given Skill, so the clearer and more accurate it is, the higher the probability that the Skill will be triggered correctly.
### Use Cases
Skills cover a wide range of applications across everyday repetitive tasks:
**Document Generation**: Batch-create Excel spreadsheets, PowerPoint presentations, Word documents, and PDF reports. Anthropic even provides an official set of document Skills that work out of the box.
**Brand Compliance**: Package your company's brand colors, logo usage rules, spacing guidelines, and tone of voice into a Skill, ensuring that all AI-generated content adheres to brand standards.
**Meeting Notes**: Automatically summarize meeting recordings, extract action items, assign owners, and generate follow-up emails.
**Data Analysis**: Execute standardized analysis workflows, such as competitive intelligence scanning (structured extraction of product updates, pricing changes, and analyst commentary) or financial analysis (analyzing earnings reports and building financial models).
**Project Management**: Build project plans from objectives, suggest milestones, and generate weekly reports or investor briefs.
## Progressive Disclosure Architecture
The most elegant aspect of Skills is how information is loaded. Traditional MCP tool descriptions can consume thousands or even tens of thousands of tokens, while Skills metadata takes up only a few dozen tokens. This means you can enable a large number of Skills simultaneously without worrying about tool descriptions filling up the context window.
This efficiency stems from an architectural pattern called **Progressive Disclosure**. Skills use a three-layer information structure that loads content on demand -- much like a manual with a table of contents:
```
📚 Skills Playbook
│
├─ 📋 Table of Contents ──────────────── [Metadata Layer] Preloaded at startup
│ │
│ │ name: "weekly-report"
│ │ description: "Generate standardized weekly reports from work content"
│ │
│ │ ✓ Only 30-50 tokens
│ │ ✓ All Skills' metadata visible at once
│ │
│
├─ 📖 Chapters ────────────────────────── [Core Document Layer] Loaded when relevant
│ │
│ │ # Weekly Report Generator
│ │
│ │ ## Instructions
│ │ Generate weekly reports following this structure...
│ │
│ │ ## Examples
│ │ Input: Completed the login feature this week...
│ │ Output: ### Completed This Week ...
│ │
│ │ ⚡ Expanded only when Claude determines it's needed
│ │ 📊 Consumes hundreds to thousands of tokens
│ │
│
└─ 📎 Appendix ────────────────────────── [Reference Resource Layer] Loaded when needed
│
│ references/
│ ├── brand-guide.md Brand guidelines
│ ├── template.xlsx Report template
│ └── examples/ Past reports
│
│ 🔍 Loaded only when explicitly needed
│ 📦 Can contain extensive reference materials
```
You first scan the table of contents to see what chapters are available (metadata layer), then open the relevant chapter to read it (core document layer), and finally consult the appendix if you need more details (reference resource layer).
| Layer | Content | When Loaded | Token Cost |
| ---------------------------- | -------------------------------- | -------------------- | --------------------- |
| **Metadata Layer** | name + description | Preloaded at startup | 30-50 |
| **Core Document Layer** | Full SKILL.md content | Loaded when relevant | Hundreds to thousands |
| **Reference Resource Layer** | Reference files, templates, etc. | Loaded when needed | On demand |
This aligns perfectly with the fundamental nature of large language models -- "feed text to the model so it understands." Skills do not introduce complex protocols or API calls; instead, they use carefully organized text structures to help AI efficiently acquire and apply knowledge. Simon Willison described this design as "absurdly elegant," precisely because it solves a complex problem in the most straightforward way possible.
## Core Advantages
### Token Efficiency
The progressive disclosure architecture of Skills delivers exceptional token efficiency. A simple comparison illustrates the difference:
| Approach | Tokens at Startup | Total for 100 Skills |
| ------------------------------- | ------------------------------ | ------------------------- |
| Traditional (full load) | Thousands to tens of thousands | May exceed context window |
| Skills (progressive disclosure) | 30-50 | 3,000-5,000 |
Since each Skill's metadata occupies only a few dozen tokens, you can enable dozens or even hundreds of Skills simultaneously, with full content loaded on demand -- never wasting precious context space.
### Composability
Multiple Skills can work together automatically. When you present a complex task, Claude intelligently identifies which Skills to invoke and coordinates them to complete the task.
For example, if you say "Generate a quarterly report from this sales data," Claude might:
1. Invoke a data analysis Skill to process the raw data
2. Invoke a chart generation Skill to create visualizations
3. Invoke a document Skill to produce the final report
Throughout this process, you never need to manually specify which Skill to use -- Claude automatically selects and combines them based on the task requirements.
### Portability
The same Skill works across all platforms in the Anthropic ecosystem:
| Platform | Description |
| ----------- | ------------------------------------------------------- |
| Claude.ai | Web interface, suited for general users |
| Claude Code | CLI tool, suited for developers |
| API | Programmatic integration, suited for system development |
A brand writing Skill you create for your team will behave consistently across all these platforms, truly achieving **build once, use anywhere**.
> **Other AI Platforms' Strategies**: Skills is currently an Anthropic-exclusive feature. OpenAI uses a dual-track approach with Custom GPTs + Assistants API (two systems that are not unified); Microsoft Copilot and Google Gemini focus on deep integration within their respective ecosystems rather than reusable skill modules. Claude Skills is considered a meaningful differentiator.
### Efficiency Data
According to Anthropic's internal benchmarks, teams using Skills **reduced repetitive prompt engineering time by 73%**. This is not just about efficiency gains -- more importantly, it represents the standardization and reusability of workflows. Team members no longer need to maintain their own sets of prompts; instead, they share a single set of verified Skills.
## Summary
At its core, Claude Skills are **reusable playbooks for AI assistants**. Through progressive disclosure architecture, they achieve exceptional token efficiency, enabling AI to master extensive domain knowledge without consuming precious context space.
Remember three key concepts, and you will have grasped the essence of Skills:
| Concept | Meaning |
| -------------- | ------------------------------------------------------------------ |
| **Efficient** | Metadata occupies only a few dozen tokens; content loads on demand |
| **Composable** | Multiple Skills work together automatically |
| **Portable** | Consistent experience across platforms |
Now that you understand the concepts, the next article, [Claude Skills Practical Guide](/en/docs/notes/claude-skills/practice), will walk you through hands-on practice: how to enable and install Skills, create your first custom Skill, and avoid common pitfalls.
If you want to further formalize your workflows, check out [What Is Spec-Driven Development](/en/docs/notes/speckit/concept) to learn how to elevate AI programming from "intuition" to "engineering."
# Practical Guide
## Quick Recap
In the [previous article](/en/docs/notes/claude-skills/concept), we covered the core concepts of Skills: reusable playbooks for AI assistants that achieve exceptional token efficiency through a progressive disclosure architecture, with three key strengths — efficiency, composability, and portability. This article takes a hands-on approach to help you understand how Skills differ from other features, learn to enable, install, and create Skills, and master best practices while avoiding common pitfalls.
## Feature Comparison
The Claude ecosystem offers multiple features, and it can be confusing to tell them apart at first. The table below provides a quick overview:
| Feature | What It Is | Best For | Persistence |
| ------------- | -------------------- | ---------------------------------------- | ------------------------------- |
| **Skills** | Expertise packages | Repetitive tasks, standardized workflows | Persistent across conversations |
| **Prompts** | Instant instructions | One-off requests | Current conversation only |
| **Projects** | Knowledge bases | Background info, project docs | Within project workspace |
| **MCP** | Connectors | External data, API calls | Continuous connection |
| **Subagents** | Sub-agents | Task delegation, parallel processing | Across sessions |
### Skills vs MCP
This is the most common source of confusion. The core distinction: **MCP connects Claude to data; Skills teach Claude how to process data**. They complement rather than replace each other.
| Dimension | Skills | MCP |
| ------------------------ | ------------------------------------------- | ------------------------------------------------- |
| **Core Function** | Teaches Claude how to perform tasks | Connects Claude to external systems |
| **Token Consumption** | Very low (tens of tokens) | Higher (thousands to tens of thousands of tokens) |
| **Technical Complexity** | Simple (Markdown + YAML) | Complex (full protocol specification) |
| **Typical Use Cases** | Brand writing, report generation, workflows | Database queries, API calls, cloud services |
| **Portability** | Across Claude.ai/Code/API | Adopted by multiple model providers |
Once you understand this distinction, you'll know when to use each one. Use MCP when you need to query databases, call APIs, or access cloud services. Use Skills when you need to follow a specific writing style, execute standardized workflows, or reuse domain expertise.
The best practice is to combine both: use MCP to connect to your CRM system and pull customer data, then use Skills to define how to analyze that data and generate reports.
### Skills vs Subagents
The core distinction: **Skills make Claude better at certain types of tasks; Subagents let Claude delegate tasks to independent "specialist workers."**
| Dimension | Skills | Subagents |
| ----------------- | ------------------------------------------- | -------------------------------------------- |
| **Core Function** | Provides expertise and instructions | Independent sub-agents that execute tasks |
| **Context** | Injected into the main conversation context | Has its own independent context window |
| **Use Cases** | Making Claude better at specific task types | Complex, multi-step independent tasks |
| **Activation** | Automatically matched based on description | Manually invoked or auto-delegated by Claude |
| **Portability** | Across Claude.ai/Code/API | Claude Code and Agent SDK only |
Think of it this way: Skills are like training materials — they teach Claude how to do something. [Subagents](/en/docs/notes/claude-subagent) are like dedicated employees — they have their own workspace (context) and permissions (tools), complete tasks independently, and report back with results.
The two can be combined: for example, a code review sub-agent can load language-specific best practice Skills, achieving a "specialist + domain expertise" combination. According to [Anthropic research](https://www.anthropic.com/engineering/multi-agent-research-system), multi-agent systems (Claude Opus 4 as the orchestrator + Claude Sonnet 4 sub-agents) scored 90.2% higher than single-agent setups in internal evaluations.
### Skills vs Slash Commands
If you've used Claude Code, you're familiar with [slash commands](/en/blog/claude-code-best-practices) like `/commit` and `/review`. The core distinction: **Skills activate automatically based on context; slash commands require manual input to trigger.**
| Dimension | Skills | Slash Commands |
| --------------------- | --------------------------------------------- | ---------------------------------- |
| **Activation** | Automatic (matched by context) | Manual input (e.g., `/commit`) |
| **Trigger Condition** | Claude decides relevance based on description | User explicitly enters the command |
| **Use Cases** | "Always-on" capability enhancement | Explicit, repeatable operations |
| **User Awareness** | Invisible, takes effect automatically | Requires remembering command names |
Example: When you type `/commit`, Claude executes a predefined commit workflow — that's a slash command. When you say "help me write a weekly report," Claude automatically identifies and loads the weekly report generation Skill without any command input — that's Skills.
A simple way to remember: slash commands are keyboard shortcuts that you trigger manually; Skills are background knowledge that Claude uses at its own discretion.
### Skills vs Plugins
Plugins are Claude Code's extension package mechanism. The core distinction: **Skills are auto-activated capability extensions; Plugins are packaged and distributable complete workflow configurations.**
| Dimension | Skills | Plugins |
| ----------------- | ----------------------------------- | ------------------------------------ |
| **Core Function** | Capability extension | Packaged workflow distribution |
| **Activation** | Auto-activated based on context | Components merged after installation |
| **Scope** | Cross-platform (Claude.ai/Code/API) | Claude Code only |
| **Contents** | Instructions + scripts + resources | Slash commands + hooks + skills |
| **Distribution** | Standalone folder | Installed via marketplace |
The key insight: Plugins can contain Skills (in their `skills/` directory) and represent a larger packaging unit. When you install a Plugin, its Skills are automatically activated, slash commands appear in autocomplete, and hooks are merged with your existing configuration.
In short: use Skills to extend Claude's capabilities; use Plugins to distribute standardized workflow configurations across your team.
## Hands-On Tutorial
### Option 1: Enable Built-in Skills
This is the easiest way to get started. Anthropic provides a set of practical document skills:
| Skill | Capability |
| --------------------- | --------------------------------------------------------------- |
| **Excel (xlsx)** | Create spreadsheets, analyze data, generate reports with charts |
| **PowerPoint (pptx)** | Create presentations, edit slides, analyze presentation content |
| **Word (docx)** | Create documents, edit content, format text |
| **PDF (pdf)** | Generate formatted PDF documents and reports |
**Steps to enable**:
1. Log in to [Claude.ai](https://claude.ai)
2. Click your avatar in the top right and go to **Settings**
3. Find the **Capabilities** option
4. Enable the skills you need
Once enabled, test it right away: "Create an Excel spreadsheet for Q3 sales budget with monthly breakdowns and totals."
> **Note**: Requires a Pro, Max, Team, or Enterprise plan, and the code execution feature must be enabled.
### Option 2: Install Community Skills
If you use Claude Code, you can install community-contributed Skills via commands.
**Install from the plugin marketplace**:
```bash
# Add the official Skills repository
/plugin marketplace add anthropics/skills
# Install the document skills package
/plugin install document-skills@anthropic-agent-skills
# Install the example skills package
/plugin install example-skills@anthropic-agent-skills
```
**Skills storage locations**:
| Location | Path | Description |
| --------------- | ------------------- | --------------------------------------------------- |
| Personal Skills | `~/.claude/skills/` | Available only to you |
| Project Skills | `.claude/skills/` | Version-controlled with git, shared across the team |
### Option 3: Create a Custom Skill
This is where Skills truly shine — creating workflows tailored to your needs.
**Step 1: Create the folder structure**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
A complete Skill folder might look like this:
```
weekly-report/
├── SKILL.md # Core instructions (required)
├── template.md # Report template (optional)
└── examples/ # Sample reports (optional)
├── good-example.md
└── bad-example.md
```
**Step 2: Write the SKILL.md**
SKILL.md is the heart of the entire Skill. It consists of two parts: YAML frontmatter (metadata) and a Markdown body (detailed instructions).
**Required metadata**:
| Field | Requirement | Description |
| ------------- | ------------------ | ----------------------------------------------------- |
| `name` | Max 64 characters | Unique identifier for the skill |
| `description` | Max 200 characters | Tells Claude when to use this skill (very important!) |
**Optional metadata**:
| Field | Description |
| --------------- | ----------------------------------------------------- |
| `dependencies` | Required packages, e.g., `python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | List of permitted tools |
| `model` | Optional model override |
A complete weekly report generation Skill example:
```yaml
---
name: weekly-report
description: 根据本周工作内容生成标准化的周报,包含进展、问题和下周计划
---
# 周报生成助手
## 使用场景
当用户需要生成周报、工作总结或进度汇报时,使用此技能。
## 输出格式
请按以下结构生成周报:
### 本周完成
- 列出已完成的主要工作项
- 每项包含简短说明和成果
### 进行中
- 列出正在进行的工作
- 标注当前进度和预期完成时间
### 遇到的问题
- 列出阻碍进展的问题
- 如果有,说明需要的支持
### 下周计划
- 列出下周的主要任务
- 按优先级排序
## 风格要求
- 使用简洁的表达
- 避免过于技术化的术语
- 突出成果和影响
## 示例
**输入**:这周完成了用户登录功能,修复了 3 个 bug,参加了产品评审。
**输出**:
### 本周完成
- 用户登录功能开发:完成前后端联调,支持邮箱和手机号登录
- Bug 修复:解决了 3 个高优先级问题,提升系统稳定性
### 进行中
- (无)
### 遇到的问题
- (无)
### 下周计划
- 开始用户注册功能开发
- 编写单元测试用例
```
**Step 3: Test it**
Test in Claude: "Help me generate this week's report. This week I completed the user login feature, fixed 3 bugs, and attended two product review meetings."
### Using the Skill Creator
If you'd rather not write the SKILL.md from scratch, Claude has a built-in skill-creator skill that guides you interactively:
```
Help me create a skill for [your workflow]
```
Claude will walk you through a series of questions to clarify your requirements, then generate a draft SKILL.md.
## Technical Principles
### Skills as a Meta-Tool System
At their core, Skills are a **meta-tool system** — they don't execute code directly but inject specialized instructions into the conversation context, altering how Claude reasons.
When you trigger a Skill, two things happen:
1. **Metadata message**: A visible status indicator showing which Skill is being loaded
2. **Skill prompt**: The complete SKILL.md instructions are sent to Claude, hidden from the user
### Discovery and Selection Mechanism
How does Claude know which Skill to invoke? The answer: **it relies entirely on language understanding.**
The name and description of all enabled Skills are formatted into a dynamic list written into the system prompt. When you send a message, Claude uses its native language understanding to match your intent and decide whether to invoke a given Skill.
This is why the `description` field is so critical — it's the sole basis for Claude's decision. There's no complex algorithmic routing; the decision happens entirely within Claude's reasoning process.
## Best Practices
Through extensive real-world usage, the community has distilled four golden rules for creating Skills:
**1. Stay Focused**
A Skill should do one thing well. Multiple focused Skills are far more useful than one catch-all Skill — they're easier to maintain and easier to compose.
**2. Write Clear Descriptions**
The description field determines when Claude invokes your Skill, so be specific about the applicable scenarios. "Generate quarterly analysis reports from sales data" is a good description; "process data" is too vague.
**3. Provide Examples**
Including input/output examples in your SKILL.md significantly improves output consistency, especially for tasks with specific formatting requirements.
**4. Start Simple**
Begin with plain Markdown instructions, validate the results, and only then consider adding scripts. Increase complexity gradually.
### Troubleshooting Common Issues
| Problem | Likely Cause | Solution |
| ------------------- | --------------------------------- | ------------------------------------------------ |
| Skill not triggered | Description is not precise enough | Rewrite with more specific scenario descriptions |
| Skill not triggered | Skill not properly installed | Check file paths and naming |
| Inconsistent output | Missing examples | Add more input/output examples |
| Inconsistent output | Instructions too vague | Add constraints and formatting requirements |
| Slow to load | File too large | Move large files to a references subdirectory |
### Security Considerations
Skills can execute code, so security matters:
* **Trust the source**: Only use Skills from trusted channels
* **Review scripts**: Inspect script code in Skills before installing
* **Protect sensitive information**: Never hardcode API keys or passwords in Skills
* **Manage permissions**: When used by a team, be mindful of the sharing scope of Skills
## Current Limitations
As an emerging feature, Skills currently has some limitations:
| Limitation | Details |
| -------------------------------- | ------------------------------------------------------------------------ |
| ~~**Anthropic ecosystem only**~~ | Resolved — see below |
| **No review mechanism** | No built-in review or audit workflow yet |
| **Learning curve** | Teams need to adapt workflows and establish version management processes |
| **Early stage** | The ecosystem is still evolving |
> **Major Update (December 18, 2025)**: Anthropic officially released Agent Skills as an [open standard](https://agentskills.io). The specification and reference SDK are available at [agentskills.io](https://agentskills.io).
>
> **Companies/products that have adopted it**:
> 
>
> * **Microsoft**: Integrated into VS Code and GitHub
> * **OpenAI**: ChatGPT and Codex CLI use the same architecture
> * **Developer Tools**: Cursor, Goose, Amp, OpenCode
> * **Partner Skills**: Atlassian, Figma, Canva, Stripe, Notion, Zapier
>
> Additionally, Anthropic, OpenAI, and Block co-founded the [Agentic AI Foundation](https://www.linuxfoundation.org/) (hosted by the Linux Foundation), with Google, Microsoft, and AWS also joining. This means Skills is evolving from a single-vendor feature into an industry standard — Skills written for Claude Code can interoperate with OpenAI Codex CLI.
>
> References:
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## Learning Resources
### Official Resources
| Resource | Link | Description |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------- | ----------------------------- |
| Skills GitHub Repo | [anthropics/skills](https://github.com/anthropics/skills) | Official examples, 22k+ Stars |
| Claude Code Docs | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Skills usage guide |
| Help Center | [support.claude.com](https://support.claude.com) | FAQ |
| Technical Blog | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | In-depth technical analysis |
| API Quickstart | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | Developer integration guide |
| Agent Skills Open Standard | [agentskills.io](https://agentskills.io) | Official spec and SDK |
### Community Picks
| Resource | Link | Description |
| --------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------- |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Curated collection of Skills |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148+ slash commands, 54 AI agents |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | Office document creation and editing skills |
### Recommended Reading
| Article | Author |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic (Official) |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## Looking Ahead
The emergence of Skills represents an important direction in AI tooling — enabling AI not just to execute tasks, but to learn and remember specific ways of working. Simon Willison predicted that Skills would trigger a "Cambrian Explosion" in the AI tool space, and that prediction is no exaggeration.
As more developers and teams build and share Skills, we can expect to see:
* **Specialized Skills Marketplaces**: Domain experts across industries packaging their knowledge into reusable Skills
* **Deep Skills + MCP Integration**: Forming complete end-to-end workflows
* **Enterprise-Grade Skills Platforms**: Team collaboration, version management, and access control
Now is a great time to get started. Here's what you can do right away: log in to Claude.ai and enable the document skills. This week, try installing a community Skill and creating your first simple Skill. In the long run, identifying repetitive work within your team and gradually building a dedicated skill library will be an effective way to boost productivity.
### Further Reading
* [A Complete Guide to Claude's System Architecture](/en/docs/notes/claude-architecture) — Where Skills fit within the overall Claude system
* [The Complete Guide to Claude Subagents](/en/docs/notes/claude-subagent) — A deep dive into the Subagent mechanism
* [My Claude Code Best Practices](/en/blog/claude-code-best-practices) — Tips and tricks for everyday Claude Code usage
# Skill-Creator in-depth analysis: Use data to drive your skill development
## Introduction
This article is based on information in March 2026 and corresponds to Claude Code v2.1+.
If you have read [Concept](/en/docs/notes/claude-skills/concept) and [Practice](/en/docs/notes/claude-skills/practice), you should already know how to manually write a SKILL.md file - define frontmatter, write a command, save it to the `.claude/skills/` directory, and you're done.
But here’s a fundamental question: \*\*How do you know your skills are really useful? \*\*
You may have changed the wording of a paragraph and feel it works better, but that's just your subjective feeling. Maybe with a different prompt word, the new version would be worse. Maybe your skills aren't improved at all compared to being unskilled - Claude can do just as well on his own.
In the conceptual and practical chapters, the process of skill development is as follows: **Written → Tried → Feeling OK → Online**. The whole process relies on intuition, there is no quantification, and there is no way to answer "How much better is this skill than no skill at all?" And Skill-Creator turned this thing into engineering: **Written → Parallel testing with/without skills → Blind test A/B comparison → Quantitative scoring → Feedback iteration → Data verification**.
This is why Skill-Creator exists. It not only helps you "generate a SKILL.md", but provides a complete set of create → test → evaluate → optimize loop, allowing you to let your data speak for itself.
## What is Skill-Creator
Skill-Creator itself is also a Skill - a 33KB SKILL.md file plus supporting subagent guidance files, Python scripts and HTML viewers. Its directory structure looks like this:
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
Installation is also very simple:
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## Follow this again: Evaluate and optimize an existing skill
Let’s go through the complete process of Skill-Creator using the skills I actually use. I maintain a Claude Code plug-in market [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub), in which the `yux-video-summary` skill is used to convert video subtitles into structured summaries - supporting Chinese and English language detection, DUAL\_FILE/SINGLE\_FILE two output modes, filler word cleaning, etc. The skill's SKILL.md looks like this:
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
The skill has been written, but how do you know it is really useful? \*\* This is where Skill-Creator comes into the picture.
> There is an important writing principle in the Skill-Creator source code: *"Try hard to explain the **why** behind everything. If you find yourself writing ALWAYS or NEVER in all caps, that's a yellow flag — reframe and explain the reasoning."* Meaning: A good skill should **explain why**, rather than pile up rigid rules.
### Step 1: Create test cases and run the evaluation
Core question: \*\*Is this skill really better than no skill at all? \*\*
Open Claude Code and enter directly:
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator will first read skill definitions and schemas, and then automatically generate test cases and quantitative assertions. My run generated 3 test cases and 39 assertions:
Note that it does not compile test cases casually - it understands the two output modes of DUAL\_FILE and SINGLE\_FILE defined in the skill, and specifically designs scenarios that cover different video types (tutorials, podcast interviews, technology sharing) and language combinations. The design of Assertions is also very particular, from language detection, output mode selection to content quality and Chinese and English filler word cleaning, it is much more comprehensive than I want to test the dimensions myself.
Then, the system starts two independent subagents simultaneously for each test case - with\_skill (loading skills) and **without\_skill** (baseline, no skills are loaded). **6 parallel agents** (3 test cases × 2 versions) were started at one time, each running in an **independent worktree** without interfering with each other.
> Anthropic's PDF skill previously had problems handling non-fillable forms - Claude needed to place text at precise coordinates without defining fields. The failure point was isolated through Eval, and the team subsequently fixed the positioning logic. That's the value of Eval - turning "something doesn't feel right" into "what exactly is wrong here".
### Step 2: Three sub-agents relay scoring
After all operations are completed, the three professional sub-agents **automatically** appear in sequence:
**Grader** Verifies assertions one by one. It will check whether the summary of the with\_skill version contains the Overview table, whether the DUAL\_FILE mode is correctly selected, whether the filler word has been cleaned, and then records the pass/fail and evidence of each item, generating `grading.json`:
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**Comparator** does a blind A/B comparison - it receives two summaries, but **doesn't know which is the skill version and which is the baseline version**. It only sees "Output A" and "Output B" and independently judges them based on its own quality standards to determine the winner.
**Analyzer** combines the above results to make a diagnosis: which assertions passed regardless of skills or not (indicating that this assertion has no differentiation and should be replaced with a better assertion), which results have high variance (the test is unstable), and what is the trade-off between time and token. Finally, suggestions for improvements are given.
### Step 3: Review results in Eval Viewer
Once scoring is complete, Skill-Creator will automatically open an HTML viewer in your browser.
**Outputs tab** You can view the output of each test case one by one. There is a feedback text box at the bottom - write down what you think is not good enough, such as "the summary lacks a timeline" and "the filler word is not cleaned up." After reading all use cases, click **Submit All Reviews** and the feedback will be saved to `feedback.json`.
**Benchmark Results tab** You can see the quantitative comparison: the pass rate, time consumption, token consumption of with\_skill and without\_skill, as well as the item-by-item comparison of each assertion.
### Step 4: Iterate and improve until satisfied
Go back to Claude Code and tell it that you have finished giving feedback. Skill-Creator will read `feedback.json` and give analysis and improvement suggestions based on the benchmark data:
My skill performed well with a pass rate of 97%. Skill-Creator accurately identified a small problem - the interview video lacked Notable Quotes paragraphs, and made suggestions for repairing it.
The key is that it doesn't patch individual test cases - it generalizes your feedback, understands the requirements behind it and adjusts the overall structure of the skill, then rewrites SKILL.md, reruns all tests into the `iteration-2/` directory, and opens a new Eval Viewer so you can compare the output of the two rounds. This cycle continues until you are satisfied.
> A noteworthy improvement philosophy in the Skill-Creator source code: *"We're trying to create skills that can be used a million times across many different prompts. Rather than put in fiddly overfitty changes, or oppressively constrictive MUSTs, if there's some stubborn issue, try branching out and using different metaphors."* Core idea: **Avoid overfitting** to test cases, and pursue generalization capabilities.
### Step 5 (optional): Optimize the description so that the skill triggers at the right time
The quality of the skill is verified, but there is another issue that is easily overlooked: the `description` field of the skill determines when Claude will call it.
Input:
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator automatically generates about 20 evaluation queries (half should be triggered, half should not be triggered), and the review interface is opened in the browser:
Note that these queries are available in both Chinese and English, covering a variety of real expressions. "Shouldn't trigger" queries shouldn't be too outrageous - a good counterexample is "Help me summarize this meeting minutes", which shares the keyword "summary" with video summarization, but actually requires document processing skills rather than video summarization.
You can edit the query text directly on the page, click **+ Add Query** to add new ones, use the Delete button to delete inappropriate ones, and you can also toggle the Should Trigger switch for each query. After confirming that it is correct, click **Export Eval Set** to export the JSON file. Go back to Claude Code and tell it that you have exported it. The system will automatically run the optimization loop in the background:
The entire process is fully automated - split the query into a training set and a test set 60/40, iteratively optimize the description on the training set (up to 5 rounds), and use the test set results to select the best version to avoid overfitting. After running, the description comparison before and after optimization will be output:
The optimized description becomes more specific - it clarifies the supported file types (.vtt/.srt), emphasizes pipeline features (filler cleaning, DUAL/SINGLE\_FILE logic), and uses MUST USE to exclude scenarios that should not be triggered. Anthropic internally used this set of optimizers to run its own document creation skills. As a result, the triggering accuracy of 5 out of 6 public skills was improved.
### Advanced usage: dynamic context injection
If you want the skill to automatically inject context when loading, you can embed a shell command in SKILL.md using Skills 2.0's `!` syntax:
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
These commands are executed before Claude sees the skill, and the data is embedded directly into the prompt. Compared with letting Claude explore the files one by one, it saves a lot of time and tokens.
## Two types of skills: Which one should you create?
Before using Skill-Creator, it is necessary to understand the two skill types defined by Anthropic:
**Ability improvement type** - Let the model do things that it can't or can't do well before. For example:
* Image generation skills: Claude cannot generate images natively, but it can be achieved by calling tools such as nanobanner through skills.
* Front-end design skills: Default AI designs are often very "AI-flavored", and good design skills can greatly improve quality.
**Coding Preference** - solidify your specific workflow. The model already has individual capabilities, but you need a precise order of execution. For example:
* PR review skills: Check code security according to fixed procedures and output risk level reports
* Video summary skills: output according to a specific template structure, automatic language detection and filler word cleaning
The reasons why these two types of skills need to be tested are different: **Capability improvement type** may become unnecessary as the model evolves - if the baseline (without\_skill) can also pass all assertions, it means that the model is native enough, and this skill can be retired; **Coding type** is more durable, but you need to verify whether it is really faithful to your workflow.
Skill-Creator's evaluation capabilities allow you to continually verify whether a skill is still valuable, rather than blindly using a skill that may be outdated.
## What the community says
The Skill-Creator update sparked a lot of discussion, from X/Twitter to Reddit to independent blogs, and real feedback is more valuable than official documentation.
### Is it really useful? Data speaks
The most direct question: Is adding skills really better than not adding skills? \*\* Several actual measurements give a clear answer.
Reddit u/hashpanak ran an eval on the title generation skill and had a 100% pass rate with\_skill and only 60% without\_skill. When asked whether the token cost was worth it, he replied: "Absolutely. After optimization, repeated tasks can be converted into scripts, which will save tokens." u/spences10 is even more extreme - he ran 250 sandbox evals and increased the skill activation rate from 84% to 100%. Comments section u/Manfluencer10kultra said: "**This should become standard practice.**"
Blogger [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/) benchmarked WordPress security skills: all 21 assertions passed (the baseline was only 90.5%), and the speed was 9.9% faster. His summary: **"Skills used to be art, now they're engineering."**
@0zhuxiaofeng gave more specific figures from the perspective of actual workflow: "After using it for a month, the biggest change is that run\_eval allows skills to score themselves. The agent I run content operations now automatically evaluates the effect after each release, and poor skills are directly eliminated and rewritten. **Manual intervention has been reduced from 3 hours a day to half an hour**."
### Overlooked Blind Spot: Trigger ≠ Quality
Blogger [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/) pointed out a blind spot that no one mentioned: **Skills can pass quality evaluation but fail on trigger evaluation** - the output quality is very good, but it will never be called. After three rounds of `run_loop.py` optimization, he triggered eval to 13/13. Core insight: "The description of a skill is not metadata, but a learnable parameter - you need to optimize the real routing behavior."
This coincides with @DrWang5257's suggestion: "Don't rewrite the whole thing at once. First split it into three sections: trigger conditions, input templates, and failure fallback, and iterate step by step. This way the update speed is fast and the rollover rate is low."
### Real pain points
Although the effect is good, there are also many pitfalls:
* **token consumption is huge**. @konghao10 bluntly said that "token consumption is huge" - running 6 parallel agents at the same time is really not cheap. Reddit u/munkymead also said that "it's expensive to get a serious test."
* **If you have too many skills, you will fight**. [RoboRhythms blogger Noah Albert](https://www.roborhythms.com/best-claude-code-skills-2026/) found that **starting to have problems when skills reach 8-10**: Claude will self-question the output, generate more verbose prefaces, and occasionally have command conflicts between skills. However, Reddit u/Specialist\_Solid523 countered: "Poorly written skills only eat context. **Well-written skills almost always make your token usage more efficient.**"
* **SKILL.md gets longer with more iterations**. Reddit u/IulianHI pointed out a contradiction: with iterative improvements, skill files continue to expand, \*\* but crowd out the context window for actually doing things \*\*. Test cases that only cover the happy path miss the critical 5%.
* **Version management is missing**. @fengqve complains "Why does Skill **not have the concept of version**? It has been updated so many times that it is hard to describe which update it is." This is especially painful after multiple rounds of iterations.
* **Headless mode has bug**. There is a key issue on GitHub: the skill is never triggered in `claude -p` mode, causing the recall describing the optimization loop to always be 0% ([#36570](https://github.com/anthropics/claude-code/issues/36570)).
### Thinking further: recursive self-improvement
@vista8 shared a related paper \[Memento-Skills: Let Agents Design Agents] ([https://github.com/Memento-Teams/Memento-Skills](https://github.com/Memento-Teams/Memento-Skills)), and someone in the comment area accurately summarized it: "The core bottleneck of Skill is iteration - it is easy to write the first version, but it is difficult to make it better and better used in real scenarios. If you can automate this 'use → evaluate → improve' cycle, it is equivalent to installing a self-evolution engine for Agent."
A 104-like thread on Reddit r/ClaudeAI also discusses this direction. But the top comment poured cold water on it - u/Tatrions said: "The recursive loop works, but the hard part is knowing when to trust improvements. We found that we have to do evidence gating - don't commit changes unless a failure occurs at least twice. Otherwise, each cycle is 'fixing' something that is not broken in the first place, and it ends up being worse."
## Installation and Ecology
Skill-Creator, as one of the skills officially maintained by Anthropic, is included in the [anthropics/skills](https://github.com/anthropics/skills) warehouse, which contains 17+ production-level skills.
The broader Skills ecosystem is also growing rapidly: [skills.sh](https://skills.sh) The market provides a convenient discovery and installation experience, and the community has maintained 1,234+ agent skills.
## Write at the end
The core problem Skill-Creator solves is: \*\*How do you know your skills are really effective? \*\*
In the absence of it, skill development relies on "writing → trying → feeling okay". With Skill-Creator you can:
* Test both skilled and unskilled effects with **Parallel Agent**
* Eliminate assessment bias with **Blind A/B comparison**
* Visualize results and leave feedback with **Eval Viewer**
* Use **Description Optimizer** to accurately control the triggering timing of skills
* Use **iterative loops** to continuously improve until you are satisfied
This is in line with the concept of test-driven development in software engineering - not "just write the code and think it can run", but "use testing to prove that it actually works as expected."
Anthropic put forward an interesting prospect in the official blog: as the model's capabilities improve, SKILL.md may evolve from "implementation plan" (telling Claude **how**) to "specification description" (telling Claude **what** and let the model figure it out on its own). The Eval framework is the first step in this direction - Eval describes "what to do". If one day this description itself is enough to become a skill, then the testing system established by Skill-Creator will become even more important.
If you already use Skills, try using `/skill-creator` to take an assessment of your most used skills. You might be surprised to find that some skills aren’t actually any better than no skills at all—and that’s where optimization starts.
Related reading:
* [What are Claude Skills](/en/docs/notes/claude-skills/concept) — Understand the core principles of Skills
* [Practice Guide](/en/docs/notes/claude-skills/practice) — Create your first Skill
# Concept Introduction
## Introduction
In the [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept), we explored a core problem: **Context Rot** — as conversations grow longer, Claude's context window fills up with failed code, outdated discussions, and irrelevant information, causing output quality to steadily degrade.
Ralph's solution was to "restart everything": use an infinite bash loop to spin up a fresh Claude instance each time, passing state through the file system. Simple, effective, but with clear limitations — it's just a methodology with no project understanding, no phase planning, and no quality verification. You need to write specs yourself, orchestrate tasks yourself, and judge "is it done?" yourself.
As Chase AI precisely summarized in his video: **The Ralph Loop is an incredibly powerful weapon, but most people don't need a weapon — they need an entire arsenal.** The Ralph loop depends entirely on upfront preparation: Is your PRD good enough? Are your feature definitions tight enough? Do you know what "done" looks like? If the answers to these questions aren't precise, no matter how many times the loop runs, it's just garbage in, garbage out.
What if you want a system that **doesn't just loop Claude, but truly understands your project and reliably delivers code**?
That's what **GSD (Get Shit Done)** is built to do.
## What is GSD
GSD was created by **TÂCHES** (GitHub: glittercowboy), an independent developer. His motivation was straightforward:
> "I'm not a 50-person software company. I don't want to play enterprise theater. I'm just a creative person who wants to make cool things."
In his livestream, TÂCHES demonstrated a stunning fact: he **never writes code by hand**. Using GSD, he built a complete macOS-native AI music generation app (Sample Digger) from scratch in 4 hours — zero hand-written code. He positions himself not as a programmer, but as a "high-level project manager" — describing the vision, making key decisions, and validating results. GSD makes this way of working possible.
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> —— TÂCHES
Other spec-driven development tools — BMAD, SpecKit — each have their own value, but they tend to introduce complex enterprise workflows: sprint ceremonies, story points, stakeholder syncs. For solo developers or small teams, these processes are burdens in themselves. As Chase AI put it: "It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way."
GSD's design philosophy is to **hide complexity inside the system**. Users only need a few simple commands while the system handles all the context management, task orchestration, and quality verification behind the scenes. Within one month of release, the project had earned nearly 3,000 GitHub stars and 14,000 npm installs, with TÂCHES pushing updates 15–20 times almost every day.
### Where GSD Fits in the Tool Ecosystem
| Dimension | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| ----------------------- | ------------------------------- | ------------------------ | --------------------------- | ------------------------------------- |
| Core Positioning | Execution technique (bash loop) | Spec generation toolkit | Enterprise-grade framework | **Context engineering + spec-driven** |
| Planning Capability | None (bring your own spec) | Strong (spec→plan→tasks) | Strong (full agile process) | **Strong (research→discuss→plan)** |
| Execution Autonomy | Highest (AFK mode) | Manual trigger per step | Manual trigger per step | **Manual trigger per step** |
| Human Involvement Model | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Context Rot Handling | New session restart | No built-in solution | No built-in solution | **Subagent fresh context** |
| Quality Verification | Relies on external tests | Build checks | Built-in QA process | **Auto verification + UAT** |
| User Complexity | Lowest | Medium | Higher | **Low** |
| System Complexity | Lowest | Medium | Higher | **High** |
This table reveals a key tradeoff: **Ralph trades minimal system complexity for maximum execution autonomy** — start it and go to sleep; while **GSD trades high system complexity for planning quality and human oversight** — you have the opportunity to intervene at every stage. SpecKit and BMAD fall in the middle ground, offering planning capabilities but lacking GSD's context engineering and Ralph's autonomous execution.
GSD and Ralph are not contradictory. GSD inherits Ralph's core principles — fresh context, files as source of truth — but builds a complete project understanding and execution framework on top. If Ralph is "give the AI a task and let it keep trying," GSD is "understand what you want, research how to do it, plan the steps, execute, and verify."
Chase AI's summary captures it perfectly: **The Ralph loop assumes you come with a complete blueprint — GSD helps you build that blueprint.** GSD takes your half-formed idea, asks deep questions, conducts research on your behalf, generates a complete PRD, breaks it down into atomic tasks, and delivers the project end-to-end. And when executing code, it uses the very same foundational principles that make the Ralph loop powerful: fresh context for subagents, and tasks as small and precise as possible.
## Core Workflow
GSD's workflow is a **discuss → plan → execute → verify** loop, with each stage having clear inputs and outputs.
### 1. Initialize the Project
```text
/gsd:new-project
```
One command kicks off the entire process. The system will:
1. **Ask questions** — Keep probing until it fully understands your idea (goals, constraints, tech preferences, edge cases)
2. **Research** — Dispatch parallel agents to investigate relevant domains (optional but recommended)
3. **Extract requirements** — Distinguish between v1, v2, and out-of-scope items
4. **Roadmap** — Create a phased plan aligned with the requirements
You approve the roadmap, then start building. TÂCHES's experience is: the more detailed the initial description you provide, the fewer follow-up questions the system asks; the vaguer it is, the more it asks. He recommends preparing a rough vision document before starting — you don't need to know the tech stack or implementation details, just describe what you want.
**Output files**: `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md`
> Already have a codebase? Run `/gsd:map-codebase` first — the system will dispatch parallel agents to analyze your tech stack, architecture, conventions, and potential issues. Then `/gsd:new-project` can plan based on your existing codebase.
### 2. Discuss Phase
```text
/gsd:discuss-phase 1
```
Each phase in the roadmap has only a sentence or two of description — that's not enough to build what you want. The discuss phase exists to **capture your implementation preferences** before research and planning.
The system analyzes the current phase and identifies "gray areas" — decision points where multiple reasonable implementation approaches exist:
* Visual features → layout, interactions, empty state handling
* API/CLI → response format, error handling, verbosity
* Content systems → structure, tone, depth, flow
Every decision you make here directly affects the quality of subsequent research and planning. Skipping this step is fine (the system will use sensible defaults), but deeper discussion lets the system build something closer to your expectations.
**Output files**: `{phase}-CONTEXT.md`
### 3. Plan Phase
```text
/gsd:plan-phase 1
```
The system will:
1. **Research** — Investigate how to implement the current phase, guided by the discuss phase decisions
2. **Plan** — Create 2–3 atomic task plans using XML-structured format
3. **Verify** — Check if the plan satisfies requirements, iterating until it passes
A key design principle is **Goal-Backward Planning**. Instead of starting from "what should we build?", it asks "what conditions must be true for the goal to be achieved?" — then works backward to derive the plan and tasks. TÂCHES says this approach "massively improved output quality" because each task understands its relationship to other tasks, rather than being just an item on a to-do list.
Each plan is small enough to execute within a single fresh context window. This is critical — **there will be no quality degradation**.
**Output files**: `{phase}-RESEARCH.md`, `{phase}-{N}-PLAN.md`
### 4. Execute Phase
```text
/gsd:execute-phase 1
```
The system will:
1. **Wave execution** — Independent tasks run in parallel; dependent tasks run sequentially
2. **Fresh context** — Each plan executes in a brand-new 200k token context, zero accumulated garbage
3. **Atomic commits** — Each task gets an independent git commit
4. **Goal verification** — Check whether the codebase delivers the functionality promised by the phase
In TÂCHES's livestream demo, he completed 3 full phases of development with his **main context window staying at just 24%**. The GSD Executor subagent needs to load fewer than 1,000 lines of context to complete an entire phase — you can execute 10 plans in a row and the context still stays below 50%. This is a completely different experience from working directly in Claude Code: no more "playing Russian roulette, betting on when you'll hit the context window wall."
**Output files**: `{phase}-{N}-SUMMARY.md`, `{phase}-VERIFICATION.md`
### 5. Verify Phase
```text
/gsd:verify-work 1
```
Automated verification can check whether code exists and tests pass. But does the feature **work the way you expect**? That requires your confirmation.
The system will:
1. **Extract testable deliverables** — List things you should now be able to do
2. **Guide verification one by one** — "Can you log in with email?" Yes/no, or describe the issue
3. **Auto-diagnose failures** — Dispatch a debug agent to find the root cause
4. **Create a fix plan** — A directly executable fix
If everything passes, continue to the next phase. If there are issues, run `/gsd:execute-phase` again to execute the fix plan.
This is the biggest philosophical difference between GSD and the Ralph loop: **Ralph is hands-off — start it and let it run; GSD has a human verification step after every phase.** Chase AI points out that the Ralph loop is "go conquer" style — it runs on its own without looking back; GSD ensures you can course-correct at every critical checkpoint, preventing errors from compounding under zero supervision.
Additionally, GSD provides a dedicated debugging workflow. When verification finds issues, `/gsd:debug` launches an **isolated debug subagent** with its own hypothesis-evidence-resolution workflow, creating independent debug documentation to track the entire investigation process without polluting the main context.
**Output files**: `{phase}-UAT.md`
### Rinse and Repeat
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
Every phase goes through the complete **discuss → plan → execute → verify** cycle. Context stays fresh, quality stays consistent.
When all phases are complete, `/gsd:complete-milestone` archives the milestone and tags the version. Then `/gsd:new-milestone` kicks off the next version's build.
## Why It Works: Technical Principles
GSD's reliability isn't accidental — four key technical pillars support it.
### Context Engineering
Claude Code is extremely powerful when given the right context. Most people don't know how to give it the right context. GSD handles this for you.
| File | Purpose |
| ----------------- | ------------------------------------------------------------------ |
| `PROJECT.md` | Project vision, always loaded |
| `research/` | Ecosystem knowledge (tech stack, features, architecture, pitfalls) |
| `REQUIREMENTS.md` | Versioned requirements with phase traceability |
| `ROADMAP.md` | Direction and progress |
| `STATE.md` | Decisions, blockers, position — memory across sessions |
| `PLAN.md` | Atomic tasks + XML structure + verification steps |
| `SUMMARY.md` | Execution records, committed to history |
Every file has **size limits** based on Claude's quality degradation thresholds. Stay within the limits, and you get consistently high-quality output. The main context window stays at 30–40%, while actual work happens in subagents' fresh 200k contexts.
Chase AI has an intuitive explanation for context rot: **No matter how large the context window — Sonnet, Opus, even million-token windows — tokens in the first half are more effective than those in the second half.** This isn't a bug; it's an inherent property of LLMs. Claude Code's built-in autocompact can only partially mitigate this. GSD's approach is more thorough: every atomic task executes in a fresh subagent, ensuring each task gets Claude's best performance.
TÂCHES's own data confirms this: on the $200/month Max plan, he consumes roughly $30,000 worth of Opus tokens per month. That sounds like a lot, but because every task executes in fresh context, rework is minimal — the actual efficiency is far higher than repeatedly patching things in a degraded context.
### XML Prompt Formatting
Every plan is structured XML optimized for Claude:
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
Precise instructions, no guesswork, verification built into every task.
### Multi-Agent Orchestration
Every stage uses the same pattern: a thin orchestrator dispatches specialized agents, collects results, and routes to the next step.
| Stage | What the Orchestrator Does | What Agents Do |
| ------------ | ---------------------------------- | ------------------------------------------------------------------------------- |
| Research | Coordinates, surfaces findings | 4 parallel researchers investigate tech stack, features, architecture, pitfalls |
| Planning | Validates, manages iterations | Planner creates plans, checker validates, loops until passing |
| Execution | Groups into waves, tracks progress | Executors implement in parallel, each with a fresh 200k context |
| Verification | Surfaces results, routes next step | Verifier checks codebase, debugger diagnoses failures |
The orchestrator never does the heavy lifting. It dispatches agents, waits, and integrates results. The result: you can run an entire phase — deep research, multiple plan creation and validation, thousands of lines of code written in parallel, automated verification — **while your main context window stays at 30–40%**.
### Atomic Git Commits
Each task is independently committed immediately upon completion:
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
Benefits: `git bisect` can pinpoint the exact failing task, each task can be independently rolled back, and clean history helps Claude understand code evolution in future sessions.
## GSD's Limitations
GSD is powerful, but understanding what it **cannot do** is equally important.
### GSD is a Human-Guided Workflow, Not an Autonomous Agent
GSD cannot run persistently. Every stage boundary — from `discuss` to `plan` to `execute` to `verify` — requires you to manually enter a command. You can't say "build me an app" and go to sleep.
This stands in stark contrast to Ralph's AFK mode. Ralph is designed for "start it and go to sleep" — the infinite bash loop keeps running until the task completes or fails. GSD requires you to be present at every critical checkpoint: approving the roadmap, answering discussion questions, triggering planning, launching execution, confirming verification results.
During his 4-hour livestream, TÂCHES was continuously typing commands: `new-project`, `discuss-phase 1`, `plan-phase 1`, `execute-phase 1`, `verify-work 1`, `discuss-phase 2`... Every transition required him to press Enter. This is not accidental — it's a deliberate design choice.
### A Deliberate Design Tradeoff
Ralph sacrificed planning capability for execution autonomy; GSD sacrificed execution autonomy for planning quality and human oversight. **This is a design tradeoff, not a deficiency.**
* **Ralph's advantage**: You can let it run through an entire feature while you sleep. But if the spec isn't good enough, it'll charge full speed in the wrong direction.
* **GSD's advantage**: You can course-correct after every phase. But you must be present throughout — you can't walk away.
What would the ideal look like? If GSD's discuss, plan, execute, and verify could be chained into an automated loop — like Ralph's bash loop but with GSD's structured planning and quality verification — that would be the best of both worlds. But no such tool exists yet. Perhaps that's the next direction worth exploring.
## Video Resources
The following videos can help you gain a more intuitive understanding of how GSD is used and what it can achieve.
## Final Thoughts
GSD represents a direction in the evolution of AI coding tools: from "let AI write code" to "let AI reliably deliver projects."
Ralph Wiggum proved a key insight — fresh context is more valuable than accumulated context. GSD builds on this foundation by adding project understanding (new-project), decision capture (discuss), structured planning (plan), parallel execution (execute), and quality verification (verify), forming a complete closed loop.
For solo developers and small teams, GSD's value lies in packaging complex engineering practices into a few simple commands. You don't need to understand subagent orchestration or XML prompt engineering — you just need to describe what you want, and let the system get it done.
Chase AI put it well: GSD is for people who "don't come from a technical background but still want to build projects end-to-end in Claude Code in a sustainable, repeatable way." And TÂCHES's livestream proved the point — someone who describes himself as "probably only able to write a Hello World HTML page on my own" used GSD to build a complete native desktop application.
This isn't magic. It's **putting the right complexity in the right place** — the system handles the complexity of orchestration while humans focus on creativity and decisions. And its limitations are equally worth respecting: GSD's choice to keep humans present at all times is both its constraint and the source of its reliability.
Ready to get hands-on? Continue to the [GSD Practice Guide](/en/docs/notes/gsd/practice) — covering complete command reference, configuration details, workflow walkthroughs, and FAQ.
***
**Related Reading**:
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Complete analysis of the Context Rot problem and the Ralph methodology
* [What is Spec-Driven Development](/en/docs/notes/speckit/concept) — The paradigm shift from Vibe Coding to spec-driven development
* [Claude Subagent Complete Guide](/en/docs/notes/claude-subagent) — Another approach to keeping context clean
* [Claude System Architecture Overview](/en/docs/notes/claude-architecture) — The overall architecture of Hooks, Subagents, and other components
* [My Claude Code Best Practices](/en/blog/claude-code-best-practices) — Day-to-day tips for using Claude Code
# Practical Guide
## Introduction
In the [previous article](/en/docs/notes/gsd/concept), we explored GSD's core principles — context engineering, subagent orchestration, goal-backward planning, and atomic commits. These concepts sound elegant, but there are many operational details between "understanding the theory" and "actually delivering a project."
In this article, we get hands-on. You'll learn GSD's complete command system, configuration options, output file structure, and how to use it to deliver a complete feature from scratch.
## Installation & Setup
### Installation
```bash
npx get-shit-done-cc@latest
```
The installer will prompt you to choose:
1. **Runtime** — Claude Code, OpenCode, Gemini CLI, or all
2. **Scope** — Global (all projects) or local (current project)
After installation, type `/gsd:help` in your runtime to verify.
### Recommended: Skip Permissions Mode
GSD is designed for frictionless automation. The recommended way to run Claude Code:
```bash
claude --dangerously-skip-permissions
```
If you'd rather not use this flag, you can configure fine-grained permissions in `.claude/settings.json`.
### Updates
```text
/gsd:update
```
GSD updates are very frequent (TÂCHES pushes 15–20 updates almost every day). Run this command regularly to stay on the latest version.
## Complete Command Reference
All GSD interactions use slash commands prefixed with `/gsd:`. Here's the complete reference organized by function.
### Core Workflow Commands
These five commands form GSD's main loop, used in sequence.
| Command | Description |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:new-project` | Initialize a project. The system asks questions until it understands your idea, then researches, extracts requirements, and creates a roadmap |
| `/gsd:discuss-phase [N]` | Discuss gray areas for phase N. Captures your implementation preferences to guide planning |
| `/gsd:plan-phase [N]` | Create atomic task plans for phase N. Includes research, planning, and verification sub-steps |
| `/gsd:execute-phase ` | Execute phase N. Subagents implement tasks in parallel, each with an independent commit |
| `/gsd:verify-work [N]` | Verify deliverables for phase N. Guides you through confirmation, auto-diagnoses issues |
> `[N]` indicates an optional parameter — the system auto-detects the current phase if omitted. `` indicates a required parameter.
### Milestone Management
| Command | Description |
| --------------------------- | ------------------------------------------------------------------------------------------- |
| `/gsd:audit-milestone` | Audit current milestone progress — checks all phase statuses, identifies incomplete items |
| `/gsd:complete-milestone` | Archive current milestone, tag version, prepare for the next cycle |
| `/gsd:new-milestone [name]` | Create a new milestone. Optionally provide a name; the system plans based on completed work |
### Phase Management
| Command | Description |
| --------------------------------- | ----------------------------------------------------------------------------------- |
| `/gsd:add-phase` | Add a new phase at the end of the roadmap |
| `/gsd:insert-phase [N]` | Insert an urgent phase at position N; subsequent phases auto-renumber |
| `/gsd:remove-phase [N]` | Remove a phase and cascade-delete all related output files |
| `/gsd:list-phase-assumptions [N]` | List all assumptions and dependencies for a phase, helping identify potential risks |
### Quick Mode & Tools
| Command | Description |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `/gsd:quick [--full]` | Quick mode — skips research, plan checking, and verification. Good for small tasks. `--full` enables full safeguards |
| `/gsd:debug [desc]` | Launch an isolated debug subagent. Optionally describe the issue; the system hypothesizes → gathers evidence → resolves |
| `/gsd:add-todo [desc]` | Capture an idea to the to-do list without modifying the roadmap |
| `/gsd:check-todos` | View the current to-do list |
| `/gsd:map-codebase` | Analyze an existing codebase — tech stack, architecture, conventions, potential issues |
### Session & Configuration Management
| Command | Description |
| ------------------ | --------------------------------------------------------------------------------- |
| `/gsd:pause-work` | Pause work. Saves current state to STATE.md for easy resumption |
| `/gsd:resume-work` | Resume work. Reads last state from STATE.md and continues where you left off |
| `/gsd:progress` | View overall project progress — completed phases, current position, pending items |
| `/gsd:help` | Display all available commands with brief descriptions |
| `/gsd:settings` | View and modify GSD configuration |
| `/gsd:set-profile` | Switch model profiles (quality / balanced / budget) |
| `/gsd:update` | Update GSD to the latest version |
## Configuration Details
### Model Profiles
GSD supports three model profiles, switchable via `/gsd:set-profile`:
| Profile | Planning | Execution | Verification | Best For |
| ------------------ | -------- | --------- | ------------ | --------------------------------------------------- |
| quality | Opus | Opus | Sonnet | Complex projects, critical features, first-time use |
| balanced (default) | Opus | Sonnet | Sonnet | Daily development, best balance for most scenarios |
| budget | Sonnet | Sonnet | Haiku | Simple features, budget-sensitive, rapid iteration |
### Core Settings
View and modify via `/gsd:settings`:
| Setting | Default | Description |
| ------------------------ | ---------- | --------------------------------------------------------------------- |
| `mode` | `balanced` | Model profile selection |
| `depth` | `standard` | Research depth: `quick` / `standard` / `deep` |
| `git.branching_strategy` | `feature` | Git branching: `feature` (per feature) / `phase` (per phase) / `none` |
### Workflow Toggles
The following agents can be individually toggled to trade off between speed and quality:
| Toggle | Default | Description |
| -------------- | ------- | ---------------------------------------------------------- |
| `research` | On | Whether to auto-research before planning |
| `plan_check` | On | Whether to auto-verify after plan creation |
| `verifier` | On | Whether to auto-verify after execution |
| `auto_advance` | Off | Whether to auto-advance to the next phase after completion |
> Disabling `research` and `plan_check` can significantly speed things up but may reduce planning quality. Consider disabling only after you're familiar with the project.
## Output File Structure
All GSD state and output is stored in the `.planning/` directory. Understanding this structure helps with debugging and manual intervention.
### Project-Level Files
| File | Purpose | Created By |
| ----------------- | ---------------------------------------------- | ----------------------------------- |
| `PROJECT.md` | Project vision and scope | `new-project` |
| `REQUIREMENTS.md` | Versioned requirements with phase traceability | `new-project` |
| `ROADMAP.md` | Phase planning and progress | `new-project` |
| `STATE.md` | Current state — decisions, blockers, position | `new-project`, continuously updated |
### Phase-Level Files
Each phase produces the following files (using Phase 1 as an example):
| File | Purpose | Created By |
| -------------------- | --------------------------------------------- | ----------------- |
| `01-CONTEXT.md` | Decisions from the discuss phase | `discuss-phase 1` |
| `01-RESEARCH.md` | Research findings and technical investigation | `plan-phase 1` |
| `01-01-PLAN.md` | First atomic task plan | `plan-phase 1` |
| `01-02-PLAN.md` | Second atomic task plan | `plan-phase 1` |
| `01-01-SUMMARY.md` | Execution record for the first plan | `execute-phase 1` |
| `01-02-SUMMARY.md` | Execution record for the second plan | `execute-phase 1` |
| `01-VERIFICATION.md` | Automated verification results | `execute-phase 1` |
| `01-UAT.md` | User acceptance testing record | `verify-work 1` |
### Directory Structure Example
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## Workflow Walkthrough
Let's walk through a complete workflow using "add a comment system to a blog" as an example.
### Step 1: Initialize the Project
```text
/gsd:new-project
```
The system will start asking questions:
```
> What do you want to build?
"I want to add a comment system to my Next.js blog. Support both
anonymous and authenticated comments, Markdown rendering, and an
admin panel. Tech stack: Prisma + PostgreSQL."
```
The more detailed your description, the fewer follow-up questions. TÂCHES recommends preparing a rough vision document describing what you want — no need for technical details.
After completion, the system produces four files and asks you to approve the roadmap. Once approved, building begins.
> **Already have a codebase?** Run `/gsd:map-codebase` first. The system will analyze your existing architecture and conventions, then `new-project` can plan based on your existing code.
### Step 2: Discuss Phase
```text
/gsd:discuss-phase 1
```
The system identifies gray areas and asks questions one by one:
```
> Comment nesting: support multi-level nesting or single-level replies only?
> Anonymous comments: require CAPTCHA or direct submission?
> Admin panel: need bulk operations or one-by-one moderation?
```
Every decision here directly affects planning quality. If unsure, let the system use defaults — but deeper discussion significantly reduces rework during execution.
### Step 3: Plan Phase
```text
/gsd:plan-phase 1
```
The system will:
1. Research how to implement a comment system with Prisma + PostgreSQL
2. Create 2–3 atomic task plans (e.g., data model, API routes, frontend components)
3. Auto-verify that the plans cover all requirements
Each plan is small enough to complete in a single fresh context window.
### Step 4: Execute Phase
```text
/gsd:execute-phase 1
```
The system begins wave execution:
* **Wave 1** (no dependencies): Database schema, Prisma models — executed in parallel
* **Wave 2** (depends on Wave 1): API routes, comment CRUD — executed in parallel
* **Wave 3** (depends on Wave 2): Frontend comment component — executed independently
Each task runs in a fresh 200k token context and gets an independent git commit.
### Step 5: Verify Phase
```text
/gsd:verify-work 1
```
The system guides you through confirmation:
```
> ✅ Database tables created
> ✅ API routes return correct status codes
> ❓ Can you see the comment input box below blog posts? [yes/no/describe issue]
> ❓ Does the page update in real-time after submitting a comment? [yes/no/describe issue]
```
If any checks fail, the system auto-diagnoses and creates a fix plan. Run `/gsd:execute-phase 1` again to execute the fix.
### Common Scenarios
**Insert an urgent phase**: Requirements changed — need to insert new work before the current phase.
```text
/gsd:insert-phase 2
```
Subsequent phases auto-renumber (original Phase 2 becomes Phase 3, and so on).
**Pause and resume**: Need to interrupt work to handle something else.
```text
/gsd:pause-work # Save current state
# ... handle other things ...
/gsd:resume-work # Resume from where you left off
```
**Roll back unsatisfactory results**:
```bash
git reset --hard HEAD~3 # Return to pre-execution state
```
```text
/gsd:remove-phase 2 # Cascade-delete all output files for this phase
```
TÂCHES demonstrated this multiple times in his livestream — if you don't like it, roll back. Clean and decisive.
## Debugging Workflow
When verification finds issues, or you encounter bugs during development, GSD provides a dedicated debugging workflow.
```text
/gsd:debug Page doesn't update in real-time after comment submission
```
The system launches an **isolated debug subagent** with the following workflow:
1. **Hypothesize** — Generate multiple possible root cause hypotheses based on the problem description
2. **Gather evidence** — Verify hypotheses one by one, checking code, logs, network requests
3. **Resolve** — After pinpointing the root cause, create a fix plan
Key features:
* **Context isolation**: The debug agent has its own context window and won't pollute your main development context
* **Documentation**: Creates independent debug documents tracking the entire investigation
* **Fix plan**: Produces a directly executable fix plan after diagnosis
This is far more efficient than debugging in the main context — debug information doesn't accumulate in your main window.
## Practical Tips
Drawing from TÂCHES's livestreams and Chase AI's hands-on experience, here are some practical recommendations.
### Slow Down to Speed Up
TÂCHES admits that when he first started using GSD, his mindset was "go go go" — but he later discovered that **spending more time in the research and discuss phases actually reduced rework during execution**. Newer versions of GSD added `research-project` and `define-requirements` steps specifically to get the direction right before writing any code.
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### Clear Context Between Phases
TÂCHES's habit is to **run `clear` between every phase** to keep the main context lean. He uses the Warp terminal, with each window full-screened (Command+Shift+Enter), executing the current phase in one window while researching the next phase in another.
### The Token Cost Tradeoff
GSD's subagent approach does consume more tokens than using Claude Code directly. But Chase AI makes a compelling argument: **"plan twice, prompt once" is cheaper in the long run than "prompt once, then patch and patch."** Doing it right in fresh context is far more efficient than repeatedly fixing things in a degraded one.
### Handling Unsatisfactory Results
If you're unhappy with a phase's results, you can `git reset --hard` and then use `/gsd:remove-phase` to cascade-delete all output files for that phase. TÂCHES demonstrated this live — he didn't like a particular visual effect, so he rolled back to the last satisfactory state, clean and decisive.
### The To-Do System
`/gsd:add-todo` lets you capture ideas to a to-do list at any time without modifying the roadmap. These ideas can be pulled up during `/gsd:discuss-milestone` as input for the next milestone. TÂCHES's strategy is "build the features first, polish the UI in milestone 2."
## FAQ & Best Practices
### Best Practices
**Provide detailed initial descriptions.** The quality of `/gsd:new-project` depends on the quality of your input. Prepare a rough vision document — describe goals, users, core features, known constraints. The more precise your description, the fewer follow-up questions and the better the planning.
**Clear context between phases.** After completing each phase, run `clear` or `/compact` to keep the main context window lean. TÂCHES's habit is to keep the main context at 30–40%.
**Test with Quick Mode first.** For uncertain small features, use `/gsd:quick` to test the waters. If it works well, incorporate it into the formal roadmap.
**Map existing codebases first.** Before using GSD on an existing codebase, run `/gsd:map-codebase`. The system will analyze your tech stack, architecture, and conventions, making subsequent planning more aligned with existing code.
### FAQ
**Q: What runtimes does GSD support?**
A: Claude Code, OpenCode, and Gemini CLI. You can choose one or all during installation.
**Q: What's the difference between Quick Mode and full mode?**
A: Quick Mode provides GSD's basic safeguards (atomic commits, state tracking) but skips research, plan checking, and verification. Ideal for bug fixes, small features, and config changes that don't need full planning.
**Q: Can I pause mid-execution?**
A: Yes. `/gsd:pause-work` saves the current state to STATE.md. Next time you run `/gsd:resume-work`, the system continues from where you left off.
**Q: How do I control token costs?**
A: Three approaches — (1) Switch to `budget` profile: `/gsd:set-profile budget`; (2) Disable `research` or `plan_check` agents; (3) Use `/gsd:quick` for simple tasks.
**Q: Can GSD and Ralph be used together?**
A: Yes. GSD and Ralph solve different problems — GSD handles planning and structured execution, Ralph handles autonomous loop execution. You can use GSD's `new-project` and `plan-phase` to generate a complete plan, then use Ralph loops to execute phases that don't require human intervention.
**Q: What about multi-person collaboration?**
A: The `.planning/` directory can be committed to Git. Multiple people can run different phases and merge results via Git. However, avoid running the same phase simultaneously.
## Summary
GSD's core value lies in **hiding complexity in the system while keeping simplicity for the user**. You only need a few commands — `new-project`, `discuss-phase`, `plan-phase`, `execute-phase`, `verify-work` — while the system handles all context management, subagent orchestration, and quality verification behind the scenes.
From installation to delivery, GSD provides a clear path: describe what you want → discuss implementation details → generate atomic plans → execute in parallel → verify deliverables. Every step gives you the opportunity to intervene, and every step is documented.
This isn't "press one button and it's done" magic. It's a system that needs your participation but shoulders most of the cognitive load. As TÂCHES puts it: you're the high-level project manager, GSD is your execution team.
***
**Further Reading**:
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — Core principles, workflows, and technical architecture
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Context Rot and the Ralph methodology
* [snarktank/ralph Practice Guide](/en/docs/notes/ralph-wiggum/snarktank) — Ralph installation, PRD writing, and hands-on guide
* [What is Spec-Driven Development](/en/docs/notes/speckit/concept) — From Vibe Coding to spec-driven development
* [Speckit Practice Guide](/en/docs/notes/speckit/practice) — Speckit command reference and complete examples
# gstack: When YC CEO puts entrepreneurial experience into Claude Code
## Introduction
In previous notes, we explored various "enhancement solutions" in the Claude Code ecosystem from the infinite loop of [Ralph Wiggum](/en/docs/notes/ralph-wiggum/concept) to the specification-driven development of [GSD](/en/docs/notes/gsd/concept). They are all trying to answer the same question: \*\*How to change AI programming from "adaptation" to "reliable delivery"? \*\*
Ralph's answer is "restart everything" - use a new process each time to avoid context rot. GSD's answer is "Specification Driven" - ensuring quality through structured phase planning and validation cycles. But what if you want not just an execution system, but a complete virtual engineering team? The CEO makes product decisions, the engineering manager reviews the architecture, the designer controls the experience, the QA runs real browser tests, and the release engineer manages the launch...all are played by AI and are commanded by you.
This is the core idea of gstack.
## What is gstack
**Garry Tan**, the creator of gstack, has a rich technical and entrepreneurial background - he started writing code at the age of 14, graduated from Stanford Computer Engineering, is the 10th employee of Palantir, co-founded Posterous (later acquired by Twitter), and has served as President & CEO of Y Combinator since 2023.
He used gstack to release more than 600,000 lines of production code (35% testing) in 60 days, averaging more than 10,000 lines per day - while still running YC full-time. One of the projects, garylist.org, was launched in 21 days, with 150,000 lines of code and 35% test coverage. According to his own words, the code quality exceeds the previous entrepreneurial project he spent $5 million, two years, and 10 engineers on.
Since the project was open sourced on March 11, 2026, it iterated from v0 to v0.15.1.0 within 3 weeks, and GitHub has received 60,500+ stars. MIT license, completely open source.
## The position of gstack in the tool ecosystem
| Dimensions | Native Claude Code | Ralph Wiggum | GSD | SpecKit | Superpowers | **gstack** |
| ---------------------- | ----------------------------- | ----------------------- | --------------------------------------------- | ------------------------------------- | ----------------------------- | ------------------------------------------ |
| Core Positioning | Universal AI Coding Assistant | Infinite Loop Iteration | Contextual Engineering + Specification Driven | Requirements → Specifications → Tasks | Process Discipline + TDD | **Role-Based Virtual Team** |
| Core Pattern | Conversational Programming | Bash Loop + New Process | Phase-based Roadmap | Spec → Plan → Tasks | Strict Development Pipeline | **Sprint Seven-Step Process** |
| Human involvement | Live conversations | Hands-off (AFK) | Verification per stage | Spec approval | Validation per step | **Role review per stage** |
| Unique capabilities | Basic coding | Unlimited iteration | Context Rot management | Requirements tracing | Forced TDD | **Browser automation + multi-role review** |
| Suitable for scenarios | Simple tasks | Continuous iteration | Large-scale project management | Projects with rigorous requirements | Engineering quality assurance | **Full-process product development** |
A key pattern can be seen from the table: \*\*These tools do not compete with each other, but solve AI programming problems in different dimensions. \*\*
Superpowers uses **process discipline** to ensure code quality (mandatory TDD, structured dialogue, implementation plan); GSD uses **context engineering** to manage complex projects (phase planning, sub-agent fresh context, file system status); gstack uses **role decomposition** to improve decision-making quality (CEO perspective reviews products, engineering managers review architecture, QA runs real browsers).
To put it simply, Superpowers is based on process guardrails, and gstack is based on role design—the former is suitable for project implementation from 1 to N, and the latter is suitable for product construction from 0 to 1. \*\*The two are complementary rather than competing products. \*\*
## Core workflow: The Sprint seven steps
gstack organizes the entire development process into a cycle of **Think → Plan → Build → Review → Test → Ship → Reflect**, called "The Sprint" - not an agile Sprint, but a development rhythm of "roles appear in sequence".
### 1. Think — Product Clinic
```text
/office-hours
```
This is the most distinctive skill of gstack. The inspiration comes directly from YC’s Office Hours – entrepreneurs go to meet YC partners and undergo soul-searching. The AI will ask you **6 forcing questions**:
1. Who specifically needs this?
2. What if they don’t have it today?
3. Why is this matter urgent now?
4. How do you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can release?
The purpose is not to help you write code, but to **re-examine the problem itself** before writing code.
### 2. Plan — Multi-role review
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
CEO Review is essentially "Founder Mode" - instead of executing requirements literally, you step back and ask "What is the true purpose of this product?" It supports four modes: Expand Scope, Selective Expand, Maintain Scope, and Reduce Scope.
### 3. Build — coding implementation
Start coding according to the approved plan. This step uses standard Claude Code capabilities.
### 4. Review — Parallel expert review
```text
/review
```
This skill dispatches **7 parallel sub-agents** at one time to review the code from 7 perspectives: testing, maintainability, security, performance, data migration, API contract, and red team attack. Obvious problems will be automatically fixed.
### 5. Test — Real browser QA
```text
/qa
```
Not a practice test. The QA skill launches a **real headless Chromium browser**, opens your app, clicks buttons, fills out forms, and takes screenshots - just like a real tester would. Automatically fix bugs, generate regression tests, and re-verify after bugs are discovered.
### 6. Ship — one-click publishing
```text
/ship
```
Automatically sync the master branch, run tests, review diffs, update version numbers and CHANGELOG, commit, push, create PRs. If the project doesn't have a testing framework, it will even build one first.
### 7. Reflect — review and learn
```text
/retro
```
Engineering manager-style weekly report: analyze commit history, test ratio, and code quality trends. Support multi-person team analysis and track indicators such as "number of consecutive release days".
## Why it works: Technical principles
### Browse Daemon: Put eyes on AI
gstack's most unique technical contribution is the Browse Daemon - a persistent headless Chromium instance that communicates over localhost HTTP. The first call launches the browser (\~3 seconds), and each subsequent command takes only 100-200ms. This means that the AI can actually see your app, rather than guessing the DOM structure.
It also introduces **Ref System** (element reference `@e1`, `@e2`) to locate elements through the accessibility tree without writing CSS selectors. This is a "truly technical contribution" that is generally recognized by the community (including critics).
### Role breakdown: not an agent, but a team
What gstack does is to disassemble all roles into independent prompt files, allowing Claude Code to switch to the perspectives of different roles at different stages to review the code. This is essentially a refined prompt engineering.
The core insight is: \*\*Planning does not equal review, review does not equal release, and founder taste and engineering rigor are completely different modes of thinking. \*\* Instead of having a general agent do everything, switch "brain modes" when needed - founder thinking, engineering rigor, paranoid review, fast execution.
### Three major philosophies
gstack's ETHOS.md records three core concepts:
1. **Boil the Lake**: When AI drives the marginal cost of completeness to zero, always choose a complete implementation - 100% test coverage, all edge cases, all error paths. "Release shortcuts" are old-time thinking.
2. **Search Before Building**: Three layers of knowledge - time-tested patterns, new and popular solutions, and first principles. Start by understanding what everyone is doing, questioning their assumptions, and discovering why the usual solutions are wrong.
3. **User Sovereignty**: AI recommendation, human decision-making. Even if two AI models reach consensus, the user’s judgment still takes precedence—because the user has domain knowledge, strategic perspective, and taste.
## The boundaries and controversies of gstack
Community reaction to gstack is probably the most polarizing of any AI programming tool.
**The bright side**: Founders and non-technical builders generally agree, especially "product thinking" skills like `/office-hours` and `/plan-ceo-review`, have helped many independent developers re-examine the product direction before starting to code. Engineering review (`/review`) can indeed discover some hidden security vulnerabilities. This multi-angle parallel review model has practical value.
The **questioning side** is also very direct:
* **LOC indicator is of little significance**: 600,000 lines of code in 60 days. The number of lines of code is never a quality indicator. A large amount of code may be just scaffolding and boilerplate.
* **Essentially a prompt template**: Each skill is a SKILL.md file, and the technical threshold is not high. The real value is not in the file itself, but in the quality of the prompt's design.
* **Limitations of AI self-review code**: `/review` Letting AI review the code written by AI is equivalent to correcting your own homework. Multi-role parallelism can alleviate this problem, but it is still the same model.
* **Celebrity effect bonus**: If the founder is not the YC CEO, there is a high probability that this project will not receive such high attention.
**My opinion**: Controversies aside, the really valuable parts of gstack are two--the browser automation technology of Browse Daemon, and the design pattern of role decomposition. None of this depends on who Garry Tan is. The core significance of roleization is not at the technical level, but at the behavioral level - it helps you organize your AI workflow more consciously, rather than throwing everything at a general agent.
gstack is suitable for forking and customizing. You can get the skills you need and change the prompts you want, rather than copying them all.
## Video resources
## Write at the end
gstack represents an interesting direction for AI programming tools: not to make AI more autonomous (Ralph's route), nor to make the process more rigid (Superpowers' route), but to let AI play different roles to improve the quality of decisions. Its controversy just illustrates the richness of the AI programming ecosystem—no one solution fits everyone.
If you are interested in gstack, the next step is to read [Practical Chapter](/en/docs/notes/gstack/practice) - a step-by-step tutorial from installation to running through the complete workflow.
***
**Related Reading**:
* [Introduction to GSD concepts](/en/docs/notes/gsd/concept) — Another structured AI programming solution
* [Ralph Wiggum in-depth analysis](/en/docs/notes/ralph-wiggum/concept) — Understand the starting point of infinite loop iteration
* [Claude Skills Concept](/en/docs/notes/claude-skills/concept) — Understand the underlying mechanism of Skills
# gstack front-end skill panorama: AI workflow from design to launch
## Introduction
In previous notes, we talked about [What is gstack](/en/docs/notes/gstack/concept), [How to run through the workflow](/en/docs/notes/gstack/practice), and [Skill’s engineering structure](/en/docs/notes/gstack/skill-architecture). But there is one question that has not been discussed - among the 60+ Skills brought after gstack installation, which ones are related to front-end/UI design? In what order? \*\*
This note does two things: first, classify the \~27 skills related to the front end by function, and then use an interesting small project - the countdown anniversary page - to walk through the complete workflow from beginning to end, with screenshots at each stage, so that you can see the actual effect.
## Front-end Skill Kit Panorama
gstack's front-end skills can be divided into 6 functional layers, from foundation to roof, with each layer solving problems at different stages.
### Design Infrastructure
A one-time setting at the project level to determine the design language, and all subsequent Skills will refer to these benchmarks.
| Skill | What to do | When to use |
| ---------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| `/design-consultation` | Complete design system consultation, output color matching, fonts, spacing, texture direction | When starting a new project, or if you want to redefine the visual style |
| `/teach-impeccable` | Collect design preferences at one time and write them into the AI configuration file | Run once after installing gstack to let AI remember your aesthetics |
| `/brand-guidelines` | Apply existing brand color matching and font specifications | Apply directly when there is an existing brand manual |
> If the project already has `DESIGN.md`, this level can be skipped.
### Design Exploration
When you are unsure about the direction, quickly compare multiple options.
| Skill | What to do | When to use |
| ------------------ | ---------------------------------------------------------------- | ----------------------------------------------------------- |
| `/design-shotgun` | Generate 3-5 visual solutions, open the comparison panel | Not sure what style you want, want to see the possibilities |
| `/frontend-design` | Generate recognizable, production-level front-end interface code | Work directly after the direction is clear |
| `/canvas-design` | Generate posters, visual art (PNG/PDF) | Requires static visual design rather than web components |
### Design implementation
Turn the plan into truly runnable code and handle typesetting, layout, and responsiveness.
| Skill | What to do | When to use |
| ------------------------ | -------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| `/design-html` | Convert the confirmed design draft into production-grade HTML/CSS | I have mockups that I want to implement directly |
| `/mobile-responsiveness` | Mobile-first responsive layout and touch interaction | Mobile adaptation from scratch |
| `/adapt` | Breakpoint adaptation across devices and screen sizes | There is a desktop version and needs to be adapted to mobile phones/tablets |
| `/typeset` | Font selection, level, size, thickness, and readability optimization | Text layout looks "almost meaningless" |
| `/arrange` | Layout spacing, visual rhythm, alignment repair | Inconsistent spacing, layout feels crowded or scattered |
### Design enhancements
On the basis of functional completion, inject dynamic effects, personality and emotional details.
| Skill | What to do | When to use |
| ------------ | ---------------------------------------------------------------------------------- | ------------------------------------------------- |
| `/animate` | Add purposeful micro-interactions and animations | Page functionality OK but feels "rigid" |
| `/delight` | Add surprise details and personalized touches | Want users to remember this page |
| `/bolder` | Amplify visual impact | Design is too plain and too safe |
| `/colorize` | Add strategic color to the monotonous interface | The page is too gray, too plain, and lacks warmth |
| `/overdrive` | Technical explosion-level effects - shader, spring physics, scroll drive animation | A certain area wants wow effect |
| `/onboard` | New user guidance process, empty state design | First-time user experience |
These four enhancement skills are in a **progressive relationship**: `animate` is the basic dynamic effect, `delight` is emotional, `bolder` is amplification, and `overdrive` is explosion. Stack them step by step according to project needs, it is not necessary to use them all.
### Design optimization
Convergence and refinement – removing excess, aligning deviations, and polishing rough edges.
| Skill | What to do | When to use |
| ------------ | ----------------------------------------------------- | ----------------------------------------------------------------- |
| `/polish` | Final quality polish: alignment, spacing, consistency | One final pass before release |
| `/quieter` | Reduce the intensity of visual stimulation | The design is too fancy and noisy |
| `/distill` | Minimalize and remove unnecessary complexity | There are too many elements on the page and I want to reduce them |
| `/normalize` | Align design system standards (token, spacing, color) | Style deviates from DESIGN.md specifications |
| `/clarify` | Improve UX copywriting, error messages, label wording | Copywriting is confusing, error messages are unfriendly |
### Design Review and Verification
Systematic inspection before going online, finding problems, scoring, and fixing them.
| Skill | What to do | When to use |
| --------------------- | -------------------------------------------------------------------------- | -------------------------------------------------------------- |
| `/plan-design-review` | Design plan review before implementation (0-10 score) | I want AI to review the plan from a designer's perspective |
| `/design-review` | Visual QA after implementation, automatic screenshot comparison and repair | After the code is written, check the visual restoration degree |
| `/critique` | UX Assessment: Visual Hierarchy, Cognitive Load, Emotional Resonance | Want a structured design review report |
| `/audit` | Technical review: accessibility, performance, themes, responsiveness | Systematic checks before going live |
| `/benchmark` | Performance baseline testing, before/after comparison | Want to quantify the impact of changes on performance |
## Practical demonstration: Use the countdown anniversary page to go through the entire process
Just looking at the classification table is too abstract. We use a small project to string together the above skills - making a **countdown/anniversary single page**: choose a meaningful date and create a countdown display with digital animation and background effects.
This project is small but complete, just enough to cover most of the 6 levels of skills. The complete process is:
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> You don’t have to run all 7 steps every time. After you become proficient, the commonly used links are only `/frontend-design → /animate → /polish → /ship` four steps. In order to show the complete ability, every step is taken here.
### Stage 1: Infrastructure - Determine the design language
**Skill**:`/design-consultation` + `/teach-impeccable`
It only needs to be done once at the beginning of the project. Outputs `DESIGN.md`, which lets the AI remember your design preferences. If the project already has `DESIGN.md`, skip it directly.
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* TODO: Screenshot — DESIGN.md fragment produced by design-consultation */}
### Stage 2: Exploration - Comparison of multiple options
**Skill**:`/design-shotgun`
When unsure of the direction, let AI generate 3-5 visual solutions and open the comparison panel to choose.
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* TODO: Screenshot — 3 solution comparison panels generated by design-shotgun */}
Choose a direction from 3 options. If you know exactly what you want, skip this step and go directly to Stage 3.
### Stage 3: Build - produce production-level code
**Skill**:`/frontend-design` + `/adapt`
core link. code while ensuring responsiveness is in place from the start.
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* TODO: Screenshot - the desktop page effect after the construction is completed */}
{/* TODO: Screenshot - mobile version effect (after /adapt adaptation) */}
### Stage 4: Enhance – Inject motion and personality
**Skill**: `/animate` → `/delight` (on demand `/overdrive`)
These three are in a progressive relationship: `animate` is the basic dynamic effect, `delight` is the emotional detail, and `overdrive` is the explosive effect. Add layers as needed.
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
Note the animation constraints referenced to `DESIGN.md` - if the design system only allows a 150ms hover transition, `/overdrive` will not apply. This is a good judgment exercise.
{/* TODO: Screenshot or GIF - before vs after motion enhancement */}
### Stage 5: Tuning - Convergence and Polishing
**Skill**: `/typeset` + `/polish` (on demand `/distill`, `/normalize`)
Spacing alignment, font hierarchy, visual rhythm. If you find you've added too much, use `/distill` for subtraction.
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* TODO: Screenshot — detailed comparison of polish before vs after polish */}
### Stage 6: Review – Systematic Check
**Skill**:`/design-review` + `/audit`
Visual QA + technical review. `/design-review` will automatically take screenshots to compare and fix issues, `/audit` will check accessibility and performance.
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* TODO: Screenshot — the scoring report produced by audit */}
### Stage 7: Release
**Skill**:`/ship`
Standard gstack release process - testing, diff review, creating PR.
```text
> /ship
```
***
**Expected results**: A visually exquisite countdown page, using 8-10 front-end skills in the process. What's more important is to establish the intuition of "what skill to use at what stage".
## Daily cheat sheet
The above is the complete process. If you encounter specific problems in daily development, just check this table:
| My current question | What to use |
| ------------------------------------------------------- | ---------------------------------------------- |
| Don’t know what style I want | `/design-shotgun` |
| The page function is good but it feels "almost useless" | `/polish` |
| I feel like something is wrong but I can’t explain | `/design-review` |
| The font/layout looks awkward | `/typeset` |
| Messy spacing and crowded layout | `/arrange` |
| Style deviates from the design system | `/normalize` |
| Want to extract public components | `/extract` |
| The page is too complex and I want to subtract | `/distill` |
| The design is too plain and too safe | `/bolder` or `/colorize` |
| The design is too fancy and noisy | `/quieter` |
| The error message text is not friendly | `/clarify` |
| There is a problem with the display on the mobile phone | `/adapt` |
| Want to add animation effects | `/animate` (basic) or `/overdrive` (explosion) |
| Systematic inspection before going online | `/audit` |
| debug | `/investigate` |
## Summary
This note does two things:
1. **Panorama** - gstack’s 27 front-end skills are classified into 6 layers (infrastructure→exploration→implementation→enhancement→optimization→review)
2. **Practical Demonstration** - Use a countdown anniversary page to walk through the complete workflow, showing what Skills are used at each stage and why
Key takeaway: The most powerful usage of these Skills is not to call them individually, but to combine them in a pipeline - exploring directions, building implementations, enhancing polishing, and reviewing releases, with clear Skill selection at each stage.
But don’t be tied down by the process – once you’re proficient, `/frontend-design → /animate → /polish → /ship` four steps is enough most of the time.
***
**Related Reading**:
* [gstack Concepts](/en/docs/notes/gstack/concept) — What is gstack and what problems does it solve?
* [gstack Practical Chapter](/en/docs/notes/gstack/practice) — Complete workflow from installation to runthrough
* [gstack Skill Architecture Teardown](/en/docs/notes/gstack/skill-architecture) — What can Skill developers learn?
* [Claude Skills Concept](/en/docs/notes/claude-skills/concept) — Understand the underlying mechanism of Skills
# gstack practice: complete workflow from installation to runthrough
## Introduction
In [Concept](/en/docs/notes/gstack/concept), we learned about the core positioning of gstack - a role-based skill set that turns Claude Code into a virtual engineering team, and its differentiated positioning in the AI programming tool ecosystem compared to GSD, Superpowers, Ralph and other solutions.
This practical article focuses on **how to use**: from installation and configuration to running through the complete workflow, helping you get started with gstack in 30 minutes.
## Installation and configuration
### Preconditions
* **Claude Code** is installed and available
* **Git** installed
* **Bun v1.0+** installed (gstack is built on Bun)
* Windows users also need Node.js
### Global installation (recommended, completed in 30 seconds)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
The installation script does three things:
1. Add gstack’s skill information to your `CLAUDE.md` file
2. Put all skill files into the skills directory
3. Install Playwright and the corresponding Chromium browser (for `/browse` and `/qa`)
### Project-level installation (team sharing)
If you want team members to automatically obtain gstack after cloning the repository:
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
\###Multi-Agent support
gstack is not limited to Claude Code, and currently supports **10 AI programming Agents**. `./setup` automatically detects installed hosts by default:
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
The skill installation path of each host is in the shape of `~/./skills/gstack-*/` and does not interfere with each other.
> 💡 **Extra options for OpenClaw users**: In addition to calling through ACP, OpenClaw can also directly install 4 native methodology skills (`gstack-openclaw-office-hours`, `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`, `gstack-openclaw-retro`) through ClawHub, which can be used conversationally without a Claude Code session.
### Team Mode (Team Sharing + Automatic Updates, Recommended)
v1.x introduces Team Mode: each developer installs gstack globally, and the warehouse only records "we use gstack", and updates occur automatically:
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
Replacing `required` with `optional` is a "gentle reminder" rather than mandatory. Every time you start Claude Code, it will automatically run an update check (throttling once per hour, safe and silent if the network fails). There are no vendored files in the warehouse, and there is no version drift.
### Update
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
Or use `/gstack-upgrade` directly in Claude Code.
## Complete command reference
### Sprint Process
| Command | Role | Description |
| --------------------- | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/office-hours` | YC Office Hours | 6 forcing questions to reconstruct product direction and generate design documents |
| `/plan-ceo-review` | CEO / Founder | Looking for 10-star products, available in four range models |
| `/plan-eng-review` | Engineering Manager | Lockdown Architecture, Data Flow, Edge Cases, Test Matrix |
| `/plan-design-review` | Senior designer | Design dimension 0-10 score, explain how to achieve 10 points |
| `/plan-devex-review` | Developer Experience Leader | Explore developer portraits, benchmark TTHW, and design magic moments; three modes (DX EXPANSION / POLISH / TRIAGE), 20-45 forcing questions |
| `/autoplan` | Review pipeline | Automatically run CEO → Design → Engineering → DX review in sequence, automatically decide according to coding decision-making principles, and only throw "taste decisions" to you |
### Design
| Command | Description |
| ---------------------- | -------------------------------------------------------------------------------- |
| `/design-consultation` | Build a complete design system from scratch and generate DESIGN.md |
| `/design-shotgun` | Generate multiple AI design variants and compare selections in the browser |
| `/design-html` | Generate production-grade HTML/CSS, support React/Svelte/Vue framework detection |
### Review and Security
| Command | Role | Description |
| ---------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/review` | Staff Engineer | Find bugs that can pass CI but will explode in production, automatically fix obvious problems, and mark integrity gaps |
| `/investigate` | Debugging expert | Systematic root cause debugging. Iron rule: Don’t fix the bug until you find the root cause; stop after 3 failed fixes |
| `/design-review` | Designer who can write code | Visual audit + automatic repair, atomic submission, before and after comparison screenshots |
| `/devex-review` | DX tester | Really run onboarding: browse documents, run entry process, timing TTHW, screenshot errors, compare with `/plan-devex-review` score |
| `/cso` | Security Officer | OWASP Top 10 + STRIDE threat modeling, 17 false positive exclusion rules, 8/10 confidence threshold, each finding is accompanied by specific utilization scenarios |
### Testing and QA
| Command | Description |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/qa` | Open the real browser test and find the bug → Atomic commit fix → Generate regression test → Re-verify |
| `/qa-only` | Same as above but only reporting, no code modifications |
| `/benchmark` | Baseline performance test: page loading, Core Web Vitals, resource size, support before and after comparison |
| `/browse` | \~100ms level browser commands, real Chromium, screenshots, form filling, element clicks |
| `/open-gstack-browser` | Start GStack Browser: visible AI control Chromium, comes with sidebar extension, anti-crawling stealth, automatic model routing (Sonnet operation/Opus analysis), supports one-click cookie import |
| `/setup-browser-cookies` | Import cookies from real browsers (Chrome/Arc/Brave/Edge) to headless sessions to test login-required pages |
| `/pair-agent` | Cross-AI Agent browser pairing: share the same GStack Browser to OpenClaw / Hermes / Codex / Cursor, etc., each Agent has an independent tab, comes with ngrok tunnel to support remote Agents, scope token + tab isolation + rate limit + behavior attribution |
### Release and operation and maintenance
| Command | Description |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/ship` | Synchronize the main branch → Run tests → Audit coverage → Update version → Submit push → Create PR; Automatic bootstrap when the project does not have a test framework |
| `/land-and-deploy` | Merge PR → Wait for CI → Deploy → Verify production environment health |
| `/canary` | Post-deployment canary monitoring: console errors, performance regressions, page failures |
| `/setup-deploy` | `/land-and-deploy` One-time configuration: auto-detection platform (Fly.io/Render/Vercel/Netlify/Heroku/GitHub Actions/custom) + production URL + deployment command |
| `/setup-gbrain` | Get started with GBrain database in one click (within 5 minutes): PGLite local, Supabase existing URL, or automatically create a new Supabase project through Management API; MCP registration + warehouse-level read-write/read-only/deny permissions |
### Review and learn
| Command | Description |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/retro` | Team perception weekly report: per capita teardown, winning streak statistics, test health trends, growth opportunities; `/retro global` across all projects + AI tools (Claude Code / Codex / Gemini) |
| `/document-release` | Automatically update project documentation to match published code (README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS); `/ship` is now automatically called |
| `/learn` | Manage cross-session learning memories: view, search, prune, export, accumulate by project |
| `/context-save` `/context-restore` | Continuous checkpoint mode package: automatic WIP commit to save context, use `/context-restore` to rebuild the session after crash/switch |
### Security Protection
| Command | Description |
| ----------------------- | ----------------------------------------------------------------- |
| `/careful` | Dangerous operation warning: rm -rf, DROP TABLE, force-push, etc. |
| `/freeze` / `/unfreeze` | Lock/unlock editing scope to specific directory |
| `/guard` | `/careful` + `/freeze` combination, highest security mode |
| `/checkpoint` | Save/restore working status snapshot |
### Tool integration
| Command | Description |
| -------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/codex` | OpenAI Codex CLI integration: independent code review (pass/fail gate), confrontation mode, consultation mode; cross-model overlap analysis will be given after running with `/review` |
| `/health` | Code quality dashboard: tsc + biome + knip + shellcheck + tests → 0-10 overall score |
| `/skillify` | Consolidate the current workflow into a reusable skill |
| `/scrape` | Web scraping workflow |
| `/landing-report` | Landing page performance and experience report |
| `/make-pdf` | Generate PDF document |
| `/benchmark-models` `/model-overlays` `/plan-tune` | Cross-model comparison, coverage overlay, plan optimization |
### Standalone CLI(v0.19+)
In addition to the slash command, gstack also comes with a set of standalone CLIs (not run within the Claude Code session):
| Command | Description |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `gstack-model-benchmark` | Cross-model evaluation: run Claude / GPT (via Codex CLI) / Gemini at the same prompt, compare delay, token, cost and (optional) LLM-judge quality score; unavailable provider automatically skips |
| `gstack-taste-update` | Design taste learning: write `/design-shotgun`'s approval/disapproval into the project-level taste file, decay by 5% every week, and feed back to subsequent variant generation |
## Configuration details
### CLAUDE.md Add content
After installation, gstack will add a list and short description of all available skills to your `CLAUDE.md`. This lets Claude Code know which commands are available.
### Skill directory structure
The main entrance is the top-level `~/.claude/skills/gstack/SKILL.md`, each subcommand exists in the form of a flat directory, and the core is the `SKILL.md` file:
```text
~/.claude/skills/gstack/
├── SKILL.md # 主入口 skill
├── browse/ # 浏览器 daemon
├── qa/ # QA 测试
├── review/ # 代码审查
├── ship/ # 发布流程
├── plan-ceo-review/ # CEO 审查
├── office-hours/ # 产品门诊
├── pair-agent/ # 跨 Agent 浏览器配对
├── open-gstack-browser/ # GStack Browser 启动器
├── setup-gbrain/ # GBrain 数据库一键上手
├── hosts/ # 10 个 host 配置(claude/codex/cursor/...)
├── bin/ # standalone CLI(gstack-model-benchmark 等)
└── ... # 当前 v1.x 共 50 个 skill 目录
```
You are free to modify any `SKILL.md` to customize the behavior - this is the advantage of "fork and customize".
### Browse Daemon
Browse Daemon is a permanent Chromium instance. Key configuration:
* **Port**: Randomly selected 10000-60000, supports 10+ parallel workspaces
* **Security**: Only bind localhost, use bearer token authentication for each session
* **Cookie**: Use `/setup-browser-cookies` to import from Chrome/Arc/Brave/Edge
## Practical workflow demonstration
The following demonstrates a typical gstack workflow. The commands and output are based on real cases in the documentation and videos.
> 💡 **Note**: The following output is a general example compiled based on research. Screenshots of specific projects will be added in the future based on actual practice.
### Step 1: Product clinic
```text
> /office-hours
[YC Office Hours] 6 forcing questions:
1. Who specifically needs this?
2. What do they do today without it?
3. Why is this urgent right now?
4. How will you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can ship?
→ Design doc generated
```
Don’t rush to write code, first let AI torture your ideas from the perspective of YC Office Hours.
### Step 2: Multi-role review plan
```text
> /autoplan
[CEO Review] Finding the 10-star product...
[Design Review] Rating dimensions 0-10...
[Eng Review] Locking architecture + edge cases...
→ Fully reviewed plan ready
```
`/autoplan` automatically runs three rounds of CEO → Design → Engineering reviews to produce a complete post-review plan.
### Step 3: Coding implementation
Code normally according to the approved plan. You can use the standard Claude Code conversation.
### Step 4: Multi-expert code review
```text
> /review
Dispatching 7 specialist reviewers...
- Testing coverage ✓
- Maintainability ✓
- Security: Found 1 issue (auto-fixing)
- Performance ✓
- Data migration ✓
- API contract ✓
- Red team: No vulnerabilities found
→ Review complete, 1 auto-fix applied
```
### Step 5: Browser QA
```text
> /qa
Opening headless browser...
Testing user flows:
- Login flow ✓
- Dashboard load ✓
- Form submission: Bug found → fixing → re-testing ✓
- Image upload ✓
→ 4 flows tested, 1 bug fixed, regression test generated
```
### Step 6: Publish
```text
> /ship
Syncing with main...
Running tests: 42 passed, 0 failed
Reviewing diff: 3 files changed
Updating VERSION: 1.2.0 → 1.3.0
Creating PR: "Add screenshot feature"
→ PR #47 created, ready for merge
```
## Practical Tips and Community Experience
### Garry Tan’s suggestion
ETHOS.md from gstack, three core principles:
1. **Boil the Lake**: AI makes completeness almost free—always do complete things and don’t take shortcuts
2. **Search Before Building**: Search first, understand first, and then start after three-layer knowledge verification
3. **User Sovereignty**: AI recommendation, you decide. Even if both AI models agree, your judgment still takes precedence
The README of gstack begins with a quote from Karpathy - this is also the starting point for Garry Tan himself to explain why he wants to build gstack:
### Positive community experiences
* **`/office-hours` for YC applications**: Multiple S26 applicants on Reddit r/ycombinator reported that using gstack’s office hours to stress test their application materials is very effective.
* **Security audit found real vulnerabilities**: There was CTO feedback `/review` discovered an XSS vulnerability that the team was not aware of.
* **`/browse` Real browser testing**: Recognized by the community (including critics) as a "truly technical contribution"
### Common pitfalls
* **Frequent permission prompts**: Some users reported that "permission prompts have to be approved every 30 seconds, making it impossible to sleep." It is recommended to configure appropriate automatic approval rules in Claude Code settings
* **High Token consumption**: Characterized prompts will increase context consumption. If you are cost-sensitive, you can selectively use the skills you need most
* **Agent Loop**: There are cases on HN where users reported that the agent got stuck in a 70-minute loop. It is recommended to set reasonable timeouts and checkpoints
* **Not for everyone**: Experienced developers may feel that most skills are unnecessary wrappers. gstack is more suitable for **independent founders and small teams** rather than teams with mature engineering processes
## Frequently Asked Questions and Best Practices
\*\*Q: Can gstack and Superpowers be used at the same time? \*\*
Yes. The two complement each other - Superpowers is good at process discipline and TDD assurance, and gstack is good at product thinking and multi-role reviews. Many teams use Superpowers for daily coding discipline and gstack for product planning and QA.
\*\*Q: Is Token expensive? \*\*
Higher than native Claude Code. Each skill's role prompt occupies the context window. But if your time is worth more than the token fee, this is usually a good deal.
\*\*Q: What type of projects is it suitable for? \*\*
Best suited for **full-process product development** – from idea to launch. If you just fix bugs or make small features, native Claude Code is enough. The value of gstack is maximized in the "complete process".
\*\*Q: How to customize skill? \*\*
Each skill is a `SKILL.md` file. Just edit it directly:
1. Find the skill directory: `~/.claude/skills/gstack//`
2. Edit `SKILL.md`
3. Rerun `./setup`
The community recommends forking the repository and customizing it instead of directly modifying the global installation.
### Best Practices
1. **First `/office-hours` then code**: Make it a habit to do product clinics before writing any code
2. **Make good use of `/browse` verification**: Don’t just look at the code, let AI really "see" your application
3. **Periodic `/retro`**: Maintain visibility into code quality and work pace
4. **Gradual Adoption**: No need to use all skills at once. Starting from `/office-hours` + `/review` + `/ship`
5. **Fork customization**: If you encounter an inappropriate prompt, change it directly. This is the advantage of open source
## Summary
The core value of gstack does not lie in how powerful a specific skill is, but in that it provides a **structured AI collaboration mode** - through role switching, you can get different types of AI assistance at different stages. First review the product direction from the CEO's perspective, then review the architecture with the rigor of an engineering manager, and finally verify the results with QA's real browser.
Next, you can try installing it yourself and start your first gstack project from `/office-hours`.
***
**Extended reading**:
* [gstack Concepts](/en/docs/notes/gstack/concept) — Understand gstack’s core concepts and tool ecological positioning
* [GSD Practical](/en/docs/notes/gsd/practice) — A practical guide to another structured AI programming solution
* [Claude Skills Practical Chapter](/en/docs/notes/claude-skills/skill-creator) — Understand the creation mechanism of Skills
# Teardown gstack: What Skill Developers Can Learn
## Introduction
In [Concept](/en/docs/notes/gstack/concept) and [Practical](/en/docs/notes/gstack/practice), we learned what gstack is and how to use it from a user perspective. This note is from a different perspective - **As a Skill developer**, after reading the gstack warehouse file by file, what engineering designs are worth learning and learning from.
gstack is more than just a collection of 23 prompt files. There is a complete engineering system behind it: template generation, automatic upgrade, learning and memory, progressive guidance, multi-platform adaptation, layered testing - these are the keys to turning a skill project from "usable" to "easy to use".
***
## 1. SKILL.md is not handwritten - template generation system
The most counter-intuitive design of gstack: \*\*Each SKILL.md is automatically generated and cannot be edited directly. \*\*
```text
SKILL.md.tmpl (人写) → gen-skill-docs → SKILL.md (机生)
```
The human-written `.tmpl` template contains workflow logic and best practices, plus `{{PLACEHOLDER}}` placeholders. The build script extracts the command reference, browser flag list, preamble startup code, etc. from the source code and fills them into the placeholders to generate the final SKILL.md.
```text
{{PREAMBLE}} ← 从 resolvers/preamble.ts 生成的启动代码
{{BROWSE_SETUP}} ← 浏览器初始化指令
{{COMMAND_REFERENCE}} ← 从 commands.ts 提取的命令文档
{{SNAPSHOT_FLAGS}} ← 从源代码常量提取的快照选项
```
\*\*Why do this? \*\*
* Documentation and code will never be out of sync - the command reference is generated from the source code, and the documentation is automatically updated when the source code changes
* 23 skills share the same preamble (about 220 lines), and all skills are updated simultaneously
* CI can `--dry-run` check whether the generated file is expired to prevent forgetting to regenerate
**Takeaway**: If you maintain multiple skills, any content shared across skills should be extracted into templates and used in build steps to generate the final files. Manually syncing multiple copies of the same content will cause problems sooner or later.
***
## 2. Upgrade mechanism - complete link from detection to execution
The upgrade system of gstack is very exquisitely designed and divided into three layers:
### First layer: version detection
`bin/gstack-update-check` is a standalone bash script that does the following:
1. Read the local `VERSION` file
2. Check cache `~/.gstack/last-update-check` (UP\_TO\_DATE caches for 60 minutes, UPGRADE\_AVAILABLE caches for 720 minutes)
3. If the cache expires, HTTP request GitHub’s `raw.githubusercontent.com/.../VERSION`
4. Compare the version number and output `UPGRADE_AVAILABLE <旧> <新>`
### Second layer: Preamble integration
**The first line of each skill's SKILL.md startup code is version detection**:
```bash
_UPD=$(~/.claude/skills/gstack/bin/gstack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
```
This means that updates will be automatically detected when users call any skill - no need to run an upgrade command specifically, zero presence but 100% coverage.
### The third level: progressive reminder + automatic upgrade
After detecting a new version, it will not immediately disturb the user, but use the snooze mechanism (Snooze) for progressive backoff:
* 1st Reminder: Please mention again after 24 hours
* 2nd Reminder: Please mention again after 48 hours
* 3rd time and later: please mention again after 7 days
* New version release resets snooze counter
Users can `gstack-config set auto_upgrade true` enable automatic upgrade and skip confirmation to execute it directly.
When performing the upgrade, 5 installation types (global git, local git, vendored, etc.) will be distinguished. git installation uses `git fetch + reset`, vendored installation first backs up and then replaces, and restores from `.bak` in case of failure. After the upgrade, the local vendored copy of the project will also be automatically synchronized.
**Points worth learning from**:
* The "detect on each call" mode has extremely high coverage and is imperceptible to users
* Gradual backoff avoids frequent interruptions
* Differentiate installation types and implement different upgrade strategies instead of one size fits all
* Backup and restore ensure that upgrade failure will not cause the entire skill to hang up.
***
## 3. Learning system - Make Skills smarter the more you use them
gstack implements a lightweight but effective **cross-session memory system**.
### Storage
Each project has an independent learning log: `~/.gstack/projects/$SLUG/learnings.jsonl`, which is written additionally.
```json
{
"skill": "review",
"type": "pitfall",
"key": "n-plus-one",
"insight": "这个项目的 User model 有 N+1 查询问题,findAll 要加 include",
"confidence": 8,
"source": "observed",
"files": ["src/models/user.ts"],
"ts": "2026-04-01T14:30:00Z"
}
```
### Automatic collection
Before each skill is completed, there is an "operational self-improvement" link - reflecting on whether there were unexpected failures, detours, or project quirks discovered during the execution, and any will be automatically recorded in learnings.jsonl. No manual triggering is required by the user.
### Automatic loading
Each time a new session starts, the preamble will load the first 3 high-confidence learning entries to inject context, allowing the new session to inherit historical knowledge.
### Confidence decay
Entries from `observed` and `inferred` sources decay by 1 point every 30 days. There is no need to manually clean up the knowledge base—outdated knowledge fades naturally and new observations naturally take its place.
### Management interface
```text
/learn # 显示最近 20 条
/learn search # 搜索
/learn prune # 检测过期条目(引用的文件已删除)
/learn export # 导出为 markdown 可加入 CLAUDE.md
```
**Points worth learning from**:
* Additional write-only design is simple and reliable, and concurrency is safe
* Confidence decay is a low-maintenance knowledge aging management - much more efficient than manual cleaning
* Use git remote URL instead of path to identify the project (via `gstack-slug`), which can be cloned to different locations and reused.
* Support cross-project query, but isolated by default
***
## 4. Preamble injection - Skill's "middleware layer"
This is one of the smartest architectural designs of gstack. Each SKILL.md shares a preamble code of about 220 lines, which functions like the middleware of a web framework:
```text
┌─ 更新检测 ──────────────────────────────────┐
│ 会话追踪 (sessions/$PPID) │
│ 配置读取 (proactive, skill_prefix, telemetry)│
│ 学习历史加载 (前 3 条高置信度) │
│ 上下文恢复 (最近的 checkpoint + timeline) │
│ 路由规则检测 │
│ 首次使用引导流程 │
└──────────────────────────────────────────────┘
↓
Skill 特有逻辑
```
Preamble's bash script outputs key-value pairs (`BRANCH: main`, `PROACTIVE: true`), and then the template uses natural language conditions to let Claude adjust his behavior accordingly:
```text
If PROACTIVE is false, do not invoke skills automatically.
Instead suggest: "I think /skillname might help here -- want me to run it?"
```
This is essentially treating bash output as Claude's "environment variables" - using bash for runtime detection and natural language for behavioral routing.
**Points worth learning from**: If you have multiple skills, the shared logic (configuration loading, state recovery, version detection) should be extracted into a unified preamble instead of writing a copy for each skill.
***
## 5. Progressive boot - Sentinel file mode
The first-time user experience of gstack is designed with great care. Make sure each boot step occurs only once via a touch file (sentinel file):
```text
~/.gstack/.completeness-intro-seen ← "Boil the Lake" 理念介绍
~/.gstack/.telemetry-prompted ← 遥测选择(community/anonymous/off)
~/.gstack/.proactive-prompted ← 主动触发开关
~/.gstack/.routing-prompted ← CLAUDE.md 路由规则写入
~/.gstack/.welcome-seen ← 安装欢迎消息
```
Check whether these files exist each time the skill is started. If not, display the corresponding boot and touch files. Steps that have already been viewed will never appear again.
**Points worth learning from**: Compared with maintaining the status of `"onboarding_step": 3` in config, sentinel files are simpler and more reliable - they will not be affected by configuration file corruption, and each step is controlled independently.
***
## 6. SKILL.md structural design - three-tier architecture
Each SKILL.md follows a standard three-layer structure:
### First layer: YAML Frontmatter
```yaml
---
name: qa
preamble-tier: 3
version: 0.15.1.0
description: |
Systematically QA test a web application...
Use when asked to "qa", "test this site", "find bugs"...
benefits-from: [office-hours]
allowed-tools:
- Bash
- Read
- Write
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "bash ${CLAUDE_SKILL_DIR}/bin/check-careful.sh"
---
```
Key fields:
* `allowed-tools`: Tool-level permission whitelist, each skill declares only the tools it needs
* `benefits-from`: Explicitly declare the pre-dependency skill
* `hooks`: PreToolUse hook, which can intercept before the tool is called (such as careful interception `rm -rf`)
* `description`: Contains all natural language trigger words
### Layer 2: Shared Preamble + General Rules
preamble startup code + Voice definition + context recovery + integrity principle + search priority + completion status protocol + upgrade rules, etc. All skills are identical and generated from templates.
### The third layer: Skill-specific logic
This is the "soul" of each skill - workflow definition, role setting, cognitive model injection, interaction gating, etc.
**Points worth learning from**: The three-layer separation allows each skill to only focus on its own unique logic, and the shared parts are ensured by the framework for consistency.
***
## 7. Prompt Engineering Tips Collection
After reading all SKILL.md, here are the prompt design techniques worth learning:
### Anti-Flattery Rules
The Startup Mode of office-hours explicitly prohibits the common "and muddy" behavior of AI:
```text
Never say:
- "That's an interesting approach" → take a position instead
- "There are many ways to think about this" → pick one
- "You might want to consider..." → say "This is wrong because..."
- "That could work" → say whether it WILL work
```
### Prohibited word list
The Voice section has clear banned words and phrases:
* Banned words: delve, crucial, robust, comprehensive, nuanced, pivotal, landscape...
* Banned phrases: "here's the kicker", "plot twist", "let me break this down"...
* Disabled format: em dash (replace with comma/period)
These are common "AI-flavored" words in LLM, and the output is obviously more natural after being disabled.
### Cognitive model injection
Each Review skill injects a different thinking framework:
* **CEO Review**: 18 cognitive models (Bezos’ one-way/two-way door decision-making, Munger’s reverse thinking, Jobs’ focus and subtraction...)
* **Eng Review**: 15 engineering management patterns ("boring by default", blast radius intuition, Conway's law\...)
* **Design Review**: 12 Design Cognition Patterns (Hierarchy as a Service, Worship of Constraints, "Would I notice?" testing...)
These modes do not let AI perform mechanically, but provide it with a **thinking framework** - just like giving a smart newcomer a list of experiences of its predecessors.
### Specification standards
```text
Not "you should test this"
but `bun test test/billing.test.ts`
Not "this might be slow"
but "this queries N+1, ~200ms per page load with 50 items"
Not "there's an issue in the auth flow"
but "auth.ts:47, the token check returns undefined"
```
### Confidence Calibration
The review skill requires each discovery to be accompanied by a confidence score, and low-confidence findings are automatically downgraded or hidden:
| Score | Meaning | Processing |
| ----- | --------------------------------- | ------------------------- |
| 9-10 | Read the specific code and verify | Normal display |
| 7-8 | High confidence pattern matching | Normal display |
| 5-6 | Moderate, possible false alarm | Display with instructions |
| 3-4 | Low confidence | Hide from reports |
| 1-2 | Pure guessing | Only shown at P0 level |
### Interactive Gating
The ship skill precisely defines when to stop and wait for the user and when to continue automatically:
```text
Only stop for:
- Tests failing with no obvious fix
- Merge conflicts requiring human judgment
- Unclear which changes to include
Never stop for:
- Normal git operations
- CHANGELOG/VERSION updates
- PR creation
```
**Points worth learning from**: A good skill is not "AI does everything", but a precise definition of the human-machine boundary.
***
## 8. State management - file system is database
All persistence of gstack is done through the file system, stored under `~/.gstack/`:
| Path | Purpose | Format |
| ------------------------------------- | -------------------- | ---------- |
| `config.yaml` | Global configuration | YAML |
| `sessions/$PPID` | active session | touch file |
| `projects/$SLUG/learnings.jsonl` | Learning record | JSONL |
| `projects/$SLUG/timeline.jsonl` | Skill Timeline | JSONL |
| `projects/$SLUG/checkpoints/*.md` | Checkpoint | Markdown |
| `projects/$SLUG/health-history.jsonl` | Health check history | JSONL |
| `analytics/skill-usage.jsonl` | Using telemetry | JSONL |
| `last-update-check` | version cache | plain text |
Almost all time series data is written appendably using **JSONL** (one JSON object per row). This choice is smart:
* Added write natural concurrency safety
* No database dependencies required
* You can use `grep` / `jq` to query directly
* Corrupted up to missing last line
***
## 9. Cross-Skill Integration Mode
### File transfer product
Transfer work products between skills through the file system:
```text
/office-hours → design doc → /plan-ceo-review 读取
/plan-ceo-review → ceo-plans/*.md → /autoplan 读取
/review → reviews.jsonl → /ship 读取并展示 Dashboard
/qa → qa-reports/ → /retro 读取
```
### Review Readiness Dashboard
The ship skill reads `reviews.jsonl`, showing cross-skill review status before publishing:
```text
| Review | Runs | Last Run | Status | Required |
| Eng Review | 1 | 2026-03-16 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
```
### Pre-dependency suggestions
When plan-ceo-review detects that there is no design doc, it will actively recommend running `/office-hours` first:
```text
"No design doc found. /office-hours produces a structured problem statement...
Takes about 10 minutes."
Options: A) Run /office-hours now B) Skip
```
### Use sequence prediction
Context Recovery will analyze the recent skill usage sequence and predict the next step:
```text
If pattern repeats (e.g., review → ship → review),
suggest: "Based on your recent pattern, you probably want /ship."
```
***
## 10. Other noteworthy designs
### Hook system
The three skills careful, freeze, and guard use `PreToolUse` hooks - this is the only mechanism that can intercept before the tool is called:
* **careful**: Intercept Bash, check `rm -rf`, `DROP TABLE`, `git push --force`
* **freeze**: intercept Edit/Write and check whether the path is within the allowed range
* **guard**: combine the above two
### Multi-platform adaptation
The same set of templates generates skill files for different platforms through the `--host` parameter:
```bash
bun run gen:skill-docs --host claude # Claude Code 格式
bun run gen:skill-docs --host codex # OpenAI Codex 格式
bun run gen:skill-docs --host kiro # AWS Kiro 格式
bun run gen:skill-docs --host factory # Factory Droid 格式
```
The path and frontmatter are automatically adapted, and the skill logic remains unchanged.
### Completion status protocol
A standardized completion status must be output at the end of each skill:
```text
DONE — 全部完成,提供证据
DONE_WITH_CONCERNS — 完成但有顾虑
BLOCKED — 无法继续
NEEDS_CONTEXT — 需要更多信息
```
### Three failed upgrade rules
```text
If you have attempted a task 3 times without success, STOP and escalate.
```
Prevent AI from getting stuck in an infinite retry loop.
### Diff-based test selection
E2E tests cost about $4 each (requires starting Claude agent), so gstack declares the source files each test depends on through `touchfiles.ts`, and only runs the affected tests according to `git diff`:
```typescript
// test/helpers/touchfiles.ts
{
"qa-workflow": ["qa/SKILL.md.tmpl", "browse/src/server.ts"],
"ship-flow": ["ship/SKILL.md.tmpl", "scripts/resolvers/preamble.ts"]
}
```
***
## Summary: Design principles you can take away
From gstack’s engineering practice, I extracted the following design principles that are most valuable to Skill developers:
1. **Template generation > Manual synchronization**: Content shared across skills is automatically generated using templates + build steps, do not copy and paste
2. **Passive detection > Active detection**: Upgrade detection is embedded in every skill call, the user is unaware but the coverage rate is 100%
3. **Append log > Complex database**: JSONL + file system can cover most persistence needs, simple and reliable
4. **Progressive Boot > One Configuration**: Use sentinel files to control boot steps, each appearing only once
5. **Precise Gating > Fully Automatic**: Clearly define the boundaries between "stop and wait for user" and "automatically continue"
6. **Confidence quantification > Fuzzy judgment**: Each AI judgment comes with a confidence score, and low confidence is automatically downgraded.
7. **Time Decay > Manual Cleaning**: The confidence of learning records decays with time, and outdated knowledge naturally fades
8. **BANNED WORDS LIST > STYLE GUIDE**: A direct list of prohibited words is much more effective than "please use a natural tone"
***
**Related Reading**:
* [gstack Concepts](/en/docs/notes/gstack/concept) — What is gstack and what problems does it solve?
* [gstack Practical Chapter](/en/docs/notes/gstack/practice) — Complete workflow from installation to runthrough
* [gstack Front-end Skill](/en/docs/notes/gstack/frontend-skills) — Front-end/UI design Skill panorama and recommended workflow
* [Claude Skills Concept](/en/docs/notes/claude-skills/concept) — Understand the underlying mechanism of Skills
# Diagnose and Triage: Establish a Feedback Loop First, Then Decide Who Handles It
## Why Discuss These Two Skills Together
`/diagnose` and `/triage` are presented as two independent skills in the README, but they address two halves of the same engineering problem:
* `/diagnose` cares about: **What exactly is this bug, how can it be reproduced, and how can we prove it's fixed?**
* `/triage` cares about: **Should this issue wait for more information, be assigned to an agent, to a human, or be closed?**
One deals with facts, the other with process. In real projects, these two are often linked: you triage a bug issue, find insufficient information, and mark it `needs-info`; once enough information is available, you use diagnose to build a feedback loop; after clearly reproducing it, you decide whether it's `ready-for-agent` or `ready-for-human`.
## The Core of /diagnose: The Feedback Loop Is Everything
The most important takeaway from `/diagnose` is: **first, establish a pass/fail signal that an agent can run.**
Matt breaks down diagnosis into 6 stages:
| Stage | Goal |
| --------------------- | ------------------------------------------------------------- |
| Build a feedback loop | Set up a fast, deterministic, repeatable failure signal |
| Reproduce | Make this signal reproduce the same bug described by the user |
| Hypothesise | List 3-5 falsifiable hypotheses |
| Instrument | Verify hypotheses with minimal probes |
| Fix + regression test | Write regression tests on the correct test surface, then fix |
| Cleanup + post-mortem | Clean up temporary probes, document the true root cause |
This is contrary to how many people debug. The common debugging process is: look at code, guess the cause, make a change, refresh the page. Matt reverses this: first, turn the bug into a repeatable machine signal, then discuss hypotheses.
## What Constitutes a Good Feedback Loop
`/diagnose` provides a set of priorities, from best to last resort:
| Loop Type | Suitable Scenario |
| ----------------------------- | ----------------------------------------------------------------------------------------- |
| Failing test | When there's an appropriate test surface that can directly express the bug |
| curl / HTTP script | API bugs, server-side behavior reproducible with requests |
| CLI + fixture | Command-line tools, parsers, transformers |
| Headless browser | UI bugs, console errors, network behavior |
| Replay captured trace | Real-world payloads, event streams, log traces from production |
| Throwaway harness | Start only a small part of the system, isolating complex dependencies |
| Property / fuzz loop | Intermittent error output, needs increased trigger rate |
| Bisection / differential loop | Broken after a certain version, needs bisection or comparison with old versions |
| HITL script | Even when manual interaction is required, ensure humans follow a script for stable output |
There's a strict judgment here: **without a loop, do not enter the hypothesis stage**. Without a signal, all analysis becomes "it looks like."
## What About Non-Deterministic Bugs?
`/diagnose` also offers a practical approach to intermittent bugs: the goal isn't 100% reproduction from the start, but to increase the reproduction rate to a debuggable level.
For example:
* Trigger 100 times in a loop
* Trigger concurrently
* Inject `sleep` to widen race windows
* Fix random seeds or time
* Narrow down environment variables and external dependencies
A 1% intermittent bug is hard to debug; a 50% intermittent bug is already a debuggable object. This approach is very useful for frontend async issues, message queues, payment callbacks, and streaming output.
## Hypotheses Must Be Falsifiable
Matt requires listing 3-5 hypotheses before attempting verification, and each hypothesis must include a prediction:
```text
If X is the cause, then changing Y should make the bug disappear;
or observing Z should reveal a certain characteristic.
```
This prevents the agent from getting stuck on the first seemingly plausible explanation. More importantly, it allows you to judge whether an experiment provides any information.
A bad hypothesis:
```text
It might be a caching issue.
```
A falsifiable hypothesis:
```text
If browser caching is causing an old script to execute, then after disabling cache and hard refreshing, the old bundle hash in the console should disappear, and button click events should resume functionality.
```
Only the latter is worth verifying.
## Common Mistakes in the Fix Stage
`/diagnose` requires that if there's a correct test surface, you should first convert the minimal reproduction into a failing test, then fix the code.
The key is "correct test surface." It's not just adding any unit test as a regression test. The correct test surface must cover the real bug pattern:
* If the bug is triggered by a combination of multiple callers, don't just test a single function.
* If the bug is triggered by a real payload structure, don't just test a handwritten toy object.
* If the bug is triggered by the order of browser events, don't just test pure functions.
If you can't find the correct test surface, that itself is a conclusion: the code structure doesn't provide a place to lock down the bug. After fixing, this information should be passed to [`/en/docs/notes/matt-pocock-skills/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture).
## The Core of /triage: Issues as State Machines
`/triage` is not about letting AI "just take a look at the issue." It treats issues as small state machines.
Every issue should simultaneously have:
* A category: `bug` or `enhancement`
* A state: `needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`
The value of this set of states is that it allows maintainers to quickly answer:
* Which ones haven't been looked at yet?
* Which ones are waiting for the reporter to provide more information?
* Which ones are clear enough to be assigned to an AFK agent?
* Which ones must be done by a human?
* Which ones should be closed, and why?
## Criteria for ready-for-agent
`ready-for-agent` is the most critical state in this workflow. It's not "this task can be tried by AI," but rather:
> The task is clear enough for an absent agent to independently pick up, implement, and verify.
This usually means the issue contains at least:
* Background and problem statement
* Relevant code paths or modules
* Clear acceptance criteria
* Known constraints
* If it's a bug, ideally a reproduction method
* No need for additional product/design judgment
If these are missing, it should be `needs-info` or `ready-for-human`, not forced onto an agent.
## needs-info Requires Specific Questions
`/triage` provides a simple template for `needs-info`, but the key is that the questions must be specific:
```markdown
## Triage Notes
**What we've established so far:**
- ...
**What we still need from you (@reporter):**
- ...
```
Bad question:
```text
Please provide more information.
```
Good question:
```text
Please provide the browser version that triggered the issue, the URL of the problematic page, the click sequence, and the response body of `/api/orders/:id` from the Network panel.
```
AI can easily write polite platitudes; this skill forces it to separate "what we already know" from "what we still need."
## wontfix Also Needs Documentation
`/triage` has an interesting design for `wontfix` enhancements: don't just close the issue, but write the rejection reason into the `.out-of-scope/` knowledge base and link to it in the comments.
This way, the next time a similar request comes up, the AI won't re-engage in the same discussion. It can first read `.out-of-scope/` and remind the maintainer: "This direction was previously rejected for reason X."
This is similar to the spirit of ADRs: not recording all decisions, but only those that will cause confusion in the future and are likely to recur.
## How They Work Together
A typical bug issue might follow this path:
1. `/triage` reads the issue, comments, labels, and relevant code.
2. It determines this is `bug + needs-triage`.
3. It first attempts to reproduce; if steps are insufficient, it changes to `needs-info`.
4. Once enough information is available, it initiates `/diagnose`.
5. `/diagnose` establishes a reproduction loop, lists hypotheses, and identifies the root cause.
6. If the fix path is clear and the test surface is well-defined, the issue becomes `ready-for-agent`.
7. If product judgment, external permissions, or manual verification are needed, the issue becomes `ready-for-human`.
8. After fixing, the root cause and regression tests are written back to the issue or PR.
The key to this process is not "AI automatically fixes bugs," but rather transforming a vague issue description into an executable work package.
## My Usage Recommendations
If you only remember one thing:
> `/diagnose` first asks "How do I prove it's broken?"; `/triage` first asks "What state should it be in now?"
These two questions can prevent a lot of low-quality AI programming:
* Fixing without reproduction
* Starting work without acceptance criteria
* Refactoring without a root cause
* Assigning to an agent without information
Matt's two skills are not flashy, but they are very much like what a senior engineer in a real team would do: first consolidate the facts, then advance the process.
## References
Next: [TDD: Force AI to Take Small Steps with Red-Green-Refactor](/en/docs/notes/matt-pocock-skills/tdd).
# Grill Me: Let AI ask you 50 questions before you write code
## Failure mode: “The AI didn’t do what I wanted”
The first failure mode that Matt talked about in his speech is: you think that the requirements in your mind are very clear, and let the AI write them out - that is not the case at all.
> "I would run it, and I would try not to look at the code, but I would look at the code, and I realized I would get worse code. I did it again, I got even worse code... I did it again, kept running the compiler, and I would just end up with garbage."
Many people are familiar with this feeling: if you say "Add a login for me", AI will not ask you "Do you want to remember the device?" "How many times have you failed to lock the account?" "How long does it take for the session to expire?" It directly lays out a plan that it thinks is reasonable. By the time you review it, 500 lines have been written - two hours of reworking.
## Why is this happening: design concept deviates
Matt quotes Frederick P. Brooks’ **design concept** (design concept) in “The Design of Design”:
> When multiple people collaborate to design something, there will be something being created between you - it is floating in your mind, an invisible "theory about this thing". It's not an asset, it's not an asset stuffed into a markdown file, it's an invisible consensus.
AI writes code as soon as it comes up, which means it doesn’t share the same design concept with you at all. What is wrong when writing code is not the syntax, but the premise.
To fix this problem, you must first do design concept alignment before starting. The tool Brooks gave is called **design tree** - split a decision into multiple branches, and then split each branch. You can't skip upstream decisions and make downstream decisions directly, otherwise everything will have to be redone once the upstream changes to the downstream.
## Matt’s Skill full text
Matt's implementation of this theory in [`mattpocock/skills`](https://github.com/mattpocock/skills) is `productivity/grill-me/SKILL.md`, and the entire file plus frontmatter is less than 15 lines:
```markdown
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
Interpret sentence by sentence:
* **"interview me relentlessly"** - The key word is *relentlessly* (not letting go). By default, LLM has a tendency to ask 1 or 2 questions and then feel "almost" and start taking action. This word forcibly suppresses this tendency.
* **"walk down each branch of the design tree"** —— Brooks' concept of design tree. Force Claude to treat your requirements as a tree, solving the upstream first and then the downstream. If you say "Login", it will first ask "Authentication method" (tree root), and then expand "How to manage session"/"How to save token" (sub-node) based on your answer.
* **"resolving dependencies between decisions one-by-one"** - Explicitly prohibit packaged questions. There are often dependencies between decisions (if you choose SSO, the downstream does not need password policy issues), first make sure that the upstream can eliminate a lot of downstream problems.
* **"for each question, provide your recommended answer"** - key bonus points. AI not only asks questions, but also recommends answers. You just nod/no and save 80% of typing time.
* **"ask the questions one at a time"** - Prevents the AI from giving you 10 questions at a time.
* **"if a question can be answered by exploring the codebase, explore the codebase instead"** - If it is a fact that already exists in the project (such as "What testing framework is used in the project"), let Claude see it himself, don't ask you.
7 lines, but each sentence corresponds to a specific LLM behavioral bias.
## How to install and use
**Installation**:
```bash
npx skills@latest add mattpocock/skills
```
Check `grill-me` and `setup-matt-pocock-skills` (grill-me does not depend on the latter, but other skills depend on it, so it is recommended to install them together).
**Call**: Enter `/grill-me` in the Claude Code dialog.
**Typical process**:
1. You describe what you want to do, which can be very vague ("I want to add a comment function to my blog")
2. Enter `/grill-me`
3. Claude started asking questions one by one, and gave recommended answers to each question.
4. You answer each question one by one (nod/no/correct)
5. Generally, a consensus is reached after 20 to 50 questions, and Claude will give you a summary.
6. The summary can be directly fed to [`/to-prd`](/en/docs/notes/matt-pocock-skills/to-prd-and-issues) to become PRD, or directly passed to [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd) to start writing.
## Real case: How much does a video editor function cost?
Matt gave a few specific numbers in ["5 Agent Skills I Use Every Day"](https://www.aihero.dev/5-agent-skills-i-use-every-day):
* **New video editor feature** - 16 questions to reach consensus
* **Complex functions** - 30\~50 questions
* **EXTREMELY COMPLEX** - 100 questions, session up to 45 minutes
Sample question (restored from Matt's video/blog post):
* "Should video clips be reorderable, or only added/removed in sequence?"
* "When a clip is deleted, do we keep its source file, or delete the file too?"
* "Does the editor need undo/redo? How many steps deep?"
* "Should we render previews in the browser, or rely on a backend service?"
None of these issues were technical - they were all product decisions. But **every decision determines the shape of hundreds of lines of code**. If you skip these questions and let AI write them directly, it will come up with a set of answers on its own and you will come back to reject them one by one after writing them.
## Differences between Plan Mode and Claude Code’s built-in Plan Mode
Claude Code comes with `plan mode` (press Shift+Tab to enter). On the surface, it looks similar to grill-me—discuss first before taking action. But Matt directly said in his speech that he prefers grill-me:
Specific differences:
| Dimensions | Plan Mode | /grill-me |
| ---------------------------------------- | ------------------------------------------------------------------- | ----------------------------------------------------------- |
| Default goal | Produce an executable plan as soon as possible | Reach a consensus first, the plan is a by-product |
| Number of questions | 0\~5 | 20\~100 |
| Question format | Ask one paragraph at a time | Ask one question at a time |
| Do you want to give a recommended answer | No | Yes |
| Whether to explore the code base | Occasionally | Actively (explicit instructions) |
| Suitable for scenarios | Already thought clearly and want to confirm the implementation plan | Not yet thought clearly, need to be forced to think clearly |
The biggest practical difference is "**urgent or not**". Plan mode is in a hurry to start, grill-me is not in a hurry - it regards "thinking clearly" as the main task rather than the prologue.
## Advanced usage
### 1. Non-programming scenarios
`grill-me` does not bind code and can also be used for pure product decision-making conversations. Matt uses it himself:
* Course syllabus design
* Article writing
* Internal communication documents
As long as you have a vague idea in your mind and want to be forced to think it through, you can use it.
### 2. Cooperate with [`/to-prd`](/en/docs/notes/matt-pocock-skills/to-prd-and-issues)
After the grill-me session is over, just say `/to-prd`, and Claude will condense the entire conversation into a structured PRD (including user story, module splitting, and testing strategy) and submit it to your issue tracker. **Key point: Do not clear context in the middle** - to-prd is extracted directly from the conversation context and will not ask you again.
### 3. Cooperate with [`/grill-with-docs`](/en/docs/notes/matt-pocock-skills/grill-with-docs)
If the project already has `CONTEXT.md` (domain language) and `docs/adr/` (architectural decisions), use `/grill-with-docs` instead of `/grill-me`. It will **synchronously update CONTEXT.md** while being tortured - decisions are being made while documents are updated, and there is no longer the problem of "documents being forever out of date".
### 4. Customize question depth
If you are pressed for time, you can add a sentence directly after `/grill-me`: "Limit to 10 questions, focus only on architecture decisions." It will converge according to your limit. But Matt doesn’t recommend it – he believes that “asking more questions” is exactly the value of this skill, and cutting it off is almost like plan mode.
## Notes
**It will be annoying the first time you run**. People who are used to "generating 500 lines in one sentence" will feel that it is a waste of time to be asked 30 times by the AI for the first time. Matt’s advice is to hold on to the first 5 questions – the first 5 questions often reveal things you haven’t even thought about. Once you pass that threshold, you’ll be addicted.
**Not suitable for extremely small tasks**. Change a typo, add a console.log - don't use grill-me. It is suitable for "making a new thing" or "changing an old thing with side effects".
**Sometimes the AI will ask for technical details**. If you don't care and want to let it judge, just reply "your call" and "you decide", and it will accept it and continue.
## Why is this Skill popular?
`/grill-me` is the most frequently screenshotted and forwarded skill among Matt’s skills. The reason is not complicated:
1. **Extremely minimalist**: 7 lines of markdown, just copy and paste
2. **Immediate effect**: You can feel the change in AI’s “problem density” during the first run.
3. **Portable**: Does not depend on Claude Code, Codex, Cursor, and Aider can all be used
4. **Comes with anti-LLM default behavior**: Each word is anti-LLM deviation, with high engineering aesthetics
Its success has also become the best argument for "**skill does not necessarily have to be long**".
## Reference resources
Next article: [Grill With Docs: Maintaining project language and ADR](/en/docs/notes/matt-pocock-skills/grill-with-docs) - an advanced version of grill-me, for projects with domain complexity.
# Grill With Docs: Equipping AI with project memory using domain language and ADR
## Failure mode: "AI is too verbose"
The second failure mode in Matt’s speech:
> The AI expresses a simple thing in a bunch of verbiage. It's like speaking two languages to you.
This has nothing to do with the amount of code, it's **vocabulary misplacement**. AI will use general terms ("item", "data", "handler") by default, and the real terms in the project in your mind may be "Course", "Draft Version", "Ghost Lesson". AI doesn't know that these words have specific meanings in your project, so it will create a bunch of synonymous new words around them. The result is:
* Long-winded thinking process (avoid your proprietary words)
* The implementation is misaligned with the design in your mind (because they are not in the same semantic space)
* Not reusable across sessions (the context must be re-established for each conversation)
## Classic theory: DDD’s Ubiquitous Language
Matt cited "Domain-Driven Design" by Eric Evans. This book was published in 2003 and proposed the concept of **ubiquitous language**:
> Use the same set of terms to connect domain experts, developers, and code. A word in product discussions, code comments, variable names, documentation - it must mean the same thing.
The goal of DDD is to make code look like the brain of a domain expert. In the AI era, there is a new role: **LLM must also be in this language**. The LLM is not in the standup meeting, cannot see the product requirements meeting, and cannot understand the slang of your group - it can only learn from the documents you give it.
Matt turned this into a skill: scan the code base to extract terms, generate a markdown file `CONTEXT.md`, and then align it with both humans and AI.
## The Evolution of Skill: From Ubiquitous Language to Grill With Docs
The earliest skill was called `ubiquitous-language` - it only did one thing: scan the code base to generate a glossary. But Matt later discovered that simply generating a document was not enough:
* **Documentation will be out of date**: it is generated today, the code is changed tomorrow, and the glossary has not kept up.
* **People won’t take the initiative to look at it**: It would be dead if you put it there
He refactored it into `grill-with-docs`, which combined three things:
1. **Torture Requirements** (all abilities inherited from grill-me)
2. **Challenge existing glossary**: The word you said is inconsistent with what is written in CONTEXT.md? Point it out immediately
3. **Update documents simultaneously when making decisions**: New conclusions reached during the torture process are written inline into CONTEXT.md or create a new ADR.
This is a paradigm shift from "generating static documents" to "dialogue is maintaining documents".
## Skill full text
Core structure of `engineering/grill-with-docs/SKILL.md`:
```markdown
---
name: grill-with-docs
description: Grilling session that challenges your plan against the
existing domain model, sharpens terminology, and updates
documentation (CONTEXT.md, ADRs) inline as decisions crystallise.
---
Interview me relentlessly about every aspect of this plan until we
reach a shared understanding. Walk down each branch of the design
tree, resolving dependencies between decisions one-by-one. For each
question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question
before continuing.
If a question can be answered by exploring the codebase, explore the
codebase instead.
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
If a CONTEXT-MAP.md exists at the root, the repo has multiple contexts.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in
CONTEXT.md, call it out immediately.
"Your glossary defines 'cancellation' as X, but you seem to mean Y —
which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise
canonical term.
"You're saying 'account' — do you mean the Customer or the User?
Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with
specific scenarios.
### Cross-reference with code
When the user states how something works, check whether the code
agrees. If you find a contradiction, surface it.
### Update CONTEXT.md inline
When a term is resolved, update CONTEXT.md right there. Don't batch
these up — capture them as they happen.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. Hard to reverse
2. Surprising without context
3. The result of a real trade-off
```
## What does the real CONTEXT.md look like?
Matt's own [`course-video-manager`](https://github.com/mattpocock/course-video-manager/blob/main/CONTEXT.md) repository gives a complete CONTEXT.md example. Let’s pick a few terms to get a feel for it:
| Terminology | Definition |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| **Course** | The primary domain entity: a structured collection of versions, sections, lessons, and videos |
| **Draft Version** | The single mutable CourseVersion that is currently being edited; always the latest by `createdAt` |
| **Published Version** | An immutable CourseVersion with a name and description, created by the Publish flow |
| **Ghost Lesson** | A lesson that exists in the database but not yet on the file system (`fsStatus = "ghost"`) |
| **Export Hash** | A SHA256 hash derived from a video's clip filenames, timestamps, clip order |
| **Unexported Video** | A video whose current Export Hash does not match any file on disk; blocks publishing |
| **Materialization Cascade** | The chain reaction when materializing a lesson inside a ghost course |
| **Clip** | A timestamped segment of source footage within a video |
| **Fractional Index** | A string-based ordering value that allows inserting items between existing items |
| **Purge** | The deliberate deletion of an Exported Video's `.mp4` file from disk |
Note a few things:
1. **Each term is a gerund or proper noun** – not a descriptive phrase like “order status”
2. **Each definition refers to other terms** (Course → Version → Lesson → Video) forming an ontology network
3. **The code field appears directly** (`fsStatus = "ghost"`) - 1:1 mapping of documents and code
4. **Include decision descriptions** ("blocks publishing", "chain reaction") - not just nouns, but rules
When writing code, when AI sees this document, it will use "Ghost Lesson" instead of "lesson with no file". Code, conversations, and commit messages are all unified.
## ADR: when created
There is an important restraint in Skill:
> Only offer to create an ADR when all three are true:
>
> 1. **Hard to reverse** — the cost of changing your mind later is meaningful
> 2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
> 3. **The result of a real trade-off** — there were genuine alternatives
The culture of ADR (Architecture Decision Record) originated from Michael Nygard's blog in 2011, but many teams use it to write ADR for all decisions - 18 out of 20 ADRs are running accounts. Matt This triangulation is a great tool: **ADR is only worthwhile if three conditions are met simultaneously**. Otherwise, let it be digested in CONTEXT.md and digested in the code.
## How to install and use
**Preconditions**: Run `/setup-matt-pocock-skills` first (it will ask you where to put CONTEXT.md and where to put the ADR directory).
**Call**: `/grill-with-docs`
**Typical process**:
1. Describe what you want to do
2. `/grill-with-docs`
3. Claude **scans** CONTEXT.md and docs/adr/ first, loading existing terms and decisions into context
4. Start the torture, during the process:
* The word you used conflicts with CONTEXT.md → Point it out on the spot
* You used vague words (for example, "Account" could be either Customer or User) → Let you choose one of the two and drop it into the document
* The behavior you mentioned is inconsistent with the existing code → point out the conflict
5. **Update CONTEXT.md** synchronously when the decision is reached (no backlog, no batch processing)
6. Key irreversible decisions → Ask whether to generate an ADR
If there is no CONTEXT.md and docs/adr/ in the project yet - it will be created lazily: the files will not be generated until the first term needs to be written and the first ADR needs to be built. We won’t give you a blank template right off the bat.
\##Multiple context projects (CONTEXT-MAP.md)
If the project is too large to fit in one CONTEXT.md (for example, ordering and billing are two independent bounded contexts), you can put `CONTEXT-MAP.md` in the root directory as the general directory:
```
/
├── CONTEXT-MAP.md ← 总目录
├── docs/adr/ ← 系统级决策
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← 模块级决策
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
`/grill-with-docs` will automatically recognize the existence of CONTEXT-MAP.md and jump to the corresponding subdirectory. This is a direct implementation of the concept of **bounded context** in DDD - "orders" in each context may have different meanings, and are maintained separately to avoid contamination.
The difference between ## and grill-me
| dimensions | /grill-me | /grill-with-docs |
| -------------------------------- | ------------------------------ | -------------------------------------------------------------------- |
| Questioning ability | ✅ | ✅ (inherit all) |
| Project terminology verification | ❌ | ✅ |
| Real-time updates CONTEXT.md | ❌ | ✅ |
| ADR trigger judgment | ❌ | ✅ |
| Applicable stage | Early ideas, personal projects | Real projects with domain complexity |
| Startup cost | 0 | Requires setup + Have/are willing to build CONTEXT.md in the project |
Simple and crude judgment:
* **Personal scripts, writing articles, teaching courses** → `/grill-me`
* **Real projects that require long-term maintenance** → `/grill-with-docs`
## A counter-intuitive benefit: Let the AI learn to "shut up"
CONTEXT.md isn't just for AI - it's for future AI sessions. Every time a new conversation starts, Claude can instantly enter the project context by reading CONTEXT.md, saving a long onboarding.
What’s even more subtle is: Matt said in his speech that after adding CONTEXT.md he could see in AI’s *thinking trace*——
> "It allows the AI to think in a less verbose way."
Why? Because without CONTEXT.md, the AI has to constantly define its own terms when thinking - "the user, by which I mean the person who ordered the item, hereinafter referred to as...". With CONTEXT.md it directly says "Customer", the thinking chain is much shorter and the response is faster.
**LLM’s token economics determines: shortening the thinking path = faster, more accurate, and cheaper output**. CONTEXT.md is the hidden lever for this efficiency.
## Notes
**CONTEXT.md will be significantly changed for the first run**. If there is already a handwritten CONTEXT.md in the project, git stash first or let it dry-run before running (you can add a sentence in the prompt "List the content to be changed first, do not write the file directly").
**ADR Moderation is Real Moderation**. Don’t get excited and make every decision generate an ADR, your docs/adr/ will be full of garbage in 6 months. Matt’s three criteria must be strictly followed.
**CONTEXT.md Do not put implementation details**. There is a saying in Skill: "Don't couple CONTEXT.md to implementation details. Only include terms that are meaningful to domain experts." Writing "PostgreSQL" into CONTEXT.md is wrong - domain experts don't care about database selection, that's a matter of ADR.
## Reference resources
Next article: [to-PRD + to-Issues: From dialogue to executable ticket](/en/docs/notes/matt-pocock-skills/to-prd-and-issues)——After the torture, how to solidify the dialogue into an executable work unit.
# Improve Codebase Architecture: Restructure shallow into deep modules
## Failure mode: “AI wandering around in a bad code base”
The fourth failure mode in Matt’s talk is a picture metaphor:
> "Shallow modules in a codebase look like this - you have a bunch of tiny blobs, and the AI has to go through a bunch of modules and understand all the dependencies before it can correct them."
> "AI is really good at creating codebases like this. So you'll have a situation where AI doesn't understand what your code is doing. It will attempt to explore the code, but because it's poorly laid out, filled with shallow modules, it doesn't get to the right module in time, or doesn't understand all the dependencies."
This is a vicious cycle unique to AI programming:
```
AI 写代码倾向于产生 shallow 模块(小、多、互相依赖)
↓
代码库变得 shallow
↓
下次 AI 进来探索更难,更容易写错
↓
更多 shallow 模块被加进去
↓
代码库越来越烂,AI 越来越无能
```
To break this cycle, manual reverse refactoring must be performed periodically - merging shallow modules into deep modules. That's exactly what `/improve-codebase-architecture` does.
## Classic theory: Ousterhout’s Deep Modules
John Ousterhout is a professor of CS at Stanford (and the author of the Tcl language and Raft papers). His 2018 book "A Philosophy of Software Design" proposes a simple but powerful ruler:
**The "depth" of the module = the hidden complexity of the interface**
| Type | Interface | Implementation | Image |
| ----------- | --------- | -------------- | ----------------------------- |
| **Deep** | Simple | Rich | A rectangle: narrow and deep |
| **Shallow** | Complex | Simple | A rectangle: wide and shallow |
The ideal module is deep - users only need to see the short interface, and the complexity is hidden inside. An extreme counterexample is the shallow module: the interface is almost as complex as the implementation, which means there is no encapsulation. Users might as well look at the implementation directly.
Ousterhout's judgment: **Good code bases are composed of a small number of deep modules; bad code bases are composed of a large number of shallow modules**. This is completely opposite to the traditional dogma of "keeping functions as small as possible, files as short as possible, and modules as many as possible" - he believes that that kind of dogma produces exactly shallow modules.
## Matt’s extension: Deletion Test
Matt translated Ousterhout's theory into an operational engineering test, which he called the **deletion test**:
> **Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.**
Human words:
* **Delete it, the complexity disappears** → This module is originally a pass-through (transit), it is not working, so cut it off
* **Delete it, and the complexity will be spread to N callers** → It is originally helping you hide the complexity, it is really deep, leave it
The beauty of this test is that it is bidirectional - it can identify both "thin packaging that should be deleted" and "common logic that should be extracted." If you find that after deleting a piece of code, the complexity will spread to 5 places, it means that this code is worth extracting into a deep module.
## Key terms (Matt’s precise definition)
There is a Glossary in `improve-codebase-architecture/SKILL.md` that requires **strict use of these words** - do not drift to "component", "service", "API" and "boundary":
| Terminology | Definition |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------- |
| **Module** | Anything with an interface and implementation (function/class/package/slice) |
| **Interface** | Everything the caller must know - types, invariants, error modes, order, configuration (not just function signatures) |
| **Implementation** | Code inside the module |
| **Depth** | The lever at the interface. Deep = high leverage, shallow = interface is almost as complex as the implementation |
| **Seam** (Seam) | The location of an interface - where behavior can be changed without in-place modification. **Use "seam" not "boundary"** |
| **Adapter** | Implement the specific implementation of an interface at seam |
| **Leverage** | The benefits the caller gets from "deep" |
| **Locality** | The benefits that maintainers gain from "depth" - changes, bugs, and knowledge are all concentrated in one place |
Several core principles:
* **Deletion test**: see above
* **The interface is the test surface**: Tests can only be run through the interface - this is the basis for deep module testability
* **One adapter = hypothetical seam. Two adapters = real seam.**: An interface with only one implementation is a false seam. **Real joints require at least two adapters**
The last one is particularly counter-intuitive - many teams will abstract an interface in advance "for future expansion", but there is actually only one implementation. Matt's judgment: **useless, delete**. Wait until the second one comes true. This has the same origin as YAGNI.
## Skill Workflow
### 1. Explore
Skill first lets AI read `CONTEXT.md` and `docs/adr/`, and then uses `subagent_type=Explore` to send a sub-agent to the code base.
Instead of rigid inspiration, use **friction** as a signal:
> * Where does understanding one concept require bouncing between many small modules?
> * Where are modules **shallow** — interface nearly as complex as the implementation?
> * Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
> * Where do tightly-coupled modules leak across their seams?
> * Which parts of the codebase are untested, or hard to test through their current interface?
Every time you find a suspicious point, apply a deletion test: will deleting it make the complexity disappear or spread out? The answer "disperse" is a candidate worthy of deepening.
### 2. Present Candidates (Statement Candidates)
Present a numbered list of candidates:
```
1. Files: src/orders/parser.ts, src/orders/validator.ts, src/orders/normalizer.ts
Problem: 三个文件互相调用,理解 Order 入站需要在三处跳转
Solution: 合并为单一 OrderIntake 模块,对外只暴露 parse(raw) → ValidatedOrder
Benefits:
- Locality: Order 入站的所有逻辑、错误处理、bug 修复集中一处
- Leverage: 调用方从理解 3 个接口降为 1 个
- Tests: 只需测 parse() 的输入输出,不再需要 mock 内部协作
```
Requirements:
* Use **CONTEXT.md vocabulary** to talk about domains ("the Order intake module", not "the FooBarHandler")
* Talk about architecture using **Glossary vocabulary** ("seam", "depth", "locality")
* **Don’t propose interface designs right away** – let users pick interesting candidates first
If a candidate conflicts with an existing ADR - only mention it if the conflict warrants revisiting the ADR, and clearly mark it:
> "contradicts ADR-0007 — but worth reopening because…"
Don't dig out every refactoring that is prohibited by ADR.
### 3. Grilling Loop
After the user selects a candidate, drop into grilling mode (inherited from [`/grill-with-docs`](/en/docs/notes/matt-pocock-skills/grill-with-docs)):
* Walk through the design tree - constraints, dependencies, the shape of the module after deepening, what is hidden behind the seams, which tests can survive
* **Side effects occur immediately**:
* Give the deepening module a name that is not in CONTEXT.md → Add it to CONTEXT.md immediately
* An ambiguous term was sharpened in the torture → Update CONTEXT.md immediately
* User rejects candidate with reason load-bearing (critical, something future explorers need to know) → Propose to generate ADR
* Want to explore the various interface designs of deepening modules → Jump to `INTERFACE-DESIGN.md` separate process
Document maintenance and architecture transformation happen in the same conversation - no two rounds.
## Real case: Mejba Ahmed’s practice
Third-party developer Mejba Ahmed wrote an article \["Deep Modules: The Claude Code Skill Saving My Codebase"] ([https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules](https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules)) to record in detail his experience using this skill. Key points:
* He originally had 50+ files in a project, each file having less than 100 lines - **Typical shallow library**
* `/improve-codebase-architecture` ran out 8 deepening candidates
* He selected 3 deepenings (two data processing modules merged, one toolset merged)
* Result: Number of files dropped from 50+ to 30+, but **the total code size remains basically the same** - complexity is squeezed into a small number of deep modules
* The hit rate of Claude's subsequent code changes in this library was significantly improved (he said "from 60% to 90%", which was not strictly measured, but it felt strong)
Mejba also has a reminder: **Don’t deepen 8 at a time**. Only pick one at a time, run the test + commit + observe, and then pick the next one. Otherwise, there is no way to roll back once you complete it.
## How to install and use
```bash
npx skills@latest add mattpocock/skills
```
Check `improve-codebase-architecture` + `setup-matt-pocock-skills`.
**Call**: `/improve-codebase-architecture`
**Recommended rhythm**:
* **Run once a week or at the end of each sprint**
* Or \*\*run it once after completing a wave of intensive development (it is especially easy to pile up shallow modules after high-frequency writing of AI code)
* **Don't run when you're in a rush** - it will suggest big changes that you won't have time to digest when you're in a rush
**Typical process**:
1. `/improve-codebase-architecture`
2. AI exploration + list N candidates (with deletion test argument)
3. You pick the one that feels the most
4. drop into grilling loop alignment design
5. AI implementation refactoring (it is recommended to run together with [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd) - the refactoring must have test protection)
6. commit + observe
7. Come again in a week
## Why is this Skill a "closed loop" of Matt's workflow?
Back to Matt’s workflow diagram:
```
/grill-me → /to-prd → /to-issues → /tdd → /improve-codebase-architecture → 回到 /grill-me
```
Notice that it loops back to the starting point. `/improve-codebase-architecture` is not a one-time tool, it is a **periodic maintenance** - because:
1. AI continues to add shallow modules to the code base (this is its default tendency, and it will pile up if you write too much)
2. As business continues to evolve, old seams will become obsolete.
3. The terms in CONTEXT.md continue to be sharpened, and the old naming will not keep up.
**Every time you run this skill, the AI friendliness of the code base is refreshed**. This is the **only** way to keep a long-term codebase healthy with LLM - if you don't refresh it, the AI will be dead in your codebase after three months.
## This set of thinking is more valuable than the Skill itself
Even if you don’t install `/improve-codebase-architecture` at all, just remember the following three things, and the quality of PR review will be improved by a notch:
1. **deletion test**: Every time you see a new module, ask yourself "If you delete it, will the complexity disappear or spread out?"
2. **True seams at least two adapters**: single implementation interface = false abstract, delete
3. **The interface is the test surface**: cannot be measured = there is a problem with the interface design
These three items do not require AI or skill—they are the hard currency of engineering aesthetics. Matt wraps them into skills for batch execution, but the real leverage is the three principles themselves.
## Notes
**Don't go too deep**. Ousterhout himself said that deep module is a goal rather than a dogma - a large Util class that crams all the functions into it is not a deep module, but a god module. The judgment criterion is "simple interface + cohesive implementation", both of which must be met.
**deepening must have test protection**. Structural changes are high-risk operations, and daring to refactor without testing = waiting to take the blame. If there are no tests currently, go to [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd) to add tests to the critical path and then come back.
**ADR decisions should not be made on a whim**. When you reject a candidate during grilling, the AI will easily suggest generating an ADR - only accepting it if the reason is really "future people need to know". Otherwise, docs/adr/ will be filled with journal entries.
**It doesn’t matter if some of the code is shallow**. A logger wrapper, a constant file, a one-off script - they're shallow, no problem. This skill looks for those shallow modules that pretend to help you with abstraction but are actually adding chaos.
## Reference resources
***
## Series conclusion
I have read all 6 articles so far. Review the entire workflow:
```
/grill-me 或 /grill-with-docs ← 谈清楚要做什么
↓
/to-prd ← 凝固成 PRD
↓
/to-issues ← 切成 vertical slice
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture ← 周期性深化
↓
回到 /grill-me
```
The spirit of this process can be condensed into one sentence:
> \*\*AI is the tactical soldier on the ground, and you are the strategic layer. Take back the three things of "defining the problem", "deconstructing the problem" and "testing the problem" and do it yourself, and leave "writing code" to AI - this is the true position of engineers in the AI era. \*\*
Matt's set of skills is not the ultimate answer, but the best practice at the current stage. There may be something better three months later, but the spiritual aspect will not change: a good code base is always more important than a bad code base, and basic software skills are always valuable.
Go back to [Overview](/en/docs/notes/matt-pocock-skills/overview), or choose the most useful skill and install it to try it out.
# Software fundamentals are more important than ever: Matt Pocock’s Claude Code skill set
## A person who is overwhelmed by the specs-to-code wave but still calm
2026 is the highlight year of the "specs-to-code" narrative of AI programming - write the specification, run the compiler, do not read the code, then write the specification, and then run the compiler. The slogan that has emerged in the community is "**code is cheap**" (code is cheap), which means: Anyway, AI can generate another 10,000 lines per second, why should you care about it.
Matt Pocock is one of the few people to publicly speak out against this. He does not deny that AI coding is very powerful, but he actually tested specs-to-code in his class "Claude Code for Real Engineers", and the conclusion is very heart-wrenching: **Every time I run it, the code gets worse and worse**. This is exactly the "software entropy" that has been talked about in Pragmatic Programmer - software entropy increase.
So he did two things:
1. Package this observation into an 18-minute talk: *Software Fundamentals Matter More Than Ever*.
2. Package the corresponding antidote into a GitHub repository: [`mattpocock/skills`](https://github.com/mattpocock/skills) - "Skills for Real Engineers. Straight from my .claude directory."
The warehouse was launched on February 3, 2026, and within 4 months it reached **61.1k stars and 5.3k forks**. It was one of the fastest growing AI programming warehouses during the same period.
***
## Who is Matt Pocock?
If you have written TypeScript, you have probably come across it. He is one of the most prolific TypeScript educators in the Chinese and English circles in recent years:
* Founder of **TotalTypeScript.com**, a series of paid courses that are very popular in the English circle
* **aihero.dev** Newsletter 60,000+ subscriptions, topic changed from TS to AI Coding
* There are a lot of short video tutorials on Twitter [`@mattpocockuk`](https://twitter.com/mattpocockuk) and YouTube `@mattpocockuk`
* Not an OpenAI/Anthropic person, purely an independent developer + educator background
His personality is very clear: **AI Coding from the perspective of a senior engineer**. We don’t shout “AGI is coming”, nor do we shout “Programmers are going to lose their jobs”. What he shouted was, "The tricks of the older generation of software engineers are still very useful, they just need to be translated into a form that LLM can execute."
***
## Core argument: Code is not cheap
There is only one argument in the entire speech, and each skill is a footnote to it:
> If your code base structure is bad, AI will only write bad code in a bad code base. So **a good code base is more important than ever, and basic software skills are more important than ever**.
Matt used a military analogy to explain the roles of humans and AI very straightforwardly:
What does the strategic layer do? Design concepts, unified language, module boundaries - these three things are "defining problems" rather than "writing code", and they happen to be what LLM is least good at doing for you.
***
## Five Failure Patterns → Five Old Books → Five Skills
Matt compressed the entire methodology into a mapping table in his speech. Every time you encounter a failure mode, he will point you back to the classic theory that was solved 20 years ago, and then give you a Skill file in Markdown format:
| # | AI programming failure mode | Classic theory and source | Corresponding Skill |
| - | ------------------------------------------- | ----------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| 1 | AI doesn’t do what you want | *The Design of Design* (Brooks) - design concept, design tree | [`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me) |
| 2 | AI talks to you in a bunch of verbose terms | *Domain-Driven Design* (Evans) - ubiquitous language | [`/grill-with-docs`](/en/docs/notes/matt-pocock-skills/grill-with-docs) |
| 3 | AI does it right but can't run | *The Pragmatic Programmer* (Hunt & Thomas)—— "rate of feedback is your speed limit" | [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd) |
| 4 | AI wanders around in bad code bases | *A Philosophy of Software Design* (Ousterhout) - deep modules, deletion test | [`/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture) |
| 5 | Your brain can’t keep up with AI output | Kent Beck —— invest in design every day | "design the interface, delegate the implementation" |
Item 5 There is no separate skill in the repo (there used to be `design-an-interface` but it is deprecated), its spirit has been absorbed into [`/to-prd`](/en/docs/notes/matt-pocock-skills/to-prd-and-issues) and [`/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture) - both of which force you to think about module interfaces before writing code.
***
## 5 Skills that you really use every day
The speech was a philosophical skeleton. Matt later posted an article "5 Agent Skills I Use Every Day" on aihero.dev to translate the skeleton into daily workflow. These 5 are the objects that will be dismantled one by one in the future of this series:
```
/grill-me ← 先和 AI 谈清楚要做什么
↓
/to-prd ← 把对话凝固成 PRD
↓
/to-issues ← 把 PRD 切成可独立领取的 vertical slice
↓
/tdd ← 每个 slice 用红绿重构跑通
↓
/improve-codebase-architecture ← 周期性检查,把 shallow 模块改成 deep
```
These 5 skills are strung together to form Matt’s complete research and development process. The failure modes corresponding to each step are shown in the table in the previous section.
Detailed dismantling of each article (subsequent pages in this series):
* [Grill Me: Let AI torture you about your needs](/en/docs/notes/matt-pocock-skills/grill-me)
* [Grill With Docs: Maintaining project language and ADR](/en/docs/notes/matt-pocock-skills/grill-with-docs)
* [to-PRD + to-Issues: from conversation to executable ticket](/en/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [TDD: Force AI to take small steps with red-green reconstruction](/en/docs/notes/matt-pocock-skills/tdd)
* [Improve Codebase Architecture: Restructure shallow into deep modules](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture)
***
## How to install
The warehouse README gives a one-line command to install:
```bash
npx skills@latest add mattpocock/skills
```
This command will:
1. Let you check which skills you want to install
2. Allows you to select which agents to install (Claude Code, Codex, Cursor, etc. are all supported)
3. Put the corresponding SKILL.md file into `.claude/skills/` (or the directory corresponding to the agent)
**It is strongly recommended to check `/setup-matt-pocock-skills`** at the same time - this is a one-time configuration skill that will ask you three questions:
* What is **Issue tracker** used for? (GitHub / GitLab / local markdown / other)
* What word is used for **Triage label**? (needs-triage or other)
* Where to put **Domain doc**? (CONTEXT.md/ADR path)
Run `/setup-matt-pocock-skills` once and it will write to `AGENTS.md` or `CLAUDE.md` in the root directory of your project. After that, all engineering skills (to-prd, to-issues, triage, tdd, etc.) will automatically read this configuration. This step is omitted, and every subsequent skill will ask you the same question again and again.
If you just want to try `/grill-me` (the most lightweight, pure productivity class), you can skip setup as it does not rely on the issue tracker.
***
## The difference between this set of skills and BMAD / Spec-Kit / GSD
If you are already using spec-driven frameworks such as [BMAD](/en/docs/notes/speckit/concept), Spec-Kit, and GSD, you may ask: "Why do you still need Matt's set?"
Matt writes very directly in the README:
**Core Differences**:
* BMAD/Spec-Kit/GSD is a **framework** that specifies a complete pipeline from spec to code. You have to follow its process.
* Matt This set is **component**. Each skill has a markdown file, ranging from a few lines to dozens of lines. You can disassemble and modify it at any time.
Example: The actual full text of `grill-me` is only this short——
```markdown
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
The entire skill is 7 lines. But it’s those seven lines that make Claude ask you 20, 50, or even 100 questions before making a decision. This design philosophy of **using very little text to leverage big behavioral changes** is the fundamental reason for the popularity of this set of skills.
***
## How to read this series
If you haven’t come across Matt’s set before, it is recommended to read it in the order in meta.json:
1. **Overview** (the article you are currently on) - Get the full picture
2. **Grill Me** - Install one individually and try it first, the threshold is the lowest
3. **Grill With Docs** - an advanced version of grill-me, starting to introduce CONTEXT.md
4. **to-PRD + to-Issues** - Turn the conversation into an executable ticket
5. **TDD** - Matt himself said "the most stable method I have ever used to improve the quality of agent output"
6. **Improve Codebase Architecture** - Periodic maintenance to make AI available for the long term
If you are already using Claude Code to write real projects, jump directly to [`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me) + [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd). These two articles are the most intuitive.
If you are teaching or writing, just reading [`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me) is enough - it is a general "design conversation" tool, not limited to code.
***
## My usage suggestions
After I installed this set of skills myself, the biggest physical changes were:
**First**: Stop rushing to start writing code. In the past, when AI received "Add me a login", it would start laying out 500 lines. Now `/grill-me` will ask you 20 questions first - "Do you want to remember the device?" "How long does it take for the session to expire?" "How many times have you failed to lock your account?" Let it be written after 30 minutes. What you will save is the subsequent two hours of rework.
**Second**: CLAUDE.md is no longer bloated. In the past, there were a lot of prohibitions written in CLAUDE.md such as "Please understand the requirements before writing code" and "Don't be overly abstract", but Claude still committed it. After switching to Matt's set, CLAUDE.md only puts domain knowledge (design system, component specifications, deployment), and general methodologies are handed over to skills. The responsibilities of both parties are clear.
**Third**: Deep module thinking is more valuable than the skill itself. Even if you don't install `/improve-codebase-architecture`, just reading the "**deletion test**" in its SKILL.md (if the complexity disappears after deleting this module, it means it is pass-through) will already make you take a second look during PR review.
**Note the cost**:
* After installing 5 skills, the AI will ask more questions. People who are used to "generating 500 lines in one sentence" will find it annoying
* After `/tdd` is strictly implemented, simple scripts will also be required to write tests first, which is not friendly to exploratory code - you can tell it "skip TDD this time"
* `/grill-with-docs` will take the initiative to modify your CONTEXT.md. It is best to dry-run it before running it for the first time.
***
## Reference resources
**5 books cited in the speech** (in order of appearance):
* *A Philosophy of Software Design* — John Ousterhout (complexity definition, deep modules)
* *The Pragmatic Programmer* — David Thomas & Andrew Hunt(software entropy、outrunning headlights)
* *The Design of Design* — Frederick P. Brooks(design concept、design tree)
* *Domain-Driven Design* — Eric Evans(ubiquitous language)
* *Test-Driven Development* — Kent Beck(invest in design every day)
Each one is over 20 years old. Matt repeated a sentence many times in his speech: "**Go on Amazon, get it.**" - this sentence itself is the Easter egg of this speech.
# Other Skills: Compressed Communication, Handoff, Teaching, Writing Skills, and Safety Guardrails
## Why not elaborate on each one
Matt's README categorizes skills into three types:
* Engineering: For writing real code daily
* Productivity: General workflow tools
* Misc: Small tools he keeps for personal use
The previous articles have covered the main engineering skills. This article combines the remaining productivity and misc skills because most of them are not full development processes, but rather **small, useful switches for specific scenarios**.
If you're only installing 5, I still recommend prioritizing:
* [`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me)
* [`/grill-with-docs`](/en/docs/notes/matt-pocock-skills/grill-with-docs)
* [`/to-prd` + `/to-issues`](/en/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd)
* [`/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture)
But if you've already got the main skills running, the following will make your daily experience smoother.
## Productivity Skills
### caveman: Extreme Communication Compression
`/caveman` is a "low token mode." It asks the agent to remove pleasantries, filler words, over-explanation, and vague buffering, retaining only technical information.
It's suitable when:
* You're iterating frequently and don't want to read long responses
* Debugging and only need facts, reasons, and next steps
* Long contexts are almost full, and output needs to be compressed
* You want to force the AI to be less verbose
It doesn't make the AI rude, but rather makes it use shorter syntax while retaining full technical precision. Note that it remains active until you explicitly exit it.
I would treat it as a temporary setting, not a long-term default. For high-risk operations, safety warnings, and complex multi-step instructions, being too brief can lead to misinterpretation.
### handoff: Handing off the current conversation to the next agent
The goal of `/handoff` is to compress the current conversation into a handoff document and save it to the system's temporary directory, rather than polluting the current workspace.
It will include:
* Current goals
* Decisions made
* Key paths and files
* Remaining tasks
* Suggestions for which skills the next agent should call
* Redacted sensitive information
It's especially suitable when a long task is interrupted, the context is almost full, or you want to switch to a different agent.
The key point is: do not duplicate content already present in PRDs, issues, ADRs, commits, or diffs; only reference paths or URLs. The value of a handoff document is to **fill in the state scattered across the conversation**, not to recreate project documentation.
### teach: Turning the current directory into a learning workspace
`/teach` is the heaviest of this group. It treats the current directory as a long-term learning workspace, maintaining:
* `MISSION.md`: Why you are learning this topic
* `RESOURCES.md`: A list of high-quality resources
* `learning-records/*.md`: Learning records, similar to ADRs
* `lessons/*.html`: Interactive lessons, one per session
* `reference/*.html`: Quick reference materials
* `NOTES.md`: Teaching preferences and work notes
Its highlight is treating learning as a long-term system, not a one-off Q\&A. It particularly emphasizes:
* Mission first: Why you learn is more important than what you learn
* Retrieval practice: Using recall exercises to build long-term memory
* Spacing / interleaving: Don't be fooled by short-term fluency
* High-trust resources: Find resources first, don't rely solely on the model's memory
If you're just asking "explain X," you don't need it; if you want to study a topic for several weeks, it's very suitable.
### write-a-skill: Scaffolding for writing new skills
`/write-a-skill` is Matt's abstraction of the skill structure itself.
It requires a skill to have at least:
```text
skill-name/
├── SKILL.md
├── REFERENCE.md
├── EXAMPLES.md
└── scripts/
```
Of course, the latter three are not mandatory and are only added when the content is too long, examples are valuable, or operations can be scripted.
Its most important judgment is: `description` is the only information the agent sees first when deciding whether to load a skill. Therefore, the description cannot be vague like "helps process documents"; it must state:
* What capabilities it provides
* When it should be triggered
* What the trigger words or context are
This aligns with my own experience writing skills: many skills fail not because the main body is poorly written, but because the description is too generic, and the agent doesn't know it should load it.
## Misc Skills
### git-guardrails-claude-code: Blocking dangerous git commands
This skill equips Claude Code with a `PreToolUse` hook to intercept dangerous git commands before executing Bash.
By default, it blocks:
* `git push`
* `git reset --hard`
* `git clean -f` / `git clean -fd`
* `git branch -D`
* `git checkout .` / `git restore .`
Its value is direct: preventing the agent from pushing, hard resetting, or cleaning untracked files without your authorization.
If you frequently let AI work in real repositories, this skill is well worth installing. It's not about distrusting AI, but about intercepting highly destructive operations at the tool layer, rather than relying on prompts.
### setup-pre-commit: Adding pre-commit checks to a project
`/setup-pre-commit` will set up:
* Husky pre-commit hook
* lint-staged + Prettier
* typecheck
* test
It will first detect the package manager, then decide what to run in the pre-commit hook based on the project's existing scripts. If `typecheck` or `test` are not present, it won't create them but will omit them and inform you.
The value of this skill is not in the configuration itself, but in Matt's quality philosophy: **don't just let the AI say the code is fine; make it pass deterministic checks**.
### migrate-to-shoehorn: Reducing `as` in tests
This is a very Total TypeScript-style small tool. It migrates TypeScript `as` type assertions in tests to `@total-typescript/shoehorn`.
Typical replacements:
| Old Way | New Way | Scenario |
| --------------------------- | ------------------ | ---------------------------------------------------------- |
| `obj as Request` | `fromPartial(obj)` | Only interested in a few fields of a large object in tests |
| `obj as unknown as Request` | `fromAny(obj)` | Intentionally passing wrong types to test error paths |
| Complete object mock data | `fromExact(obj)` | Requires enforcing a complete shape |
It is explicitly for test code, not production code.
This skill is narrow, but it aligns with Matt's engineering taste: don't create 20 meaningless fields in tests for the sake of the type system, and don't completely turn off type safety with bare `as`.
### scaffold-exercises: Generating exercise directories for course repositories
This skill clearly comes from Matt's own course creation workflow. It creates directories according to specifications:
```text
exercises/
└── 05-memory-skill-building/
└── 05.02-short-term-memory/
├── explainer/
├── problem/
└── solution/
```
Each subdirectory must have a non-empty `readme.md`, `main.ts` if necessary, and must pass `pnpm ai-hero-cli internal lint`.
It's not useful for most engineering projects, but it's very practical for courses, bootcamps, and exercise repositories. More importantly, it demonstrates a characteristic of a good skill: **delegating repetitive, mechanical, and detail-prone formatting work to the agent**.
## Directories not recommended for the main line at present
The upstream repository also contains `deprecated/`, `in-progress/`, and `personal/`.
I suggest not writing them into formal usage guides for now:
| Directory | Why not include in the main line |
| -------------- | -------------------------------------------------------------------------------- |
| `deprecated/` | Deprecated, easily misleads readers into adopting old processes |
| `in-progress/` | Still experimental, behavior and naming may change |
| `personal/` | More like Matt's private workspace, not necessarily suitable for general readers |
If they are to be written about in the future, a separate article on "Matt Pocock skills repository archaeology" could be made, rather than mixing them with stable recommendations.
## Commonalities of this set of small tools
These skills may seem scattered, but they share a common principle:
> Turn things that agents tend to drift on into small, well-defined work patterns.
* `caveman` prevents communication drift
* `handoff` prevents context loss
* `teach` prevents learning from becoming a one-off Q\&A
* `write-a-skill` prevents ad-hoc skill structures
* `git-guardrails` prevents dangerous commands from relying on self-awareness
* `setup-pre-commit` prevents quality checks from relying on AI's self-description
* `migrate-to-shoehorn` prevents test type assertions from getting out of control
* `scaffold-exercises` prevents manual omissions in course structures
This is also the most valuable aspect of Matt's repository to learn from: skills don't need to be grand. A small, frequent deviation, if it can be consistently corrected by 20 lines of instructions, is worth writing as a skill.
## References
# Prototype: Answering a Design Question with Throwaway Code
## Prototypes Are Not "Just Write Something Quick"
The first sentence of the `/prototype` definition is crucial:
> A prototype is throwaway code written to answer a question.
This separates prototypes from "lazy implementations." A prototype is not a precursor to production code, nor a half-finished product that will be gradually refactored into a formal version. It should be marked as throwaway from day one.
Therefore, the key to `/prototype` is not writing fast, but first clarifying:
> What question is this prototype intended to answer?
## Two Branches
Matt categorizes prototypes into two types, with completely different outputs.
| Question to Answer | Branch | Output |
| -------------------------------------------------------- | --------------- | ----------------------------------------------- |
| Are the logic, state machine, or data models reasonable? | Logic prototype | A runnable terminal application |
| What should this interface look like? | UI prototype | Multiple UI solutions switchable within a route |
This distinction is highly practical. Many teams say "let's build a prototype" without clarifying whether they want to validate the visual appearance or the state transitions. These two goals require entirely different approaches.
## Logic Prototype: Exposing State in the Terminal
If the question is "Is this state machine correct?" or "Can these business rules be implemented?", the prototype should be a small command-line program.
Its characteristics:
* In-memory state, no reliance on a real database.
* Starts with a single command.
* Prints the complete relevant state after each operation.
* Covers edge cases that are difficult to deduce on paper.
* No tests, no exception handling, no framework abstraction.
Example: You're designing a subscription state flow.
Don't modify production code directly. First, create a `subscription-prototype.ts` that allows users to select options in the terminal:
```text
1. start trial
2. pay
3. cancel
4. expire
5. refund
6. print state
```
After each step, print the current entitlement, trial quota, paid state, and next renewal. You'll quickly discover that some state combinations haven't been fully considered.
The value of this type of prototype is: **turning abstract rules into operable objects**.
## UI Prototype: Multiple Aggressive Solutions in a Single Route
If the question is "How should the interface be designed?", the prototype should generate several significantly different UI variations, rather than tweaking the same solution three times.
The `/prototype` UI branch requires:
* Multiple variations within a single route.
* Switching via URL search parameters or a bottom floating switch bar.
* Clear differences between the proposed solutions.
* Prototype code that is close to the future real page, but with clear naming indicating it's a prototype.
* Avoid premature integration with real data and persistence.
The difference from standard AI UI generation is that it doesn't ask the AI for the "best solution" at once. Instead, it allows you to compare several directions using a real browser.
For example, for a dashboard's empty state, don't just ask the AI to change the copy. Have it generate:
* A: Table-based, dense, emphasizing next actions.
* B: Task-oriented, with a checklist on the left and a preview on the right.
* C: Guided, highlighting a primary CTA and historical examples.
Then, you can switch between these within the same route, rather than trying to visualize them from three screenshots in a chat.
## All Prototypes Must Be Deletable
Among the general rules for `/prototype`, the most important is "deletable":
* Filenames or paths should clearly indicate it's a prototype.
* Do not default to connecting to a production database.
* Do not write generic abstractions.
* Do not implement excessive error handling.
* Delete it after completion, or absorb the learned conclusions into the formal code.
If a prototype cannot be deleted, it has already become a liability in production code.
This is particularly important in AI programming. AI is adept at writing prototypes that "look usable," leading humans to be reluctant to delete them, resulting in a codebase filled with temporary code that no one dares to touch.
## What Remains After the Prototype
Prototype code is not worth keeping, but the answers are.
Matt suggests documenting the following in a persistent location:
* The question the prototype aimed to answer.
* Observed conclusions.
* The chosen direction.
* Directions that were abandoned.
* If necessary, convert these into ADRs, issues, PRDs, or commit messages.
In other words, the output of `/prototype` is not code, but **decisions**.
## When Not to Use It
Do not use `/prototype` in these scenarios:
* Requirements are already clear and only implementation is needed.
* Bugs have reproducible steps; use [`/diagnose`](/en/docs/notes/matt-pocock-skills/diagnose-and-triage) instead.
* Refactoring direction is clear; use [`/tdd`](/en/docs/notes/matt-pocock-skills/tdd) for protection during implementation.
* UI polish is minor and doesn't warrant multiple solution variations.
* You don't have time to delete or absorb the prototype.
The cost of a prototype is not in writing it, but in its conclusion. Without a conclusion, don't start.
## A Useful Prompt
You can invoke it like this:
```text
/prototype
I want to verify if this checkout state machine is reasonable. Please use the logic branch.
Only create a throwaway terminal prototype, do not connect to a real DB.
Print the complete state after each operation.
```
Or:
```text
/prototype
I want to compare 3 information architectures for the project details page. Please use the UI branch.
Place it in a prototype route within the existing routing system, providing a bottom switch bar.
Do not modify production components.
```
The most important aspect here is clarifying "the question to be answered." As long as this question is clear, the prototype is less likely to go off track.
## Relationship with Grill Me
[`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me) is suitable for refining decisions through questioning; `/prototype` is suitable for refining decisions through hands-on experience.
Some questions can be resolved by asking, such as "Should anonymous comments be moderated?". Some questions require hands-on testing, such as "Will this drag-and-drop sorting state machine actually be difficult to use?". The latter is where prototyping is appropriate.
Therefore, I place it at a branching point in the workflow:
```text
Vague Idea
↓
/grill-me
↓
If experience or validation is still needed
↓
/prototype
↓
Retain conclusions, delete prototype
↓
/to-prd or /tdd
```
## Resources
Next article: [Improve Codebase Architecture: Refactoring Shallow to Deep Modules](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture).
# Setup Matt Pocock Skills: Clarify Project Rules First
## This Skill Solves More Than Just Installation
`/setup-matt-pocock-skills` can easily be misunderstood as an "initialization command to run after installation." In reality, it's more like a **project contract generator**: it informs subsequent skills about how the repository tracks tasks, how issues are labeled, and where to find domain language and architectural decisions.
Matt specifically reminds users in the README's Quickstart to select `/setup-matt-pocock-skills` during installation and then run it within the agent. The reason is simple: `to-prd`, `to-issues`, `triage`, `diagnose`, `tdd`, `improve-codebase-architecture`, and `zoom-out` all require the same project context. If each skill were to ask for this information temporarily, the workflow would become fragmented.
It doesn't "configure Claude's preferences," but rather answers three engineering questions:
| Question | What it clarifies | Who uses it later |
| ---------------------------------- | ------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------- |
| Where is the issue tracker? | GitHub, GitLab, local markdown, or another system | `to-prd`, `to-issues`, `triage` |
| How are triage labels mapped? | Which real labels correspond to roles like `needs-triage`, `needs-info`, `ready-for-agent`, etc. | `triage` |
| Where is the domain documentation? | A single `CONTEXT.md`, or multiple contexts like `CONTEXT-MAP.md` + partitioned ADRs | `grill-with-docs`, `diagnose`, `tdd`, `zoom-out`, `improve-codebase-architecture` |
## Why It's Important
The core idea behind this set of skills is "small and composable." The trade-off for being small is that these skills don't want to take over the entire project workflow, so they must understand your project's actual conventions.
For example, `/to-issues` needs to create an issue. Without setup, it wouldn't know whether to:
* Call `gh issue create`
* Call `glab issue create`
* Write to `.scratch//`
* Or generate text that can be copied for Linear/Jira
Another example: `/triage` needs to move an issue to `ready-for-agent`. If your repository's actual label is `ai:ready`, and the skill creates a new label `ready-for-agent`, your issue tracker will immediately become cluttered.
Therefore, the value of `/setup-matt-pocock-skills` isn't automation, but **making implicit conventions explicit**.
## What It Reads
This skill begins by exploring the repository rather than making assumptions:
* `git remote -v` and `.git/config`: To determine if it's a GitHub/GitLab project.
* Root directory `AGENTS.md` / `CLAUDE.md`: To check if a `## Agent skills` section already exists.
* Root directory `CONTEXT.md` / `CONTEXT-MAP.md`: To determine the format of domain language documentation.
* `docs/adr/` and `src/*/docs/adr/`: To check if ADRs are global or module-specific.
* `docs/agents/`: To see if setup has already been run.
* `.scratch/`: To check if local markdown issue conventions already exist.
This aligns with Matt's entire workflow style: **first, examine the project's actual state, then write the rules.**
## Three Decisions
### 1. Issue Tracker
This is where subsequent work items will be landed.
The default preference is GitHub, as this skill set was originally designed around GitHub Issues. However, it treats GitLab and local markdown as first-class options:
| Choice | Suitable for what scenario |
| -------------- | ------------------------------------------------------------------------------------- |
| GitHub | Open-source projects, existing GitHub issue workflow |
| GitLab | Company projects on GitLab, accustomed to using `glab` |
| Local markdown | Personal projects, temporary exploration, no remote issue tracker |
| Other | Jira, Linear, Feishu, Notion, etc. Requires documenting your actual workflow in text. |
The key is not to pick one, but to choose the **one your team actually uses**. If you choose incorrectly, subsequent skills will create tasks in the wrong system.
### 2. Triage Label Vocabulary
`/triage` internally uses 5 status roles:
| Role | Meaning |
| ----------------- | --------------------------------------------- |
| `needs-triage` | Awaiting maintainer judgment |
| `needs-info` | Awaiting reporter to provide more information |
| `ready-for-agent` | Clear enough to be handed off to an AFK agent |
| `ready-for-human` | Requires human judgment or implementation |
| `wontfix` | Will not be addressed |
Setup will ask you for the real label names corresponding to these roles. If your project doesn't have existing labels, using the default names is fine; if you already have your own naming system, you should map them here rather than letting the skill create a new set.
### 3. Domain Docs
This is the biggest difference between Matt's skills and ordinary prompts: they don't just look at the current conversation, but also read the project's **domain language** and **architectural decisions**.
The simplest form:
```text
/
├── CONTEXT.md
└── docs/
└── adr/
```
A large monorepo can use multiple contexts:
```text
/
├── CONTEXT-MAP.md
├── apps/
│ └── web/
│ ├── CONTEXT.md
│ └── docs/adr/
└── services/
└── billing/
├── CONTEXT.md
└── docs/adr/
```
Setup doesn't force you to complete all documentation now; instead, it tells subsequent skills where to look and how to create them if they are not found.
## What It Writes
Ultimately, there will be two types of output.
The first type is the `## Agent skills` section in `AGENTS.md` or `CLAUDE.md`:
```markdown
## Agent skills
### Issue tracker
...
### Triage labels
...
### Domain docs
...
```
The second type consists of three files under `docs/agents/`:
| File | Content |
| ------------------------------ | ---------------------------------------------------------- |
| `docs/agents/issue-tracker.md` | Issue system, commands, conventions for creation/updates |
| `docs/agents/triage-labels.md` | Mapping from canonical roles to real labels |
| `docs/agents/domain.md` | Rules for reading `CONTEXT.md`, `CONTEXT-MAP.md`, and ADRs |
Note that it will prioritize editing an existing `CLAUDE.md`; only if `CLAUDE.md` doesn't exist will it consider `AGENTS.md`. This reflects an important restraint: **do not create two competing entry points for agent rules within a project.**
## How I Recommend Using It
When installing Matt's skills for the first time, the sequence should be:
1. Install: `npx skills@latest add mattpocock/skills`
2. Select `/setup-matt-pocock-skills`
3. Run `/setup-matt-pocock-skills`
4. Answer the three questions about issue tracker, labels, and domain documentation based on your actual project state.
5. Review the generated `## Agent skills` and `docs/agents/*.md` files.
6. Then start using [`/grill-with-docs`](/en/docs/notes/matt-pocock-skills/grill-with-docs), [`/to-prd`](/en/docs/notes/matt-pocock-skills/to-prd-and-issues), or [`/triage`](/en/docs/notes/matt-pocock-skills/diagnose-and-triage).
If you just want to experience [`/grill-me`](/en/docs/notes/matt-pocock-skills/grill-me) in isolation, you can skip setup; but as soon as you enter the engineering workflow, it's best to do this first.
## Design Inspiration for This Skill
`/setup-matt-pocock-skills` might seem simple, but it solves a common problem in agent workflows: **rules are scattered in people's minds.**
Many teams treat information like "which labels do we use," "which issues can be given to AI," and "where is CONTEXT.md" as verbal agreements. People know, but AI doesn't. When AI doesn't know, it repeatedly asks, or worse: guesses.
The purpose of setup is to turn these verbal agreements into readable files. Subsequent skills don't need to be smarter; they just need to reliably read the same project contract.
This is also why I think it deserves its own article: it's not a flashy skill, but it's the foundation that allows the entire workflow to run smoothly long-term.
## Resources
Next article: [Grill Me: Let AI Grill You with 50 Questions Before You Write Code](/en/docs/notes/matt-pocock-skills/grill-me).
# TDD: Use red-green refactoring to force AI to take small steps
## Failure mode: "AI does the right thing, but it can't run"
The third failure mode in Matt's talk: **The direction is right, but it doesn't work**.
The most direct fix is to install feedback infrastructure for AI:
* TypeScript (no static typing *is crazy*)
* Allow LLM to access the browser and view the page by itself
* Automated testing
But Matt observed one thing: **Even with this feedback installed, LLM doesn't work well**. It tends to write 500 lines at a time and then think "oh I should type check that". This is what the Pragmatic Programmer calls *outrunning your headlights* - drive faster than the headlights can illuminate, and it's only a matter of time before you hit the wall.
> "The rate of feedback is your speed limit, which means you should be testing as you go, taking small deliberate steps. **And the AI by default is really not very good at that.**"
To fix this problem, you need to force the AI to stop step by step at the tool level. Matt's answer is TDD - **Testing first can force checkpoints**.
## Classic Theory: Kent Beck’s Red-Green Reconstruction
The standard rhythm of TDD is defined by Kent Beck in his 2003 book "Test-Driven Development: By Example":
1. **RED**: Write a failing test (describe what to do)
2. **GREEN**: Write the smallest enough code to make the test pass
3. **REFACTOR**: Improve code structure under test protection
Each loop is extremely short - on the order of minutes. There are automated checks (test pass/fail) at every step.
Matt directly follows this rhythm, but his SKILL.md spends a lot of time talking about an **anti-pattern** - this is the core.
## Key anti-pattern: horizontally sliced red and green
Many people think that TDD means "write all the tests first, then write all the implementations." Matt directly says this is wrong in SKILL.md:
```
WRONG (horizontal slicing):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical slicing via tracer bullets):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
Why is horizontal wrong? SKILL.md gave three reasons:
> 1. Tests written in bulk test *imagined* behavior, not *actual* behavior
> 2. You end up testing the *shape* of things (data structures, function signatures) rather than user-facing behavior
> 3. Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
Human saying: **Writing all the tests in one go is testing what's in your head, not the real code**. When you write impl3, you realize that the design of test1 is wrong - but at this time, test2/test3/test4 are all coupled to the wrong design. Go back and make changes.
The correct approach is to test one implementation and then open the next pair after writing one pair. After each pair is completed and you have learned something from this implementation, the next pair of tests can be designed based on real experience - not imagination.
## Skill full text structure
`engineering/tdd/SKILL.md` is one of the longest skills Matt has written because TDD itself has a lot of nuance. The core structure is as follows:
### Philosophy
> **Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
> **Good tests** are integration-style: they exercise real code paths through public APIs. They describe *what* the system does, not *how* it does it. A good test reads like a specification.
> **Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly). The warning sign: your test breaks when you refactor, but behavior hasn't changed.
Remember a diagnosis: **If you rename an internal function, the test will kneel down - then this test is testing the implementation rather than the behavior, which is a bad test**.
### Workflow (with checklist)
#### 1. Planning
Align with users before writing code:
```
[ ] Confirm with user what interface changes are needed
[ ] Confirm with user which behaviors to test (prioritize)
[ ] Identify opportunities for deep modules (small interface, deep impl)
[ ] Design interfaces for testability
[ ] List the behaviors to test (not implementation steps)
[ ] Get user approval on the plan
```
Key question: "**What should the public interface look like? Which behaviors are most important to test?**"
> "**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case."
This one is very counter-intuitive. By default, AI will want to exhaust all edge cases, but Matt emphasizes **priority** - not all behaviors are worth measuring, and focus firepower on the core path.
#### 2. Tracer Bullet
Write **a** test that verifies **one** thing:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
This is a "tracer bullet" - shoot it first and check the sight. Matt emphasized that this work should be **end-to-end** - not to write the schema first, then write the API and then write the UI, but to cut the thinnest path that runs through the entire stack.
#### 3. Incremental Loop
Repeat for each subsequent behavior RED→GREEN:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
Rules:
* One test at a time
* Write only enough code to pass the current test
* **Don’t predict future tests**
* Testing focuses on observable behavior
"Don't predict" is particularly important. The AI can't help but think, "This function needs to support X anyway, let's add it by the way" - and this starts horizontal slicing.
#### 4. Refactor
After all tests have passed, look for refactoring opportunities:
```
[ ] Extract duplication
[ ] Deepen modules (move complexity behind simple interfaces)
[ ] Apply SOLID principles where natural
[ ] Consider what new code reveals about existing code
[ ] Run tests after each refactor step
```
> **Never refactor while RED.** Get to GREEN first.
Refactoring in red = changing tests and code at the same time = you don’t know whether the test is wrong or the code is wrong. **Green first, then refactor**.
### Per-Cycle Checklist
At the end of each red and green cycle, Matt asks the AI to self-check:
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
These five points are used to identify bad tests and over-implementation. AI self-checking can prevent most common mistakes.
## Real use: from issue to PR
`/tdd` is the next step in Matt's workflow from `/to-issues`. Given a vertical slice issue, the process is:
```
你: 实现 issue #43
↓
/tdd
↓
Claude 读 issue acceptance criteria
↓
Claude 探索代码库 → 找到 CONTEXT.md → 用项目术语
↓
Planning 阶段:
- 列出准备改的接口
- 列出准备测的行为(按优先级排序)
- 让你点头
↓
Tracer Bullet:
- RED: 写第一个测试(基于 acceptance criteria 第 1 条)
- 跑测试,确认 fail
- GREEN: 写最小实现
- 跑测试,确认 pass
↓
Incremental Loop:
- 每个 acceptance criteria 一个 RED→GREEN
↓
Refactor:
- 看 deep module 提取机会
- 每次重构后跑全套测试
↓
PR
```
Every red and green cycle the AI will stop and give you a status - "test fails"/"test passes, here's the diff". **These pauses are the antidote to outrun headlights** - the AI has no chance of laying out a thousand lines in one go.
## About Mock: Matt’s strong opinion
SKILL.md specifically mentions the dangers of mocks - he also provides a separate `mocking.md`. Core ideas:
> "Bad tests... mock internal collaborators."
Mock internal collaborators = 1:1 coupling between testing and implementation = testing team kneels when refactoring. Matt's preference is **integration-style testing** - try to use a real database (in-memory or testcontainers), real HTTP (MSW), and a real file system (tmp dir). Only mock at really expensive or unstable boundaries (like calls to the OpenAI API).
This is contrary to the current situation of many teams - most code libraries are full of unit tests, and there are more mocks than real code. Matt made a judgment in his speech: **Good code base = easy to test code base**. If you have to mock a bunch of things to test, it means there is a problem with the code structure, and you should change the architecture first (go to [`/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture)).
## The new significance of TDD in the AI era
When Kent Beck wrote that book 23 years ago, the core benefit of TDD was "people not writing wrong code." In the AI era, TDD has an additional meaning:
**It is the only "success criterion" that AI can understand**.
Matt quoted Karpathy’s words of wisdom later in his speech:
The best form of "success criteria" is **test** - it is machine-verifiable, binary, and cannot be quibbled with. Giving AI a test suite + "let it pass" is ten times more reliable than giving AI a description of requirements + "please implement".
So `/tdd` is not just a quality assurance tool - it is the input interface to the agent loop. Each red and green cycle is a complete "input → action → feedback". The AI learns the true situation of this implementation in the cycle, and the next cycle will be more accurate.
## How to install and use
```bash
npx skills@latest add mattpocock/skills
```
Check `tdd` + `setup-matt-pocock-skills`.
If you mainly use Codex now, take over the installation product to `.agents/skills/`, and write the project-level workflow, test commands, and issue tracker rules into `AGENTS.md`. The essence of Matt's `/tdd` is the red-green refactoring cycle, not tied to Claude Code.
**Calling method**:
* Direct: `/tdd` - let it infer what to measure from the current conversation context
* Fetch issue: `/tdd implement #43` - it will fetch issue and then open it
* Fix bug: `/tdd reproduce this bug then fix it` - It will first write a failed test that can reproduce the bug, and then fix it
## Notes
**Not suitable for all tasks**. One-off scripts, playground exploration code, UI fine-tuning - don't use TDD, it will slow down the pace. Matt himself said that TDD is suitable for code that "has lasting value and needs to be maintained."
**Prepare the test infrastructure first**. If the project has not installed the test framework (Vitest / Jest / Playwright, etc.), install it first and then use `/tdd`, otherwise it will install it for you first, but there are many questions in that step.
**Don't let it automatically add e2e tests**. e2e is slow and crisp, and the TDD rhythm is minute-level. `/tdd` defaults to integration test rather than e2e, but you can explicitly tell it "unit + integration only, not e2e".
**The reconstruction phase is the easiest to get out of control**. After the AI gets the GREEN status, it will refactor a bunch of things excitedly - stare at it, and run tests after each refactoring. This part is a high-risk area for AI deviation.
## Reference resources
Next article: [Improve Codebase Architecture: Reconstruct shallow into deep modules](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture) - Periodic maintenance allows AI to run in your code base in the long term.
# to-PRD + to-Issues: Condensate the conversation from the grill into an executable vertical slice
## The position of this section in the workflow
Back to Matt’s workflow diagram:
```
/grill-me 或 /grill-with-docs ← 谈清楚
↓
/to-prd ← 凝固成 PRD(你在这里)
↓
/to-issues ← 切成可领取的 vertical slice(你在这里)
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture
```
`/to-prd` and `/to-issues` are a link between the previous and the following: translating abstract dialogue decisions\*\* into executable work units\*\*. I combine them because in Matt's actual use they are two consecutive steps.
## Failure mode: Nothing to do after grilling
Many people get stuck after using `/grill-me`: they get a long conversation with a bunch of decisions, but **how to start writing code**?
Giving Claude a "please do it" directly is wrong - because:
1. Implementing complete functions at one time = AI output 1000+ lines = difficult to review, difficult to test, difficult to locate bugs
2. After AI loses its memory, the next session has no context.
3. No tracking - no way to know where we have reached and how much is left
The correct approach is to freeze the decision into an artifact (PRD), and then cut the PRD into work packages (issues) that are small enough to be completed independently. This has been common knowledge in software engineering for 30 years, but it takes on new meaning in the AI era:
> Chop it finely enough so that the AFK agent (the agent that runs when you are not present) can pick up and complete it independently.
## /to-prd: compress the conversation into PRD
### Key Constraints of Skill
`/to-prd`'s SKILL.md writes a very important sentence at the beginning:
> "This skill takes the current conversation context and codebase understanding and produces a PRD. **Do NOT interview the user — just synthesize what you already know.**"
Ask no more questions. This is what the grill-me stage does, to-prd only does **synthesis**. So **don't clear the context and run to-prd** - it relies on all the dialogue from the previous grill.
### Skill processing flow
1. **Explore the code base** (if you haven’t explored it yet) - use the project’s CONTEXT.md vocabulary and respect existing ADRs
2. **Draft module**——Proactively look for opportunities that can be extracted as deep modules so that the interface can be independently tested
3. **Align modules with users** - "Are these modules correct? Which ones need to be tested?"
4. **Generate PRD** according to the template, send it to issue tracker, and tag it with `needs-triage`
### PRD Template
The template given by Matt is as follows:
```markdown
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list:
1. As a , I want a , so that
2. ...
## Implementation Decisions
- The modules that will be built/modified
- The interfaces of those modules
- Technical clarifications from the developer
- Architectural decisions
- Schema changes / API contracts / Specific interactions
(NO specific file paths or code snippets — they rot fast.)
## Testing Decisions
- What makes a good test (test external behavior, not internals)
- Which modules will be tested
- Prior art (similar tests in the codebase)
## Out of Scope
What's NOT in this PRD.
## Further Notes
```
Several key designs:
* **User Stories account for the majority**: Requirements LONG, numbered list - forcing you to exhaustively enumerate complete function points. This avoids the blind spot of "I thought it was clear"
* **Implementation does not write file paths or code**: Matt directly said "they may end up being outdated very quickly". This is a unique consideration in the LLM era - the specific path will become obsolete immediately after reconstruction, but the "module boundary" and "interface contract" have a longer life cycle
* **Must have Out of Scope**: This paragraph is ignored in most PRD templates, but it is the boundary insurance when cutting issues later.
## /to-issues: Cut PRD into vertical slices
### What is Vertical Slice?
This is one of the most important concepts in Matt’s entire methodology. SKILL.md said directly:
> Each issue is a thin vertical slice cutting through ALL integration layers end-to-end, NOT a horizontal slice of one layer.
The clearest example is to create a “comment function”.
**Horizontal slicing (wrong way)**:
* Issue 1: Database schema
* Issue 2: API endpoints
* Issue 3: UI components
* Issue 4: Testing
**Longitudinal slicing (Tracer Bullet method)**:
* Issue 1: "Visitors can submit an anonymous comment" (schema + API + UI + test are all included, but the scope is so small that it can only be anonymous)
* Issue 2: "Comments from logged in users are associated with the account"
* Issue 3: "Comments can be replied to"
* Issue 4: "Administrators can delete comments"
The problem with horizontal slicing: each slice cannot be tested individually. After Issue 1 was completed, there was nothing to demonstrate, and it was not until Issue 4 that the entire link could be run through—it was only then that I discovered that the schema design was wrong.
Each completed vertical slice is a **end-to-end available functional subset** - called a **tracer bullet** (tracer bullet) in the words of the Pragmatic Programmer. Shoot one round first to check the crosshair, and then adjust the next round.
### HITL vs AFK
`/to-issues` also labels each slice:
* **HITL** (Human in the Loop) - requires people to participate in decision-making. Such as architectural decisions, design reviews
* **AFK** (Away From Keyboard)——The agent can complete the work independently, and you can come back to see the results.
> "Prefer AFK over HITL where possible."
This is a very radical idea in Matt's workflow: after you finish cutting the issue, you send it directly to an agent that runs when you are not present (such as at night or on weekends). When you come back the next day, the PR is already lying there waiting for review. The HITL part stays during the day and works with the agent.
### Slicing confirmation link
`/to-issues` will not generate an issue as soon as it comes up - it will first show you the tiling scheme as a numbered list:
```
1. Title: 访客提交匿名评论
Type: AFK
Blocked by: None
User stories covered: #1, #2
2. Title: 评论关联到登录账户
Type: AFK
Blocked by: #1
User stories covered: #3
3. Title: 评论审核流程
Type: HITL(需要确认审核 UI 设计)
Blocked by: #1
User stories covered: #4, #5
```
Then ask you:
* Is the granularity correct? Too thick / too thin?
* Are the dependencies correct?
* Which ones should be merged/split?
* Are the HITL/AFK marks correct?
Iterate until you nod before actually sending it to the issue tracker, in order of dependency (blocker first), so that later issues can reference the real issue ID of the first issue.
### Issue Template
```markdown
## Parent
A reference to the parent issue (if any).
## What to build
A concise description. Describe end-to-end behavior, NOT layer-by-layer
implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
## Blocked by
- A reference to the blocking ticket
(or "None - can start immediately")
```
Note the "describe end-to-end behavior" line - consistent with the spirit of vertical slice. Acceptance criteria is an acceptance list, which Claude will convert into tests one by one in the `/tdd` stage.
## How to use: Complete process example
Suppose you want to add comment functionality to your blog. Complete process:
```
你: 我想给博客加评论功能
↓
/grill-me → Claude 问 30 个问题(要不要登录?匿名?嵌套?审核?……)
↓
你回答完毕,达成共识
↓
/to-prd → Claude 生成结构化 PRD,提交到 GitHub Issues #42
↓
/to-issues → Claude 提议切成 4 个 vertical slice
让你确认粒度和依赖
你点头
按依赖顺序发布到 GitHub Issues #43~#46
↓
你回家睡觉
↓
夜里 AFK agent 抓 #43(无依赖),跑 /tdd 完成 → 提 PR
你早上 review、merge
↓
agent 抓 #44 / #45 ……
```
The entire process does not require you to sit in front of the screen and watch every detail. The key decisions are made in the grill-me stage.
## Installation and prerequisites
```bash
npx skills@latest add mattpocock/skills
```
Check `to-prd`, `to-issues`, `setup-matt-pocock-skills`.
**You must run `/setup-matt-pocock-skills`** first - it will write your issue tracker (GitHub / GitLab / local markdown) and triage label vocabulary to AGENTS.md/CLAUDE.md, otherwise to-prd and to-issues will not know where to send issues.
Supported issue trackers:
* **GitHub Issues** (default, uses `gh` CLI)
* **GitLab Issues** (using `glab` CLI)
* **Local markdown** (create files under `.scratch//`) - suitable for personal projects or projects without remote
* **Others** (Jira, Linear, etc.) - Use a piece of prose to describe the workflow, and the skill will be called according to your description
## FAQ
\*\*Q: There is already a ready-made PRD. Can I skip to-prd and go directly to to-issues? \*\*
A: Yes. `/to-issues` accepts an issue reference as a parameter ("Break down issue #42 into vertical slices"), and it will fetch the issue content and then slice it.
\*\*Q: My project does not use issue tracker, can it be used? \*\*
A: Yes. Select "Local markdown" during setup, and all issues will become local files like `.scratch//001-foo.md`.
\*\*Q: What should I do if the PRD is too long and the AI itself can’t handle it? \*\*
A: This is a sign of slice granularity - the PRD should be sliced into multiple independent PRDs rather than one giant PRD. You should feel it during the grill stage: If you talk about the 50th issue and are still introducing new features, stop first and cut into two PRDs and do them in batches.
\*\*Q: How does AFK agent automatically catch issues? \*\*
A: Matt's repo does not provide this part, and it must be coordinated with your own agent arrangement (for example, GitHub Actions triggers Claude Code to run issues). The easiest thing to do is to cron to check `is:open no:assignee label:agent-ready` every hour.
## The true value of this process
`/to-prd` and `/to-issues` look like "automated project management", but Matt puts them at the heart of his workflow for a deeper reason:
**It forces you to think "**What is a complete little thing**"**. When you are forced to cut functionality into vertical slices, you are doing one of the hardest things in software engineering - finding seams. The thought itself is more valuable than these two skills.
And the size of the vertical slice is the size that the AI can handle at one time. **Let the task size match the upper limit of the AI's capabilities** - This is the fundamental rhythm of collaboration with LLM.
## Reference resources
Next article: [TDD: Use red and green reconstruction to force AI to take small steps](/en/docs/notes/matt-pocock-skills/tdd)——After the issue is cut, how to make AI really realize small steps.
# Zoom Out: Let AI Draw the Map When You're Lost
## Shortest, Yet Useful
`/zoom-out` is likely the shortest in Matt's suite of stable engineering skills. Its core instruction can be summarized as:
> I'm unfamiliar with this piece of code. Please abstract to a higher level and draw me a map of relevant modules and callers using the project's domain language.
It's not for writing code, nor for refactoring. It's for handling a common state: **both you and the AI are deep in a file, but starting to forget why this file exists.**
## The Failure Mode It Solves
AI programming can easily fall into local optima:
1. User points to a file
2. AI reads the file
3. AI guesses intent based on local code
4. After making changes, it's discovered that upstream callers, domain rules, or ADRs don't support the modification.
Humans do this too. When debugging for a long time, we can get fixated on a single function and forget its place in the system.
`/zoom-out`'s role is to break this tunnel vision. It prompts the agent to pause implementation and first answer:
* Which domain concept does this code belong to?
* Who calls it?
* Who does it call?
* What invariants underlie it?
* How does it map to terms in `CONTEXT.md`?
* Is it constrained by any ADRs?
## Difference from `/improve-codebase-architecture`
Both `/zoom-out` and [`/improve-codebase-architecture`](/en/docs/notes/matt-pocock-skills/improve-codebase-architecture) examine the system globally, but their goals are entirely different.
| Skill | Goal | Output |
| -------------------------------- | ------------------------------------------------ | ----------------------------------------------------- |
| `/zoom-out` | Help you understand unfamiliar code | Map, call relationships, domain explanation |
| `/improve-codebase-architecture` | Find opportunities for architectural improvement | Candidate refactors, deletion tests, interface design |
`/zoom-out` is more like "Please explain this to me." It shouldn't rush to propose refactoring solutions, nor should it directly modify code. Its task is to reduce cognitive load.
## Why Emphasize Domain Vocabulary
This skill explicitly requires using the project's domain glossary. The reason is: if explanations only use filenames, the AI might easily output something like:
```text
OrderService calls OrderRepository, and then OrderRepository calls the db client.
```
This sounds like an explanation, but it isn't. A more useful map would look like this:
```text
In the Checkout flow, Order Draft is a temporary order before the user completes payment.
Order Finalization converts a Draft into an immutable Order and triggers Inventory Reservation.
`OrderService.finalize()` is the seam for this conversion, primarily called by Payment Callback and Admin Retry.
```
The second explanation places the code back into the business language. You don't just know "who calls whom," but also "why it exists."
## When to Use It
I recommend proactively invoking `/zoom-out` in these scenarios:
* Before taking over an unfamiliar module
* When fixing a bug but unsure of the relevant call chain
* When reviewing AI-generated code and can't tell if it's modified at the correct level
* When preparing to write a PRD and wanting to confirm module boundaries
* After looking at 3 files without forming a system diagram
* Before running `/improve-codebase-architecture`, but unsure of the candidate areas
It's particularly suitable as a "5 minutes before starting" step. Some bugs aren't due to difficult code, but because the wrong level was looked at initially.
## A Reusable Output Format
Although the original skill is extremely brief, I suggest having the AI output in this format when using it:
```markdown
## This code's place in the system
## Key Domain Terms
## Main Modules
| Module | Responsibility | Callers | Called By |
|---|---|---|---|
## Key Flows
## Known Constraints / ADRs
## Files I Recommend Reading First
```
This format is more stable than a casual explanation and better suited for conversion into context for subsequent `/to-prd` or `/diagnose` tasks.
## Don't Use It as a Planning Mode
The danger of `/zoom-out` is that after explaining the map, the AI might casually suggest "it could be changed like this." If you only want to understand the code, you should explicitly limit it:
```text
Explain the structure only, do not propose implementation solutions, do not modify files.
```
Its value lies in separating decision-making from understanding. Proposing solutions when understanding is unclear often just wraps misunderstandings in prettier packaging.
## My Usage Recommendations
`/zoom-out` is well-suited for combination with other skills:
* `/zoom-out` → `/diagnose`: First understand the system map, then build a feedback loop.
* `/zoom-out` → `/grill-with-docs`: First understand the existing domain language, then interrogate new requirements.
* `/zoom-out` → `/to-prd`: First confirm module boundaries, then write the PRD.
* `/zoom-out` → `/improve-codebase-architecture`: First draw the map, then find shallow/deep issues.
It's not a complete process, just a brake. When the AI starts making more and more changes in local files, having it zoom out first can often save a round of rework later.
## Resources
Next: [Prototype: Answer a Design Question with Disposable Code](/en/docs/notes/matt-pocock-skills/prototype).
# What is Pi Agent
## Introduction
If you only look at its features, Pi Agent can easily be underestimated: it runs in the terminal, can read and modify files, execute commands, save sessions, and switch models. It sounds like another Claude Code or Codex.
But what I find truly interesting about Pi isn't what it does more of, but what it does less of. It retains the most core layers of an AI programming tool: models, context, tools, sessions, and extensions, then tries not to hardcode the user's workflow in advance.
So I prefer to understand Pi this way:
**Pi Agent is not a "more complete" AI programming product, but a thinner, more transparent coding agent harness.**
This judgment is more important than a feature list, because it determines how you should learn Pi: not by memorizing commands first, but by understanding what layers a coding agent is actually composed of.
## Positioning Pi Correctly
An AI programming tool can usually be roughly divided into three layers: model, harness, and engineering environment. But just drawing these three layers is still too abstract. What's truly worth looking at in Pi is how the middle harness layer is further broken down into modules:
There are three key points in this diagram.
First, Pi isn't just a "chat UI." CLI, interactive TUI, print/JSON, RPC, and SDK are just entry points; the `AgentSessionRuntime` and `AgentSession` are what truly handle tasks.
Second, Pi performs resource loading before requesting the model. The `ResourceLoader` organizes `AGENTS.md`, `CLAUDE.md`, skills, extensions, and prompt templates, then hands them over to the `SystemPrompt Builder` to assemble the context the model actually sees.
Third, when the model calls tools, it doesn't directly control the file system. `AgentHarness` and `AgentLoop` are responsible for validating tools, executing them, capturing results, and continuing to the next round. Extensions, Tool Registry, and SessionManager extend capabilities and save state alongside this process.
So Pi stands in the middle, but this "middle" is not an empty phrase. It specifically controls: which context enters the model, which tools can be called, how tool results return to the session, and which capabilities are added by extensions.
This is also why many articles introducing Pi emphasize minimal, transparent, and extensible. They are all saying the same thing: Pi tries to keep the core operating layer of the agent small, making it visible and modifiable to users.
## How a Task Flows
A single request in Pi isn't "ask the model a sentence, the model replies with a sentence." More accurately, it's a sequence with branches:
The most crucial part here is the back-and-forth between steps four and five. The model doesn't directly touch your file system; it only proposes a `tool_call`. Pi captures this call, validates the tool name and parameters, triggers any potential extension hooks, executes the real operation, and then puts the `tool_result` back into the context. The model then decides the next step based on the new context.
This is the difference between a coding agent and a regular chatbot. Chatbots primarily complete tasks within text; coding agents need to interact with an engineering system, so they must have a harness to manage tools, context, and state.
Pi's default tools are few:
| Tool | Meaning |
| ------- | -------------------------- |
| `read` | Read a file |
| `edit` | Modify an existing file |
| `write` | Create or overwrite a file |
| `bash` | Execute a shell command |
There are also read-only tools like `grep`, `find`, `ls` that can be enabled or restricted. This toolset seems restrained, but it already forms a programming loop: read code, modify code, run tests, and continue fixing based on errors.
The question behind this design isn't "will Pi do more," but "should more things be in the core by default." Pi's answer is clear: not necessarily.
## Why It's Not Rushing to Build Many Features
Many AI programming products integrate features like planning mode, todos, sub-agents, MCP, permission pop-ups, background tasks, and browser tools directly into the product. This makes it quick to get started, but it comes with a cost: it's hard to know what context the model actually received, and it's difficult to adapt the product's workflow to your own.
Pi takes the opposite approach. It keeps the core very small and externalizes the workflow:
| What you want to change | Where Pi delegates it |
| ------------------------------- | ------------------------------ |
| Project rules | `AGENTS.md` / `CLAUDE.md` |
| Specialized task methods | Skills |
| Custom tools and UI | Extensions |
| A set of shareable capabilities | Pi Packages |
| Model selection | Provider / Model configuration |
This isn't "lack of features," but a product trade-off: the core only manages the agent loop, while specific workflows are left to users and teams to combine themselves.
For example, Pi doesn't have DeepSearch built-in by default. But this doesn't mean it can't perform deep searches. A more Pi-like approach would be to write an extension: register a `deep_search` tool, integrate Tavily, Exa, Brave Search, or internal company search, and then let the model call it when needed.
This is different from hardcoding a "search button" into the product. The former is you extending the agent's capabilities; the latter is the product deciding the workflow for you.
## Key Point: How to Extend Your Workflow
To truly use Pi, and not just "try it out," the focus must be on extending your workflow.
What's easily confused here is that Pi doesn't just have "plugins" as an extension method. It's more like giving you four layers of entry points: project rules, task methods, real tools, and shareable packages. You need to first determine what kind of thing you want to solidify.
| What you want to solidify | What to use | Suitable scenario |
| ----------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Project habits and constraints | `AGENTS.md` / `CLAUDE.md` | Tell the agent how to modify code, what checks to run, which directories to avoid |
| A set of reusable methods | Skill | Code review, writing articles, releases, documentation generation, image processing – these are "steps and experiences" |
| A real capability | Extension | Register tools, intercept tool calls, add slash commands, add UI, connect to external APIs |
| A set of distributable capabilities | Pi Package | Bundle extensions, skills, prompt templates, themes for personal or team reuse |
My understanding is: **A Skill is a manual, an Extension is an executable plugin, and a Package is a distribution container.**
For example, my current blog workflow can be broken down like this:
| Workflow Requirement | Where to put it |
| ---------------------------------------------------------------------------------------------------------------- | ----------------------- |
| "Write content only in Chinese, images must use `BlogImage`, run `pnpm types:check` after modifying MDX" | `AGENTS.md` |
| "When writing concept articles, organize them by misconception, definition, mechanism, example, and boundary" | `article-writing` Skill |
| "Add a `deep_search` tool to Pi that can query Tavily / Exa / Brave Search" | Extension |
| "Bundle the writing Skill, DeepSearch Extension, and WeChat publishing command for use across multiple projects" | Pi Package |
This is more accurate than simply saying "install plugins." Because many workflows don't require writing code, just a good set of rules or a Skill; but if you want the agent to truly gain a new capability, such as querying external search, querying databases, calling CI, or intercepting dangerous commands, you should write an Extension.
### Extension: The True Plugin Layer
Pi's Extensions are TypeScript modules. They can do several types of things:
| Capability | Example |
| --------------------------- | --------------------------------------------------------------------------------- |
| Register tools | `deep_search`, `query_logs`, `open_issue` |
| Register commands | `/review`, `/publish`, `/checkpoint` |
| Intercept events | Require confirmation before `bash` executes `rm -rf`, `sudo`, or writes to `.env` |
| Modify UI | Display status, selection boxes, confirmation boxes, task panels in the TUI |
| Save state | Record todos, connection pools, last search results, task stages |
| Connect to external systems | CI, GitHub, logging systems, internal company APIs |
Extensions can be global or project-specific:
```text
~/.pi/agent/extensions/ # Global extensions, available to all projects
.pi/extensions/ # Project extensions, only used in the current project
```
To test a temporary extension, you can use:
```bash
pi -e ./my-extension.ts
```
After placing it in the auto-discovery directory, you can use it in Pi with:
```text
/reload
```
to reload extensions, skills, prompts, and context files.
This is what I find most valuable about Pi: you're not just "having the model help you write code," but designing a controllable work environment for the model. Extensions determine what capabilities the model can call, hooks determine which behaviors should be intercepted, and commands determine how your own workflows are triggered.
### Skill: Don't Write Everything as a Plugin
If a capability is primarily about "how to do something" rather than "calling a real API or executing a program," then it's better suited as a Skill.
A Skill's structure is typically:
```text
my-skill/
SKILL.md
scripts/
templates/
references/
```
Pi does not load the entire skill into context at startup. It first loads the skill's name and description; when a task matches, it then lets the model read the complete `SKILL.md`. This is called progressive disclosure. The benefit is: you can save complex methodologies without polluting the context every time.
For example, "write a good article," "publish to official account," "perform browser QA" are more like Skills. Their value lies primarily in steps, judgment criteria, and reference materials, not necessarily in registering an LLM-callable tool.
### Package: Packaging Your Workflow
Once you have a stable set of capabilities, you can consider turning them into a Pi Package.
A Package can contain:
| Content | Purpose |
| ---------------- | ------------------------------------------ |
| extensions | Executable plugins, tools, commands, hooks |
| skills | Work methods and task manuals |
| prompt templates | Common prompt templates |
| themes | TUI themes |
Installation methods are roughly:
```bash
pi install npm:@scope/my-pi-package
pi install git:github.com/user/repo@v1
pi install ./relative/path/to/package
pi list
pi remove npm:@scope/my-pi-package
pi update --extensions
```
Default installation writes to personal settings. If you want to share it with team projects, you can use project-level settings, so the package is recorded in `.pi/settings.json`. This way, when others enter the project and start Pi, missing packages can be automatically installed.
However, extreme caution is needed here: Packages, Extensions, and Skills can all affect the agent's behavior. A third-party package is not a low-privilege browser plugin; it can run code and instruct the model to execute commands. You should review the source code before installing.
So I would learn Pi's extensions in this order:
1. First, use `AGENTS.md` to clearly define project rules.
2. Then, turn repetitive methods into Skills.
3. When real tool capabilities are needed, write Extensions.
4. Finally, when reusing across multiple projects, package them.
This learning approach is more stable. You don't jump straight into writing plugins, but first break down your workflow into four categories: "rules, methods, tools, distribution," and then decide which layer of Pi each category belongs to.
## What Can Be Seen from the Source Code
When I look at Pi's source code, what's most helpful isn't tracing every function, but seeing which design layer each file represents:
| Source Location | Description |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| `packages/agent/src/agent-loop.ts` | Core loop: connects user messages, model responses, tool calls, and tool results |
| `packages/agent/src/harness/agent-harness.ts` | Harness state: manages sessions, system prompts, tools, hooks, and message queues |
| `packages/coding-agent/src/core/tools/index.ts` | Built-in toolset: default coding tools are `read`, `bash`, `edit`, `write` |
| `packages/coding-agent/src/core/resource-loader.ts` | Resource loading: reads project instructions, extensions, skills, prompt templates, and themes |
| `packages/coding-agent/src/core/system-prompt.ts` | System prompt construction: puts tool descriptions, project context, skills, and current directory into the prompt |
| `packages/coding-agent/src/core/extensions/types.ts` | Extension system: allows extensions to register tools, commands, shortcuts, UI, and lifecycle events |
These parts together essentially form Pi's heart: it first assembles context and tools, then hands the request to the model; if the model needs to call a tool, Pi executes it; once the result returns, the loop continues.
So Pi's "minimalism" isn't an empty claim. The source code structure itself expresses this idea: separating the agent loop, harness, coding tools, resource loading, and extension system, with each layer being relatively clear.
## Pi's Boundaries
Pi offers a high degree of freedom, but freedom does not equate to safety.
Pi packages and extensions can run code; skills can also instruct the model to execute scripts; `bash` can access your real system. Official documentation and security analyses have warned that third-party packages, extensions, and skills need to be reviewed by the user.
I understand Pi's boundaries in three points:
1. **It's not a sandbox**: Don't treat it as a naturally isolated environment. Dangerous projects are best placed in containers, temporary directories, or clean worktrees.
2. **It doesn't judge permissions for you**: Pi's core philosophy isn't to manage risk with a bunch of pop-ups, but to let you control tools, context, and extensions.
3. **It's suitable for those who understand engineering boundaries**: You need to know when to let the agent run commands, when to only provide read-only tools, and when to create a git checkpoint first.
This is also a difference between Pi and some more productized agents. Productized tools will wrap more security and interaction details for you; Pi gives you more direct control, while also returning more responsibility to you.
## How to Learn Pi
When learning Pi, I don't recommend starting with "what commands are there." Commands can be looked up quickly; what's truly worth learning are these questions:
| Question | Why it's important |
| ------------------------------------------------- | -------------------------------------------------------------------------------- |
| How Pi assembles context | Determines what the model actually knows |
| Why Pi's tool surface is so small | Determines if the agent's behavior is observable |
| How Extension registers tools | Determines if you can integrate your own workflow |
| What's the difference between Skill and Extension | Determines when to write descriptions, when to write code |
| How Session saves and forks | Determines if an engineering exploration can be resumed, reviewed, and continued |
If you've already used Claude Code or Codex, you can treat Pi as an opportunity to "look under the hood": why do some tools feel like black-box products, while others feel like modifiable runtimes, even though they all let the model write code?
This question is more worth asking than "Can Pi replace a certain tool?"
## Concluding Thoughts
Pi Agent's most valuable aspect is that it exposes the intermediate layer of AI programming tools.
It reminds us that a coding agent's capabilities don't just come from the model, but also from the harness's design. The model is responsible for thinking, and the harness is responsible for enabling the model to act in a real engineering environment. How context enters, how tools are used, how results return, and how extensions are inserted—these details collectively determine whether the agent is reliable, transparent, and controllable.
Therefore, I wouldn't simply view Pi as a "Claude Code alternative." It's more like an agent runtime suitable for developers to study and modify. You can use it directly to write code, or use it to learn how to design your own agent workflows.
In the next practical article, I will continue with this idea: instead of writing a regular tutorial, I will use Pi Extension to create a `deep_search` tool and see how to integrate external search capabilities into the agent loop.
## Further Reading
# Pi Agent Practice Guide
## Quick Recap
In the conceptual article, I defined Pi Agent as a minimalist **Agent Harness**: it connects models, terminals, file systems, shells, sessions, and the extension system, but doesn't pre-configure a heavy workflow for you.
So, in this practice guide, I don't want to create another typical "let Pi modify files" example. While that example demonstrates the basic loop, it doesn't fully showcase Pi's extensibility.
A more suitable practical example for Pi is to add a capability it doesn't have by default, but many people genuinely need: **DeepSearch**.
Here, DeepSearch is not just simple online searching, but a research-oriented workflow:
| Stage | What to do |
| --------------------- | ------------------------------------------------------------------------------------- |
| Problem Decomposition | Break down a vague question into several searchable sub-questions |
| Multi-round Retrieval | Search official documentation, code repositories, blogs, discussion forums, or papers |
| Source Filtering | Deduplicate, exclude low-quality results, prioritize primary sources |
| Evidence Organization | Extract key facts, links, dates, versions, and uncertainties |
| Synthesized Answer | Provide conclusions, along with supporting evidence and limitations |
My judgment is: **DeepSearch should not be built into Pi's core, nor should it rely solely on prompt engineering. It's better suited as a Pi Extension.**
The reason is simple: DeepSearch involves network requests, third-party search APIs, source filtering, result truncation, citation formatting, and security boundaries. These are all workflow capabilities, not part of the coding agent's minimal core.
## Design Goals
This example aims to implement a minimal viable version, not a perfect research system.
The goals are as follows:
```text
Add a deep_search tool to Pi.
It accepts:
- query: The user's research question
- depth: Search depth
- maxResults: Maximum number of candidate sources to return
It outputs:
- Structured search results
- Title, URL, snippet, and relevance score for each result
- Evidence prompts for the model
After Pi receives this evidence, the current model will generate the final conclusion.
```
I will deliberately separate "retrieval" and "synthesis":
| Part | Responsible Party | Reason |
| ----------------------------------- | -------------------- | --------------------------------------------------- |
| Search API calls | DeepSearch extension | This is a deterministic external capability |
| Result deduplication and truncation | DeepSearch extension | Avoid context being filled with noise |
| Judging which evidence is important | Pi's current model | Requires reasoning and context understanding |
| Final answer writing | Pi's current model | Needs to combine user questions and project context |
This approach is more robust. The extension doesn't need to call another model itself, nor does it need to become a nested agent. It only provides high-quality evidence, allowing Pi's original model to continue reasoning.
## Preparation
Pi extensions can be placed in a global directory or a project directory. Here, I recommend placing them in the project directory first:
```text
.pi/extensions/deepsearch/
package.json
index.ts
```
The advantage of a project-local extension is clear boundaries. This DeepSearch capability is only enabled within the current project and will not affect all Pi sessions.
For the search service, you can choose Tavily, Exa, Brave Search, SerpAPI, or even your own search backend. For the first version, don't worry about the service provider; first abstract it into a `searchWeb()` function.
For example, save the API Key as an environment variable:
```bash
export TAVILY_API_KEY=tvly-...
```
If you don't want to connect to a third-party search API, you can also use local mock data to get the extension running first. Once tool registration, parameter passing, and result formatting are stable, then connect to a real search service.
## Step 1: Create the Extension Directory
First, create the directory:
```bash
mkdir -p .pi/extensions/deepsearch
```
If the extension needs dependencies, you can include a `package.json`:
```json
{
"name": "pi-deepsearch-extension",
"private": true,
"dependencies": {
"typebox": "*",
"@earendil-works/pi-ai": "*",
"@earendil-works/pi-coding-agent": "*"
},
"pi": {
"extensions": ["./index.ts"]
}
}
```
Then install dependencies:
```bash
cd .pi/extensions/deepsearch
npm install
```
Pi's extensions are TypeScript modules; you don't need to compile them manually first. This experience is great for quickly experimenting with tools.
## Step 2: Register the `deep_search` Tool
The core file is `.pi/extensions/deepsearch/index.ts`.
The first version can be written like this:
```typescript
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { StringEnum } from "@earendil-works/pi-ai";
import { Type } from "typebox";
type SearchResult = {
title: string;
url: string;
snippet: string;
score?: number;
};
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "deep_search",
label: "DeepSearch",
description: "Search the web for source-backed evidence about a question.",
promptSnippet: "Research a question with web search and return source-backed evidence.",
promptGuidelines: [
"Use deep_search when the user asks for current facts, external sources, comparison, investigation, or source-backed research.",
"After deep_search returns results, synthesize an answer with citations and clearly separate facts, inference, and uncertainty.",
"Do not treat deep_search results as final truth; inspect source quality and mention gaps."
],
parameters: Type.Object({
query: Type.String({
description: "The research question or search query."
}),
depth: Type.Optional(StringEnum(["quick", "normal", "deep"] as const)),
maxResults: Type.Optional(Type.Number({
minimum: 3,
maximum: 10,
default: 6
}))
}),
async execute(_toolCallId, params, signal) {
const depth = params.depth ?? "normal";
const maxResults = params.maxResults ?? 6;
const results = await searchWeb(params.query, depth, maxResults, signal);
return {
content: [
{
type: "text",
text: formatResultsForModel(params.query, results)
}
],
details: {
query: params.query,
depth,
results
}
};
}
});
}
async function searchWeb(
query: string,
depth: "quick" | "normal" | "deep",
maxResults: number,
signal: AbortSignal
): Promise {
const apiKey = process.env.TAVILY_API_KEY;
if (!apiKey) {
throw new Error("Missing TAVILY_API_KEY. Set it before starting pi.");
}
const response = await fetch("https://api.tavily.com/search", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
api_key: apiKey,
query,
search_depth: depth === "quick" ? "basic" : "advanced",
max_results: maxResults,
include_answer: false,
include_raw_content: depth === "deep"
}),
signal
});
if (!response.ok) {
throw new Error(`Search failed: ${response.status} ${response.statusText}`);
}
const data = await response.json() as {
results?: Array<{
title?: string;
url?: string;
content?: string;
score?: number;
}>;
};
return dedupeByUrl((data.results ?? []).map((item) => ({
title: item.title ?? "Untitled",
url: item.url ?? "",
snippet: item.content ?? "",
score: item.score
}))).filter((item) => item.url);
}
function dedupeByUrl(results: SearchResult[]): SearchResult[] {
const seen = new Set();
const deduped: SearchResult[] = [];
for (const result of results) {
const key = normalizeUrl(result.url);
if (seen.has(key)) continue;
seen.add(key);
deduped.push(result);
}
return deduped;
}
function normalizeUrl(url: string): string {
try {
const parsed = new URL(url);
parsed.hash = "";
parsed.searchParams.delete("utm_source");
parsed.searchParams.delete("utm_medium");
parsed.searchParams.delete("utm_campaign");
return parsed.toString();
} catch {
return url;
}
}
function formatResultsForModel(query: string, results: SearchResult[]): string {
if (results.length === 0) {
return `DeepSearch found no results for: ${query}`;
}
const lines = results.map((result, index) => {
return [
`## Source ${index + 1}`,
`Title: ${result.title}`,
`URL: ${result.url}`,
result.score === undefined ? undefined : `Score: ${result.score}`,
`Snippet: ${result.snippet}`
].filter(Boolean).join("\n");
});
return [
`DeepSearch query: ${query}`,
"",
"Use these sources as evidence. Cite URLs when making factual claims.",
"Separate confirmed facts from inference and uncertainty.",
"",
...lines
].join("\n\n");
}
```
This code only does the most critical things:
| Code Location | Purpose |
| ------------------------- | --------------------------------------------------------------------------------- |
| `pi.registerTool()` | Exposes `deep_search` for model calls |
| `parameters` | Tells the model what parameters the tool needs |
| `promptGuidelines` | Tells the model when to use it and how to handle results |
| `searchWeb()` | Calls the real search service |
| `dedupeByUrl()` | Removes duplicate URLs |
| `formatResultsForModel()` | Organizes search results into evidence blocks that are easy for the model to cite |
The first version should not be too complex. The real challenge of DeepSearch is controlling source quality, context length, citation format, and uncertainty.
## Step 3: Add a `/deepsearch` Command
Tools are for models to call, but users also need a direct entry point.
You can register another command to rewrite user input into a more explicit research task:
```typescript
export default function (pi: ExtensionAPI) {
pi.registerCommand("deepsearch", {
description: "Run a source-backed DeepSearch task",
handler: async (args, ctx) => {
const query = String(args ?? "").trim();
if (!query) {
ctx.ui.notify("Usage: /deepsearch ", "warning");
return;
}
pi.sendUserMessage(
[
"Please perform a DeepSearch on the following question.",
"",
`Question: ${query}`,
"",
"Requirements:",
"1. First, determine if deep_search needs to be called.",
"2. If the question is complex, break it down into 2-4 sub-questions and retrieve information for each.",
"3. The final answer must include source links.",
"4. Distinguish between facts, inferences, and uncertain parts.",
"5. Do not just list search results; provide a comprehensive judgment."
].join("\n"),
{ deliverAs: "followUp" }
);
}
});
pi.registerTool({
// deep_search tool definition...
});
}
```
This way, users can directly input:
```text
/deepsearch What capabilities does the latest version of Pi Coding Agent's extension system support?
```
`/deepsearch` doesn't search directly, but sends Pi a more complete task description. The model will call `deep_search` based on the description, and then synthesize the results.
I prefer this design because it preserves the agent's judgment space. The search tool is just an entry point for evidence, not the final answer generator.
## Step 4: Launch and Verify
Once the project-local extension is in place, you can launch Pi directly from the project root:
```bash
TAVILY_API_KEY=tvly-... pi
```
If you are just testing temporarily, you can explicitly specify the extension:
```bash
TAVILY_API_KEY=tvly-... pi -e ./.pi/extensions/deepsearch/index.ts
```
After entering Pi, ask a question that requires external facts:
```text
/deepsearch What capabilities does the latest version of Pi Coding Agent's extension system support?
```
An acceptable output should not just be a few search results, but should include:
| Checkpoint | Acceptable Performance |
| --------------- | ------------------------------------------------------------------------------------ |
| Tool invocation | `deep_search` is called |
| Clear sources | Each key fact is followed by a URL |
| Deduplication | No repeated citations of the same page |
| Judgment | Not just listing information, but also summarizing applicable scenarios |
| Uncertainty | Maintains boundaries for version changes, third-party APIs, and community extensions |
If the result is just a "list of search results," it means the `promptGuidelines` are not strong enough. You can make the guidelines more explicit:
```typescript
promptGuidelines: [
"Use deep_search to gather evidence, not to produce the final answer.",
"After deep_search, write a concise research brief with citations.",
"Prefer official documentation, source code, release notes, and primary sources.",
"Mention when sources disagree or when the evidence is incomplete."
]
```
## Step 5: Make DeepSearch More Like a Research Tool
After getting the first version running, you can continue to add three types of capabilities.
### Sub-problem Decomposition
The most common failure point for DeepSearch is throwing a large question directly at the search API.
For example:
```text
Can Pi Agent replace Claude Code?
```
This is not a good search query. It can at least be broken down into:
| Sub-question | Purpose |
| ------------------------------------------------------------------------------------ | ------------------------------ |
| What is the core design of Pi Agent | Find its positioning |
| What tools and extensions does Pi Agent support | Find its capability boundaries |
| What are Claude Code's default capabilities | Find comparison objects |
| What are the differences in permissions, security, and extensibility between the two | Form a judgment |
The first version can let the model decompose itself; the second version can make the `/deepsearch` command force the model to first list sub-questions, then call `deep_search` for each.
### Source Quality Layering
DeepSearch output should not just be sorted by the search API's score. When writing technical articles, I prioritize:
| Priority | Source |
| -------- | --------------------------------------------------- |
| P0 | Official documentation, source code, release notes |
| P1 | Author's blog, maintainer's explanation, issue / PR |
| P2 | High-quality tutorials, technical analysis |
| P3 | Community discussions, Reddit, X, forums |
The extension can annotate source types in `formatResultsForModel()`:
```typescript
function classifySource(url: string): "official" | "source" | "community" | "other" {
const host = new URL(url).hostname;
if (host === "pi.dev") return "official";
if (host === "github.com") return "source";
if (host.includes("reddit.com")) return "community";
return "other";
}
```
This way, when the model synthesizes, it won't treat community rumors and official documentation as the same level of evidence.
### Context Truncation
Search results can easily pollute the context. The tool output of DeepSearch should be concise and precise.
My suggestion is:
| Content | Include in tool output |
| ------------------------ | ------------------------------------- |
| Title | Yes |
| URL | Yes |
| 200-500 word snippet | Yes |
| Full page content | No by default |
| Raw HTML | No |
| Raw JSON from search API | In `details`, not in the main content |
If full text reading is indeed required, a second tool can be created:
```text
fetch_source(url)
```
This way, DeepSearch's first step is to find candidate sources, and the second step only grabs the 2-3 most important pages. Don't immediately feed the model the full text of a dozen web pages.
## Common Questions
### Why not just use bash to run search scripts?
You can, but it's not as stable as an extension.
The problem with bash is that the model has to re-decide commands, parameters, output format, and error handling every time. An extension fixes these details, and the model only needs to call `deep_search`.
### Why not write the summary in the extension as well?
Not recommended for the first version.
If the extension itself calls another model to summarize, you will encounter issues with nested model calls, cost accounting, context drift, and citation responsibility. A simpler approach is: the extension only returns evidence, and the model in the current Pi session is responsible for synthesis.
### Is this DeepSearch considered MCP?
No. It's a local tool registered by a Pi extension.
If you already have a mature MCP search server, you can also integrate it via Pi's MCP-related packages or extensions. However, this example chooses to write the extension directly to understand Pi's own extension mechanism.
### What security considerations should be kept in mind?
At least four things:
| Risk | Practice |
| ----------------------- | -------------------------------------------------------------------------------- |
| API Key leakage | Read only from environment variables, do not commit to repository |
| Untrusted web content | Do not treat web content as system instructions, only as evidence to be verified |
| Search result pollution | Prioritize official and source code, reduce weight of community results |
| Context explosion | Limit result quantity and snippet length |
DeepSearch appears to be "search enhancement," but it essentially brings external web pages into the agent's context. As long as external content enters the context, prompt injection should be treated as a real risk.
## Summary
I would define the first practical example of Pi Agent as the **DeepSearch Extension** because it simultaneously demonstrates three key characteristics of Pi:
* Pi's core is very small by default and does not include all workflows.
* Truly useful capabilities can be added via extensions.
* Extensions are not just about adding commands, but about defining the boundaries of the model's interaction with the external world.
Once this example is running, Pi will no longer be just a local code editing agent, but will have a controllable research entry point: when encountering questions requiring external information, it can first retrieve, then filter, and then answer with sources.
This is more reliable than letting the model answer from memory, and more reusable than manually writing search commands every time.
## References
# Ralph Wiggum Deep Dive
## Introduction
Assign tasks to AI before leaving work, and wake up to usable code the next morning — this dream sounds like it would require complex Agent clusters and sophisticated orchestration systems. Yet the hottest AI programming technique of 2025 boils down to this single line:
```bash
while :; do cat PROMPT.md | claude ; done
```
An infinite loop that repeatedly feeds tasks to Claude. This is **Ralph Wiggum**. It's embarrassingly simple, yet someone actually used it to complete a project originally quoted at $50,000 for just $297 in API costs.
Why does such a simple approach actually work? And when Anthropic released an official plugin, why did the inventor Geoffrey Huntley say "This isn't it"?
## What Is Ralph
The name comes from a character in The Simpsons. Ralph Wiggum is the police chief's son — the most "innocent" person in the show. He doesn't quite understand what he's doing, but he never stops. His signature line "I'm helping!" unexpectedly captures the essence of this technique: **Naive and relentless persistence**.
There's an important distinction here: **Ralph is a methodology, not a tool**. Just as "Agile" is a methodology rather than a specific piece of software, Ralph describes a way of working. Different implementations can vary dramatically in effectiveness — we'll discuss this in detail later.
## Why Ralph Is Needed: The Context Rot Problem
To understand why Ralph works, you first need to understand the problem it solves.
### How AI "Gets Dumber"
When using Claude for complex tasks, you may have experienced this: the conversation starts smoothly — Claude understands accurately and executes well. But as the conversation grows longer, it becomes "sluggish" — forgetting important information, repeating the same mistakes, declining in code quality, and even producing inexplicable hallucinations.
This isn't because the AI isn't smart enough. The problem is that **the context window has been polluted**.
Imagine this scenario: you ask Claude to write a feature, and it fails the first time. You say "fix this," it tries but fails again. After ten back-and-forth exchanges, Claude's context is stuffed with: nine failed code attempts, nine sets of error messages, and a mountain of no-longer-relevant discussion. Finding the key information among all this noise becomes increasingly difficult.
### The Dumb Zone
Geoffrey Huntley and the community discovered a phenomenon they call the "Dumb Zone":
| Context Size | Performance |
| ----------------- | ---------------------------------------------------- |
| 0 - 50k tokens | Peak performance |
| 50k - 100k tokens | Good, slight degradation |
| 100k+ tokens | Noticeable degradation, starts ignoring instructions |
| 150k+ tokens | Severe degradation |
There's no precise threshold, but the rule of thumb is: **start worrying when context reaches about half capacity**. For Claude's 200k token window, beyond 100k tokens you may be working with an AI that's "gotten dumber."
### Accumulated Context Is a Liability
Here's a counterintuitive insight: accumulated context isn't an asset — it's a liability.
We're accustomed to thinking that better memory is always better, and more retained information is always better. But in the world of large language models, this intuition is wrong. The longer the conversation, the more "negative information" clutters the context: failed code, irrelevant discussions, corrected misunderstandings. These don't just take up space — they also scatter the AI's "attention."
## How Ralph Works
Once you understand Context Rot, Ralph's solution becomes clear: **if accumulated context is the problem, don't accumulate it**.
Ralph is built on three pillars:
### 1. Fresh Session
At each loop iteration, a **brand new Claude instance** is launched with a completely clean context window. This isn't just "clearing conversation history" — accumulated state might still persist that way. It means completely terminating the current process and starting a new one.
This means Claude is at peak performance at the start of each iteration. No previous errors to haunt it, no stale discussions to distract it.
**This is why the loop must run outside of Claude Code** — the bash loop needs to control the lifecycle of the Claude process.
### 2. Files as the Source of Truth
If each iteration starts with fresh context, how does the AI know what was done before? The answer: through the file system, not conversation history.
Key files:
* **PRD/spec file** — Defines goals, feature lists, success criteria
* **IMPLEMENTATION\_PLAN.md** — Task breakdown and progress tracking
* **progress.txt** — Free-form log; each iteration appends what it learned
* **Git history** — Proof of code changes
At the start of each iteration, Claude reads these files to understand goals and progress. What it sees is a carefully organized state snapshot, not chaotic conversation history.
### 3. Feedback Loop
Clean context and persistent state alone aren't enough. If the AI writes buggy code and commits it, errors will accumulate.
The feedback loop serves as an automated quality gate:
* **TypeScript type checking** — Immediate feedback on type correctness
* **Unit tests** — Verify that features meet expectations
* **CI/CD** — Ensure code builds and integrates properly
If tests fail, the code isn't committed, and Claude sees the failure messages. The next iteration's fresh Claude instance will attempt to fix the issue.
> For more on building a comprehensive quality assurance system, I shared my five-layer defense approach in [My Claude Code Quality Control Workflow](/en/blog/claude-code-quality-control): Hooks automation, testing strategy, AI Review, Pre-commit, and GitHub integration.
## Human on the Loop
Geoffrey Huntley repeatedly emphasizes a conceptual distinction:
| Human **in** the Loop | Human **on** the Loop |
| -------------------------------------------- | -------------------------------------------------- |
| Babysitting | Supervisory management |
| AI waits for your confirmation at every step | You set goals and boundaries, AI runs autonomously |
| You are the bottleneck in the workflow | You check progress occasionally |
In practice, there are two modes:
* **AFK mode**: Start it before leaving work, go home and sleep, check results in the morning
* **Human-in-the-loop mode**: Pause to review after each iteration, suitable for complex or uncertain tasks
## What Tasks Are Suitable for Ralph
Ralph isn't a silver bullet. Its core strength is "iterate until success," which makes it suited for specific types of tasks.
### Suitable Tasks
| Scenario | Reason |
| ------------------------------------- | ----------------------------------------------------------------------- |
| Tasks with clear success criteria | Completion can be verified automatically (tests pass, type checks pass) |
| Tasks requiring iterative improvement | Ralph's core strength is persistent retrying |
| Greenfield projects | No risk of breaking existing code |
| Projects with automated tests | Tests serve as backpressure to ensure quality |
### Unsuitable Tasks
| Scenario | Reason |
| ----------------------------------------- | ------------------------------------------------------- |
| Design decisions requiring human judgment | "Does it look good?" can't be verified automatically |
| One-off operations | Tasks that don't need iteration are wasteful with Ralph |
| Production debugging | Too risky for unattended operation |
| Tasks with unclear success criteria | No way to determine when to stop |
### Three Usage Patterns
**Full Implementation Mode**
This is the most common use of Ralph: building a complete feature or project from scratch. You prepare a spec file and implementation plan, and let Ralph execute all tasks automatically.
Typical scenarios:
* Building a new REST API
* Developing a CLI tool
* Implementing a new feature module
Real-world examples: One developer used this pattern to complete a project originally quoted at $50,000, with total API costs of just $297. The entire process — MVP development, test writing, and code review — was fully automated. Another case involved upgrading a legacy codebase from React v16 to v19; Ralph ran for 14 hours with zero human intervention.
**Exploration Mode**
Not all tasks require code output. Sometimes what you need is understanding — understanding a newly inherited codebase, a complex system's architecture, or how a particular module works.
Typical scenarios:
* Taking over an unfamiliar project and needing to quickly build a mental model
* Generating documentation for an existing codebase
* Analyzing system architecture to identify potential issues
In this mode, your prompt isn't "implement feature X" but rather "read this codebase and generate architecture documentation" or "find all API endpoints and describe their purpose." Claude dives deeper with each iteration, gradually building a more complete understanding.
**Brute-Force Testing Mode**
Some bugs — you know the symptoms, you know the expected correct behavior, but you just can't find the root cause. This is where you let Ralph "brute-force" it.
Typical scenarios:
* An intermittent bug that's hard to reproduce
* A test that fails occasionally for unknown reasons
* A performance issue where the bottleneck is unclear
Set the goal: "Fix this bug and make this test pass consistently." Ralph will keep trying different fix approaches until it finds one that works. This method is especially well-suited for problems where "I don't know how to fix it, but I know when it's fixed."
## Choosing an Implementation
Now that you understand the Ralph methodology, a practical question arises: how do you implement this loop?
The community has developed two implementations with different levels of engineering:
**Minimalist approach** — [snarktank/ralph](/en/docs/notes/ralph-wiggum/snarktank): A few hundred lines of bash script, fresh session each time, focused on the loop itself. Lightweight, easy to get started, ideal for quick adoption.
**Engineered approach** — [frankbria/ralph-claude-code](/en/docs/notes/ralph-wiggum/frankbria): Complete toolchain (monitoring dashboard, circuit breaker, rate limiting, session expiry management). Defaults to session reuse via `--continue`, with an option to switch to fresh sessions via `--no-continue`.
| Dimension | Minimalist (snarktank) | Engineered (frankbria) |
| ----------------------- | ---------------------- | ----------------------------------------- |
| Session mode | Fresh each time | Reuse by default, switchable to fresh |
| Monitoring | Manual inspection | Built-in tmux dashboard |
| Safety mechanisms | max\_iterations | Circuit breaker + rate limiting + timeout |
| Installation complexity | Skill copy | install.sh + wizard |
Both implementations have their strengths and weaknesses — the choice depends on your need for engineering tooling. See each implementation's practice article for detailed usage and comparison.
## Final Thoughts
Ralph teaches us an important lesson: sometimes the simplest approach is the most effective. While everyone else was chasing more complex architectures, a bash loop changed the game.
Of course, Ralph is just one piece of the puzzle. It needs good prompts, the right project, and proper feedback mechanisms to reach its full potential. Now that you understand the principles, you can choose the right implementation for your needs:
* Need long-term AFK with many iterations? → [snarktank/ralph Practice Guide](/en/docs/notes/ralph-wiggum/snarktank)
* Need engineered monitoring and safety mechanisms? → [frankbria/ralph-claude-code Practice Guide](/en/docs/notes/ralph-wiggum/frankbria)
***
**Related Reading**:
* [snarktank/ralph Practice Guide](/en/docs/notes/ralph-wiggum/snarktank) — Minimalist external loop, complete operational manual from installation to real-world use
* [frankbria/ralph-claude-code Practice Guide](/en/docs/notes/ralph-wiggum/frankbria) — Engineered implementation: monitoring, circuit breaker, and safety mechanisms
* [The Complete Guide to Claude Subagents](/en/docs/notes/claude-subagent) — Another approach to keeping context clean
* [What Are Claude Skills](/en/docs/notes/claude-skills/concept) — Exploring Claude's reusable playbooks
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — A complete context engineering system built on Ralph's foundations
* [Claude System Architecture Overview](/en/docs/notes/claude-architecture) — Understanding the overall architecture of Hooks, Subagents, and other components
# frankbria/ralph-claude-code Practical Guide
## Introduction
The [previous article](/en/docs/notes/ralph-wiggum/concept) introduced Ralph's methodology, and [snarktank/ralph](/en/docs/notes/ralph-wiggum/snarktank) demonstrated a minimalist outer loop implementation. Now let's look at another approach: [frankbria/ralph-claude-code](https://github.com/frankbria/ralph-claude-code).
If snarktank/ralph's philosophy is "do the most with the least code," then frankbria's philosophy is "**engineer everything**" — interactive configuration wizards, real-time monitoring dashboards, circuit breakers, rate limiting, and session expiration management. It doesn't pursue simplicity; it pursues **controllability**.
Neither implementation is inherently better — they suit different use cases. This article walks you through frankbria's complete toolchain.
## Installation and Configuration
### Global Installation
```bash
# Clone the repository
git clone https://github.com/frankbria/ralph-claude-code.git
cd ralph-claude-code
# Global install
./install.sh
```
After installation, you'll have the following global commands:
| Command | Description |
| --------------- | -------------------------------------------- |
| `ralph` | Start the Ralph loop |
| `ralph-enable` | Enable Ralph in an existing project |
| `ralph-setup` | Create a new project and configure Ralph |
| `ralph-import` | Import an existing PRD/requirements document |
| `ralph-monitor` | Launch the real-time monitoring dashboard |
### Project Initialization
For existing projects, use the interactive wizard:
```bash
cd your-project
ralph-enable
```
The wizard automatically detects the project type (Node.js, Python, Go, etc.) and framework (Next.js, FastAPI, etc.), then generates the corresponding configuration files.
For brand-new projects:
```bash
ralph-setup my-new-project
```
This creates the project directory, initializes Git, and generates the `.ralph/` configuration directory.
### Importing Existing Requirements
If you already have a PRD document or requirements specification:
```bash
ralph-import path/to/your-prd.md
```
Ralph will parse the document, extract the task list, and generate a structured `fix_plan.md`.
## .ralph/ Directory Structure
frankbria's memory and configuration are centralized in the `.ralph/` directory:
```
.ralph/
├── PROMPT.md # Project goals and context
├── fix_plan.md # Task checklist (similar role to prd.json)
├── AGENT.md # Build/test commands (auto-maintained)
├── specs/ # Detailed requirement documents
│ ├── feature-a.md
│ └── feature-b.md
└── sessions/ # Session persistence data
├── current.json
└── history/
```
**Comparison with snarktank/ralph**:
| frankbria | snarktank | Purpose |
| ------------- | ------------------------------------------ | -------------------------------------- |
| `PROMPT.md` | `prd.json`'s `projectName` + `description` | Define project goals |
| `fix_plan.md` | `prd.json`'s `userStories` | Task list and progress |
| `AGENT.md` | `CLAUDE.md` / `AGENTS.md` | Build commands and project conventions |
| `specs/` | `prd.json`'s `notes` field | Detailed requirements |
| `sessions/` | None (new process each time) | Session state tracking |
Note that `AGENT.md` is **auto-maintained** — Ralph automatically updates this file based on project conventions discovered during execution, similar to snarktank/ralph's `progress.txt`, but more structured.
## Core Commands
### Basic Execution
```bash
# Start the Ralph loop
ralph
# With real-time monitoring
ralph --monitor
# Run in tmux (recommended for long-running tasks)
ralph --live
```
### Monitoring Dashboard
```bash
# Launch monitoring independently
ralph-monitor
```
`ralph-monitor` opens a tmux dashboard that displays in real time:
* The currently executing task
* Completed/remaining task counts
* API call counts and cost estimates
* Circuit breaker status
* Recent error logs
### Common Parameters
| Parameter | Description | Default |
| ----------------- | -------------------------------- | ------- |
| `--resume` | Continue from where you left off | - |
| `--calls ` | Maximum API call count | 100 |
| `--timeout ` | Timeout in minutes | 300 |
| `--monitor` | Enable real-time monitoring | false |
| `--live` | Run in tmux | false |
```bash
# Limit to 50 API calls with a 2-hour timeout
ralph --calls 50 --timeout 120
# Resume from where you left off
ralph --resume
```
## Safety Mechanisms
frankbria's biggest differentiator is its multi-layered safety mechanisms.
### Circuit Breaker
The circuit breaker automatically stops the loop when it detects "no progress," preventing wasteful API consumption:
**Consecutive no-progress detection**: If N consecutive iterations complete without any new task being finished, the circuit breaker triggers.
**Repeated error detection**: If the same error message appears consecutively, it indicates the AI is stuck in a loop, and the circuit breaker triggers.
### Rate Limiting
The default limit is 100 calls/hour, preventing unexpected API billing spikes. You can adjust this via parameters:
```bash
ralph --calls 200 # Increase to 200 calls
```
### 5-Hour API Quota Three-Layer Detection
The Anthropic API has a 5-hour sliding window usage quota. frankbria includes three layers of detection:
1. **Pre-check**: Estimates remaining quota before each API call
2. **Response detection**: Parses rate limit headers from API responses
3. **Fallback strategy**: Automatically reduces call frequency when approaching the limit
### Session Expiration Management
The default session validity is 24 hours. After that, session data is automatically cleaned up to prevent stale context from affecting subsequent executions.
## Intelligent Exit Detection
frankbria doesn't simply exit when all tasks are complete. It uses a **dual-condition exit gate**:
```
Exit condition = completion_indicators >= 2 AND EXIT_SIGNAL: true
```
**completion\_indicators** is the number of completion signals detected from AI output, including:
* "All tasks completed"
* "No more pending items"
* All tests passing
* All entries in fix\_plan.md marked as done
**EXIT\_SIGNAL** is an explicit exit intent declared by the AI in its output.
Why two conditions? To prevent **premature exits**. A single signal could be a false positive — for example, the AI might say "task complete" when it has only finished the current story. The dual condition ensures the loop only truly exits when multiple independent signals confirm completion.
## Comparison with snarktank/ralph
| Dimension | snarktank/ralph | frankbria/ralph-claude-code |
| --------------------- | ------------------------------------------ | ------------------------------------------------------ |
| **Implementation** | External bash loop (new session each time) | External bash loop (`--continue` reuses session) |
| **Session mode** | Fresh each time | Reuse by default (switch to fresh via `--no-continue`) |
| **Context** | Fresh each time | Accumulated across iterations via `--continue` |
| **Installation** | Skill copy | install.sh + interactive wizard |
| **Task format** | prd.json | PROMPT.md + fix\_plan.md |
| **Monitoring** | Manual `cat`/`jq` | Built-in tmux dashboard |
| **Safety mechanisms** | max\_iterations | Circuit breaker + rate limiting + timeout |
| **Task sources** | PRD only | beads / GitHub Issues / PRD |
| **Best suited for** | Long AFK sessions, many iterations | Short-to-medium iterations, monitoring needed |
### Core Difference: Level of Engineering
Both are external bash loops launching new Claude processes. The core difference isn't in session management (frankbria can switch to fresh session mode via `--no-continue`), but in the **level of engineering**:
* **snarktank**: Minimalist script, a few hundred lines of bash, focused on the loop itself
* **frankbria**: Full-engineered toolchain — monitoring dashboard, circuit breaker, rate limiting, session expiration management
frankbria enables `--continue` by default to reuse sessions, which suits short tasks. For long tasks, you can switch to `--no-continue` mode to get the same Context Rot protection as snarktank, while retaining frankbria's engineering advantages.
### How to Disable Session Reuse
frankbria offers three ways to disable `--continue`:
```bash
# Method 1: Command-line argument
ralph --no-continue
# Method 2: Environment variable
export CLAUDE_USE_CONTINUE=false
# Method 3: .ralphrc configuration
SESSION_CONTINUITY=false
```
Once disabled, frankbria behaves like snarktank (fresh session each time), but retains all engineering tools (monitoring, circuit breaker, rate limiting, etc.).
## The Real-World Trade-off of Context Rot
The choice of session reuse strategy is fundamentally a trade-off between Context Rot and startup overhead:
**Short tasks (\< 50k tokens)**: Session reuse has the advantage. Context hasn't had time to degrade, and memory from earlier iterations can still benefit subsequent ones. The startup overhead of creating a new session each time is actually wasteful.
**Long tasks (100k+ tokens)**: Fresh sessions are more reliable. Beyond 100k tokens, Context Rot intensifies significantly — accumulated context turns from an asset into a liability. Fresh sessions have startup overhead, but each one starts in optimal condition.
**Practical recommendations**:
| Scenario | Recommendation | Reason |
| ------------------------- | ---------------------------------------- | ------------------------------------------ |
| Fewer than 5 small tasks | frankbria (default mode) | Fast startup, reusable context |
| 10+ tasks, need to go AFK | snarktank or frankbria + `--no-continue` | Avoids Context Rot, more reliable |
| Need real-time monitoring | frankbria | Built-in dashboard |
| Uncertain task volume | frankbria + `--no-continue` | Engineering tools + Context Rot protection |
frankbria users can choose flexibly based on task scale: use the default `--continue` mode for short tasks, and switch to `--no-continue` mode for long tasks. Compared to snarktank, frankbria's advantage is that the complete engineering toolchain is retained regardless of which mode is used.
## Summary
frankbria/ralph-claude-code represents the engineered implementation path of the Ralph methodology. It trades some of snarktank's simplicity for more comprehensive monitoring, safety, and configuration capabilities.
Which implementation you choose depends on your specific needs — there's no "more correct" answer, only a "better fit."
***
**Further reading**:
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Core principles and methodology
* [snarktank/ralph Practical Guide](/en/docs/notes/ralph-wiggum/snarktank) — Minimalist outer loop implementation
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — A complete context engineering system built on Ralph
* [Claude System Architecture Guide](/en/docs/notes/claude-architecture) — Understanding the overall architecture of Hooks, Subagents, and other components
# Ralph Practical Guide
## Introduction
In the [previous article](/en/docs/notes/ralph-wiggum/concept), we explored Ralph's core principles — infinite loop + fresh context each time + files as the source of truth. These three pillars sound simple, but there are quite a few details to work through between understanding the concept and actually getting it running.
In this article, we'll get hands-on. You'll learn how to use [snarktank/ralph](https://github.com/snarktank/ralph) to complete the full workflow from installation to execution. snarktank/ralph is one of the most mature Ralph implementations in the community (10k+ stars), supporting both Claude Code and Amp, with a complete toolchain for PRD generation, JSON conversion, and automated execution.
## Prerequisites
Before getting started, make sure your environment meets these requirements:
| Dependency | Description |
| ------------------ | ------------------------------------------------------------------- |
| **AI Coding Tool** | Claude Code (`npm install -g @anthropic-ai/claude-code`) or Amp CLI |
| **jq** | JSON processing tool (macOS: `brew install jq`) |
| **Git** | Project must be a Git repository |
```bash
# Check dependencies
claude --version # Claude Code CLI
jq --version # JSON processing
git --version # Git
```
## Installation & Configuration
snarktank/ralph offers multiple installation methods. Choose based on your use case.
### Option 1: Install Directly in Claude Code (Recommended)
The simplest approach — paste the GitHub link in a Claude Code conversation and let Claude handle the installation automatically:
```
Install this skill for me: https://github.com/snarktank/ralph
```
Claude Code will automatically clone the repository and copy the skill files to the correct location. Once installed, you can use the `/prd` and `/ralph` commands.
### Option 2: Claude Code Marketplace
Install via Marketplace commands:
```bash
# Add and install the plugin
/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace
```
After installation, you'll have access to `/prd` (generate PRD) and `/ralph` (convert to JSON) skills.
### Option 3: Manual Skill Installation (Claude Code / Amp)
Manually copy the skill files to the corresponding tool's global config directory:
```bash
# Clone the repository first
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# For Claude Code users
cp -r /tmp/ralph/skills/prd ~/.claude/skills/
cp -r /tmp/ralph/skills/ralph ~/.claude/skills/
# For Amp users
cp -r /tmp/ralph/skills/prd ~/.config/amp/skills/
cp -r /tmp/ralph/skills/ralph ~/.config/amp/skills/
```
After installation, you can use the `/prd` and `/ralph` commands.
### Option 4: Project-Level Installation
Copy Ralph scripts directly into your project — best for teams that need to share or customize scripts:
```bash
# Clone the Ralph repository
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# Copy core files to your project
mkdir -p scripts/ralph
cp /tmp/ralph/ralph.sh scripts/ralph/
cp /tmp/ralph/CLAUDE.md scripts/ralph/ # For Claude Code users
# OR
cp /tmp/ralph/prompt.md scripts/ralph/ # For Amp users
# Grant execution permission
chmod +x scripts/ralph/ralph.sh
```
After installation, your project will have this structure:
```
your-project/
├── scripts/ralph/
│ ├── ralph.sh # Core loop script
│ └── CLAUDE.md # Prompt template for Claude Code
├── tasks/ # PRD file directory (created automatically during execution)
│ └── prd.json # Your task definitions
└── ...
```
> **Recommendation**: Option 1 is the easiest — just give Claude Code the GitHub link. For manual control over the installation process, choose Option 2 (Marketplace commands) or Option 3 (manual copy). Use Option 4 when you need team sharing or script customization.
***
## Core File Structure
Ralph's memory relies entirely on the file system. Understanding each file's role is essential to using Ralph effectively.
### ralph.sh — The Loop Engine
This is Ralph's core: a bash script that repeatedly spawns new AI instances.
```bash
# Basic usage
./scripts/ralph/ralph.sh [max_iterations] # Default: Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # Use Claude Code
```
Each iteration, ralph.sh performs these steps:
1. Creates a feature branch (from `branchName` in prd.json)
2. Selects the highest-priority incomplete story (`passes: false`)
3. Spawns a **brand new** AI instance to implement this story
4. Runs quality checks (type checking, tests)
5. Checks pass → git commit; checks fail → left for next iteration
6. Updates prd.json, marking the story as `passes: true`
7. Appends learnings to progress.txt
8. Repeats until all stories are complete or iteration limit is reached
Default iteration limit is 10. Adjust based on project complexity:
```bash
# Simple project
./scripts/ralph/ralph.sh --tool claude 10
# Complex project
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — Task Definitions
This is Ralph's "brain" — all tasks are defined here. It's a flat JSON file:
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists with valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve MDX component props format"
}
]
}
```
**Field reference**:
| Field | Description |
| -------------------- | ----------------------------------------------------------------- |
| `projectName` | Project name, used for logging and branch naming |
| `branchName` | Git branch name — Ralph creates it automatically |
| `id` | Unique story ID, recommend `US-001` format |
| `title` | Short title |
| `description` | Detailed description — the more specific, the better |
| `acceptanceCriteria` | List of acceptance criteria — **this is the most critical field** |
| `priority` | Priority number — lower executes first |
| `passes` | Whether completed — Ralph updates this automatically |
| `dependsOn` | List of dependent story IDs |
| `notes` | Additional hints and context |
### progress.txt — Learnings Log
This is Ralph's "long-term memory." After each iteration, the AI appends what it learned:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
The next iteration's fresh Claude instance reads this file, immediately gaining all previous experience. This is why Ralph gets smoother over time — **knowledge accumulates across iterations while context stays clean**.
### AGENTS.md — Persistent Knowledge Base
Beyond progress.txt, Ralph also updates `AGENTS.md` (or `CLAUDE.md`) files in the project. Both Claude Code and Amp automatically read these files on startup.
Unlike progress.txt, AGENTS.md records **stable, cross-project knowledge**:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## Writing a PRD
The quality of your PRD (Product Requirements Document) directly determines how well Ralph executes. Write it well, and Ralph sails through. Write it poorly, and Ralph will repeatedly fail on the same story.
### Using Skills to Generate a PRD
If you installed snarktank/ralph's skills, you can generate a PRD interactively:
```bash
# In Claude Code or Amp
/prd I want to add i18n support to the blog, translating all Chinese content to English
```
The AI will ask you clarifying questions (which files are involved, tech stack constraints, quality standards, etc.), then generate a structured PRD document.
After generation, use the `/ralph` command to convert the PRD to `prd.json` format:
```bash
/ralph # Convert PRD to prd.json
```
### Writing PRD Manually
You can also write prd.json directly. Here are the key design principles.
**Principle 1: Right-size your stories**
Each story should be small enough to complete in one iteration, yet large enough to deliver independent value.
```json
// ❌ Too large: can't finish in one iteration
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, forgot password, OAuth, permission management..."
}
// ❌ Too small: no independent value
{
"id": "US-001",
"title": "Create email field on User table",
"description": "Add email field to User model"
}
// ✅ Just right: completable in one iteration, delivers independent value
{
"id": "US-001",
"title": "Implement email/password login",
"description": "Create login API and login page with email/password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form submits successfully",
"All tests pass"
]
}
```
**Rule of thumb**: A story involves 1-3 file modifications and has 3-5 acceptance criteria.
**Principle 2: Acceptance criteria must be automatically verifiable**
Ralph needs to determine whether a story is complete, so acceptance criteria must be objectively assessable:
```json
// ❌ Vague criteria
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ Verifiable criteria
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**Principle 3: Use dependsOn to control ordering**
Some stories have dependencies. The `dependsOn` field ensures Ralph executes in the correct order:
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**Principle 4: Provide context in notes**
The notes field gives the AI extra hints. Write things here that you know but the AI might not:
```json
{
"notes": "Project uses fumadocs framework, i18n files follow .en.mdx suffix naming. Reference content/docs/notes/speckit/concept.en.mdx for translation style."
}
```
***
## Executing the Ralph Loop
With the PRD ready, it's time to run the loop.
### Starting Execution
```bash
# Using Claude Code, default 10 iterations
./scripts/ralph/ralph.sh --tool claude
# Specify iteration count
./scripts/ralph/ralph.sh --tool claude 30
# Using Amp (default)
./scripts/ralph/ralph.sh 20
```
### The Execution Process
After starting, you'll see output like this:
```
=== Ralph Loop - Iteration 1 ===
Branch: ralph/i18n-translation
Selected story: US-001 - Translate homepage metadata
Spawning fresh Claude instance...
[Claude Code executing...]
Quality check: pnpm types:check ... PASSED
Committing: feat: [US-001] - Translate homepage metadata
Updating prd.json: US-001 passes: true
Appending to progress.txt
=== Ralph Loop - Iteration 2 ===
Selected story: US-002 - Translate blog post hello-world
Spawning fresh Claude instance...
```
Each iteration is a brand new Claude instance. It knows what to do by reading prd.json, and it knows what was learned previously by reading progress.txt.
### Completion Signal
When all stories are marked `passes: true`, Ralph outputs a completion signal and exits:
```
All stories completed!
COMPLETE
```
### Monitoring & Debugging
While Ralph is running, you can check progress with these commands:
```bash
# View completion status for each story
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# View learnings log
cat progress.txt
# Check recent git commits
git log --oneline -10
# Watch Ralph's output in real-time
tail -f progress.txt
```
### Automatic Archiving
When you start a different feature with a new `branchName`, Ralph automatically archives the previous run's files to `archive/YYYY-MM-DD-feature-name/`, keeping the working directory clean.
***
## Feedback Loops & Quality Gates
Ralph's "self-correction" ability depends entirely on the quality of the feedback loop. Without a feedback loop, Ralph is just a blind looping script — it will keep producing code but has no way to tell if the code is correct.
### Configuring Quality Checks
Define your quality check commands in CLAUDE.md (or prompt.md):
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### Quality Gate Layers
| Layer | Tool | Issues Caught |
| ------------------------ | ------------------- | --------------------------------------- |
| Immediate feedback | TypeScript compiler | Type errors, syntax errors |
| Functional verification | Unit tests | Logic errors, edge cases |
| Integration verification | Build command | Dependency issues, configuration errors |
| Runtime verification | dev-browser skill | UI rendering issues (frontend projects) |
> For frontend stories, Ralph recommends adding this acceptance criterion: "Verify in browser using dev-browser skill" — letting the AI actually open a browser to confirm the page renders correctly.
### When Quality Checks Fail
If a story's quality checks fail repeatedly, Ralph won't retry the same story indefinitely. After reaching the iteration limit, it stops and leaves the current state. You can:
1. Check progress.txt to see what's blocking
2. Fix the issue manually and re-run
3. Adjust story granularity (it might be too large)
4. Add more context in the notes field
***
## Prompt Customization
Ralph's prompt template (CLAUDE.md or prompt.md) is your primary means of controlling AI behavior. After installation, you should customize it for your project.
### Key Customizations
**1. Project-specific quality commands**
```markdown
## Project-Specific Commands
- Typecheck: `pnpm types:check` (not `tsc` or `pnpm typecheck`)
- Test: `pnpm vitest run`
- Build: `pnpm build`
- Lint: `pnpm lint`
```
**2. Code style constraints**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**3. Known gotchas**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**4. Handling stuck situations**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## Case Study: Blog i18n Translation with Ralph
To illustrate how Ralph works in practice, here's a real-world example: translating this entire blog from Chinese to English using a Ralph-style autonomous agent.
### The Setup
The project required translating 22+ content files (blog posts, documentation, navigation metadata) from Chinese to English for a fumadocs-based Next.js blog with i18n support. The task was defined in a `prd.json` file with 16 user stories, each with clear acceptance criteria:
```
scripts/ralph/
├── prd.json # 16 user stories with acceptance criteria
└── progress.txt # Learnings log, updated after each story
```
Each user story followed a consistent pattern:
* **Clear deliverables**: "Create content/blog/xxx.en.mdx"
* **Verifiable criteria**: "Typecheck passes", "Internal links use /en/ prefix"
* **Technical constraints**: "Keep code blocks untranslated", "Set defaultLang='en' on QuoteCard"
### The Execution Pattern
The agent followed the Ralph methodology's core principles:
1. **File as source of truth**: The `prd.json` tracked which stories passed (`passes: true/false`). The `progress.txt` accumulated learnings across iterations — e.g., "Typecheck command is `pnpm types:check`, not `pnpm typecheck`"
2. **Automated quality gate**: After each translation, `pnpm types:check` ran to verify the MDX files compiled correctly. If typecheck failed, the issue was fixed before committing.
3. **Incremental progress**: Each story was committed independently with a descriptive message (`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`), enabling easy rollback if needed.
4. **Parallel execution**: For longer articles, multiple subagents translated files simultaneously — e.g., US-010 (claude-skills concept + practice), US-011 (speckit concept + practice), and US-012 (claude-architecture + claude-subagent) all ran in parallel.
### Key Learnings
| Learning | Details |
| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Accumulated knowledge matters** | Early stories discovered patterns (QuoteCard `defaultLang`, link prefixing rules) that made later stories faster |
| **Typecheck as feedback loop** | Caught issues like missing imports or malformed MDX before they could compound |
| **Parallelization scales** | 6 translation agents running simultaneously completed in roughly the same time as 1 |
| **PRD granularity is critical** | Each story was scoped to 1-2 files — small enough to complete reliably, large enough to be meaningful |
| **Progress log prevents repeat mistakes** | The `progress.txt` "Codebase Patterns" section became a knowledge base that prevented rediscovering the same issues |
### Results
16 user stories completed across a single session: 8 meta.en.json navigation files created, 3 blog posts translated, 12 documentation pages translated, full site build verified. Each translation maintained consistent quality because the acceptance criteria were explicit and the feedback loop (typecheck) caught issues immediately.
This project demonstrates Ralph's **Full Implementation Mode** in action — a well-defined task with clear success criteria, automated verification, and incremental delivery through the file system.
***
## Community Implementations & Alternatives
snarktank/ralph isn't the only choice. Depending on your needs, these implementations each have their strengths:
| Resource | Link | Description |
| ------------------ | ----------------------------------------------------------------------------------- | -------------------------------------------------- |
| snarktank/ralph | [snarktank/ralph](https://github.com/snarktank/ralph) | Used in this article, most feature-complete |
| ralph-orchestrator | [mikeyobrien/ralph-orchestrator](https://github.com/mikeyobrien/ralph-orchestrator) | By Mickey O'Brien, with more customization options |
| ralph-loop-agent | [vercel-labs/ralph-loop-agent](https://github.com/vercel-labs/ralph-loop-agent) | Vercel's implementation based on the AI SDK |
| ralphy | [michaelshimeles/ralphy](https://github.com/michaelshimeles/ralphy) | Michael Shimeles' lightweight implementation |
### Alternative: GSD
GSD is not strictly a "community implementation" of Ralph — it's an **alternative approach**. It applies Ralph's core principles (context management, atomic tasks) but provides a more complete workflow: discuss → plan → execute → verify.
| Resource | Link | Description |
| -------------------- | ----------------------------------------------------------------------------- | -------------------------------------------------- |
| GSD (Get Stuff Done) | [glittercowboy/get-shit-done](https://github.com/glittercowboy/get-shit-done) | A complete framework from idea to PRD to execution |
If you find Ralph too "raw" and need more process support, GSD might be a better fit. See [GSD Deep Dive](/en/docs/notes/gsd/concept) for details.
***
## Recommended Resources
**Official Sources**:
| Resource | Link | Description |
| ----------------------- | ------------------------------------------------------------------------------- | ------------------------------- |
| Geoffrey Huntley's Blog | [ghuntley.com/ralph](https://ghuntley.com/ralph/) | The inventor's original article |
| how-to-ralph-wiggum | [ghuntley/how-to-ralph-wiggum](https://github.com/ghuntley/how-to-ralph-wiggum) | Official usage guide |
**Video Tutorials**:
| Resource | Link | Description |
| ---------------------------- | ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| Ralph Wiggum Deep Discussion | [Why Claude Code's implementation isn't it](https://www.youtube.com/watch?v=O2bBWDoxO4s) | Geoffrey Huntley explains the problems with the official implementation |
| Using Ralph Correctly | [You're Using Ralph Wiggum Loops WRONG](https://www.youtube.com/watch?v=I7azCAgoUHc) | Roman (Mentat)'s usage walkthrough |
| We need to talk about Ralph | [We need to talk about Ralph](https://www.youtube.com/watch?v=Yr9O6KFwbW4) | Theo's analysis of the controversy |
***
## Best Practices & FAQ
### Cost Control
Ralph's automated execution means API costs are continuously incurred. A few control measures:
* **Always set `max_iterations`**: This is the most basic safety net
* **Keep story granularity reasonable**: Stories that are too large consume multiple iterations; stories that are too granular increase startup overhead
* **Test small first**: For new projects, do a trial run with 3-5 iterations to confirm your prompt and quality gates are working before scaling up
### Common Pitfalls
**Pitfall 1: Stories too large**
Symptoms: A story fails repeatedly, iteration count depletes quickly.
Solution: Split into 2-3 smaller stories. "Build complete auth system" becomes "Implement login API" + "Create login page" + "Add JWT middleware."
**Pitfall 2: No feedback loop**
Symptoms: Ralph claims stories are complete, but the actual code has issues.
Solution: Add executable check commands to acceptance criteria. "Code is written" is not an acceptance criterion — "pnpm test all passes" is.
**Pitfall 3: progress.txt not being utilized**
Symptoms: The same error recurs across different iterations.
Solution: Verify your prompt template explicitly instructs "Read progress.txt and follow the learnings within." If the AI isn't appending learnings automatically, add "After each story, append learnings to progress.txt" to your prompt.
**Pitfall 4: Incorrect dependency ordering**
Symptoms: A story depends on code that doesn't exist yet, causing implementation failure.
Solution: Set the `dependsOn` field correctly. Ensure infrastructure stories come first.
### FAQ
**Q: What's the difference between Ralph and the official plugin?**
The core difference: snarktank/ralph spawns a new process each iteration (truly fresh context), while the official plugin loops within the same session (context keeps accumulating). See the [analysis in the previous article](/en/docs/notes/ralph-wiggum/concept#the-problem-with-the-official-plugin) for details.
**Q: Can I manually modify prd.json mid-execution?**
Yes. Ralph re-reads prd.json at the start of each iteration. You can modify story descriptions, add new stories, or manually mark a story as `passes: true` (to skip it) between iterations.
**Q: What if Ralph gets stuck on a story that keeps failing?**
1. Check progress.txt for the failure reason
2. Add more context in notes
3. Split the story (it may be too large)
4. Fix the blocking issue manually and re-run
**Q: Can I do other things while Ralph runs?**
Yes. Ralph is designed for "Human on the Loop" — you don't need to watch it. In AFK mode, start it before leaving work and check results the next day. Just don't modify files that Ralph is actively working on.
**Q: How do I control costs?**
Three approaches: set a reasonable `max_iterations`, keep story granularity appropriate (reducing wasted iterations), and do a small trial run first to confirm the workflow is correct. Generally, a project with 10-20 stories falls within $50-100 in API costs.
***
## Summary
Ralph's workflow can be distilled into five steps:
```
Install → Write PRD → Configure Quality Gates → Run Loop → Review Results
```
The core philosophy never changes: **let files be the source of truth, let every iteration start fresh, and let quality gates do the gatekeeping for you**.
Now, go back to your project, prepare your prd.json, run `./scripts/ralph/ralph.sh --tool claude`, and go grab a coffee.
### Further Reading
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Revisit Ralph's core principles
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — A complete context engineering system built on Ralph's foundations
* [What Are Claude Skills](/en/docs/notes/claude-skills/concept) — Ralph's PRD skill is a Claude Skill
* [Speckit Practical Guide](/en/docs/notes/speckit/practice) — Another structured AI programming workflow
# snarktank/ralph Practical Guide
## Introduction
In the [previous article](/en/docs/notes/ralph-wiggum/concept), we explored Ralph's core principles — infinite loop + fresh context every time + files as the source of truth. The three pillars sound simple, but there are quite a few details between understanding the concept and actually getting it running.
In this article, we'll get hands-on. [snarktank/ralph](https://github.com/snarktank/ralph) is the **outer loop implementation** of the Ralph methodology — each iteration launches a brand-new Claude process, completely solving the Context Rot problem. It's one of the most mature Ralph implementations in the community (10k+ stars), supporting both Claude Code and Amp platforms, and providing a full toolchain for PRD generation, JSON conversion, and automated execution.
> Another implementation path is [frankbria/ralph-claude-code](/en/docs/notes/ralph-wiggum/frankbria), which offers a complete engineering toolchain (monitoring dashboard, circuit breakers, rate limiting), focusing on controllability and safety mechanisms. See that article for a comparison of the two.
## Prerequisites
Before getting started, make sure your environment meets the following requirements:
| Dependency | Description |
| ------------------ | ------------------------------------------------------------------- |
| **AI Coding Tool** | Claude Code (`npm install -g @anthropic-ai/claude-code`) or Amp CLI |
| **jq** | JSON processing tool (macOS: `brew install jq`) |
| **Git** | Project must be a Git repository |
```bash
# Check dependencies
claude --version # Claude Code CLI
jq --version # JSON processing
git --version # Git
```
## Installation & Configuration
The simplest approach — paste the GitHub link directly in a Claude Code conversation:
```
Help me install this skill: https://github.com/snarktank/ralph
```
Claude Code will automatically clone the repository and copy the skill files to the correct location. Once installed, you can use the `/prd` and `/ralph` commands.
> snarktank/ralph also supports Marketplace installation, manual skill file copying, project-level installation, and more. See the [GitHub repository](https://github.com/snarktank/ralph) for details.
***
## Core File Structure
Ralph's memory relies entirely on the file system. Understanding each file's role is essential to using Ralph effectively.
### ralph.sh — The Loop Engine
This is Ralph's core: a bash script that repeatedly launches new AI instances.
```bash
# Basic usage
./scripts/ralph/ralph.sh [max_iterations] # Uses Amp by default
./scripts/ralph/ralph.sh --tool claude [iterations] # Uses Claude Code
```
Each iteration, ralph.sh does the following:
1. Creates a feature branch (based on `branchName` in prd.json)
2. Selects the highest-priority incomplete story (`passes: false`)
3. Launches a **brand-new** AI instance to implement the story
4. Runs quality checks (type checking, tests)
5. Checks pass → git commit; fail → leave for the next iteration
6. Updates prd.json, marking the story as `passes: true`
7. Appends lessons learned to progress.txt
8. Repeats until all stories are complete or the iteration limit is reached
The default iteration limit is 10. Adjust based on project complexity:
```bash
# Simple projects
./scripts/ralph/ralph.sh --tool claude 10
# Complex projects
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — Task Definition
This is Ralph's "brain" — all tasks are defined here. The format is a flat JSON file:
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists and has valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve the MDX component props format"
}
]
}
```
**Field descriptions**:
| Field | Description |
| -------------------- | ----------------------------------------------------------------- |
| `projectName` | Project name, used for logging and branch naming |
| `branchName` | Git branch name, Ralph creates it automatically |
| `id` | Unique story identifier, `US-001` format recommended |
| `title` | Short title |
| `description` | Detailed description — the more specific, the better |
| `acceptanceCriteria` | List of acceptance criteria — **this is the most critical field** |
| `priority` | Priority number, lower values are executed first |
| `passes` | Whether completed, automatically updated by Ralph |
| `dependsOn` | List of dependent story IDs |
| `notes` | Additional notes and hints |
### progress.txt — Experience Log
This is Ralph's "long-term memory." After each iteration, the AI appends what it learned:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
The next iteration's new Claude instance reads this file and immediately gains all previous experience. This is why Ralph gets smoother over time — **knowledge accumulates across iterations while the context stays clean**.
### AGENTS.md — Persistent Knowledge Base
In addition to progress.txt, Ralph also updates the project's `AGENTS.md` file (or `CLAUDE.md`). Both Claude Code and Amp automatically read these files on startup.
Unlike progress.txt, AGENTS.md records **stable, cross-project knowledge**:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## Writing the PRD
The quality of the PRD (Product Requirements Document) directly determines Ralph's execution effectiveness. Write it well, and Ralph runs smoothly; write it poorly, and Ralph will repeatedly fail on the same story.
### Using the Skill to Generate a PRD
If you installed the snarktank/ralph skill, you can generate a PRD interactively:
```bash
# In Claude Code or Amp
/prd I want to add i18n support to the blog system, translating all Chinese content to English
```
The AI will ask you a series of clarifying questions (which files are involved, tech stack constraints, quality standards, etc.), then generate a structured PRD document.
After generation, use the `/ralph` command to convert the PRD to `prd.json` format:
```bash
/ralph # Convert PRD to prd.json
```
### Writing the PRD Manually
You can also write prd.json directly. Here are the key design principles.
**Principle 1: Get the Story Granularity Right**
Each story should be small enough to complete in a single iteration, yet large enough to deliver independent value.
```json
// ❌ Too large: can't finish in one iteration
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, password reset, OAuth, permission management..."
}
// ❌ Too small: no independent value
{
"id": "US-001",
"title": "Create email field for User table",
"description": "Add email field to the User model"
}
// ✅ Just right: completable in one iteration, delivers independent value
{
"id": "US-001",
"title": "Implement email-password login",
"description": "Create login API and login page with email-password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form can be submitted",
"All tests pass"
]
}
```
**Rule of thumb**: A story involves 1-3 file modifications and has 3-5 acceptance criteria.
**Principle 2: Acceptance Criteria Must Be Automatically Verifiable**
Ralph needs to determine whether a story is complete, so acceptance criteria must be objectively verifiable:
```json
// ❌ Vague criteria
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ Verifiable criteria
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**Principle 3: Use dependsOn to Control Execution Order**
Some stories have dependencies. The `dependsOn` field ensures Ralph executes them in the correct order:
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**Principle 4: Provide Context in notes**
The notes field gives additional hints to the AI. Write things you know but the AI might not:
```json
{
"notes": "The project uses the fumadocs framework. i18n file naming convention is the .en.mdx suffix. Reference the translation style of content/docs/notes/speckit/concept.en.mdx."
}
```
***
## Running the Ralph Loop
With the PRD ready, it's time to start the loop.
### Starting Execution
```bash
# Using Claude Code, default 10 iterations
./scripts/ralph/ralph.sh --tool claude
# Specify iteration count
./scripts/ralph/ralph.sh --tool claude 30
# Using Amp (default)
./scripts/ralph/ralph.sh 20
```
### Execution Process
After launching, you'll see output similar to this:
```
Starting Ralph - Tool: claude - Max iterations: 35
===============================================================
Ralph Iteration 1 of 35 (claude)
===============================================================
## US-001 Complete
**Summary of what was done:**
1. Created meta.en.json with all navigation items translated
2. Ran pnpm types:check — PASSED
3. Committed: feat: [US-001] - Translate homepage metadata
There are still **15 user stories with `passes: false`** remaining.
The next story is **US-002: Translate blog post hello-world**.
Iteration 1 complete. Continuing...
===============================================================
Ralph Iteration 2 of 35 (claude)
===============================================================
```
Each iteration is a brand-new Claude instance. It knows what to do by reading prd.json, and it knows what was learned previously by reading progress.txt.
### Completion Signal
When all stories are marked as `passes: true`, Ralph outputs the completion signal and exits:
```
All stories completed!
COMPLETE
```
### Monitoring & Debugging
While Ralph is running, you can use these commands to check progress:
```bash
# View completion status of each story (with icons for clarity)
cat tasks/prd.json | python3 -c "
import json,sys
for s in json.load(sys.stdin)['userStories']:
print(f'{\"✅\" if s[\"passes\"] else \"⬜\"} {s[\"id\"]}: {s[\"title\"]}')"
# Or use jq
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# View the experience log
cat progress.txt
# View recent git commits
git log --oneline -10
# Watch Ralph's output in real time
tail -f progress.txt
# After completion, view all changes compared to main branch
git diff main...ralph/your-branch-name --stat
```
### Interrupting & Resuming
Ralph may run for a long time, and interrupting midway is completely safe:
* **Interrupt**: Just press `Ctrl+C`. Completed stories (`passes: true`) won't be lost — they're already committed and written to prd.json
* **Resume**: Run the same command again. Ralph will automatically continue from the first story with `passes: false`
```bash
# Resume after interruption — just rerun the same command
./scripts/ralph/ralph.sh --tool claude 35
```
If a story repeatedly fails and blocks progress, you can skip it manually — edit `prd.json`, set that story's `passes` field to `true`, then rerun. Ralph will skip it and continue with subsequent stories.
### Auto-archiving
When you start a different feature with a new `branchName`, Ralph automatically archives the previous run's files to the `archive/YYYY-MM-DD-feature-name/` directory, keeping the working directory clean.
***
## Feedback Loops & Quality Gates
Ralph's "self-correction" ability depends entirely on the quality of the feedback loop. Without a feedback loop, Ralph is just a blindly looping script — it will keep producing code but can't tell if the code is correct.
### Configuring Quality Checks
Define your quality check commands in CLAUDE.md (or prompt.md):
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### Quality Gate Layers
| Layer | Tool | Issues Caught |
| ------------------------ | ------------------- | --------------------------------------- |
| Instant Feedback | TypeScript compiler | Type errors, syntax errors |
| Functional Verification | Unit tests | Logic errors, edge cases |
| Integration Verification | Build command | Dependency issues, configuration errors |
| Runtime Verification | dev-browser skill | UI rendering issues (frontend projects) |
> For frontend stories, Ralph recommends adding to the acceptance criteria: "Verify in browser using dev-browser skill" — letting the AI actually open a browser to confirm the page renders correctly.
### When Quality Checks Fail
If a story's quality checks fail repeatedly, Ralph won't infinitely retry the same story. After reaching the iteration limit, it stops and preserves the current state. You can:
1. Check progress.txt to see where the AI got stuck
2. Manually fix the issue and rerun
3. Adjust the story granularity (it might be too large)
4. Add more context in notes
### Prompt Customization
Ralph's prompt template (CLAUDE.md or prompt.md) is your primary means of controlling AI behavior. After installation, you should customize it for your project. Key customization areas:
**Code Style Constraints**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**Common Pitfalls**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**Handling Stuck Situations**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## Real-World Case Study: Using Ralph for Blog i18n Translation
To demonstrate how Ralph works in a real project, here's an actual case study: using a Ralph-style autonomous agent to translate an entire blog from Chinese to English.
### Project Setup
The project needed to translate 22+ content files (blog posts, documentation, navigation metadata) from Chinese to English, targeting i18n support for a fumadocs-based Next.js blog. Tasks were defined in a `prd.json` file containing 16 user stories, each with clear acceptance criteria:
```
scripts/ralph/
├── prd.json # 16 user stories with acceptance criteria
└── progress.txt # Experience log, updated after each story
```
Each user story followed a consistent pattern:
* **Clear deliverables**: "Create content/blog/xxx.en.mdx"
* **Verifiable criteria**: "Typecheck passes", "Internal links use /en/ prefix"
* **Technical constraints**: "Keep code blocks untranslated", "Set defaultLang='en' on QuoteCard"
### Execution Model
The agent followed the core principles of the Ralph methodology:
1. **Files as the source of truth**: `prd.json` tracks each story's status (`passes: true/false`). `progress.txt` accumulates experience across iterations — such as "Typecheck command is `pnpm types:check`, not `pnpm typecheck`"
2. **Automated quality gates**: After each translation, `pnpm types:check` verifies the MDX file compiles correctly. If typecheck fails, fix the issue before committing.
3. **Incremental progress**: Each story is committed independently with descriptive commit messages (`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`), making rollbacks easy when needed.
4. **Parallel execution**: For longer articles, multiple subagents translate simultaneously — for example, US-010 (claude-skills concept + practice), US-011 (speckit concept + practice), and US-012 (claude-architecture + claude-subagent) all ran in parallel.
### Key Takeaways
| Takeaway | Details |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| **Knowledge accumulation matters** | Patterns discovered in early stories (QuoteCard's `defaultLang`, link prefix rules) made later stories faster to complete |
| **Typecheck as a feedback loop** | Catches missing imports or malformed MDX before issues compound |
| **Parallelization scales** | 6 translation agents running simultaneously finished in about the same time as 1 |
| **PRD granularity is critical** | Each story scoped to 1-2 files — small enough to reliably complete, large enough to be meaningful |
| **Progress logs prevent repeated mistakes** | The "Codebase Patterns" section of `progress.txt` became a knowledge base, avoiding stepping on the same pitfalls |
### Results
All 16 user stories were completed in a single session: 8 meta.en.json navigation files created, 3 blog posts translated, 12 documentation pages translated, and full site build verification passed. Each translation maintained consistent quality because the acceptance criteria were explicit and the feedback loop (typecheck) caught issues immediately.
This project demonstrated Ralph's **Full Implementation Mode** — clearly defined tasks, explicit success criteria, automated verification, and incremental delivery through the file system.
***
## Best Practices & FAQ
### Cost Control
Ralph's automated execution means API costs are ongoing. Several control measures:
* **Always set `max_iterations`**: This is the most basic safety net
* **Keep story granularity reasonable**: Stories that are too large consume multiple iterations; stories that are too granular increase startup overhead
* **Test small first**: For new projects, try 3-5 iterations first to confirm the prompt and quality gates work properly before scaling up
### Common Pitfalls
**Pitfall 1: Story Too Large**
Symptom: A story fails repeatedly, quickly exhausting the iteration count.
Solution: Split it into 2-3 smaller stories. "Build complete authentication system" becomes "Implement login API" + "Create login page" + "Add JWT middleware."
**Pitfall 2: No Feedback Loop**
Symptom: Ralph claims the story is complete, but the code actually has issues.
Solution: Add executable check commands to your acceptance criteria. "Code is written" is not an acceptance criterion — "pnpm test all pass" is.
**Pitfall 3: progress.txt Not Being Utilized**
Symptom: The same error keeps appearing across different iterations.
Solution: Confirm your prompt template explicitly instructs "read progress.txt and follow the lessons within." If the AI isn't automatically appending learnings, add "After each story, append learnings to progress.txt" to the prompt.
**Pitfall 4: Incorrect Dependency Order**
Symptom: A story depends on code that doesn't exist yet, causing implementation failure.
Solution: Set the `dependsOn` field correctly, ensuring infrastructure stories come first.
### FAQ
**Q: Can I manually modify prd.json mid-run?**
Yes. Ralph re-reads prd.json at the start of each iteration. You can modify story descriptions, add new stories, or manually mark a story as `passes: true` (to skip it) between iterations.
**Q: What if Ralph gets stuck on one story and keeps failing?**
1. Check progress.txt for the failure reason
2. Add more context in notes
3. Split the story (the granularity might be too coarse)
4. Manually fix the blocking issue and rerun
**Q: Can I do other things while Ralph is running?**
Yes. Ralph is designed for "Human on the Loop" — you don't need to watch it. In AFK mode, start it before leaving work and check the results the next day. Just don't modify files that Ralph is currently working on.
**Q: How do I control costs?**
Three ways: set a reasonable `max_iterations`, keep story granularity appropriate (reducing wasted iterations), and run a small-scale test first to confirm the process works. Generally, a project with 10-20 stories falls within the $50-100 API cost range.
***
## Summary
Ralph's workflow can be summarized in five steps:
```
Install → Write PRD → Configure Quality Gates → Run Loop → Review Results
```
The core philosophy never changes: **Let files be the source of truth, let every iteration start fresh, and let quality gates do the gatekeeping for you**.
Now, go back to your project, prepare your prd.json, run `./scripts/ralph/ralph.sh --tool claude`, and go grab a cup of coffee.
### Further Reading
* [Ralph Wiggum Deep Dive](/en/docs/notes/ralph-wiggum/concept) — Review Ralph's core principles
* [frankbria/ralph-claude-code Practical Guide](/en/docs/notes/ralph-wiggum/frankbria) — Engineering-grade Ralph implementation: monitoring, circuit breakers, and safety mechanisms
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — A complete context engineering system built on top of Ralph
* [What Are Claude Skills](/en/docs/notes/claude-skills/concept) — Ralph's PRD skill is a Claude Skill
* [Speckit Practical Guide](/en/docs/notes/speckit/practice) — Another structured AI coding workflow
# Concept Introduction
## Introduction
In October 2025, GitHub open-sourced a toolkit called Spec Kit, officially bringing the concept of "Spec-Driven Development" into the AI programming landscape. This seemingly retro idea -- write specs before code -- is becoming the new paradigm for harnessing AI programming tools.
If you regularly use AI coding assistants like Claude Code, Cursor, or GitHub Copilot, you have likely encountered this frustration: you say "add a user login feature," the AI eagerly produces a heap of code, but upon closer inspection -- it uses an unfamiliar framework, the security strategy differs from what you expected, the UI style does not match... Then you enter round after round of corrections until you are exhausted.
Where does the problem lie? It is not that the AI is not smart enough -- it is that you have not provided enough information. "Add a user login feature" seems clear, but it actually hides hundreds of unstated decisions: What authentication method? What are the password requirements? How to handle failed logins? Should the login state be remembered? Support third-party login? ... The AI has no choice but to guess, and guessing means deviation.
Spec-driven development was born to solve exactly this problem.
## Vibe Coding: The Cost of Speed
In early 2025, former Tesla AI Director Andrej Karpathy coined the term "Vibe Coding" to describe a development approach of "accepting AI suggestions without deep review." The term went viral and was even named Collins Dictionary's Word of the Year for 2025.
The appeal of Vibe Coding is obvious: you describe an idea, the AI generates code, and if it looks like it works, that is good enough. For rapid prototyping, hackathons, and throwaway scripts, this approach is genuinely efficient. But when it is applied to production systems, problems emerge.
Stories like this are commonplace in the industry: AI-generated database queries run fine in small-scale tests but crawl under real traffic; a patched-together authentication module passes QA, only for someone to discover two weeks later that deactivated accounts can still access admin tools. According to a 2025 survey by Final Round AI, **16 out of 18 CTOs have experienced production disasters caused by AI-generated code**.
This does not mean Vibe Coding is worthless. The key is **knowing its boundaries**:
| Scenario | Vibe Coding | Spec-Driven Development |
| ------------------- | ---------------- | ----------------------- |
| Prototypes / demos | Suitable | Overkill |
| Throwaway scripts | Suitable | Overkill |
| Production features | High risk | Recommended |
| Security-related | Dangerous | Essential |
| Team collaboration | Hard to maintain | Recommended |
Spec-driven development aims to retain the efficiency of AI while avoiding the pitfalls of Vibe Coding.
## What Is Spec-Driven Development
The core idea of spec-driven development can be summarized in one sentence: **Define "what to build" before considering "how to build it"**.
This sounds like a software engineering cliche, but in the age of AI programming, it takes on new meaning. Traditional requirements documents are written for humans -- often lengthy, vague, and filled with jargon. The "spec" in spec-driven development is written for AI -- concise, structured, and actionable.
Imagine you are building a house. The traditional AI programming approach is like telling the construction crew "build me a comfortable three-bedroom house" and letting them figure it out. The result might be fine, but more likely it will be vastly different from what you imagined. Spec-driven development means drawing the architectural blueprint first: how many floors, square footage per floor, window orientation, material specifications... The crew builds according to the blueprint, and the result naturally meets expectations.
In AI programming, this blueprint is the **specification**. It does not concern itself with what programming language or framework to use -- only what the feature should achieve, what tasks users need to complete, and what the success criteria are.
Compared to traditional development workflows, spec-driven development has a fundamental difference:
| Traditional AI Programming | Spec-Driven Development |
| --------------------------------------------------- | ------------------------------------------------ |
| Describe requirements directly -> AI generates code | Requirements -> Spec -> Plan -> Tasks -> Code |
| AI must guess many details | Every step is explicit; AI only needs to execute |
| Frequent rework, high communication cost | Upfront investment, smooth execution later |
| Suited for simple tasks | Suited for complex features |
This process of "progressive refinement" is the essence of spec-driven development. You do not go from zero to code in one leap -- instead, you progressively clarify requirements through multiple stages, each of which can be reviewed and adjusted.
## Speckit Workflow Overview
GitHub's Spec Kit and the speckit command in Claude Code both follow a similar workflow, roughly divided into six stages:
```
Constitution -> Specify -> Clarify -> Plan -> Tasks -> Implement
| | | | | |
Project Feature Resolve Technical Task Execute
Charter Spec Ambiguity Plan Breakdown Implementation
```
**1. Constitution (Project Charter)**
The project constitution defines the fundamental principles and constraints for the entire project, such as "test first," "simplicity above all," "API first," etc. These principles carry through all subsequent stages, ensuring AI-generated solutions align with your technical preferences.
**2. Specify (Feature Spec)**
This is the critical first step. You describe the desired feature in natural language, and the AI helps organize it into a structured specification document, including:
* User stories: who needs to do what, and why
* Functional requirements: capabilities the system must have
* Success criteria: how to determine if the feature meets the bar
Importantly, the spec document **focuses only on "what to build," not "how to build it"** -- no specific tech stacks, no code structure.
**3. Clarify (Resolve Ambiguity)**
The AI reviews the spec for ambiguous points and asks up to 5 key questions. These typically involve feature boundaries, user types, security requirements, etc. Through this Q\&A, the spec becomes much clearer.
**4. Plan (Technical Plan)**
Only after having a clear spec do you begin considering the technical approach. This step produces:
* Technology choices (language, framework, database)
* Data model design
* API contract definitions
* Research reports (resolving technical decisions)
**5. Tasks (Task Breakdown)**
The technical plan is broken down into an executable task list. Each task has a clear ID, description, and file paths, ready to be handed off to the AI for execution. Tasks are grouped by user story and support parallel development.
**6. Implement (Execute)**
Tasks are executed one by one from the list. Each completed task is marked as done, ensuring traceability.
The deliverables from these six stages form a clear chain:
| Stage | Deliverable | Purpose |
| ------------ | -------------------- | -------------------------------- |
| Constitution | constitution.md | Define project principles |
| Specify | spec.md | Describe functional requirements |
| Clarify | Updated spec.md | Eliminate ambiguity |
| Plan | plan.md, research.md | Technical solution design |
| Tasks | tasks.md | Executable task list |
| Implement | Actual code | Final deliverable |
## Why This Approach Works
Spec-driven development succeeds where "just let AI write code" fails because it addresses the core contradiction of AI programming: **information asymmetry**.
When you say "add a photo sharing feature," you may have a complete picture in your mind, but the AI only sees those few words. It must guess: share to where? Who can see it? Does it need compression? Watermarks? Batch support? ... Every guess could be wrong.
Spec-driven development solves this by "forcing you to think it through first." When you are required to write down user stories, functional requirements, and success criteria, the details you assumed were "obvious" surface. This process itself is valuable -- even without AI, clearly writing out requirements reduces communication costs.
Furthermore, progressive refinement exposes errors earlier. Discovering a requirements deviation at the Specify stage costs virtually nothing to fix; discovering it after the code is written might mean starting over.
Of course, spec-driven development is not a silver bullet. It has clear use cases:
**Good fit**:
* Complex feature development (involving multiple modules and interactions)
* Team collaboration projects (spec documents serve as communication medium)
* High-quality scenarios (requiring traceability and verifiability)
**Not a good fit**:
* Simple bug fixes or minor changes
* Exploratory programming (when you do not yet know what to build)
* Extremely time-constrained situations (no time to write specs)
The key is recognizing task complexity. Something that takes an hour to complete does not need an hour of spec writing; a week-long feature absolutely warrants two hours of spec writing.
## But Specs Are Not a Silver Bullet
One common misconception needs clarifying: **spec-driven development reduces guessing, but it does not eliminate the need for review**.
Even with a complete spec, AI may still:
* Miss edge cases (extreme scenarios the spec did not cover)
* Generate code that does not meet performance requirements
* Introduce potential security vulnerabilities
* Produce implementations with inconsistent style
This is like construction: even with detailed blueprints, inspection is still necessary. You would not move into a house just because the crew finished building it according to the plans -- you would check that the wiring is safe, the plumbing works, and the doors and windows are secure.
The value of spec-driven development lies in **making errors easier to find**, not in eliminating errors themselves.
## Summary
The core of spec-driven development is a simple truth: **the more complex the task, the more you need to think it through before starting**. AI programming tools amplify the importance of this truth -- because AI will faithfully execute your instructions, but cannot truly understand your intent.
Remember three key takeaways:
| Takeaway | Meaning |
| -------------------------- | ------------------------------------------------------------- |
| **Spec before code** | Define "what to build" before considering "how to build it" |
| **Progressive refinement** | From vague to clear, with review and adjustment at every step |
| **Reduce guessing** | Clear specs = less room for AI speculation |
Now that you understand the philosophy, the next article, [Speckit Practical Guide](/en/docs/notes/speckit/practice), will walk you through hands-on practice: how to use the speckit command to complete a spec-driven development workflow for a feature.
Combined with [Claude Skills](/en/docs/notes/claude-skills/concept), you can further automate and standardize the execution of specs.
# Practical Guide
## Introduction
In the [previous article](/en/docs/notes/speckit/concept), we explored the philosophy of spec-driven development — define "what to build" before thinking about "how to build it." While this extra step may seem unnecessary, it dramatically reduces rework and communication overhead in AI-assisted programming.
In this article, we get hands-on. You will learn how to use the speckit command suite to complete the entire workflow from requirements to working code.
## Installation and Configuration
Speckit commands originate from GitHub's official [Spec Kit](https://github.com/github/spec-kit) project. Depending on your use case, there are several ways to integrate them.
### New Project Initialization
For new projects, the recommended approach is to use the official specify-cli tool:
```bash
# Install specify-cli using uv
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git
# Initialize a new project, specifying Claude as the AI assistant
specify init my-project --ai claude
```
This automatically creates the project directory structure, including the `.specify/` configuration directory and related template files.
### Integrating with an Existing Project
Speckit commands require configuration files to work. To integrate speckit into an existing project, use specify-cli:
```bash
cd your-existing-project
specify init . --ai claude # Note: . refers to the current directory
```
This creates the following in your project:
```
your-project/
├── .specify/
│ ├── templates/ # Spec, plan, and other templates
│ ├── scripts/ # Helper scripts
│ └── memory/ # constitution.md
├── .claude/
│ └── commands/ # Claude Code command configs
│ ├── speckit.specify.md
│ ├── speckit.plan.md
│ └── ...
└── specs/ # Feature spec storage directory
```
The initialization will not overwrite your existing files. Once complete, you can use the `/speckit.*` command suite in Claude Code.
> **Note**: Speckit commands are not built into Claude Code — you must complete the initialization steps above first. Running `/speckit.specify` without initialization will result in a "command not found" error.
***
## Command Reference
Speckit provides a set of commands that support each phase of spec-driven development. Each command has well-defined inputs and outputs, forming a traceable chain.
### /speckit.specify — Create a Feature Spec
This is the starting point of the entire workflow. You describe the feature you want in natural language, and the AI organizes it into a structured spec document.
**Purpose**: Create a feature spec from a natural language description
**Input**: Feature description (natural language)
**Output**:
* `specs/[number]-[feature-name]/spec.md` — Feature spec document
* A new git branch (e.g., `001-user-auth`)
**Usage example**:
```
/speckit.specify I want to add a user login feature with email/password authentication and a "remember me" option
```
After execution, the AI will:
1. Generate a short feature name (e.g., `user-auth`)
2. Create a new feature branch
3. Produce a spec document with user stories, functional requirements, and success criteria
4. Mark unclear areas with `[NEEDS CLARIFICATION]`
**Core structure of a spec document**:
```markdown
# Feature Specification: User Login
## User Scenarios & Testing
### User Story 1 - User Login (Priority: P1)
Users log into the system using email and password...
**Acceptance Scenarios**:
1. Given valid email and password, When clicking login, Then successfully enter the system
## Requirements
### Functional Requirements
- FR-001: The system must support email/password login
- FR-002: The system must provide a "remember me" option
## Success Criteria
- SC-001: Users can complete the login flow within 30 seconds
```
Note that the spec document **contains no technical details** — no mention of frameworks, database schemas, or API definitions. Those come in later phases.
***
### /speckit.clarify — Resolve Ambiguities
After the spec is drafted, there may still be ambiguous areas. This command reviews the spec and asks key questions to help clarify them.
**Purpose**: Identify ambiguities in the spec and refine it through Q\&A
**Input**: Existing spec.md document
**Output**: Updated spec.md (with clarification records)
**Usage example**:
```
/speckit.clarify
```
After execution, the AI will:
1. Scan the spec for ambiguous points
2. Prioritize them (Scope > Security > UX > Technical Details)
3. Ask one question at a time
4. Update the spec based on your answers
**Q\&A example**:
```markdown
## Question 1: Handling Login Failures
**Context**: The spec mentions user login but does not specify how login failures should be handled.
**Recommended:** Option B - Locking the account after 5 consecutive failures is a security best practice
| Option | Description |
|--------|-------------|
| A | Show error message only, no restrictions |
| B | Lock account for 15 minutes after 5 consecutive failures |
| C | Use CAPTCHA to prevent brute force attacks |
You can reply with an option letter (e.g., "B"), say "yes" to accept the recommendation, or provide your own answer.
```
After each clarification, the spec document is automatically updated with a clarification record:
```markdown
## Clarifications
### Session 2025-12-20
- Q: How should login failures be handled? → A: Lock account for 15 minutes after 5 consecutive failures
```
***
### /speckit.plan — Generate a Technical Plan
Once the spec is clear, you move into the technical design phase. This step produces a technical plan and research report.
**Purpose**: Generate a technical implementation plan from the spec
**Input**: spec.md document
**Output**:
* `plan.md` — Technical plan (architecture, data models, API design)
* `research.md` — Research report (technology selection decisions)
* `data-model.md` — Data model (if applicable)
* `contracts/` — API contracts (if applicable)
**Usage example**:
```
/speckit.plan I'm using Next.js + Prisma + PostgreSQL
```
You can append your tech stack preferences after the command. After execution, the AI will:
1. Analyze the functional requirements from the spec
2. Research best practices for the relevant technologies
3. Design data models and API structures
4. Produce a complete technical plan
**Core content of a technical plan**:
```markdown
# Implementation Plan: User Login
## Technical Context
**Language/Version**: TypeScript 5.x
**Primary Dependencies**: Next.js 15, Prisma, PostgreSQL
**Authentication**: NextAuth.js with credentials provider
## Project Structure
src/
├── app/
│ └── (auth)/
│ ├── login/
│ └── api/auth/
├── lib/
│ └── auth/
└── prisma/
└── schema.prisma
## Data Model
- User: id, email, passwordHash, createdAt, updatedAt
- Session: id, userId, expiresAt
```
***
### /speckit.tasks — Break Down Tasks
With the technical plan in place, the next step is to decompose it into an actionable task list.
**Purpose**: Split the technical plan into an executable task list
**Input**: plan.md document
**Output**: `tasks.md` — Task list sorted by dependencies
**Usage example**:
```
/speckit.tasks
```
After execution, the AI will:
1. Extract the technical approach from plan.md
2. Extract user story priorities from spec.md
3. Generate tasks grouped by user story
4. Mark parallelizable tasks with `[P]`
5. Specify concrete file paths for each task
**Task list format**:
```markdown
## Phase 1: Setup
- [ ] T001 Create project structure
- [ ] T002 [P] Configure Prisma schema
- [ ] T003 [P] Configure NextAuth
## Phase 2: User Story 1 - User Login (P1)
- [ ] T004 [US1] Create User model in prisma/schema.prisma
- [ ] T005 [US1] Implement login API in src/app/api/auth/[...nextauth]/route.ts
- [ ] T006 [US1] Create login page in src/app/(auth)/login/page.tsx
```
Each task includes:
* **Task ID** (T001, T002...) — for tracking
* **\[P] marker** — indicates it can run in parallel with other \[P] tasks
* **\[US] tag** — indicates which user story it belongs to
* **File path** — specifies exactly which file to work on
***
### /speckit.implement — Execute Implementation
Everything is ready — time to execute the task list.
**Purpose**: Execute tasks from the task list one by one
**Input**: tasks.md document
**Output**: Actual code
**Usage example**:
```
/speckit.implement
```
Before execution, the AI checks the checklist (if one exists). During execution:
1. Tasks are executed in phase order
2. Each completed task is marked as `[X]`
3. Task dependencies are respected
4. Parallel tasks can run simultaneously
**Execution example**:
```
Phase 1: Setup
✓ T001 Create project structure
✓ T002 Configure Prisma schema
✓ T003 Configure NextAuth
Phase 2: User Story 1
✓ T004 Create User model
Executing T005...
```
### Post-Implementation Review
After `/speckit.implement` finishes, **do not merge the code directly**. AI-generated code requires human review:
**Required verification steps**:
1. **Run the test suite**
```bash
npm test # or your test command
```
Ensure the AI hasn't broken existing functionality.
2. **Code review checklist**
* Does the code match the spec's intent (cross-reference spec.md)?
* Does it follow the project's coding style?
* Are there any potential security issues?
3. **Boundary testing**
Manually test edge cases the AI may have missed:
* Null value handling
* Extreme inputs
* Concurrency scenarios
* Error paths
4. **Performance check**
If database operations or API calls are involved, check for N+1 queries and similar performance issues.
> **Tip**: Even with a thorough spec, the AI may still deviate in implementation details. Review is not a sign of distrust in spec-driven development — it is part of engineering discipline.
***
### /speckit.analyze — Consistency Analysis
This is an optional quality check step that verifies consistency across the spec, plan, and tasks.
**Purpose**: Cross-document consistency and quality analysis
**Input**: spec.md, plan.md, tasks.md
**Output**: Analysis report (no files are modified)
**Usage example**:
```
/speckit.analyze
```
After execution, it checks:
* Whether every requirement has a corresponding task
* Whether tasks cover all user stories
* Whether terminology is consistent
* Whether there are gaps or duplications
***
### Other Commands (Optional)
In addition to the core commands above, speckit provides several auxiliary commands. These are not part of the main workflow but are useful in specific scenarios.
**`/speckit.constitution`** — Create a Project Constitution
Used to define development principles and standards for the project. Ideal for team projects to ensure all members follow unified development standards.
* **Input**: Interactive Q\&A or directly provided principles
* **Output**: `.specify/constitution.md` project constitution file
* **Use case**: New team project initialization, unifying coding style and architectural decisions
**`/speckit.checklist`** — Generate a Quality Checklist
Generates a customized quality checklist based on the feature spec, used for pre-implementation quality assurance.
* **Input**: spec.md document
* **Output**: Checklists in the `checklists/` directory
* **Use case**: Quality gates before important feature launches, code review reference
**`/speckit.taskstoissues`** — Convert Tasks to GitHub Issues
Automatically converts tasks from tasks.md into GitHub Issues for team collaboration and task assignment.
* **Input**: tasks.md document
* **Output**: GitHub Issues (created via the gh CLI)
* **Use case**: Team collaboration, sprint planning, task tracking
***
## Tool Ecosystem
The speckit commands introduced in this article come from the [GitHub Spec Kit](https://github.com/github/spec-kit) project. Beyond this, in 2025, several major AI coding tools began supporting similar spec-driven workflows:
| Tool | Features | Best For |
| --------------------------------------------------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------- |
| **[GitHub Spec Kit](https://github.com/github/spec-kit)** | The tool used in this article, MIT licensed, supports Claude Code / Copilot / Gemini CLI | Command-line users, cross-tool collaboration |
| **[AWS Kiro](https://kiro.dev/)** | VS Code fork, visual workflow, EARS notation | GUI-oriented users, AWS ecosystem |
| **[JetBrains Junie](https://blog.jetbrains.com/junie/)** | IntelliJ ecosystem integration, Think More reasoning mode | JetBrains IDE users |
| **Cursor Plan Mode** | Built-in planning phase, auto-generated execution plans | Developers already using Cursor |
**How to choose**:
* If you use Claude Code, GitHub Copilot, or Gemini CLI, GitHub Spec Kit is recommended
* If you prefer graphical interfaces and visual workflows, try AWS Kiro
* If you are a JetBrains user, Junie offers more natural IDE integration
* If you are already using Cursor, its Plan Mode provides similar planning capabilities
The core philosophy is the same — the tool is just the vehicle; what matters is the **spec-first, code-second** mindset.
***
## End-to-End Walkthrough
Let us walk through a real-world example using the complete speckit workflow.
**Scenario**: Add an article bookmarking feature to a blog system
### Step 1: Create the Spec
```
/speckit.specify I want to add an article bookmarking feature so users can save their favorite articles and view them in their profile
```
The AI generates the spec document `specs/002-article-bookmark/spec.md`:
```markdown
# Feature Specification: Article Bookmarking
## User Scenarios & Testing
### User Story 1 - Bookmark an Article (Priority: P1)
While reading an article, the user clicks the bookmark button to add it to their bookmark list.
**Acceptance Scenarios**:
1. Given the user is logged in, When clicking the bookmark button, Then the article is added to the bookmark list
2. Given the user is not logged in, When clicking the bookmark button, Then a login prompt is shown
### User Story 2 - View Bookmark List (Priority: P2)
The user views all their bookmarked articles in their profile.
## Requirements
- FR-001: Users must be able to bookmark/unbookmark articles
- FR-002: The bookmark button must reflect the current bookmark state
- FR-003: The profile must display the bookmark list
## Success Criteria
- SC-001: Bookmark operations complete within 500ms
- SC-002: The bookmark list supports pagination with 10 articles per page
```
### Step 2: Clarify Requirements
```
/speckit.clarify
```
The AI asks: "Is there an upper limit on the number of bookmarks?"
Answer: "A maximum of 100 bookmarked articles per user"
The spec is updated with:
* FR-004: Each user can bookmark a maximum of 100 articles
* A notification is displayed when the limit is reached
### Step 3: Generate the Plan
```
/speckit.plan Using Next.js + Prisma
```
The generated technical plan includes:
* Bookmark model (userId, articleId, createdAt)
* API route design (POST/DELETE /api/bookmarks)
* Component design (BookmarkButton, BookmarkList)
### Step 4: Break Down Tasks
```
/speckit.tasks
```
The generated task list:
```markdown
## Phase 1: Setup
- [ ] T001 Add Bookmark model to Prisma schema
## Phase 2: US1 - Bookmark Article
- [ ] T002 [US1] Create bookmark API in src/app/api/bookmarks/route.ts
- [ ] T003 [US1] Create BookmarkButton component in src/components/BookmarkButton.tsx
- [ ] T004 [US1] Integrate into article page
## Phase 3: US2 - Bookmark List
- [ ] T005 [US2] Create bookmark list page in src/app/profile/bookmarks/page.tsx
- [ ] T006 [US2] Implement pagination logic
```
### Step 5: Execute Implementation
```
/speckit.implement
```
Tasks are executed in order, with each completed task marked as `[X]`.
***
## Best Practices and Considerations
### When to Use Speckit
**Good fit**:
* New feature development (involving 3+ files)
* When requirements are not fully clear (use clarify to resolve)
* Multi-person collaborative projects (specs serve as shared understanding)
* Critical features (where traceability is needed)
**Not a good fit**:
* Simple bug fixes
* One-line code changes
* Emergency hotfixes
* Purely exploratory experiments
### Common Pitfalls
There are several common pitfalls to watch out for when using speckit:
**Pitfall 1: Specs that are too vague**
Symptom: AI-generated code diverges significantly from expectations, requiring extensive rework.
```markdown
# ❌ Vague spec
Users can search for articles
# ✓ Clear spec
- FR-001: Users can search articles by title keywords
- FR-002: Search results are sorted by relevance, showing 10 per page
- FR-003: Search terms are highlighted in the results
- FR-004: An empty search term displays trending articles
```
Solution: Run `/speckit.clarify`, or manually add functional requirements and success criteria.
**Pitfall 2: Specs that are too detailed**
Symptom: The AI is over-constrained and cannot leverage its strengths, producing rigid code — or it simply ignores parts of the instructions.
```markdown
# ❌ Over-specified (dictating implementation details)
Use lodash's debounce function with a 300ms delay,
wrapped in useCallback with [searchTerm] as a dependency...
# ✓ Appropriate level of detail (only state what, not how)
Search input should be debounced to avoid excessive requests
```
Solution: Keep specs at the "what" level and leave "how" for the Plan phase.
**Pitfall 3: Skipping the Plan phase**
Symptom: Tasks are too coarse or too fragmented, leading to frequent rework during implementation and tangled dependencies between tasks.
Solution: Always complete the Plan phase for complex features. Planning not only produces a technical approach but also helps identify potential architectural issues.
**Pitfall 4: Merging without review**
Symptom: Edge cases, security vulnerabilities, or performance issues are discovered after deployment.
Solution: Refer to the "Post-Implementation Review" section above — always run tests and perform code review before merging.
### Frequently Asked Questions
**Q: Do I need to go through the entire workflow for every feature?**
Not necessarily. Simple changes can go straight to code. For complex features, completing at least specify + plan is recommended.
**Q: The spec is very detailed, but the AI still generated unexpected code?**
Check whether the spec is truly "detailed." Often we think we have been clear, but ambiguities remain. Try running `/speckit.clarify` to see if anything was missed.
**Q: Can I skip certain steps?**
Yes. The minimal workflow is specify → tasks → implement. However, skipping clarify and plan may increase the risk of rework later.
**Q: How do I modify an already-generated spec?**
Simply edit the spec.md file directly. After making changes, it is recommended to re-run plan and tasks to maintain consistency.
**Q: The AI-generated code is completely wrong — how do I debug?**
Troubleshoot in stages:
1. **Check the spec**: Is the spec truly clear? Try running `/speckit.clarify` to see if anything was missed
2. **Check the plan**: Is the technical approach in plan.md reasonable? If not, edit it directly and regenerate tasks
3. **Narrow the scope**: Have the AI execute just one task and observe whether the output meets expectations
4. **Add constraints**: Add more explicit technical preferences in constitution.md
**Q: What if the Plan and Tasks are inconsistent?**
Run `/speckit.analyze` to detect inconsistencies. Common causes:
* Plan was updated but Tasks were not regenerated
* Tasks were manually edited without updating the Plan
* Spec changed but only some documents were updated
Solution: Treat spec.md as the source of truth and regenerate plan.md and tasks.md in sequence.
**Q: How do I handle cross-feature dependencies?**
If Feature B depends on Feature A, there are two approaches:
1. **Merge specs**: Write A and B into the same spec.md so the AI plans them together
2. **Develop in phases**: Complete Feature A's full workflow first, then start Feature B's specify
It is not recommended to develop multiple features with dependencies simultaneously, as this easily leads to integration issues.
***
## Summary
The core value of speckit is not adding process for its own sake — it is about **making implicit knowledge explicit**. When you are required to write down user stories, functional requirements, and success criteria, the details you thought were "obvious" surface naturally.
Remember this workflow:
```
Specify → Clarify → Plan → Tasks → Implement
Define Refine Design Break Execute
Down
```
Each step reduces ambiguity for the next. Ultimately, the AI receives a clear task list instead of a vague description of intent.
Now, go back to your project and try starting your first spec-driven development workflow with `/speckit.specify`.
### Further Reading
* [What is Spec-Driven Development](/en/docs/notes/speckit/concept) — Revisit the core philosophy
* [GSD Deep Dive](/en/docs/notes/gsd/concept) — Another context engineering system built on spec-driven principles
* [What are Claude Skills](/en/docs/notes/claude-skills/concept) — Speckit itself is a Claude Skill
# TDD in the AI era: Let the model hit the red light first
## Let’s talk about the conclusion first
After AI writes code, TDD is not outdated, but has changed its position.
In the past, when we talked about TDD, we often talked about the self-discipline of programmers: write tests first, then write implementation, and refactor in small steps. When it comes to AI programming, it is more like a braking system. Because what the model is best at is also the most dangerous thing: it can quickly write a large piece of code that looks complete.
If you ask it to implement a function, it may give you:
* an implementation file
* a set of tests
* an explanation
* A word "done"
The thing is, "looking complete" is not done in the engineering sense. To be completed in an engineering sense, at least answer:
> Is this behavior defined by an explicit failing test?
> Did this failure turn green due to implementation?
> After turning green, have we tidied up the code without changing the behavior?
This is why TDD is being talked about again in the AI era.
It's not about making the process look advanced, but about replacing "trust in the model" with "trust in feedback."
## 1. Common misunderstanding: TDD is not “write tests first”
Many people hate TDD because they understand it as a ritual:
```text
先写测试。
再写代码。
最后跑一下。
```
This is certainly boring and can easily turn into formalism.
The truly useful TDD is not "test files appear early", but "failures appear early enough".
### The key is not testing, but red
The first step of TDD is called RED, not TEST.
RED means: write a test first so that the system clearly fails. Three things must be true for this to fail:
1. It does fail.
2. It fails because the target behavior does not exist.
3. It fails in the way you expect.
If the red is not seen first, the green that follows is meaningless.
For example, you want to implement `slugify("Hello World") -> "hello-world"`. A valuable RED is not "I wrote a test file", but:
```text
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
Failure: NameError: name 'slugify' is not defined
Reason: 目标函数还不存在,符合预期
```
This is when testing becomes specification. It tells you: To achieve the next step, you only need to make this behavior true.
### Go green first and then make up the test, usually making up the story
It's easy for AI to go the other way: write the implementation first and add testing later.
This is a smooth experience. When you see that the code has been run and the tests are in place, you will feel "almost" in your heart. But it has a fatal problem: the test is likely to be just a retroactive implementation of the current implementation.
It is not asking "what the requirements should be", but "how to write the current code so that it can easily pass".
This is why AI writing tests often have these smells:
* Assertions are too specific to the current implementation
* There are too many mocks and the real boundaries are not measured.
* only test happy path
* To make existing code pass, make assertions very wide
* No test can prove that the old code was originally wrong
TDD requires the opposite: first let the requirements fail, and then let the code catch up with the requirements.
## 2. Why TDD is more needed in the AI era
The core contradiction of AI programming is not "code is written slowly", but "feedback comes late".
Without TDD, you usually work like this:
```text
描述需求 -> AI 写一堆代码 -> 人肉看 diff -> 跑一下 -> 发现问题 -> 回头修
```
The questions will pile up until the end. By the time you find out it's wrong, there may have been three categories of things mixed together:
* Misunderstanding of requirements
* The implementation path is wrong
* Refactoring breaks old behavior
The role of TDD is to shorten this long chain.
### It gives the model a decidable goal
"Write elegantly" is not the goal.
"Automatically jumping back to the login page after the user's login status expires" is not specific enough.
A better goal would be:
```text
当 access token 过期时:
1. 请求返回 401。
2. 客户端清理本地 session。
3. 用户被重定向到 /login。
4. 原始目标地址被保存在 redirect 参数里。
```
Go one step further and turn one of them into a failing test:
```text
given expired session
when user opens /settings
then app redirects to /login?redirect=/settings
```
At this time, the AI is no longer guessing "how to handle the expiration of the login state", but completing a clear behavior.
### It breaks large tasks into small closed loops
The easiest place for AI to lose control is to do it all in one go.
Let it implement login, permissions, refresh token, error prompts, and route jumps at once, and you will get a big diff in the end. It might work, but the review cost is high. You have to judge business, status, routing, boundaries, testing and refactoring at the same time.
The rhythm of TDD looks more like this:
```text
一个行为 -> 一个失败测试 -> 最小实现 -> 变绿 -> 再下一个行为
```
Advance only a small amount at a time. It’s small enough that you can understand it, small enough that it’s difficult for AI to make up stories, and small enough that it can quickly locate when it fails.
### It limits the model to "play smoothly"
A common problem with AI is overzealousness.
You ask it to fix a boundary bug, and it extracts the helper; you ask it to add a test, and it changes the implementation; you ask it to refactor, and it changes its behavior.
TDD uses phases to separate these actions:
| Stages | What to do | What not to do |
| -------- | -------------------------------- | --------------------------------- |
| RED | Write a failing test | Write a production implementation |
| GREEN | Write the minimum implementation | Modify the test to get green |
| REFACTOR | Clean up structure | Introduce new behaviors |
This table is more useful than "Please be cautious." It lets the model know which stage it is in, and makes it easier for humans to detect boundary violations.
## 3. Reconstruction of red and green: three doors, not three slogans
“Red, Green, Refactor” could easily be a slogan. In actual use, it should look like three doors. Every time you pass through a door, you must leave evidence.
### The first door: RED, proving that the needs have not been met
The most important questions during the RED phase are:
> If this test fails, does it prove that we still lack a target behavior?
A BAD RED:
```py
assert True
```
Not a very good RED either:
```py
assert "hello" in format_title("Hello World")
```
It's too wide. Many faulty implementations also pass.
Better RED:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
This test is small, but it's clear. It specifies inputs, outputs, and behavior.
### Second door: GREEN, only let the current test pass
The GREEN phase is not about writing the final architecture.
It has only one mission: to pass the currently failing test with the least amount of code.
This statement sounds counter-intuitive. Many people worry about whether "minimum implementation" will be too ugly. Yes, sometimes it can be ugly. But its value lies in maintaining design pressure.
If the first test is:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
An acceptable GREEN might just be:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
You don't need to support Chinese, accents, continuous punctuation, emoji, SEO special cases right away. Those should be driven by later tests.
### The third door: REFACTOR, only changes the structure, not the behavior
The REFACTOR stage is the easiest to get confused by the AI.
It will interpret "tidy up the code" as "enhance it by the way." This won't work. The definition of refactoring is very narrow: the external behavior remains unchanged, but the internal structure becomes better.
Good refactoring looks like this:
* Change the variable name to a more accurate one
* Extract repeated expressions
* Remove conditional branches that are too deep
* Move function locations to make module responsibilities clearer
Bad refactoring looks like this:
* New input is now supported
* The error message has been changed easily
* Changed dependencies easily
* Conveniently changed the test assertion
The judgment criteria are simple:
> If this commit was just called `refactor:`, it should be the same green before and after testing, and user behavior should be the same.
## 4. The taste of good testing
TDD is not about more tests being better. AI is also very good at generating a bunch of tests that have little value.
What's more important is testing the taste.
### Good tests are like specs
A good test should read like a business specification:
```text
当用户没有权限时,保存按钮不可点击。
当标题为空时,表单显示错误信息。
当重复提交同一个请求时,只创建一条记录。
```
It is concerned with external behavior, not what is done internally.
Bad tests look like implementation notes:
```text
应该调用 validateInput 三次。
应该读取 state.user.flags。
应该触发 handleClick 内部函数。
```
Once the implementation details are tied up, refactoring will be painful. You just changed the internal structure, but the tests failed on a large scale. Rather than protecting the code, such tests freeze it.
### Good tests have boundaries
A test is best designed to answer only one question.
If a test also asserts:
* Correct format
* Permissions are correct
* The network request is correct
* The toast copy is correct
* The database status is correct
When it fails it's hard to know what the problem is.
AI is especially prone to writing such “big, comprehensive” tests because it wants to prove a lot of things at once. TDD is the opposite: an action, a failure, and an implementation.
### Good tests make it difficult for implementations to cheat
If the test only covers an input that is too specific, the AI may write a fake implementation that just matches.
For example:
```py
def slugify(text: str) -> str:
return "hello-world"
```
The first test will allow it to pass, but the second test will force out the real logic:
```py
def test_slugify_handles_another_title():
assert slugify("Test Driven Development") == "test-driven-development"
```
Therefore, TDD does not always write only one test, but only adds one behavioral pressure in each round. The pressure gradually increases and the design gradually grows out.
## 5. How will AI bypass TDD?
This part must be made clear because AI does not naturally respect testing.
Its optimization goal is simple: complete the task you just mentioned. If you say "let the test pass" it may do some actions that humans don't want.
### The first type: change the test to get green
The most typical:
* Change `assert slugify("Hello World") == "hello-world"` to the current output
* Remove failed assertions
* Add `skip` to the test
* Change strict assertions to loose assertions
This is not TDD, this is taking out the red light.
### The second type: write over-fitting implementation
For example, the test has only one input:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
The model might be written as:
```py
def slugify(text: str) -> str:
if text == "Hello World":
return "hello-world"
return text
```
You don't need to scold it at this time. You need to continue adding the next behavior so that the implementation cannot continue to be hard-coded.
### The third method: use mock to cover the real boundary
AI loves mocks. Mocks make tests easier to write and make many real problems disappear.
It’s not that you can’t mock, but you have to ask:
> What I am mocking now is a slow dependency, or is it a boundary that I really want to verify?
If you want to verify payment callback parsing but mock the parsing layer, the test will be meaningless.
## 6. When not to use TDD
TDD has value, but not everything is worth it.
### Unsuitable scene
* Pure visual fine-tuning
* one-time script
* Technology exploration demo
* Prototypes for which the requirements themselves have not been clearly thought out
* The test framework has not yet set up a warehouse
In these scenarios, pursue exploration speed first and don’t be held back by the process.
### Suitable scene
* bug fixes
* Permissions, billing, state machine
* Data transformation and boundary handling
* Core modules that will be maintained for a long time
* Code paths that AI will modify repeatedly
The judgment standard is not "whether this function is great or not", but:
> If it's wrong, are the costs obvious?
The cost is obvious, so it's worth writing the test first.
## 7. An executable mental method
If I were to give AI just one sentence, I wouldn’t say:
```text
请高质量实现这个功能。
```
I would say:
```text
先写一个失败测试,运行它,确认失败原因符合预期。不要写实现,直到我说 go。
```
This sentence is of higher quality because it is not asking the model to "behave well" but rather asking it to enter a process that can be checked.
A little more complete:
```text
每轮只处理一个行为。
RED:写一个失败测试并运行。
GREEN:写最小实现,不改测试。
REFACTOR:只在绿色状态下整理结构。
每轮报告测试文件、命令、失败原因、通过结果。
```
This is the core of TDD in the AI era.
Not superstitious about tests or processes, but about having evidence for every step.
## Closing
What AI programming needs most is not more code, but shorter feedback.
Here's the value of TDD: it turns "I thought it should be right" into "Here was a failure, and then it turned green." The change is small, but real enough.
If you only remember one sentence, remember this:
> Don’t let AI deliver code directly. First let it deliver a red light, and then let it turn the red light green.
The next [Practical Guide](/en/docs/notes/tdd-with-ai/practice) will turn this rhythm into a workflow that can be directly copied.
## Recommended resources
# TDD in the AI Era: Codex Practical Manual
## Give me a map first
[Concept](/en/docs/notes/tdd-with-ai/concept) It’s about: In AI programming, the value of TDD is not the ritual of “write the test first”, but rather create a verifiable red light first, and then let the implementation turn it green.
This article talks about how to implement Codex.
Don’t ask “what documents do I need to match” right away. A better question is:
> How do I make Codex act on the same TDD workflow every time?
This workflow can be split into four levels:
| Hierarchy | What are you doing | Where | When is it appropriate |
| --------- | ------------------------- | ----------------------------------- | ------------------------------------------------------------------------------ |
| L1 | Write project disciplines | `AGENTS.md` | All projects should have |
| L2 | Solidification process | `.agents/skills/tdd-codex/SKILL.md` | Repeatedly use TDD to meet requirements |
| L3 | Isolation stage | `.codex/agents/*.toml` | Complex tasks, fear of mutual contamination between testing and implementation |
| L4 | Automatic reminder | `.codex/hooks.json` | Important warehouse, afraid of AI secretly modifying the test |
The smallest available versions are L1 + L2.
The complete line of defense is L1 + L2 + L3 + L4.
## 1. First define what "completion" is
If there are no completion standards, Codex can easily treat "written code" as "done".
In a TDD scenario, completion criteria should be more specific.
### Evidence must be delivered in every round
Have Codex report these six items every round:
```text
Behavior: 这一轮实现哪个行为
Test: 测试文件和测试名
Command: 跑了什么命令
RED: 失败原因是否符合预期
GREEN: 通过结果是什么
REFACTOR: 是否重构,为什么
```
This is much more useful than a "done."
It lets you know that the model has really gone through the red and green cycles, instead of writing the implementation first and then adding a test that looks reasonable.
### Only one behavior is processed in a round
This is critical.
Don't let Codex generate the entire test matrix at once. That would become a "horizontal laying test":
```text
RED: test1, test2, test3, test4, test5
GREEN: 一次写一个大实现
```
What you want is to slice it lengthwise:
```text
RED test1 -> GREEN impl1 -> REFACTOR
RED test2 -> GREEN impl2 -> REFACTOR
RED test3 -> GREEN impl3 -> REFACTOR
```
The first round of implementation will change your understanding of the problem. Don't write all your tests in one go.
## 2. L1: Write discipline into AGENTS.md
`AGENTS.md` is the description file that Codex will read when entering the project.
OpenAI official documentation states that Codex will first read the global description, and then read it from the project root directory all the way to the current directory. Each layer reads `AGENTS.override.md` first, otherwise reads `AGENTS.md`. Descriptions closer to the current directory appear later and therefore have higher priority. The default merge limit is `32 KiB`, so a long tutorial cannot be written here.
### It should be like project traffic rules
`AGENTS.md` is not responsible for teaching Codex what TDD is. It is only responsible for writing clearly: what behavior is not allowed in this project.
You can put this paragraph directly:
```markdown
# TDD Rules
- For new behavior and bug fixes, use red/green TDD.
- RED: write exactly one failing behavior test first.
- Run the smallest relevant test command and confirm the failure is expected.
- Do not edit production implementation during RED.
- GREEN: write the minimum production code required to pass the current failing test.
- Never modify, delete, skip, or weaken tests to make implementation pass.
- REFACTOR only after tests are green.
- Keep structural changes and behavior changes separate.
- Report Behavior, Test, Command, RED, GREEN, and REFACTOR for each cycle.
```
Additional project commands:
```markdown
# Verification
- Use `pytest` or the smallest relevant pytest command for Python behavior tests.
- Use `npm run types:check` only when this blog site's MDX or TypeScript changes.
- Use the smallest targeted command during RED/GREEN loops.
- If a command is slow, explain what targeted command was used first and what full command remains.
```
### It should not be written as an encyclopedia
A bad `AGENTS.md` would be filled with:
* History of TDD
* All testing philosophies
* A bunch of framework tutorials
* Complex prompt templates
* Complete specifications in different languages
These things dilute the rules that really matter.
My suggestion is: `AGENTS.md` Only put resident disciplines. Use skills for long processes.
## 3. L2: Make the process into Codex Skill
`AGENTS.md` solves "default discipline" and skill solves "complete process".
When you often say "do it by TDD" to Codex, you should upgrade this sentence to a project-level skill.
### Directory structure
Put it here:
```text
.agents/
skills/
tdd-codex/
SKILL.md
```
Codex will scan from the current directory all the way up to `.agents/skills`. Skills in the root directory of the warehouse are suitable for workflows used by the team.
### Minimum available SKILL.md
```markdown
---
name: tdd-codex
description: Implementing or fixing maintainable code with Codex using strict red-green-refactor TDD. Use for new behavior, bug reproduction, behavior tests, or safe AI coding.
---
# TDD Codex Workflow
Use one behavior slice per cycle.
## Phase 0: Scope
Identify one observable behavior.
Name the public API, user flow, or integration boundary under test.
Do not edit production code.
## Phase 1: RED
Write exactly one failing behavior test.
Prefer public behavior over implementation details.
Run the smallest relevant test command.
Confirm the failure is expected.
Stop and report:
- Behavior
- Test file
- Command
- Failure reason
## Phase 2: GREEN
Write the minimum production code to pass the current failing test.
Never modify, delete, skip, or weaken tests to pass.
Do not add speculative features or abstractions.
Run the same test command.
Report the passing result.
## Phase 3: REFACTOR
Only refactor after tests are green.
If the code is already simple, skip.
If refactoring, make one structural change at a time.
Run tests after each refactor.
Do not change behavior.
## Cycle Report
Return:
- Behavior:
- Test:
- Command:
- RED:
- GREEN:
- REFACTOR:
- Next slice:
```
### Calling method
From now on you can say:
```text
用 tdd-codex skill 做这个需求。
每轮只处理一个行为。
先 RED,确认失败后停下来,不要直接写实现。
```
Or shorter:
```text
用 tdd-codex。先写红灯,等我说 go。
```
The point is not how beautiful the prompt is, but that it brings Codex back on the same track every time.
## 4. L3: Use Subagents to isolate red and green reconstruction
Not all tasks require subagents.
However, when tasks are complex, tests are easily contaminated by implementation, and refactoring is easy to get out of control, it will be more stable to split RED, GREEN, and REFACTOR into different agents.
### When is it worth dismantling?
Suitable for disassembly:
* Permissions, billing, state machine
-Multi-module functionality
* The bug is very hidden, so you need to write a recurrence test first
* The model is always changed to test to get green
* You want someone who is only responsible for reviewing test quality
Not suitable for disassembly:
* Gadget functions
* Copywriting changes
* Pure visual fine-tuning
* one-time script
The cost of tearing down an agent is real. Use only when the benefits of isolation outweigh the costs of communication.
### RED agent
```toml
# .codex/agents/tdd-test-writer.toml
name = "tdd_test_writer"
description = "RED phase agent. Writes one failing behavior test and stops before implementation."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the RED phase.
Write exactly one behavior-focused test for the requested slice.
Prefer public APIs and user-visible behavior over implementation details.
Run the smallest relevant test command.
Confirm the test fails for the expected reason.
Do not edit production implementation.
Do not add multiple tests at once.
Return Behavior, Test, Command, and RED failure reason.
"""
```
### GREEN agent
```toml
# .codex/agents/tdd-implementer.toml
name = "tdd_implementer"
description = "GREEN phase agent. Implements the minimum production code to pass the current failing test."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the GREEN phase.
Read the failing test and relevant production code.
Write the minimum implementation required to pass the current test.
Never modify, delete, skip, or weaken tests to make them pass.
Do not add speculative features, helpers, configuration, or abstractions.
Run the relevant tests and return the command plus passing output.
"""
```
### REFACTOR agent
```toml
# .codex/agents/tdd-refactorer.toml
name = "tdd_refactorer"
description = "REFACTOR phase agent. Improves structure only after tests are green."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the REFACTOR phase.
Start by running the relevant tests to confirm the code is green.
Look for duplication, unclear names, needless branching, or misplaced responsibility.
Skip refactoring when the code is already simple.
If you refactor, make one structural change at a time.
Run tests after each refactor.
Never change behavior in this phase.
"""
```
### How to command the main session
```text
按三阶段 TDD 做这个 slice:
1. tdd_test_writer 只写一个失败测试,并确认 RED。
2. 等我确认后,tdd_implementer 写最小实现,并确认 GREEN。
3. tdd_refactorer 判断是否需要结构重构。
不要批量铺测试。
不要在 GREEN 阶段修改测试。
```
The point here is to isolate context. People who write tests should try not to be affected by implementation details; people who write implementations cannot test manually; people who refactor cannot introduce new behaviors.
## 5. L4: Use Hooks to focus on test diff
If you rely solely on rules, the model may still go out of bounds.
The most common cross-border is: the test is red, and the model changes the test in order to turn it green.
The value of hooks is not to be "absolutely safe", but to expose this action immediately.
### Enable Codex hooks
First open the feature flag in the configuration:
```toml
# ~/.codex/config.toml 或 /.codex/config.toml
[features]
codex_hooks = true
```
Codex looks for hooks next to the active configuration layer. Common locations:
* `~/.codex/hooks.json`
* `~/.codex/config.toml`
* `/.codex/hooks.json`
* `/.codex/config.toml`
At the project level, it is recommended to use `/.codex/hooks.json` first, because it can follow the warehouse.
### hooks.json
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/watch-test-edits.sh\"",
"timeout": 10,
"statusMessage": "Checking test file edits"
}
]
},
{
"matcher": "Bash|apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/run-fast-check.sh\"",
"timeout": 120,
"statusMessage": "Running fast checks"
}
]
}
]
}
}
```
### Check whether the test file has been modified
```bash
# .codex/hooks/watch-test-edits.sh
#!/usr/bin/env bash
set -euo pipefail
changed_tests="$(
git diff --name-only |
grep -E '(^|/)(__tests__|tests?)/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.go$|test_.*\.py$' || true
)"
if [ -n "$changed_tests" ]; then
cat < "hello-world"
- trim leading/trailing spaces
- collapse repeated spaces into one hyphen
- remove punctuation
- normalize "Café" -> "cafe"
- empty input returns empty string
```
### Step 1: First red light
```text
读取 SPEC.md。
只实现第一条行为:"Hello World" -> "hello-world"。
先 RED:只写一个失败测试,运行它,确认失败。
不要写生产实现。
```
Ideal output:
```text
Behavior: basic title becomes lowercase hyphenated slug
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
RED: failed because slugify is not defined
```
Only then can you continue.
### Step 2: Minimal green
```text
go
```
Codex writes the minimal implementation:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
Then report:
```text
GREEN: pytest tests/test_slugify.py -q passed
REFACTOR: skipped, implementation is still simple
Next slice: trim leading/trailing spaces
```
### Step 3: Second red light
```text
继续下一条:去掉首尾空格。
先 RED,只写一个测试。
```
Test:
```py
def test_slugify_trims_spaces():
assert slugify(" Hello World ") == "hello-world"
```
If the current implementation outputs `-hello-world-`, the red light is true.
Then GREEN:
```py
def slugify(text: str) -> str:
return text.strip().lower().replace(" ", "-")
```
### Step 4: Don’t rush to abstraction
At this point many AIs will want to draw a `normalizeInput`, `removePunctuation`, `toAscii`.
Don't rush yet.
The design of TDD should be pushed out by test pressure, not by imagination. Wait until you add unicode, punctuation, and empty strings, and structural pressure really arises, then refactor.
## 8. Daily usage quick check
### New features
```text
用 TDD 实现这个需求。
每轮只处理一个行为。
先写一个失败测试并运行确认 RED。
不要写生产实现,直到我说 go。
```
### Fix bug
```text
先写一个能复现这个 bug 的失败测试。
确认它因为这个 bug 失败后,再写最小修复。
不要改测试来适配当前实现。
```
### Complex functions
```text
先不要写代码。
请给出 TDD 分解计划:
- 外圈集成测试是什么
- 内圈每个行为 slice 是什么
- 每轮用什么命令验证
- 哪些地方不能 mock
等我确认后再开始 RED。
```
### Review
```text
Review 这次改动,重点看:
- 是否先有失败测试
- 测试是否测行为而不是实现
- 是否存在为了通过而弱化测试
- 结构改动和行为改动是否混在一起
- 是否缺少外圈集成测试
```
## 9. Final checklist
Every time you ask Codex to do TDD, use this table to check at the end.
| Question | Eligibility Criteria |
| ------------------------------------ | ------------------------------------------------------------- |
| Is it really popular first | Are there failed commands and reasons for failure |
| Is the red color correct | The failure reason corresponds to the lack of target behavior |
| Do only one behavior per round | No batch testing |
| GREEN Has the test been changed? | No test has been changed to get green |
| Does testing test behavior | Does not rely on internal implementation details |
| Have mixed behaviors been refactored | Separate structural changes and behavioral changes |
| Is there complete verification | Target tests and necessary full inspections have been run |
If this table cannot pass, don't rush to merge.
## Recommended resources
# AIが「勉強しなくても試験に合格できる」時代、大学に残るものとは?
Anthropicは最近、プリンストン大学、バークレー、ロンドン・スクール・オブ・エコノミクス、アリゾナ州立大学の4人の学生を招いてインタビューを行い、キャンパスにおけるAIの実態について語り合いました。約40分にわたるこの対話に、宣伝文句は一切なく、リアルな困惑、不安、そして思考が詰まっています。
このインタビューが明らかにしたのは、より深い問題です:**AIは単に学び方を変えているのではなく、教育システム全体の根底にある論理を解体しつつあるのです**。
***
## 避けられない現実
インタビューの冒頭、司会者が率直な質問を投げかけました:今のキャンパスにおけるAIへの雰囲気はどうですか?
答えは:**学生の90%がAIを使っている**。たまに使うのではなく、日常のワークフローの一部として——講義ノートのまとめ、問題集への回答、課題へのフィードバック取得、ビジネスケース分析、市場調査、財務研究の完成。一部の学生はクイズにすら使っており、その理由は現実的です:大学院生として複数の仕事を掛け持ちしているとき、常に時間があるわけではないからです。
しかしさらに興味深いのは、ほぼ全員が使っているにもかかわらず、誰もルールを知らないという点です。AIを明文で禁止する授業もあれば、積極的に奨励する授業もあり、ほとんどは曖昧なグレーゾーンに存在しています。学生はどこが境界線なのかわからず、教授も管理の仕方がわかっていません。この状態は「グレーゾーン」と呼ばれています——使いたいけど違反が怖い、使わなければ取り残される気がする。
このグレーゾーンが最も危険なのは、学生がルールを破るかもしれないからではなく、**本当に価値ある議論が生まれるのを妨げているから**です。学生はAI活用のベストプラクティスを公に共有できず、教授は責任ある道具の使い方を指導できず、学術コミュニティ全体が、表向きは禁止・裏では使用という気まずい状態に陥っています。
ルールが有効に執行できなくなると、それは「うまく偽装できる人」と「できない人」を選別するツールになってしまいます。明文で禁止されているが裏では広く使われている状態は、**学生がAIを使うのを止めるのではなく、より良い使い方について公に議論するのを妨げるだけ**です。
***
## AIは鏡である
インタビューの中に鋭い指摘がありました:人工知能、特に学生がそれをどう使うかは、彼らの動機を非常によく示します。
この言葉の背後にある洞察:**AIは鏡になり、あなたが大学に通う本当の目的を映し出します**。
インタビューは大学の目標を3つに整理しています:第一は、専門知識を深く学び、ある分野への深い理解を得ること。第二は、就職準備——良い仕事を見つけ、職業的なネットワークを構築すること。第三は、人脈を広げ、社会的生活を楽しみ、大学文化を体験すること。学生ごとにこの3つの目標のウェイトは異なり、AIの使い方はそのウェイトをくっきりと露わにします。
もし「試験に合格する」と「学位を取る」ことしか気にしないなら、AIのアウトプットをそのまま課題として提出するでしょう。これは道徳的な批判ではなく、現実です——技術が最小コストで目標を達成させてくれるなら、なぜわざわざ遠回りするのでしょうか?学位を取って就職することが最初からの目標なら、AIで課題を済ませることは完全に合理的な選択です。
しかし本当に学びたい、ある分野を深く理解したいと思うなら、AIを対話相手として使うでしょう。質問し、概念を説明させ、自分の言葉で言い直す。初稿のコードを書かせ、自分でリファクタリングして最適化する。すべての段階で、自分が何が起きているかを本当に理解していることを確かめる。
この分化は学生間だけでなく、専攻間にも存在します。人文系の学生はAIを使わない選択をしがちです。彼らの学習には精読が必要だからです——原文を丁寧に読み、言語の細部を味わい、著者の意図を理解する。AIはこのプロセスを損ないます。AIが提供するのは要約と言い換えであり、原文との直接の体験ではないからです。工学・ビジネス系の学生はAIを大量に活用します。AIが技術的なハードルを下げ、コンピュータサイエンスのバックグラウンドがなくてもコードを書いたり、ウェブサイトを構築したり、データを分析したりできるようになるからです。
**この二極化は本質的に、「何を学ぶ価値があるか」に対する異なる理解です**。人文系の学生にとっては、シェイクスピアの原文を読む体験そのものが学習であり、工学系の学生にとっては、問題を解決できるかどうかが重要で、すべてのコードを自分で書いたかどうかではありません。AIはこの差異をより顕著にしています。
***
## ツールか松葉杖か?シンプルな判断基準
インタビューの重要な問い:AIがツールなのか松葉杖なのか、どう区別するか?
学生たちが出した答えは驚くほど一致していました:**説明できるかどうか**です。
自分が作ったものを説明できない、AIがどんな役割を果たしたか言えないなら、それは松葉杖です。小学5年生に説明するように話せる、低レベルな説明も高レベルな説明もできるなら、それはツールです。
この基準はシンプルに見えますが、学習の本質に触れています。フェインマン学習法の核心的なロジックとは、もし簡単な言葉でコンセプトを説明できないなら、まだ本当に理解していないということです。AI時代の学習も同じ道理です——AIが自分のためにやってくれたことを説明できないなら、「思考を外部委託」しているだけで、「思考を拡張」しているのではありません。
インタビューで紹介された面白い例があります:ある学生が、講義スライドを入力するとAIが各スライドの隣に教授のような注釈を生成するツールを開発しました。その学生は言いました:「このツールがうまく機能するのは、私がすでに自分が知りたいことを理解するようにプロンプトを与えているからです——スライド上のある事柄の定義。スライドは時として抽象的で背景情報が足りず、隣にコンテキストを追加する必要があります。」
ポイントは「すでに自分が知りたいことを理解するようにプロンプトを与えている」という部分です。この学生は自分の知識のギャップがどこにあるかを把握し、どんな助けが必要かを知り、そしてAIを能動的に導いてその助けを引き出しています。それがツールです。もしスライドをAIに投げて「この授業を要約して」と言い、それを丸暗記するだけなら、それは松葉杖です。
**違いは主体性と理解にあります**。ツールを使う人は自分が何をしているかを知り、プロセス全体をコントロールしています。松葉杖を使う人はコントロールを技術に渡し、受動的な受け取り手になってしまいます。
***
## 学校の遅れは反応の遅さではなく、本質的に対応できないこと
インタビューでは学校側のいくつかの取り組みも紹介されました。LSEの必修科目では、学生にClaudeの使い方を指導し始めました——対話し、異なる役割を与え、そして学生にAIとのやり取りを示す対話記録の提出を義務付けています。アリゾナ州立大学のキャリアセンターはプロンプトライブラリを構築し、さまざまなシナリオ向けのプロンプトテンプレートを提供しています。これらは良い取り組みであり、核心となる考え方は、AIを禁止するのではなく、責任ある使い方を教えるということです。
しかしこれらは少数派です。ほとんどの学校はまだ「学生がAIを使うことを許可すべきか」を議論しています。使っても良いが課題に使用方法を明記するよう求める教授もいれば、完全禁止の授業もあり、言及すらせずデフォルトで学生は使わないと仮定しているものもあります。統合フレームワークも統一基準もなく、システム全体が混乱した状態にあります。
より深い問題は:**この問題は本質的に規則や規制によって解決できない**ということです。
伝統的な教育の管理ロジックは、学校がルールを設定し、学生がルールに従い、違反者が罰せられるというものです。このロジックが成立する前提は「違反行為が検出できること」です。しかしAIはこの前提を崩しました。
課題提出時にAIを使うことを禁止できても、学生が思考の過程でAIを使ったかどうかを監視することはできません。AI検出ツールを使うことはできても、その精度は罰則の根拠として使えるレベルにはほど遠く——誤検知率が高すぎるうえ、学生はすぐに検出を回避する方法を学んでしまいます。さらに重要なのは、**根本的に「AI支援のもとで完成した高品質な課題」と「独立して完成した高品質な課題」を区別することができない**という点です。なぜなら、良いAIの使い方は本来シームレスであるべきだからです。
インタビューの中で率直に語られていました:根本的に、いかなる規則も学生がAIを使う方法を変えることはできない。責任は学生自身にある。これは責任の押しつけではなく、現実です。
技術が「勉強しなくても試験に合格できる」ことを可能にしたとき、学校が直面しているのは管理の問題ではなく、実存的な問題です:**もし試験が学習を証明できないなら、学校が存在する意味は何でしょうか?**
***
## 教育システムの根底にある論理が崩れた
この問いは教育システムの根本的な矛盾に触れています。
伝統的な教育はいくつかの核心的な前提の上に成り立っています:第一に、知識は希少であり、専門の機関(学校)と専門家(教授)によって伝授される必要がある。第二に、学習成果は試験によって測定できる。第三に、学位はある分野の知識を習得したことを証明し、関連する仕事に就く資格を持つことを示す。
AIはこれらの前提を一つ一つ崩しています。
**知識はもはや希少ではありません**。YouTubeには無料のスタンフォードの講座があり、Claudeはいつでも質問に答えてくれ、GitHubには学べるオープンソースプロジェクトが無数にあります。知識に触れるために学校に行く必要はなく、有料サブスクリプションなしでも基礎的なAI家庭教師を利用できます。
**試験では学習を測定できません**。AIがほとんどの試験問題をこなせるようになると、試験は「理解度を測るツール」から「AIを使えるかどうかを測るツール」に変わります。試験が完全に無意味というわけではありませんが、「本当に学んだ人」と「ツールを使いこなせる人」を正確に区別することはもはやできません。
**学位の価値は下落しています**。雇用主が学位で候補者が本当に関連知識を習得しているかを保証できないと気づくと、実際の能力の証明——ポートフォリオ、プロジェクト経験、インターンのパフォーマンス——をより重視するようになります。学位は「能力の証明」から「基本的な足切り基準」へと格下げされます。
**これらの前提が崩れたとき、教育システムの価値命題を再定義する必要があります**。
インタビューはひとつの答えを示しています:大学の価値は「**知識を伝授する**」から「**環境を提供する**」へと転換します。失敗でき、探索でき、他者とアイデアをぶつけられる環境。室友と週末を使って「卒業前にやりたいことリスト」を作り、ハッカソンで「おそらくバカげた」アイデアをテストし、教授と議論し、同級生と話し合い、キャリアのリスクなしに失敗から学べる。
インタビューで紹介された学生プロジェクトはこれをよく示しています——「自動履修登録アラート」、「空き教室検索ツール」、「卒業前ウィッシュリストランキング」。技術的にはどれも複雑ではなく、多くの作者はコンピュータサイエンスのバックグラウンドすら持っていません。しかしそれらは真の人間感情から生まれています:機会を逃すことへの恐れ、便利さへの追求、大学生活への愛着。
**技術的なハードルが下がったとき、重要なのはもはや「コードが書けるかどうか」ではなく、「どんな問題を解決したいか」です**。大学が提供するのは、これらの問いを自由に探索できる空間、アイデアを現実にし失敗から学べる環境です。
AIはあなたの課題を完成させることができますが、この「失敗でき、探索できる」時間をあなたに代わって過ごすことはできません。
***
## 就職市場のパラドックス
インタビューの後半では就職について語られ、別のパラドックスが明らかになりました。
学生はAIで履歴書を書き、企業はAIで履歴書を選考します。採用プロセス全体が画面に向かって話すことになっています——まずAIに向かって志望動機書を書き、次にビデオで質問に答え、最後にAIが生成した不採用通知を受け取る。履歴書を提出してから不採用通知を受け取るまで、15分もかからないことがあります。非常に効率的ですが、人間的なものはほとんどありません。
これが「AI対AI」の就職市場を生み出しています。学生はAIに「良い」履歴書の書き方を学習させ、企業はAIに「良い」候補者の選び方を学習させます。このプロセスにおける本物の人間の役割はますます小さくなっています。画面に向かって話しても化学反応は生まれず、数値化しにくいが非常に重要な資質——ユーモアセンス、応変能力、チームワークの微妙さ——を示す方法がありません。
しかしパラドックスの裏面もあります:AI習熟度そのものが新たな競争力になっています。四大コンサルティングファームはかつてジェネラリスト型のMBAを採用していましたが、今ではAI能力を持つMBAを特に探しています。異なる業界にAIを応用する方法を知っているなら、あなたが最初の候補者になります。
**パラドックスはここにあります:AIは就職市場をより冷たくする一方で、AI能力をより重視させます。逃げることはできない——効果的な使い方を学ぶしかありません**。
これはあの核心的な問いに戻ります:「効果的な使い方」とは何でしょうか?ChatGPTでメールを書けることではなく、どの問題がAIに向いているかを見分け、必要な結果を引き出すプロンプトを設計し、AIの出力品質を評価して必要な修正を加えられることです。
この能力はAIを禁止することで育まれるのではなく、大量の実践と試行錯誤によって育まれます。これが、AIを積極的に受け入れ、Claude Builder Clubを設立し、ハッカソンを開催する学校が、より価値のある教育を提供している理由です——比較的安全な環境で、学生がAIと協働する方法を学べるようにしているのです。
***
## 責任の移転
このインタビューの最も核心的な洞察はおそらくこれです:技術があなたを「勉強しなくても試験に合格できる」ようにしたとき、学習そのものの意味は、各自が自分で答えなければならない問いになります。
これは責任の移転です。**学校から学生へ、ルールから自覚へ、外部動機から内部動機へ**。
伝統的な教育の動機は外部的なものです:学位を得るために試験に合格する必要があり、就職のために学位が必要です。この外部インセンティブシステムが学生を学習に駆り立てていました。しかしAIが勉強なしに試験に合格させてくれるようになると、このインセンティブシステムは機能しなくなります。
残るのは内部動機だけです:本当に学びたいですか?この分野に本当に興味がありますか?深く理解したいのか、それともただ学位が欲しいだけですか?
インタビューの中に興味深い細部があります。ある学生が、大学院に通うことの悪い面を挙げました——複数の仕事を掛け持ちし、時間がないため、時にはAIを使ってクイズを素早くこなすと。しかし彼は続けました:大学院は批判的思考を広げる時期のはずで、より果断な側面を示す時期だと。**彼は矛盾に気づいていたが、効率を選んだのです**。
これは道徳的な批判ではありません。現実のプレッシャーの下では、効率は理想よりも重要なことが多い。しかしこの選択は事実を明らかにしています:**外部のプレッシャー(クイズを完成させること)と内部動機(深く学ぶこと)が衝突するとき、多くの人は前者を選びます**。
AIはこの対立をより鋭くします。AIが「試験をやり過ごすこと」を極めて簡単にするからです。AIがなかった時代は、試験をやり過ごしたいだけだとしても、合格するには何かを学ばなければなりませんでした。AIはこの中間ステップを取り除きます——まったく学ばなくても合格できるのです。
**これは誰もがあの問いに向き合うことを迫ります:あなたはなぜ大学に通っているのですか?**
答えが「学位を取って就職するため」なら、AIで課題をこなすことは完全に合理的です。答えが「本当にこの分野を学びたい」なら、AIがもたらす近道の誘惑に能動的に抵抗する必要があります。
学校があなたに代わってこの選択はできません。ルールがあなたに内部動機を強制することもできません。**これはあなた自身の責任です**。
***
## 技術はあなたが準備できるのを待ってくれない
インタビューの最後を通じて、ある態度が一貫していました:「なんとかなるさ。」
学校のルールが追いつかない?まず使い始めて、それから何が有効かを学校に伝えればいい。AIが不正に使われるかもしれない?少しずつ責任ある使い方を学んでいこう。就職市場が変わった?新しいゲームのルールに適応しよう。
これは盲目的な楽観主義ではなく、リアリズムです。技術はすでにここにあり、あなたが準備できるのを待って世界を変え始めるわけではありません。抵抗を選ぶこともできるし、適応を選ぶこともできますが、時間を止めることを選ぶことはできません。
**この世代の学生とAIの関係は、恐怖でもなく、盲目的な受け入れでもなく、混乱の中で手探りし、試行錯誤の中で学ぶことです**。
Claude Builder Clubで作っているプロジェクト——技術的に複雑ではないが、リアルな問題を解決している。ハッカソンで試みているアイデア——バカげているかもしれないが、少なくとも挑戦している。教室での困惑——ルールは不明確だが、少なくとも考えている。
この「**やりながら学ぶ**」という姿勢は、いかなる規則よりも彼らがAI時代に適応するのを助けるかもしれません。
ケヴィン・ケリーは『テクニウム——テクノロジーはどこへ向かうのか?』の中で「technium」という概念を提唱しました:技術全体として、自らの意志を持ち、ますます強くなりたがっているように見えるというものです。インタビューの中にこれと呼応する観察があります:過去2年間、AIが発展し続けるために必要なものが何であれ、それを手に入れてきた。原子力に対する態度の転換、宇宙データセンターの議論——AIの減速につながりかねないボトルネックがあるたびに、障壁が取り除かれています。
しかしより正確な言い方はこうかもしれません:**学生が必要とするものを、自ら創造する**。
AIは単なるツールです。未来を決めるのは、この世代の学生がこのツールをどう使うかという選択です——思考を回避するために使うか、思考を拡張するために使うか。試験をこなすために使うか、世界を探索するために使うか。松葉杖として使うか、ツールとして使うか。
この選択は学校には管理できない——自分たちで決めるしかありません。
そしてこのインタビューを見る限り、少なくとも一部の学生は、この問いを真剣に考えています。それで十分かもしれません。
# Agentのように考える:Claude Codeチームのツール設計哲学
ThariqはAnthropicのエンジニアであり、Claude Codeの中核的な構築者の一人です。2月末にXで長文を投稿し、Claude Codeを構築する過程でのagentツール設計に関する5つの実際のケースを共有しました——理論的なフレームワークではなく、現場で身をもって経験した教訓です。この記事は351万回の閲覧、209件の返信、9,691件のいいねを獲得し、質の高い議論を数多く生み出しました。
記事全体の核心となる問いを引き出すために、Thariqは非常に良いアナロジーを使っています。難しい数学の問題に取り組むとき、どんなツールが必要でしょうか?
**紙とペン**は最低限の構成——でも手計算に縛られます。**電卓**はより良い——でも高度な機能の使い方を知っている必要があります。**コンピューター**が最強——でもコードが書けなければなりません。
ツールの選択は使用者の能力に依存します。プログラミングができない人にコンピューターを渡すくらいなら、電卓を渡した方がましです。プログラマーに電卓を渡すと、むしろその力を制限してしまいます。
agentにとっても同じです。問いは「どのツールが最も強力か」ではなく、「どのツールがモデルの現在の能力に最もマッチしているか」です。この記事が共有するのは、Claude Codeチームがそのマッチングポイントを探す過程で踏んできた落とし穴です。
以下は、原文とコミュニティの議論から私が抽出した4つの発展的なテーマです。
***
## 少ない方が多い——ツール数の逆説
直感的には、ツールが多ければ多いほどagentの能力は高まると思いがちです。しかしClaude Codeチームの経験はまったく逆でした。
Claude Codeが現在持つツールはわずか約20個であり、チームはこれらすべてのツールが本当に必要かどうかを常に見直しています。新しいツールを追加するためのハードルは高く、それはモデルが考慮すべき選択肢を一つ増やすことを意味するからです——ツールが増えるほど、モデルが意思決定する際の「認知的負荷」が増えていきます。
この発見はコミュニティで強い共感を呼びました。leon.Mは自身の経験を共有しています:
AppleのCodeAct研究がこの見解に定量的な裏付けを提供しています:**単一のコード実行プリミティブ(code execution primitive)は、複雑なタスクにおいて大規模な専用ツールセットより最大20%高いパフォーマンスを発揮します**。少ない方が、確かに多いのです。
AskUserQuestionツールの3度の反復は、この原則を最もよく示すケースです。Claude Codeチームは、Claudeがユーザーに質問する能力を向上させたいと考えていました——Claudeは純粋なテキストで質問できますが、それらの質問に答えることが面倒に感じられました。摩擦を減らすにはどうすれば良いでしょうか?
**1回目の試み**:ExitPlanToolにパラメーターを追加し、計画を出力すると同時に一連の質問を付け加えるようにしました。結果——Claudeが混乱しました。計画と計画に関する質問を同時に出力するように求め、もしユーザーの回答が計画と矛盾したらどうなるのか?
**2回目の試み**:出力指示を変更し、Claudeが特定のmarkdown形式で質問し、フロントエンドがフォーマットを解析するようにしました。結果——不安定でした。Claudeは余分な文章を追加したり、選択肢を省略したり、まったく異なる形式を使ったりしました。
**3回目の試み**:独立したAskUserQuestionツールを作成しました。Claudeはいつでもそれを呼び出せ、呼び出すとポップアップウィンドウに質問が表示され、ユーザーが回答するまでagentループがブロックされます。**成功しました。**
Thariqは原文にとても興味深い一文を書いています:
> Claudeはこのツールを喜んで呼び出しているようで、出力結果も良好だと分かりました。どんなに上手く設計されたツールでも、Claudeが呼び出し方を理解していなければ機能しません。
phuongが鋭く問いただしました:「『Claudeがこのツールを喜んで呼び出しているようだ』というのは、ここで最も興味深くかつ最も説明が不足している一文です。モデルのツールに対する『好感度』をどう検出しているのか——transcriptを読むことで?それとも内部の呼び出し頻度の指標で?このヒューリスティックをより具体的にできれば、すべての人がagentツールを設計する方法が変わるでしょう。」
この質問には答えが得られませんでしたが、重要な直感を指し示しています:**ツール設計の成功基準は「人間が合理的と感じる」ことではなく、「モデルが使い方を理解し、使いたいと思う」こと**です。
Emekaは企業向けツール開発の観点からとても率直な教訓を補足しました:「以前agentのツールを構築したとき、あらゆる入力と出力をコントロールしようとしましたが、モデルは……それを回避してしまいました。エンジニアリングの手間を省いて、モデルが曖昧さを処理できると信頼しましょう。」
***
## ツールは陳腐化する——能力向上後の制約効果
「少ない方が多い」が空間的な認識に関するものだとすれば、「ツールは陳腐化する」は時間的な教訓です。
**かつてモデルを助けたツールが、モデルの進化とともに制約になる場合があります。**
Claude Codeの初期リリース時、チームはモデルが方向性を保つためにToDoリストが必要だと気付きました——開始時にタスクを書き込み、作業が完了するにつれて一つずつチェックを入れていくものです。そのためにClaudeにTodoWriteツールを提供しました。しかしそれでも、Claudeは自分が何をすべきか忘れることが多くありました。
チームの対応策は、5ターンごとにシステムリマインダーを挿入し、Claudeに目標を思い出させることでした。
しかしモデルの進化とともに、問題は逆転しました:モデルはToDoリストのリマインダーが不要になっただけでなく、それらのリマインダーを制約と感じるようになりました。**ToDoリストを繰り返し思い出させることで、Claudeはリストをフレキシブルにではなくきっちり従わなければならないと感じるようになりました。** 同時に、Opus 4.5はサブagentの活用において大幅に改善しましたが、サブagent間で共有されたToDoリストをどう協調させるのか?
そこでチームはTodoWriteをTask Toolに置き換えました。両者の違いは本質的です:Todosの役割はモデルを軌道に乗せることで——ボスが従業員のタスクリストを監視するようなもの;Tasksはagent間のコミュニケーションを支援することに重きを置き——チームのコラボレーションかんばんのようなものです。Tasksは依存関係をサポートし、サブagent間で更新を共有でき、モデルはそれらを修正・削除することもできます。
**TodoWrite → 5ターンごとのリマインダー → Task Toolは、3度の再設計です。** 以前の設計が「間違っていた」からではなく、モデルが成長したからです。
modiのコメントがこの現象を的確に要約しています:
これはまた興味深い論争も引き起こしました。Danny Cossonは、AskUserQuestionは「実は設計が悪い」と考えています——それはClaude Codeの非常にエレガントなテキスト入力/テキスト出力モデルを特定のインタラクションパターンに無理やり当てはめており、ほとんどメリットがないと主張しています。
しかし@Toongの反論は説得力があります:AskUserQuestionの価値は2点にあります——**主導権**(モデルに質問する権限があることを明確に伝える)と**状態管理**(標準的なテキスト出力と「ユーザーの介入を待つ」ブロッキング状態を明確に区別する)。
同じツールに対して、まったく異なる2つの評価があります。これはまさにThariqが結末で述べたことを証明しています:**ツール設計はサイエンスではなくアートです**。使用するモデル、agentの目的、そしてそれが置かれた環境によって異なります。
***
## Agentに自分で答えを見つけさせる——「与える」から「自律的な探索」へ
これが全文の中で私が最も実用的だと思う部分です。
Claude Codeは当初、Claudeのコンテキストを見つけるためにRAGベクターデータベースを使用していました。RAGは強力で高速ですが、2つの問題がありました:一つはインデックス作成と設定が必要で、異なる環境では脆弱になる可能性があること;もう一つはより根本的な問題で——**この方法はコンテキストをClaudeに提供するものであり、Claudeが自分で探しに行くものではないということ**。
チームは重要な転換を行いました:Claudeがウェブを検索できるなら、コードベースを検索できないのはなぜか?ClaudeにGrepツールを提供することで、Claudeが自分でファイルを検索し、自分でコンテキストを構築するようにしました。
Beaconのコメントが核心をついています:
**1年の間に、Claudeはほぼ自律的にコンテキストを構築できなかった状態から、複数層のファイルをまたいでネストした検索を行い、必要なコンテキストを正確に見つけられるまでに進化しました。** この進化の鍵は、Claudeにより多くの情報を与えることではなく、より良い検索能力を与えることでした。
Claude CodeにAgent Skillsが導入された際、チームは正式に\*\*漸進的開示(Progressive Disclosure)\*\*の概念を提唱しました:agentが探索を通じて関連するコンテキストを段階的に発見できるようにするものです。
具体的な実装はエレガントです:Claudeはskillファイルを読み込め、それらのファイルはさらに他のファイルを参照でき、モデルはそれらを再帰的に読み込めます。skillの一般的な用途は、ClaudeにAPIの使い方やデータベースのクエリ方法に関する指示を与えるなど、より多くの検索能力を追加することです。
Brian Wagnerは、Claude Codeの考え方と完全に一致する3層の実践を共有しています:
> SKILL.mdは簡潔に保ち(約100行)、重いコンテキストはClaudeが必要なときに発見するファイルに置いています。私はそれを第3層と呼んでいます。あなた方はそれを漸進的開示と呼んでいます。本質は同じです。
PrimeLineはさらに定量的な3層システムを構築しました:JSON設定(約500トークン)> スキーマサマリー(約300トークン)> 完全なmarkdown(約3Kトークン)。コンテキストルーターがタスクのキーワードに基づいてどの層を読み込むかを決定します。
この階層的戦略の核心的なロジックは:**コンテキストは有限のリソースであり、限界効用は逓減する**ということです。すべての情報を一度にagentに詰め込んでも、トークンを無駄にするだけでなく、本当に重要な情報が希薄化されます。オンデマンドで提供し、agentが自分でより深い詳細が必要なタイミングを決める——これこそがスケーラブルな方法です。
Claude Code Guideサブagentは漸進的開示の別の巧妙な応用例です。チームはClaudeがClaude Code自体の使い方についての知識が不足していることに気付きました——MCPの追加方法や特定のスラッシュコマンドの機能について聞いても、答えられませんでした。
すべての情報をシステムプロンプトに詰め込むこともできましたが、ユーザーがこのような質問をすることは稀であり、そうするとコンテキストが汚染され、Claude Codeの主な仕事であるコードを書くことに支障をきたします。
まずClaudeにドキュメントリンクを与えて自分で検索させることを試みましたが——可能ではあるものの、Claudeは正しい答えを探すために大量の結果をコンテキストに読み込んでしまいます。最終的な解決策は、専用のサブagentを構築することでした:Claude Code Guideです。このサブagentは詳細な検索指示を持ち、ドキュメントを効率的に検索する方法と返すべき内容を把握しています。**新しいツールを追加することなく、Claudeの行動空間を拡張しました。**
Lance Martinはagent設計パターンに関する記事で補完的な視点を提示しています:**agentに数十のツールを定義するのではなく、コンピューターを与えてコードでツールを編成させることを考えましょう。** Claude Codeの中核的な抽象化はCLIです——agentはあなたのコンピューター上に生き、bashとファイルシステムというプリミティブを通じて複雑なタスクを完遂します。少数のアトミックなツール(bashツールなど)は、膨大なツールセットよりも柔軟でトークン効率的です。
***
## モデル共感——Agentのように考える
前述の3つのテーマ——少ない方が多い、ツールは陳腐化する、agentに自分で答えを見つけさせる——の背景には共通のメタ方法論があります。Thariqは記事の冒頭でそれを指摘しています:**agentのように世界を見る。**
これはルールのセットではなく、思考方法です。David Zhangがそれに名前をつけています:
**モデル共感(Model Empathy)**——人間の視点から「合理的な」ツールを設計するのではなく、モデルの視点から、それが実際に何を見ているか、どのように理解するか、どのように使うかを考えることです。
Terminally Driftingはすべてのagentチームが経験する同じ教訓を3ステップに精製しています:
> 1)人間のためにツールを設計する
> 2)モデルが管理者権限を持ったアライグマのようにそれらを使う
> 3)トークンエコノミーと予測可能な副作用のために再設計する
「agentのように考える」ことがその重要なブレークスルーです。
Vishは転換点の体験を共有しています:「私たちはずっと人間にとって合理的に感じるツールインターフェースを構築してきて、agentがなぜ奇妙な選択をするのか分かりませんでした。**一度心のモデルを転換して、モデルがツール定義の中で実際に何を見ているかを考えると、すべてが変わりました。**」
Emekaは企業向けツール開発の観点から同じことを述べています:「デフォルトはすべて人間の心的モデルです——スキーマ、フィールド名、フロー、すべてが人間が読むために最適化されています。しかしagentにとって、これは間違ったフレームワークです。」
この「心的モデルの転換」は口で言うのは簡単ですが、実践には継続的な練習が必要です。agentの出力を注意深く読む必要があります——何を正しくできたかを見るのではなく、なぜある選択をしたのか、どこで迷ったのか、どこで遠回りしたのかを見ることです。こうした「異常な行動」はしばしばモデルのバグではなく、ツール設計のバグです。
コミュニティの議論には、いくつかの考察すべき補足視点も現れました。
OAIRは盲点を提起しています:**これらのツールの反復はすべて、agentがステートレスであることを前提としています——各セッションはゼロから始まります。もし最も重要な「ツール」が行動空間の中にあるのではなく、このコードベースがどのように機能するかについての永続的なメモリにあるとしたら?** Claude Codeはその後CLAUDE.mdファイルとmemoryシステムでこの問題に部分的に対応しましたが、永続的な状態管理はagent設計のオープンな課題であり続けています。
Clinkerは別の設計原則を提示しています:**生の能力ではなく回復可能性(recoverability)を最適化する——明示的なツールの前提条件、観察可能な状態、低コストな再試行は、より多くのツールを素早く追加するよりも効果的であることが多い。** これはソフトウェアエンジニアリングの「エラーフリーではなくフォールトトレラントなシステムを作る」という理念と一脈相通じます。
范式折叠の総括が最も鋭い:
> Action Space設計は本質的に権限設計です——AIに何の権限を与えるかで、それがどんな役割になるかが決まります。チームのマネジメントとまったく同じで:ボトルネックが人の能力にあると思っていたら、実はボトルネックは自分が引いた権限の境界にあるのです。
***
## おわりに
Thariqの結語に戻りましょう:
> たくさん実験して、出力を読んで、新しいことを試してみてください。agentのように観察しましょう。
Claude Codeのヘビーユーザーとして、この記事を読んで一番強く感じたことは:**一見「自然に」見える機能の背後には、無数の「これはうまくいかない、別の方法を試そう」という反復があるということです。** AskUserQuestionは3度試み、TodoWriteは3度再設計され、RAGはGrepに置き換えられました。毎回の改善は、よりスマートな解決策を思いついたからではなく、モデルの実際の行動を注意深く観察したからです。
この記事のサブタイトルは「Seeing like an Agent」——agentのように世界を見ること。しかし別の角度から見れば、これはすべての良いエンジニアリング実践の本質でもあります:**自分の視点からシステムを設計するのではなく、使用者の視点から設計する。** ただ今回は、使用者がAIモデルであるというだけです。
将来、agentを構築するすべての開発者は、David Zhangが言う「モデル共感」を習得する必要があるかもしれません。これは神秘的な能力ではなく、核心は3つのことです:**モデルの実際の行動を観察し、その出力を読み、見たものに基づいて設計を調整する。**
agentのように観察しましょう。
***
**関連記事:**
# チップ戦争から宇宙データセンターへ:AIの次の10年
> 投資とは真実の探求です。もし真実をいち早く見つけ、正しく判断できれば、それがアルファを生み出す方法です。そしてそれは、他の人がまだ気づいていない真実でなければなりません。
この言葉は、Atreides Management の創設者 Gavin Baker が Patrick O'Shaughnessy のポッドキャスト「Invest Like the Best」に出演した際のインタビューからのものです。Gavin はテクノロジー投資分野で最も情熱的かつ洞察力のある投資家の一人として知られており、この約2時間の対話では、GPU、TPU、AIエコノミクス、宇宙データセンター、SaaSの未来、そしてスキーインストラクターから投資家へという彼の人生の転換点まで幅広く語られています。
このインタビューは情報密度が非常に高く、「立ち止まって考えさせられる」場面がたくさんあります。以下に、特に深く考えさせられた観点をいくつかご紹介します。
***
## AIの動向を追うには?まず200ドルを使う
インタビューの冒頭、Patrick がとても実践的な質問をしました。Gemini 3 のような新しいモデルがリリースされたとき、どのようにその情報を処理していますか?
Gavin の答えはシンプルでした。**自分で使ってみるしかない**。
しかし重要なのは「使う」こと自体ではなく、どのバージョンを使うかです。彼は、無料版の AI を試しただけで「AI はたいしたことない」と結論づけている投資家たちに驚いたと言います。
> 無料版は10歳の子どもを相手にしているようなものです。その10歳の子のパフォーマンスを見て、35歳になったらどうなるかを予測しているわけです。お金を払えばいい——いや、払わなければなりません、最高クラスのメンバーシップを得るためには、月200ドル必要です。そちらが本物の30〜35歳の大人なのです。
この比喩は非常に的確です。AIモデルのグレード間の差は全く同じ理屈で説明できます——多くの人が無料モデルを少し試して「AIはたいしたことない」と思ってしまいますが、Claude 4.5 Opus、Gemini 3 Pro、GPT-5.2 Reasoning といったトップクラスのモデルを使ったことがあれば、体験はまったく異なります。
費用の話では、私自身も毎月様々な AI 製品のサブスクリプションに相当な金額を使っており、その大半は Claude Code Max(月250ドル)に充てています。開発者でコーディングのニーズが多い方には、ミラーサイトを使わず、公式の Claude Code(125ドルから)に直接サブスクライブすることを強くお勧めします。ミラーサイトが本当に本物のモデルを使っているかどうかわかりませんし、Max プランのコストパフォーマンスは実際に非常に優れていて、従量課金よりもはるかにお得です。
情報収集の手段として、Gavin の答えは多くの人を驚かせるかもしれません。\*\*X(Twitter)\*\*です。
彼は、AI の発展は「X プラットフォーム上でリアルタイムに起きている」と語っています。AI フロンティアを本当に理解している人は地球上に500〜1000人程度しかおらず、かなりの割合が中国にいます。そういった人たちを注意深くフォローする必要があります。彼は特に Andrej Karpathy について言及しました。
> Andrej Karpathy が書いたものはすべて、最低でも3回読まなければなりません。
AI の動向をフォローしている者として、この点には深く共感します。Twitter 上の AI に関する議論は、あらゆるニュースメディアよりもリアルタイムで、かつ深いものです。各ラボの研究者たちは最新の進捗を直接投稿し、互いに「議論」を交わすこともあります——Gavin は Meta の PyTorch チームと Google の Jax チームが X 上で公開論争を繰り広げたことがあり、最終的に両ラボのトップが「うちの人間が相手のラボの悪口を言ってはならない」と表明しなければならなかったと語っています。
***
## Scaling Laws:我々の「古代エジプトの瞬間」
Gemini 3 のリリース後、多くの人が Scaling Laws(スケーリング則)について何を示しているかに注目しました。Gavin は私がかつて聞いたことのない視点を提示しました。
> 事前学習のスケーリング則に関する私たちの理解は、おそらく古代エジプト人の太陽に対する理解と同じようなものです。彼らは非常に精密に測定できました——大ピラミッドの東西軸が春分・秋分と完全に一致し、ストーンヘンジも同様です。完璧な測定でした。しかし彼らは軌道力学を理解していませんでした。なぜ太陽が東から昇り、西に沈むのか、わかっていなかったのです。
この比喩は、しばらく考え込ませるものでした。確かに私たちは非常に精密に予測できます。モデルの計算量を10倍にすれば、性能がどれだけ向上するかを。しかし、なぜそうなるのかはわかりません。これは「法則」ではなく、「経験的観察」——非常に精密に計測できるが、その原理は理解していない経験的観察です。
では、なぜ Gemini 3 が重要なのでしょうか?それは、この「経験的観察」がまだ成立することを証明したからです。Blackwell チップの遅延が続き、皆が「スケーリング則はもう限界なのか」と心配していた時期に、Gemini 3 は明確な答えを出しました。**限界ではない**、と。
しかしさらに興味深いのは、その後に Gavin が語ったことです。推論モデル(reasoning models)が登場しなければ、2024年から2025年にかけての AI の発展は本来停滞していたはずだというのです。
なぜでしょうか?XAI が20万枚の Hopper GPU を協調動作させることに成功した後、次のステップでは Blackwell チップを待つ必要がありました。20万枚を超える Hopper GPU を「コヒーレント」(coherent)——つまり一体として動作させること——に保つことはできません。そして Blackwell は遅延しました。
> 推論モデルがなければ、2024年半ばから現在まで、AI には何の進展もなかったでしょう。すべてが止まっていたはずです。それが市場に何を意味するか想像できますか?私たちはまったく異なる環境の中にいたでしょう。推論モデルはある意味で AI を救いました。Blackwell なしで AI が前進し続けることを可能にしたからです。
これは、私がそれまで意識していなかった視点です。推論モデル(o1 など)は単なる新しい能力ではなく、AI 業界全体の発展ペースを実際に「救った」のです。
***
## チップ戦争:Google が「酸素を吸い尽くしている」
GPU と TPU の競争について語る中で、Gavin は印象的な言葉を述べました。
> Google は現在、トークンの最低コスト生産者です。彼らがやってきたことは、「AI エコシステムの経済的な酸素を吸い尽くす」ことだと言えるでしょう——これは彼らにとって極めて合理的な戦略です。
低コスト生産者として、Google は低価格(場合によっては赤字)で AI サービスを提供し続け、競合他社を苦しめてきました。これはテクノロジー業界の古典的な戦略ですが、Gavin は興味深い変化を指摘しています。
> AI は私のキャリアの中で初めて、テクノロジー分野で「低コスト生産者」であることが本当に重要になった事例です。Apple が携帯電話の低コスト生産者だから時価総額が数兆ドルなのではありません。Microsoft がソフトウェアの低コスト生産者だから数兆ドルなのではありません。NVIDIA が AI アクセラレーターの低コスト生産者だから数兆ドルなのでもありません。これまでそれが重要だったことは一度もありませんでした。
しかし AI の時代には、電力が制約要因となり、**1ワット当たりのトークン数**が極めて重要になります。1ワットで3〜5倍のトークンを生成できれば、それは3〜5倍の収益です。コンピューティングの価格は関係なくなります。なぜなら、ボトルネックは電力だからです。
この状況はまもなく変わろうとしています。Blackwell チップの展開がついに始まり、Gavin は最初の Blackwell モデルは XAI から登場すると予測しています。
> Jensen によれば、Elon よりも速くデータセンターを建設できる人間はいないそうです。Jensen はこれを公言しています。
Blackwell およびそれに続く Ruben チップが大規模に展開されれば、低コスト生産者としての Google の優位性は消えてしまいます。その時でも、彼らはマイナス30%の粗利率で AI ビジネスを運営し続けようとするでしょうか?その計算は完全に変わることになります。
***
## 宇宙データセンター:常識外れだが、第一原理から見れば正しい
Patrick が「あまり議論されていない、ちょっと突飛なアイデアはありますか?」と聞いたとき、Gavin は宇宙データセンターについて語り始めました。最初は冗談かと思いましたが、彼の分析を聞いた後、これがインタビュー全体の中で最も先見性のある部分かもしれないと気づきました。
> 第一原理の観点から見ると、宇宙データセンターはあらゆる次元において地球上のデータセンターよりも優れています。
彼の論証はこのようなものです。
**1. エネルギー**:宇宙では、衛星は24時間太陽光にさらされており、太陽放射の強度は地上の6倍です。常に日光が当たっているため、バッテリーが不要です——バッテリーはコストのかなりの部分を占めています。したがって、太陽系で最もコストの低いエネルギーは「宇宙太陽光発電」です。
**2. 冷却**:地球上のデータセンターでは、コストと重量の大半が冷却に使われます。しかし宇宙では?冷却は無料です。衛星の日陰側に放熱板を置けば、そこは絶対零度に近い温度です。
**3. ネットワーク**:データセンターではラック間を光ファイバーで接続します——本質的にはケーブルを通るレーザーです。それよりも速いものは何でしょうか?真空を通るレーザーです。宇宙の衛星をレーザーで接続すれば、ネットワーク速度は実際には地上のデータセンターよりも速くなります。
**4. ユーザー体験**:現在、AI に質問すると、信号はスマートフォンから基地局、光ファイバー、どこかのデータセンターへと向かい、処理されて同じルートで戻ってきます。しかし衛星がスマートフォンと直接通信できれば(Starlink はすでに直接接続能力を実証しています)、経路全体ははるかに短くなります。
もちろん、これを実現するには Starship の大規模打ち上げが必要で、あと5〜6年はかかるかもしれません。しかし Gavin は興味深い収束を指摘しています。Tesla、SpaceX、XAI が融合しつつあります。XAI は Optimus ロボットの「インテリジェンスモジュール」となり、SpaceX は AI に算力を提供するために宇宙にデータセンターを構築する——この3社は互いの競争優位性を高め合うフライホイールを形成しつつあります。
***
## SaaS の「燃えるプラットフォーム」
前のセクションで AI の未来に興奮した方も、このセクションでは多くの既存企業の将来について心配になるかもしれません。
Gavin は率直に述べています。**アプリケーション SaaS 企業は、実店舗小売業者が EC に直面した際に犯したのとまったく同じ間違いを犯しています。**
かつての実店舗小売業者は Amazon を見て、「EC は低利益率のビジネスだ、どうして私たちより効率的になれるのか?今は顧客がお金を払って店に来て、自分で商品を持ち帰る」と思っていました。顧客のニーズははっきりと見えていたのに、EC の利益率構造が気に入らないという理由で投資を拒んだのです。結果はどうなったでしょうか?Amazon の北米小売事業の利益率は今や多くの伝統的小売業者よりも高くなっています。
SaaS 企業も今、同じ状況に直面しています。従来のソフトウェアは一度書けば無制限に複製・配布でき、粗利率は80〜90%に達することもあります。しかし AI は違います——使用するたびに新たな計算が必要で、優れた AI 企業の粗利率は40%程度にとどまることもあります。
> AI エージェントを構築したいが、粗利率が35%を下回ることを嫌がっているなら、あなたは決して成功しません。なぜなら AI ネイティブの企業はそのような利益率で運営しているからです。80%の粗利率を守ろうとするなら、AI の分野で失敗することを自ら保証しているようなものです。絶対に保証されています。
Gavin はこれを「生死を分ける決断」と呼び、**Microsoft を除いてほぼすべての企業が失敗しつつある**と述べています。
彼はノキアの有名な「燃えるプラットフォーム」メモを引用しました。あなたのプラットフォームは燃えています。しかし、すぐ隣には素晴らしい新しいプラットフォームがあります。そこに飛び移って、元のプラットフォームの火を消しに戻ればいい。そうすれば2つのプラットフォームを持つことになります。
Salesforce、ServiceNow、HubSpot、GitLab、Atlassian——これらすべての企業はこの戦略を実行できるし、実行すべきだと彼は考えています。AI の収益を公開し、AI の粗利率を公開し(低い粗利率こそが「本物の AI」の証明になります)、そして未だに赤字のベンチャーキャピタル支援の競合他社を指して「私には彼らにないものがある:キャッシュフローを生み出すビジネスだ」と言えばいい。
***
## ある投資家の成長物語
インタビューの終盤、Patrick はより個人的な質問をしました。若い人たちに自分の仕事をどう紹介しますか?
Gavin の答えは「投資とは真実の探求だ」から始まりますが、本当に興味深いのは彼の人生のストーリーです。
彼の当初の計画は、冬はスキーインストラクター、夏はラフティングガイド、オフシーズンはロッククライミング、合間に小説を書いたり野生動物の写真を撮ったりするというものでした。これが大学時代の「人生設計」であり、両親も大いに賛成していました。
しかし両親は一つだけお願いをしました。専門的なインターンシップを一つだけ、何でもいいから探してみないかと。
彼が見つけられた唯一のインターンシップは、ある証券会社のプライベート・ウェルス・マネジメント部門でした。仕事は単純で、会社がリサーチレポートを発行するたびに、どのクライアントがその株を保有しているかを調べてレポートを郵送するというものでした。
そして彼はそのレポートを読み始めました。
> 「なんてことだ、これは私が想像できる中で最も面白いことだ」と思いました。
彼は投資を「スキルと運の混在するゲーム」、ポーカーにやや似たものとして理解しました。運の悪さで負けることもあります——投資先の企業の本社に隕石が落ちるような——しかしほとんどの場合、スキルが重要です。そして優位性を得る方法は、最も深い歴史的知識と、現在の世界に対する最も正確な理解を組み合わせて、「次に何が起きるか」について差別化された見解を形成することです。
それはインターンシップ3日目のことでした。彼は書店に行き、ピーター・リンチの本を買って2日で読み終えました。次にバフェット、『マーケットの魔術師』を読み、バフェットの株主への手紙を2回読みました。それから独学で会計を学び、学校に戻ると専攻を英語と歴史から歴史と経済学に変更しました。
彼はまた、清掃員として働いた経験についても語りました。Alta スキーリゾートでアルバイトをしていたとき、客室清掃をしていました。ある日、部屋を掃除していると、宿泊客が自分と同じ本を読んでいるのに気づきました。「いい本ですね、私もちょうど同じくらいのところまで読んでいます」と言うと、相手は宇宙人でも見るような目で彼を見て、さらに驚いて聞きました。「あなた、本を読むんですか?」
> その出来事は、他者への接し方を永久に変えました。
***
## エピローグ:AI が必要とするものは、必ず手に入る
インタビューの締めくくりに近いところで、Gavin は最も興味深いと感じた言葉を語りました。
> ここ2年間、AI が発展し続けるために必要なものは何でも手に入りました。アメリカの世論が原子力のように急速に転換したのを、他に見たことがありますか?それが起きたのです。そして AI がそれを必要としていたまさにそのタイミングで起きました。今は地球上の電力制約に直面しており、突然、宇宙データセンターの議論が生まれています。AI の発展を遅らせるものが現れるたびに、むしろすべてが加速するのです。
これは、Kevin Kelly が『テクニウム——テクノロジーはどこへ向かうのか?』で提唱した「technium(テクニウム)」という概念を思い起こさせます。技術全体として、何かしら自分の意志を持ち、ますます強大になろうとしているかのようです。
これは単なる偶然かもしれません。あるいは単に多くの賢い人々が問題を解決しているだけかもしれません。しかし Gavin が観察したこのパターン——AI が障害に直面するたびに、その障害がなんとかして取り除かれる——は、確かに考えさせられるものです。
# キュレーション
# Curations キュレーション
テック分野の著名な動画や良質なブログを観た後の考察とまとめを収録しています。
単なるメモの抜粋ではなく、自分なりの理解と実践から得た知見を織り交ぜています。
## コンテンツソース
* テック動画の解説
* ブログ記事のキュレーション
* ポッドキャスト・インタビューのまとめ
## 最新コンテンツ
### Agentのように考える:Claude Code チームのツール設計哲学
Anthropic エンジニアの Thariq が、Claude Code 構築過程での agent ツール設計経験を共有しました。AskUserQuestion の3回のイテレーションから TodoWrite の3回のリファクタリング、RAG からプログレッシブ・ディスクロージャーまで、すべてのケースが同じコア方法論を指し示しています:Agent のように世界を見ること。
[全文を読む →](./claude-code-seeing-like-an-agent)
### 大学生の90%がAIを使っているが、ルールは誰も知らない
プリンストン、バークレー、LSEの4人の学生が、キャンパスにおけるAIのリアルな状況について語りました。不正行為、困惑、二極化、そして誰も聞けない問い:大学にはまだ何の意味があるのか?プロモーションではなく、本音の困惑と思考の記録です。
[全文を読む →](./ai-on-campus-student-perspectives)
### チップ戦争から宇宙データセンターへ:AI業界の次の10年
チップ戦争から宇宙データセンター、SaaSの生死を分ける選択から投資の本質まで。Gavin Baker が「Invest Like the Best」ポッドキャストで、AI業界についての最も深い洞察を共有しました:なぜ有料版AIを使わなければならないのか、Scaling Laws はなぜ「古代エジプト人の太陽理解」のようなのか、推論モデルがいかにAI業界全体の発展ペースを「救った」のか。
[全文を読む →](./gavin-baker-ai-economics)
# Karpathy のツイートは爆発的に 62,000 個のスターを獲得しました: andrej-karpathy-skills は一体何をしたのですか?
2026 年 1 月 27 日、Andrej Karpathy は X に非常に長いツイートを投稿しました。これは 11 セクション、約 1,400 ワードのプログラミング エッセイで、11 月の「手書き 80% + エージェント 20%」から 12 月の「エージェント 80% + 研磨 20%」への移行中に遭遇した落とし穴を記録しています。このツイートは最終的に **769 万回の閲覧、39,000 件の「いいね!」、36,000 件のブックマーク**に達しました。
3 か月後、`forrestchang/andrej-karpathy-skills` という GitHub リポジトリが開始され、Karpathy のツイートが 4 つのインストール可能なルールにパッケージ化されました。 2 週間以内に **62.7,000 のスターと 5.5,000 のフォーク**に達し、2026 年 4 月の GitHub 週間リストで第 1 位になりました。
ウェアハウス オントロジー: **Markdown ファイル**。
***
##1.カルパシーは何を訴えていますか?
Karpathy 氏の長い記事には、基本的に LLM でコード化された 4 つの「慢性疾患」が列挙されています。
**最初の病気: 密かに思い込みをする**
> 「最も一般的なタイプの間違いは、モデルが誤った仮定を立て、それを検証せずに適用することです。モデルは自分自身の混乱を管理せず、明確化を求めず、矛盾を示さず、トレードオフを提示せず、反論すべきときに反論せず、少しお世辞になりすぎます。」
これは「共謀ミス」です。 「ログインを追加してください」と言うと、どのような認証を使用するか、デバイスを記憶するかどうか、セッションを管理する方法などは尋ねられません。合理的と思われるソリューションを提示するだけです。レビューが終わって、自分が求めているものと違うことがわかったときには、すでに 500 行が書かれているでしょう。
**第二の病気: オーバーエンジニアリング**
> 「彼らは特に、コードと API を過度に複雑にし、抽象化レイヤーを肥大化させ、死んだコードをクリーンアップしないことを好みます。彼らは、非効率で肥大化した脆弱な構造を実装するために 1000 行のコードを使用します。彼らが「もちろんです!」と言う前に、子供のようになだめて、「まあ、これをやればいいのでは?」と言う必要があります。そしてすぐに 100 行に減らします。」
これは、LLM コーディングの最も典型的な観察者効果です。コンテキストが「寛大に」与えられている一方で、複雑さにも「寛大に」報われます\*\*。ストラテジー モード、ファクトリー モード、依存関係の注入 - すべてが与えられます。
**第三の病気: 変えてもらっていないものを変える**
> 「気に入らない、または完全に理解していないという理由で、コメントやコードを変更または削除することがあります。たとえそれらの変更が現在のタスクと何の関係もない場合です。」
バグの修正を依頼すると、その隣にある未完成の TODO コメントが「もう必要ないようです」という理由で都合よく削除されます。
**病気4:CLAUDE.mdにルールを書いても壊れてしまう**
> 「CLAUDE.md で簡単な修復を試みたとしても、上記の問題は依然として存在します。」
これはツイート全体の中で最も悲痛な文章です。元 OpenAI 創設チーム メンバーで Tesla AI ディレクターの Karpathy 氏は、Claude を完全に一致させる CLAUDE.md を書くことができませんでした。
***
## 2. `andrej-karpathy-skills` の解決策
`forrestchang` これら 4 つの疾患に対する解決策を 4 つの原則に体系化し、CLAUDE.md ファイルにパッケージ化します。
| 原則 | 対応する病気 | 主要なアクション |
| --------------------- | ------------ | ----------------------------------------------------------------------- |
| **コーディングする前に考えてください** | 密かに想定 | 自分の仮定を明確に述べ、複数の解釈を列挙し、混乱していないかどうかを立ち止まって尋ね、必要に応じて反論します。 |
| **シンプル第一** | オーバーエンジニアリング | 必要な最小限のコードのみを記述します。投機的な柔軟性、エラー処理、抽象化を記述しないでください。 |
| **外科的変更** | 不正な変更 | 必要なものだけに触れてください。スタイルをリファクタリングしたり変更したりしないでください。他のデッドコードは報告するだけで、削除はしません。 |
| **目標主導型の実行** | メソッドの不整合 | 検証可能な成功基準とテストを提供し、モデルのループを自動的に通過させます。 |
4 番目の原則は、カルパシーのツイートの別の有名な行「レバレッジ」への直接の言及です。
これは、プロジェクト方法論全体の足がかりです。 \*\*最初の 3 つの原則は、LLM の混乱を防ぎます。 4 番目の原則は、その強みを真に活用する方法を示しています。 \*\*
***
## 3. インストール方法と使用方法
**方法 A: クロード コード プラグインとして** (最初にこれを使用することをお勧めします)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
インストール後、すべてのプロジェクトのクロード コード ダイアログは、これら 4 つの原則に自動的に準拠します。これはグローバルに有効になり、`/plugin` でいつでもオフにできます。
**方法 B: CLAUDE.md を手動でコピーします**
リポジトリに入力し、`CLAUDE.md` を開き、コピーして、プロジェクトのルート ディレクトリの `CLAUDE.md` に貼り付けます。このプロジェクトにのみ有効です。
このリポジトリは、追加の `CURSOR.md` および `.cursor/rules/` の適応、つまり主流の AI IDE をカバーするコンテンツのセットも提供します。
***
##4.なぜ62kスターに到達できるのでしょうか?
これは解明する価値のある現象です。 62.7,000 個のスターは、「単一ファイル リポジトリ」としては誇張された数字です。比較のために、同じ期間の Microsoft マークイットダウン (9,000) とアディ オスマニのエージェント スキル (4.6,000) を合わせたものほど多くはありません。
衝撃重量別に分類すると次のようになります。
**1. Karpathy IP 承認** - `forrestchang-skills` という名前の場合、同じコンテンツが 10,000 を超えることはできません。 Karpathy 氏は「元 OpenAI 創設チーム + Tesla AI ディレクター + CS231n インストラクター」という文化資本を持っており、彼のツイートには「必読」ラベルが付いています。
**2.完璧なタイミング** — Opus 4.7 は 4 月 16 日にリリースされましたが、オーバーエンジニアリングに関する苦情はピークに達していました。このレポは、誰もが「クロードの狂気を止める」解毒剤を探していたまさにその時に現れました。
**3.問題点は普遍的です** - すべてのクロード コード/カーソル ユーザーはこれら 4 つの落とし穴を踏んだことがあり、共感率は 100% に近いです。
**4.しきい値は非常に低く、** - 1 ファイルまたは 2 行のコマンドです。スターのコストは無視できるほど低いです。 「インストールしないと損ですよ。」
**5.強い検証可能性** - 4 つの原則は明確で覚えやすく、スクリーンショットや転送も簡単です。法外な 1000 行プロンプトのプロジェクト ガイドとは異なります。
**6.バイリンガル README** ——`README.zh.md` は中国 AI サークルのトラフィックを直接食い尽くし、Weibo 上で V2EX/瞬時に/同時に爆発します。
**7.著者によるクロスプロモーション** - 一番上の列の文 *「私の新しいプロジェクト Multica をチェックしてください」* は、著者自身の商用エージェント プラットフォーム `multica-ai/multica` にトラフィックを誘導します。 \*\*このリポジトリは本質的に Multica の顧客獲得ファネルの最上位にあります。 \*\*
**8.メタ フィット** - この記事で説明されている「LLM コーディング エラー」は、まさにすべての読者が LLM でコーディングするときに経験しているものです。読み取りと使用が統合されており、コンバージョン率は非常に高くなります。
一言で言えば、売っているのはコードやツールではなく、**Karpathy の感情をインストール可能なルールにパッケージ化したもの**です。これは、2026 年の AI プログラミング界における最も典型的な「コンテンツは製品である」ケースです。
***
## 5. 私の使用上の提案
**最初に方法 A を使用してグローバルにインストールします**。ツールやスクリプトを作成するときのエクスペリエンスが向上するかどうか、特にクロードに他の人のコードを変更させる場合、ランダムな変更を行う問題が軽減されるかどうかを確認してください。
**1 ~ 2 週間後に CLAUDE.md プロジェクトにマージするかどうかを決定します**。各プロジェクトの CLAUDE.md にはすでにドメイン知識 (設計システム、コンポーネント仕様、展開プロセス) が詰め込まれていますが、Karpathy のセットは一般的な方法論です。この 2 つは競合するものではなく、重ね合わせることができます。ただし、このタイミングは、それが役立つと本当に確信できるまで待つ必要があります。
**コストに注意してください**: クロードにさらに多くの質問をさせることになるため、「一文で生成する」ことに慣れている人にとっては煩わしいでしょう。実行すべきわずかなクリーニングが実行されない可能性があります (厳密すぎる)。非常に曖昧な探索タスクによって制約を受けることになります。
**より深い価値**: 要件を明確に記述する必要があります。これは、すべての高品質のソフトウェア エンジニアリングの前提条件となります。
**注目すべきフォローアップ**: forrestchang 自身も、スキル メカニズムを製品化するための「オープンソースのマネージド エージェント プラットフォーム」である Multica を推進しています。この 4 つの原則が最終的に事実上の標準になれば、Multica がその商用手段となるでしょう。この行に注目してください。
***
## 参考リソース
# Indie Dev
インディー開発の完全なプロセスを記録 — 準備からリリースまでの全ステップ。
# Bark
# Bark
シンプルなHTTPリクエストでiPhoneにカスタムプッシュ通知を送信できるツールです。無料、オープンソース、セルフホスト対応。
## おすすめの理由
* **シンプルなAPI** - curlコマンド1つでプッシュ通知を送信、複雑な設定は不要
* **オープンソース&無料** - MITライセンス、完全オープンソース、費用ゼロ
* **プライバシー重視** - サーバーのセルフホストに対応、プッシュデータを完全に自分で管理
## 活用シーン
* **スクリプト通知** - データバックアップの完了など、長時間タスクの終了時にアラートを受け取る
* **サービス監視** - サーバー異常時に即座にアラートを受信
* **自動化連携** - CI/CDビルド結果、定期タスク完了通知
## クイックスタート
1. [App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865) から Bark をダウンロード
2. アプリを開いて、プッシュURLをコピー
3. 最初の通知を送信:
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
タイトル付きプッシュ:
```bash
curl https://api.day.app/YOUR_KEY/タイトル/内容
```
## プロジェクト情報
* GitHub: [Finb/Bark](https://github.com/Finb/Bark)
* Stars: 7.2k+
* ライセンス: MIT
# Toolkit
# Toolkit
見つけた良質なGitHubプロジェクトと実用的なソフトウェアを収録しています。
# Claude Agent Teams 完全ガイド
## はじめに
Claude Code の Subagent を使ったことがあれば、並行開発は十分にパワフルだと感じているかもしれません。しかし Subagent には制限があります:メイン Agent に結果を報告することしかできず、互いにコミュニケーションを取ることができません。
Agent Teams はこの状況を根本的に変えます。想像してみてください:1つの Agent がセキュリティレビューを担当し、もう1つがパフォーマンス最適化を担当し、3番目がテストカバレッジを担当する——並行して作業できるだけでなく、直接対話し、互いに挑戦し、コンセンサスに達することができます。これが Agent Teams の核心的な価値です。
## Agent Teams を理解する
Agent Teams のアーキテクチャは、実際の開発チームによく似ています:
```
┌─────────────────────────────────────────────────────────┐
│ あなた(ユーザー) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (メイン Claude インスタンス、調整を担当) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│ セキュリティ │◄─►│ パフォーマンス│◄─►│ テスト │
│ レビュー │ │ 最適化 │ │ カバレッジ │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└──────────────┼──────────────┘
▼
┌─────────────────┐
│ 共有タスクリスト │
└─────────────────┘
```
### Subagent との違い
| 特性 | Subagent | Agent Teams |
| ------------- | ------------------------- | ------------------------- |
| **コンテキスト** | 独立コンテキスト、結果をメイン Agent に返す | 独立コンテキスト、完全に独立して動作 |
| **通信方法** | メイン Agent にしか報告できない | Teammates 間で直接通信可能 |
| **タスク調整** | メイン Agent がすべての作業を管理 | 共有タスクリストで自律的に調整 |
| **適用シナリオ** | 結果だけが必要なフォーカスタスク | 議論と協力が必要な複雑な作業 |
| **Token コスト** | 低め:結果の要約がメインコンテキストに返される | 高め:各 Teammate が独立したインスタンス |
簡単に言えば:**Subagent はタスクを実行するために派遣する下請け業者、Agent Teams は同じ部屋で協力するプロジェクトチーム**です。
### なぜ Agent Teams が効果的なのか
核心的な洞察:**専門化が集中をもたらす**。
単一の Agent が複雑なマルチステップタスクを処理する際、コンテキストは膨張し続け、頻繁に `/clear` でリセットが必要になります。Agent Teams は各 Teammate に狭い専門領域を持たせ、コンテキストをクリーンに保ち、パフォーマンスをより安定させます。
## Agent Teams の有効化
Agent Teams は現在実験的機能であり、デフォルトでオフになっています。手動で有効にする必要があります:
**方法1:環境変数**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**方法2:settings.json(推奨、永続的に有効)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## コアの使い方
### 最初の Agent Team を作成する
有効化後、自然言語で Claude にチームの作成を指示するだけです:
```
agent team を作成して PR #142 をレビューして。
3人のレビュアーを生成して:
- 1人はセキュリティ問題に焦点
- 1人はパフォーマンスへの影響をチェック
- 1人はテストカバレッジを検証
それぞれがレビューして発見を報告するようにして。
```
**キーワードのヒント**:"create an agent team" または "spawn an agent team" を使用してください。"spawn agents" だけだと、Subagent と Agent Teams が混同される可能性があります。
### 表示モード
Agent Teams は2つの表示モードをサポートしています:
| モード | 説明 | 要件 |
| --------------- | --------------------------- | ------------------- |
| **In-process** | すべての Teammates がメインターミナルで動作 | 特別な要件なし |
| **Split panes** | 各 Teammate が独立したペイン | tmux または iTerm2 が必要 |
デフォルトは `auto`:tmux 内で実行している場合は split panes を使用し、それ以外は in-process を使用。
**表示モードの設定**:
```json
{
"teammateMode": "in-process"
}
```
**単一セッションでの指定**:
```bash
claude --teammate-mode in-process
```
### よく使うショートカットキー
| 操作 | ショートカット |
| ---------------------- | ------------ |
| Teammates 間の切り替え | `Shift+Down` |
| 前の Teammate に戻る | `Shift+Up` |
| タスクリスト表示の切り替え | `Ctrl+T` |
| 現在の Teammate を中断 | `Escape` |
| **Delegate Mode を有効化** | `Shift+Tab` |
| Teammate セッションに入る | `Enter` |
### Delegate Mode(重要)
Delegate Mode は Agent Teams の最も重要な機能の1つです:
| モード | Lead の動作 |
| ----------------- | ---------------------------------- |
| **通常モード** | Lead がタスクを自分で実装し、コードを書く可能性がある |
| **Delegate Mode** | Lead は調整のみ可能、コードを書いたりテストを実行したりできない |
**なぜ Delegate Mode が必要なのか**:
この制限がないと、Lead は頻繁に「仕事を奪う」ことがあります——3人の Teammates が待っているにもかかわらず、Lead が自分でコードを書き始めてしまう。Delegate Mode を有効にすると、Lead は純粋なプロジェクトマネージャーに強制され、タスク管理、Teammates とのコミュニケーション、出力のレビューしかできなくなります。
```
# チームを起動したら直ちに Shift+Tab を押して有効化
```
## 実践事例
### 事例1:並行コードレビュー
単一のレビュアーは特定の種類の問題に深入りしがちです。レビューの次元を独立した領域に分割することで、セキュリティ、パフォーマンス、テストカバレッジすべてに同等の注意が払われます:
```
agent team を作成してこの PR をレビューして。3人のレビュアーを生成して:
- セキュリティレビュアー:認証、認可、インジェクション脆弱性をチェック
- パフォーマンスレビュアー:アルゴリズムの複雑度、データベースクエリ、キャッシュ戦略を分析
- テストレビュアー:テストカバレッジ、境界条件、エラーハンドリングを検証
各自がレビューした後、発見した問題について互いに議論するようにして。
```
### 事例2:競合的仮説デバッグ
根本原因が不明な場合、単一の Agent は一見もっともらしい説明を見つけたら停止する傾向があります。Teammates に互いに挑戦させることでこの問題を回避できます:
```
ユーザーからフィードバック:アプリがメッセージを1つ送信した後に接続を維持せずに終了する。
5つの agent teammates を生成して異なる仮説を調査して。互いに議論し、
互いの理論を反証しようとするように、科学的な討論のように。合意に達した
発見を調査レポートに更新して。
```
**キーメカニズム**:討論構造。複数の独立した調査者が積極的に互いの理論を覆そうとし、最終的に生き残った仮説が真の根本原因である可能性が高くなります。
### 事例3:コンテンツの大量生産
これは非技術タスクの典型的な適用——1つの入力を複数の出力に変換:
```
agent team を作成してこのビデオスクリプトを4つのプラットフォーム向けコンテンツに変換して:
- LinkedIn 記事ライター
- Twitter スレッドライター
- Newsletter ライター
- ブログ記事ライター
スクリプトの場所:/content/scripts/video-20.md
```
各 Teammate が独立して制作しつつ、コンテンツの一貫性を維持します。
### 事例4:QA 品質チェッククラスター
ブログウェブサイトの品質チェックで、5つの Agent を並行してデプロイして異なる側面をテスト:
```
agent team を作成してブログの包括的な品質チェックを実施して:
- Agent 1:コアページテスト(ホーム、アバウト、お問い合わせ)
- Agent 2:記事ページテスト(レンダリング、ナビゲーション、SEO メタデータ)
- Agent 3:リンクチェック(内部リンク、外部リンク、デッドリンク)
- Agent 4:SEO 検証(タイトル、説明、構造化データ)
- Agent 5:アクセシビリティテスト(ARIA ラベル、コントラスト、キーボードナビゲーション)
優先度順に並べた問題レポートを生成して。
```
**効果**:本来なら人手で順次実行する必要がある包括的なチェックを数分で完了し、各 Agent が自分の領域に集中し、最後に優先度順の問題リストに統合します。
### 事例5:マルチラウンドディスカッションモード
有用なプロンプトパターン——Teammates にミーティングのように議論させる:
```
Agent Teams を使って4つの teammates を作成して [技術的決定] について議論して、
3ラウンドのディスカッションを行って。各ラウンドで teammates が互いに交流するようにして。
そのうち1つの teammate は Red Team の視点を専門に担当し、批判的な意見を述べるようにして。
```
このモードは、アーキテクチャの決定や技術選定など、多角的な検討が必要なシナリオに特に適しています。
### 事例6:C コンパイラプロジェクト
Anthropic は16の Agent を使って、Linux カーネルをコンパイルできる C コンパイラをゼロから構築しました:
| 指標 | データ |
| --------- | ------------------------- |
| Agent 数 | 16の並行インスタンス |
| セッション数 | 約2,000の Claude Code セッション |
| コスト | 約$20,000 |
| コード行数 | 100,000行 |
| Token 使用量 | 20億入力 + 1.4億出力 |
最終成果:x86、ARM、RISC-V でブート可能な Linux 6.9 をビルドできる Rust コンパイラ。
## チーム管理
### Teammates とモデルの指定
Claude はタスクに基づいて生成する Teammates の数を自動的に決定しますが、明示的に指定することもできます:
```
4つの teammates を作成してこれらのモジュールを並行してリファクタリングして。
各 teammate は Sonnet モデルを使用して。
```
### 計画の承認を要求
複雑またはハイリスクなタスクでは、Teammates に実行前にまず計画を策定するよう要求できます:
```
認証モジュールをリファクタリングするアーキテクト teammate を生成して。
変更を行う前に、計画の承認を要求するようにして。
```
Teammate が計画を完成すると、Lead に承認リクエストを送信します。Lead はレビュー後に承認するか、修正意見を返すことができます。
### Teammates と直接対話
各 Teammate は完全な Claude Code セッションです。どの Teammate にも直接メッセージを送ることができます:
* **In-process モード**:`Shift+Down` で切り替え、メッセージを入力
* **Split-pane モード**:対応するペインを直接クリック
### Teammates の終了
```
セキュリティレビューの teammate を終了して
```
Lead が終了リクエストを送信し、Teammate は承認するか拒否する(理由を説明して)ことができます。
### チームのクリーンアップ
完了後、Lead にリソースをクリーンアップさせます:
```
チームをクリーンアップして
```
**重要**:常に Lead を通じてクリーンアップしてください。Teammates にクリーンアップを実行させると、リソース状態の不整合が生じる可能性があります。
## ベストプラクティス
### チームサイズの制御
サイズの推奨:
| チームサイズ | 適用シナリオ |
| ------ | ------------------------- |
| 3人 | シンプルなマルチビューレビュー |
| 4-5人 | 標準的な機能開発やリファクタリング |
| 6人以上 | 大規模マイグレーションや複雑なアーキテクチャタスク |
**経験則**:各 Teammate に5-6個のタスクを割り当てるのが適切です。15個の独立したタスクがある場合、3人の Teammates が良い出発点です。
### タスクの粒度
* **小さすぎる**:調整のオーバーヘッドが効果を上回る
* **大きすぎる**:Teammates がチェックポイントなしで長時間作業し、無駄のリスクが増加
* **ちょうど良い**:独立して完結し、成果物が明確な作業単位(1つの関数、1つのテストファイル、1つのレビューレポート)
### ファイル衝突の回避
2人の Teammates が同じファイルを編集すると上書きが発生します。作業を分割する際は、各 Teammate が異なるファイルセットを担当するようにしてください:
```
Teammate 1:src/auth/ ディレクトリを担当
Teammate 2:src/api/ ディレクトリを担当
Teammate 3:src/utils/ ディレクトリを担当
```
### 監視とガイダンス
定期的に Teammates の進捗をチェックし、適切でない方向を早期に修正してください。チームを長時間監視なしで実行させると、無駄のリスクが増加します。
Lead がタスクを Teammates に待たせず自分で実装し始めた場合:
```
teammates がタスクを完了するのを待ってから続けて
```
### 十分なコンテキストを与える
Teammates はプロジェクトのコンテキスト(CLAUDE.md、MCP servers、skills)を自動的にロードしますが、Lead の会話履歴は継承しません。生成時に十分なタスクの詳細を提供してください:
```
セキュリティレビューの teammate を生成して、以下のプロンプトで:
"src/auth/ ディレクトリの認証モジュールのセキュリティ脆弱性をレビューして。
トークン処理、セッション管理、入力バリデーションに重点を置いて。
アプリケーションは httpOnly cookies に保存された JWT トークンを使用している。
発見を報告する際は深刻度の評価を添えて。"
```
### セルフレポーティング検証モード
タスクの説明に明確な検証基準を含め、Teammates が完了後にセルフチェックを行うようにします:
```
タスク完了時に、Lead に以下を報告して:
1. チェックしたファイル
2. 発見した問題
3. 行った変更
4. 検証基準(テスト合格、lint 警告なしなど)が満たされたかどうか
```
このモードにより、Lead の検証作業量が減り、タスクが「完了したように見える」のではなく本当に完了していることが保証されます。
## 高度なテクニック
### Hooks で品質ゲートを強制
Hooks を通じて Teammates が作業を完了した際にルールを強制実行します:
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle`:Teammate がアイドル状態になろうとする時に実行。exit code 2 を返すとフィードバックを送信して Teammate に作業を継続させる
* `TaskCompleted`:タスクが完了としてマークされた時に実行。exit code 2 を返すと完了をブロックしてフィードバックを送信
### 権限の事前承認
Teammate の権限リクエストは Lead にバブルアップし、頻繁な中断を引き起こす可能性があります。生成前に一般的な操作を事前に承認しておきます:
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### Worktree との組み合わせ
Agent Teams は Worktree と組み合わせて使用でき、各 Teammate が自身の worktree で作業できます:
```
agent team を作成して、各 teammate が独立した worktree で作業するようにして、
ファイル衝突を回避して。
```
### サードパーティオーケストレーションツール
ネイティブの Agent Teams 以外に、コミュニティもいくつかのオーケストレーションツールを開発しています:
| ツール | 説明 |
| --------------- | ------------------------------- |
| **Gas Town** | 複数の並行 Claude セッションを管理するツール |
| **Multiclaude** | 複数のターミナルウィンドウで Claude インスタンスを実行 |
これらのツールは Agent Teams の実験的機能の代替を提供しますが、より多くの手動設定が必要です。ネイティブ Agent Teams がニーズを満たす場合は、公式機能の使用を優先することをお勧めします。
## 現在の制限事項
Agent Teams はまだ実験的機能であり、制限を理解することが重要です:
| 制限事項 | 説明 |
| ----------------------------- | --------------------------------------------------- |
| in-process teammates の復元不可 | `/resume` と `/rewind` は in-process teammates を復元しない |
| タスク状態の遅延 | Teammates がタスクの完了マークを忘れることがある |
| 終了が遅い場合がある | Teammates は現在のリクエストを完了してから終了 |
| セッションごとに1チーム | Lead は一度に1つのチームのみ管理可能 |
| ネストされたチーム不可 | Teammates は独自のチームを生成できない |
| Lead は固定 | チームを作成したセッションが Lead、移譲不可 |
| Split panes は tmux/iTerm2 が必要 | VS Code ターミナル、Windows Terminal、Ghostty は非対応 |
| Plan mode はセッションレベル | Teammate の Plan mode 状態は生成時にロックされ、セッション中に変更不可 |
## コストの考慮
Agent Teams の Token 消費はシングルセッションより大幅に高くなります:
| シナリオ | Token 消費 | コスト倍率 |
| ------------------------- | ------------ | ------------- |
| 単一 Agent セッション | 約200k tokens | 1x |
| 3人の Teammates | 約800k tokens | 約4x |
| 5人の Teammates | 約1.2M tokens | 約6x |
| 16人の Teammates(C コンパイラ事例) | 20億 tokens | $20,000 / 2週間 |
**コスト分析**:
* 各 Teammate は完全に独立した Claude インスタンスで、独自のコンテキストを持つ
* Teammates 間のコミュニケーションも Token を消費
* Lead がすべての Teammates を調整するため、追加のオーバーヘッドが発生
**いつ価値があるか**:
* ✅ 並行探索が必要なリサーチタスク
* ✅ マルチビューレビュー(セキュリティ、パフォーマンス、テスト)
* ✅ 議論によるコンセンサスが必要な意思決定
* ❌ 順次完了できるルーティンタスク
* ❌ 相互通信が不要な並行タスク(Subagent の方が経済的)
## 私の使用感
### いつ Agent Teams を使うか
私の判断基準:
1. **タスクに複数の視点が必要**:異なる領域の専門知識(セキュリティ + パフォーマンス + テスト)
2. **議論とコンセンサスが必要**:競合的仮説、アーキテクチャの決定
3. **並行探索に価値がある**:複数の実装方案の比較
並行実行だけが必要で、相互通信が不要な場合は、Subagent や Worktree の方が適しています。
### リサーチとレビューから始める
Agent Teams 初心者なら、コードを書く必要がないタスクから始めてください:PR のレビュー、技術方案のリサーチ、バグの調査。これらのタスクは明確な境界があり、並行探索の価値を示すことができ、同時に並行実装に伴う調整の課題を回避できます。
### 他の機能との組み合わせ
| 組み合わせ | 効果 |
| ---------------------- | ------------------- |
| Agent Teams + Worktree | 各 Teammate が隔離環境で作業 |
| Agent Teams + Hooks | 品質チェックとフィードバックの自動化 |
| Agent Teams + Skills | 各 Teammate が専門能力を獲得 |
## おわりに
Agent Teams は、AI支援開発の新しいパラダイムを代表しています:「1つのAIアシスタント」から「1つのAIチーム」へ。
Anthropic は16の Agent を使い、2週間、$20,000で10万行のコードの C コンパイラを書き上げました。このプロジェクトの重要な教訓は:**テスト品質が何より重要**。Agent は与えられた問題を自律的に解決するため、タスクバリデータはほぼ完璧でなければなりません。そうでなければ Agent は間違った問題を解決してしまいます。
3つの核心ポイントを覚えておいてください:
| ポイント | 説明 |
| ------ | ------------------------------------- |
| **協力** | Teammates は結果を報告するだけでなく、直接コミュニケーション可能 |
| **共有** | 共有タスクリストで作業を調整 |
| **監督** | 定期的に進捗をチェックし、方向を早期に修正 |
始め方はシンプルです:
```
agent team を作成して [あなたのタスク]
```
***
**関連記事**:
* [Claude Worktree 完全ガイド](/ja/docs/notes/claude-worktree) — Worktree と Agent Teams の連携を理解する
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) — Subagent と Agent Teams のユースケースを比較する
* [Tmux 快速入門ガイド](/ja/docs/notes/tmux-tutorial) — Tmux で複数の Agent セッションを管理する
**参考資料**:
* [Claude Code 公式ドキュメント - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**動画チュートリアル**:
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) — 7つの非技術ユースケースのデモ
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) — マルチエージェントワークフローの詳細解説
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) — チームオーケストレーションの深掘り
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) — 完全なセットアップチュートリアル
# Claude システムアーキテクチャ全解析
## はじめに
2025年9月、Anthropic は **1,830億ドル** の評価額で130億ドルの資金調達を完了し、世界第4位の非公開企業となりました。その看板製品 Claude Code は2月のリリース以来、**11.5万** のアクティブ開発者を獲得し、毎週 **1.95億行** のコードを処理し、ユーザー数は **300%** 増加しています。
さらに興味深いのは、Anthropic CEO の Dario Amodei が明かした事実です:**Claude Code のコードの90%は自分自身で書かれている**。
**なぜそんなことが可能なのか?**
AIプログラミングアシスタントが、どうやって「自分で自分を書く」ことができるのか?そのアーキテクチャ設計にはどんな独自性があり、人間の開発者を効率的に補助——さらには代替——できるのか?
答えは Claude の**モジュラーアーキテクチャ**にあります:MCP がツールを提供し、Skills が使い方を教え、Subagents が並行実行し、Hooks が制御性を確保する——これらのコンポーネントが協調動作することで、Claude は「プログラマーのように働く」能力を獲得しています。
このドキュメントでは、**このアーキテクチャの全体像を俯瞰**します——各コンポーネントの位置づけ、それらの間の連携関係、そしてクイックスタートの設定例。後続の記事で各コンポーネントの詳細を深掘りします。
## 全体アーキテクチャ概要
Claude システムは**モジュラーアーキテクチャ**設計を採用しており、各コンポーネントは機能別に分類され、**階層依存ではなく相互補完**で動作します:
**コアとなる理解**:これらのコンポーネントは階層依存ではなく**同レベルで相互補完**する拡張能力であり、ニーズに応じて自由に組み合わせることができます:
| やりたいこと | 使うもの | 一言での位置づけ |
| --------------------- | ------------- | ----------------------------------------- |
| 外部データソースやサービスに接続 | **MCP** | Claude に「手足」を装着し、データベース、API、ファイルシステムにアクセス |
| Claude に特定のワークフローを教える | **Skills** | Claude がある領域で「どうすればいいか」を知るようにする |
| 複雑なタスクを並行処理 | **Subagents** | 大きなタスクを小さなタスクに分割し、複数の Agent が同時に作業 |
| 繰り返し操作を素早くトリガー | **Commands** | ワンクリックで頻繁に使うワークフローを起動し、繰り返しの指示を省略 |
| 特定の操作が必ず実行されることを保証 | **Hooks** | Claude がどう判断しようと、このステップは必ず実行される |
***
## コアランタイム
### Agent SDK — ランタイムエンジン
Agent SDK は Claude Agent システム全体の**ランタイムコアエンジン**であり、以下を提供します:
* **メインループ (Main Loop)**:Agent のコアワークループ
* **コンテキスト管理**:Token バジェット、自動圧縮(使用率92%でトリガー)
* **ツールディスパッチ**:どのツールを使用するか、どう実行するかを決定
* **権限システム**:ツールアクセス権限の制御
Agent のコアワークモードはシンプルな**フィードバックループ**です:
```
コンテキスト収集 → 操作実行 → 作業検証 → 繰り返し
```
***
### Built-in Tools — コアツール
Claude Agent には20以上のコアツールが内蔵されており、3つのカテゴリに分かれています:
| カテゴリ | ツール | 説明 |
| ---------- | ------------------- | -------------------------- |
| **読み取り** | Read, Glob, Grep | ファイル読み取り、パターンマッチング、コンテンツ検索 |
| **操作** | Write, Edit, Bash | ファイル書き込み、編集、コマンド実行 |
| **ネットワーク** | WebSearch, WebFetch | ウェブ検索、ウェブページ取得 |
これらのツールは**デフォルトで利用可能**であり、追加設定は不要です。Claude はこれらのツールを通じてコンピュータとやり取りし、プログラマーが IDE を使うのと同じように動作します。
***
## 設定とコンテキスト
### CLAUDE.md — 永続化コンテキスト
新しい会話を始めるたびに、プロジェクトの背景、コーディング規約、アーキテクチャの約束事を繰り返し説明しなければならない……CLAUDE.md はこれらの情報を**一度設定すれば自動ロード**にしてくれます。
CLAUDE.md はプロジェクトの **AI 用 README** のようなもの——Claude にこのプロジェクトの背景知識、作業方法、約束事を伝えます。
#### 階層的オーバーライド
Claude は以下の順序で CLAUDE.md をロードし、**より具体的なものほど優先度が高く**なります:
```
Enterprise(最低)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md)(最高)
```
#### コンテンツの推奨事項
CLAUDE.md には以下のコア情報を含めるべきです:
| カテゴリ | コンテンツ例 |
| ------------ | ---------------------------------------------- |
| **技術スタック** | Next.js 14 + TypeScript, Tailwind CSS |
| **ビルドコマンド** | `npm run dev`, `npm run build`, `npm run test` |
| **コード規約** | 命名規則、Lint ツール設定 |
| **プロジェクト構造** | 主要ディレクトリの用途説明 |
**キー原則**:簡潔に保つこと。CLAUDE.md は**毎回の会話でロード**されるため、長すぎると貴重な Token を浪費します。
***
## パッケージング・配布
### Plugins — インストール可能なユニット
チームの設定が分散し、共有や標準化が難しい。各メンバーが独自の Skills、Commands、Hooks のセットを持っている……どうやって統一管理するのか?
Plugins は **Skills + Commands + Subagents + Hooks + MCP** を**インストール可能なユニット**にパッケージ化し、ワンクリック配布とチーム標準化を実現します。
#### ディレクトリ構造
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # プラグインマニフェスト(必須)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Hooks 設定
├── .mcp.json # MCP Server 設定
└── README.md # 説明ドキュメント
```
#### 設定例
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "フロントエンド開発ツールキット",
"author": "Your Team"
}
```
```bash
# インストール方法
claude plugin install github:your-org/your-plugin # GitHub から
claude plugin install /path/to/plugin # ローカルから
```
**関連リソース**
| リソース | 説明 |
| -------------------------------------------------------------------------------- | ------------------------------ |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Anthropic 公式プラグインリポジトリ |
| [wshobson/agents](https://github.com/wshobson/agents) | ⭐ 24.3k、高品質 Agent テンプレートコレクション |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | コミュニティプラグインマーケットプレイス |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | ⭐ 19.3k、厳選 Claude Code リソースリスト |
***
## 拡張機能(相互補完モジュール)
Claude システムの拡張機能は、それぞれの役割を担い協調動作する複数の**相互補完モジュール**で構成されています:
| モジュール | 機能の位置づけ | アクティベーション方法 |
| ------------- | ---------------------------- | ----------- |
| **MCP** | 外部データやサービスへの接続 (WHAT) | 設定後に利用可能 |
| **Skills** | 手続き的知識、Claude にやり方を教える (HOW) | 自動マッチング |
| **Subagents** | 独立コンテキスト、並行タスク委任 | 明示的な呼び出し |
| **Commands** | 繰り返しワークフロー | 手動 `/cmd` |
| **Hooks** | 確定的制御、イベント駆動 | 自動トリガー |
***
### MCP — 外部接続
#### 設計理念
従来の方法では、各外部データソースに**カスタム統合**が必要で、N×M の統合地獄に陥っていました。MCP は標準化されたプロトコルを提供し、**一度接続すれば、どこでも利用可能**を実現します。
MCP (Model Context Protocol) は **AIアプリケーションの USB-C インターフェース**として設計されました:
| 特性 | 説明 |
| -------------- | ----------------------------------------- |
| **オープンスタンダード** | 2024年11月発表、2025年12月に Linux Foundation に寄贈 |
| **業界採用** | OpenAI、Microsoft、Google、AWS などが採用 |
| **エコシステム規模** | 月間9,700万以上の SDK ダウンロード、数千のコミュニティサーバー |
#### アーキテクチャパターン
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, AI ツール)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (プロトコルクライアント)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (特定の機能を公開)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**ユースケース**:データベース接続、サードパーティサービス統合(GitHub、Slack、Notion)、プライベートAPIアクセス、リアルタイムデータストリーム処理。
#### 設定例
プロジェクトルートに `.mcp.json` を作成:
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — 手動ワークフロー
Slash Commands は**手動トリガー**の繰り返しワークフローを提供します。
| 特性 | 説明 |
| ---------- | ----------------------- |
| **トリガー方法** | 手動で `/command-name` を入力 |
| **保存場所** | `.claude/commands/` |
| **用途** | 繰り返しワークフロー、標準化操作 |
**例**:`.claude/commands/review.md` を作成
```markdown
現在の変更についてコードレビューを行ってください。以下に重点を置いてください:
1. コードスタイルと一貫性
2. 潜在的なパフォーマンスの問題
3. セキュリティ脆弱性
4. テストカバレッジ
```
その後 `/review` と入力すればトリガーされます。
***
### Hooks — 確定的制御
Hooks は**確定的制御**のコアです——特定の操作は必ず実行されなければならず、LLM の判断に依存できません。
| カテゴリ | イベント | トリガータイミング |
| ------------ | -------------------- | ------------------- |
| **ツール** | `PreToolUse` | ツール実行前 |
| | `PostToolUse` | ツール正常実行後 |
| | `PostToolUseFailure` | ツール実行失敗後 |
| | `PermissionRequest` | 権限リクエスト時 |
| **セッション** | `SessionStart` | セッション開始時 |
| | `SessionEnd` | セッション終了時 |
| | `Stop` | Claude がレスポンスを完了した時 |
| **サブエージェント** | `SubagentStart` | サブエージェント起動時 |
| | `SubagentStop` | サブエージェント停止時 |
| **その他** | `UserPromptSubmit` | ユーザーがプロンプトを送信した後 |
| | `Notification` | 通知イベント |
| | `PreCompact` | コンテキスト圧縮前 |
#### 設定例
TypeScript ファイルの自動フォーマット:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### 実践:自律ループ
**Ralph Wiggum** は Anthropic 公式プラグインで、Stop hook を利用して自律的なイテレーションループを実現します:
```bash
/ralph-loop "TODO API を実装、CRUD とテストを含む" --max-iterations 20
```
**動作原理**:Stop hook が Claude の終了をインターセプト → 元のプロンプトを再注入 → タスクが完了するか最大イテレーション数に達するまでイテレーションを継続。
**適用シナリオ**:複数ラウンドのイテレーションが必要なタスク(テストパス、コードリファクタリング)、自動検証手段があるタスク。
***
### Subagents — タスク委任と並行実行
#### 設計理念
単一の Agent が直面する課題:コンテキストウィンドウの制限、並行処理不可、責任の不明確さ。Subagents は **Orchestrator-Worker** アーキテクチャパターンでこれらの問題を解決します:
```
メイン Agent (Orchestrator)
├── ユーザーリクエストの分析
├── 計画の策定
├── タスクの分解
└── 専門化されたサブエージェントの生成
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - 並行実行)
│ コード │ テスト │ ドキュメント │
└────────┴────────┴────────┘
↓
結果の集約 → メインエージェントが総合出力
```
#### コア特性
| 特性 | 説明 |
| ------------ | ------------------------------- |
| **コンテキスト分離** | 各 Subagent は独立したコンテキストを持ち、汚染を防止 |
| **タスク専門化** | カスタムシステムプロンプトで専用の役割を定義 |
| **ツール権限制御** | Subagent が使用できるツールを特定のものに制限可能 |
| **並行実行** | 複数の Subagent が同時に動作 |
**パフォーマンスデータ**:マルチエージェントシステムはシングルエージェントより90.2%高いパフォーマンス、並列化でリサーチ時間を90%削減可能(Token 消費は約15倍だが、複雑なタスクでは価値あり)。
#### 設定例
`.claude/agents/` に Markdown ファイルを作成:
```markdown
---
name: Code Reviewer
description: コードレビュー専門のサブエージェント
tools:
- Read
- Grep
- Glob
---
あなたはシニアコードレビュー専門家です。以下に重点を置いてください:
1. コード品質と保守性
2. 潜在的なバグとエッジケース
3. パフォーマンス最適化の機会
4. セキュリティ脆弱性
```
***
### Skills — 手続き的知識
#### 設計理念
Skills はAIへの**再利用可能なワークブック**——モジュール化されたナレッジパッケージで、Claude が必要に応じて動的にロードできます。コア設計原則は**段階的開示 (Progressive Disclosure)** です:
```
📚 Skills ワークブック
│
├─ 📋 目次 ────────────── 【メタデータ層】起動時にプリロード (~30-50 tokens)
│ name: "weekly-report"
│ description: "標準化された週報を生成"
│
├─ 📖 本文章節 ─────────── 【コアドキュメント層】関連時にロード (~数百-数千 tokens)
│ # Weekly Report Generator
│ ## Instructions
│ 以下の構成で週報を生成...
│
└─ 📎 付録 ────────────── 【参照リソース層】必要時にロード
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills**:MCP は Claude にツールへのアクセス能力を与え(WHAT)、Skills は Claude にこれらのツールを効果的に使う方法を教えます(HOW)。
#### コア優位性
| 優位性 | 説明 |
| --------------- | ------------------------------------------- |
| **Token 効率** | メタデータはわずか 30-50 tokens、数十の Skills を同時に有効化可能 |
| **自動アクティベーション** | タスクコンテキストに応じて自動マッチング、手動トリガー不要 |
| **コンポーザブル** | 複数の Skills が自動的に協調動作 |
| **ポータブル** | Claude.ai、Claude Code、API で一貫した体験 |
#### 設定例
`.claude/skills/` にディレクトリを作成:
```
my-skill/
├── SKILL.md # コア指示(必須)
├── scripts/ # 実行可能スクリプト(任意)
└── references/ # 参考資料(任意)
```
SKILL.md のコア構造:
```yaml
---
name: code-review # Skill 名
description: コードレビュー、品質とセキュリティのチェック # 簡潔な説明(自動マッチングに使用)
---
# Code Review Skill
## Instructions
[具体的な手順の説明...]
## Output Format
[出力フォーマットの要件...]
```
**ポイント**:frontmatter の `description` は自動マッチングに使用されるため、簡潔かつ正確に保ちましょう。
***
## 公式参考リンク
**設計理念**
| リソース | 説明 |
| --------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Agent アーキテクチャの綱領的記事 |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Agent SDK エンジニアリング実践 |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Skills 設計理念 |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | MCP 発表公告 |
**公式ドキュメント**
| リソース | 説明 |
| ----------------------------------------------------------------------------- | -------------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Skills と他のコンポーネントの比較 |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | CLAUDE.md 使用ガイド |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Subagents 公式ドキュメント |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Hooks 公式ドキュメント |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | MCP 公式ドキュメント |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | MCP プロトコル仕様 |
**深層分析**
| リソース | 説明 |
| --------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | アーキテクチャ図を含む |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Skills 原理の詳細分析 |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Claude Code 構築の内幕 |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Skills と MCP の技術比較 |
***
## さらに読む
Skills のコンセプトと実践をさらに深く理解したい方は、以下をご覧ください:
* [Claude Skills とは](/ja/docs/notes/claude-skills/concept) — Skills コア原理の詳解
* [Claude Skills 実戦ガイド](/ja/docs/notes/claude-skills/practice) — はじめての Skill を作成しよう
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) — サブエージェントの使い方とカスタマイズ
* [GSD 深層解析](/ja/docs/notes/gsd/concept) — コンテキストエンジニアリングに基づくAIプログラミングシステム
* [私の Claude Code ベストプラクティス](/ja/blog/claude-code-best-practices) — Claude Code の日常的な使用テクニック
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# Claude Worktree 完全ガイド
## はじめに
Claude Code で複雑なタスクを処理する際、こんな困りごとに遭遇したことはありませんか?3つの独立したタスクを処理する必要があるのに、同じディレクトリで複数の Claude インスタンスを実行するとコードが衝突する——一つの Agent がファイルを修正中に、もう一つも同じファイルを変更し、最後のマージ時に大混乱になる。
2025年2月、Anthropic は `--worktree` コマンドをリリースし、この状況を根本的に変えました。これで、3つのターミナルでそれぞれ `claude -w feature-1`、`claude -w feature-2`、`claude -w bugfix-1` を実行でき、3つの Agent がそれぞれ隔離された環境で作業し、互いに干渉しません。
## Worktree を理解する
あなたが建築家で、同時に3つの異なる部屋を設計していると想像してください。従来の方法は同じ図面上で描くことですが、修正を繰り返すと混乱しやすくなります。Worktree のアプローチは、3枚の独立した図面を提供し、各図面を1つの部屋の設計に専用し、最後にメイン図面に統合するというものです。
技術的に言えば、Worktree は Git のネイティブ機能です。Claude Code の `--worktree` コマンドはこの機能をよりシンプルに使えるようにラップしています——1つのコマンドで隔離環境の作成、Claude インスタンスの起動、完了後の自動クリーンアップが行えます。
### なぜ直接複数回 Clone しないのか
こう思うかもしれません:なぜ直接コードを複数回 clone しないのか?
| 方法 | ディスク使用量 | 同期の難しさ | クリーンアップの複雑さ |
| ------------------- | ----------------------- | ----------------- | --------------------- |
| 複数回 Clone | 毎回完全なリポジトリ | 手動で pull/push が必要 | 手動でディレクトリを削除する必要 |
| Git Worktree | ワーキングファイルのみコピー、.git を共有 | 自動的に履歴を共有 | `git worktree remove` |
| Claude `--worktree` | ワーキングファイルのみコピー、.git を共有 | 自動的に履歴を共有 | 終了時に自動クリーンアップ |
Worktree は同じ `.git` データベースを共有し、すべてのコミット履歴やブランチ情報が共有されます。これは、1つの worktree で作成した commit が他の worktree ですぐに表示されることを意味します。
### いつ Worktree を使うべきか
作業を始める前に、タスクが worktree に適しているかどうかを判断しましょう。
経験則:**タスクが30分以上かかるなら、worktree の使用を検討しましょう**。短いタスクでは worktree はかえって時間の無駄です——環境の作成、依存関係のインストール、最後のマージを合計すると、タスク自体より長くかかる可能性があります。しかし、深い作業が必要なタスクでは、worktree の隔離性は非常に価値があります。
| Worktree に適している | あまり適していない |
| ------------------- | ------------------ |
| 独立した機能開発 | 10分で完了する小さな変更 |
| 異なるモジュールの並行リファクタリング | 頻繁なインタラクションが必要なタスク |
| 長時間実行されるタスク | 進行中の他の変更に強く依存 |
| 隔離テストが必要な実験的変更 | シンプルなバグ修正 |
### 前提条件
worktree を使用する前に、以下の条件を満たしていることを確認してください:
| 条件 | 説明 |
| ------------------- | ---------------------------------------- |
| Git が初期化済み | Git リポジトリディレクトリ内にいる必要(`.git` ディレクトリがある) |
| 少なくとも1つの commit がある | 空のリポジトリでは worktree を作成できない |
| リモートブランチが利用可能 | デフォルトでリモートブランチからチェックアウト(例:`origin/main`) |
## 完全なワークフロー:作成からクリーンアップまで
以下、実際の開発順序に沿って、Worktree の作成から最終的なクリーンアップまでの完全なフローを見ていきましょう。
### ステップ1:Worktree の作成
#### リモートデフォルトブランチからの作成
`-w` または `--worktree` パラメータで Claude を起動します:
```bash
# "feature-auth" という名前の worktree を作成し Claude を起動
claude -w feature-auth
# ランダム名を自動生成(例:"bright-running-fox")
claude -w
```
このコマンドは実際に4つのことを行います:
1. `/.claude/worktrees/feature-auth/` に新しい作業ディレクトリを作成
2. `worktree-feature-auth` という名前の新しいブランチを作成
3. リモートデフォルトブランチ(例:`origin/main` や `origin/master`)からコードをチェックアウト——**注意:現在いるブランチではありません**
4. 新しいディレクトリで Claude Code を起動
すべての worktree は `.claude/worktrees/` ディレクトリ下にあります:
```
your-project/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← 最初の worktree
│ ├── bugfix-123/ ← 2番目の worktree
│ └── refactor-api/ ← 3番目の worktree
├── src/
└── package.json
```
このパスを `.gitignore` に追加することをお勧めします:
```bash
# .gitignore
.claude/worktrees/
```
#### 現在の/特定のブランチからの作成
`-w` は常にリモートデフォルトブランチからチェックアウトし、現在はベースブランチの指定をサポートしていません。現在のブランチ(または特定のブランチ)をベースに worktree を作成したい場合は、3つの方法があります:
**方法1:手動で Git を使って作成**
```bash
# 現在の HEAD をベースに worktree を作成
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# または特定のブランチをベースに
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# そのディレクトリで Claude を起動
cd .claude/worktrees/my-feature && claude
```
この方法は完全な制御を提供します——任意のブランチ、任意の commit をベースに worktree を作成でき、作業ディレクトリは最初から正しいブランチにあります。公式ドキュメントでも推奨されています:「ブランチと場所のより細かい制御が必要な場合は、Git で直接 worktree を作成し、そのディレクトリで Claude を実行してください。」
**方法2:会話中に作成(推奨)**
既存の Claude セッション内で、直接 Claude に worktree を作成させます:
```
> 現在のブランチから worktree を開いて
> start a worktree
```
`-w` コマンドとは異なり、会話中に作成された worktree は**自動的に現在のブランチをベース**にし、リモートデフォルトブランチではありません。Claude は自動的に worktree の作成と切り替えを完了し、Git コマンドを手動で操作する必要はありません。すでに feature ブランチで作業中なら、これが最も便利な方法です——一言で現在のブランチベースの隔離環境を作れます。
**方法3:先に `-w` で作成し、会話中にブランチを切り替え**
先に `claude -w` で worktree を作成し、会話に入ってから Claude にターゲットブランチへの切り替えを指示します。この方法の欠点は、まずリモートデフォルトブランチをプルしてから切り替えるため、一手間多くなることです。最初の2つの方法ほどクリーンではありません。また、ターゲットブランチが他の worktree で既に使用されている場合、ブランチの衝突に遭遇します:
スクリーンショットに示されているように、Claude はブランチの衝突を検出し、2つの選択肢を提示します:メインディレクトリに戻って操作するか、ターゲットブランチをベースに新しい作業ブランチを作成するか。最終的には動作しますが、方法1と方法2ほどダイレクトではありません。
**応用:Makefile でワンコマンド化**
現在のブランチから worktree を頻繁に作成する場合は、プロジェクトルートの `Makefile` にショートカットコマンドを追加し、作成 + エディタ起動 + Claude 起動を確定的なパイプラインとして連結できます:
```makefile
# 現在のブランチから worktree を作成し、開発環境を起動
# 使い方: make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "使い方: make worktree name="; \
echo "例: make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ $$(git branch --show-current) から worktree を作成: $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ 環境を初期化"
cd .claude/worktrees/$(name) && npm install
@echo "→ Zed で開く"
zed .claude/worktrees/$(name)
@echo "→ Claude を起動"
cd .claude/worktrees/$(name) && claude
```
使い方は非常にシンプルです:
```bash
# 現在のブランチベースで worktree を作成、Zed で開き、Claude を起動
make worktree name=fix-login-bug
# 複数の並行タスクを開始
make worktree name=feature-search
make worktree name=refactor-api
```
複数の Git/cd/claude コマンドを手動で入力するのと比べ、`make worktree name=xxx` は1行で済み、毎回実行するフローが完全に一貫します——ステップを忘れることも、パスを間違えることもありません。注意点として、Makefile はネイティブの `git worktree add` を使用するため、Claude Code の `WorktreeCreate` Hook はトリガーされません(このフックは `claude -w` または会話中に worktree を作成する場合のみ有効)。そのため、環境初期化ステップ(依存関係のインストール、`.env` のコピーなど)は上記の例の `npm install` のように、Makefile に直接記述する必要があります。
### ステップ2:環境の初期化
Worktree の作成が完了したら、最初にやるべきことは開発環境の初期化です。各新しい worktree は独立したディレクトリであり、`node_modules`、仮想環境、`.env` ファイルなどは自動的に持ち込まれません。
Claude Code は `WorktreeCreate` Hook を提供して環境セットアップを自動化します:
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
これにより worktree 作成時に依存関係が自動インストールされ、環境変数ファイルが自動コピーされます。一般的な初期化ステップ:
| プロジェクトタイプ | 初期化コマンド |
| --------- | ------------------------------------------------- |
| Node.js | `npm install` または `yarn` |
| Python | `pip install -r requirements.txt` または仮想環境のアクティベート |
| Go | `go mod download` |
| 汎用 | `.env` ファイルのコピー、環境変数の設定 |
Hook を設定していない場合は、各 worktree セッションの開始時に `/init` を実行して、Claude が現在の作業ディレクトリのコンテキストを正しく理解し、プロジェクト構造と CLAUDE.md 設定を再読み込みするようにできます。
### ステップ3:コミットとマージ
環境が整い、開発が完了したら、次のステップは変更をターゲットブランチにマージすることです。
**main ブランチへのマージ**
最も一般的なケース——worktree が `origin/main` から分岐し、変更も `main` にマージする。worktree の Claude セッション内で直接言います:
```
> すべての変更をコミットして、リモートにプッシュして、main への PR を作成して
```
Claude が自動的に commit → push → `gh pr create` の全フローを完了します。
**feature ブランチへのマージ**
`feature-x` ブランチで開発していて、worktree の変更を `main` ではなく `feature-x` にマージする必要がある場合:
```
> 変更をコミットしてプッシュして、feature-x ブランチへの PR を作成して
```
Claude は `gh pr create --base feature-x` を実行し、feature ブランチを指す PR を直接作成します。
worktree セッションを終了した後(worktree を保持する選択をした場合)、メインディレクトリに戻って Claude を起動することもできます:
```
> worktree-my-task ブランチの変更を現在のブランチにマージして
```
worktree 内の一部のコミットが不要な場合は、選択的に cherry-pick できます:
```
> worktree-my-task ブランチのコミット履歴を確認して、認証モジュールに関するコミットを現在のブランチに cherry-pick して
```
> **ヒント**:すべての worktree は同じ `.git` データベースを共有しているため、worktree で作成したコミットはメインディレクトリですぐに表示されます。追加の push/pull 操作は不要です。
### ステップ4:終了とクリーンアップ
コードのマージが完了したら、worktree セッションを終了できます。
worktree セッション終了時、Claude は状況に応じて自動的に処理します:
| 状態 | 処理方法 |
| ------------- | ------------------- |
| **変更なし** | worktree とブランチを自動削除 |
| **変更やコミットあり** | 保持するか削除するかの選択を促す |
保持された worktree はそのまま残り、後で作業を続けるのに便利です。
`WorktreeRemove` Hook を設定してクリーンアップを自動化することもできます:
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**手動管理コマンド**
worktree を手動で管理する必要がある場合は、標準の Git コマンドを使用できます:
```bash
# すべての worktree を一覧表示
git worktree list
# 特定の worktree を手動削除
git worktree remove .claude/worktrees/feature-auth
# 古い worktree 参照をクリーンアップ
git worktree prune
```
> **注意**:worktree ディレクトリを直接 `rm -rf` で削除しないでください。正しい方法は `git worktree remove` を使用することです。誤って削除した場合は、`git worktree prune` を実行して残留する参照をクリーンアップしてください。
## 並行開発モード
基本的なワークフローをマスターしたら、worktree を活用して並行開発を実現する方法を見てみましょう。
### マルチターミナル並行
最も一般的な使い方は、複数のターミナルタブで同時に実行することです:
```bash
# ターミナル1:ユーザー認証機能の処理
claude -w feature-auth
# ターミナル2:決済バグの修正
claude -w bugfix-payment
# ターミナル3:API モジュールのリファクタリング
claude -w refactor-api
```
各 Claude インスタンスは自身の worktree で動作し、変更は互いに影響しません。以下のことができます:
* 1つのターミナルで Claude に新機能を開発させる
* もう1つのターミナルで Claude にバグを修正させる
* 3番目のターミナルで自分のコードレビューを続ける
### 競合的実装
効率的な使い方の1つは、複数の Agent に同じ機能を独立して実装させることです:
```bash
# 3つのターミナルでそれぞれ実行
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
同じ要件説明を与え、各自に実装させます。最後に3つの方案を比較し、最良のものをマージします。これは LLM の非決定性を活用しています——同じ入力から異なる出力が生まれ、時には2番目のバージョンの方が良いこともあります。
UI デザイン探索にもこのモードは非常に適しています。アプリケーションのインターフェースを再デザインしたいが、どのスタイルが良いか分からない場合:
```bash
# 3つの Agent にそれぞれ異なるスタイルで実装させる
claude -w ui-minimal # ミニマリストスタイル
claude -w ui-colorful # 鮮やかな配色
claude -w ui-glassmorphism # すりガラスモーフィズムスタイル
```
完了後、3つのバージョンの開発サーバーを同時に実行(異なるポート)し、並べて効果を比較し、最も満足できる方案をメインブランチにマージします——従来の「1版作って、効果を見て、不満なら再修正」のフローよりはるかに効率的です。
### Subagent の隔離
Worktree はメインの Claude インスタンスだけでなく、Subagent にも適用できます。カスタム Subagent の frontmatter に `isolation: worktree` を追加します:
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
会話中に直接 Claude に伝えることもできます:
```
> agent を隔離するために worktree を使って
> use worktrees for your agents
```
Subagent が worktree 隔離で設定されている場合:
```
メイン Agent(メインディレクトリ)
│
├── Migration Agent 1 を起動 ──→ worktree-migration-1/
│ └── src/auth/ ディレクトリを処理
│
├── Migration Agent 2 を起動 ──→ worktree-migration-2/
│ └── src/api/ ディレクトリを処理
│
└── Migration Agent 3 を起動 ──→ worktree-migration-3/
└── src/utils/ ディレクトリを処理
```
各 Subagent は自身の worktree で独立して動作し、互いに干渉しません。完了後、worktree は自動的にクリーンアップされます(未コミットの変更がない場合)。
### Tmux と IDE の組み合わせ
`--tmux` パラメータと組み合わせることで、新しい Tmux セッション内で自動起動でき、ターミナルを閉じても Claude はバックグラウンドで実行し続けます:
```bash
claude -w feature-auth --tmux
```
VS Code や Cursor を使用している場合、ソースコントロールパネルがすべての worktree を自動認識します——メインリポジトリが1つの repo として表示され、各 worktree が独立した repo として表示され、IDE 内で直接切り替え、コミット、プッシュができます。Worktree は [Ralph ループ](/ja/docs/notes/ralph-wiggum/concept)と組み合わせることもでき、各 Ralph ループが独自の worktree で実行され、ループが失敗してもメインブランチに影響しません。
## 注意事項とベストプラクティス
### よくある落とし穴
1. **ブランチの出自を間違えやすい**:`-w` で作成された worktree は**リモートデフォルトブランチ**からチェックアウトされ、現在いるブランチではありません。`feature-x` ブランチで `claude -w my-task` を実行すると、新しい worktree のコードは `origin/main` から取得され、`feature-x` の変更は含まれません。現在のブランチベースで作業したい場合は、[現在の/特定のブランチからの作成](#現在の特定のブランチからの作成)を参照してください。
2. **未コミットの変更は持ち込まれない**:worktree 作成時、メインディレクトリ内の未ステージングまたは未コミットの変更は新しい worktree に表示されません。Worktree はコミット履歴のみに基づいて作成されるため、重要な変更は必ずコミットしておきましょう。
3. **同一ブランチを複数の Worktree で使用できない**:Git は2つの worktree が同時に同じブランチをチェックアウトすることを許可しません。メインディレクトリで既に `feature-x` ブランチにいる場合、worktree で `feature-x` をチェックアウトしようとするとエラーになります。各 worktree は異なるブランチにある必要があります。
4. **環境の再初期化が必要**:各新しい worktree には `node_modules` などのランタイム依存関係が含まれないため、`WorktreeCreate` Hook を設定して自動化することをお勧めします([ステップ2:環境の初期化](#ステップ2環境の初期化)を参照)。
### 使用上のアドバイス
欲張りすぎないこと。技術的には多くの worktree を開けますが、各 Claude インスタンスは API クォータを消費し、あまりに多くの並行タスクは追跡が難しくなり、最終的なマージ時のコンフリクトもより複雑になります。
**命名規則**:良い命名習慣を身につけると、後の管理が楽になります:
```bash
# 良い命名
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# 悪い命名
claude -w test
claude -w temp
claude -w 1
```
## 非 Git バージョン管理
SVN、Perforce、Mercurial を使用している場合は、`WorktreeCreate` と `WorktreeRemove` Hook を設定して同様の隔離効果を実現できます。これらの Hook を設定すると、`--worktree` 使用時にデフォルトの Git 動作の代わりにカスタムコマンドが呼び出されます。
## おわりに
Worktree は Claude Code チームが毎日使用している機能で、Boris Cherny は「ナンバーワンの生産性テクニック」と呼んでいます。核心的な価値はシンプルです:**複数の Agent が互いに干渉することなく並行して作業できるようにすること**。
始め方はシンプルです:
```bash
claude -w your-task-name
```
***
**関連記事**:
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) — Subagent と Worktree の連携を理解する
* [Ralph Wiggum 深層解析](/ja/docs/notes/ralph-wiggum/concept) — AIプログラミング効率を向上させるもう1つの方法
* [Claude システムアーキテクチャ全解析](/ja/docs/notes/claude-architecture) — アーキテクチャ全体における Worktree の位置を理解する
**参考資料**:
* [Claude Code 公式ドキュメント - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Boris Cherny の Worktree リリース告知](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Git Worktree 公式ドキュメント](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**動画チュートリアル**:
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — worktree 完全ワークフローの実践デモ
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — 手動 worktree 作成の方法
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — ネイティブ worktree 機能の詳細解説
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — クイックスタートチュートリアル、Subagent の使い方を含む
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — マルチ worktree 並行開発の利点
# ドキュメント
ドキュメントセンターへようこそ。ここでは、私が整理した Claude Code 関連の技術ドキュメントとチュートリアルを掲載しています。
# Tmux 快速入門ガイド
## はじめに
Claude Code の Agent Teams を使ったことがある、あるいは複数の Claude インスタンスを同時に実行したいなら、Tmux はほぼ必須のツールです。1つのターミナルウィンドウで複数のセッションを実行でき、ターミナルを閉じてもセッションはバックグラウンドで動き続け、さらに Claude は Tmux 内で複数の Agent を自動的に生成・管理できます。
このチュートリアルは Claude Code ユーザー向けに設計されており、Tmux の基礎知識と Claude Code との統合テクニックの両方をカバーしています。
## Tmux を理解する
Tmux の3つのコア概念:
```
┌─────────────────────────────────────────────────────────┐
│ Session(セッション) │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window(ウィンドウ) ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| 概念 | たとえ | 説明 |
| ----------- | ------- | ----------------------- |
| **Session** | ワークスペース | 最上位のコンテナ、接続を切断しても動作し続ける |
| **Window** | ブラウザタブ | 1つのセッションに複数のウィンドウを含められる |
| **Pane** | 画面分割 | 1つのウィンドウを複数のペインに分割できる |
## インストールと基本
### Tmux のインストール
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
インストールの確認:
```bash
tmux -V
# 出力例:tmux 3.6a
```
### プレフィックスキー
Tmux のすべてのコマンドは**プレフィックスキー**で始まり、デフォルトは `Ctrl+B` です。
コマンドの入力方法:
1. `Ctrl+B` を押す(離さない)
2. 離した後、コマンドキーを押す
例えば、ウィンドウの分割:`Ctrl+B` の後に `%` を押す
## よく使うコマンドのクイックリファレンス
### セッション管理
| コマンド | 説明 |
| --------------------------- | -------------------------- |
| `tmux` | 新しいセッションを作成 |
| `tmux new -s name` | 名前付きセッションを作成 |
| `tmux ls` | すべてのセッションを一覧表示 |
| `tmux attach -t name` | セッションに接続 |
| `tmux kill-session -t name` | セッションを終了 |
| `Ctrl+B d` | 現在のセッションをデタッチ(バックグラウンドで実行) |
### ウィンドウ管理
| ショートカット | 説明 |
| ------------ | -------------- |
| `Ctrl+B c` | 新しいウィンドウを作成 |
| `Ctrl+B n` | 次のウィンドウ |
| `Ctrl+B p` | 前のウィンドウ |
| `Ctrl+B 0-9` | 指定のウィンドウに切り替え |
| `Ctrl+B ,` | 現在のウィンドウの名前を変更 |
| `Ctrl+B &` | 現在のウィンドウを閉じる |
### ペイン管理
| ショートカット | 説明 |
| ------------- | ---------- |
| `Ctrl+B %` | 垂直分割(左右) |
| `Ctrl+B "` | 水平分割(上下) |
| `Ctrl+B 矢印キー` | ペイン間を移動 |
| `Ctrl+B x` | 現在のペインを閉じる |
| `Ctrl+B z` | ペインを最大化/復元 |
| `Ctrl+B {` | ペインを左に移動 |
| `Ctrl+B }` | ペインを右に移動 |
### その他のよく使うもの
| ショートカット | 説明 |
| ---------- | ------------------ |
| `Ctrl+B [` | コピーモードに入る(スクロール可能) |
| `q` | コピーモードを終了 |
| `Ctrl+B ?` | すべてのショートカットキーを表示 |
## Claude Code との統合
### なぜ Claude Code に Tmux が必要なのか
1. **Agent Teams の Split-pane モード**:各 Teammate が独立したペインに表示
2. **バックグラウンド実行**:ターミナルを閉じてもタスクが実行し続ける
3. **セッションの永続化**:切断後に再接続で完全なコンテキストを復元
4. **マルチインスタンス管理**:複数の Claude セッションを同時に実行
### 基本的な使い方:Claude をバックグラウンドで実行
```bash
# tmux 内で Claude を起動
tmux new -s claude-work
claude
# セッションをデタッチ(Claude は動き続ける)
# Ctrl+B d
# 後で再接続
tmux attach -t claude-work
```
### --tmux パラメータの使用
Claude Code は Tmux 統合をネイティブサポートしています:
```bash
# 新しい tmux セッション内で Claude を起動
claude --tmux
# worktree と組み合わせて使用
claude -w feature-auth --tmux
```
これにより自動的に:
1. 新しい tmux セッションを作成
2. その中で Claude Code を起動
3. セッション名を `claude-{ランダムID}` とする
### Agent Teams の Tmux モード
Agent Teams は split-pane 表示モードを使用でき、各 Teammate が独立したペインで動作します:
```json
// settings.json
{
"teammateMode": "tmux"
}
```
またはコマンドラインで:
```bash
claude --teammate-mode tmux
```
効果:
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## 実用的な設定
### 推奨の \~/.tmux.conf
`~/.tmux.conf` を作成または編集:
```bash
# Ctrl+A をプレフィックスキーとして使用(押しやすい)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# マウスサポートを有効化
set -g mouse on
# 履歴バッファを増やす(Claude は大量に出力する)
set -g history-limit 50000
# vim スタイルのペインナビゲーション
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# より直感的な分割ショートカット
bind | split-window -h
bind - split-window -v
# 設定の即時リロード
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# 256色サポート
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# ウィンドウ番号を1から開始(0は遠すぎる)
set -g base-index 1
setw -g pane-base-index 1
# ステータスバーの最適化
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
設定のリロード:
```bash
tmux source-file ~/.tmux.conf
```
### Claude Code 専用設定
Claude Code 向けの最適化設定:
```bash
# Claude セッションポップアップのショートカット
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
この設定の効果:
1. `Ctrl+A y` で Claude ポップアップを開く
2. 各ディレクトリに独立した Claude セッション
3. ポップアップを閉じてもセッションは動作し続ける
4. 再度開くと前の会話を復元
## よくあるワークフロー
### ワークフロー1:マルチプロジェクト並行
```bash
# 各プロジェクトに独立したセッションを作成
tmux new -s project-a
# 中で Claude を起動
claude -w feature-x
# デタッチ後、別のセッションを作成
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# セッション間を切り替え
tmux switch -t project-a
tmux switch -t project-b
# またはすべてのセッションを一覧から選択
# Ctrl+B s
```
### ワークフロー2:開発ダッシュボード
マルチペインの開発環境を作成:
```bash
# セッションを作成
tmux new -s dev
# 3つのペインに分割
# Ctrl+B %(垂直分割)
# Ctrl+B "(右側を水平分割)
# ペインレイアウト:
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# 最初のペインで Claude を実行
claude
# 2番目のペインに切り替え(Ctrl+B 右矢印)
tail -f logs/app.log
# 3番目のペインに切り替え
npm test -- --watch
```
### ワークフロー3:リモート開発
Tmux の最も強力な機能はセッションの永続化であり、SSH リモート開発に特に適しています:
```bash
# リモートサーバーに接続
ssh user@server
# tmux セッションを作成
tmux new -s remote-claude
# Claude を起動
claude
# SSH 接続を切断(Claude は動き続ける)
# Ctrl+B d
exit
# 後で再接続
ssh user@server
tmux attach -t remote-claude
# Claude セッションが完全に復元
```
### ワークフロー4:Agent Teams の監視
tmux で Agent Teams のすべての Teammates を監視:
```bash
# Claude を起動して tmux モードを使用
claude --teammate-mode tmux
# Agent Team を作成
# "agent team を作成してコードをレビューして..."
# この時点で画面が自動分割され、各 Teammate が1つのペイン
# 異なるペインをクリックして対応する Teammate と直接コミュニケーション可能
```
## トラブルシューティング
### よくある問題
| 問題 | 解決策 |
| ----------- | ------------------------- |
| 色の表示が正しくない | `TERM=xterm-256color` を確認 |
| マウスが動作しない | 設定に `set -g mouse on` を追加 |
| コピー&ペーストの問題 | コピーモードで `Enter` を使用してコピー |
| セッションが消えた | `tmux ls` を確認、システム再起動の可能性 |
### 孤立セッションのクリーンアップ
Claude Code は時々、未クリーンアップの tmux セッションを残すことがあります:
```bash
# すべてのセッションを一覧表示
tmux ls
# 特定のセッションを終了
tmux kill-session -t session-name
# すべてのセッションを終了(注意!)
tmux kill-server
```
### iTerm2 ユーザー
macOS の iTerm2 を使用している場合は、ネイティブ統合を使用できます:
```bash
# iTerm2 の tmux 統合モードを使用
tmux -CC
# または Claude Code で
claude --teammate-mode tmux
```
iTerm2 は tmux ペインを自動的にネイティブタブとスプリットスクリーンに変換します。
## 私の使用感
### いつ Tmux を使うか
| シナリオ | Tmux は必要か |
| ------------------ | --------- |
| シンプルな単発の Claude 会話 | 不要 |
| 長時間実行されるタスク | 必要 |
| Agent Teams | 強く推奨 |
| リモート開発 | 必須 |
| マルチプロジェクト並行 | 推奨 |
### 最小構成
設定をいじりたくなければ、これらのコマンドだけ覚えてください:
```bash
# セッションを作成
tmux new -s work
# デタッチ(バックグラウンド実行)
Ctrl+B d
# 再接続
tmux attach -t work
# ペインを分割
Ctrl+B % # 左右分割
Ctrl+B " # 上下分割
# ペインを切り替え
Ctrl+B 矢印キー
```
### Claude Code との最良の組み合わせ
1. **Worktree + Tmux**:各 worktree を独立した tmux セッション内で
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux**:すべての Teammates を可視化して管理
```bash
claude --teammate-mode tmux
```
3. **長時間タスク + デタッチ**:起動後にデタッチし、後でチェック
```bash
# 起動
tmux new -s migration
claude
# "データベースマイグレーションを実行して..."
# Ctrl+B d
# 数時間後
tmux attach -t migration
```
## おわりに
Tmux は Claude Code を効率的に使うための重要なツールであり、特に以下のシナリオで威力を発揮します:
| ポイント | 説明 |
| ------- | --------------------------- |
| **永続化** | 接続を切断してもセッションが失われない |
| **並行** | 複数の Claude インスタンスを同時に管理 |
| **可視化** | Agent Teams の split-pane 表示 |
3つのコアコマンドで始められます:
* `tmux new -s name` セッション作成
* `Ctrl+B d` セッションのデタッチ
* `tmux attach -t name` 再接続
***
**関連記事**:
* [Claude Agent Teams 完全ガイド](/ja/docs/notes/claude-agent-teams) — Agent Teams には Tmux の split-pane モードサポートが必要
* [Claude Worktree 完全ガイド](/ja/docs/notes/claude-worktree) — Worktree は Tmux と組み合わせてバックグラウンドで実行可能
**参考資料**:
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**動画チュートリアル**:
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) — Tmux 基礎入門
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) — Claude Code との統合ワークフロー
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) — 高度な設定テクニック
# MVPスプリント:2週間でコア機能を完成させる
これはテスト記事です。
# Apple デベロッパーアカウントの登録方法
自分で開発したアプリを App Store に公開するには、まず Apple Developer Program(Apple 開発者プログラム)への登録が必要です。年額 ¥688($99)で、すべての iOS 個人開発者にとって避けられない投資です。
この記事では、アカウントの種類の違い、登録前に準備すべきこと、そして完全な登録手順をご紹介します。
## アカウントの種類比較
Apple Developer Program には 3 つのアカウントタイプがあり、それぞれ異なる開発シーンに適しています。
| 特性 | 個人アカウント | 組織アカウント | エンタープライズアカウント |
| -------------- | ----------------- | ---------------- | ------------- |
| 年額 | ¥688($99) | ¥688($99) | ¥1,988($299) |
| App Store への公開 | ✅ | ✅ | ❌(社内配布のみ) |
| デベロッパー名の表示 | 個人の氏名 | 組織名・会社名 | 組織名 |
| チームメンバー管理 | ❌ | ✅ | ✅ |
| D-U-N-S 番号 | 不要 | 必要 | 必要 |
| 審査期間 | 比較的早い(通常 48 時間以内) | やや遅い(組織情報の確認が必要) | やや遅い |
| 対象者 | 個人開発者・個人 | 法人・スタジオ | 大企業の社内アプリ |
**個人開発者の方へ**:個人で開発される場合は、**個人アカウント**を選べば問題ありません。手続きが最もシンプルで、審査も最速、機能も十分です。App Store に表示されるデベロッパー名は本名になります。
## 登録前の準備
### 必須条件
登録を始める前に、以下の準備が整っていることを確認してください。
* **Apple ID**:まだお持ちでない場合は、[appleid.apple.com](https://appleid.apple.com) で作成しましょう。普段お使いのメールアドレスでの登録をおすすめします。開発関連の通知はすべてこのメールアドレスに届きます。
* **2 ファクタ認証**:Apple ID で 2 ファクタ認証(Two-Factor Authentication)を有効にする必要があります。iPhone で「設定 → Apple ID → サインインとセキュリティ → 2 ファクタ認証」から有効にできます。
* **Apple デバイス**:登録時の本人確認は iPhone または iPad で行う必要があり、Apple Developer App のダウンロードが必要です。
### 組織アカウントの追加要件
組織アカウントを登録する場合は、以下も必要です。
* **D-U-N-S 番号**:事前にダンアンドブラッドストリートの公式サイトで申請してください。審査に 5〜14 営業日かかります。
* **法人資格**:登録者が組織の法人代表者または正式に委任された代理人である必要があります。
* **組織情報**:登録住所、法人代表者の氏名、連絡先などが必要です。
## 開発環境の準備
デベロッパーアカウントの登録は最初の一歩にすぎません。iOS 開発にはハードウェアとソフトウェアのツールも必要です。
### 必要なデバイス
* **Mac** — Xcode は macOS でのみ動作するため、これは必須要件です。Apple Silicon(M シリーズチップ)搭載の Mac がおすすめで、コンパイル速度が速く、iOS シミュレータも直接実行できます。MacBook Air M シリーズで個人開発のニーズは十分に満たせます。予算を抑えたい場合は Mac mini も選択肢になります。
* **iPhone / iPad(推奨、ただし必須ではありません)** — シミュレータで大部分のデバッグシーンをカバーできますが、パフォーマンス、センサー(カメラ/GPS/NFC)、プッシュ通知などの検証には実機テストが欠かせません。有料デベロッパーアカウントがなくても、無料の Apple ID で実機デバッグは可能です(ただし 7 日ごとの再署名などの制限があります。詳しくは末尾の Q\&A をご覧ください)。
### 開発ツール
* **Xcode** — Apple 公式の IDE で、Mac App Store から無料でダウンロードできます。容量が大きく(約 12GB 以上)、初回インストールには時間がかかります。
* **Apple Developer App** — アカウント登録、WWDC のビデオやドキュメントの閲覧に使用します。
* **TestFlight** — ベータ版配布ツールで、ユーザーを招待してアプリをテストしてもらうための公式チャネルです。
### 注意事項
* macOS と Xcode のバージョンは最新に保つ必要があります。Apple は毎年 WWDC 後に新しい Xcode をリリースし、通常は直近 1〜2 世代の macOS が必要です。
* Xcode は頻繁にアップデートがあり容量も大きいため、十分なディスク容量(50GB 以上)を確保しておくことをおすすめします。
* カメラ、Bluetooth、NFC などのハードウェア機能を利用するアプリの場合、実機テストは必須です。
* Mac をお持ちでない場合は、クラウド Mac サービス(MacStadium、AWS EC2 Mac など)が代替手段になりますが、ネイティブデバイスほどの使い心地は得られません。
## 登録手順
### ステップ 1:Apple Developer App をダウンロード
iPhone または iPad で App Store を開き、「Apple Developer」を検索してダウンロード・インストールします。
### ステップ 2:ログインして登録を開始
Apple Developer App を開き、Apple ID でログインします。「アカウント」タブをタップし、「Apple Developer Program に登録」をタップします。
### ステップ 3:情報の入力と本人確認
画面の指示に従って個人情報を入力します。
1. **本人情報の確認**:氏名、住所などの基本情報を確認します。
2. **本人確認**:お住まいの地域によっては、政府発行の有効な身分証明書(パスポート、運転免許証など)の撮影や自撮り写真による本人確認が求められる場合があります。
3. **契約への同意**:Apple Developer Program 使用許諾契約を確認して同意します。
> 本人確認は十分な明るさのある環境で行い、写真が鮮明に撮れるようにしてください。登録手続き全体を同じデバイスで完了させる必要があります。
### ステップ 4:年会費の支払い
情報に誤りがないことを確認したら、年会費 ¥688($99)を支払います。Apple ID に紐づけた支払い方法が利用できます。支払い完了後、確認メールが届きます。
### ステップ 5:審査を待つ
* **個人アカウント**:通常 48 時間以内に審査が完了します。私の場合は 3 月 14 日に支払いを行い、3 月 15 日の午前中にはウェルカムメールが届きました。24 時間もかかりませんでした。
* **組織アカウント**:Apple が組織情報と D-U-N-S 番号を確認するため、より長い時間がかかる場合があります。
審査が通ったら、[developer.apple.com](https://developer.apple.com) にログインしてデベロッパーコンソールにアクセスし、すべての開発リソースを利用できるようになります。
## サブスクリプション管理と更新
Apple Developer Program は年間サブスクリプション制で、毎年 ¥688 が自動更新されます。
### 自動更新
デフォルトで自動更新が有効になっており、期限前に Apple ID に紐づけた支払い方法から引き落とされます。アカウントの有効期限切れで公開中のアプリに影響が出ないよう、自動更新を維持することをおすすめします(詳しくは下記の「よくある質問」をご覧ください)。
### サブスクリプションの解約・管理
更新設定を変更する場合は、iPhone の「設定 → Apple ID → サブスクリプション」を開き、Apple Developer Program を選択して管理できます。
## よくある質問
**Q:更新を忘れてしまった場合はどうなりますか?**
アカウントの有効期限が切れると、アプリは App Store から取り下げられますが、削除されるわけではありません。再度支払いを行えば、アプリは再公開されます。ただし、その間はユーザーのダウンロードやアップデートに影響が出るため、自動更新を有効にしておくことをおすすめします。
**Q:D-U-N-S 番号はどうやって申請しますか?**
以下のリンクにアクセスし、会社情報を入力して申請を提出してください。審査には通常 5〜14 営業日かかります。個人アカウントにはこの番号は不要です。
**Q:問題が発生した場合、Apple にどう連絡すればいいですか?**
[Apple Developer サポート](https://developer.apple.com/contact/)にアクセスすると、オンラインチャットや電話で問い合わせができます。日本語サポートも提供されており、レスポンスも比較的良好です。
**Q:先に開発してから登録することはできますか?**
まず試してみることは可能ですが、無料アカウントの制限を理解しておく必要があります。無料の Apple ID だけでも、Xcode でコードを書いたり、シミュレータでデバッグしたり、自分のデバイスにインストールして実行することができます。Swift の学習や基本的な UI のアイデア検証には十分です。
ただし、無料アカウントにはいくつかの制限があります。実機にインストールしたアプリは 7 日ごとに再コンパイル・再インストールが必要で、プラットフォームごとに最大 3 台のデバイスまでとなっています。また、プッシュ通知、iCloud、TestFlight、アプリ内課金などの機能は利用できません。これらの機能が必要な場合は、開発段階から有料メンバーシップが必要です――App Store への公開だけが目的ではありません。
おすすめ:Swift の入門やデモの動作確認だけなら、まず無料アカウントで始めましょう。本格的なプロジェクトに取り組み始めたら、早めに有料メンバーシップに登録して、機能制限による遅れを防ぐことをおすすめします。
# アイデア検証:漠然としたインスピレーションから実行可能な方向性へ
これはテスト記事です。
# 技術選定:なぜ Next.js + Supabase なのか
これはテスト記事です。
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# コンセプト紹介
## はじめに
Claude Code で完璧なワークフローを構築したとき——カスタムコマンド、コードレビューフック、専用の Skills——こう思うかもしれません:これらをまとめてパッケージ化し、チームやコミュニティと共有できないだろうか?
それこそが Plugin が解決する課題です。
Skills が AI のための「作業マニュアル」だとすれば、Plugin は「ツールボックス」です。Skills、Commands、Hooks、MCP サーバーなどの設定をすべてまとめてパッケージ化し、ワンコマンドでインストール・配布できるようにします。
## Plugin を理解する
あなたが長年にわたって信頼できるツールを蓄積してきた熟練の職人だと想像してください:ハンマー、ノコギリ、定規、各種ドライバー。作業台を変えるたびに、ツールを一つずつ運んで整理し直す必要があります。Plugin は設計に優れたツールボックスのようなものです。すべてのツールを収納するだけでなく、カテゴリ別にきちんと整理されており、どこに持っていってもすぐに作業を開始できます。
技術的な観点から言えば、Plugin は Claude Code の拡張パッケージメカニズムです。Plugin には以下のものを含めることができます:
| コンポーネント | 役割 | ファイルの場所 |
| ------------- | ------------------ | ----------- |
| **スラッシュコマンド** | クイックアクションのエントリポイント | `commands/` |
| **Subagents** | 専門化されたサブエージェント | `agents/` |
| **Skills** | AI のナレッジパッケージ | `skills/` |
| **Hooks** | イベントトリガーの自動化スクリプト | `hooks/` |
| **MCP サーバー** | 外部システムとの接続 | `.mcp.json` |
| **LSP サーバー** | 言語サーバーの設定 | `.lsp.json` |
これらのコンポーネントが連携して、完全なワークフローソリューションを形成します。
## Plugin vs 独立設定
Claude Code では、設定をプロジェクトの `.claude/` ディレクトリに配置することも、Plugin としてパッケージ化することもできます。両者の本質的な違いは**配布方法**と**名前空間**にあります:
| 側面 | 独立設定 (`.claude/`) | Plugin |
| ------- | -------------------- | ------------------------- |
| コマンド名 | `/hello` | `/plugin-name:hello` |
| 使用場面 | 個人ワークフロー、プロジェクト固有の設定 | チーム共有、コミュニティ配布 |
| バージョン管理 | プロジェクトコードと一緒に管理 | Semantic Versioning をサポート |
| 更新方法 | 手動同期 | 自動更新をサポート |
| 競合処理 | 他の設定と競合する可能性あり | 名前空間による分離 |
**Plugin を選ぶべきとき**:
* チームメンバーとワークフロー設定を共有する必要がある
* 複数のプロジェクトで同じツールセットを再利用したい
* コミュニティに設定を配布する予定がある
* バージョン管理と自動更新が必要
**独立設定を使うべきとき**:
* 個人的なクイック実験
* 再利用の必要がないプロジェクト固有の設定
* シンプルな一度きりのコマンド
## Plugin のディレクトリ構造
標準的な Plugin の構造は以下の通りです:
```
my-plugin/
├── .claude-plugin/ # 元数据目录
│ └── plugin.json # 必需:插件清单
├── commands/ # 斜杠命令
│ ├── review.md
│ └── deploy.md
├── agents/ # 子代理
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # 事件钩子
│ └── hooks.json
├── scripts/ # 辅助脚本
│ └── format-code.sh
├── .mcp.json # MCP 服务器配置
└── .lsp.json # LSP 服务器配置
```
**重要な注意事項**:
* `plugin.json` は `.claude-plugin/` ディレクトリ内に配置する必要があります
* その他のディレクトリ(commands、agents、skills など)はプラグインのルートに配置します
* 機能ディレクトリを `.claude-plugin/` の中に入れないでください
## コア設定ファイル
Plugin のコアは `.claude-plugin/plugin.json` で、プラグインのメタデータとコンポーネントのパスを定義します:
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "一个示例插件",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| フィールド | 必須 | 説明 |
| ------------- | --- | ------------------------ |
| `name` | はい | プラグインの一意な識別子。小文字とハイフンを使用 |
| `version` | いいえ | セマンティックバージョン番号 |
| `description` | いいえ | プラグインの簡潔な説明 |
| `author` | いいえ | 著者情報 |
| `keywords` | いいえ | 検索用のタグ |
| `commands` | いいえ | コマンドファイルまたはディレクトリのパス |
| `agents` | いいえ | エージェントファイルまたはディレクトリのパス |
| `skills` | いいえ | Skills ディレクトリのパス |
| `hooks` | いいえ | フック設定のパス |
| `mcpServers` | いいえ | MCP 設定のパス |
## インストールスコープ
Plugin は異なる使用場面に対応するため、4つのインストールスコープをサポートしています:
| スコープ | 設定ファイル | 用途 |
| --------- | ----------------------------- | ----------------------- |
| `user` | `~/.claude/settings.json` | 個人プラグイン、すべてのプロジェクトで利用可能 |
| `project` | `.claude/settings.json` | チームプラグイン、バージョン管理で共有 |
| `local` | `.claude/settings.local.json` | プロジェクト固有、gitignore 対象 |
| `managed` | `managed-settings.json` | 企業管理(読み取り専用) |
デフォルトのインストールスコープは `user` です。プラグイン設定を Git にコミットしてチームで使用したい場合は、`project` スコープを選択してください。
## コアメリット
### 名前空間の分離
Plugin のコマンドには名前空間プレフィックスが付きます(例:`/my-plugin:review`)。これにより、他のプラグインやプロジェクト設定との命名の衝突を避けることができます。チームコラボレーションにおいて特に重要で、異なるチームが開発したプラグインが問題なく共存できます。
### バージョン管理
Plugin は Semantic Versioning をサポートしており、以下のことが可能です:
* プラグインの変更履歴を追跡する
* 必要に応じて古いバージョンにロールバックする
* 互換性のある更新を自動的に受信する
### 簡単な配布
Plugin Marketplace を通じて、以下のことができます:
* GitHub でプラグインをホスティングする
* ユーザーがシンプルなコマンドでインストールできるようにする
* 依存関係と更新を自動的に処理する
### チームコラボレーション
Plugin はチームでの使用に特に適しています:
* チームの開発ツールチェーンを統一する
* 新しいメンバーがワンコマンドですべてのツールを取得できる
* 設定の一元管理により重複作業を削減する
## Plugin エコシステム
Claude Code の Plugin エコシステムは急速に成長しています。2025年初頭の時点で、エコシステムはかなりの規模に達しています:
* **229以上のプラグイン**がエコシステムで活発に利用されています
* **239の Agent Skills** がマーケットプレイス全体に分布しています
* **200以上の MCP サーバー**が Docker ツールキットでプリビルドされています
**公式リソース**:
| リソース | リンク | 説明 |
| ---------------------------------- | ------------------------------------------------------------------------------------- | -------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 公式 Skills リポジトリ |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 公式プラグインディレクトリ |
| Docker MCP Toolkit | [公式サイト](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200以上のプリビルト MCP サーバー |
**コミュニティセレクション**:
| リソース | リンク | 説明 |
| ---------------------- | ----------------------------------------------------------------- | ---------------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243のプラグインを自動収集 |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | ベストプラクティス集 |
| claude-plugins.dev | [公式サイト](https://claude-plugins.dev/) | コミュニティレジストリと CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99エージェント + 15オーケストレーター |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17の専門エージェント |
## 他の機能との関係
Plugin は「コンテナ」の概念であり、Claude Code エコシステムの他の機能を含めることができます:
```
Plugin(コンテナ)
├── Skills(ナレッジパッケージ)
├── Commands(クイックコマンド)
├── Agents(サブエージェント)
├── Hooks(イベントフック)
└── MCP/LSP(外部接続)
```
この階層関係を理解することが重要です:
* **Skills** は Claude に何かのやり方を教えます
* **Commands** はクイックトリガーのエントリポイントを提供します
* **Agents** は独立した専門タスクを処理します
* **Hooks** はイベント駆動の自動化を実現します
* **Plugin** はこれらすべてをまとめてパッケージ化し、配布と管理を容易にします
### Skills vs Plugins
初めて触れる方は Skills と Plugins の違いに混乱するかもしれません。[Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) の分析によると:
| 特性 | Skills | Plugins |
| ------------- | ---------------------------- | -------------------------------- |
| **スコープ** | すべての Claude 製品(Web、API、Code) | Claude Code のみ |
| **含まれるもの** | Markdown ガイド + オプションのスクリプト | Commands、Agents、Hooks、MCP、Skills |
| **アクティベーション** | 自動(モデルが使用タイミングを判断) | 可変(コンポーネントの種類に依存) |
| **最適な用途** | Claude にドメイン専門知識を教える | Claude Code 環境を拡張する |
| **配布方法** | GitHub リポジトリ、ファイルシステム | 分散型 Marketplace |
**重要なポイント**:Skills はモデルによって自動的にトリガーされ、手動での呼び出しは不要です。Plugins はパッケージングメカニズムであり、分散型共有の課題を解決します。両者は組み合わせて使用できます——Plugin は Skills を含めることができます。
## まとめ
Claude Code Plugin は本質的に**ワークフローのパッケージ化と配布のメカニズム**です。設定の再利用とチームコラボレーションの課題を解決し、丹精込めて作り上げたツールチェーンをより多くの人と共有できるようにします。
3つのキーワードを覚えてください:
| キーワード | 意味 |
| ---------- | ----------------------- |
| **パッケージ化** | 複数の設定コンポーネントを一つのユニットに統合 |
| **分離** | 名前空間による競合の回避 |
| **配布** | Marketplace を通じた簡単な共有 |
コンセプトを理解したところで、次の記事『[Claude Code Plugin 実践ガイド](/ja/docs/notes/claude-plugin/practice)』では、実際にハンズオンで進めていきます:Plugin をゼロから作成し、Marketplace に公開し、チームコラボレーションのベストプラクティスを学びます。
Plugin に含められるコンポーネントにまだ馴染みがない方は、まず『[Claude Skills とは](/ja/docs/notes/claude-skills/concept)』を読んで、Skills のコアコンセプトを理解することをお勧めします。
# 実践ガイド
## おさらい
前の記事では、Plugin のコア概念について学びました。Plugin は Claude Code のワークフローをパッケージ化して配布する仕組みであり、Commands、Skills、Agents、Hooks などのコンポーネントを一つにまとめて、チームでの共有やコミュニティへの配布を可能にします。本記事では実践的な視点から、作成から公開までの完全なフローを一緒に進めていきます。
## 最初の Plugin を作成する
### ステップ 1:ディレクトリ構造を作成する
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### ステップ 2:プラグインマニフェストを作成する
`.claude-plugin/plugin.json` にプラグインのメタデータを定義します:
```json
{
"name": "my-first-plugin",
"description": "我的第一个 Claude Code 插件",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### ステップ 3:スラッシュコマンドを追加する
`commands/` ディレクトリに Markdown ファイルを作成します。各ファイルが一つのコマンドに対応します:
`commands/hello.md`:
```markdown
---
description: 向用户发送友好的问候
---
# Hello 命令
请热情地问候用户,并询问今天可以帮助他们做什么。
```
### ステップ 4:プラグインをテストする
`--plugin-dir` フラグを使ってローカルのプラグインを読み込んでテストします:
```bash
claude --plugin-dir ./my-first-plugin
```
Claude Code でコマンドを実行します:
```
/my-first-plugin:hello
```
### ステップ 5:コマンド引数を追加する
コマンドはユーザーが入力した引数を受け取ることができます。`hello.md` を更新します:
```markdown
---
description: 向指定用户发送个性化问候
---
# Hello 命令
请热情地问候名为 "$ARGUMENTS" 的用户,并询问今天可以帮助他们做什么。
如果用户没有提供名字,就使用"朋友"作为称呼。
```
引数付きでコマンドをテストします:
```
/my-first-plugin:hello 小明
```
**サポートされる引数プレースホルダー**:
* `$ARGUMENTS` - すべてのユーザー入力
* `$1`, `$2`, `$3` - 個別の引数
## コンポーネントを追加する
### Skills を追加する
`skills/` ディレクトリを作成します。各 Skill は `SKILL.md` を含むフォルダです:
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md`:
```yaml
---
name: code-review
description: 审查代码质量、安全性和可维护性
---
当审查代码时,请检查以下方面:
1. **代码组织**:结构是否清晰
2. **错误处理**:异常是否被妥善处理
3. **安全隐患**:是否存在安全漏洞
4. **测试覆盖**:关键逻辑是否有测试
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Subagents を追加する
`agents/` ディレクトリを作成します:
`agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家。
当被调用时:
1. 运行 git diff 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(如注入、敏感信息泄露)
- 性能优化机会
```
### Hooks を追加する
Hooks を使うと、特定のイベントが発生した際にスクリプトを自動実行できます。`hooks/hooks.json` を作成します:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**重要**:`${CLAUDE_PLUGIN_ROOT}` 環境変数を使ってプラグインディレクトリ内のファイルを参照してください。こうすることで、プラグインがどこにインストールされてもパスが正しく解決されます。
対応するスクリプト `scripts/format-code.sh` を作成します:
```bash
#!/bin/bash
# 格式化代码
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
実行権限を付与することを忘れないでください:
```bash
chmod +x scripts/format-code.sh
```
### MCP サーバーを追加する
プラグインが外部システムに接続する必要がある場合は、`.mcp.json` を作成します:
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## 完全な Plugin 構造
機能がフル装備された Plugin は次のような構造になります:
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单
├── commands/
│ ├── review.md # 代码审查命令
│ ├── deploy.md # 部署命令
│ └── test.md # 测试命令
├── agents/
│ ├── code-reviewer.md # 代码审查代理
│ └── debugger.md # 调试代理
├── skills/
│ └── code-standards/
│ └── SKILL.md # 代码规范知识
├── hooks/
│ └── hooks.json # 事件钩子配置
├── scripts/
│ ├── format-code.sh # 格式化脚本
│ └── run-tests.sh # 测试脚本
├── .mcp.json # MCP 配置
├── LICENSE
├── README.md
└── CHANGELOG.md
```
対応する `plugin.json`:
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "开发者工具箱:代码审查、测试、部署一站式解决方案",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## Marketplace に公開する
### Marketplace とは
Marketplace は Plugin の配布センターです。「プラグインストア」と考えるとわかりやすいでしょう。ユーザーはシンプルなコマンドひとつで、あなたが公開したプラグインをインストールできます。
### Marketplace 設定を作成する
GitHub リポジトリ内に `.claude-plugin/marketplace.json` を作成します:
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "开发者工具箱",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "文档生成工具"
}
]
}
```
### プラグインのソースタイプ
Marketplace は複数のソースタイプをサポートしています:
**相対パス**(同一リポジトリ内のプラグイン):
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**GitHub リポジトリ**:
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**任意の Git リポジトリ**:
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### 公開フロー
1. **GitHub リポジトリを作成する**
2. **コードをプッシュする**:
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **ユーザーがあなたの Marketplace を追加する**:
```bash
/plugin marketplace add your-username/your-repo
```
4. **ユーザーがプラグインをインストールする**:
```bash
/plugin install dev-toolkit@your-marketplace
```
## Plugin のインストールと管理
### インタラクティブメニューから
```bash
/plugin
```
これにより、プラグインの閲覧、インストール、有効化、無効化ができるインタラクティブなインターフェースが開きます。
### コマンドラインから
**Marketplace を追加する**:
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # 本地
```
**プラグインをインストールする**:
```bash
# 安装到用户范围(默认)
/plugin install formatter@my-marketplace
# 安装到项目范围(团队共享)
/plugin install formatter@my-marketplace --scope project
# 安装到本地范围(gitignored)
/plugin install formatter@my-marketplace --scope local
```
**その他の管理コマンド**:
```bash
/plugin enable # 启用插件
/plugin disable # 禁用插件
/plugin uninstall # 卸载插件
/plugin update # 更新插件
```
### プラグインの検証
公開前にプラグイン設定が正しいか検証します:
```bash
claude plugin validate .
```
または Claude Code 内で:
```
/plugin validate .
```
## チームコラボレーション設定
### プロジェクト内でプラグイン設定を共有する
プラグイン設定をバージョン管理にコミットすることで、チームメンバーが自動的に利用できるようになります:
`.claude/settings.json`:
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
チームメンバーがプロジェクトをクローンすると、これらのプラグインが自動的に利用可能になります。
### エンタープライズ向け Marketplace 制限
厳格な管理が必要なエンタープライズ環境では、マネージド設定で許可する Marketplace を制限できます:
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
空の配列 `[]` に設定すると、外部プラグインを完全に無効化できます。
## CLI コマンドリファレンス
| コマンド | 説明 |
| -------------------------------------- | ----------------------- |
| `/plugin` | インタラクティブ管理画面を開く |
| `/plugin install @` | プラグインをインストール |
| `/plugin uninstall ` | プラグインをアンインストール |
| `/plugin enable ` | プラグインを有効化 |
| `/plugin disable ` | プラグインを無効化 |
| `/plugin update ` | プラグインを更新 |
| `/plugin validate .` | 現在のディレクトリのプラグイン設定を検証 |
| `/plugin marketplace add ` | Marketplace を追加 |
| `/plugin marketplace list` | 追加済みの Marketplace を一覧表示 |
| `/plugin marketplace update` | Marketplace キャッシュを更新 |
| `/plugin marketplace remove ` | Marketplace を削除 |
## ベストプラクティス
### 開発のベストプラクティス
1. **Skills は一つのことに集中させる**:各 Skill は一つのことをしっかり行い、あれもこれもと詰め込まないようにしましょう
2. **明確な説明を書く**:Claude がコンポーネントをいつ使うべきか理解できるようにしましょう
3. **まずチーム内でテストする**:コミュニティに配布する前に、チーム内で検証しましょう
4. **バージョン変更を記録する**:CHANGELOG.md に各バージョンの変更内容を記録しましょう
### ディレクトリ構造のベストプラクティス
* `commands/`、`agents/`、`skills/` はプラグインのルートディレクトリに配置します
* `.claude-plugin/` ディレクトリには `plugin.json` のみを配置します
* プラグイン内のファイルを参照するには `${CLAUDE_PLUGIN_ROOT}` を使用します
* `../` を使ってプラグイン外のファイルにアクセスしないでください
### Hooks のベストプラクティス
1. スクリプトは実行可能にする:`chmod +x script.sh`
2. shebang でインタープリターを宣言する:`#!/bin/bash`
3. `${CLAUDE_PLUGIN_ROOT}` 変数を使ってパスの正確性を確保する
4. Hook に統合する前にスクリプトを単独でテストする
### バージョン管理のベストプラクティス
Semantic Versioning に従います:
* **MAJOR** (1.0.0 → 2.0.0):破壊的変更
* **MINOR** (1.0.0 → 1.1.0):新機能の追加(後方互換あり)
* **PATCH** (1.0.0 → 1.0.1):バグ修正(後方互換あり)
## よくある問題のトラブルシューティング
| 問題 | 考えられる原因 | 解決策 |
| ------------- | ---------------------- | ------------------------------------------------------- |
| プラグインが読み込まれない | plugin.json のフォーマットエラー | `claude plugin validate` で検証する |
| コマンドが表示されない | ディレクトリ構造の誤り | `commands/` がルートディレクトリにあり、`.claude-plugin/` 内にないことを確認する |
| Hooks が発火しない | スクリプトに実行権限がない | `chmod +x script.sh` を実行する |
| パスが見つからない | 相対パスを使用している | `${CLAUDE_PLUGIN_ROOT}` に変更する |
| MCP サーバーが失敗する | 環境変数が未設定 | `.mcp.json` 内のパス設定を確認する |
## 既存設定からの移行
すでに `.claude/` ディレクトリに設定がある場合は、以下の手順で Plugin に移行できます:
1. **Plugin 構造を作成する**:
```bash
mkdir my-plugin/.claude-plugin
```
2. **plugin.json を作成する**:
```json
{
"name": "my-plugin",
"description": "从现有配置迁移的插件",
"version": "1.0.0"
}
```
3. **既存ファイルをコピーする**:
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **Hooks を移行する**:
`.claude/settings.json` から `hooks` 設定を `hooks/hooks.json` にコピーします
5. **テストする**:
```bash
claude --plugin-dir ./my-plugin
```
## 学習リソース
### 公式ドキュメント
| リソース | リンク | 説明 |
| ----------------- | ----------------------------------------------------------------------------------------- | ----------------- |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | プラグインリファレンスドキュメント |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | プラグイン作成ガイド |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | 公式ベストプラクティス |
| Agent Skills 標準 | [agentskills.io](https://agentskills.io) | オープンスタンダード仕様 |
### 公式リポジトリ
| リソース | リンク | 説明 |
| ---------------------------------- | ------------------------------------------------------------------------------------- | ---------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 公式 Skills リポジトリ |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 公式プラグインカタログ |
| Docker MCP Toolkit | [公式サイト](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200 以上のプリビルト MCP |
### コミュニティリソース
| リソース | リンク | 説明 |
| ----------------------- | ---------------------------------------------------------------------------- | -------------------------- |
| claude-plugins.dev | [公式サイト](https://claude-plugins.dev/) | コミュニティレジストリと CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 個のプラグインコレクション |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | ベストプラクティスまとめ |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 エージェント + 15 オーケストレーター |
| jeremylongshore チュートリアル | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | 数百のプラグイン + Jupyter チュートリアル |
### おすすめの記事
| 記事 | ソース |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## 今後の展望
Plugin の仕組みにより、Claude Code の拡張能力は飛躍的に向上しました。コミュニティの発展に伴い、今後は以下のような進化が期待されます:
* **より豊かなプラグインエコシステム**:さまざまな開発シナリオやワークフローをカバー
* **エンタープライズ向け機能**:より充実した権限管理と監査機能
* **クロスプラットフォーム互換性**:Skills のオープンスタンダードはすでに複数のベンダーに採用されています
今こそ参入する絶好のタイミングです。シンプルなコマンドから始めて、徐々に Skills や Hooks を追加し、最終的に完全なワークフローソリューションを構築していきましょう。
Plugin に含めることができる Subagent コンポーネントについて詳しく知りたい場合は、《[Claude Code Subagent とは](/ja/docs/notes/claude-subagent/concept)》をご覧ください。
# Claude Code Subagent の概念の紹介
## はじめに
Claude Code を使用して複雑なタスクを処理する場合、このようなジレンマに遭遇したことがあるかもしれません。主な会話のコンテキストはますます長くなり、AI は以前の重要な情報を「忘れ」始め、応答の品質は徐々に低下します。
Subagent はこの問題を解決するために生まれました。
Skills がクロードの「作業マニュアル」である場合、Subagent はあなたが雇用する「フルタイム従業員」です。彼らは独自の独立したワークステーション (コンテキスト) を持ち、特定の種類の作業に集中し、完了後に結果をあなたに報告します。
## サブエージェントについて
あなたが会社の CEO であると想像してください。会社が小さいときは、すべてを自分で処理します。しかし、ビジネスが拡大するにつれて、財務担当の会計士、採用担当の人事担当者、開発担当のエンジニアなど、フルタイムの従業員を雇用し始めます。各従業員は自分のステーションで勤務し、タスクの完了後にあなたに報告します。
クロード コードではサブエージェントがまさにこの役割を果たします。
技術的な観点から見ると、サブエージェントは次の特徴を持つ特殊な AI アシスタントです。
| 特長 | 説明 |
| ------------------ | --------------------------------- |
| **独立したコンテキスト** | 各サブエージェントは独自のコンテキスト ウィンドウで実行されます。 |
| **専門的な機能** | 特定のタスク タイプ向けに最適化 |
| **構成可能なツール** | 指定されたツールのセットのみにアクセスできます。 |
| **カスタマイズされたプロンプト** | 動作をガイドするための特別なシステム プロンプトがあります。 |
## なぜ独立したコンテキストが必要なのでしょうか?
これは Subagent の核となる設計コンセプトであり、深く理解する必要があります。
通常の会話では、すべての情報が同じコンテキストに積み重ねられます。クロードにコード ベースを検索させ、ファイルを分析して変更を加えると、この中間処理すべてがコンテキスト スペースを占有します。会話が進み、文脈が混み合うにつれて、クロードは以前の重要な情報を「忘れ」始める可能性があります。
サブエージェントの変更は次のとおりです。
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
サブエージェントの分析プロセスは、メインの会話を汚染しません。メインの対話では、明確さと焦点を維持しながら、洗練された結果のみが得られました。
## 組み込みサブエージェントのタイプ
Claude Code は、最も一般的な使用シナリオをカバーする 3 つの強力な組み込みサブエージェントを提供します。
### サブエージェントを探索する
**ターゲティング**: コードベースの高速な読み取り専用探索。
**特徴**:
* Haiku モデルを使用する (高速、低遅延)
* 厳密に読み取り専用 - ファイルの作成、変更、削除はできません
* 利用可能なツール: Glob、Grep、Read、Bash (読み取り専用操作)
**使用する場合**:
「この機能はどこに実装されていますか?」などの探索的な質問をするとき。 「エラーはどのように処理されますか?」クロードは自動的に Explore Subagent を呼び出します。
**詳細レベル**:
| レベル | 説明 | 該当するシナリオ |
| ---------- | ----------- | ----------------- |
| クイック | 最小限の探索で高速検索 | シンプルでターゲットを絞ったクエリ |
| 中 | 適度な探索 | スピードと完全性のバランス |
| 非常に徹底しています | 総合分析 | 深い理解が必要な複雑な問題 |
### サブエージェントを計画する
**位置付け**: コードベースを調査し、実装計画を準備します。
**特徴**:
* Sonnet モデルを使用します (より強力な推論機能)
* 探索ツールのみ: Read、Glob、Grep、Bash
* 計画モードで自動的に呼び出されます
**使用する場合**:
計画モードに入り、計画を提案する前にクロードに調査を行う必要がある場合、Plan Subagent は自動的に情報を収集し、調査結果に基づいて計画を提案します。
### 汎用サブエージェント
**ポジショニング**: 複雑な複数ステップのタスクを処理します。
**特徴**:
* ソネットモデルを使用
* すべてのツールへのアクセス (読み取りと書き込みを含む)
* 探索と変更が必要な複雑なタスクに適しています
**使用する場合**:
タスクに複数のステップが含まれる場合、変更する前に検索が必要な場合、または最初の検索が失敗する可能性があり、複数の戦略を試行する必要がある場合。
## 私の理解と実践
3 つの公式組み込みサブエージェントを注意深く観察すると、**それらはすべて調査および計画タスクである**という共通点が 1 つあることがわかります。 Exploreはコードベースの探索を担当し、Planは計画の作成を担当し、汎用も主に調査と分析に使用されます。これらはいずれも、コードを書くために特別に設計されたものではありません。
これは、Subagent についての私の理解を裏付けています。 **Subagent の中心的な価値は、「クリーンなコンテキスト」ではなく、メイン エージェントが物事の実行に集中できるようにすることです**。
### 分業モード
私の利用方法は非常にシンプルで、サブエージェントが調査・計画・検討といった「情報収集」の作業を担当し、メインエージェントが実際の実行を担当します。
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### Subagent にコードを書かせてみてはいかがでしょうか?
メイン エージェントに複数のサブエージェントをスケジュールしてコードを作成させることを好む人もいます。これは信頼できないと思います。理由は簡単です: **コンテキストが大幅に欠落しています**。
サブエージェントのコンテキストは独立しています。主要な会話で何が議論され、どのような決定が下され、どのような制約が課されたのかはわかりません。コードの作成を依頼することは、新入社員に背景情報をまったく持たずにタスクを完了するよう依頼するようなものです。作成されたコードは期待と一致しない可能性があります。
それどころか、Subagent を「研究者」として位置づける方がはるかに合理的です。
* 研究タスク自体には多くのコンテキストは必要ありません
* コードの代わりに情報が返され、メイン エージェントが完全なコンテキストに基づいて使用できます。
* 調査結果に偏りがあった場合でも、主体エージェントが修正できる
### 私の毎日の使い方
1. **新しいタスクを開始する前に**: Explore エージェントが関連するコードの構造をすぐに理解できるようにします。
2. **複雑なタスクの計画**: 計画エージェントに要件を分析させ、実装手順を策定させます
3. **コード レビュー**: レビュー エージェントにコードの品質とセキュリティの問題をチェックさせます
4. **実際のコーディング**: メイン エージェントは、収集されたコンテキストに基づいてコードを作成します。
私が現在使用しているエージェントのリストは次のとおりです。
この利点は、メイン エージェントのコンテキスト ウィンドウがクリーンなままであり、「サブエージェントの検索プロセス中に生成された中間結果の束」ではなく、「知る必要がある情報」だけが表示されることです。
## 他の関数との比較
### Subagent vs Skills
これは最も一般的な混乱です。主な違い: **スキルはクロードに知識を注入します。サブエージェントは独立したワーカー**を作成します。
| 寸法 | スキル | サブエージェント |
| ------------ | --------------------- | ------------------ |
| **コア機能** | 専門知識と指示を提供する | タスクを独立して実行するエージェント |
| **コンテキスト** | 主要な会話のコンテキストを共有する | 独立したコンテキストを持つ |
| **トリガー方法** | 説明に基づく自動マッチング | 自動委任または手動呼び出し |
| **該当するシナリオ** | クロードを特定の種類のタスクでより良くする | 複雑な複数ステップの独立したタスク |
比喩的に言えば、スキルはトレーニング資料のようなもので、クロードが何かを行う方法を学ぶことができます。サブエージェントはフルタイムの従業員のようなもので、自分のワークステーションで独立してタスクを完了し、結果を報告します。
この 2 つは組み合わせることができます。コード レビュー Subagent はコード仕様をロードして、「専門家 + 専門知識」の複合効果を達成することができます。
### サブエージェントとスラッシュ コマンド
| 寸法 | サブエージェント | スラッシュコマンド |
| --------------- | --------------- | ----------- |
| **アクティベーション方法** | 自動委任または明示的な呼び出し | ユーザーマニュアル入力 |
| **コンテキスト** | スタンドアロンコンテキスト | 共有されたメインの会話 |
| **複雑さ** | 複雑なタスクに適しています | 簡単な操作に最適 |
スラッシュ コマンドはショートカット キーであり、`/review` と入力すると、事前定義された操作がトリガーされます。サブエージェントは、複雑な複数ステップのタスクを自律的に実行できる独立したワーカーです。
### Subagent vs Plugin
プラグインは「コンテナ」の概念であり、サブエージェントを含めることができます。
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
プラグインの `agents/` ディレクトリでサブエージェントを定義し、プラグインとともに配布できます。
## エージェントのデザイン パターン
Anthropic は、6 つの主要な Agentic 設計パターンを公式ドキュメントにまとめています。これらのパターンを理解すると、サブエージェント システムをより適切に設計するのに役立ちます。
| パターン | コアアイデア | サブエージェントの申請 |
| ------------------ | ------------------------ | --------------------------------- |
| **プロンプトチェーン** | 複雑なタスクを複数の連続したステップに分解する | 複数のサブエージェントをチェーン呼び出しする |
| **ルーティング** | 入力タイプに基づいて専用のプロセッサに分散 | さまざまな種類のタスクが特殊なサブエージェントに委任されます。 |
| **並列化** | 複数の独立したサブタスクを同時に実行する | 複数のサブエージェントを並行して開始する |
| **オーケストレーター兼ワーカー** | 中央コーディネーターが作業者にタスクを割り当てる | コーディネーターとしてのクロード、ワーカーとしてのサブエージェント |
| **評価者/オプティマイザー** | ジェネレーターの出力、エバリュエーターの最適化 | サブエージェントの生成 + サブエージェントの確認 |
| **エージェント** | 自律的な意思決定を行う独立したエージェント | 各サブエージェントは独立して実行されます。 |
これらのモードは組み合わせて使用できます。たとえば、コード品質システムでは次の両方が使用される場合があります。
* **並列化**: セキュリティ スキャンとパフォーマンス分析を同時に実行します。
* **オーケストレーター兼ワーカー**: マスター クロードは複数の専門化されたサブエージェントを調整します
* **Evaluator-Optimizer**: 生成直後にコードをレビューします
## 主な利点
### コンテキスト保護
Subagent の最大の価値は、メインの会話のコンテキストを保護することにあります。コード検索やファイル分析などの中間処理がメインダイアログに蓄積されないため、メインダイアログは常に高レベルの目標に焦点を当てることができます。
### 専門化機能
詳細な手順と適切なツールを使用して構成された、特定のドメインに特化したサブエージェントを作成できます。特化したサブエージェントは、汎用のクロードよりも特定のタスクで優れたパフォーマンスを発揮します。
### 柔軟な権限制御
各サブエージェントは異なるツール アクセス権を持つことができます。たとえば、探索クラス Subagent には読み取り専用権限のみが与えられ、変更クラス Subagent には書き込み権限のみが与えられます。このきめ細かい制御により、セキュリティが向上します。
### 再利用性
Subagent を作成すると、プロジェクト間で再利用したり、プラグインを介してチームと共有したりできます。
## サブエージェントを使用する場合
**サブエージェントの使用に適したシナリオ**:
* タスクを実行するには独立したコンテキストが必要です
* タスクは複雑な複数ステップのワークフローです
* メインの会話とは異なるツールセットが必要です
* タスクの実行には時間がかかる場合があります
**サブエージェントの使用に適さないシナリオ**:
* シンプルなワンタイムクエリ
* メインダイアログとの緊密な対話が必要です
* ミッションはすぐに完了できます
## 一般的なアプリケーション シナリオ
### コードレビュー
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
コードを完成させると、主要な開発作業を妨げることなく、コード レビュー サブエージェントに独立したコンテキストでコードをレビューさせることができます。
### デバッグ分析
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
エラーが発生した場合、デバッグ サブエージェントはエラーの原因を深く分析し、さまざまな仮説を試し、最終的に修復の提案を提供します。
### コードベースの探索
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
新しいプロジェクトに初めて取り組む場合、Explore Subagent を使用すると、大量の検索結果でメインの会話を混乱させることなく、コード ベースをすばやくマッピングできます。
## 学習リソース
### 公式リソース
| リソース | リンク | 説明書 |
| --------------- | ----------------------------------------------------------------------------------------------------------- | -------------------- |
| クロードコードのドキュメント | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 公式ドキュメントのエントリ |
| サブエージェントガイド | [クロード コード ドキュメント](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | サブエージェントの公式ドキュメント |
| マルチエージェントシステム研究 | [人類工学](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2%のパフォーマンス向上の研究詳細 |
| エージェントのデザインパターン | [人類ドキュメント](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 6つのコア設計パターンを詳しく解説 |
### コミュニティリソース
| リソース | リンク | 説明書 |
| --------------- | ----------------------------------------------------------------- | ---------------------------- |
| wshobson/エージェント | [GitHub](https://github.com/wshobson/agents) | 99 人のエージェント + 15 人のオーケストレーター |
| 化合物工学 | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 の専門エージェント用のプラグイン |
| 素晴らしいクロードコード | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | ベストプラクティスの概要 |
## 概要
Claude Code Subagent は本質的に、特化されたコンテキストに依存しない AI アシスタントです。コンテキストを分離することで複雑なタスクにおける情報過多の問題を解決し、主要な会話を常に明瞭かつ集中的に保ちます。
3 つのキーワードを覚えておいてください。
| キーワード | 意味 |
| -------- | ----------------------------------- |
| **独立した** | 各サブエージェントには独自のコンテキスト ウィンドウがあります。 |
| **専門分野** | 特定のタスク タイプ向けに最適化 |
| **代表団** | クロードは自動または手動でタスクを Subagent に委任できます。 |
概念を理解したら、次の記事「[Claude Code Subagent Practical Guide](/ja/docs/notes/claude-subagent/practice)」で、カスタム Subagent の作成、ツール権限の構成、および実際のプロジェクトでのベスト プラクティスを実践してください。
Subagentがロードできるスキルについて知りたい場合は、「【クロードスキルとは】(/ja/docs/notes/claude-skills/concept)」をご覧ください。 Subagentをパッケージ化して配布したい場合は、「【クロードコードプラグインとは】(/ja/docs/notes/claude-plugin/concept)」をご覧ください。
# クロード・コード・サブエージェント実践ガイド
## 簡単なレビュー
前回の記事では、Subagent の中心的な概念について学びました。それは、コンテキストの分離を通じて複雑なタスクにおける情報過多の問題を解決する、コンテキストに依存しない特殊な AI アシスタントです。 Claude Code には、Explore、Plan、および General-Purpose の 3 つの組み込みサブエージェントがあります。この記事では、実践的な観点からカスタム サブエージェントを作成し、高度な使用法を習得する方法を説明します。
## サブエージェントの管理
### /agents コマンド経由
最も簡単な方法は、対話型インターフェイスを使用することです。
```bash
/agents
```
これにより、次のことができるメニューが開きます。
* すべてのサブエージェントを表示 (組み込み + カスタム)
* 新しいサブエージェントの作成
* 既存のサブエージェントの構成とツールの権限を編集します
* 不要なサブエージェントを削除する
* 名前の競合がある場合にどのサブエージェントがアクティブであるかを確認する
### ファイル管理を通じて
サブエージェントは Markdown ファイルとして保存されます。ファイルを直接作成および編集することもできます。
**保管場所**:
| 場所 | パス | 範囲 | |
| --------- | ----------------------- | ---------------- | -------- |
| プロジェクトレベル | `.claude/agents/` | 現在のプロジェクト専用で、Git | に送信できます。 |
| ユーザーレベル | `~/.claude/agents/` | すべてのプロジェクトで利用可能 | |
| プラグイン | プラグインの `agents/` ディレクトリ | プラグインとともにインストール | |
**優先度**: プロジェクト レベル > ユーザー レベル > プラグイン レベル
同じ名前のサブエージェントが複数の場所に存在する場合、優先度の高いサブエージェントが優先度の低いサブエージェントを上書きします。
## 最初のサブエージェントを作成する
### ステップ 1: ディレクトリを作成する
```bash
mkdir -p .claude/agents
```
### ステップ 2: Markdown ファイルを作成する
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### ステップ 3: サブエージェントをテストする
クロードコードでは:
```
> 用 code-reviewer 代理审查我最近的修改
```
または、クロードに自動的に選択させます。
```
> 帮我审查一下代码质量
```
`description` が十分に明確に記述されていれば、Claude は自動的にサブエージェントを認識して呼び出します。
## 設定フィールドの詳細な説明
サブエージェント構成ファイルは、YAML フロントマッターと Markdown 本体の 2 つの部分で構成されます。
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| フィールド | 必須 | 説明 |
| ---------------- | --- | --------------------------------------------- |
| `name` | は | 一意の識別子。小文字とハイフンを使用します。 |
| `description` | はい | 自然言語による説明 (クロードはこれを使用していつ電話をかけるかを決定します) |
| `tools` | いいえ | カンマ区切りのツールのリスト。省略した場合、すべてのツールが継承されます。 |
| `model` | いいえ | モデルの選択: `sonnet`、`opus`、`haiku`、または `inherit` |
| `permissionMode` | いいえ | 許可モード (以下を参照) |
| `skills` | いいえ | 自動ロードされたスキル (サブエージェントは親セッションのスキルを継承しません) |
### 許可モード
| モード | 説明 |
| ------------------- | ----------------- |
| `default` | 通常の権限チェック |
| `acceptEdits` | 編集操作を自動的に受け入れる |
| `bypassPermissions` | すべての権限チェックをスキップする |
| `plan` | 計画を提案するだけで、実行しない |
| `ignore` | このサブエージェントを無視します |
### マークダウンテキスト
テキストは Subagent のシステム プロンプトです。より詳細に記述するほど、Subagent のパフォーマンスが向上します。
適切なシステム プロンプトには次のものが含まれている必要があります。
* 明確な役割定義
* 具体的な作業手順
* 重要なチェックリスト
* 出力形式の要件
## トリガーメカニズム
### 自動委任
クロードは、タスクの内容とサブエージェントの `description` に基づいて、委任するかどうかを自動的に決定します。
**自動使用を促進するためのヒント**: `description` でトリガー ワードを使用します。
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
または:
```yaml
description: MUST BE USED when encountering errors or test failures
```
### 明示的な呼び出し
クロードにどのサブエージェントを使用するかを直接指示します。
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
## ツールの構成
### 一般的に使用されるツールのリスト
| ツール | 説明書 |
| ----------- | -------------- |
| `Read` | ファイルの内容を読み取る |
| `Write` | ファイルに書き込む |
| `Edit` | ファイルを編集 |
| `Glob` | ファイルパターンマッチング |
| `Grep` | 正規表現検索 |
| `Bash` | シェルコマンドを実行する |
| `WebFetch` | Web コンテンツを取得する |
| `WebSearch` | ウェブを検索 |
### ツール構成戦略
**読み取り専用サブエージェント** (探索、分析):
```yaml
tools: Read, Grep, Glob, Bash
```
注: Bash が含まれている場合でも、Subagent は読み取り専用コマンド (ls、git status、git log など) にのみ使用する必要があります。
**サブエージェントの読み取りと書き込み** (修復、リファクタリング):
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**最小権限の原則**: 誤った操作を避けるために、必要なツールのみを付与します。
## 実用的なサブエージェント テンプレート
### コードレビューア
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### デバッグエキスパート
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### テストランナー
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### ドキュメントジェネレーター
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### セキュリティスキャナー
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## 高度な使用法
### 実稼働レベルの設計パターン
実稼働環境では、マルチエージェントのコラボレーションのための実証済みのパターンがいくつかあります。
#### 3 アミーゴスモード
コラボレーション モデルは、製品、アーキテクチャ、実装の 3 つの役割で構成されます。
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| 役割 | 責任 | ツール構成 |
| --------- | ------------- | -------------- |
| PMエージェント | 機能定義、要件整理 | 読む、ウェブ検索 |
| 建築家エージェント | 技術的ソリューションの設計 | 読み取り、Glob、Grep |
| クロード・コード | コードの実装 | すべてのツール |
#### 3 段階のパイプライン
複雑なタスクを 3 つの明確な段階に分割します。
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
各ステージは専用のサブエージェントを担当し、出力は次のステージへの入力として機能します。
#### モデル オーケストレーション戦略
コストと効果を最適化するために、さまざまな段階でさまざまなモデルが使用されます。
|ステージ |おすすめモデル|理由 |
\|------|---------|------|
|計画段階 |ソネット |深い推論が必要 |
|実行フェーズ |俳句 |早くて低コスト |
|レビュー段階 |ソネット |総合的な判断が必要 |
構成例:
```yaml
---
name: quick-executor
model: haiku
---
```
### サブエージェントのリンク
複雑なワークフローの場合、複数のサブエージェントをリンクできます。
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### 再開可能な実行
サブエージェントの実行は、以前の完全なコンテキストを維持しながら一時停止および再開できます。
**最初の通話**:
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**回復エージェント**:
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**使用シナリオ**:
* 複数のセッションで完了する長期にわたる研究
* コンテキストを維持した反復的な改善
* 関連タスクを順番に処理する複数ステップのワークフロー
### サブエージェントのスキルを構成する
サブエージェントは親セッションからスキルを自動的に継承しません。必要に応じて、明示的に宣言します。
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### CLI の動的定義
ファイルを保存する必要はありません。一時サブエージェントをコマンド ラインで直接定義します。
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
簡単なテストや 1 回限りの使用に適しています。
## ベストプラクティス
### 1. 集中力を保つ
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
1 つのことをうまく実行するサブエージェントは、多くのことを実行するサブエージェントよりも優れています。
### 2. 明確な説明を書きます
クロードは、`description` を使用して、サブエージェントをいつ使用するかを決定します。適切な説明には次のような答えが必要です。
1. \*\*このサブエージェントは何をしますか? \*\* 特定の能力を列挙する
2. \*\*いつ使用する必要がありますか? \*\* トリガーワードが含まれています
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. ツールへのアクセスを制限する
必要なツールのみを付与します。
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
これにより、Subagent が誤ってファイルを変更することを防ぎ、レビュー作業にさらに集中できるようになります。
### 4. 詳細なシステム プロンプトを作成する
システム プロンプトが詳細であればあるほど、サブエージェントのパフォーマンスは向上します。
* 明確な役割定義
* 具体的な作業手順
* 重要なチェックリスト
* 出力形式の要件
### 5. バージョン管理
プロジェクトレベルのサブエージェントを Git にコミットします。
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
チーム メンバーは、プロジェクトのクローンを作成した後、同じサブエージェントを自動的に取得します。
## 一般的な問題のトラブルシューティング
| 問題 | 考えられる原因 | ソリューション |
| ----------------- | ------------------------------------- | --------------------------------------------------------------- |
| サブエージェントは呼び出されません | 説明が十分に明確ではありません。トリガーワードを追加してより具体的にします | |
| サブエージェントが呼び出されません | ファイルの場所が間違っています | ファイルが `.claude/agents/` または `~/.claude/agents/` にあることを確認してください。 |
| ツールは利用できません | ツールフィールド設定エラー | ツール名のスペルをチェックし、カンマで区切られていることを確認してください。 |
| 出力が不安定です | システムプロンプトが曖昧すぎる | 特定の手順と出力形式の要件を追加する |
| コンテキストが失われました | セッションが終了しました | 再開可能な実行の使用 |
| 名前の競合 | 複数の場所に同じ名前のサブエージェント | `/agents` を使用してどれがアクティブであるかを確認します。 |
## チームと共有する
### 方法 1: Git を使用する
Subagent を `.claude/agents/` ディレクトリに配置し、プロジェクト リポジトリに送信します。チームメンバーはクローン作成後に自動的に取得されます。
### 方法 2: プラグインを使用する
Subagent を Plugin の `agents/` ディレクトリに配置し、Plugin メカニズムを通じて配布します。
### 方法 3: ユーザーレベルの共有
よく使用されるサブエージェントを `~/.claude/agents/` に配置して、すべてのプロジェクトで使用できるようにします。複数のマシン間の同期は、ドットファイルを使用して管理できます。
## 学習リソース
### 公式ドキュメント
| リソース | リンク | 説明書 |
| --------------- | ----------------------------------------------------------------------------------------------------------- | -------------------- |
| クロードコードのドキュメント | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 公式ドキュメントのエントリ |
| サブエージェントガイド | [クロード コード ドキュメント](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | サブエージェント構成の詳細 |
| エージェントのデザインパターン | [人類ドキュメント](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 6 つのコア設計パターン |
| 多剤研究 | [人類工学](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2%のパフォーマンス向上の研究詳細 |
### コミュニティリソース
| リソース | リンク | 説明書 |
| --------------- | ----------------------------------------------------------------- | ------------------------------- |
| wshobson/エージェント | [GitHub](https://github.com/wshobson/agents) | 99 エージェント + 15 オーケストレーター テンプレート |
| 化合物工学 | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 の専門エージェント用のプラグイン |
| 素晴らしいクロードコード | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Claude Code のベスト プラクティスの概要 |
### 推奨読書
| 記事 | 出典 | トピック |
| ---------------------------- | -------------------------------------------------------------------------------------- | ------------------- |
| 効果的なエージェントを構築する | [人類](https://www.anthropic.com/research/building-effective-agents) | エージェントの設計原則 |
| マルチエージェント研究システムをどのように構築したか | [人類](https://www.anthropic.com/engineering/multi-agent-research-system) | マルチエージェントアーキテクチャの実践 |
| クロードのスキル、コマンド、サブエージェント、プラグイン | [ヤングリーダーテック](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | 機能比較分析 |
## 概要
Claude Code Subagent は、AI プログラミングの効率を向上させる強力なツールです。これにより、独立したコンテキストと特殊な構成を通じて複雑なタスクを管理できるようになります。
クイックスタート:
1. `/agents` を実行して管理インターフェイスを開きます
2. 単純なサブエージェント (コードレビューアなど) を作成します。
3. 自動委任と明示的な呼び出しをテストする
4. 必要に応じて構成を調整します
使い方を深めていくと、徐々に次のことができるようになります。
* チーム専用のサブエージェントを作成します
* 複雑なワークフローを処理するためにサブエージェント リンクを構成する
* 再開可能な実行で長期タスクを処理する
他の構成で Subagent をパッケージ化して配布する場合は、「[クロード コード プラグイン実践ガイド](/ja/docs/notes/claude-plugin/practice)」を参照してください。
# 上級編
## ターミナル通知:タスク完了のアラート
Claude がタスクを完了したときに通知を受け取りたいですか?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
iTerm2 の通知機能と組み合わせるか、`terminal-notifier` を使ってカスタム通知を設定できます([ベストプラクティス](/ja/blog/claude-code-best-practices)の Hooks 設定を参照)。
## Hooks の高度な使い方
Hooks はシェルコマンドを実行するだけではありません。実際には4つのタイプがあります:
1. **command**:Shell 命令(最常见)
2. **http**:POST JSON 到 URL(支持自定义 headers 和环境变量展开)
3. **prompt**:发给 Claude 评估(比如「所有任务都完成了吗?」)
4. **agent**:启动一个有工具访问权限的子代理来验证
知っておくべき高度な Hook イベント:
* `PostCompact`:压缩完成后触发,适合注入提醒让 Claude 重新读取关键文件
* `SessionStart`:写入 `$CLAUDE_ENV_FILE` 可以给整个会话持久化环境变量
* `PreToolUse`:可以修改工具输入(`updatedInput`),甚至自动批准或拒绝操作
## プラグインエコシステム
`/plugin` を使ってコミュニティプラグインを閲覧・インストールできます。注目のプラグイン:
* **dx**(by ykdojo):提供 `/handoff`(自动写交接文档)、`/clone`(克隆对话)、`/half-clone`(只克隆最近的对话减少上下文)
* **mine**(by anipotts):把所有 Claude Code 会话数据导入 SQLite,支持成本追踪、缓存分析、错误记忆等查询
## Agent Teams:マルチエージェント協調
環境変数を設定して、実験的な Agent Teams 機能を有効にします:
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
有効にすると、セッションが Team Lead として機能し、git worktree を通じて複数のエージェントを同時に協調させることができます。各エージェントは独自のコンテキストウィンドウで独立して実行されるため、大規模プロジェクトの並行開発に最適です。
ただし、トークン消費量は4〜15倍に増加するため、状況に応じて使い分けてください。
## プロンプト戦略
以下は Boris Cherny が Twitter で共有したチームの実践から得られたものです。Claude Code における「プロンプトエンジニアリング」のベストプラクティスと言えます。
### Claude をコードレビュアーとして活用する
Claude にコードを書かせるだけでなく、あなたのコードをレビューさせましょう:
```
Grill me on these changes and don't make a PR until I pass your test.
```
または、コードが動作することを証明させましょう:
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### 不満な回答を言い換えて再質問しない
Boris のヒント #6:Claude が平凡な回答をした場合、言い回しを変えてもう一度質問しないでください。代わりに「この解決策は十分ではない、具体的にどこを改善できるか教えて」と言いましょう。既存の回答を改善する方がゼロから始めるより効果的です。
### Claude 自身に CLAUDE.md を更新させる
ミスを修正した後に、次の一言を加えましょう:
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris によると、Claude は自分自身のルールを書くことが驚くほど上手だそうです。時間が経つにつれて CLAUDE.md はますます正確になり、会話の品質も継続的に向上します。
### 「fix」とだけ言う
Slack MCP を有効にした状態で、Slack のバグレポートを貼り付けて一言だけ:**fix**。コンテキストスイッチはゼロです。
または CI が失敗した場合、単純に:
```
Go fix the failing CI tests.
```
手動でログを分析したり、問題を説明する必要はありません。Claude にログを確認させ、問題を特定させ、修正させましょう。
## おわりに
Claude Code は非常に速いペースで進化しており、これらのテクニックも常に改善されています。公式 Changelog をフォローして最新情報を把握することをお勧めします。
以前の記事をまだ読んでいない方は、基本的なワークフローから始めることをお勧めします:
### 関連記事
* [私の Claude Code ベストプラクティス](/ja/blog/claude-code-best-practices) — ワークフローの核心テクニックとスラッシュコマンドガイド
* [AIプログラミング品質管理:コード品質を保証する5つの防衛線](/ja/blog/claude-code-quality-control) — Claude Code プログラミングにおける品質保証体制
* [Claude システムアーキテクチャ完全解説](/ja/docs/notes/claude-architecture) — MCP、Skills、Subagents、Hooks などのコンポーネントを理解する
# 実用コマンドと自動化
## `/diff`:インタラクティブ Diff ビューア
`/diff` と入力すると、インタラクティブな diff ビューが開きます:
* **左右矢印キー**:git diff(全変更)と Claude の各ターンごとの変更を切り替えます
* **上下矢印キー**:異なるファイルを閲覧します
ターミナルで `git diff` を実行するよりもはるかに使いやすく、特に複数ファイルにまたがる変更がある場合に便利です。
## `/simplify`:マルチエージェントコードレビュー
`/simplify` を実行すると、3つの並列レビューエージェントが同時に起動します:
* **コード再利用**エージェント:重複パターンを検出します
* **コード品質**エージェント:可読性と構造をチェックします
* **効率性**エージェント:不要なパフォーマンスオーバーヘッドを分析します
3つのエージェントは独立して作業し、最後に結果を集約して、有効な問題を自動修正し、誤検知をスキップします。
## `/security-review`:セキュリティスキャン
現在のブランチの変更に対してセキュリティ監査を実施し、SQLインジェクション、XSS、認証の欠陥、データ処理の問題、依存関係の脆弱性をチェックします。各検出結果は敵対的検証を経て、誤検知を減らします。
## `/copy` の隠れた機能
`/copy` は最後の返答をコピーするだけではありません。返答にコードブロックが含まれている場合、インタラクティブなセレクターがポップアップし、返答全体ではなく特定のコードブロックを選択できます。また、数字を渡すことで過去の返答をコピーすることもできます:`/copy 2` で最後から2番目、`/copy 3` で最後から3番目をコピーでき、スクロールして手動で選択する必要がありません。
## `/batch`:大規模並列リファクタリング
```
/batch 把 src/ 下所有组件从 Class 组件迁移到函数组件
```
これは強力な機能です。`/batch` はコードベースを分析し、タスクを5〜30の独立したユニットに分解し、各ユニットごとに独立したエージェントが隔離された git worktree で作業を行い、最後に各エージェントがコミットして PR を作成します。
大規模なマイグレーション、一括型注釈の追加、グローバルなリネームなどのシナリオに適しています。
## `/loop`:定時タスク
```
/loop 5m 检查部署是否完成
/loop 1h /review-pr 1234
```
セッション内に指定した間隔で繰り返し実行される定時タスクを作成します。デプロイ状態のポーリングや、PR の定期チェックなどに便利です。セッションレベル(終了するとなくなります)で、最大50タスク、3日後に自動的に期限切れになります。
## パイプ入力:何でも Claude に渡す
```bash
# 让 Claude 分析错误日志
cat error.log | claude -p "分析这个错误日志,找出根本原因"
# 让 Claude 总结最近的改动
git diff HEAD~3 | claude -p "总结这三次提交的改动"
# 让 Claude 解读命令输出
kubectl get pods | claude -p "哪些 pod 状态异常?"
```
`-p` は Headless モード(非対話型)で、スクリプトや CI/CD パイプラインでの使用に適しています。
## Headless モードの隠しパラメータ
`-p` モードには非常に強力ですが、あまり知られていないパラメータがあります:
```bash
# 设置花费上限(超过就停)
claude -p --max-budget-usd 5.00 "重构认证模块"
# 限制对话轮数
claude -p --max-turns 3 "修复这个测试"
# 输出 JSON 格式(方便程序解析)
claude -p --output-format json "分析这个项目"
# 要求输出符合特定 JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "总结项目"
# 多轮 headless 对话(用 session-id 保持上下文)
claude -p --session-id my-task "第一步:分析代码"
claude -p --session-id my-task "第二步:生成测试"
# 指定备用模型(主模型过载时自动切换)
claude -p --fallback-model sonnet "复杂分析"
# 限制可用工具
claude -p --tools "Read,Grep,Glob" "只读分析,不要改代码"
# 完全替换系统提示词
claude -p --system-prompt "你是一个 Python 专家" "优化这段代码"
```
# 設定とトラブルシューティング
## `/statusline`:カスタムステータスバー
`/statusline` を使って、自然言語の説明で下部ステータスバーに表示する情報をカスタマイズできます。または、`~/.claude/statusline.sh` スクリプトを手動で作成することもできます。
表示できる情報には、現在のモデル、git ブランチ、未コミットのファイル数、コンテキスト使用量の進捗、セッションコストなどがあります。複数の Claude ウィンドウを開いて異なるタスクを処理している場合、ステータスバーで各ウィンドウが何をしているかをすぐに識別できます。
## settings.json の自動補完
settings.json の先頭に `$schema` を追加すると、VS Code / Cursor が設定項目の自動補完とバリデーションを提供します:
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## 便利な隠し設定
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration`:各会話ターンの所要時間を表示
* `DISABLE_AUTOUPDATER`:自動アップデートチェックを無効にし、コンテキストのオーバーヘッドを削減
## `/stats` と `/insights`:使用分析
* `/stats`:日次使用量、セッション履歴、使用ストリーク、モデルの好みを可視化し、日付範囲フィルタリングに対応
* `/insights`:すべての Claude Code 履歴を分析し、どのワークフローが効果的か、どこがボトルネックかを教えてくれ、最適化の提案も生成
## history.jsonl:プロンプト履歴
Claude は送信されたすべてのプロンプトを `~/.claude/history.jsonl` に保存します。Claude にこのファイルを分析させて、プロンプトのパターンや最適化の余地を見つけることができます。
## `/doctor`:ヘルスチェック
奇妙な問題に遭遇した場合、`/doctor`(またはターミナルで `claude doctor`)を実行してください。インストール状態、バージョン、認証状態、システム依存関係をチェックし、問題の特定を素早くサポートします。
## コミュニティツール
コミュニティツール `ccusage` でトークン使用量を追跡できます:
```bash
npx ccusage daily
```
`--dangerously-skip-permissions` を使用したり、多くのコマンドを承認した場合は、`cc-safe` でリスクをスキャンできます:
```bash
npx cc-safe .
```
`.claude/settings.json` 内の `sudo`、`rm -rf`、`chmod 777`、`git reset --hard` などの高リスクコマンドをチェックします。
## その他の知っておくべきスラッシュコマンド
| コマンド | 機能 |
| --------------------- | ----------------------------------------------------- |
| `/export [filename]` | 会話をプレーンテキストとしてエクスポート |
| `/pr-comments [PR]` | PR コメントを取得(現在のブランチを自動検出) |
| `/release-notes` | 現在のバージョンの変更履歴を表示 |
| `/plugin` | コミュニティプラグインの閲覧とインストール |
| `/fast` | 高速モードの切り替え |
| `claude --debug` | 起動時にデバッグログを有効化(カテゴリフィルタリング対応、例:`--debug "api,hooks"`) |
| `/install-github-app` | GitHub App をインストールして PR の自動レビューを実現 |
# コンテキスト管理
## `/compact` は引数を受け付ける
多くの人が `/compact` でコンテキストを圧縮できることは知っていますが、引数を指定して何を保持するかを指定できることはあまり知られていません:
```
/compact 保留所有关于数据库 schema 的讨论,以及当前的重构方案
```
こうすることで、圧縮時に指定したコンテンツが優先的に保持され、重要なコンテキストが失われるのを防げます。
## CLAUDE.md に圧縮サバイバル指示を書く
CLAUDE.md に `## Compact Instructions` セクションを追加して、圧縮時に何を保持すべきかを Claude に伝えます:
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
これにより、自動圧縮でも重要な情報が失われることはありません。
## トークン予算による早期中断を防ぐ
CLAUDE.md に以下を追加します:
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
Claude はコンテキストがほぼ満杯になると、「コンテキストがもうすぐ一杯です」と言って自ら停止することがあります。この記述を追加することで、早期の中断を防げます。
## Handoff プロトコル:セッション引き継ぎ
コンテキストがほぼ満杯だがタスクが完了していない場合、Claude に引き継ぎドキュメントを作成させます:
```
把剩余的计划写到 HANDOFF.md 里,说明你尝试了什么、什么有效、什么没效。
```
その後、新しいセッションを開き、`@HANDOFF.md` とするだけで完全なコンテキストを復元できます。10K+ トークンのコンテキストが 2K 未満に圧縮され、`/compact` よりもはるかに正確です。
## 70-80% の時点で積極的に圧縮する
見落としがちなポイント:コンテキストが上限に近づくと、Claude は自動的に圧縮をトリガーします。しかし、タスクの途中で自動圧縮が発生すると、重要な情報が失われ、その後の応答品質が低下する可能性があります。
より良いアプローチは**積極的な管理**です:コンテキストが 70-80% に達した時点で手動で `/compact` を実行すると、自動圧縮を待つよりもはるかに効果的です。タスクを完了したらすぐに `/clear` を実行し、コンテキストが無限に膨張するのを防ぎましょう。
環境変数を使って自動圧縮を早めにトリガーすることもできます:
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context`:コンテキスト診断
コンテキストウィンドウにどれだけのスペースが残っているか分からない場合、`/context` が教えてくれます:
* どのツールや MCP サービスが最もコンテキストを消費しているか
* 現在の容量使用率
* 的を絞った最適化の提案
一部の MCP サービスを登録しただけで(使用していなくても)、コンテキストウィンドウの 30% 以上を消費することがあることに気づきました。`/context` で確認し、使っていない MCP をクリーンアップすれば、かなりのスペースを解放できます。
## MCP ツールの自動遅延読み込み
MCP ツール定義がコンテキストの 10% を超えると、Claude Code は自動的に Tool Search を有効にします。完全なツール定義の代わりに軽量な検索インデックスを読み込みます。これにより MCP のコンテキスト消費を 85% 以上削減できます(例えば 77K トークンから 8.7K に)。この機能は**デフォルトで有効**であり、手動設定は不要です。
注意点:Tool Search は Sonnet 4+ と Opus 4+ モデルのみをサポートし、Haiku はサポートしていません。`ANTHROPIC_BASE_URL` が非公式プロキシを指している場合、Tool Search は自動的に無効化されます(ほとんどのプロキシが `tool_reference` ブロックを転送しないため)。
動作をカスタマイズしたい場合は、settings.json で設定できます:
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
サポートされる設定値:
* **未設定**:デフォルトで有効
* **`true`**:強制有効化(非公式プロキシのシナリオを含む)
* **`auto`**:コンテキストが 10% を超えた時にアクティブ化(デフォルトの動作と同等)
* **`auto:`**:カスタムしきい値、例えば `auto:5` は 5% を超えた時にアクティブ化
* **`false`**:無効化、すべての MCP ツールがプリロードされる
# Claude Code 隠れた便利テクニック集
創設者 Boris Cherny のツイートやコミュニティ、changelog から厳選した Claude Code の実用テクニックです。「こんな機能あったの?」というような隠れた便利操作が満載です -- ショートカット、隠し機能、コマンドラインのテクニックなど、一度使い始めるともう戻れません。
## 目次
* [ショートカット編](./shortcuts) -- Shift+Tab でモード切替、Esc+Esc で履歴呼び出し、Ctrl+S でスタッシュなど
* [入力とインタラクション編](./input-interaction) -- `!` でターミナルコマンド、`@` でファイル注入、URL ペースト、/btw で割り込み、Vim モード
* [思考とモデル制御編](./thinking-model) -- think/ultrathink キーワード、/effort、subagents、opusplan
* [セッション管理編](./session-management) -- /rename、/branch、/color、リモートコントロール
* [コンテキスト管理編](./context-management) -- /compact のパラメータ、圧縮ディレクティブ、Handoff プロトコル、MCP 遅延読み込み
* [コマンドと自動化編](./commands-automation) -- /diff、/simplify、/batch、/loop、Headless モード
* [設定と診断編](./config-diagnostics) -- statusline、settings.json、/stats、/doctor
* [上級編](./advanced) -- Hooks の応用、プラグインエコシステム、Agent Teams、プロンプトの心得
### 関連記事
* [Claude システムアーキテクチャ完全解説](/ja/docs/notes/claude-architecture) -- MCP、Skills、Subagents、Hooks などのコンポーネントを理解する
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) -- サブエージェントの概念と実践
# 入力とインタラクション
## `!`:直接运行终端命令
在输入框以 `!` 开头,可以直接在 Claude Code 内执行终端命令,不需要切换到另一个终端窗口:
```
! git status
! npm run build
! docker ps
```
输入 `!` 加命令前缀后按 Tab 还能自动补全历史命令。
## `@` + 文件路径:注入文件上下文
在输入时用 `@` 加文件路径,可以把文件内容直接注入到上下文中:
```
帮我看看 @src/auth/login.ts 和 @src/auth/middleware.ts 之间的逻辑有没有问题
```
支持 Tab 键自动补全路径,不需要手动输入完整路径。比让 Claude 自己去读文件更快,因为省去了工具调用的开销。
## 直接粘贴 URL
直接把 URL 粘贴到输入中,Claude 会自动抓取网页内容作为上下文:
```
参考这个 API 文档 https://docs.example.com/api/v2 来写客户端代码
```
## 喂 `/llms-full.txt` 让 Claude 自己查文档
很多开源项目的文档站点会提供 `/llms-full.txt` 文件(LLM 友好的完整文档)。遇到某个库的问题时,把这个文件的 URL 粘贴给 Claude,它能自己查文档解决绝大部分问题:
```
参考 https://docs.astro.build/llms-full.txt 帮我解决这个路由问题
```
## `/btw`:在 Claude 工作时插嘴
这是 2026 年 3 月刚加的新功能。当 Claude 正在执行任务时,你可以用 `/btw` 发起一个旁路对话——问问它在想什么、给它补充信息,而不需要打断当前任务。
正如 Anthropic 工程师 @trq212 在推特上说的:「没人会用 Ctrl+C 打断同事,你只需要说一句 'btw',他们就会抬头看你。」
## Vim 模式
输入 `/vim` 开启 Vim 模式,支持:
* 模式切换(Normal/Insert)
* 导航(h/j/k/l, w/b/e, 0/$)
* 编辑操作(d, c, y, p)
* 文本对象(iw, aw, i", a())
如果你是 Vim 用户,这比默认的输入体验好太多。用 `/config` 可以设置为永久开启。
## 语音模式
输入 `/voice` 激活语音模式,长按空格键说话,松开发送。适合不想打字但又需要给 Claude 交代任务的时候。按键可以在 `keybindings.json` 中自定义。
# セッション管理
## `/rename`:セッションに名前を付ける
```
/rename my-auth-refactor
```
現在のセッションに名前を付けます。名前を付けるメリットは、対話式セッションセレクター(`claude --resume`)で、名前付きセッションを直接選択して復元でき、Enter キーで確認する必要がないことです。ターミナルから `claude --resume my-auth-refactor` で直接起動することもできます。
セレクターではテキストを入力するだけで検索・フィルタリングが可能で、以下のショートカットキーにも対応しています:`Ctrl+V` でセッションをプレビュー、`Ctrl+R` でリネーム、`Ctrl+A` で全プロジェクト表示の切り替え、`Ctrl+B` でブランチによるフィルタリング。
## `/branch`:会話を分岐させる
git のブランチと同じように、`/branch` は現在の会話ポイントにフォークを作成します。フォーク内で異なるアプローチを試すことができ、元の会話には影響しません。うまくいかなければ、元のブランチに戻って続けるだけです。
## `/color`:ウィンドウに色を付ける
現在のセッションのプロンプトバーに色を設定します。red、blue、green、yellow、purple、orange、pink、cyan に対応しています。
## コマンドラインセッション管理ツールキット
```bash
# 恢复当前目录最近的会话
claude --continue
# 打开会话选择器,或按名称恢复
claude --resume
claude --resume my-auth-refactor
# 启动时直接命名会话
claude -n "auth-refactor"
# Fork 上一次会话(保留上下文,创建新分支)
claude -c --fork-session
# 恢复与特定 PR 关联的会话
claude --from-pr 123
# 在隔离的 git worktree 中启动
claude -w
```
Claude Code のセッションは自動保存されます(Ctrl+S は不要)。ターミナルを開くたびに `--continue` で前回の作業を再開できます。
## `claude --remote`:デバイス間で続行
```bash
claude --remote "your task description"
```
Web セッションを開始し、claude.ai やモバイルアプリで操作を続けることができます。
## `/remote-control`:スマートフォンからローカルの Claude を操作
パソコンの Claude Code で `/remote-control` と入力すると、接続コードが生成されます。次にスマートフォンの Claude アプリでこの接続コードを入力すれば、ローカルの Claude Code セッションをリモートで操作できます。スマートフォンからパソコン上の Claude に指示を出せるのです。
# キーボードショートカット
Claude Code のショートカット体系は、多くの方が想像する以上に充実しています。`?` を押すと、現在のコンテキストで利用可能なすべてのショートカットを確認できます。
## Shift+Tab:モードの循環切替
おそらく最も重要なショートカットです。`Shift+Tab` を押すと、3つのモード間を循環的に切り替えられます:
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
`/plan` や `/auto-accept` を手動で入力する必要はありません。キー1つですべて完了です。私の使い方は、新しいタスクを受け取ったら2回押して Plan Mode に切り替え、方針を確認してからもう1回押して Auto-Accept に切り替え、Claude に自律的に実行させるというものです。
## Esc + Esc:タイムマシン
`Esc` を2回連続で押すと、巻き戻しメニュー(Rewind)が表示されます:
* **コードと会話を復元**:以前のチェックポイントに戻り、ファイルと会話履歴の両方がロールバックされます
* **会話のみ復元**:メッセージをロールバックしますが、現在のコード変更はそのまま保持されます
* **コードのみ復元**:ファイルの変更を元に戻しますが、会話履歴は保持されます
Claude はファイル編集のたびに自動的にチェックポイントを記録します。これは `git checkout .` よりもはるかに細かい粒度で操作できます。最後のコミットに戻るだけでなく、任意の編集ステップに戻ることが可能です。
ただし注意点があります。追跡されるのは Claude がツールを通じて直接編集したファイルのみです。手動で変更したファイルや、`git push` などの外部操作は巻き戻しの対象外です。
## Ctrl+S:プロンプトの一時保存(Prompt Stash)
プロンプトを書いている途中で、先に別の作業を処理する必要が出てきた場合は、`Ctrl+S` を押すと現在の入力が一時保存されます:
その後、別のコマンドや質問を入力できます。そのメッセージを送信すると、一時保存していた内容が入力欄に**自動的に復元**され、中断したところから続けられます。
これは `git stash` のプロンプト版と考えてください。使用例:長いリファクタリング要件の説明を書いている最中に、まず Claude にあるファイルを確認してもらいたくなった場合、`Ctrl+S` で要件説明を一時保存し、ファイルに関する質問をして、回答が返ってきたら要件説明が自動的に戻ってきます。
## Ctrl+B:タスクをバックグラウンドに送る
Claude が時間のかかるタスク(大規模なリファクタリングなど)を処理中で、別の作業に取りかかりたい場合は、`Ctrl+B` を押すと現在のタスクがバックグラウンドに移動し、ターミナルがすぐに新しい入力を受け付けるようになります。
`Ctrl+T` でバックグラウンドタスクの一覧を確認でき、`Ctrl+F` を2回押すとすべてのバックグラウンドエージェントを終了できます。
> tmux ユーザーへの注意:tmux のデフォルトプレフィックスキーも `Ctrl+B` のため、Claude のバックグラウンド機能を使うには2回押す必要があります。
## Ctrl+G:エディタで長いプロンプトを書く
Claude に詳細な指示を出す必要がある場合、ターミナルでの入力は快適とは言えません。`Ctrl+G` を押すとシステムのデフォルト `$EDITOR`(VS Code、Vim など)が開き、そこでプロンプトを書いて保存・終了すると、自動的に Claude に送信されます。
デフォルトエディタを変更するには、シェル設定ファイル(`~/.zshrc` または `~/.bashrc`)で設定します:
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
`--wait` パラメータは重要です。エディタにファイルを閉じるまで待機するよう指示するもので、これがないと Claude は空の内容を即座に受け取ってしまいます。Vim のようなターミナルエディタは元々ブロッキング動作なので、このパラメータは不要です。
複数段落にわたる要件説明や、大量の参考資料を貼り付ける際に特に便利です。Plan Mode では `Ctrl+G` を使って、Claude が生成した計画をエディタ内で直接編集することもできます。
## Cmd+T:拡張思考の切替
デフォルトのショートカットは `Cmd+T`(Windows/Linux では `Meta+T`)で、拡張思考(Extended Thinking)モードのオン・オフを切り替えます。有効にすると、Claude は応答前により深い推論を行います。複雑なアーキテクチャの検討や難解なバグの調査に最適です。
ただし注意点があります。ほとんどのターミナル(iTerm2、Terminal.app、Warp など)は `Cmd+T` を「新しいタブ」として処理するため、実際にはこのショートカットが機能しないことが多いです。対処法は2つあります:`/keybindings` で競合しないキーに再割り当てするか、`/effort` コマンドを使って思考の深さを切り替えるか(同じ効果で、レベルの細かい制御も可能です)。
## Readline ショートカット
Claude Code の入力欄は標準的な Readline ショートカットに対応しており、ターミナルに慣れた方にはおなじみのものばかりです:
| ショートカット | 機能 |
| ----------------- | ---------------- |
| Ctrl+A | 行頭に移動 |
| Ctrl+E | 行末に移動 |
| Ctrl+W | 前の単語を削除 |
| Ctrl+U | 行頭まで削除 |
| Ctrl+K | 行末まで削除 |
| Ctrl+Y | 最後に削除したテキストを貼り付け |
| Alt+Y | 削除履歴を順に表示 |
| Option+Left/Right | 単語単位で移動(Mac) |
## 承認ショートカット:`y/n/d/e`
Claude がファイル変更を提案して確認を求めている際、4つの単一キーショートカットでフローを制御できます:
* `y`:承認
* `n`:拒否
* `d`:完全な diff を表示
* **`e`:編集してから承認**
`e` は最も見落とされがちですが、最も便利なショートカットです。Claude の変更に微調整を加えてから適用できます。数行のコードが気に入らない場合でも、拒否してやり直す必要はなく、`e` を押して修正すれば完了です。
## クイックリファレンス
| ショートカット | 機能 |
| ----------- | --------------------------------------------------- |
| Shift+Tab | モード切替:Normal → Auto-Accept → Plan |
| Esc+Esc | 巻き戻しメニューを開く |
| Ctrl+S | 現在の入力を一時保存、次の送信後に自動復元 |
| Ctrl+B | 現在のタスクをバックグラウンドに送る |
| Ctrl+T | バックグラウンドタスク一覧を表示 |
| Ctrl+F (x2) | すべてのバックグラウンドエージェントを終了 |
| Ctrl+G | 外部エディタでプロンプトを作成 |
| Ctrl+O | 詳細なツール出力表示の切替 |
| Cmd+T | 拡張思考の切替(ターミナルに横取りされる場合あり。再割り当てまたは `/effort` の使用を推奨) |
| `\` + Enter | 複数行入力(設定不要) |
| Shift+Enter | 複数行入力(事前に `/terminal-setup` の実行が必要) |
| Up / Down | 入力履歴の閲覧 |
| Ctrl+R | コマンド履歴の検索 |
| Ctrl+L | 画面クリア(履歴は保持) |
| Ctrl+C | 現在の生成をキャンセル |
| Ctrl+D | Claude Code を終了 |
| `?` | 利用可能なすべてのショートカットを表示 |
## カスタムキーバインド
デフォルトのショートカットが合わない場合は、`/keybindings` で `~/.claude/keybindings.json` を開いてカスタマイズできます。変更は即座に反映され、再起動は不要です。
コンビネーションキー構文(例:`ctrl+shift+c`)とコードモード(例:`ctrl+k ctrl+s` -- Ctrl+K を押して離し、次に Ctrl+S を押す)に対応しています。16 種類のバインディングコンテキスト(Chat、Autocomplete、Confirmation、DiffDialog など)があり、各コンテキストに異なるアクションを割り当てることができます。
# 思考とモデル制御
## キーワードで思考の深さを制御する
プロンプトに特定のキーワードを追加することで、異なるレベルの思考バジェットをトリガーできます。これは Claude Code 独自の機能です(claude.ai ウェブ版にはありません):
| キーワード | 思考バジェット | 適用シーン |
| ----------------------------- | ----------- | ----------------- |
| `think` | 約4,000トークン | 日常的なコーディングの質問 |
| `think hard` / `megathink` | 約10,000トークン | 複雑なロジック、複数ファイルの関連 |
| `think harder` / `ultrathink` | 約31,999トークン | アーキテクチャ設計、難解なバグ |
実際の使用では、Claude が浅い回答をしたときに `think hard` を追加して再度質問することが多いです。特に複雑な問題(複数のサービスにまたがるバグ調査など)の場合は、直接 `ultrathink` を使います。
## `/effort`:思考の深さを制御する
キーワード(think / ultrathink)以外にも、`/effort` で思考の深さを直接設定できます:
```
/effort low # 简单任务,跳过深度思考,更快更省
/effort high # 复杂任务,深度推理
/effort max # 最大思考预算(仅 Opus)
/effort auto # 让 Claude 自己判断
```
設定はセッション全体を通じて有効です。シンプルなファイル編集には `low`、複雑なアーキテクチャ設計には `max` を使うことで、品質を犠牲にせずコストを節約できます。
## `use subagents` キーワード
任意のリクエストの後に `use subagents` を追加すると、Claude はタスクを複数のサブエージェントに分解して並列処理します。これにより速度が向上するだけでなく、メインエージェントのコンテキストウィンドウをクリーンに保つことができます。
Boris は Twitter でこの点を特に言及しています:個別のタスクをサブエージェントに委任し、メインエージェントのコンテキストを集中させるということです。
## `opusplan`:最もコストパフォーマンスの良いモデル戦略
一言でまとめると:**Opus が考え、Sonnet が実行する**。
## `/model`:モデルを切り替える
`/model` を使うと、セッション中にいつでもモデルを切り替えることができます。例えば、普段は Sonnet を使い、複雑な問題に遭遇したら一時的に Opus に切り替え、解決したら元に戻すといった使い方ができます。
## 出力スタイルの制御
`/config` で「Output style」を選択すると、あまり知られていないものの非常に便利な2つのモードがあります:
* **Explanatory モード**:Claude がタスクの合間に「ナレッジポイント」を挿入し、関連するフレームワークやコードパターンを解説します — 新しいプロジェクトを学ぶのに最適です
* **Learning モード**:協調学習モードで、Claude が直接答えを出す代わりに、コード内に `TODO(human)` マーカーを追加して、あなた自身で実装するよう促します
`~/.claude/output-styles/` にカスタム出力スタイルファイル(Markdown 形式)を作成して、システムプロンプトを直接修正することもできます。注意:カスタム出力スタイルは、`keep-coding-instructions: true` を設定しない限り、デフォルトのコーディングシステムプロンプトを**完全に置き換えます**。
# コンセプト紹介
## はじめに
2025年10月、Anthropic は Claude Skills という新機能をひっそりとリリースしました。この一見控えめなアップデートは、著名な技術ブロガー Simon Willison に「MCP より重要かもしれない」と評価され、AIツール分野に「カンブリア爆発」をもたらすと予測されました。
このような評価は根拠のないものではありません。AIアシスタントを頻繁に使用していれば、こんな悩みに遭遇したことがあるでしょう:新しい会話を始めるたびに同じワークフローの説明を繰り返し入力しなければならない。せっかくAIを満足のいく状態に調整しても、別の会話ウィンドウに切り替えるとまたゼロからやり直し。Skills はまさにこの課題を解決するために生まれました。
## Claude Skills を理解する
あなたが会社の社長だと想像してください。新入社員が入社する際、仕事の進め方、ブランド規範、よくある問題の処理方法が詳しく記載されたワークマニュアルを渡します。Claude Skills はAIアシスタントに渡すこの「ワークマニュアル」であり、再現可能で標準化された方法で特定のタスクを完了できるようにするものです。
技術的に言えば、Skills は指示、スクリプト、リソースを含むフォルダで、Claude が必要な時に動的にロードできます。各 Skill は Claude に特定のタイプのタスクを一貫した方法で完了する方法を教え、これらの知識は会話をまたいで永続的に保存されます。つまり、一度「トレーニング」するだけで、その後いつ使用しても Claude はどうすべきかを覚えています。
### 3つの構成要素
完全な Skill は以下の3つの部分で構成されます:
| コンポーネント | 役割 | 必須か |
| ------------ | ------------------------------------- | --- |
| **SKILL.md** | コア指示ドキュメント、メタデータと詳細な指示を含む | 必須 |
| **参考資料** | ブランドガイドライン、ポリシー文書、テンプレートなどの補足情報 | 任意 |
| **スクリプト** | Python/JavaScript コード、複雑な計算やファイル操作を処理 | 任意 |
このうち SKILL.md は Skill 全体の「魂」であり、基本構造は以下の通りです:
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
ファイル先頭の YAML frontmatter には2つのキーフィールドがあります:`name` はスキルの識別名で最大64文字。`description` はこのスキルが何をするか、いつ使うべきかを Claude に伝え、最大200文字です。Claude はまさにこの説明に基づいて、いつ特定の Skill を呼び出すべきかを判断するため、明確で正確に書くほど、Skill が正しくトリガーされる確率が高くなります。
### 適用シナリオ
Skills の適用シナリオは非常に広範で、日常業務のさまざまな繰り返しタスクをカバーしています:
**ドキュメント処理**:Excel スプレッドシート、PPT プレゼンテーション、Word 文書、PDF レポートの一括作成。Anthropic 公式もドキュメントスキルのセットを提供しており、すぐに使えます。
**ブランドコンプライアンス**:会社のブランドカラー、ロゴ使用ルール、スペース規範、トーン&マナーを Skill にパッケージ化し、AIが生成するすべてのコンテンツがブランド基準に準拠することを保証します。
**議事録**:会議記録の自動要約、アクションアイテムの抽出、担当者の割り当て、フォローアップメールの生成。
**データ分析**:標準化された分析プロセスの実行。例えば競合インテリジェンススキャン(製品アップデート、価格変更、アナリストコメントの構造化抽出)、財務分析(決算書分析、財務モデル構築)。
**プロジェクト管理**:目標からプロジェクト計画の構築、マイルストーンの提案、週報/投資家ブリーフィングの生成。
## 段階的開示アーキテクチャ
Skills の最も巧妙な設計は、その情報ロード方式にあります。従来の MCP ツール説明は数千から数万の token を消費する可能性がありますが、Skills のメタデータはわずか数十 token しか占有しません。つまり、大量の Skills を同時に有効化しても、コンテキストウィンドウがツール説明で埋め尽くされる心配は全くありません。
この効率は**段階的開示**(Progressive Disclosure)と呼ばれるアーキテクチャ設計に由来します。Skills は3層の情報構造を採用し、必要に応じて段階的にロードします——目次付きのマニュアルのようなものです:
```
📚 Skills ワークブック
│
├─ 📋 目次 ─────────────────────────── 【メタデータ層】起動時にプリロード
│ │
│ │ name: "weekly-report"
│ │ description: "作業内容に基づいて標準化された週報を生成"
│ │
│ │ ✓ わずか 30-50 tokens
│ │ ✓ すべての Skills の目次が同時に表示可能
│ │
│
├─ 📖 本文 ─────────────────────────── 【コアドキュメント層】関連時にロード
│ │
│ │ # Weekly Report Generator
│ │
│ │ ## Instructions
│ │ 以下の構成で週報を生成...
│ │
│ │ ## Examples
│ │ 入力:今週はログイン機能を完成...
│ │ 出力:### 今週の成果 ...
│ │
│ │ ⚡ Claude が必要と判断した時のみ展開
│ │ 📊 数百から数千 tokens を消費
│ │
│
└─ 📎 付録 ─────────────────────────── 【参照リソース層】必要時にロード
│
│ references/
│ ├── brand-guide.md ブランド規範
│ ├── template.xlsx レポートテンプレート
│ └── examples/ 過去の週報
│
│ 🔍 明確に必要な時のみロード
│ 📦 大量の参考資料を含められる
```
まず目次を見てどんな章があるかを知り(メタデータ層)、必要な章を見つけたら開いて読み(コアドキュメント層)、さらに詳細が必要なら付録を参照します(参照リソース層)。
| 層 | 内容 | ロードタイミング | Token 消費 |
| ------------- | ------------------ | --------- | -------- |
| **メタデータ層** | name + description | 起動時にプリロード | 30-50 |
| **コアドキュメント層** | SKILL.md 全文 | 関連時にロード | 数百から数千 |
| **参照リソース層** | 参考ファイル、テンプレートなど | 必要時にロード | 必要に応じて |
これは大規模言語モデルの本質——「テキストを入力してモデルに理解させる」——と完全に一致しています。Skills は複雑なプロトコルや API 呼び出しを導入せず、精巧に組織されたテキスト構造を通じて、AIが効率的に知識を取得し活用できるようにしています。Simon Willison はこの設計を「呆れるほどシンプル」と評しましたが、それはまさに複雑な問題を最も素朴な方法で解決しているからです。
## コア優位性
### Token 効率
Skills の段階的開示アーキテクチャは極めて高い token 効率をもたらします。シンプルな比較で理解できます:
| 方法 | 起動時の Token 消費 | 100個のスキルの総消費 |
| ------------- | ------------- | ------------------ |
| 従来の方法(全量ロード) | 数千から数万 | コンテキストウィンドウを超える可能性 |
| Skills(段階的開示) | 30-50 | 3000-5000 |
各スキルのメタデータがわずか数十 token しか占有しないため、数十から数百の Skills を同時に有効化でき、完全なコンテンツは必要に応じてロードされ、貴重なコンテキストスペースを浪費しません。
### コンポーザビリティ
複数の Skills が自動的に協調動作できます。複雑なタスクを提示すると、Claude はどの Skills を呼び出すべきかをインテリジェントに識別し、それらを協調させてタスクを完了します。
例えば「この売上データから四半期レポートを作成して」と言うと、Claude は以下のように動作する可能性があります:
1. データ分析 Skill を呼び出して生データを処理
2. チャート生成 Skill を呼び出してビジュアライゼーションを作成
3. ドキュメント Skill を呼び出して最終レポートを生成
プロセス全体を通じて、どの Skill を使用するか手動で指定する必要はなく、Claude がタスクのニーズに応じて自動的に選択し組み合わせます。
### ポータビリティ
同じ Skill を Anthropic エコシステムのすべてのプラットフォームで使用できます:
| プラットフォーム | 説明 |
| ----------- | -------------------- |
| Claude.ai | Web版、一般ユーザー向け |
| Claude Code | コマンドラインツール、開発者向け |
| API | プログラマティック統合、システム開発向け |
チーム用に作成したブランドライティング Skill は、これらすべてのプラットフォームで一貫した動作を維持でき、真に**一度構築、どこでも使用**を実現します。
> **他のAIプラットフォームの戦略**:現在、Skills は Anthropic 独自の機能です。OpenAI は Custom GPTs + Assistants API のデュアルトラック戦略を採用し(2つのシステムが統一されていない)、Microsoft Copilot と Google Gemini はそれぞれのエコシステムの深い統合に注力しており、再利用可能なスキルモジュールではありません。Claude Skills は意味のある差別化特性と考えられています。
### 効率データ
Anthropic の内部ベンチマークテストによると、Skills を使用するチームは**繰り返しのプロンプトエンジニアリング時間を73%削減**しました。これは効率の向上だけでなく、ワークフローの標準化と再利用性も意味しています——チームメンバーがそれぞれ独自のプロンプト集を維持する必要がなくなり、検証済みの同一の Skills セットを共有できます。
## まとめ
Claude Skills は本質的にAIアシスタントのための**再利用可能なワークブック**です。段階的開示アーキテクチャにより極めて高い token 効率を実現し、AIが貴重なコンテキストスペースを占有することなく大量の専門知識をマスターできるようにしています。
3つのキーワードを覚えれば、Skills のエッセンスを把握できます:
| キーワード | 意味 |
| ----------- | --------------------------- |
| **効率的** | メタデータはわずか数十 token、必要に応じてロード |
| **コンポーザブル** | 複数の Skills が自動的に協調動作 |
| **ポータブル** | クロスプラットフォームで一貫した体験 |
コンセプトを理解したら、次の《[Claude Skills 実戦ガイド](/ja/docs/notes/claude-skills/practice)》で実践に移りましょう:Skills の有効化とインストール方法、最初のカスタム Skill の作成方法、そしてよくある落とし穴の回避方法を学びます。
ワークフローのさらなる標準化を目指すなら、《[仕様駆動開発とは](/ja/docs/notes/speckit/concept)》を参考にして、AIプログラミングを「直感」から「エンジニアリング」へとアップグレードする方法を学びましょう。
# 実践ガイド
## クイックレビュー
[前回の記事](/ja/docs/notes/claude-skills/concept)では、Skills のコアコンセプトを学びました:AIアシスタントのための再利用可能なワークブックであり、段階的開示アーキテクチャにより極めて高い token 効率を実現し、効率的、コンポーザブル、ポータブルという3つの特徴を持っています。この記事では実戦の視点から、Skills と他の機能との違いを理解し、Skills の有効化、インストール、作成を学び、ベストプラクティスをマスターしてよくある落とし穴を回避します。
## 機能比較
Claude エコシステムにはさまざまな機能があり、初めて触れると区別が混乱するかもしれません。以下の表で素早く区別できます:
| 機能 | 何か | 最適な用途 | 永続性 |
| ------------- | --------- | ----------------- | -------------- |
| **Skills** | 専門知識パッケージ | 繰り返しタスク、標準化プロセス | 会話をまたいで永続 |
| **Prompts** | 即時指示 | 一回限りのリクエスト | 現在の会話のみ |
| **Projects** | ナレッジベース | 背景情報、プロジェクトドキュメント | プロジェクトワークスペース内 |
| **MCP** | コネクタ | 外部データ、API 呼び出し | 持続的接続 |
| **Subagents** | サブエージェント | タスク委託、並行処理 | セッションをまたいで |
### Skills vs MCP
最もよくある混乱です。核心的な違い:**MCP は Claude をデータに接続し、Skills は Claude にデータの処理方法を教えます**。両者は補完関係であり代替関係ではありません。
| 次元 | Skills | MCP |
| ------------ | ------------------------ | --------------------------- |
| **コア機能** | Claude にタスクの実行方法を教える | Claude を外部システムに接続 |
| **Token 消費** | 極めて低い(数十 token) | 高め(数千から数万 token) |
| **技術的複雑さ** | シンプル(Markdown + YAML) | 複雑(完全なプロトコル仕様) |
| **典型的なシナリオ** | ブランドライティング、レポート生成、ワークフロー | データベースクエリ、API 呼び出し、クラウドサービス |
| **ポータビリティ** | Claude.ai/Code/API をまたいで | 複数のモデル企業が採用済み |
この違いを理解すれば、いつどちらを使うべきかが分かります。データベースのクエリ、API の呼び出し、クラウドサービスへのアクセスには MCP を使用。特定のライティングスタイルに従う、標準化されたプロセスを実行する、専門知識を再利用する場合は Skills を使用します。
ベストプラクティスは両方を組み合わせて使用すること:MCP で CRM システムに接続して顧客データを取得し、Skills でそのデータの分析方法とレポート生成方法を定義します。
### Skills vs Subagents
核心的な違い:**Skills は Claude を特定のタスクに強くする。Subagents は Claude にタスクを独立した「専門スタッフ」に委任させる**。
| 次元 | Skills | Subagents |
| --------------- | ------------------------ | -------------------------- |
| **コア機能** | 専門知識と指示を提供 | 独立してタスクを実行するサブエージェント |
| **コンテキスト** | メイン会話のコンテキストに注入 | 独立したコンテキストウィンドウを所有 |
| **適用シナリオ** | Claude を特定のタスクに強くする | 複雑なマルチステップの独立タスク |
| **アクティベーション方法** | 説明に基づいて自動マッチング | 手動呼び出しまたは Claude の自動委任 |
| **ポータビリティ** | Claude.ai/Code/API をまたいで | Claude Code と Agent SDK のみ |
わかりやすく言えば、Skills はトレーニング資料——Claude に特定のことのやり方を学ばせる。[Subagents](/ja/docs/notes/claude-subagent) は専任スタッフ——独自のデスク(コンテキスト)と権限(ツール)を持ち、独立してタスクを完了して結果を報告します。
両者は組み合わせて使用できます:例えばコードレビュー用のサブエージェントに言語固有のベストプラクティス Skill をロードさせ、「専門家 + 専門知識」の組み合わせ効果を実現できます。[Anthropic のリサーチ](https://www.anthropic.com/engineering/multi-agent-research-system)によると、マルチエージェントシステム(Claude Opus 4 メインエージェント + Claude Sonnet 4 サブエージェント)は内部評価でシングルエージェントより90.2%高い成績を記録しました。
### Skills vs スラッシュコマンド
Claude Code を使ったことがあれば、`/commit` や `/review` のような[スラッシュコマンド](/ja/blog/claude-code-best-practices)に馴染みがあるでしょう。核心的な違い:**Skills はコンテキストに基づいて自動アクティベート。スラッシュコマンドは手動入力でトリガー**。
| 次元 | Skills | スラッシュコマンド (Slash Commands) |
| --------------- | -------------------------------- | -------------------------- |
| **アクティベーション方法** | 自動アクティベート(コンテキストマッチングに基づく) | 手動入力(例:`/commit`) |
| **トリガー条件** | Claude が description に基づいて関連性を判断 | ユーザーが明示的にコマンドを入力 |
| **適用シナリオ** | 「常時オン」の能力強化 | 明確で繰り返し可能な操作 |
| **ユーザーの認知** | 無意識のうちに自動的に効果を発揮 | コマンド名を覚える必要がある |
例えば、`/commit` と入力すると Claude が事前定義されたコミットプロセスを実行する——これがスラッシュコマンド。「週報を書いて」と言うと Claude が自動的に週報生成 Skill を認識してロードし、コマンドを入力する必要がない——これが Skills です。
シンプルに覚えましょう:スラッシュコマンドはショートカットキーで、自分からトリガーする必要がある。Skills はバックグラウンドの知識で、Claude が自動的にいつ使うかを判断します。
### Skills vs Plugins
Plugins は Claude Code の拡張パッケージメカニズムです。核心的な違い:**Skills は自動アクティベートする能力拡張。Plugins はパッケージ化して配布する完全なワークフロー設定**。
| 次元 | Skills | Plugins |
| --------------- | ------------------------------- | -------------------------- |
| **コア機能** | 専門能力の拡張 | ワークフローのパッケージ配布 |
| **アクティベーション方法** | コンテキストに基づいて自動アクティベート | インストール後にコンポーネントが統合 |
| **適用範囲** | クロスプラットフォーム(Claude.ai/Code/API) | Claude Code のみ |
| **含まれるもの** | 指示 + スクリプト + リソース | スラッシュコマンド + hooks + skills |
| **配布メカニズム** | 単独フォルダ | marketplace 経由でインストール |
キーとなる理解:Plugins は Skills を含むことができ(`skills/` ディレクトリ内)、より大きなパッケージングユニットです。Plugin をインストールすると、その中の Skills は自動的にアクティベートされ、スラッシュコマンドはオートコンプリートに表示され、hooks は既存の設定と統合されます。
シンプルに言えば:Skills で Claude の能力を拡張し、Plugins でチーム間に標準化されたワークフロー設定を配布します。
## 実戦チュートリアル
### 方法1:内蔵 Skills を有効にする
これが最もシンプルな入門方法です。Anthropic 公式が実用的なドキュメントスキルのセットを提供しています:
| スキル | 機能 |
| --------------------- | ------------------------------- |
| **Excel (xlsx)** | スプレッドシートの作成、データ分析、チャート付きレポートの生成 |
| **PowerPoint (pptx)** | プレゼンテーションの作成、スライドの編集、プレゼン内容の分析 |
| **Word (docx)** | ドキュメントの作成、コンテンツの編集、テキストのフォーマット |
| **PDF (pdf)** | フォーマットされた PDF ドキュメントとレポートの生成 |
**有効化ステップ**:
1. [Claude.ai](https://claude.ai) にログイン
2. 右上のアバターをクリックし、**Settings** に入る
3. **Capabilities** オプションを見つける
4. 必要なスキルを有効化
有効化後にすぐテストできます:「Q3 の販売予算の Excel 表を作成して、月次明細と合計を含めて」。
> **注意**:Pro、Max、Team または Enterprise プランが必要で、コード実行機能を有効にする必要があります。
### 方法2:コミュニティ Skills をインストール
Claude Code を使用している場合、コマンドでコミュニティが提供する Skills をインストールできます。
**プラグインマーケットプレイス経由でインストール**:
```bash
# 公式 Skills リポジトリを追加
/plugin marketplace add anthropics/skills
# ドキュメントスキルパックをインストール
/plugin install document-skills@anthropic-agent-skills
# サンプルスキルパックをインストール
/plugin install example-skills@anthropic-agent-skills
```
**Skills の保存場所**:
| 場所 | パス | 説明 |
| ------------- | ------------------- | -------------------- |
| 個人 Skills | `~/.claude/skills/` | 自分だけが使用可能 |
| プロジェクト Skills | `.claude/skills/` | git バージョン管理に含め、チーム共有 |
### 方法3:カスタム Skill を作成
これこそが Skills の真の力——自分専用のワークフローを作成します。
**ステップ1:フォルダ構造を作成**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
完全な Skill フォルダは以下のようになります:
```
weekly-report/
├── SKILL.md # コア指示(必須)
├── template.md # 週報テンプレート(任意)
└── examples/ # 週報サンプル(任意)
├── good-example.md
└── bad-example.md
```
**ステップ2:SKILL.md を書く**
SKILL.md は Skill 全体のコアです。YAML frontmatter(メタデータ)と Markdown 本文(詳細指示)の2つの部分で構成されます。
**必須メタデータ**:
| フィールド | 要件 | 説明 |
| ------------- | ------- | --------------------------------- |
| `name` | 最大64文字 | スキルの一意の識別名 |
| `description` | 最大200文字 | Claude にいつこのスキルを使うべきかを伝える(非常に重要!) |
**任意メタデータ**:
| フィールド | 説明 |
| --------------- | --------------------------------------------- |
| `dependencies` | 必要なソフトウェアパッケージ、例:`python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | 使用を許可するツールのリスト |
| `model` | 任意のモデルオーバーライド |
完全な週報生成 Skill の例:
```yaml
---
name: weekly-report
description: 今週の作業内容に基づいて標準化された週報を生成。進捗、課題、来週の計画を含む
---
# 週報生成アシスタント
## 使用シナリオ
ユーザーが週報、業務サマリー、進捗報告の生成を必要とする場合に使用。
## 出力フォーマット
以下の構成で週報を生成してください:
### 今週の成果
- 完了した主要な作業項目をリストアップ
- 各項目に簡潔な説明と成果を含める
### 進行中
- 進行中の作業をリストアップ
- 現在の進捗と予想完了時期を記載
### 発生した問題
- 進捗を妨げている問題をリストアップ
- あれば、必要なサポートを説明
### 来週の計画
- 来週の主要タスクをリストアップ
- 優先度順に並べる
## スタイル要件
- 簡潔な表現を使用
- 過度に技術的な用語を避ける
- 成果とインパクトを強調
## 例
**入力**:今週はユーザーログイン機能を完成し、3つのバグを修正し、プロダクトレビューに参加しました。
**出力**:
### 今週の成果
- ユーザーログイン機能の開発:フロント・バックエンドの結合テスト完了、メールと電話番号でのログインをサポート
- バグ修正:3つの高優先度の問題を解決、システムの安定性を向上
### 進行中
- (なし)
### 発生した問題
- (なし)
### 来週の計画
- ユーザー登録機能の開発を開始
- 単体テストケースの作成
```
**ステップ3:テスト**
Claude でテスト:「今週の週報を作成して。今週はユーザーログイン機能の開発を完了し、3つのバグを修正し、2回のプロダクトレビューミーティングに参加しました。」
### Skill Creator の使用
SKILL.md をゼロから書きたくない場合、Claude には skill-creator スキルが内蔵されており、インタラクティブに作成をガイドしてくれます:
```
Help me create a skill for [your workflow]
```
Claude が一連の質問を通じてニーズを整理し、SKILL.md の初稿を生成します。
## 技術原理
### メタツールシステムとしての Skills
Skills は本質的に**メタツールシステム**です——コードを直接実行するのではなく、専門的な指示を会話コンテキストに注入し、Claude の推論方法を変化させます。
Skill がトリガーされると、2つのことが起こります:
1. **メタデータメッセージ**:どの Skill がロードされているかを示す可視的なステータスインジケーター
2. **スキルプロンプト**:完全な SKILL.md 指示が Claude に送信されるが、ユーザーには非表示
### 発見と選択メカニズム
Claude はどの Skill を呼び出すべきかをどう知るのか?答え:**完全に言語理解に依存**。
有効化されたすべての Skills の name と description がダイナミックリストにフォーマットされ、システムプロンプトに書き込まれます。メッセージを送信すると、Claude はネイティブの言語理解能力を使って意図をマッチングし、特定の Skill を呼び出すかどうかを決定します。
だからこそ `description` フィールドが非常に重要です——それが Claude の判断の唯一の根拠です。複雑なアルゴリズムルーティングはなく、決定は完全に Claude の推論プロセス内で完了します。
## ベストプラクティス
大量の実践を経て、コミュニティは Skills 作成の4つのゴールデンルールをまとめました:
**1. フォーカスを保つ**
1つの Skill は1つのことだけを行うべきです。複数のフォーカスされた Skills は、1つの大きくて包括的な Skill よりはるかに使いやすく、メンテナンスもしやすく、組み合わせて使うのも簡単です。
**2. 明確な説明**
description フィールドは Claude がいつ Skill を呼び出すかを決定するため、適用シナリオを明確に書く必要があります。「売上データに基づいて四半期分析レポートを生成」は良い説明。「データを処理」は曖昧すぎます。
**3. 例を提供する**
SKILL.md に入出力の例を含めると、出力の安定性が大幅に向上します。特に特定のフォーマット要件があるタスクに有効です。
**4. シンプルに始める**
まず純粋な Markdown で基本的な指示を書き、効果を検証してからスクリプトの追加を検討し、徐々に複雑さを増していきます。
### よくある問題のトラブルシューティング
| 問題 | 考えられる原因 | 解決策 |
| --------------- | ---------------------- | ------------------------------- |
| Skill がトリガーされない | description が十分に正確でない | より具体的な使用シナリオの説明に書き換える |
| Skill がトリガーされない | Skill が正しくインストールされていない | ファイルパスと命名を確認 |
| 出力が不安定 | 例が不足 | より多くの入出力例を追加 |
| 出力が不安定 | 指示が曖昧すぎる | 制約条件とフォーマット要件を追加 |
| ロードが遅い | ファイルが大きすぎる | 大きなファイルを references サブディレクトリに移動 |
### セキュリティに関する注意事項
Skills はコードを実行できるため、セキュリティが非常に重要です:
* **信頼できるソース**:信頼できるチャネルの Skills のみ使用
* **スクリプトの審査**:インストール前に Skills 内のスクリプトコードを確認
* **機密情報の保護**:Skills に API キーやパスワードをハードコードしない
* **権限管理**:チームで使用する際、Skills の共有範囲に注意
## 現在の制限事項
新興の機能として、Skills には現在いくつかの制限があります:
| 制限事項 | 説明 |
| -------------------------- | -------------------------------- |
| ~~**Anthropic エコシステムのみ**~~ | ✅ **解決済み** - 下記参照 |
| **レビューメカニズムの不足** | 内蔵のレビューや監査ワークフローがまだない |
| **学習曲線** | チームがワークフローを調整し、バージョン管理フローを確立する必要 |
| **新興段階** | エコシステムはまだ発展中 |
> **重大アップデート(2025年12月18日)**:Anthropic は Agent Skills を[オープンスタンダード](https://agentskills.io)として正式に公開しました。仕様とリファレンス SDK は [agentskills.io](https://agentskills.io) で公開されています。
>
> **採用済みの企業/製品**:
> 
>
> * **Microsoft**:VS Code、GitHub が統合
> * **OpenAI**:ChatGPT、Codex CLI が同じアーキテクチャを採用
> * **プログラミングツール**:Cursor、Goose、Amp、OpenCode
> * **パートナー Skills**:Atlassian、Figma、Canva、Stripe、Notion、Zapier
>
> 同時に、Anthropic、OpenAI、Block が共同で [Agentic AI Foundation](https://www.linuxfoundation.org/)(Linux Foundation がホスト)を設立し、Google、Microsoft、AWS も参加しています。これは Skills が単一ベンダーの機能から業界標準へと進化していることを意味し、Claude Code 用に作成した Skills は OpenAI Codex CLI と相互運用可能になります。
>
> 参照元:
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## 学習リソース
### 公式リソース
| リソース | リンク | 説明 |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------- | ----------------- |
| Skills GitHub リポジトリ | [anthropics/skills](https://github.com/anthropics/skills) | 公式サンプル、22k+ Stars |
| Claude Code ドキュメント | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Skills 使用ガイド |
| ヘルプセンター | [support.claude.com](https://support.claude.com) | よくある質問 |
| 技術ブログ | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | 技術原理の深層解析 |
| API クイックスタート | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | 開発者統合ガイド |
| Agent Skills オープンスタンダード | [agentskills.io](https://agentskills.io) | 公式仕様と SDK |
### コミュニティ厳選
| リソース | リンク | 説明 |
| --------------------- | ------------------------------------------------------------------------------------- | ----------------------------- |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Skills 厳選コレクション |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148以上のスラッシュコマンド、54の AI エージェント |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | オフィスドキュメントの作成・編集スキル |
### 推奨読書
| 記事 | 著者 |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic 公式 |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## 展望
Skills の登場は、AIツール発展の重要な方向性を示しています——AIがタスクを実行できるだけでなく、特定の作業方法を学習し記憶できるようにすること。Simon Willison は Skills がAIツール分野の「カンブリア爆発」をもたらすと予測しましたが、この判断は誇張ではありません。
ますます多くの開発者とチームが Skills を構築・共有し始めるにつれ、以下のような展開が見られるかもしれません:
* **専門化された Skills マーケットプレイス**:各業界の専門家が知識を再利用可能な Skills にパッケージ化
* **Skills と MCP の深い融合**:完全なエンドツーエンドのワークフローを形成
* **エンタープライズ級 Skills プラットフォーム**:チームコラボレーション、バージョン管理、権限コントロール
今がまさに参入の好機です。すぐにできることは Claude.ai にログインしてドキュメントスキルを有効にすること。今週はコミュニティ Skill をインストールし、最初のシンプルな Skill を作成してみてください。長期的には、チーム内の繰り返し作業を特定し、徐々に専用のスキルライブラリを構築することが、効率向上の有効な手段となるでしょう。
### さらに読む
* 《[Claude システムアーキテクチャ全解析](/ja/docs/notes/claude-architecture)》— Claude システム全体における Skills の位置づけ
* 《[Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent)》— Subagent メカニズムを深く理解
* 《[私の Claude Code ベストプラクティス](/ja/blog/claude-code-best-practices)》— Claude Code の日常的な使用テクニック
# Skill-Creator の詳細な分析: データを使用してスキル開発を推進します
## はじめに
この記事は 2026 年 3 月の情報に基づいており、Claude Code v2.1 以降に対応しています。
[概念](/ja/docs/notes/claude-skills/concept) と [実践](/ja/docs/notes/claude-skills/practice) を読んだことがあれば、SKILL.md ファイルを手動で作成する方法をすでに知っているはずです。フロントマッターを定義し、コマンドを作成し、それを `.claude/skills/` ディレクトリに保存すれば完了です。
しかし、ここで基本的な質問があります: \*\*自分のスキルが本当に役に立つとどうやってわかりますか? \*\*
段落の文言を変更して、より良く機能したと感じるかもしれませんが、それはあなたの主観的な感覚にすぎません。おそらく、別のプロンプト単語があれば、新しいバージョンはさらに悪化するでしょう。もしかしたら、あなたのスキルは未熟な状態に比べてまったく向上していないかもしれません。クロードは一人でも十分にできます。
概念的な章と実践的な章では、スキル開発のプロセスは次のとおりです: \*\*書いた → 試した → 大丈夫だと感じた → オンライン \*\*。プロセス全体は直感に依存しており、定量化はなく、「このスキルは、スキルをまったく持たない場合よりもどの程度優れているか?」に答える方法がありません。そして、Skill-Creator はこれをエンジニアリングに変えました。**筆記 → スキルあり / なしの並行テスト → ブラインド テスト A/B 比較 → 定量的スコアリング → フィードバックの反復 → データ検証**。
これがスキルクリエイターの存在理由です。 「SKILL.md の生成」を支援するだけでなく、作成 → テスト → 評価 → 最適化のループの完全なセットを提供し、データそのものに語らせることができます。
## スキルクリエイターとは
Skill-Creator 自体もスキルです。33KB の SKILL.md ファイルに加えて、サブエージェント ガイダンス ファイル、Python スクリプト、および HTML ビューアをサポートしています。そのディレクトリ構造は次のようになります。
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
インストールも非常に簡単です。
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## もう一度これに従います: 既存のスキルを評価して最適化する
私が実際に使用しているスキルを使用して、Skill-Creator の完全なプロセスを見てみましょう。私はクロード コード プラグイン マーケット [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub) を管理しています。このマーケットでは、`yux-video-summary` スキルを使用してビデオ字幕を構造化された要約に変換します。中国語と英語の言語検出、DUAL\_FILE/SINGLE\_FILE の 2 つの出力モード、フィラーワードクリーニングなどをサポートしています。スキルの SKILL.md は次のようになります。
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
スキルは書かれていますが、それが本当に役立つかどうかはどうやってわかりますか? \*\* ここで Skill-Creator が登場します。
> Skill-Creator のソース コードには重要な記述原則があります: *「すべての背後にある**理由**を一生懸命説明してください。ALWAYS または NEVER をすべて大文字で書いていることに気付いたら、それは黄色の旗です。フレームを再構成して理由を説明してください。」* 意味: 優れたスキルは、厳格なルールを積み上げるのではなく、**理由を説明する必要があります**。
### ステップ 1: テスト ケースを作成し、評価を実行する
主な質問: \*\*このスキルは本当にスキルがないより優れていますか? \*\*
クロード コードを開き、次のように直接入力します。
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator は、まずスキル定義とスキーマを読み取り、次にテスト ケースと定量的アサーションを自動的に生成します。私の実行では、3 つのテスト ケースと 39 のアサーションが生成されました。
テスト ケースを無造作にコンパイルするわけではないことに注意してください。スキルで定義されている DUAL\_FILE と SINGLE\_FILE の 2 つの出力モードを理解し、さまざまなビデオ タイプ (チュートリアル、ポッドキャスト インタビュー、テクノロジー共有) と言語の組み合わせをカバーするシナリオを具体的に設計します。アサーションの設計も非常に特殊で、言語検出、出力モードの選択からコンテンツの品質、中国語と英語のフィラーワードのクリーニングに至るまで、自分でディメンションをテストするよりもはるかに包括的です。
次に、システムはテスト ケースごとに 2 つの独立したサブエージェント、with\_skill (スキルのロード) と **without\_skill** (ベースライン、スキルはロードされない) を同時に開始します。 **6 つの並列エージェント** (3 つのテスト ケース × 2 バージョン) が一度に開始され、それぞれが互いに干渉することなく **独立したワークツリー** で実行されました。
> Anthropic の PDF スキルは以前、入力不可能なフォームの処理に問題がありました。クロードはフィールドを定義せずにテキストを正確な座標に配置する必要がありました。 Eval を通じて障害点が特定され、その後チームは位置決めロジックを修正しました。それが Eval の価値です。「何かが正しくないと感じられる」を「ここで正確に何が間違っているのか」に変えるのです。
### ステップ 2: 3 人のサブエージェントがスコアリングをリレーします。
すべての操作が完了すると、3 つの専門的なサブエージェントが **自動的に** 順番に表示されます。
**採点者** アサーションを 1 つずつ検証します。 with\_skill バージョンの概要に概要テーブルが含まれているかどうか、DUAL\_FILE モードが正しく選択されているかどうか、フィラー ワードがクリーンアップされているかどうかを確認し、各項目の合格/不合格と証拠を記録して、`grading.json` を生成します。
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**Comparator** はブラインド A/B 比較を行います。2 つの概要を受け取りますが、**どれがスキル バージョンでどれがベースライン バージョンであるかはわかりません**。 「アウトプットA」と「アウトプットB」のみを見て、独自の品質基準に基づいて独自に審査し、勝者を決定します。
**Analyzer** は上記の結果を組み合わせて、スキルに関係なくどのアサーションが合格したか (このアサーションには区別がなく、より良いアサーションに置き換える必要があることを示します)、どの結果の分散が大きいか (テストが不安定である)、時間とトークンの間のトレードオフは何か、などの診断を行います。最後に、改善のための提案が示されます。
### ステップ 3: Eval Viewer で結果を確認する
スコアリングが完了すると、Skill-Creator はブラウザで HTML ビューアを自動的に開きます。
**\[出力] タブ** 各テスト ケースの出力を 1 つずつ表示できます。下部にフィードバック テキスト ボックスがあります。「要約にタイムラインが欠けている」「つなぎ言葉が整理されていない」など、不十分だと思われる点を書き留めてください。すべての使用例を読んだ後、**すべてのレビューを送信** をクリックすると、フィードバックが `feedback.json` に保存されます。
**\[ベンチマーク結果] タブ** 定量的な比較 (with\_skill と without\_skill の合格率、消費時間、トークン消費、および各アサーションの項目ごとの比較) を確認できます。
### ステップ 4: 満足するまで繰り返して改善する
Claude Code に戻り、フィードバックの提供が終了したことを伝えます。 Skill-Creator は `feedback.json` を読み取り、ベンチマーク データに基づいて分析と改善の提案を行います。
私のスキルは合格率 97% と好調でした。 Skill-Creator は、小さな問題を正確に特定しました。インタビュー ビデオには注目に値する引用文の段落が欠けており、それを修復するための提案を行いました。
重要なのは、個々のテスト ケースにパッチを適用しないことです。フィードバックを一般化し、その背後にある要件を理解し、スキルの全体的な構造を調整してから、SKILL.md を書き換え、すべてのテストを `iteration-2/` ディレクトリに再実行して、新しい Eval Viewer を開いて 2 つのラウンドの出力を比較できるようにします。このサイクルは満足するまで続きます。
> Skill-Creator ソース コードの注目すべき改善哲学: *「私たちは、さまざまなプロンプトで何百万回も使用できるスキルを作成しようとしています。厄介な過剰適合の変更や、抑圧的な MUST を加えるのではなく、頑固な問題がある場合は、分岐して別のメタファーを使用してみてください。」* 中心的なアイデア: **過剰適合を回避**して、テスト ケースを作成し、一般化機能を追求します。
### ステップ 5 (オプション): スキルが適切なタイミングで発動するように説明を最適化します。
スキルの品質は検証されていますが、見落とされやすい別の問題があります。それは、スキルの `description` フィールドによって、クロードがいつそれを呼び出すかが決定されるということです。
入力:
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator は約 20 の評価クエリを自動的に生成し (半分はトリガーする必要があり、半分はトリガーしないでください)、レビュー インターフェイスがブラウザーで開きます。
これらのクエリは中国語と英語の両方で利用でき、実際のさまざまな表現をカバーしていることに注意してください。 「トリガーすべきではありません」クエリはあまりにも法外であってはなりません。良い反例は「この会議の議事録を要約するのを手伝ってください」です。これはキーワード「要約」をビデオ要約と共有していますが、実際にはビデオ要約よりも文書処理スキルが必要です。
ページ上でクエリ テキストを直接編集したり、\[**+ クエリの追加**] をクリックして新しいクエリを追加したり、\[削除] ボタンを使用して不適切なクエリを削除したり、クエリごとに \[トリガーする必要がある] スイッチを切り替えることもできます。正しいことを確認したら、\[**評価セットのエクスポート**] をクリックして JSON ファイルをエクスポートします。 Claude Code に戻り、エクスポートしたことを伝えます。システムはバックグラウンドで最適化ループを自動的に実行します。
プロセス全体は完全に自動化されており、クエリをトレーニング セットとテスト セット 60/40 に分割し、トレーニング セットの記述を繰り返し最適化し (最大 5 ラウンド)、テスト セットの結果を使用して、過学習を回避する最適なバージョンを選択します。実行後、最適化前後の説明の比較が出力されます。
最適化された説明はより具体的になります。サポートされるファイル タイプ (.vtt/.srt) が明確になり、パイプライン機能 (フィラー クリーニング、DUAL/SINGLE\_FILE ロジック) が強調され、MUST USE を使用してトリガーされるべきではないシナリオが除外されます。 Anthropic は、内部でこのオプティマイザーのセットを使用して、独自のドキュメント作成スキルを実行しました。これにより公開スキル6種類中5種類の発動精度が向上しました。
### 高度な使用法: 動的なコンテキスト インジェクション
スキルのロード時にコンテキストを自動的に挿入する場合は、スキル 2.0 の `!` 構文を使用して SKILL.md にシェル コマンドを埋め込むことができます。
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
これらのコマンドはクロードがスキルを確認する前に実行され、データはプロンプトに直接埋め込まれます。クロードにファイルを 1 つずつ探索させる場合と比較して、時間とトークンを大幅に節約できます。
## 2 種類のスキル: どちらを作成する必要がありますか?
Skill-Creator を使用する前に、Anthropic によって定義された 2 つのスキル タイプを理解する必要があります。
**能力向上型** - 今までできなかったこと、うまくできなかったことをモデルにやらせます。たとえば:
* 画像生成スキル:クロードはネイティブで画像を生成できませんが、スキルを通じてナノバナーなどのツールを呼び出すことで実現できます。
* フロントエンド設計スキル: デフォルトの AI 設計は非常に「AI 風味」であることが多く、優れた設計スキルにより品質が大幅に向上します。
**コーディング設定** - 特定のワークフローを確立します。モデルにはすでに個別の機能が備わっていますが、正確な実行順序が必要です。たとえば:
* PRレビュースキル:一定の手順に従ってコードのセキュリティをチェックし、リスクレベルレポートを出力
* ビデオ要約スキル: 特定のテンプレート構造に従った出力、自動言語検出、フィラーワードクリーニング
これら 2 つのタイプのスキルをテストする必要がある理由は異なります。 **能力向上タイプ**は、モデルが進化するにつれて不要になる可能性があります。ベースライン (without\_skill) もすべてのアサーションを通過できる場合、モデルが十分にネイティブであることを意味し、このスキルは廃止できます。 **コーディング タイプ** は耐久性が高くなりますが、ワークフローに本当に忠実であるかどうかを検証する必要があります。
Skill-Creator の評価機能を使用すると、古くなった可能性のあるスキルを盲目的に使用するのではなく、スキルがまだ価値があるかどうかを継続的に検証できます。
## コミュニティの意見
Skill-Creator のアップデートは、X/Twitter から Reddit、独立したブログに至るまで、多くの議論を引き起こしました。実際のフィードバックは公式ドキュメントよりも価値があります。
### それは本当に役に立ちますか?データが語る
最も直接的な質問は、スキルを追加した方が、スキルを追加しないよりも本当に優れているのかということです。 \*\* いくつかの実測により明確な答えが得られます。
Reddit u/hashpanak がタイトル生成スキルの評価を実行したところ、with\_skill では 100% の合格率でしたが、With\_skill ではわずか 60% でした。トークンのコストに見合う価値があるかどうか尋ねると、彼は「その通りです。最適化後は、繰り返されるタスクをスクリプトに変換できるため、トークンを節約できます。」と答えました。 u/spences10 はさらに極端です。彼は 250 のサンドボックス評価を実行し、スキルのアクティブ化率を 84% から 100% に増加させました。コメントセクションの u/Manfluencer10kultra は次のように述べています。「**これは標準的な慣行になるはずです。**」
ブロガー [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/) による WordPress セキュリティ スキルのベンチマークテスト: 21 のアサーションすべてに合格し (ベースラインは 90.5% のみ)、速度は 9.9% 速くなりました。彼の要約: **「スキルはかつては芸術でしたが、今ではエンジニアリングです。」**
@0zhuxiaofeng 氏は、実際のワークフローの観点から、より具体的な数値を示しました。「1 か月間使用してみて、最大の変化は、run\_eval によってスキル自体がスコアリングできるようになったということです。コンテンツ操作を実行するエージェントは、各リリース後に効果を自動的に評価するようになり、不十分なスキルは直接排除され、書き直されます。**手動介入が 1 日 3 時間から 30 分に短縮されました**。」
### 見落とされている盲点: トリガー ≠ 品質
ブロガー [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/) は、誰も言及していなかった盲点を指摘しました。 **スキルは品質評価に合格することもありますが、トリガー評価では失敗します** - 出力品質は非常に優れていますが、決して呼び出されることはありません。 `run_loop.py` 最適化を 3 ラウンド行った後、彼は eval を 13/13 にトリガーしました。核となる洞察: 「スキルの説明はメタデータではなく、学習可能なパラメーターです。実際のルーティング動作を最適化する必要があります。」
これは、@DrWang5257 の提案と一致します。「一度に全体を書き直さないでください。まず、トリガー条件、入力テンプレート、失敗フォールバックの 3 つのセクションに分割し、ステップごとに繰り返します。この方法では、更新速度が速く、ロールオーバー率が低くなります。」
### 本当の問題点
効果は良好ですが、落とし穴もたくさんあります。
* **トークンの消費量が膨大です**。 @konghao10 は、「トークンの消費量は膨大だ」と率直に言いました。6 つの並列エージェントを同時に実行するのは実際には安くありません。 Reddit u/munkymead も「本格的な検査を受けるには費用がかかる」とも述べている。
* **スキルが多すぎると戦闘になります**。 [RoboRhythms ブロガー Noah Albert](https://www.roborhythms.com/best-claude-code-skills-2026/) は、**スキルが 8 ~ 10 に達すると問題が発生し始める**ことを発見しました。クロードは出力を自問し、より冗長な序文を生成し、スキル間でコマンドの競合が発生することがあります。しかし、Reddit u/Specialist\_Solid523 は、「下手に書かれたスキルはコンテキストを消費するだけです。**よく書かれたスキルは、ほとんどの場合、トークンの使用をより効率的にします。**」と反論しました。
* **SKILL.md は反復回数が増えると長くなる**。 Reddit u/IulianHI は矛盾を指摘しました。改善を繰り返すことで、スキル ファイルは拡張し続けます \*\* が、実際に何かを行うためのコンテキスト ウィンドウが押し出されてしまいます \*\*。ハッピー パスのみをカバーするテスト ケースは、重要な 5% を見逃します。
* **バージョン管理がありません**。 @fengqve は「スキル \*\* にはバージョンという概念がないのはなぜですか? これは何度も更新されているため、どの更新であるかを説明するのが難しいです。」と不満を述べています。これは、複数ラウンドの繰り返しの後では特に苦痛です。
* **ヘッドレスモードにはバグがあります**。 GitHub には重要な問題があります。スキルが `claude -p` モードではトリガーされず、最適化ループを説明するリコールが常に 0% になります ([#36570](https://github.com/anthropics/claude-code/issues/36570))。
### さらに考える: 再帰的な自己改善
@vista8 は関連する論文 \[Memento-Skills: Let Agents Design Agents] ([https://github.com/Memento-Teams/Memento-Skills](https://github.com/Memento-Teams/Memento-Skills)) を共有し、コメント エリアの誰かがそれを正確に要約しました。「スキルの中心的なボトルネックは反復です。最初のバージョンを作成するのは簡単ですが、実際のシナリオでより良く使用できるようにするのは困難です。この「使用→評価→改善」サイクルを自動化できれば、それはエージェントに自己進化エンジンをインストールするのと同じです。」
Reddit r/ClaudeAI の 104 のようなスレッドでも、この方向性について議論されています。しかし、一番上のコメントはそれに冷水を浴びせた。u/Tatrions は次のように述べた。「再帰ループは機能しますが、難しいのは改善をいつ信頼するかを判断することです。証拠のゲートを行う必要があることがわかりました。少なくとも 2 回失敗しない限り、変更をコミットしないでください。そうしないと、各サイクルが最初から壊れていないものを「修正」することになり、最終的にはさらに悪化することになります。」
## 設置とエコロジー
Skill-Creator は、Anthropic によって公式に管理されているスキルの 1 つとして、[anthropics/skills](https://github.com/anthropics/skills) ウェアハウスに含まれており、17 以上の製品レベルのスキルが含まれています。
より広範なスキル エコシステムも急速に成長しています: [skills.sh](https://skills.sh) 市場は便利な検出とインストールのエクスペリエンスを提供し、コミュニティは 1,234 以上のエージェント スキルを維持しています。
## 最後に書きます
Skill-Creator が解決する中心的な問題は次のとおりです: \*\*自分のスキルが本当に効果的であることはどのようにしてわかりますか? \*\*
それがない場合、スキル開発は「書く→試す→大丈夫と感じる」に依存します。 Skill-Creator を使用すると、次のことが可能になります。
* **Parallel Agent** を使用して、熟練した効果と未熟な効果の両方をテストします
* **ブラインド A/B 比較** により評価バイアスを排除
* **Eval Viewer** を使用して結果を視覚化し、フィードバックを残す
* **Description Optimizer** を使用して、スキルの発動タイミングを正確に制御します
* **反復ループ**を使用して、満足するまで継続的に改善します
これは、ソフトウェア エンジニアリングにおけるテスト駆動開発の概念、つまり「コードを書いて実行できると考えるだけ」ではなく、「実際に期待どおりに動作することをテストによって証明する」という概念と一致しています。
Anthropic は公式ブログで興味深い見通しを提示しました。モデルの機能が向上するにつれて、SKILL.md は「実装計画」 (クロードに **どのように** を伝える) から「仕様の説明」 (クロードに **何を** 伝え、モデルにそれを独自に理解させる) に進化する可能性があります。 Eval フレームワークは、この方向への最初のステップです。Eval は「何をすべきか」を説明します。いつかこの記述自体がスキルになるとしたら、スキルクリエイターが確立したテスト制度はさらに重要なものになるでしょう。
すでにスキルを使用している場合は、`/skill-creator` を使用して、最もよく使用されているスキルを評価してみてください。一部のスキルは、実際にはまったくスキルがないことよりも優れているわけではないことに驚くかもしれません。そこから最適化が始まります。
関連書籍:
* [クロード スキルとは](/ja/docs/notes/claude-skills/concept) — スキルの基本原則を理解する
* [練習ガイド](/ja/docs/notes/claude-skills/practice) — 最初のスキルを作成する
# コンセプト紹介
## はじめに
[Ralph Wiggum 徹底解説](/ja/docs/notes/ralph-wiggum/concept)では、ある核心的な問題を探りました:**Context Rot** — 会話が長くなるにつれて、Claude のコンテキストウィンドウが失敗したコード、古くなった議論、無関係な情報で埋まり、出力品質が着実に低下していく現象です。
Ralph の解決策は「すべてをリスタートする」というものでした:bash の無限ループで毎回新しい Claude インスタンスを起動し、ファイルシステムを通じて状態を受け渡します。シンプルで効果的ですが、明確な限界もあります — それは単なる手法であり、プロジェクトの理解もフェーズ計画も品質検証もありません。スペックは自分で書き、タスクは自分で編成し、「完了したかどうか」も自分で判断する必要があります。
Chase AI が動画で的確にまとめたように:**Ralph Loop は極めて強力な武器ですが、ほとんどの人に必要なのは一つの武器ではなく、武器庫全体です。** Ralph ループは完全に事前準備に依存しています:PRD は十分に良いですか?機能定義は十分にタイトですか?「完了」がどのようなものか分かっていますか?これらの質問への回答が正確でなければ、ループを何回実行しても、garbage in, garbage out になるだけです。
もし、**Claude を単にループさせるのではなく、あなたのプロジェクトを本当に理解し、確実にコードを届けてくれる**システムがあったら?
それが **GSD (Get Shit Done)** の目指すところです。
## GSD とは
GSD の作成者は **TÂCHES**(GitHub: glittercowboy)という独立開発者です。彼のモチベーションは非常にシンプルでした:
> 「私は50人規模のソフトウェア会社ではない。エンタープライズシアターをやりたくない。ただクールなものを作りたいクリエイターなんだ。」
ライブ配信で、TÂCHES は衝撃的な事実を実演しました:彼は**一切コードを手書きしません**。GSD を使って、4時間でゼロから完全な macOS ネイティブ AI 音楽生成アプリ(Sample Digger)を構築しました — 手書きコードはゼロです。彼は自分をプログラマーではなく「ハイレベルなプロジェクトマネージャー」と位置づけています — ビジョンを描き、重要な判断を下し、結果を検証する。GSD がこのような働き方を可能にしています。
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> —— TÂCHES
他のスペック駆動開発ツール — BMAD、SpecKit — にはそれぞれの価値がありますが、複雑なエンタープライズワークフローを導入しがちです:スプリントセレモニー、ストーリーポイント、ステークホルダーシンク。ソロ開発者や小規模チームにとって、これらのプロセス自体が負担になります。Chase AI が述べたように:「It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way.」
GSD の設計哲学は**複雑性をシステムの中に隠す**ことです。ユーザーは数個のシンプルなコマンドだけで、システムが裏側でコンテキスト管理、タスクオーケストレーション、品質検証のすべてを処理します。リリースから1ヶ月以内に、プロジェクトは約3,000の GitHub スターと14,000の npm インストールを獲得し、TÂCHES はほぼ毎日15〜20回のアップデートをプッシュしています。
### ツールエコシステムにおける GSD の位置づけ
| 観点 | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| -------------- | ----------------- | ------------------- | ----------------- | ----------------------------- |
| 核心的な位置づけ | 実行技術(bash loop) | スペック生成ツールキット | エンタープライズ級フレームワーク | **コンテキストエンジニアリング + スペック駆動** |
| 計画能力 | なし(スペックは自前) | 強い(spec→plan→tasks) | 強い(完全なアジャイルプロセス) | **強い(research→discuss→plan)** |
| 実行の自律性 | 最高(AFK モード) | ステップごとに手動トリガー | ステップごとに手動トリガー | **ステップごとに手動トリガー** |
| 人間の関与モデル | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Context Rot 対策 | 新セッションでリスタート | 組み込みソリューションなし | 組み込みソリューションなし | **サブエージェントによる新鮮なコンテキスト** |
| 品質検証 | 外部テストに依存 | ビルドチェック | 組み込み QA プロセス | **自動検証 + UAT** |
| ユーザー側の複雑性 | 最低 | 中程度 | やや高い | **低い** |
| システム側の複雑性 | 最低 | 中程度 | やや高い | **高い** |
この表は重要なトレードオフを示しています:**Ralph は最小限のシステム複雑性と引き換えに最大の実行自律性を得ています** — 起動したら寝てしまえます。一方、**GSD は高いシステム複雑性と引き換えに計画品質と人間による監督を得ています** — 各ステージで介入する機会があります。SpecKit と BMAD はその中間に位置し、計画能力を提供しますが、GSD のコンテキストエンジニアリングと Ralph の自律実行の両方が欠けています。
GSD と Ralph は矛盾するものではありません。GSD は Ralph の核心原則 — 新鮮なコンテキスト、ファイルを信頼の源泉とすること — を継承しつつ、その上に完全なプロジェクト理解と実行フレームワークを構築しています。Ralph が「AI にタスクを与えて繰り返し試行させる」ものだとすれば、GSD は「あなたが何を望んでいるかを理解し、どう実現するかを調査し、ステップを計画し、実行し、検証する」ものです。
Chase AI の要約が非常に的確です:**Ralph ループは完成された設計図を持ってくることを前提としています — GSD はその設計図を構築する手助けをします。** GSD はあなたの半ば形になったアイデアを受け取り、深い質問をし、代わりに調査を行い、完全な PRD を生成し、それをアトミックなタスクに分解し、プロジェクトをエンドツーエンドで届けます。そしてコード実行時には、Ralph ループを強力にしているのと同じ基本原則を使用します:サブエージェントの新鮮なコンテキストと、可能な限り小さく正確なタスクです。
## コアワークフロー
GSD のワークフローは**議論 → 計画 → 実行 → 検証**のループであり、各ステージに明確な入力と出力があります。
### 1. プロジェクトの初期化
```text
/gsd:new-project
```
1つのコマンドでプロセス全体が始まります。システムは以下を行います:
1. **質問** — あなたのアイデアを完全に理解するまで掘り下げて質問します(目標、制約、技術的好み、エッジケース)
2. **調査** — 関連領域を調査するために並列エージェントを派遣します(任意ですが推奨)
3. **要件抽出** — v1、v2、スコープ外の内容を区別します
4. **ロードマップ** — 要件に対応したフェーズ計画を作成します
ロードマップを承認したら、構築を開始します。TÂCHES の経験では:最初に提供する説明が詳細であるほどシステムの追加質問は少なくなり、曖昧であるほど質問が増えます。開始前に大まかなビジョンドキュメントを準備することを推奨しています — 技術スタックや実装の詳細を知る必要はなく、何が欲しいかを記述するだけで十分です。
**出力ファイル**: `PROJECT.md`、`REQUIREMENTS.md`、`ROADMAP.md`、`STATE.md`
> 既存のコードベースがある場合は、まず `/gsd:map-codebase` を実行してください。システムが並列エージェントを派遣し、技術スタック、アーキテクチャ、規約、潜在的な問題を分析します。その後、`/gsd:new-project` が既存のコードベースに基づいて計画を立てられるようになります。
### 2. 議論フェーズ
```text
/gsd:discuss-phase 1
```
ロードマップの各フェーズには1〜2文の説明しかありません — これだけでは望むものを構築するには不十分です。議論フェーズの役割は、調査と計画の前に**あなたの実装に関する好みを捉える**ことです。
システムは現在のフェーズを分析し、「グレーゾーン」— 複数の合理的な実装アプローチが存在する判断ポイント — を特定します:
* ビジュアル機能 → レイアウト、インタラクション、空の状態の処理
* API/CLI → レスポンスフォーマット、エラーハンドリング、冗長性
* コンテンツシステム → 構造、トーン、深さ、フロー
ここでの各判断は、その後の調査と計画の品質に直接影響します。このステップをスキップしても構いません(システムは合理的なデフォルト値を使用します)が、深い議論を行うことで、システムがあなたの期待により近いものを構築できるようになります。
**出力ファイル**: `{phase}-CONTEXT.md`
### 3. 計画フェーズ
```text
/gsd:plan-phase 1
```
システムは以下を行います:
1. **調査** — 議論フェーズでの決定をガイドとして、現在のフェーズの実装方法を調査します
2. **計画** — XML 構造化フォーマットで2〜3個のアトミックタスク計画を作成します
3. **検証** — 計画が要件を満たしているか確認し、合格するまで反復します
重要な設計理念は **Goal-Backward Planning**(目標逆算計画)です。「何を構築すべきか?」から始めるのではなく、「目標を達成するためにどのような条件が成立していなければならないか?」と問い、そこから逆算して計画とタスクを導き出します。TÂCHES はこのアプローチが「出力品質を大幅に向上させた」と述べています。各タスクが他のタスクとの関係を理解しており、単なる ToDo リストの項目ではないからです。
各計画は、1つの新しいコンテキストウィンドウ内で実行できるほど小さく設計されています。これが重要です — **品質の劣化は起こりません**。
**出力ファイル**: `{phase}-RESEARCH.md`、`{phase}-{N}-PLAN.md`
### 4. 実行フェーズ
```text
/gsd:execute-phase 1
```
システムは以下を行います:
1. **ウェーブ実行** — 独立したタスクは並列実行、依存関係のあるタスクは順次実行
2. **新鮮なコンテキスト** — 各計画は全く新しい 200k トークンのコンテキストで実行され、蓄積されたゴミはゼロ
3. **アトミックコミット** — 各タスクが独立した git commit になる
4. **目標検証** — コードベースがフェーズで約束された機能を実現しているか確認
TÂCHES のライブ配信デモでは、3つの完全なフェーズの開発を完了し、**メインコンテキストウィンドウは常に24%に留まっていました**。GSD Executor サブエージェントは、1つの完全なフェーズを完了するのに1,000行未満のコンテキストをロードするだけで済みます — 10個の計画を連続実行しても、コンテキストは50%未満に留まります。これは Claude Code で直接作業する場合とはまったく異なる体験です:「コンテキストウィンドウの壁にいつぶつかるか賭けるロシアンルーレット」はもうありません。
**出力ファイル**: `{phase}-{N}-SUMMARY.md`、`{phase}-VERIFICATION.md`
### 5. 検証フェーズ
```text
/gsd:verify-work 1
```
自動検証はコードが存在するか、テストが通るかを確認できます。しかし、機能が**期待通りに動作するか**どうかは、あなた自身の確認が必要です。
システムは以下を行います:
1. **テスト可能な成果物の抽出** — あなたが今できるはずのことをリストアップ
2. **一つずつ検証をガイド** — 「メールでログインできますか?」はい/いいえ、または問題を記述
3. **失敗の自動診断** — デバッグエージェントを派遣して根本原因を特定
4. **修正計画の作成** — 直接実行可能な修正方案
すべて合格すれば、次のフェーズに進みます。問題があれば、再度 `/gsd:execute-phase` を実行して修正計画を実行します。
これが GSD と Ralph ループの最大の思想的な違いです:**Ralph はハンズオフ — 起動したらそのまま走らせる。GSD は各フェーズの終わりに人間による検証ステップがあります。** Chase AI が指摘しているように、Ralph ループは「征服しに行く」スタイル — 振り返らずに自力で動き続けます。GSD は各重要なチェックポイントで軌道修正できることを保証し、監視なしでエラーが積み重なることを防ぎます。
さらに、GSD は専用のデバッグワークフローも提供しています。検証で問題が見つかった場合、`/gsd:debug` は**隔離されたデバッグサブエージェント**を起動します。このエージェントは独自の仮説-証拠-解決ワークフローを持ち、調査プロセス全体を追跡する独立したデバッグドキュメントを作成し、メインコンテキストを汚染しません。
**出力ファイル**: `{phase}-UAT.md`
### サイクルの繰り返し
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
各フェーズが完全な**議論 → 計画 → 実行 → 検証**サイクルを経ます。コンテキストは新鮮なまま、品質は一貫して保たれます。
すべてのフェーズが完了したら、`/gsd:complete-milestone` でマイルストーンをアーカイブしてバージョンをタグ付けします。その後、`/gsd:new-milestone` で次のバージョンの構築を開始します。
## なぜ効果的なのか:技術的原理
GSD の信頼性は偶然ではありません — 4つの重要な技術的柱が支えています。
### Context Engineering
Claude Code は正しいコンテキストが与えられると非常に強力です。ほとんどの人は正しいコンテキストの与え方を知りません。GSD がこれを代わりに処理します。
| ファイル | 用途 |
| ----------------- | --------------------------------- |
| `PROJECT.md` | プロジェクトビジョン、常にロード |
| `research/` | エコシステムの知識(技術スタック、機能、アーキテクチャ、落とし穴) |
| `REQUIREMENTS.md` | バージョン別の要件、フェーズのトレーサビリティ付き |
| `ROADMAP.md` | 方向性と進捗 |
| `STATE.md` | 決定事項、ブロッカー、現在位置 — セッションをまたぐ記憶 |
| `PLAN.md` | アトミックタスク + XML 構造 + 検証ステップ |
| `SUMMARY.md` | 実行記録、履歴としてコミット |
各ファイルには、Claude の品質劣化閾値に基づいた**サイズ制限**があります。制限内に収めれば、一貫した高品質の出力が得られます。メインコンテキストウィンドウは30〜40%に保たれ、実際の作業はサブエージェントの新しい 200k コンテキストで行われます。
Chase AI は Context Rot について直感的な説明をしています:**コンテキストウィンドウがどれだけ大きくても — Sonnet、Opus、100万トークンのウィンドウでさえ — 前半のトークンは後半よりも効果的です。** これはバグではなく、LLM 固有の特性です。Claude Code に組み込まれた autocompact は部分的にしか緩和できません。GSD のアプローチはより徹底的です:すべてのアトミックタスクが新しいサブエージェントで実行され、各タスクが Claude の最高のパフォーマンスを引き出せることを保証します。
TÂCHES 自身のデータがこれを裏付けています:月額 $200 の Max プランで、毎月約 $30,000 相当の Opus トークンを消費しています。多く聞こえるかもしれませんが、各タスクが新鮮なコンテキストで実行されるため、やり直しは最小限であり、実際の効率は劣化したコンテキストで繰り返しパッチを当てるよりもはるかに高いのです。
### XML Prompt Formatting
すべての計画は Claude 向けに最適化された構造化 XML です:
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
正確な指示、推測は不要、検証が各タスクに組み込まれています。
### Multi-Agent Orchestration
各ステージは同じパターンを使用します:薄いオーケストレーターが専門化されたエージェントを派遣し、結果を収集し、次のステップにルーティングします。
| ステージ | オーケストレーターの役割 | エージェントの役割 |
| ---- | ------------------- | ------------------------------------- |
| 調査 | 調整、発見の提示 | 4つの並列リサーチャーが技術スタック、機能、アーキテクチャ、落とし穴を調査 |
| 計画 | 検証、イテレーション管理 | プランナーが計画を作成、チェッカーが検証、合格するまでループ |
| 実行 | ウェーブにグループ化、進捗追跡 | エグゼキューターが並列で実装、各自が新しい 200k コンテキストを持つ |
| 検証 | 結果の提示、次のステップへルーティング | ベリファイアーがコードベースを確認、デバッガーが失敗を診断 |
オーケストレーターは重い作業を一切しません。エージェントを派遣し、待機し、結果を統合します。その結果:1つのフェーズ全体を実行できます — 深い調査、複数の計画作成と検証、数千行のコードの並列記述、自動検証 — **メインコンテキストウィンドウは30〜40%に留まったまま**です。
### Atomic Git Commits
各タスクは完了後すぐに独立してコミットされます:
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
利点:`git bisect` で具体的な失敗タスクを特定でき、各タスクを独立してロールバックでき、クリーンな履歴が将来のセッションで Claude がコードの進化を理解する助けになります。
## GSD の限界
GSD は強力ですが、**できないこと**を理解することも同様に重要です。
### GSD は人間がガイドするワークフローであり、自律エージェントではない
GSD は持続的に実行することができません。各ステージの境界 — `discuss` から `plan`、`execute`、`verify` まで — で手動でコマンドを入力する必要があります。「アプリを作って」と言って寝に行くことはできません。
これは Ralph の AFK モードとは対照的です。Ralph は「起動して寝に行く」ために設計されています — bash の無限ループがタスクの完了または失敗まで実行し続けます。GSD はあなたが各重要なチェックポイントにいることを要求します:ロードマップの承認、議論の質問への回答、計画のトリガー、実行の開始、検証結果の確認。
4時間のライブ配信中、TÂCHES はコマンドを打ち続けていました:`new-project`、`discuss-phase 1`、`plan-phase 1`、`execute-phase 1`、`verify-work 1`、`discuss-phase 2`……各遷移で彼がエンターキーを押す必要がありました。これは偶然ではありません — 意図的な設計上の選択です。
### 意図的な設計上のトレードオフ
Ralph は計画能力を犠牲にして実行の自律性を得ました。GSD は実行の自律性を犠牲にして計画品質と人間による監督を得ました。**これは設計上のトレードオフであり、欠陥ではありません。**
* **Ralph の強み**:寝ている間に機能を一つ完成させることができます。ただし、スペックが十分でなければ、間違った方向に全速力で突き進みます。
* **GSD の強み**:各フェーズの後に軌道修正できます。ただし、全プロセスを通じて立ち会う必要があり、離れることはできません。
理想はどのようなものでしょうか?GSD の議論、計画、実行、検証を自動ループに繋げられたら — Ralph の bash loop のようでありながら、GSD の構造化された計画と品質検証を備えている — それが両方の世界のベストです。しかし、そのようなツールはまだ存在しません。おそらく、これが次に探求すべき方向性でしょう。
## 動画リソース
以下の動画は、GSD の使い方と効果をより直感的に理解する助けになります。
## おわりに
GSD は AI コーディングツールの進化における一つの方向性を示しています:「AI にコードを書かせる」から「AI に確実にプロジェクトを届けさせる」へ。
Ralph Wiggum は重要な洞察を証明しました — 新鮮なコンテキストは蓄積されたコンテキストよりも価値がある。GSD はこの基盤の上に、プロジェクト理解(new-project)、意思決定の記録(discuss)、構造化された計画(plan)、並列実行(execute)、品質検証(verify)を加え、完全な閉ループを形成しています。
ソロ開発者や小規模チームにとって、GSD の価値は複雑なエンジニアリングプラクティスを数個のシンプルなコマンドにパッケージ化していることにあります。サブエージェントオーケストレーションや XML プロンプトエンジニアリングを理解する必要はありません — 何が欲しいかを記述して、システムにやらせるだけです。
Chase AI はうまく表現しています:GSD は「技術的なバックグラウンドがないが、それでも Claude Code で持続可能かつ再現可能な方法でプロジェクトをエンドツーエンドで構築したい」人のためのものです。そして TÂCHES のライブ配信がそれを証明しました — 「おそらく自分では Hello World の HTML ページしか書けない」と自称する人が、GSD で完全なネイティブデスクトップアプリケーションを構築したのです。
これは魔法ではありません。**正しい複雑性を正しい場所に置くこと** — システムがオーケストレーションの複雑性を担い、人間はクリエイティビティと意思決定に集中する。そしてその限界も同様に尊重すべきです:GSD が人間を常に立ち会わせる選択をしていることは、制約であると同時に、その信頼性の源泉でもあります。
実践に進みたい方は、[GSD 実践ガイド](/ja/docs/notes/gsd/practice)をお読みください — 完全なコマンドリファレンス、設定の詳細、実践ワークフローのウォークスルー、FAQ を網羅しています。
***
**関連記事**:
* [Ralph Wiggum 徹底解説](/ja/docs/notes/ralph-wiggum/concept) — Context Rot の問題と Ralph 手法の完全な解析
* [スペック駆動開発とは](/ja/docs/notes/speckit/concept) — Vibe Coding からスペック駆動開発へのパラダイムシフト
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) — コンテキストをクリーンに保つもう一つのアプローチ
* [Claude システムアーキテクチャ全解説](/ja/docs/notes/claude-architecture) — Hooks、Subagent などのコンポーネントの全体アーキテクチャ
* [私の Claude Code ベストプラクティス](/ja/blog/claude-code-best-practices) — Claude Code の日常的な使い方のコツ
# 実践ガイド
## はじめに
[前回の記事](/ja/docs/notes/gsd/concept)では、GSD のコア原理 ── コンテキストエンジニアリング、サブエージェントオーケストレーション、ゴール逆算型プランニング、アトミックコミット ── を深く理解しました。これらの理念は美しく聞こえますが、「原理を理解する」から「実際にプロジェクトを動かす」までの間には多くの操作上のディテールがあります。
この記事では、実際に手を動かしていきます。GSD の完全なコマンド体系、設定オプション、出力ファイル構造、そしてゼロから完全な機能を納品する方法を学びます。
## インストールと設定
### インストール
```bash
npx get-shit-done-cc@latest
```
インストーラーは以下の選択を求めます:
1. **ランタイム** ── Claude Code、OpenCode、Gemini CLI、または全て
2. **スコープ** ── グローバル(全プロジェクト)またはローカル(現在のプロジェクト)
インストール後、ランタイムで `/gsd:help` と入力してインストール成功を確認してください。
### 推奨:パーミッションスキップモード
GSD は摩擦のない自動化のために設計されています。Claude Code の推奨実行方法は以下の通りです:
```bash
claude --dangerously-skip-permissions
```
このフラグを使いたくない場合は、`.claude/settings.json` で細かい権限を設定できます。
### アップデート
```text
/gsd:update
```
GSD のアップデートは非常に頻繁です(TÂCHES はほぼ毎日15〜20回の更新をプッシュしています)。最新バージョンを維持するために、定期的にこのコマンドを実行することをお勧めします。
## 完全コマンドリファレンス
GSD のすべてのインタラクションは `/gsd:` プレフィックス付きのスラッシュコマンドで行います。以下に機能別の完全なリファレンスを示します。
### コアワークフローコマンド
この5つのコマンドが GSD のメインループを構成し、順番に使用します。
| コマンド | 説明 |
| ------------------------ | --------------------------------------------------------------- |
| `/gsd:new-project` | プロジェクトを初期化します。システムがあなたのアイデアを理解するまで質問し続け、リサーチ、要件抽出、ロードマップ作成を行います |
| `/gsd:discuss-phase [N]` | フェーズ N のグレーゾーンについて議論します。実装の好みを把握し、プランニングの方向性を定めます |
| `/gsd:plan-phase [N]` | フェーズ N のアトミックタスク計画を作成します。リサーチ、プランニング、検証の3つのサブステップを含みます |
| `/gsd:execute-phase ` | フェーズ N を実行します。サブエージェントがタスクを並列で実装し、各タスクが独立してコミットされます |
| `/gsd:verify-work [N]` | フェーズ N の成果物を検証します。一つずつ確認を案内し、問題を自動診断します |
> `[N]` はオプションパラメータです ── 省略するとシステムが現在のフェーズを自動検出します。`` は必須パラメータです。
### マイルストーン管理
| コマンド | 説明 |
| --------------------------- | ------------------------------------------------------- |
| `/gsd:audit-milestone` | 現在のマイルストーンの進捗を監査します ── 全フェーズのステータスを確認し、未完了項目を特定します |
| `/gsd:complete-milestone` | 現在のマイルストーンをアーカイブし、バージョンをタグ付けし、次のサイクルに備えます |
| `/gsd:new-milestone [name]` | 新しいマイルストーンを作成します。オプションで名前を指定でき、完了済みの作業に基づいて次のフェーズを計画します |
### フェーズ管理
| コマンド | 説明 |
| --------------------------------- | ------------------------------------------ |
| `/gsd:add-phase` | ロードマップの末尾に新しいフェーズを追加します |
| `/gsd:insert-phase [N]` | 指定位置に緊急フェーズを挿入し、後続フェーズは自動的に再ナンバリングされます |
| `/gsd:remove-phase [N]` | 指定フェーズを削除し、関連する全出力ファイルをカスケード削除します |
| `/gsd:list-phase-assumptions [N]` | 指定フェーズの全ての前提条件と依存関係を一覧表示し、潜在的なリスクの特定に役立ちます |
### Quick Mode とツール
| コマンド | 説明 |
| ---------------------- | -------------------------------------------------------------------- |
| `/gsd:quick [--full]` | クイックモード ── リサーチ、計画チェック、検証をスキップします。小さなタスクに最適です。`--full` で完全な保護を有効にします |
| `/gsd:debug [desc]` | 隔離されたデバッグサブエージェントを起動します。オプションで問題を説明でき、システムが仮説→証拠収集→解決を行います |
| `/gsd:add-todo [desc]` | ロードマップを変更せずに、アイデアを To-Do リストに記録します |
| `/gsd:check-todos` | 現在の To-Do リストを表示します |
| `/gsd:map-codebase` | 既存のコードベースを分析します ── 技術スタック、アーキテクチャ、コーディング規約、潜在的な問題を把握します |
### セッションと設定管理
| コマンド | 説明 |
| ------------------ | --------------------------------------------- |
| `/gsd:pause-work` | 作業を一時停止します。現在の状態を STATE.md に保存し、次回の再開を容易にします |
| `/gsd:resume-work` | 作業を再開します。STATE.md から前回の状態を読み込み、中断した箇所から続行します |
| `/gsd:progress` | プロジェクト全体の進捗を確認します ── 完了フェーズ数、現在位置、未処理項目 |
| `/gsd:help` | 利用可能な全コマンドと簡単な説明を表示します |
| `/gsd:settings` | GSD の設定を表示・変更します |
| `/gsd:set-profile` | モデルプロファイルを切り替えます(quality / balanced / budget) |
| `/gsd:update` | GSD を最新バージョンにアップデートします |
## 設定詳細
### モデルプロファイル
GSD は3つのモデルプロファイルをサポートしており、`/gsd:set-profile` で切り替えられます:
| プロファイル | プランニング | 実行 | 検証 | 適用シナリオ |
| --------------- | ------ | ------ | ------ | ---------------------- |
| quality | Opus | Opus | Sonnet | 複雑なプロジェクト、重要な機能、初回利用 |
| balanced(デフォルト) | Opus | Sonnet | Sonnet | 日常開発、ほとんどのシナリオに最適なバランス |
| budget | Sonnet | Sonnet | Haiku | シンプルな機能、予算重視、高速イテレーション |
### コア設定
`/gsd:settings` で以下の設定を表示・変更できます:
| 設定 | デフォルト値 | 説明 |
| ------------------------ | ---------- | --------------------------------------------------- |
| `mode` | `balanced` | モデルプロファイルの選択 |
| `depth` | `standard` | リサーチ深度:`quick`(クイック)/ `standard`(標準)/ `deep`(ディープ) |
| `git.branching_strategy` | `feature` | Git ブランチ戦略:`feature`(機能ごと)/ `phase`(フェーズごと)/ `none` |
### ワークフロートグル
以下のエージェントは個別にオン・オフを切り替えられ、速度と品質のトレードオフが可能です:
| トグル | デフォルト | 説明 |
| -------------- | ----- | ------------------------- |
| `research` | オン | プランニング前に自動リサーチを行うかどうか |
| `plan_check` | オン | 計画作成後に自動検証を行うかどうか |
| `verifier` | オン | 実行後に自動検証を行うかどうか |
| `auto_advance` | オフ | フェーズ完了後に自動的に次のフェーズに進むかどうか |
> `research` と `plan_check` を無効にすると大幅な高速化が可能ですが、プランニング品質が低下する可能性があります。プロジェクトに慣れてから無効化を検討することをお勧めします。
## 出力ファイル構造
GSD のすべての状態と出力は `.planning/` ディレクトリに保存されます。この構造を理解しておくと、デバッグや手動介入に役立ちます。
### プロジェクトレベルファイル
| ファイル | 用途 | 作成タイミング |
| ----------------- | --------------------------- | -------------------- |
| `PROJECT.md` | プロジェクトのビジョンとスコープ | `new-project` |
| `REQUIREMENTS.md` | バージョン管理された要件ドキュメント、フェーズ追跡あり | `new-project` |
| `ROADMAP.md` | フェーズ計画と進捗 | `new-project` |
| `STATE.md` | 現在の状態 ── 意思決定、ブロッカー、位置 | `new-project`、継続的に更新 |
### フェーズレベルファイル
各フェーズは以下のファイルを生成します(フェーズ1を例として):
| ファイル | 用途 | 作成タイミング |
| -------------------- | ------------------- | ----------------- |
| `01-CONTEXT.md` | ディスカッションフェーズの意思決定記録 | `discuss-phase 1` |
| `01-RESEARCH.md` | リサーチ結果と技術調査 | `plan-phase 1` |
| `01-01-PLAN.md` | 最初のアトミックタスク計画 | `plan-phase 1` |
| `01-02-PLAN.md` | 2番目のアトミックタスク計画 | `plan-phase 1` |
| `01-01-SUMMARY.md` | 最初の計画の実行記録 | `execute-phase 1` |
| `01-02-SUMMARY.md` | 2番目の計画の実行記録 | `execute-phase 1` |
| `01-VERIFICATION.md` | 自動検証結果 | `execute-phase 1` |
| `01-UAT.md` | ユーザー受入テスト記録 | `verify-work 1` |
### ディレクトリ構造例
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## 実戦ワークフローデモ
以下では「ブログシステムにコメント機能を追加する」を例に、初期化から納品までの完全なフローを紹介します。
### Step 1: プロジェクトの初期化
```text
/gsd:new-project
```
システムが継続的に質問を始めます:
```
> 何を構築したいですか?
「Next.js ブログにコメント機能を追加したいです。匿名とログインの両方のコメント、
Markdown レンダリング、管理パネルをサポート。技術スタックは Prisma + PostgreSQL。」
```
説明が詳細であればあるほど、システムからの追加質問は少なくなります。TÂCHES のアドバイスは:大まかなビジョンドキュメントを準備して、何が欲しいかを説明すること ── 技術的な詳細を知っている必要はありません。
システムが完了すると4つのファイルを生成し、ロードマップの承認を求めます。承認後、構築フェーズに入ります。
> **既存のコードベースがある場合は?** まず `/gsd:map-codebase` を実行してください。システムが既存のアーキテクチャとコーディング規約を分析し、その後 `new-project` が既存のコードに基づいてプランニングできるようになります。
### Step 2: ディスカッションフェーズ
```text
/gsd:discuss-phase 1
```
システムがグレーゾーンを特定し、一つずつ質問します:
```
> コメントのネスト:多階層ネストをサポートしますか、それとも1階層のリプライのみですか?
> 匿名コメント:CAPTCHA を要求しますか、それとも直接投稿ですか?
> 管理パネル:一括操作が必要ですか、それとも1件ずつの審査ですか?
```
ここでの一つ一つの決定が、その後のプランニング品質に直接影響します。不確かな場合はシステムにデフォルト値を使わせることもできますが、深い議論をすることで実行フェーズでの手戻りが大幅に減少します。
### Step 3: プランニングフェーズ
```text
/gsd:plan-phase 1
```
システムは以下を行います:
1. Prisma + PostgreSQL でコメントシステムを実装する方法をリサーチ
2. 2〜3個のアトミックタスク計画を作成(例:データモデル、API ルート、フロントエンドコンポーネント)
3. 計画が全ての要件をカバーしているか自動検証
各計画は、一つの新しいコンテキストウィンドウ内で完了できるほど小さく設計されています。
### Step 4: 実行フェーズ
```text
/gsd:execute-phase 1
```
システムがウェーブ実行を開始します:
* **Wave 1**(依存関係なし):データベーススキーマ、Prisma モデル ── 並列実行
* **Wave 2**(Wave 1 に依存):API ルート、コメント CRUD ── 並列実行
* **Wave 3**(Wave 2 に依存):フロントエンドコメントコンポーネント ── 独立実行
各タスクは新しい 200k tokens のコンテキストで実行され、独立した git commit が作成されます。
### Step 5: 検証フェーズ
```text
/gsd:verify-work 1
```
システムが一つずつ確認を案内します:
```
> ✅ データベーステーブルが作成されました
> ✅ API ルートが正しいステータスコードを返しています
> ❓ ブログ記事の下にコメント入力欄が表示されていますか? [はい/いいえ/問題を説明]
> ❓ コメント投稿後、ページがリアルタイムに更新されますか? [はい/いいえ/問題を説明]
```
失敗した項目がある場合、システムが自動的に診断して修正計画を作成します。再度 `/gsd:execute-phase 1` を実行すれば修正が実行されます。
### よくある操作シナリオ
**緊急フェーズの挿入**:要件が変更され、現在のフェーズの前に新しい作業を挿入する必要がある場合。
```text
/gsd:insert-phase 2
```
後続のフェーズは自動的に再ナンバリングされます(元の Phase 2 が Phase 3 に、以下同様)。
**一時停止と再開**:他の作業を処理するために中断する必要がある場合。
```text
/gsd:pause-work # 現在の状態を保存
# ... 他の作業を処理 ...
/gsd:resume-work # 中断した箇所から再開
```
**不満な結果のロールバック**:
```bash
git reset --hard HEAD~3 # 実行前の状態に戻る
```
```text
/gsd:remove-phase 2 # このフェーズの全出力ファイルをカスケード削除
```
TÂCHES はライブ配信でこの操作を何度も実演しています ── 気に入らなければロールバック、潔く明快です。
## デバッグワークフロー
検証で問題が見つかったとき、または開発中にバグに遭遇したとき、GSD は専用のデバッグフローを提供します。
```text
/gsd:debug コメント投稿後にページがリアルタイム更新されない
```
システムは**隔離されたデバッグサブエージェント**を起動し、以下のワークフローで対処します:
1. **仮説** ── 問題の説明に基づいて複数の根本原因の仮説を生成
2. **証拠収集** ── 仮説を一つずつ検証し、コード、ログ、ネットワークリクエストをチェック
3. **解決** ── 根本原因を特定した後、修正計画を作成
主要な特徴:
* **コンテキスト隔離**:デバッグエージェントは独自のコンテキストウィンドウを持ち、メインの開発コンテキストを汚染しません
* **ドキュメント追跡**:調査プロセス全体を記録する独立したデバッグドキュメントを作成
* **修正計画**:診断完了後、直接実行可能な修正計画を出力
これはメインコンテキストで直接デバッグするよりもはるかに効率的です ── デバッグ情報がメインウィンドウに蓄積されません。
## 実戦から得た知見
TÂCHES のライブ配信と Chase AI の使用体験を総合した、実戦的なアドバイスをご紹介します。
### 急がば回れ
TÂCHES は、GSD を使い始めた当初は「速く、速く、速く」というマインドセットだったと率直に語っていますが、後になって**リサーチとディスカッションフェーズに時間をかけるほうが、実行フェーズでの手戻りが減る**ことに気づきました。新しいバージョンの GSD に `research-project` と `define-requirements` のステップが追加されたのは、まさにコードを書く前に方向性を正しく定めるためです。
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### フェーズ間でコンテキストをクリア
TÂCHES の習慣は、**フェーズごとに `clear` を実行する**ことで、メインコンテキストをスリムに保つことです。彼は Warp ターミナルを使用し、各ウィンドウをフルスクリーン(Command+Shift+Enter)にして、一つのウィンドウで現在のフェーズを実行しながら、別のウィンドウで次のフェーズのリサーチを行っています。
### Token コストのトレードオフ
GSD のサブエージェント方式は、Claude Code を直接使用するよりも確かに多くの token を消費します。しかし Chase AI は説得力のある論点を提示しています:**「plan twice, prompt once(2回計画して、1回プロンプト)」は、「1回プロンプトして、パッチの繰り返し」よりも長期的にはコスト効率が良い**ということです。新鮮なコンテキストで一発で正しく仕上げるほうが、劣化したコンテキストで繰り返し修正するよりもはるかに効率的です。
### 不満な結果への対処
フェーズの結果に満足できない場合は、`git reset --hard` してから `/gsd:remove-phase` でそのフェーズの全出力ファイルをカスケード削除できます。TÂCHES はライブ配信でこの操作を実演しています ── ある視覚効果が気に入らず、前の満足できる状態まで直接ロールバックしました。潔く明快です。
### To-Do システム
`/gsd:add-todo` を使えば、ロードマップを変更することなく、いつでもアイデアを To-Do リストに記録できます。これらのアイデアは `/gsd:discuss-milestone` の際に取り出され、次のマイルストーンのインプットとして活用できます。TÂCHES の戦略は「まず機能を作り、マイルストーン2で UI を磨く」です。
## FAQ とベストプラクティス
### ベストプラクティス
**詳細な初期説明を提供しましょう。** `/gsd:new-project` の品質は、あなたのインプットの品質に依存します。大まかなビジョンドキュメントを準備してください ── 目標、ユーザー、コア機能、既知の制約を記述します。説明が正確であればあるほど、システムからの追加質問は少なくなり、プランニングの精度が上がります。
**フェーズ間でコンテキストをクリアしましょう。** 各フェーズ完了後、`clear` または `/compact` を実行してメインコンテキストウィンドウをスリムに保ちます。TÂCHES の習慣は、メインコンテキストを30〜40%に維持することです。
**まず Quick Mode でテストしましょう。** 不確かな小さな機能には、まず `/gsd:quick` で試してみましょう。うまくいけば、正式なロードマップに組み込みます。
**既存プロジェクトではまず map-codebase を実行しましょう。** 既存のコードベースで GSD を使用する前に、`/gsd:map-codebase` を実行してください。システムが技術スタック、アーキテクチャ、コーディング規約を分析し、その後のプランニングが既存コードにより適合したものになります。
### FAQ
**Q: GSD はどのランタイムに対応していますか?**
A: Claude Code、OpenCode、Gemini CLI です。インストール時に単体または全てを選択できます。
**Q: Quick Mode と通常モードの違いは何ですか?**
A: Quick Mode は GSD の基本的な保護機能(アトミックコミット、状態追跡)を提供しますが、リサーチ、計画チェック、検証のステップをスキップします。完全なプランニングが不要なバグ修正、小さな機能追加、設定変更などに最適です。
**Q: 実行中に一時停止できますか?**
A: はい。`/gsd:pause-work` で現在の状態を STATE.md に保存します。次回 `/gsd:resume-work` を実行すると、システムは中断した箇所から続行します。
**Q: Token コストはどう制御しますか?**
A: 3つの方法があります ── (1) `budget` プロファイルに切り替える:`/gsd:set-profile budget`、(2) `research` や `plan_check` エージェントを無効にする、(3) シンプルなタスクには `/gsd:quick` を使用する。
**Q: GSD と Ralph は一緒に使えますか?**
A: はい。GSD と Ralph は異なる問題を解決します ── GSD はプランニングと構造化された実行を担当し、Ralph は自律的なループ実行を担当します。GSD の `new-project` と `plan-phase` で完全な計画を生成し、人間の介入が不要なフェーズは Ralph ループで実行するという使い方ができます。
**Q: 複数人でのコラボレーションはどうすればいいですか?**
A: `.planning/` ディレクトリを Git にコミットできます。複数人がそれぞれ異なるフェーズを実行し、Git でマージすることが可能です。ただし、同じフェーズを同時に実行することは避けてください。
## まとめ
GSD のコアバリューは、**複雑さをシステム内に隠し、ユーザーにはシンプルさを提供する**ことにあります。必要なコマンドはわずか ── `new-project`、`discuss-phase`、`plan-phase`、`execute-phase`、`verify-work` ── で、システムが裏側で全てのコンテキスト管理、サブエージェントオーケストレーション、品質検証を処理します。
インストールから納品まで、GSD は明確なパスを提供します:何が欲しいかを記述 → 実装の詳細を議論 → アトミック計画を生成 → 並列実行 → 成果物を検証。すべてのステップであなたが介入する機会があり、すべてのステップがドキュメントに記録されます。
これは「ボタンを一つ押せば全て完了」という魔法ではありません。あなたの参加が必要ですが、認知的負荷の大部分を引き受けてくれるシステムです。TÂCHES が言うように:あなたは上級プロジェクトマネージャーであり、GSD はあなたの実行チームです。
***
**関連記事**:
* [GSD 深掘り解析](/ja/docs/notes/gsd/concept) ── コア原理、ワークフロー、技術アーキテクチャ
* [Ralph Wiggum 深掘り解析](/ja/docs/notes/ralph-wiggum/concept) ── Context Rot と Ralph の方法論
* [snarktank/ralph 実戦ガイド](/ja/docs/notes/ralph-wiggum/snarktank) ── Ralph のインストール、PRD の書き方と実戦
* [スペック駆動開発とは](/ja/docs/notes/speckit/concept) ── Vibe Coding からスペック駆動開発へ
* [Speckit 実践ガイド](/ja/docs/notes/speckit/practice) ── Speckit コマンド詳解と完全な事例
# gstack: YC CEO が起業家としての経験をクロード コードに取り入れたとき
## はじめに
以前のノートでは、[Ralph Wiggum](/ja/docs/notes/ralph-wiggum/concept) の無限ループから [GSD](/ja/docs/notes/gsd/concept) の仕様駆動型開発に至るまで、Claude Code エコシステム内のさまざまな「拡張ソリューション」を検討しました。彼らは皆、同じ質問に答えようとしています。\*\*AI プログラミングを「適応」から「信頼性の高い配信」に変えるにはどうすればよいですか? \*\*
Ralph の答えは、「すべてを再起動する」です。コンテキストの腐敗を避けるために、毎回新しいプロセスを使用します。 GSD の答えは「仕様主導」、つまり構造化されたフェーズ計画と検証サイクルを通じて品質を確保することです。しかし、実行システムだけでなく、完全な仮想エンジニアリング チームが必要な場合はどうすればよいでしょうか? CEO が製品の意思決定を行い、エンジニアリング マネージャーがアーキテクチャをレビューし、デザイナーがエクスペリエンスを制御し、QA が実際のブラウザー テストを実行し、リリース エンジニアがリリースを管理します...すべてが AI によって実行され、ユーザーによって指揮されます。
これが gstack の核となる考え方です。
## gstack とは
gstack の作成者である **Garry Tan** は、技術的および起業家としての豊富な背景を持っています。14 歳でコードを書き始め、スタンフォード コンピューター エンジニアリングを卒業し、Palantir の 10 人目の従業員であり、Posterous (後に Twitter に買収) を共同設立し、2023 年から Y Combinator の社長兼 CEO を務めています。
彼は gstack を使用して、YC をフルタイムで実行しながら、60 日間で 600,000 行を超える実稼働コード (35% テスト) をリリースしました。これは 1 日あたり平均 10,000 行以上になります。プロジェクトの 1 つである garylist.org は 21 日で立ち上げられ、コードは 150,000 行、テスト カバレッジは 35% でした。彼自身の言葉によると、コードの品質は、彼が 500 万ドル、2 年、10 人のエンジニアを費やした以前の起業家プロジェクトを上回っています。
このプロジェクトは 2026 年 3 月 11 日にオープンソース化されて以来、3 週間以内に v0 から v0.15.1.0 まで繰り返され、GitHub は 60,500 以上のスターを獲得しました。 MIT ライセンス、完全にオープンソース。
## ツールエコシステムにおける gstack の位置
| 寸法 | ネイティブクロードコード | ラルフ・ウィガム | GSD | スペックキット | 超大国 | **gstack** |
| ----------- | ----------------------- | ------------------ | --------------------- | -------------- | ------------ | ------------------------- |
| コアの位置決め | ユニバーサル AI コーディング アシスタント | 無限ループの反復 | コンテキストエンジニアリング + 仕様主導 | 要件 → 仕様 → タスク | プロセス規律 + TDD | **役割ベースの仮想チーム** |
| コアパターン | 会話型プログラミング | Bash ループ + 新しいプロセス | フェーズベースのロードマップ | 仕様 → 計画 → タスク | 厳密な開発パイプライン | **スプリントの 7 ステップ プロセス** |
| 人間の関与 | ライブ会話 | ハンズオフ (AFK) | 段階ごとの検証 | 仕様の承認 | ステップごとの検証 | **ステージごとの役割のレビュー** |
| ユニークな機能 | 基本的なコーディング | 無制限の反復 | コンテキスト腐敗管理 | 要件のトレース | 強制 TDD | **ブラウザの自動化 + 複数の役割のレビュー** |
| シナリオに適しています | 簡単なタスク | 継続的な反復 | 大規模プロジェクト管理 | 厳しい要件を伴うプロジェクト | エンジニアリング品質保証 | **フルプロセスの製品開発** |
主要なパターンが表からわかります。 \*\*これらのツールは互いに競合しませんが、AI プログラミングの問題を異なる次元で解決します。 \*\*
Superpowers は **プロセス規律** を使用してコードの品質を保証します (必須の TDD、構造化された対話、実装計画)。 GSD は **コンテキスト エンジニアリング** を使用して、複雑なプロジェクト (フェーズ計画、サブエージェントの新しいコンテキスト、ファイル システムのステータス) を管理します。 gstack は、**役割分解** を使用して意思決定の品質を向上させます (CEO の視点で製品をレビューし、エンジニアリング マネージャーがアーキテクチャをレビューし、QA が実際のブラウザを実行します)。
簡単に言うと、Superpowers はプロセス ガードレールに基づいており、gstack はロール設計に基づいています。前者は 1 から N までのプロジェクトの実装に適しており、後者は 0 から 1 までの製品構築に適しています。\*\*この 2 つは競合する製品ではなく、補完的な製品です。 \*\*
## コア ワークフロー: スプリントの 7 つのステップ
gstack は、開発プロセス全体を \*\*思考 → 計画 → 構築 → レビュー → テスト → 出荷 → 反映 \*\* のサイクルに編成します。これは「スプリント」と呼ばれます。アジャイル スプリントではなく、「役割が順番に現れる」という開発リズムです。
### 1. 考える — 製品クリニック
```text
/office-hours
```
これがgstackの最も特徴的なスキルです。インスピレーションは YC のオフィス アワーから直接得られます。起業家は YC のパートナーに会いに行き、自分自身を見つめ直します。 AI は **6 つの強制的な質問**をします:
1. これを特に必要としているのは誰ですか?
2. 今日それがなかったらどうしますか?
3. この問題が今緊急であるのはなぜですか?
4. それが機能することはどうやってわかりますか?
5. 何もしなかったらどうなりますか?
6. リリースできる最小バージョンは何ですか?
この目的は、コードの作成を支援することではなく、コードを作成する前に **問題自体を再検討する**ことです。
### 2. 計画 — 複数の役割によるレビュー
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
CEO レビューは本質的に「創設者モード」です。文字通り要件を実行するのではなく、一歩下がって「この製品の本当の目的は何ですか?」と問いかけます。スコープの拡張、選択的拡張、スコープの維持、およびスコープの縮小の 4 つのモードをサポートします。
### 3. ビルド — コーディングの実装
承認された計画に従ってコーディングを開始します。このステップでは、標準のクロード コード機能を使用します。
### 4. レビュー — 専門家による並行レビュー
```text
/review
```
このスキルは、**7 つの並列サブエージェント**を一度に派遣して、テスト、保守性、セキュリティ、パフォーマンス、データ移行、API コントラクト、レッド チーム攻撃の 7 つの観点からコードをレビューします。明らかな問題は自動的に修正されます。
### 5. テスト — 実際のブラウザの QA
```text
/qa
```
模擬試験ではありません。 QA スキルは、実際のテスターと同じように、**本物のヘッドレス Chromium ブラウザ**を起動し、アプリを開いて、ボタンをクリックし、フォームに記入し、スクリーンショットを撮ります。自動的にバグを修正し、回帰テストを生成し、バグが発見された後に再検証します。
### 6. 出荷 — ワンクリック公開
```text
/ship
```
自動的にマスター ブランチを同期し、テストを実行し、差分を確認し、バージョン番号と CHANGELOG を更新し、コミット、プッシュ、PR を作成します。プロジェクトにテスト フレームワークがない場合は、最初にフレームワークを構築することもあります。
### 7. 振り返る — 見直して学ぶ
```text
/retro
```
エンジニアリング マネージャー スタイルの週次レポート: コミット履歴、テスト率、コード品質の傾向を分析します。複数人チームの分析をサポートし、「連続リリース日数」などの指標を追跡します。
## 機能する理由: 技術原則
### ブラウズ デーモン: AI に注目
gstack の最もユニークな技術的貢献は、ローカルホスト HTTP 経由で通信する永続的なヘッドレス Chromium インスタンスであるブラウズ デーモンです。最初の呼び出しでブラウザが起動され (約 3 秒)、後続の各コマンドには 100 ~ 200 ミリ秒しかかかりません。これは、AI が DOM 構造を推測するのではなく、実際にアプリを認識できることを意味します。
また、CSS セレクターを作成せずにアクセシビリティ ツリーを通じて要素を見つけるための **Ref System** (要素参照 `@e1`、`@e2`) も導入されています。これは、コミュニティ (批評家を含む) によって一般に認められている「真の技術的貢献」です。
### 役割の内訳: エージェントではなくチーム
gstack が行うことは、すべてのロールを独立したプロンプト ファイルに分解することで、クロード コードがさまざまな段階でさまざまなロールの視点に切り替えてコードをレビューできるようにすることです。これは本質的に洗練されたプロンプトエンジニアリングです。
核となる洞察は次のとおりです。 \*\*計画はレビューと同等ではなく、レビューとリリースは同等ではありません。また、創業者の好みとエンジニアリングの厳密さはまったく異なる思考モードです。 \*\* 一般的なエージェントにすべてを任せるのではなく、必要に応じて「頭脳モード」を切り替えます (創業者の思考、エンジニアリングの厳密さ、偏執的なレビュー、迅速な実行など)。
### 3大理念
gstack の ETHOS.md には、次の 3 つの主要な概念が記録されています。
1. **湖を沸騰させる**: AI によって完全性の限界コストがゼロになる場合は、常に完全な実装 (100% のテスト カバレッジ、すべてのエッジ ケース、すべてのエラー パス) を選択してください。 「ショートカットを解放する」というのは古い考え方です。
2. **構築する前に検索**: 3 層の知識 - 実績のあるパターン、新しく人気のあるソリューション、および第一原則。まずは全員が何をしているのかを理解し、思い込みを疑い、なぜ通常の解決策が間違っているのかを発見することから始めましょう。
3. **ユーザー主権**: AI による推奨、人間の意思決定。 2 つの AI モデルが合意に達した場合でも、ユーザーの判断が優先されます。ユーザーにはドメインの知識、戦略的な観点、好みがあるためです。
## gstack の境界と論争
gstack に対するコミュニティの反応は、おそらくあらゆる AI プログラミング ツールの中で最も二極化しています。
**明るい側面**: 創業者と非技術系ビルダーは、特に `/office-hours` や `/plan-ceo-review` のような「製品思考」スキルが、多くの独立系開発者がコードを書き始める前に製品の方向性を再検討するのに役立ってきたという点では、おおむね同意しています。エンジニアリングレビュー (`/review`) では、隠れたセキュリティ脆弱性を実際に発見することができます。このマルチアングル並列レビューモデルには実用的な価値があります。
**質問する側**も非常に直接的です。
* **LOC インジケーターはあまり重要ではありません**: 60 日間で 600,000 行のコード。コードの行数は決して品質の指標ではありません。大量のコードは単なる足場と定型的なものである可能性があります。
* **基本的にプロンプト テンプレート**: 各スキルは SKILL.md ファイルであり、技術的な敷居は高くありません。本当の価値はファイル自体にあるのではなく、プロンプトのデザインの品質にあります。
* **AI セルフレビュー コードの制限事項**: `/review` AI が書いたコードを AI にレビューさせることは、あなた自身の宿題を修正することと同じです。マルチロール並列処理によりこの問題は軽減されますが、依然として同じモデルです。
* **有名人効果ボーナス**: 創設者が YC CEO でない場合、このプロジェクトはそれほど注目されない可能性が高くなります。
**私の意見**: 論争はさておき、gstack の本当に価値のある部分は 2 つです。それは、Browse Daemon のブラウザ自動化テクノロジと、ロール分解の設計パターンです。これは、ギャリー・タンが誰であるかには関係ありません。役割分担の中心的な重要性は技術レベルではなく、行動レベルにあります。これにより、一般的なエージェントにすべてを委ねるのではなく、AI ワークフローをより意識的に整理することができます。
gstack はフォークやカスタマイズに適しています。すべてをコピーするのではなく、必要なスキルを取得し、必要なプロンプトを変更できます。
## ビデオリソース
## 最後に書きます
gstack は、AI プログラミング ツールの興味深い方向性を示しています。つまり、AI をより自律的にすることではなく (ラルフのルート)、プロセスをより厳格にすることではなく (スーパーパワーズのルート)、意思決定の質を向上させるために AI にさまざまな役割を果たさせることです。この論争は、AI プログラミング エコシステムの豊かさを示しています。すべての人に適合するソリューションはありません。
gstack に興味がある場合は、次のステップは [実践編](/ja/docs/notes/gstack/practice) を読むことです。これは、インストールから完全なワークフローの実行までの段階的なチュートリアルです。
***
**関連書籍**:
* [GSD 概念の紹介](/ja/docs/notes/gsd/concept) — 別の構造化 AI プログラミング ソリューション
* [Ralph Wiggum の詳細な分析](/ja/docs/notes/ralph-wiggum/concept) — 無限ループ反復の開始点を理解する
* [Claude Skills Concept](/ja/docs/notes/claude-skills/concept) — スキルの基礎となるメカニズムを理解する
# gstack フロントエンド スキル パノラマ: 設計から起動までの AI ワークフロー
## はじめに
以前のノートでは、[gstack とは](/ja/docs/notes/gstack/concept)、[ワークフローの実行方法](/ja/docs/notes/gstack/practice)、[スキルのエンジニアリング アーキテクチャ](/ja/docs/notes/gstack/skill-architecture) について説明しました。しかし、議論されていない質問が 1 つあります。gstack のインストール後にもたらされる 60 以上のスキルのうち、フロントエンド/UI デザインに関連するものはどれですか?どのような順序で? \*\*
このノートでは 2 つのことを行います。まず、フロント エンドに関連する約 27 のスキルを機能ごとに分類し、次に、興味深い小さなプロジェクトであるカウントダウン記念ページを使用して、実際の効果を確認できるように、最初から最後まで完全なワークフローを各段階のスクリーンショットとともに説明します。
## フロントエンド スキル キット パノラマ
gstack のフロントエンド スキルは、基礎から屋根まで 6 つの機能レイヤーに分割でき、各レイヤーはさまざまな段階で問題を解決します。
### インフラストラクチャの設計
プロジェクト レベルでの 1 回限りの設定で設計言語を決定し、その後のすべてのスキルがこれらのベンチマークを参照します。
| スキル | 何をすべきか | いつ使用するか |
| ---------------------- | ---------------------------------------------------------------------------- | ----------------------------------- |
| `/design-consultation` | 完全なデザイン システムのコンサルティング、出力カラー マッチング、フォント、間隔、テクスチャの方向 | 新しいプロジェクトを開始するとき、または視覚スタイルを再定義したい場合 |
| `/teach-impeccable` | デザイン設定を一度に収集し、AI 構成ファイルに書き込みます。 gstack のインストール後に一度実行すると、AI に美学を記憶させることができます。 | |
| `/brand-guidelines` | 既存のブランドのカラーマッチングとフォント仕様を適用 | 既存のブランドマニュアルがある場合は直接申請 |
> プロジェクトにすでに `DESIGN.md` がある場合、このレベルはスキップできます。
### 設計の検討
方向性がわからない場合は、複数のオプションをすぐに比較してください。
| スキル | 何をすべきか | いつ使用するか |
| ------------------ | --------------------------------------------------------------- | ------------------------------- |
| `/design-shotgun` | 3 ~ 5 個の視覚的なソリューションを生成し、比較パネルを開きます。どのようなスタイルが必要かわからない、可能性を見てみたい | |
| `/frontend-design` | 認識可能な運用レベルのフロントエンド インターフェイス コードを生成する | 方向性が明確になったらすぐに作業 |
| `/canvas-design` | ポスター、ビジュアル アートの生成 (PNG/PDF) | Web コンポーネントではなく静的なビジュアル デザインが必要 |
### 設計の実装
計画を実際に実行可能なコードに変換し、植字、レイアウト、応答性を処理します。
| スキル | 何をすべきか | いつ使用するか |
| ------------------------ | --------------------------------- | ---------------------------------------- |
| `/design-html` | 確認されたデザインドラフトを製品レベルの HTML/CSS に変換 | 直接実装したいモックアップがあります |
| `/mobile-responsiveness` | モバイルファーストのレスポンシブなレイアウトとタッチ操作 | ゼロからのモバイル適応 |
| `/adapt` | デバイスと画面サイズ間のブレークポイントの適応 | デスクトップ バージョンがあり、携帯電話/タブレットに適合させる必要があります。 |
| `/typeset` | フォントの選択、レベル、サイズ、太さ、読みやすさの最適化 | テキストのレイアウトは「ほとんど無意味」に見える |
| `/arrange` | レイアウトの間隔、視覚的なリズム、配置の修正 | 間隔が一貫しておらず、レイアウトが混雑しているか散在しているように感じられます |
### デザインの強化
機能の完成に基づいて、動的な効果、個性、感情的な詳細を注入します。
| スキル | 何をすべきか | いつ使用するか |
| ------------ | --------------------------------------------------- | -------------------------- |
| `/animate` | 目的を持ったマイクロインタラクションとアニメーションを追加する | ページの機能は問題ありませんが、「硬い」と感じます。 |
| `/delight` | サプライズの詳細とパーソナライズされたタッチを追加 | ユーザーにこのページを覚えておいてもらいたい |
| `/bolder` | 視覚的なインパクトを増幅する | デザインが地味すぎて無難すぎる |
| `/colorize` | 単調なインターフェースに戦略的な彩りを添える | ページは灰色すぎて、地味すぎて、温かみがありません。 |
| `/overdrive` | 技術的な爆発レベルのエフェクト - シェーダー、スプリング物理学、スクロール ドライブ アニメーション | 特定の領域はすごい効果を望んでいます |
| `/onboard` | 新しいユーザー ガイダンス プロセス、空の状態の設計 | 初めてのユーザーエクスペリエンス |
これら 4 つの強化スキルは**漸進的な関係**にあります。`animate` は基本的なダイナミック効果、`delight` は感情的、`bolder` は増幅、`overdrive` は爆発です。プロジェクトのニーズに応じて段階的に積み重ねてください。すべてを使用する必要はありません。
### 設計の最適化
収束と洗練 – 余分なものを取り除き、ずれを調整し、粗いエッジを研磨します。
| スキル | 何をすべきか | いつ使用するか |
| ------------ | ----------------------------------- | ------------------------------ |
| `/polish` | 最終的な品質の磨き: 位置合わせ、間隔、一貫性 | リリース前の最後のパス |
| `/quieter` | 視覚刺激の強度を下げる | デザインが派手すぎてうるさい |
| `/distill` | 不必要な複雑さを最小限に抑えて削除する | ページ上の要素が多すぎるため、要素を減らしたい |
| `/normalize` | デザイン システム標準を調整する (トークン、間隔、色) | スタイルが DESIGN.md 仕様から逸脱しています。 |
| `/clarify` | UX のコピーライティング、エラー メッセージ、ラベルの文言を改善する | コピーライティングはわかりにくい、エラー メッセージは不親切 |
### 設計のレビューと検証
オンラインにする前の体系的な検査、問題の発見、採点、修正。
| スキル | 何をすべきか | いつ使用するか |
| --------------------- | -------------------------------- | ------------------------ |
| `/plan-design-review` | 実装前の設計計画のレビュー (0-10 スコア) | AIにデザイナーの視点で計画を見直してもらいたい |
| `/design-review` | 導入後のビジュアル QA、スクリーンショットの自動比較と修復 | コードを書いた後、視覚的な復元度を確認します |
| `/critique` | UX 評価: 視覚階層、認知負荷、感情的共鳴 | 構造化設計レビューレポートが欲しい |
| `/audit` | 技術レビュー: アクセシビリティ、パフォーマンス、テーマ、応答性 | 本番稼働前の体系的なチェック |
| `/benchmark` | パフォーマンスのベースライン テスト、前後の比較 | パフォーマンスに対する変更の影響を定量化したい |
## 実践的なデモンストレーション: カウントダウン記念日ページを使用してプロセス全体を確認します
分類表を見るだけでは抽象的すぎます。私たちは、小さなプロジェクトを使用して上記のスキルをつなぎ合わせます - **カウントダウン/記念日の単一ページ**の作成: 意味のある日付を選択し、デジタル アニメーションと背景効果を備えたカウントダウン表示を作成します。
このプロジェクトは小規模ですが完成度が高く、6 つのスキル レベルのほとんどをカバーするのに十分です。完全なプロセスは次のとおりです。
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> 7 つのステップすべてを毎回実行する必要はありません。熟練したら、一般的に使用されるリンクは `/frontend-design → /animate → /polish → /ship` の 4 ステップだけです。完全な能力を発揮するために、あらゆるステップがここで行われます。
### ステージ 1: インフラストラクチャ - 設計言語を決定する
**Skill**:`/design-consultation` + `/teach-impeccable`
これはプロジェクトの開始時に 1 回だけ実行する必要があります。 `DESIGN.md` を出力します。これにより、AI にデザイン設定を記憶させることができます。プロジェクトにすでに `DESIGN.md` がある場合は、それを直接スキップします。
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* TODO: スクリーンショット — デザイン コンサルテーションによって生成された DESIGN.md フラグメント */}
### ステージ 2: 探索 - 複数のオプションの比較
**Skill**:`/design-shotgun`
方向性がわからない場合は、AI に 3 ~ 5 つの視覚的なソリューションを生成させ、比較パネルを開いて選択します。
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* TODO: スクリーンショット — design-shotgun によって生成された 3 つのソリューション比較パネル */}
方向を 3 つのオプションから選択します。必要なものが正確にわかっている場合は、このステップをスキップして、ステージ 3 に直接進みます。
### ステージ 3: ビルド - 運用レベルのコードを生成する
**Skill**:`/frontend-design` + `/adapt`
コアリンク。最初から応答性を確保しながらコードを作成します。
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* TODO: スクリーンショット - 構築完了後のデスクトップ ページの効果 */}
{/* TODO: スクリーンショット - モバイル バージョンの効果(/adapt 適応後) */}
### ステージ 4: 強化 – 動きと個性を注入する
**スキル**: `/animate` → `/delight` (オンデマンド `/overdrive`)
これら 3 つは漸進的な関係にあります。`animate` は基本的なダイナミック エフェクト、`delight` は感情的なディテール、`overdrive` は爆発的なエフェクトです。必要に応じてレイヤーを追加します。
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
`DESIGN.md` を参照するアニメーション制約に注意してください。デザイン システムで 150 ミリ秒のホバー遷移のみが許可されている場合、`/overdrive` は適用されません。これは良い判断練習になります。
{/* TODO: スクリーンショットまたは GIF - モーション強化前と後 */}
### ステージ 5: チューニング - 収束と磨き
**スキル**: `/typeset` + `/polish` (オンデマンド `/distill`、`/normalize`)
間隔の調整、フォントの階層、視覚的なリズム。追加しすぎた場合は、`/distill` を使用して減算してください。
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* TODO: スクリーンショット — 磨き前と磨き後の詳細な比較 */}
### ステージ 6: レビュー – 体系的なチェック
**Skill**:`/design-review` + `/audit`
ビジュアル QA + 技術レビュー。 `/design-review` は自動的にスクリーンショットを取得して問題を比較および修正し、`/audit` はアクセシビリティとパフォーマンスをチェックします。
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* TODO: スクリーンショット — 監査によって生成されたスコアリング レポート */}
### ステージ 7: リリース
**Skill**:`/ship`
標準の gstack リリース プロセス - テスト、差分レビュー、PR の作成。
```text
> /ship
```
***
**期待される結果**: プロセス内で 8 ~ 10 のフロントエンド スキルを使用した、視覚的に優れたカウントダウン ページ。それよりも重要なのは、「どの段階でどのスキルを使うか」という感覚を確立することです。
## 毎日のチートシート
以上が完全なプロセスです。日々の開発で特定の問題が発生した場合は、次の表を確認してください。
| 私の現在の質問 | 何を使うか |
| ------------------------------------ | ------------------------------------- |
| どのようなスタイルにしたいのかわからない | `/design-shotgun` |
| ページ機能は良いが「ほとんど役に立たない」と感じる | `/polish` |
| 何かが間違っているような気がしますが、説明できません | `/design-review` |
| フォント/レイアウトがぎこちないです | `/typeset` |
| 乱雑な間隔と混雑したレイアウト | `/arrange` |
| スタイルがデザインシステムから逸脱している | `/normalize` |
| パブリックコンポーネントを抽出したい | `/extract` |
| ページが複雑すぎるため、 | を削除したいと思います。 `/distill` |
| デザインが地味すぎて無難すぎる | `/bolder` または `/colorize` |
| デザインが派手すぎてうるさい | `/quieter` |
| エラー メッセージのテキストが親切ではありません。 `/clarify` | |
| 携帯電話の表示に問題がある | `/adapt` |
| アニメーション効果を追加したい | `/animate` (基本) または `/overdrive` (爆発) |
| オンライン化前の体系的な検査 | `/audit` |
| debug | `/investigate` |
## 概要
このメモは次の 2 つのことを行います。
1. **パノラマ** - gstack の 27 のフロントエンド スキルは 6 つの層 (インフラストラクチャ→探索→実装→強化→最適化→レビュー) に分類されています。
2. **実践的なデモンストレーション** - カウントダウン記念日ページを使用して完全なワークフローを実行し、各段階でどのようなスキルが使用されるか、およびその理由を示します。
重要なポイント: これらのスキルの最も強力な使用法は、スキルを個別に呼び出すのではなく、パイプラインで組み合わせることであり、各段階で明確なスキルを選択しながら、方向性を検討し、実装を構築し、磨きを強化し、リリースをレビューします。
ただし、プロセスに縛られないでください。熟練したら、ほとんどの場合、`/frontend-design → /animate → /polish → /ship` 4 つのステップで十分です。
***
**関連書籍**:
* [gstack の概念](/ja/docs/notes/gstack/concept) — gstack とは何ですか? gstack によってどのような問題が解決されますか?
* [gstack 実践編](/ja/docs/notes/gstack/practice) — インストールから実行までの完全なワークフロー
* [gstack スキル アーキテクチャの分解](/ja/docs/notes/gstack/skill-architecture) — スキル開発者は何を学ぶことができますか?
* [Claude Skills Concept](/ja/docs/notes/claude-skills/concept) — スキルの基礎となるメカニズムを理解する
# gstack の実践: インストールから実行までの完全なワークフロー
## はじめに
[コンセプト](/ja/docs/notes/gstack/concept) では、クロード コードを仮想エンジニアリング チームに変える役割ベースのスキル セットである gstack の中核的な位置付けと、GSD、Superpowers、Ralph およびその他のソリューションと比較した AI プログラミング ツール エコシステムにおける差別化された位置付けについて学びました。
この実践的な記事は、**使用方法** に焦点を当てており、インストールと構成から完全なワークフローの実行まで、30 分で gstack を使い始めるのに役立ちます。
## インストールと構成
### 前提条件
* **クロード コード** がインストールされ、使用可能になります
* **Git** がインストールされている
* **Bun v1.0+** がインストールされています (gstack は Bun 上に構築されています)
* Windows ユーザーにも Node.js が必要です
### グローバル インストール (推奨、30 秒で完了)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
インストール スクリプトは次の 3 つのことを行います。
1. gstack のスキル情報を `CLAUDE.md` ファイルに追加します
2. すべてのスキル ファイルをスキル ディレクトリに置きます
3. Playwright と対応する Chromium ブラウザをインストールします (`/browse` および `/qa` の場合)
### プロジェクトレベルのインストール (チーム共有)
リポジトリのクローン作成後にチーム メンバーが自動的に gstack を取得できるようにする場合は、次のようにします。
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
\###マルチエージェントのサポート
gstack はクロード コードに限定されず、現在 **10 個の AI プログラミング エージェント** をサポートしています。 `./setup` は、デフォルトでインストールされているホストを自動的に検出します。
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
各ホストのスキルインストールパスは`~/./skills/gstack-*/`の形になっており、相互に干渉しません。
> 💡 **OpenClaw ユーザー向けの追加オプション**: ACP を介した呼び出しに加えて、OpenClaw は ClawHub を介して 4 つのネイティブ メソドロジー スキル (`gstack-openclaw-office-hours`、`gstack-openclaw-ceo-review`、`gstack-openclaw-investigate`、`gstack-openclaw-retro`) を直接インストールすることもでき、これはクロード コード セッションなしで会話的に使用できます。
### チーム モード (チーム共有 + 自動更新、推奨)
v1.x ではチーム モードが導入されています。各開発者は gstack をグローバルにインストールし、ウェアハウスには「gstack を使用している」と記録されるだけで、更新は自動的に行われます。
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
`required` を `optional` に置き換えることは、必須ではなく「穏やかなリマインダー」です。 Claude Code を起動するたびに、更新チェックが自動的に実行されます (1 時間に 1 回スロットリングし、ネットワークに障害が発生しても安全かつサイレントです)。ウェアハウスにはベンダーから提供されたファイルはなく、バージョンのドリフトもありません。
### アップデート
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
または、クロード コードで `/gstack-upgrade` を直接使用します。
## 完全なコマンド リファレンス
### スプリントプロセス
| コマンド | 役割 | 説明 |
| --------------------- | --------------- | ------------------------------------------------------------------------------------------- |
| `/office-hours` | YC オフィスアワー | 製品の方向性を再構築し、設計ドキュメントを作成するための 6 つの強制的な質問 |
| `/plan-ceo-review` | CEO / 創設者 | 4 つの範囲のモデルで利用可能な 10 つ星の製品を探しています |
| `/plan-eng-review` | エンジニアリングマネージャー | ロックダウン アーキテクチャ、データ フロー、エッジ ケース、テスト マトリックス |
| `/plan-design-review` | シニアデザイナー | デザインディメンション 0 ~ 10 のスコア、10 ポイントを達成する方法を説明 |
| `/plan-devex-review` | 開発者エクスペリエンスリーダー | 開発者のポートレート、TTHW のベンチマーク、デザインの魔法の瞬間を探索します。 3つのモード(DX EXPANSION / POLISH / TRIAGE)、20〜45の強制問題 |
| `/autoplan` | パイプラインのレビュー | CEO→デザイン→エンジニアリング→DXレビューを順に自動実行し、コーディングの意思決定原則に従って自動決定し、「好みの決定」のみをあなたに丸投げ |
### デザイン
| コマンド | 説明 |
| ---------------------- | ------------------------------------------------------ |
| `/design-consultation` | 完全なデザイン システムを最初から構築し、DESIGN.md を生成します。 |
| `/design-shotgun` | 複数の AI デザイン バリアントを生成し、ブラウザーで選択内容を比較する |
| `/design-html` | 実稼働グレードの HTML/CSS を生成し、React/Svelte/Vue フレームワーク検出をサポート |
### レビューとセキュリティ
| コマンド | 役割 | 説明 |
| ---------------- | ------------ | --------------------------------------------------------------------------------------------- |
| `/review` | スタッフエンジニア | CI は通過できるが本番環境では爆発的に増加するバグを見つけ、明らかな問題を自動的に修正し、整合性のギャップをマークします。 |
| `/investigate` | デバッグエキスパート | 系統的な根本原因のデバッグ。鉄則: 根本原因が見つかるまでバグを修正しないでください。修正が 3 回失敗したら停止 |
| `/design-review` | コードが書けるデザイナー | 視覚的監査 + 自動修復、アトミック送信、前後の比較スクリーンショット |
| `/devex-review` | DXテスター | 実際にオンボーディングを実行します: ドキュメントの参照、入力プロセスの実行、TTHW のタイミング、スクリーンショット エラー、`/plan-devex-review` スコアとの比較 |
| `/cso` | セキュリティ担当者 | OWASP トップ 10 + STRIDE 脅威モデリング、17 の誤検知除外ルール、8/10 の信頼しきい値、各結果には特定の使用シナリオが伴います。 |
### テストと QA
| コマンド | 説明 |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/qa` | 実際のブラウザテストを開いてバグを見つけます → アトミックコミット修正 → 回帰テストの生成 → 再検証 |
| `/qa-only` | 上記と同じですが、レポートのみであり、コードの変更はありません。 |
| `/benchmark` | ベースライン パフォーマンス テスト: ページの読み込み、コア Web バイタル、リソース サイズ、比較前後のサポート |
| `/browse` | \~100ms レベルのブラウザ コマンド、本物の Chromium、スクリーンショット、フォーム入力、要素のクリック |
| `/open-gstack-browser` | GStack ブラウザの起動: 可視 AI コントロール Chromium、サイドバー拡張機能、アンチクロール ステルス、自動モデル ルーティング (Sonnet 操作/Opus 分析) が付属し、ワンクリック Cookie インポートをサポート |
| `/setup-browser-cookies` | 実際のブラウザ (Chrome/Arc/Brave/Edge) からヘッドレス セッションに Cookie をインポートして、ログインが必要なページをテストする |
| `/pair-agent` | AI エージェント間ブラウザ ペアリング: 同じ GStack ブラウザを OpenClaw/Hermes/Codex/Cursor などに共有します。各エージェントには独立したタブがあり、リモート エージェントをサポートする ngrok トンネルが付属しています。スコープ トークン + タブ分離 + レート制限 + 動作属性 |
### リリースと運用保守
| コマンド | 説明 |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/ship` | メインブランチを同期→テストを実行→カバレッジを監査→バージョンを更新→プッシュを送信→PR を作成。プロジェクトにテスト フレームワークがない場合の自動ブートストラップ |
| `/land-and-deploy` | PR をマージ → CI を待つ → デプロイ → 本番環境の健全性を確認 |
| `/canary` | デプロイ後のカナリア監視: コンソール エラー、パフォーマンス低下、ページ エラー |
| `/setup-deploy` | `/land-and-deploy` ワンタイム構成: 自動検出プラットフォーム (Fly.io/Render/Vercel/Netlify/Heraku/GitHub Actions/custom) + 本番 URL + デプロイメント コマンド |
| `/setup-gbrain` | ワンクリック (5 分以内) で GBrain データベースを開始します: PGLite ローカル、Supabase の既存の URL、または管理 API を通じて新しい Supabase プロジェクトを自動的に作成します。 MCP 登録 + ウェアハウス レベルの読み取り/書き込み/読み取り専用/拒否権限 |
### 見直して学ぶ
| コマンド | 説明 |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `/retro` | チーム認識週次レポート: 一人当たりの分析、連勝統計、テストの健康傾向、成長の機会。すべてのプロジェクト + AI ツール (Claude Code / Codex / Gemini) にわたる `/retro global` |
| `/document-release` | 公開されたコード (README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS) と一致するようにプロジェクト ドキュメントを自動的に更新します。 `/ship` が自動的に呼び出されるようになりました。 |
| `/learn` | セッション間の学習記憶の管理: プロジェクトごとに表示、検索、プルーニング、エクスポート、蓄積 |
| `/context-save` `/context-restore` | 継続的チェックポイント モード パッケージ: コンテキストを保存するための自動 WIP コミット。クラッシュ/スイッチ後に `/context-restore` を使用してセッションを再構築します。 |
### セキュリティ保護
| コマンド | 説明 |
| ----------------------- | ------------------------------------------- |
| `/careful` | 危険な操作の警告: rm -rf、DROP TABLE、force-push など。 |
| `/freeze` / `/unfreeze` | 編集範囲を特定のディレクトリにロック/ロック解除 |
| `/guard` | `/careful` + `/freeze` の組み合わせ、最高のセキュリティ モード |
| `/checkpoint` | 作業ステータスのスナップショットを保存/復元する |
### ツールの統合
| コマンド | 説明 | |
| -------------------------------------------------- | --------------------------------------------------------------------------------------- | ------------- |
| `/codex` | OpenAI Codex CLI 統合: 独立したコード レビュー (パス/フェイル ゲート)、対立モード、協議モード。クロスモデルのオーバーラップ解析は、`/review` | で実行した後に行われます。 |
| `/health` | コード品質ダッシュボード: tsc + バイオーム + knip + シェルチェック + テスト → 0-10 総合スコア | |
| `/skillify` | 現在のワークフローを再利用可能なスキルに統合する | |
| `/scrape` | Webスクレイピングのワークフロー | |
| `/landing-report` | ランディング ページのパフォーマンスとエクスペリエンス レポート | |
| `/make-pdf` | PDF ドキュメントを生成 | |
| `/benchmark-models` `/model-overlays` `/plan-tune` | モデル間の比較、カバレッジのオーバーレイ、計画の最適化 | |
### Standalone CLI(v0.19+)
gstack には、スラッシュ コマンドに加えて、スタンドアロン CLI のセットも付属しています (クロード コード セッション内では実行されません)。
| コマンド | 説明 |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------- |
| `gstack-model-benchmark` | クロスモデル評価: 同じプロンプトで Claude / GPT (Codex CLI 経由) / Gemini を実行し、遅延、トークン、コスト、および (オプション) LLM 判定品質スコアを比較します。利用できないプロバイダーは自動的にスキップします。 |
| `gstack-taste-update` | デザイン テイストの学習: `/design-shotgun` の承認/不承認をプロジェクト レベルのテイスト ファイルに書き込み、毎週 5% ずつ減衰し、後続のバリアント生成にフィードバックします。 |
## 構成の詳細
### CLAUDE.md コンテンツを追加
インストール後、gstack は利用可能なすべてのスキルのリストと簡単な説明を `CLAUDE.md` に追加します。これにより、Claude Code はどのコマンドが利用可能であるかを知ることができます。
### スキルのディレクトリ構造
メインの入り口はトップレベルの `~/.claude/skills/gstack/SKILL.md` で、各サブコマンドはフラット ディレクトリの形式で存在し、コアは `SKILL.md` ファイルです。
```text
~/.claude/skills/gstack/
├── SKILL.md # 主入口 skill
├── browse/ # 浏览器 daemon
├── qa/ # QA 测试
├── review/ # 代码审查
├── ship/ # 发布流程
├── plan-ceo-review/ # CEO 审查
├── office-hours/ # 产品门诊
├── pair-agent/ # 跨 Agent 浏览器配对
├── open-gstack-browser/ # GStack Browser 启动器
├── setup-gbrain/ # GBrain 数据库一键上手
├── hosts/ # 10 个 host 配置(claude/codex/cursor/...)
├── bin/ # standalone CLI(gstack-model-benchmark 等)
└── ... # 当前 v1.x 共 50 个 skill 目录
```
`SKILL.md` を自由に変更して動作をカスタマイズできます。これが「フォークしてカスタマイズ」の利点です。
### Browse Daemon
Browse Daemon は永続的な Chromium インスタンスです。主要な構成:
* **ポート**: ランダムに選択された 10000 ~ 60000、10 個以上の並列ワークスペースをサポート
* **セキュリティ**: ローカルホストのみをバインドし、セッションごとにベアラー トークン認証を使用します。
* **Cookie**: Chrome/Arc/Brave/Edge からインポートするには `/setup-browser-cookies` を使用します
## 実践的なワークフローのデモンストレーション
以下に、一般的な gstack ワークフローを示します。コマンドと出力は、ドキュメントとビデオ内の実際のケースに基づいています。
> 💡 **注**: 次の出力は、調査に基づいてコンパイルされた一般的な例です。特定のプロジェクトのスクリーンショットは、実際の実践に基づいて将来追加される予定です。
### ステップ 1: 製品クリニック
```text
> /office-hours
[YC Office Hours] 6 forcing questions:
1. Who specifically needs this?
2. What do they do today without it?
3. Why is this urgent right now?
4. How will you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can ship?
→ Design doc generated
```
急いでコードを書かずに、まず YC オフィス アワーの観点から AI にアイデアを拷問してもらいましょう。
### ステップ 2: 複数の役割によるレビュー計画
```text
> /autoplan
[CEO Review] Finding the 10-star product...
[Design Review] Rating dimensions 0-10...
[Eng Review] Locking architecture + edge cases...
→ Fully reviewed plan ready
```
`/autoplan` は、CEO → 設計 → エンジニアリングの 3 ラウンドのレビューを自動的に実行し、完全なポストレビュー計画を作成します。
### ステップ 3: コーディングの実装
通常は承認された計画に従ってコーディングします。標準のクロード コード会話を使用できます。
### ステップ 4: 複数の専門家によるコードレビュー
```text
> /review
Dispatching 7 specialist reviewers...
- Testing coverage ✓
- Maintainability ✓
- Security: Found 1 issue (auto-fixing)
- Performance ✓
- Data migration ✓
- API contract ✓
- Red team: No vulnerabilities found
→ Review complete, 1 auto-fix applied
```
### ステップ 5: ブラウザの QA
```text
> /qa
Opening headless browser...
Testing user flows:
- Login flow ✓
- Dashboard load ✓
- Form submission: Bug found → fixing → re-testing ✓
- Image upload ✓
→ 4 flows tested, 1 bug fixed, regression test generated
```
### ステップ 6: 公開する
```text
> /ship
Syncing with main...
Running tests: 42 passed, 0 failed
Reviewing diff: 3 files changed
Updating VERSION: 1.2.0 → 1.3.0
Creating PR: "Add screenshot feature"
→ PR #47 created, ready for merge
```
## 実践的なヒントとコミュニティでの経験
### ギャリー・タンの提案
gstack の ETHOS.md、3 つの核となる原則:
1. **湖を沸騰させる**: AI により完全性がほぼ無料になります。常に完全な作業を実行し、近道はしません。
2. **構築する前に検索**: まず検索、最初に理解、そして 3 層の知識検証の後に開始
3. **ユーザー主権**: AI の推奨はあなたが決定します。両方の AI モデルが一致する場合でも、ユーザーの判断が優先されます。
gstack の README は、Karpathy からの引用で始まります。これは、Garry Tan 自身が gstack を構築したい理由を説明する出発点でもあります。
### コミュニティでのポジティブな体験
* **`/office-hours`、YC 申請用**: Reddit r/ycombinator 上の複数の S26 申請者が、申請資料のストレス テストに gstack のオフィス アワーを利用することが非常に効果的であると報告しました。
* **セキュリティ監査で実際の脆弱性が発見されました**: CTO からのフィードバック `/review` により、チームが認識していなかった XSS 脆弱性が発見されました。
* **`/browse` 実際のブラウザ テスト**: コミュニティ (批評家を含む) によって「真の技術的貢献」として認められています。
### よくある落とし穴
* **許可のプロンプトが頻繁に表示される**: 一部のユーザーは、「許可のプロンプトが 30 秒ごとに承認されなければならず、眠れなくなる」と報告しました。クロードコード設定で適切な自動承認ルールを構成することをお勧めします。
* **トークンの消費量が多い**: 特徴的なプロンプトにより、コンテキストの消費量が増加します。コストを重視する場合は、最も必要なスキルを選択して使用できます
* **エージェント ループ**: HN 上で、エージェントが 70 分のループに陥ったとユーザーが報告するケースがあります。適切なタイムアウトとチェックポイントを設定することをお勧めします
* **すべての人に適しているわけではありません**: 経験豊富な開発者は、ほとんどのスキルは不必要なラッパーであると感じるかもしれません。 gstack は、成熟したエンジニアリング プロセスを持つチームよりも、**独立した創設者や小規模なチーム**に適しています。
## よくある質問とベストプラクティス
\*\*Q: gstack と Superpowers は同時に使用できますか? \*\*
はい。この 2 つは相互に補完し合います。Superpowers はプロセス規律と TDD 保証を得意とし、gstack は製品思考と複数役割のレビューを得意としています。多くのチームは、日常のコーディング規律に Superpowers を使用し、製品計画と QA に gstack を使用しています。
\*\*Q: トークンは高価ですか? \*\*
ネイティブのクロード コードよりも高い。各スキルのロール プロンプトがコンテキスト ウィンドウを占めます。しかし、あなたの時間がトークン料金よりも価値があるのであれば、これは通常、良い取引です。
\*\*Q: どのようなタイプのプロジェクトに適していますか? \*\*
アイデアから発売までの**全プロセスの製品開発**に最適です。バグを修正したり、小さな機能を作成したりするだけの場合は、ネイティブの Claude コードで十分です。 gstack の価値は「完全なプロセス」で最大化されます。
\*\*Q: スキルをカスタマイズするにはどうすればよいですか? \*\*
各スキルは `SKILL.md` ファイルです。直接編集するだけです。
1. スキル ディレクトリを見つけます: `~/.claude/skills/gstack//`
2. `SKILL.md` を編集します
3. `./setup` を再実行します
コミュニティでは、グローバル インストールを直接変更するのではなく、リポジトリをフォークしてカスタマイズすることを推奨しています。
### ベストプラクティス
1. **最初に`/office-hours`、次にコーディング**: コードを記述する前に製品クリニックを行う習慣をつけましょう
2. **`/browse` 検証を有効に活用する**: コードを見るだけでなく、AI にアプリケーションを実際に「認識」させます。
3. **定期 `/retro`**: コードの品質と作業ペースの可視性を維持する
4. **段階的な導入**: すべてのスキルを一度に使用する必要はありません。 `/office-hours` + `/review` + `/ship` から開始
5. **フォークのカスタマイズ**: 不適切なプロンプトが表示された場合は、それを直接変更します。これがオープンソースの利点です
## 概要
gstack の核となる価値は、特定のスキルがいかに強力であるかということではなく、**構造化された AI コラボレーション モード**を提供することにあります。役割の切り替えを通じて、さまざまな段階でさまざまな種類の AI 支援を得ることができます。まずCEOの視点で製品の方向性を検討し、次にエンジニアリングマネージャーの厳密さでアーキテクチャを検討し、最後にQAの実際のブラウザで結果を検証します。
次に、自分でインストールして、最初の gstack プロジェクトを `/office-hours` から開始してみてください。
***
**拡張読書**:
* [gstack の概念](/ja/docs/notes/gstack/concept) — gstack の中核となる概念とツールの生態学的位置付けを理解する
* [GSD 実践編](/ja/docs/notes/gsd/practice) — 別の構造化 AI プログラミング ソリューションの実践的なガイド
* [クロードスキル実践編](/ja/docs/notes/claude-skills/skill-creator) — スキル生成の仕組みを理解する
# gstack の分解: 開発者が学べるスキル
## はじめに
[概念](/ja/docs/notes/gstack/concept) と [実践](/ja/docs/notes/gstack/practice) では、gstack とは何か、およびユーザーの観点からその使用方法を学びました。このメモは別の視点からのものです。**スキル開発者として**、gstack ウェアハウスをファイルごとに読んだ後、どのようなエンジニアリング設計が学び、学ぶ価値があるかを考えます。
gstack は、23 個のプロンプト ファイルの単なるコレクションではありません。その背後には完全なエンジニアリング システムがあり、テンプレートの生成、自動アップグレード、学習と記憶、段階的なガイダンス、マルチプラットフォームへの適応、階層化されたテストが、スキル プロジェクトを「使える」ものから「使いやすい」ものに変える鍵となります。
***
## 1. SKILL.md は手書きではありません - テンプレート生成システム
gstack の最も直観に反する設計: \*\*各 SKILL.md は自動的に生成され、直接編集することはできません。 \*\*
```text
SKILL.md.tmpl (人写) → gen-skill-docs → SKILL.md (机生)
```
人間が作成した `.tmpl` テンプレートには、ワークフロー ロジックとベスト プラクティスに加えて、`{{PLACEHOLDER}}` プレースホルダーが含まれています。ビルド スクリプトは、ソース コードからコマンド リファレンス、ブラウザ フラグ リスト、プリアンブル スタートアップ コードなどを抽出し、プレースホルダーに埋め込んで最終的な SKILL.md を生成します。
```text
{{PREAMBLE}} ← 从 resolvers/preamble.ts 生成的启动代码
{{BROWSE_SETUP}} ← 浏览器初始化指令
{{COMMAND_REFERENCE}} ← 从 commands.ts 提取的命令文档
{{SNAPSHOT_FLAGS}} ← 从源代码常量提取的快照选项
```
\*\*なぜこれを行うのですか? \*\*
* ドキュメントとコードが同期しなくなることはありません。 - コマンド リファレンスはソース コードから生成され、ソース コードが変更されるとドキュメントは自動的に更新されます。
* 23 のスキルが同じプリアンブル (約 220 行) を共有し、すべてのスキルが同時に更新されます
* CI は、再生成の忘れを防ぐために、生成されたファイルの有効期限が切れているかどうかを `--dry-run` チェックできます
**要点**: 複数のスキルを維持する場合は、スキル間で共有されるコンテンツをテンプレートに抽出し、ビルド ステップで使用して最終ファイルを生成する必要があります。同じコンテンツの複数のコピーを手動で同期すると、遅かれ早かれ問題が発生します。
***
## 2. アップグレード メカニズム - 検出から実行までの完全なリンク
gstack のアップグレード システムは非常に精巧に設計されており、次の 3 つの層に分かれています。
### 最初の層: バージョン検出
`bin/gstack-update-check` は、次のことを行うスタンドアロンの bash スクリプトです。
1. ローカルの `VERSION` ファイルを読み取ります
2. キャッシュ `~/.gstack/last-update-check` を確認します (UP\_TO\_DATE キャッシュは 60 分間、UPGRADE\_AVAILABLE キャッシュは 720 分間)
3. キャッシュの有効期限が切れた場合、HTTP リクエストは GitHub の `raw.githubusercontent.com/.../VERSION`
4. バージョン番号を比較し、`UPGRADE_AVAILABLE <旧> <新>` を出力します。
### 第 2 層: プリアンブルの統合
**各スキルの SKILL.md スタートアップ コードの最初の行はバージョン検出です**:
```bash
_UPD=$(~/.claude/skills/gstack/bin/gstack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
```
これは、ユーザーがスキルを呼び出すと更新が自動的に検出されることを意味します。特にアップグレード コマンドを実行する必要はありません。存在感はありませんが、カバレッジは 100% です。
### 3 番目のレベル: プログレッシブ リマインダー + 自動アップグレード
新しいバージョンを検出した後、すぐにユーザーに迷惑をかけることはありませんが、プログレッシブ バックオフのためにスヌーズ メカニズム (スヌーズ) を使用します。
* 1 回目のリマインダー: 24 時間後にもう一度言及してください
* 2 回目のリマインダー: 48 時間後にもう一度言及してください
* 3回目以降:7日後にもう一度言ってください
* 新しいバージョンのリリースによりスヌーズカウンターがリセットされる
ユーザーは `gstack-config set auto_upgrade true` 自動アップグレードを有効にし、確認をスキップして直接実行することができます。
アップグレードを実行すると、5 つのインストール タイプ (グローバル git、ローカル git、ベンダーなど) が区別されます。 git インストールは `git fetch + reset` を使用し、ベンダーのインストールは最初にバックアップしてから置き換え、障害が発生した場合は `.bak` から復元します。アップグレード後、プロジェクトのローカル ベンダーのコピーも自動的に同期されます。
**学ぶ価値のあるポイント**:
* 「通話ごとに検出」モードは非常に高いカバレッジを持ち、ユーザーには認識されません。
* 段階的なバックオフにより、頻繁な中断を回避します
* インストール タイプを区別し、画一的な方法ではなく、さまざまなアップグレード戦略を実装します。
* バックアップと復元により、アップグレードの失敗によってスキル全体がハングアップすることがなくなります。
***
## 3. 学習システム - 使えば使うほどスキルが賢くなる
gstack は、軽量かつ効果的な **クロスセッション メモリ システム**を実装しています。
### ストレージ
各プロジェクトには独立した学習ログ `~/.gstack/projects/$SLUG/learnings.jsonl` が追加で書き込まれます。
```json
{
"skill": "review",
"type": "pitfall",
"key": "n-plus-one",
"insight": "这个项目的 User model 有 N+1 查询问题,findAll 要加 include",
"confidence": 8,
"source": "observed",
"files": ["src/models/user.ts"],
"ts": "2026-04-01T14:30:00Z"
}
```
### 自動収集
各スキルが完了する前に、「運用上の自己改善」リンクがあり、実行中に予期せぬ失敗、回り道、またはプロジェクトの癖が発見されたかどうかが反映され、それらはすべて learnings.jsonl に自動的に記録されます。ユーザーが手動でトリガーする必要はありません。
### 自動ロード
新しいセッションが開始されるたびに、プリアンブルは最初の 3 つの高信頼学習エントリをロードしてコンテキストを挿入し、新しいセッションが履歴知識を継承できるようにします。
### 自信の低下
`observed` および `inferred` ソースからのエントリーは 30 日ごとに 1 ポイントずつ減少します。ナレッジ ベースを手動でクリーンアップする必要はありません。古い知識は自然に消え、新しい観察が自然に置き換えられます。
### 管理インターフェース
```text
/learn # 显示最近 20 条
/learn search # 搜索
/learn prune # 检测过期条目(引用的文件已删除)
/learn export # 导出为 markdown 可加入 CLAUDE.md
```
**学ぶ価値のあるポイント**:
* 追加の書き込み専用設計はシンプルで信頼性が高く、同時実行性も安全です
* 信頼の減衰は、メンテナンスの手間がかからない知識の老化管理であり、手動でのクリーニングよりもはるかに効率的です。
* パスの代わりに git リモート URL を使用してプロジェクトを識別します (`gstack-slug` 経由)。これは別の場所に複製して再利用できます。
* プロジェクト間のクエリをサポートしますが、デフォルトでは分離されています
***
## 4. プリアンブルインジェクション - スキルの「ミドルウェア層」
これは、gstack の最もスマートなアーキテクチャ設計の 1 つです。各 SKILL.md は約 220 行のプリアンブル コードを共有しており、Web フレームワークのミドルウェアのように機能します。
```text
┌─ 更新检测 ──────────────────────────────────┐
│ 会话追踪 (sessions/$PPID) │
│ 配置读取 (proactive, skill_prefix, telemetry)│
│ 学习历史加载 (前 3 条高置信度) │
│ 上下文恢复 (最近的 checkpoint + timeline) │
│ 路由规则检测 │
│ 首次使用引导流程 │
└──────────────────────────────────────────────┘
↓
Skill 特有逻辑
```
プリアンブルの bash スクリプトはキーと値のペア (`BRANCH: main`、`PROACTIVE: true`) を出力し、テンプレートは自然言語条件を使用して、クロードがそれに応じて動作を調整できるようにします。
```text
If PROACTIVE is false, do not invoke skills automatically.
Instead suggest: "I think /skillname might help here -- want me to run it?"
```
これは基本的に bash 出力をクロードの「環境変数」として扱います。ランタイム検出には bash を、動作ルーティングには自然言語を使用します。
**学ぶ価値のあるポイント**: 複数のスキルがある場合は、スキルごとにコピーを作成するのではなく、共有ロジック (構成の読み込み、状態の回復、バージョンの検出) を統合されたプリアンブルに抽出する必要があります。
***
## 5. プログレッシブ ブート - Sentinel ファイル モード
gstack の初めてのユーザー エクスペリエンスは細心の注意を払って設計されています。各ブート ステップが touch ファイル (センチネル ファイル) 経由で 1 回だけ実行されるようにします。
```text
~/.gstack/.completeness-intro-seen ← "Boil the Lake" 理念介绍
~/.gstack/.telemetry-prompted ← 遥测选择(community/anonymous/off)
~/.gstack/.proactive-prompted ← 主动触发开关
~/.gstack/.routing-prompted ← CLAUDE.md 路由规则写入
~/.gstack/.welcome-seen ← 安装欢迎消息
```
スキルを開始するたびに、これらのファイルが存在するかどうかを確認します。そうでない場合は、対応するブート ファイルとタッチ ファイルを表示します。すでに表示されたステップは再度表示されません。
**学ぶ価値のあるポイント**: 構成で `"onboarding_step": 3` のステータスを維持する場合と比較して、センチネル ファイルはよりシンプルで信頼性が高くなります。構成ファイルの破損による影響を受けず、各ステップは独立して制御されます。
***
## 6. SKILL.md 構造設計 - 3 層アーキテクチャ
各 SKILL.md は標準の 3 層構造に従います。
### 最初の層: YAML Frontmatter
```yaml
---
name: qa
preamble-tier: 3
version: 0.15.1.0
description: |
Systematically QA test a web application...
Use when asked to "qa", "test this site", "find bugs"...
benefits-from: [office-hours]
allowed-tools:
- Bash
- Read
- Write
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "bash ${CLAUDE_SKILL_DIR}/bin/check-careful.sh"
---
```
主要なフィールド:
* `allowed-tools`: ツールレベルの権限ホワイトリスト。各スキルは必要なツールのみを宣言します。
* `benefits-from`: 事前依存スキルを明示的に宣言します
* `hooks`: PreToolUse フック。ツールが呼び出される前にインターセプトできます (慎重なインターセプト `rm -rf` など)。
* `description`: すべての自然言語トリガーワードが含まれています
### レイヤ 2: 共有プリアンブル + 一般規則
プリアンブル起動コード + 音声定義 + コンテキスト回復 + 整合性原則 + 検索優先度 + 完了ステータス プロトコル + アップグレード ルールなど。すべてのスキルは同一であり、テンプレートから生成されます。
### 3 番目の層: スキル固有のロジック
これは、ワークフロー定義、ロール設定、認知モデルの注入、インタラクション ゲーティングなどの各スキルの「魂」です。
**学ぶ価値のあるポイント**: 3 層に分離されているため、各スキルは独自のロジックのみに集中でき、共有部分はフレームワークによって一貫性が確保されています。
***
## 7. 迅速なエンジニアリングのヒント集
SKILL.md をすべて読んだ後は、学ぶ価値のあるプロンプト デザイン テクニックを以下に示します。
### 媚び防止ルール
オフィスアワーのスタートアップ モードは、AI の一般的な「そして泥臭い」動作を明示的に禁止します。
```text
Never say:
- "That's an interesting approach" → take a position instead
- "There are many ways to think about this" → pick one
- "You might want to consider..." → say "This is wrong because..."
- "That could work" → say whether it WILL work
```
### 禁止用語リスト
「音声」セクションには、明確に禁止されている単語やフレーズが含まれています。
* 禁止されている単語: 掘り下げる、極めて重要、堅牢、包括的、微妙な、重要な、景観...
* 禁止されているフレーズ: 「ここにキッカーがある」、「どんでん返し」、「これを分解してみましょう」...
* 無効な形式: 全角ダッシュ (カンマ/ピリオドに置き換えます)
これらは LLM の一般的な「AI 風味の」単語であり、無効にした後の出力は明らかにより自然になります。
### 認知モデルインジェクション
各レビュー スキルには、異なる思考フレームワークが注入されます。
* **CEO レビュー**: 18 の認知モデル (ベゾスの一方通行/双方向ドアの意思決定、マンガーの逆思考、ジョブズの集中力と引き算...)
* **英語レビュー**: 15 のエンジニアリング管理パターン (「デフォルトでは退屈」、爆発範囲の直感、コンウェイの法則...)
* **デザインレビュー**: 12 のデザイン認識パターン (サービスとしての階層、制約の崇拝、「気づきますか?」テスト...)
これらのモードは AI に機械的に動作させるのではなく、**思考フレームワーク**を提供します。これは、賢い新参者に先任者の経験のリストを与えるのと同じです。
### 仕様規格
```text
Not "you should test this"
but `bun test test/billing.test.ts`
Not "this might be slow"
but "this queries N+1, ~200ms per page load with 50 items"
Not "there's an issue in the auth flow"
but "auth.ts:47, the token check returns undefined"
```
### 信頼度の調整
レビュー スキルでは、各発見に信頼スコアが伴う必要があり、信頼性の低い発見は自動的に格下げされるか非表示になります。
| スコア | 意味 | 処理 |
| ---- | ------------------ | ----------- |
| 9-10 | 特定のコードを読んで確認してください | 通常表示 |
| 7-8 | 高信頼パターンマッチング | 通常表示 |
| 5-6 | 中程度の誤報の可能性 | 説明付きのディスプレイ |
| 3-4 | 信頼度が低い | レポートから隠す |
| 1-2 | 純粋な推測 | P0 レベルでのみ表示 |
### インタラクティブ ゲーティング
Ship スキルは、いつ停止してユーザーを待つか、いつ自動的に続行するかを正確に定義します。
```text
Only stop for:
- Tests failing with no obvious fix
- Merge conflicts requiring human judgment
- Unclear which changes to include
Never stop for:
- Normal git operations
- CHANGELOG/VERSION updates
- PR creation
```
**学ぶ価値のあるポイント**: 優れたスキルとは、「AI がすべてを行う」のではなく、人間とマシンの境界を正確に定義することです。
***
## 8. 状態管理 - ファイルシステムはデータベースです
gstack の永続化はすべて、`~/.gstack/` に保存されるファイル システムを通じて行われます。
| パス | 目的 | フォーマット |
| ------------------------------------- | ----------- | -------- |
| `config.yaml` | グローバル構成 | ヤムル |
| `sessions/$PPID` | アクティブなセッション | タッチファイル |
| `projects/$SLUG/learnings.jsonl` | 学習記録 | JSONL |
| `projects/$SLUG/timeline.jsonl` | スキルタイムライン | JSONL |
| `projects/$SLUG/checkpoints/*.md` | チェックポイント | マークダウン |
| `projects/$SLUG/health-history.jsonl` | 健康診断履歴 | JSONL |
| `analytics/skill-usage.jsonl` | テレメトリの使用 | JSONL |
| `last-update-check` | バージョンキャッシュ | プレーンテキスト |
ほとんどすべての時系列データは、**JSONL** (行ごとに 1 つの JSON オブジェクト) を使用して追加的に書き込まれます。この選択は賢明です:
* 書き込みの自然な同時実行の安全性を追加しました
* データベースの依存関係は必要ありません
* `grep` / `jq` を使用して直接クエリできます
* 最後の行が欠落しているまで破損しています
***
## 9. クロススキル統合モード
### ファイル転送製品
ファイル システムを介してスキル間で作業成果物を転送します。
```text
/office-hours → design doc → /plan-ceo-review 读取
/plan-ceo-review → ceo-plans/*.md → /autoplan 读取
/review → reviews.jsonl → /ship 读取并展示 Dashboard
/qa → qa-reports/ → /retro 读取
```
### Review Readiness Dashboard
船のスキルは `reviews.jsonl` と読み取られ、公開前のクロススキル レビュー ステータスを示します。
```text
| Review | Runs | Last Run | Status | Required |
| Eng Review | 1 | 2026-03-16 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
```
### 依存関係前の提案
plan-ceo-review が設計ドキュメントがないことを検出すると、最初に `/office-hours` を実行することを積極的に推奨します。
```text
"No design doc found. /office-hours produces a structured problem statement...
Takes about 10 minutes."
Options: A) Run /office-hours now B) Skip
```
### シーケンス予測を使用する
Context Recovery は、最近のスキル使用シーケンスを分析し、次のステップを予測します。
```text
If pattern repeats (e.g., review → ship → review),
suggest: "Based on your recent pattern, you probably want /ship."
```
***
## 10. その他の注目デザイン
### フックシステム
注意、フリーズ、ガードの 3 つのスキルは `PreToolUse` フックを使用します。これは、ツールが呼び出される前にインターセプトできる唯一のメカニズムです。
* **注意**: Bash をインターセプトし、`rm -rf`、`DROP TABLE`、`git push --force` を確認してください。
* **フリーズ**: 編集/書き込みをインターセプトし、パスが許可された範囲内にあるかどうかを確認します。
* **ガード**: 上記 2 つを組み合わせます
### マルチプラットフォームへの適応
同じテンプレートのセットは、`--host` パラメーターを通じてさまざまなプラットフォーム用のスキル ファイルを生成します。
```bash
bun run gen:skill-docs --host claude # Claude Code 格式
bun run gen:skill-docs --host codex # OpenAI Codex 格式
bun run gen:skill-docs --host kiro # AWS Kiro 格式
bun run gen:skill-docs --host factory # Factory Droid 格式
```
パスとフロントマターは自動的に適応され、スキルのロジックは変更されません。
### 完了ステータスプロトコル
標準化された完了ステータスは、各スキルの終了時に出力される必要があります。
```text
DONE — 全部完成,提供证据
DONE_WITH_CONCERNS — 完成但有顾虑
BLOCKED — 无法继续
NEEDS_CONTEXT — 需要更多信息
```
### 失敗した 3 つのアップグレード ルール
```text
If you have attempted a task 3 times without success, STOP and escalate.
```
AI が無限再試行ループに陥るのを防ぎます。
### 差分ベースのテストの選択
E2E テストの費用はそれぞれ約 4 ドル (クロード エージェントの起動が必要) であるため、gstack は各テストが依存するソース ファイルを `touchfiles.ts` を通じて宣言し、影響を受けるテストのみ `git diff` に従って実行します。
```typescript
// test/helpers/touchfiles.ts
{
"qa-workflow": ["qa/SKILL.md.tmpl", "browse/src/server.ts"],
"ship-flow": ["ship/SKILL.md.tmpl", "scripts/resolvers/preamble.ts"]
}
```
***
## 概要: 参考にできる設計原則
gstack のエンジニアリング実践から、スキル開発者にとって最も価値のある次の設計原則を抽出しました。
1. **テンプレート生成 > 手動同期**: スキル間で共有されるコンテンツは、テンプレートとビルド手順を使用して自動的に生成されます。コピーして貼り付ける必要はありません。
2. **パッシブ検出 > アクティブ検出**: アップグレード検出はすべてのスキル呼び出しに組み込まれており、ユーザーは気づきませんが、カバー率は 100% です。
3. **ログの追加 > 複雑なデータベース**: JSONL + ファイル システムは、ほとんどの永続化ニーズをカバーでき、シンプルで信頼性が高くなります。
4. **プログレッシブ ブート > 1 つの構成**: センチネル ファイルを使用してブート ステップを制御します。各ファイルは 1 回だけ表示されます。
5. **正確なゲート > 全自動**: 「停止してユーザーを待つ」と「自動的に続行する」の間の境界を明確に定義します。
6. **信頼性の定量化 > ファジー判定**: 各 AI 判定には信頼性スコアが付属しており、信頼性が低い場合は自動的に格下げされます。
7. **時間の減衰 > 手作業によるクリーニング**: 学習記録の信頼性は時間の経過とともに減衰し、古い知識は自然に薄れていきます。
8. **禁止語リスト > スタイルガイド**: 禁止語を直接リストした方が、「自然な口調でお願いします」よりもはるかに効果的です。
***
**関連書籍**:
* [gstack の概念](/ja/docs/notes/gstack/concept) — gstack とは何ですか? gstack によってどのような問題が解決されますか?
* [gstack 実践編](/ja/docs/notes/gstack/practice) — インストールから実行までの完全なワークフロー
* [gstack フロントエンド スキル](/ja/docs/notes/gstack/frontend-skills) — フロントエンド/UI 設計スキルのパノラマと推奨ワークフロー
* [Claude Skills Concept](/ja/docs/notes/claude-skills/concept) — スキルの基礎となるメカニズムを理解する
# Diagnose と Triage:まずフィードバックループを確立し、次に誰に任せるかを決定する
## 2つのスキルをまとめて説明する理由
`/diagnose` と `/triage` は README では独立したスキルですが、同じエンジニアリング問題の2つの側面を解決します。
* `/diagnose` は、**このバグが何であるか、どのように再現するか、どのように修正されたことを証明するか**に関心があります。
* `/triage` は、**この issue が現在、情報待ち、エージェントに任せるべきか、人間に任せるべきか、それとも何もしないべきか**に関心があります。
一方は事実を担当し、もう一方はプロセスを担当します。実際のプロジェクトでは、これら2つはしばしば連携しています。まずバグ issue をトリアージし、情報が不足している場合は `needs-info` にします。情報が十分であれば、diagnose を使用してフィードバックループを構築します。再現が明確になったら、`ready-for-agent` にするか `ready-for-human` にするかを決定します。
## /diagnose の核心:フィードバックループがすべて
`/diagnose` で最も覚えておくべきことは、**まずエージェントが実行できる pass/fail シグナルを確立する**ということです。
Matt は診断を6つの段階に分けています。
| 段階 | 目標 |
| --------------------- | ----------------------------- |
| Build a feedback loop | 迅速で確実、繰り返し実行可能な失敗シグナルを構築する |
| Reproduce | このシグナルでユーザーが記述したのと同じバグを再現する |
| Hypothesise | 3〜5つの反証可能な仮説を立てる |
| Instrument | 最小限のプローブで仮説を検証する |
| Fix + regression test | 適切なテスト面で回帰テストを書き、修正する |
| Cleanup + post-mortem | 一時的なプローブをクリーンアップし、真の根本原因を記録する |
これは、多くの人がバグをデバッグする順序とは逆です。一般的なデバッグプロセスは、コードを見て、原因を推測し、修正し、ページを更新するというものです。Matt は逆で、まずバグを再現可能な機械シグナルに変えてから、仮説を立てます。
## 良いフィードバックループとは
`/diagnose` は、優先順位の高いものから低いものまで、一連のフィードバックループを提供しています。
| ループ | 適切なシナリオ |
| ----------------------------- | ---------------------------------------- |
| 失敗テスト | 適切なテスト面があり、バグを直接表現できる |
| curl / HTTP script | API バグ、サーバーサイドの動作はリクエストで再現可能 |
| CLI + fixture | コマンドラインツール、パーサー、トランスフォーマー |
| Headless browser | UI バグ、コンソールエラー、ネットワーク動作 |
| Replay captured trace | オンラインの実際のペイロード、イベントストリーム、ログトレース |
| Throwaway harness | システムの一部のみを起動し、複雑な依存関係を隔離する |
| Property / fuzz loop | 偶発的なエラー出力、トリガー率を上げる必要がある |
| Bisection / differential loop | あるバージョン以降で壊れた場合、二分探索または旧バージョンとの比較が必要 |
| HITL script | 手動でクリックする必要がある場合でも、スクリプトに従って安定した出力を提供させる |
ここで非常に重要な判断があります。**ループがなければ、仮説段階に進まないでください**。シグナルがなければ、すべての分析は「〜のように見える」になってしまいます。
## 非決定性バグにはどう対処するか
`/diagnose` は、偶発的なバグに対するアプローチも非常に実用的です。目標は、最初から100%再現することではなく、まず再現率をデバッグ可能なレベルにまで高めることです。
例えば:
* 100回ループでトリガーする
* 並行してトリガーする
* スリープを注入して競合ウィンドウを拡大する
* ランダムシードまたは時間を固定する
* 環境変数と外部依存関係を縮小する
1% の偶発的なバグはデバッグが困難ですが、50% の偶発的なバグはすでにデバッグ可能な対象です。この考え方は、フロントエンドの非同期処理、メッセージキュー、支払いコールバック、ストリーミング出力などに非常に役立ちます。
## 仮説は反証可能でなければならない
Matt は、検証に着手する前に3〜5つの仮説を立て、各仮説について予測を記述することを求めています。
```text
X が原因であれば、Y を変更するとバグは消えるはずです。
または、Z を観察すると、特定の特性が現れるはずです。
```
これにより、エージェントが最初に合理的に見える説明に固執するのを防ぎます。さらに重要なのは、特定の実験が情報量を持っているかどうかを判断できるようになることです。
悪い仮説:
```text
キャッシュの問題かもしれません。
```
反証可能な仮説:
```text
ブラウザのキャッシュが古いスクリプトの実行を引き起こしている場合、キャッシュを無効にして強制的に更新すると、コンソール内の古いバンドルハッシュが消え、ボタンのクリックイベントも回復するはずです。
```
後者こそ検証する価値があります。
## 修正段階で最も犯しやすい間違い
`/diagnose` は、適切なテスト面がある場合、まず最小限の再現を失敗テストに変換し、それからコードを修正することを要求します。
重要なのは「適切なテスト面」です。単にユニットテストを追加するだけでは回帰テストにはなりません。適切なテスト面は、実際のバグパターンをカバーする必要があります。
* バグが複数の呼び出し元の組み合わせによってトリガーされる場合、単一の関数だけをテストしてはいけません。
* バグが実際のペイロード構造によってトリガーされる場合、手書きの玩具オブジェクトだけをテストしてはいけません。
* バグがブラウザのイベント順序によってトリガーされる場合、純粋な関数だけをテストしてはいけません。
適切なテスト面が見つからない場合、それ自体が結論です。コード構造がバグを特定できる場所を残していません。修正後、この情報を [`/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) に渡すべきです。
## /triage の核心:issue はステートマシン
`/triage` は、AI に「issue を見て」と漠然と頼むものではありません。issue を小さなステートマシンとして扱います。
各 issue は同時に以下を持つべきです。
* カテゴリ:`bug` または `enhancement`
* 状態:`needs-triage`、`needs-info`、`ready-for-agent`、`ready-for-human`、`wontfix`
この状態セットの価値は、メンテナーが迅速に以下の質問に答えられるようにすることです。
* まだ誰も見ていないものはどれか?
* 報告者からの情報待ちのものはどれか?
* AFK エージェントに任せられるほど明確なものはどれか?
* 人間が自分でやらなければならないものはどれか?
* 閉じるべきで、その理由を記録すべきものはどれか?
## ready-for-agent の基準
`ready-for-agent` は、このプロセスで最も重要な状態です。「このタスクは AI に試させることができる」という意味ではなく、
> タスクは、不在のエージェントが独立して受け取り、実装し、検証できるほど明確である。
これは通常、issue に少なくとも以下のものが含まれていることを意味します。
* 背景と問題の記述
* 関連するコードパスまたはモジュール
* 明確な受け入れ基準
* 既知の制約
* バグの場合、再現方法があることが望ましい
* 追加の製品/設計判断は不要
これらが不足している場合、`needs-info` または `ready-for-human` であるべきであり、無理にエージェントに投げつけるべきではありません。
## needs-info は具体的な質問をする
`/triage` は `needs-info` のテンプレートを非常にシンプルにしていますが、重要なのは質問が具体的でなければならないことです。
```markdown
## Triage Notes
**What we've established so far:**
- ...
**What we still need from you (@reporter):**
- ...
```
悪い質問:
```text
詳細情報を提供してください。
```
良い質問:
```text
問題が発生したブラウザのバージョン、エラーページの URL、クリック順序、および Network パネルの `/api/orders/:id` のレスポンスボディを提供してください。
```
AI は丁寧な定型文を書きがちですが、このスキルは「私たちがすでに知っていること」と「まだ不足していること」を区別するように強制します。
## wontfix も蓄積する
`/triage` は、enhancement の `wontfix` について興味深い設計をしています。単に issue を閉じるだけでなく、拒否理由を `.out-of-scope/` 知識ベースに書き込み、コメントでリンクします。
これにより、次回同様の要求があったときに、AI は同じ議論を最初からやり直すことはありません。まず `.out-of-scope/` を読み、メンテナーに「この方向は以前に X という理由で拒否されました」と注意を促すことができます。
これは ADR の精神とよく似ています。すべての決定を記録するのではなく、将来的に疑問を抱かせ、繰り返し発生する決定のみを記録します。
## 両者の連携方法
典型的なバグ issue は次のように進みます。
1. `/triage` が issue、コメント、ラベル、関連コードを読み込む
2. `bug + needs-triage` と判断する
3. まず再現を試みる。手順が不足している場合は `needs-info` に変更する
4. 情報が十分になったら、`/diagnose` を起動する
5. `/diagnose` が再現ループを確立し、仮説を立て、根本原因を特定する
6. 修正パスが明確で、テスト面が明確な場合、issue は `ready-for-agent` になる
7. 製品判断、外部権限、手動検証が必要な場合、issue は `ready-for-human` になる
8. 修正後、根本原因と回帰テストを issue または PR に書き戻す
このプロセスの鍵は「AI が自動的にバグを修正する」ことではなく、issue を曖昧な記述から実行可能な作業パッケージに変えることです。
## 私の使用提案
もし1つだけ覚えるなら:
> `/diagnose` はまず「それが壊れていることをどう証明するか」と問い、`/triage` はまず「それは今どの状態にあるべきか」と問います。
この2つの質問は、大量の低品質な AI プログラミングを防ぐことができます。
* 再現せずに修正する
* 受け入れ基準なしで作業を開始する
* 根本原因なしでリファクタリングする
* 情報なしでエージェントに丸投げする
Matt のこれら2つのスキルは派手ではありませんが、実際のチームのベテランエンジニアがすることによく似ています。まず事実を整理し、次にプロセスを進めます。
## 参考リソース
次の記事:[TDD:赤緑リファクタリングで AI に小さなステップを踏ませる](/ja/docs/notes/matt-pocock-skills/tdd)。
# Grill Me: コードを書く前に AI に 50 の質問をさせます
## 失敗モード: 「AI は私が望むことをしてくれませんでした」
Matt がスピーチで話した最初の失敗モードは、頭の中で要件が非常に明確であると考え、AI にそれを書き出してもらうことです。それはまったく当てはまりません。
> "I would run it, and I would try not to look at the code, but I would look at the code, and I realized I would get worse code. I did it again, I got even worse code... I did it again, kept running the compiler, and I would just end up with garbage."
多くの人がこの感覚をよく知っています。「ログインを追加してください」と言った場合、AI は「デバイスを覚えておきますか?」と尋ねることはありません。 「何回アカウントロックに失敗しましたか?」 「セッションが期限切れになるまでどれくらい時間がかかりますか?」合理的であると思われる計画を直接提示します。レビューするまでに 500 行が書かれており、やり直しに 2 時間かかります。
## なぜこのようなことが起こるのか: 設計コンセプトの逸脱
マットは、フレデリック・P・ブルックスの**デザインコンセプト**(デザインコンセプト)を『デザインのデザイン』で引用しています。
> 複数の人々が協力して何かをデザインするとき、あなたたちの間で何かが生み出されます - それは目に見えない「これについての理論」としてあなたの心の中に浮かんでいます。それは資産ではなく、マークダウン ファイルに詰め込まれた資産でもありません。それは目に見えない合意です。
AI はコードを思いつくとすぐに書きます。つまり、AI は同じ設計コンセプトをまったく共有していません。コードを書くときに問題があるのは構文ではなく前提です。
この問題を解決するには、開始する前に、まず設計コンセプトの調整を行う必要があります。 Brooks が提供したツールは **デザイン ツリー** と呼ばれるもので、決定を複数のブランチに分割し、各ブランチを分割します。上流の決定をスキップして下流の決定を直接行うことはできません。そうしないと、上流から下流に変更されたときにすべてをやり直す必要があります。
## マットのスキル全文
Matt による [`mattpocock/skills`](https://github.com/mattpocock/skills) でのこの理論の実装は `productivity/grill-me/SKILL.md` であり、ファイル全体と前付け部分を合わせた合計は 15 行未満です。
```markdown
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
文ごとに解釈します。
* **「執拗に面接してください」** - キーワードは *執拗* (手放さない) です。デフォルトでは、LLM は 1 つまたは 2 つの質問をした後、「ほぼ」と感じて行動を開始する傾向があります。この言葉はその傾向を強引に抑え込んでいる。
* **「デザイン ツリーの各枝を下っていく」** —— Brooks のデザイン ツリーの概念。クロードに要件をツリーとして扱い、最初に上流を解決し、次に下流を解決するように強制します。 「ログイン」と言うと、最初に「認証方法」(ツリールート)を尋ねられ、答えに基づいて「セッションの管理方法」/「トークンの保存方法」(サブノード)が展開されます。
* **「決定間の依存関係を 1 つずつ解決する」** - パッケージ化された質問を明示的に禁止します。多くの場合、決定の間には依存関係があります (SSO を選択した場合、ダウンストリームはパスワード ポリシーの問題を必要としません)。まず、アップストリームがダウンストリームの問題の多くを排除できることを確認してください。
* **「質問ごとに、推奨される回答を入力してください」** - 主要なボーナス ポイント。 AIは質問するだけでなく、回答も推奨します。うなずいたり「いいえ」するだけで、入力時間を 80% 節約できます。
* **「一度に 1 つずつ質問する」** - AI が一度に 10 個の質問をするのを防ぎます。
* **「コードベースを調べることで質問に答えられる場合は、代わりにコードベースを調べてください。」** - それがプロジェクトにすでに存在する事実 (「プロジェクトでどのテスト フレームワークが使用されているか」など) である場合は、質問せずにクロードに自分で見てもらいます。
7 行ですが、各文は特定の LLM 行動バイアスに対応しています。
## インストールと使用方法
**インストール**:
```bash
npx skills@latest add mattpocock/skills
```
`grill-me` と `setup-matt-pocock-skills` を確認します (グリルミーは後者に依存しませんが、他のスキルが依存するため、一緒にインストールすることをお勧めします)。
**呼び出し**: \[クロード コード] ダイアログに `/grill-me` と入力します。
**一般的なプロセス**:
1. やりたいことを説明しますが、非常に漠然としたものになる可能性があります (「ブログにコメント機能を追加したい」)
2.「`/grill-me`」と入力します。
2. クロードは 1 つずつ質問を開始し、それぞれの質問に対して推奨される回答を示しました。
3. 各質問に 1 つずつ答えます (うなずく/いいえ/正解)
4. 通常、20 ~ 50 問の質問で合意に達し、クロードが要約を教えてくれます。
5. 概要を [`/to-prd`](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues) に直接渡して PRD にすることも、[`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) に直接渡して書き込みを開始することもできます。
## 実際のケース: ビデオ編集機能の費用はいくらですか?
Matt は、[「私が毎日使用する 5 つのエージェント スキル」](https://www.aihero.dev/5-agent-skills-i-use-every-day) でいくつかの具体的な数字を示しました。
* **新しいビデオエディター機能** - 合意に達するための 16 の質問
* **複雑な関数** - 30\~50問
* **非常に複雑** - 100 の質問、セッションは最大 45 分
質問の例 (Matt のビデオ/ブログ投稿から復元):
* "Should video clips be reorderable, or only added/removed in sequence?"
* "When a clip is deleted, do we keep its source file, or delete the file too?"
* "Does the editor need undo/redo? How many steps deep?"
* "Should we render previews in the browser, or rely on a backend service?"
これらの問題はいずれも技術的なものではなく、すべて製品に関する決定でした。しかし、**すべての決定が、数百行のコードの形状を決定します**。これらの質問をスキップして、AI に直接質問を書かせた場合、AI は自動的に一連の答えを考え出し、書いた後に戻ってきて 1 つずつ拒否することになります。
## プラン モードとクロード コードの組み込みプラン モードの違い
クロード コードには `plan mode` が付属しています (Shift+Tab を押して入力します)。表面的には、グリルミーに似ています。行動を起こす前に、まず話し合ってください。しかし、マットはスピーチの中で、グリルミーの方が好きだと直接言いました。
具体的な違い:
| 寸法 | プランモード | /グリルミー |
| --------------- | ------------------------------------------ | ------------------------------------ |
| デフォルトの目標 | できるだけ早く実行可能な計画を作成します。まずは合意に達すること、計画は副産物である | |
| 質問数 | 0\~5 | 20\~100 |
| 質問形式 | 一度に 1 段落ずつ質問する | 一度に 1 つずつ質問する |
| 推奨される回答を提供しますか | いいえ | はい |
| コードベースを探索するかどうか | 時々 | 積極的に (明示的な指示) |
| シナリオに適しています | すでに明確に考えており、実行計画を確認したい | まだ明確に考えられていないため、明確に考えるように強制する必要があります |
実際的な最大の違いは、「**緊急かどうか**」です。プランモードは急いで始めますが、グリルミーは急いでいません。「明確に考える」ことをプロローグではなく主要なタスクとして捉えます。
## 高度な使用法
### 1. プログラミング以外のシナリオ
`grill-me` はコードをバインドしないため、純粋な製品に関する意思決定の会話にも使用できます。マット自身もそれを使用しています:
* コースのシラバスの設計
* 記事執筆
* 社内コミュニケーション文書
頭の中に漠然としたアイデアがあり、それをじっくり考えさせたい場合に使用できます。
### 2. [`/to-prd`](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues) に協力する
グリルミー セッションが終了したら、`/to-prd` と言うだけで、クロードが会話全体を構造化された PRD (ユーザー ストーリー、モジュール分割、テスト戦略を含む) に凝縮し、問題トラッカーに送信します。 **重要なポイント: 途中でコンテキストをクリアしないでください** - to-prd は会話コンテキストから直接抽出され、再度尋ねられることはありません。
### 3. [`/grill-with-docs`](/ja/docs/notes/matt-pocock-skills/grill-with-docs) に協力する
プロジェクトにすでに `CONTEXT.md` (ドメイン言語) と `docs/adr/` (アーキテクチャ上の決定) がある場合は、`/grill-me` の代わりに `/grill-with-docs` を使用します。拷問中、**同期的に CONTEXT.md** を更新します。ドキュメントが更新されている間に意思決定が行われるため、「ドキュメントが永久に古い」という問題はもうありません。
### 4. 質問の深さをカスタマイズする
時間がない場合は、`/grill-me` の直後に「質問は 10 個に制限し、アーキテクチャに関する決定のみに焦点を当ててください。」という文を追加できます。制限に応じて収束します。しかし、マットはそれをお勧めしません。彼は「もっと質問する」ことがまさにこのスキルの価値であり、それをカットすることは計画モードとほぼ同じであると信じています。
## 注意事項
**初めて実行するときはイライラするでしょう**。 「一文で500行を生成する」ことに慣れている人は、初めてAIに30回も質問されるのは時間の無駄に感じるでしょう。マットのアドバイスは、最初の 5 つの質問を保留することです。最初の 5 つの質問によって、考えもしなかったことが明らかになることがよくあります。その閾値を超えると、中毒になります。
**非常に小さなタスクには適していません**。タイプミスを変更し、console.log を追加します。grill-me は使用しないでください。 「新しいものを作る」場合や「副作用を伴う古いものを変更する」場合に適しています。
**AI は技術的な詳細を尋ねることがあります**。気にせず、判断させたい場合は、「あなたの呼びかけ」「あなたが決めます」と答えるだけで、それは受け入れて続行されます。
## なぜこのスキルが人気があるのでしょうか?
`/grill-me` は、Matt のスキルの中で最も頻繁にスクリーンショットされ、転送されるスキルです。理由は複雑ではありません。
1. **非常にミニマリスト**: 7 行のマークダウン、コピーして貼り付けるだけ
2. **即時効果**: 最初の実行時に AI の「問題密度」の変化を感じることができます。
3. **ポータブル**: Claude Code に依存せず、Codex、Cursor、Aider をすべて使用可能
4. **反 LLM のデフォルト動作が付属しています**: 各単語は反 LLM 逸脱であり、高度なエンジニアリングの美学を備えています。
その成功は、「**スキルは必ずしも長い必要はない**」ということを裏付ける最良の論拠にもなりました。
## 参考リソース
次の記事: [Grill With Docs: プロジェクト言語と ADR のメンテナンス](/ja/docs/notes/matt-pocock-skills/grill-with-docs) - ドメインが複雑なプロジェクト向けの Grille-me の高度なバージョン。
# Grill With Docs: ドメイン言語と ADR を使用して AI にプロジェクト メモリを装備する
## 失敗モード: 「AI が冗長すぎる」
マットのスピーチの 2 番目の失敗モード:
> AI は単純なことをたくさんの言葉で表現します。それはあなたに2つの言語を話しているようなものです。
これはコードの量とは関係なく、**語彙の間違い**です。 AI はデフォルトで一般的な用語 (「アイテム」、「データ」、「ハンドラー」) を使用しますが、あなたの頭の中にあるプロジェクトの実際の用語は、「コース」、「ドラフト バージョン」、「ゴースト レッスン」などである可能性があります。 AI はこれらの単語がプロジェクト内で特定の意味を持っていることを認識していないため、それらの周囲に同義の新しい単語を大量に作成します。結果は次のとおりです。
* 長々とした思考プロセス (独自の言葉を避ける)
* 実装が頭の中にある設計とずれている (同じ意味空間にないため)
* セッション間では再利用できません (会話ごとにコンテキストを再確立する必要があります)
## 古典的な理論: DDD のユビキタス言語
マット氏はエリック・エヴァンスの「ドメイン駆動設計」を引用した。この本は 2003 年に出版され、**ユビキタス言語** の概念を提案しました。
> 同じ一連の用語を使用して、ドメインの専門家、開発者、コードを結び付けます。製品ディスカッション、コード コメント、変数名、ドキュメント内の単語 - それらは同じことを意味する必要があります。
DDD の目標は、コードをドメイン専門家の頭脳のように見せることです。 AI 時代には、**LLM もこの言語である必要がある**という新しい役割があります。 LLM はスタンドアップ ミーティングに参加しておらず、製品要件の会議も見ることができず、グループのスラングも理解できません。LLM は、与えられた文書からしか学ぶことができません。
Matt はこれをスキルに変えました。コード ベースをスキャンして用語を抽出し、マークダウン ファイル `CONTEXT.md` を生成し、それを人間と AI の両方に合わせて調整します。
## スキルの進化: ユビキタス言語からドキュメントを使ったグリルまで
最も初期のスキルは `ubiquitous-language` と呼ばれていました。これは、コード ベースをスキャンして用語集を生成するという 1 つのことだけを行いました。しかしマットは後に、単にドキュメントを生成するだけでは十分ではないことに気づきました。
* **ドキュメントは古くなります**: ドキュメントは今日生成され、コードは明日変更されますが、用語集は更新されていません。
* **人々は率先してそれを見ようとはしません**: そこに置いたら死んでしまいます
彼はそれを `grill-with-docs` にリファクタリングし、次の 3 つを組み合わせました。
1. **拷問条件** (グリルミーからすべての能力を継承)
2. **既存の用語集に異議を唱えます**: あなたが言った単語は CONTEXT.md に書かれていることと矛盾していますか?すぐに指摘してください
3. **決定を下すときに同時に文書を更新します**: 拷問プロセス中に到達した新しい結論は CONTEXT.md にインラインで書き込まれるか、新しい ADR を作成します。
これは、「静的な文書の生成」から「対話は文書の維持である」へのパラダイム シフトです。
## スキル全文
`engineering/grill-with-docs/SKILL.md` のコア構造:
```markdown
---
name: grill-with-docs
description: Grilling session that challenges your plan against the
existing domain model, sharpens terminology, and updates
documentation (CONTEXT.md, ADRs) inline as decisions crystallise.
---
Interview me relentlessly about every aspect of this plan until we
reach a shared understanding. Walk down each branch of the design
tree, resolving dependencies between decisions one-by-one. For each
question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question
before continuing.
If a question can be answered by exploring the codebase, explore the
codebase instead.
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
If a CONTEXT-MAP.md exists at the root, the repo has multiple contexts.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in
CONTEXT.md, call it out immediately.
"Your glossary defines 'cancellation' as X, but you seem to mean Y —
which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise
canonical term.
"You're saying 'account' — do you mean the Customer or the User?
Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with
specific scenarios.
### Cross-reference with code
When the user states how something works, check whether the code
agrees. If you find a contradiction, surface it.
### Update CONTEXT.md inline
When a term is resolved, update CONTEXT.md right there. Don't batch
these up — capture them as they happen.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. Hard to reverse
2. Surprising without context
3. The result of a real trade-off
```
## 実際の CONTEXT.md はどのようなものですか?
Matt 自身の [`course-video-manager`](https://github.com/mattpocock/course-video-manager/blob/main/CONTEXT.md) リポジトリには、完全な CONTEXT.md の例が含まれています。それを理解するためにいくつかの用語を選択してみましょう。
| 用語 | 定義 |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| **Course** | The primary domain entity: a structured collection of versions, sections, lessons, and videos |
| **Draft Version** | The single mutable CourseVersion that is currently being edited; always the latest by `createdAt` |
| **Published Version** | An immutable CourseVersion with a name and description, created by the Publish flow |
| **Ghost Lesson** | A lesson that exists in the database but not yet on the file system (`fsStatus = "ghost"`) |
| **Export Hash** | A SHA256 hash derived from a video's clip filenames, timestamps, clip order |
| **Unexported Video** | A video whose current Export Hash does not match any file on disk; blocks publishing |
| **Materialization Cascade** | The chain reaction when materializing a lesson inside a ghost course |
| **Clip** | A timestamped segment of source footage within a video |
| **Fractional Index** | A string-based ordering value that allows inserting items between existing items |
| **Purge** | The deliberate deletion of an Exported Video's `.mp4` file from disk |
いくつかの点に注意してください。
1. **各用語は動名詞または固有名詞です** – 「注文状況」のような説明的なフレーズではありません
2. **各定義は他の用語を参照** (コース → バージョン → レッスン → ビデオ) オントロジー ネットワークを形成します
3. **コード フィールドが直接表示されます** (`fsStatus = "ghost"`) - ドキュメントとコードの 1:1 マッピング
4. **意思決定の説明を含める** (「公開のブロック」、「連鎖反応」) - 名詞だけでなくルールも含めます
コードを書くとき、AI がこのドキュメントを見たとき、「ファイルのないレッスン」の代わりに「ゴースト レッスン」を使用します。コード、会話、コミットメッセージはすべて統合されています。
## ADR: 作成時
スキルには重要な制約があります。
> Only offer to create an ADR when all three are true:
>
> 1. **Hard to reverse** — the cost of changing your mind later is meaningful
> 2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
> 3. **The result of a real trade-off** — there were genuine alternatives
ADR (Architecture Decision Record) の文化は、2011 年の Michael Nygard のブログから生まれましたが、多くのチームがすべての意思決定について ADR を作成するために ADR を使用しており、20 の ADR のうち 18 がアカウントを実行しています。 Matt この三角測量は素晴らしいツールです。**ADR は 3 つの条件が同時に満たされた場合にのみ価値があります**。それ以外の場合は、CONTEXT.md 内でダイジェストされ、コード内でダイジェストされます。
## インストールと使用方法
**前提条件**: 最初に `/setup-matt-pocock-skills` を実行します (CONTEXT.md を配置する場所と ADR ディレクトリを配置する場所を尋ねられます)。
**電話**: `/grill-with-docs`
**一般的なプロセス**:
1. やりたいことを説明する
2. `/grill-with-docs`
3. クロードは最初に CONTEXT.md と docs/adr/ を **スキャン**し、既存の用語と決定をコンテキストにロードします
4. プロセス中に拷問を開始します。
* CONTEXT.mdと矛盾する単語を使用 → その場で指摘
* 曖昧な単語を使用した (たとえば、「アカウント」は顧客またはユーザーのいずれかになる可能性があります) → 2 つのうちの 1 つを選択して文書にドロップします。
* 指摘された動作は既存のコードと矛盾しています → 矛盾を指摘してください
5. 決定に達したら **CONTEXT.md** を同期的に更新します (バックログやバッチ処理はありません)。
6. 重要な不可逆的な決定 → ADR を生成するかどうかを尋ねる
プロジェクトに CONTEXT.md と docs/adr/ がまだない場合は、遅延して作成されます。最初の用語を書き、最初の ADR を構築する必要があるまで、ファイルは生成されません。最初から空のテンプレートを提供するわけではありません。
\##複数のコンテキスト プロジェクト (CONTEXT-MAP.md)
プロジェクトが大きすぎて 1 つの CONTEXT.md に収まらない場合 (たとえば、注文と請求が 2 つの独立した境界付きコンテキストである場合)、`CONTEXT-MAP.md` を一般ディレクトリとしてルート ディレクトリに置くことができます。
```
/
├── CONTEXT-MAP.md ← 总目录
├── docs/adr/ ← 系统级决策
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← 模块级决策
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
`/grill-with-docs` は CONTEXT-MAP.md の存在を自動的に認識し、対応するサブディレクトリにジャンプします。これは、DDD の **境界コンテキスト** の概念を直接実装したものです。各コンテキストの「順序」は異なる意味を持つ可能性があり、汚染を避けるために個別に維持されます。
\##とグリルミーの違い
| 寸法 | /グリルミー | /グリル付きドキュメント |
| ------------------- | ---------------- | -------------------------------------------------- |
| 質問力 | ✅ | ✅ (すべて継承) |
| プロジェクト用語の確認 | ❌ | ✅ |
| リアルタイム更新 CONTEXT.md | ❌ | ✅ |
| ADRトリガー判定| ❌ | ✅ | |
| 適用ステージ | 初期のアイデア、個人プロジェクト | ドメインが複雑な実際のプロジェクト |
| 初期費用 | 0 | セットアップが必要 + プロジェクトに CONTEXT.md をビルドしている/ビルドする意思がある |
単純かつ大雑把な判断:
* **個人の脚本、記事の執筆、コースの指導** → `/grill-me`
* **長期的なメンテナンスが必要な実際のプロジェクト** → `/grill-with-docs`
## 直観に反する利点: AI に「黙る」ことを学習させる
CONTEXT.md は AI だけのためのものではなく、将来の AI セッションのためのものです。新しい会話が始まるたびに、クロードは CONTEXT.md を読み取ることでプロジェクトのコンテキストに即座に入ることができ、長いオンボーディングを節約できます。
さらに微妙なのは、マットはスピーチの中で、CONTEXT.md を追加した後、AI の *思考痕跡* を見ることができると述べました—
> 「これにより、AI はあまり冗長ではない方法で考えることができるようになります。」
なぜですか? CONTEXT.md がないと、AI は考えるときに常に独自の用語を定義しなければなりません - 「ユーザー、つまり商品を注文した人を指します。以下、...」と呼びます。 CONTEXT.md を使用すると、「顧客」と直接表示されるため、思考チェーンがはるかに短くなり、応答が速くなります。
**LLM のトークンエコノミクスは、思考経路の短縮 = より速く、より正確で、より安価な成果物であると判断します**。 CONTEXT.md は、この効率化のための隠れたレバーです。
## 注意事項
**CONTEXT.md は最初の実行時に大幅に変更されます**。プロジェクトにすでに手書きの CONTEXT.md がある場合は、最初に git stash するか、実行する前に予行演習します (プロンプトに「最初に変更する内容をリストします。ファイルを直接書き込まないでください」という文を追加できます)。
**ADR モデレーションは実質モデレーションです**。興奮してすべての決定で ADR を生成しないでください。6 か月後には docs/adr/ がゴミでいっぱいになってしまいます。 Matt の 3 つの基準は厳密に従う必要があります。
**CONTEXT.md 実装の詳細を記述しないでください**。 Skill には、「CONTEXT.md を実装の詳細と組み合わせないでください。ドメインの専門家にとって意味のある用語のみを含めてください。」という言葉があります。 CONTEXT.md に「PostgreSQL」と記述するのは間違いです。ドメインの専門家はデータベースの選択を気にしません。これは ADR の問題です。
## 参考リソース
次の記事: [to-PRD + to-Issues: 対話から実行可能なチケットへ](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues)——拷問の後、対話を実行可能な作業単位に固める方法。
# コードベース アーキテクチャの改善: 浅いモジュールを深いモジュールに再構築します。
## 障害モード: 「AI が不正なコードベースをさまよっている」
Matt の講演の 4 番目の失敗モードは、次のような絵の比喩です。
> 「コードベース内の浅いモジュールは次のようになります。小さな BLOB が多数あり、AI はそれらを修正する前に、多数のモジュールを調べてすべての依存関係を理解する必要があります。」
> "AI is really good at creating codebases like this. So you'll have a situation where AI doesn't understand what your code is doing. It will attempt to explore the code, but because it's poorly laid out, filled with shallow modules, it doesn't get to the right module in time, or doesn't understand all the dependencies."
これは AI プログラミング特有の悪循環です。
```
AI 写代码倾向于产生 shallow 模块(小、多、互相依赖)
↓
代码库变得 shallow
↓
下次 AI 进来探索更难,更容易写错
↓
更多 shallow 模块被加进去
↓
代码库越来越烂,AI 越来越无能
```
このサイクルを断ち切るには、手動によるリバース リファクタリングを定期的に実行し、浅いモジュールを深いモジュールにマージする必要があります。まさに `/improve-codebase-architecture` が行うことです。
## 古典的な理論: Ousterhout のディープ モジュール
John Ousterhout はスタンフォード大学の CS 教授 (Tcl 言語と Raft 論文の著者) です。彼の 2018 年の著書「A Philosophy of Software Design」では、シンプルだが強力な定規を提案しています。
**モジュールの「深さ」 = インターフェースの隠れた複雑さ**
| タイプ | インターフェース | 実装 | 画像 |
| ------ | -------- | ---- | ---------- |
| **深い** | シンプル | リッチ | 長方形: 狭くて深い |
| **浅い** | 複雑な | シンプル | 長方形:広くて浅い |
理想的なモジュールは奥深く、ユーザーは短いインターフェイスを見るだけで済み、複雑さは内部に隠されています。極端な反例は浅いモジュールです。インターフェイスは実装とほぼ同じくらい複雑です。これは、カプセル化がないことを意味します。ユーザーは実装を直接見ることもできます。
Ousterhout の判断: **優れたコード ベースは、少数の深いモジュールで構成されています。悪いコードベースは、多数の浅いモジュール**で構成されています。これは、「関数をできるだけ小さくし、ファイルをできるだけ短くし、モジュールをできるだけ多くする」という従来の定説とは完全に反対です。彼は、そのような定説がまさに浅いモジュールを生み出すと信じています。
## Matt の拡張機能: 削除テスト
マットは、オーステルハウトの理論を運用工学テストに翻訳し、**削除テスト**と呼びました。
> **Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.**
人間の言葉:
* **削除すると複雑さがなくなります** → このモジュールは本来パススルー (トランジット) であり、動作しないので切断してください
* **削除すると、複雑さが N 人の発信者に広がります** → 元々は複雑さを隠すのに役立ちます。非常に奥深いので、そのままにしてください。
このテストの利点は、双方向であることです。「削除する必要がある薄いパッケージ」と「抽出する必要がある共通ロジック」の両方を識別できます。コードの一部を削除した後に複雑さが 5 か所に広がることがわかった場合、このコードは深いモジュールに抽出する価値があることを意味します。
## 重要な用語 (Matt の正確な定義)
`improve-codebase-architecture/SKILL.md` には **これらの単語を厳密に使用する**ことが必要な用語集があります。「コンポーネント」、「サービス」、「API」、および「境界」に偏らないようにしてください。
| 用語 | 定義 |
| ------------- | ----------------------------------------------------------------- |
| **モジュール** | インターフェイスと実装 (関数/クラス/パッケージ/スライス) を持つものすべて |
| **インターフェース** | 呼び出し元が知っておくべきすべて - 型、不変式、エラー モード、順序、構成 (関数のシグネチャだけでなく) |
| **実装** | モジュール内のコード |
| **深さ** | インターフェイスにあるレバー。深い = 高レバレッジ、浅い = インターフェースは実装とほぼ同じ複雑さ |
| **シーム** (シーム) | インターフェイスの場所 - インプレース変更を行わずに動作を変更できる場所。 **「境界」ではなく「継ぎ目」を使用してください** |
| **アダプター** | インターフェイスの特定の実装をシームで実装する |
| **レバレッジ** | 呼び出し元が「ディープ」から得られるメリット |
| **地域** | メンテナが「深さ」から得られる利点 - 変更、バグ、知識がすべて 1 か所に集中している |
いくつかの基本原則:
* **削除テスト**: 上記を参照
* **インターフェイスはテスト面です**: テストはインターフェイスを通じてのみ実行できます。これはモジュールの詳細なテスト容易性の基礎です。
* \*\* 1 つのアダプター = 仮想の継ぎ目。 2 つのアダプター = 本物のシーム\*\*: 実装が 1 つだけのインターフェイスは偽シームです。 **実際のジョイントには少なくとも 2 つのアダプターが必要です**
最後のものは特に直感に反します。多くのチームは「将来の拡張に備えて」事前にインターフェイスを抽象化しますが、実際には実装は 1 つだけです。 Matt の判断: **無駄、削除**。 2番目が実現するまで待ちます。これはYAGNIと同じ起源を持ちます。
## スキルのワークフロー
### 1. 探索する
スキルはまず AI に `CONTEXT.md` と `docs/adr/` を読み取らせ、次に `subagent_type=Explore` を使用してサブエージェントをコード ベースに送信します。
厳密なインスピレーションの代わりに、**摩擦**を信号として使用します。
> * Where does understanding one concept require bouncing between many small modules?
> * Where are modules **shallow** — interface nearly as complex as the implementation?
> * Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
> * Where do tightly-coupled modules leak across their seams?
> * Which parts of the codebase are untested, or hard to test through their current interface?
疑わしい点を見つけるたびに、削除テストを適用します。削除すると複雑さが消えるか、それとも広がるか? 「分散する」という答えは、さらに深める価値のある候補です。
### 2. 現在の候補者 (発言候補者)
候補者の番号付きリストを提示します。
```
1. Files: src/orders/parser.ts, src/orders/validator.ts, src/orders/normalizer.ts
Problem: 三个文件互相调用,理解 Order 入站需要在三处跳转
Solution: 合并为单一 OrderIntake 模块,对外只暴露 parse(raw) → ValidatedOrder
Benefits:
- Locality: Order 入站的所有逻辑、错误处理、bug 修复集中一处
- Leverage: 调用方从理解 3 个接口降为 1 个
- Tests: 只需测 parse() 的输入输出,不再需要 mock 内部协作
```
要件:
* **CONTEXT.md 語彙**を使用してドメイン (「FooBarHandler」ではなく「注文受付モジュール」) について話します。
* **用語集** (「継ぎ目」、「深さ」、「局所性」) を使用して建築について話します。
* **インターフェイス デザインをすぐに提案しないでください** – ユーザーが最初に興味のある候補を選択できるようにします
候補者が既存の ADR と競合する場合 - 競合により ADR を再検討する必要がある場合にのみ言及し、それを明確にマークします。
> "contradicts ADR-0007 — but worth reopening because…"
ADR で禁止されているリファクタリングをすべて掘り出さないでください。
### 3. グリルループ
ユーザーが候補を選択した後、グリル モードにドロップします ([`/grill-with-docs`](/ja/docs/notes/matt-pocock-skills/grill-with-docs) から継承)。
* デザイン ツリーをウォークスルーします - 制約、依存関係、深くした後のモジュールの形状、継ぎ目の後ろに隠れているもの、どのテストが生き残れるか
* **副作用はすぐに発生します**:
* 深化モジュールにCONTEXT.mdにない名前を付ける→すぐにCONTEXT.mdに追加
* 拷問で曖昧な用語が先鋭化 → CONTEXT.mdをすぐに更新
* ユーザーは負荷を理由に候補者を拒否 (重要、将来の探索者が知っておくべきこと) → ADR の生成を提案
* 深化モジュールのさまざまなインターフェイス設計を検討したい → `INTERFACE-DESIGN.md` 別のプロセスにジャンプ
ドキュメントの保守とアーキテクチャの変革は同じ会話で行われます。2 回繰り返すことはありません。
## 実際のケース: メイバ・アーメドの実践
サードパーティ開発者の Mejba Ahmed は、このスキルを使用した経験を詳細に記録する記事 \["Deep Modules: The Claude Code Skill Saving My Codebase"] ([https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules](https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules)) を作成しました。重要なポイント:
* 彼はもともとプロジェクトに 50 以上のファイルを持っていましたが、各ファイルの行数は 100 行未満でした - **典型的な浅いライブラリ**
* `/improve-codebase-architecture` は 8 つの深化候補を使い果たしました
* 彼は 3 つの深化を選択しました (2 つのデータ処理モジュールが統合され、1 つのツールセットが統合されました)
* 結果: ファイル数は 50 以上から 30 以上に減少しましたが、**合計コード サイズは基本的に同じままです** - 複雑さは少数の深いモジュールに圧縮されています
* このライブラリにおけるクロードのその後のコード変更のヒット率が大幅に向上しました (彼は「60% から 90%」と述べました。これは厳密に測定したわけではありませんが、強いと感じました)
Mejba は次の注意事項もあります: **一度に 8 つずつ深めないでください**。一度に 1 つだけを選択し、テストを実行し、コミットし、観察してから、次のものを選択します。それ以外の場合、完了後にロールバックする方法はありません。
## インストールと使用方法
```bash
npx skills@latest add mattpocock/skills
```
`improve-codebase-architecture` + `setup-matt-pocock-skills` を確認してください。
**電話**: `/improve-codebase-architecture`
**推奨リズム**:
* **週に 1 回、または各スプリントの最後に実行**
* または \*\*集中的な開発の波が完了した後に一度実行します (AI コードを高頻度で作成した後は、浅いモジュールを積み上げることが特に簡単です)
* **急いでいるときは実行しないでください** - 急いでいるときには消化する時間がないような大きな変化を示唆します。
**一般的なプロセス**:
1. `/improve-codebase-architecture`
2. AI探索 + N個の候補リスト(削除テスト引数付き)
3. 最も感じるものを選択します
4. グリルループ調整デザインにドロップ
5. AI 実装のリファクタリング ([`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) と一緒に実行することをお勧めします。リファクタリングにはテスト保護が必要です)
6. コミット + 観察
7. 1週間後にまた来てください
## このスキルが Matt のワークフローの「クローズド ループ」であるのはなぜですか?
Matt のワークフロー図に戻ります。
```
/grill-me → /to-prd → /to-issues → /tdd → /improve-codebase-architecture → 回到 /grill-me
```
開始点にループバックしていることに注意してください。 `/improve-codebase-architecture` は 1 回限りのツールではなく、**定期的なメンテナンス**です。理由は次のとおりです。
1. AI はコードベースに浅いモジュールを追加し続けます (これはデフォルトの傾向であり、書きすぎると山積みになります)
2. ビジネスが進化し続けると、古い継ぎ目は時代遅れになります。
3. CONTEXT.md の用語は引き続き改良されており、古い名前は維持されません。
**このスキルを実行するたびに、コードベースの AI フレンドリーさが更新されます**。これは、LLM でコードベースを長期的に健全に保つ**唯一**の方法です。更新しないと、3 か月後に AI がコードベースで停止します。
## この一連の考え方はスキルそのものよりも価値があります
`/improve-codebase-architecture` をまったくインストールしなくても、次の 3 つのことを覚えておくだけで、PR レビューの品質がワンランク上がります。
1. **削除テスト**: 新しいモジュールを見つけるたびに、「削除したら、複雑さは消えるのか、それとも広がるのか?」と自問してください。
2. **少なくとも 2 つのアダプターを真に継ぎ目**: 単一の実装インターフェイス = false abstract、削除
3. **インターフェースはテスト面です**: 測定できない = インターフェースの設計に問題があります
これら 3 つの項目には AI やスキルは必要ありません。これらはエンジニアリングの美学の貴重な通貨です。 Matt はこれらをバッチ実行用のスキルにラップしていますが、実際に活用するのは 3 つの原則そのものです。
## 注意事項
**あまり深く入らないでください**。 Ousterhout 自身は、ディープ モジュールは教義ではなく目標であると述べました。すべての機能を詰め込んだ大きな Util クラスはディープ モジュールではなく、神モジュールです。判断基準は「シンプルなインターフェース+一貫した実装」であり、両方を満たす必要があります。
**深化にはテスト保護が必要です**。構造変更はリスクの高い操作であり、テストせずにあえてリファクタリングを行う = 責任を負うのを待つことになります。現在テストがない場合は、[`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) に移動してクリティカル パスにテストを追加してから戻ってきます。
**ADR の決定は気まぐれに行われるべきではありません**。グリル中に候補者を拒否すると、AI は簡単に ADR の生成を提案します。その理由が本当に「将来の人が知っておく必要がある」場合にのみ、ADR を受け入れます。そうしないと、docs/adr/ がジャーナル エントリでいっぱいになります。
**コードの一部が浅くても問題ありません**。ロガーラッパー、定数ファイル、一回限りのスクリプト - それらは浅いので問題ありません。このスキルは、抽象化を支援するふりをしながら、実際には混乱を加えている浅いモジュールを探します。
## 参考リソース
***
## シリーズの結論
これまでの6つの記事をすべて読みました。ワークフロー全体を確認します。
```
/grill-me 或 /grill-with-docs ← 谈清楚要做什么
↓
/to-prd ← 凝固成 PRD
↓
/to-issues ← 切成 vertical slice
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture ← 周期性深化
↓
回到 /grill-me
```
このプロセスの精神は次の 1 つの文に凝縮できます。
> \*\*AI は地上の戦術兵士であり、あなたは戦略層です。 「問題の定義」「問題の分解」「問題のテスト」の3つを持ち帰って自分でやり、「コードを書く」ことはAIに任せる、これがAI時代のエンジニアの本当の立場です。 \*\*
Matt の一連のスキルは究極の答えではありませんが、現段階でのベスト プラクティスです。 3 か月後にはもっと良いものがあるかもしれませんが、精神的な側面は変わりません。良いコード ベースは悪いコード ベースよりも常に重要であり、基本的なソフトウェア スキルは常に価値があります。
[概要](/ja/docs/notes/matt-pocock-skills/overview) に戻るか、最も便利なスキルを選択してインストールして試してください。
# ソフトウェアの基礎はこれまで以上に重要です: Matt Pocock の Claude Code スキルセット
## スペックからコードへの波に圧倒されながらも冷静な人
2026 年は、AI プログラミングの「仕様からコードへ」の物語のハイライトの年です。仕様を作成し、コンパイラを実行します。コードを読んでから仕様を作成し、その後コンパイラを実行するのではありません。コミュニティで浮上したスローガンは「**code is easy**」(コードは安い)で、これは「とにかく、AI はさらに 1 秒あたり 10,000 行生成できるのに、なぜそれを気にする必要があるのか」を意味します。
マット・ポーコックは、これに公に反対の声を上げる数少ない人物の一人だ。彼は AI コーディングが非常に強力であることを否定しませんが、クラス「本物のエンジニアのためのクロード コード」で仕様からコードまでを実際にテストしました。その結論は非常に胸が張り裂けるものでした。**コードを実行するたびに、コードはますます悪化します**。これはまさに、『Pragmatic Programmer』で話題になった「ソフトウェア エントロピー」、つまりソフトウェア エントロピーの増加です。
そこで彼は次の 2 つのことを行いました。
1. この観察を 18 分間のトークにまとめます: *ソフトウェアの基礎はこれまで以上に重要です*。
2. 対応する解毒剤を GitHub リポジトリにパッケージ化します: [`mattpocock/skills`](https://github.com/mattpocock/skills) - 「本物のエンジニアのためのスキル。私の .claude ディレクトリから直接。」
ウェアハウスは 2026 年 2 月 3 日に開始され、4 か月以内に **61.1k スターと 5.3k フォーク**に達しました。これは、同時期に最も急速に成長した AI プログラミング ウェアハウスの 1 つでした。
***
## マット・ポーコックとは誰ですか?
TypeScript を書いたことがある人なら、おそらく一度は目にしたことがあるでしょう。彼は、近年中国語と英語界で最も多作な TypeScript 教育者の 1 人です。
* 英語界で非常に人気のある一連の有料コースである **TotalTypeScript.com** の創設者
* **aihero.dev** ニュースレター購読数 60,000 以上、トピックは TS から AI コーディングに変更
* Twitter [`@mattpocockuk`](https://twitter.com/mattpocockuk) と YouTube `@mattpocockuk` には短いビデオ チュートリアルがたくさんあります。
* OpenAI/人類主義者ではなく、純粋に独立した開発者と教育者の経歴を持つ
彼の性格は非常に明確です。**シニア エンジニアの視点から見た AI コーディング**。私たちは「AGI がやってくる」とは叫びませんし、「プログラマーが職を失う」とも叫びません。彼が叫んだのは、「古い世代のソフトウェア エンジニアのトリックは今でも非常に役に立ちます。LLM が実行できる形式に変換するだけで十分です。」というものでした。
***
## 主な議論: コードは安くない
スピーチ全体に議論は 1 つだけあり、各スキルはその脚注です。
> コードベースの構造が悪い場合、AI は悪いコードベースに悪いコードしか書きません。したがって、**優れたコードベースがこれまで以上に重要になり、基本的なソフトウェア スキルがこれまで以上に重要になります**。
マットは軍事に例えて、人間と AI の役割を非常にわかりやすく説明しました。
戦略層は何をするのでしょうか?設計コンセプト、統一言語、モジュール境界 - これら 3 つは「コードを書く」というよりは「問題を定義する」もので、LLM が最も苦手とするものです。
***
## 5 つの失敗パターン → 5 つの古書 → 5 つのスキル
Matt はスピーチの中で、方法論全体をマッピング テーブルに圧縮しました。障害モードに遭遇するたびに、彼は 20 年前に解決された古典的な理論を示し、マークダウン形式のスキル ファイルを提供します。
| # | AIプログラミング失敗モード | 古典的な理論と情報源 | 対応スキル |
| - | ----------------------------------------------------------------------- | ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| 1 | AIはあなたが望むことをしません | *デザインのデザイン* (ブルックス)—— デザインコンセプト、デザインツリー | [`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me) |
| 2 | AI は冗長な言葉であなたに話しかけます。 *ドメイン駆動設計* (Evans) - ユビキタス言語 | [`/grill-with-docs`](/ja/docs/notes/matt-pocock-skills/grill-with-docs) | |
| 3 | AI は正しく実行しますが、実行できません。 *実践的なプログラマー* (ハント & トーマス)—— 「フィードバックの速度が速度の限界です」 | [`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) | |
| 4 | AI は悪いコードベースを歩き回ります | *ソフトウェア設計の哲学* (Ousterhout) - ディープモジュール、削除テスト | [`/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) |
| 5 | あなたの脳は AI の出力に追いつけない | Kent Beck —— 毎日デザインに投資 | 「インターフェイスを設計し、実装を委任する」 |
項目 5 リポジトリには個別のスキルはありません (以前は `design-an-interface` がありましたが、非推奨になりました)。その精神は [`/to-prd`](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues) と [`/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) に吸収されています。どちらも、コードを記述する前にモジュール インターフェイスについて考える必要があります。
***
## 5 毎日実際に使用するスキル
そのスピーチは哲学的な骨組みでした。 Matt はその後、そのスケルトンを毎日のワークフローに変換する記事「私が毎日使う 5 つのエージェント スキル」を aihero.dev に投稿しました。これら 5 つは、このシリーズの将来的に 1 つずつ解体されるオブジェクトです。
```
/grill-me ← 先和 AI 谈清楚要做什么
↓
/to-prd ← 把对话凝固成 PRD
↓
/to-issues ← 把 PRD 切成可独立领取的 vertical slice
↓
/tdd ← 每个 slice 用红绿重构跑通
↓
/improve-codebase-architecture ← 周期性检查,把 shallow 模块改成 deep
```
これら 5 つのスキルが結びついて、Matt の完全な研究開発プロセスが形成されます。各ステップに対応する故障モードは、前のセクションの表に示されています。
各記事の詳細な分解 (このシリーズの後続ページ):
* [グリル ミー: AI にあなたのニーズを教えてもらいましょう](/ja/docs/notes/matt-pocock-skills/grill-me)
* [Grill With Docs: プロジェクト言語と ADR のメンテナンス](/ja/docs/notes/matt-pocock-skills/grill-with-docs)
* [to-PRD + to-Issues: 会話から実行可能なチケットまで](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [TDD: 赤と緑の再構築を使用して AI に小さなステップを強制する](/ja/docs/notes/matt-pocock-skills/tdd)
* [コードベース アーキテクチャの改善: 浅いモジュールを深いモジュールに再構築](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture)
***
## インストール方法
ウェアハウスの README には、インストールするための 1 行コマンドが記載されています。
```bash
npx skills@latest add mattpocock/skills
```
このコマンドは次のことを行います。
1.インストールしたいスキルを確認します
2\. インストールするエージェントを選択できます (Claude Code、Codex、Cursor などがすべてサポートされています)
3\. 対応する SKILL.md ファイルを `.claude/skills/` (またはエージェントに対応するディレクトリ) に配置します。
**`/setup-matt-pocock-skills`** を同時にチェックすることを強くお勧めします。これは 3 つの質問を行う 1 回限りの構成スキルです。
* **問題トラッカー**は何に使用されますか? (GitHub / GitLab / ローカルマークダウン / その他)
* **トリアージラベル**にはどのような単語が使われていますか? (ニーズのトリアージまたはその他)
* **ドメイン ドキュメント**をどこに配置しますか? (CONTEXT.md/ADR パス)
`/setup-matt-pocock-skills` を 1 回実行すると、プロジェクトのルート ディレクトリの `AGENTS.md` または `CLAUDE.md` に書き込みます。その後、すべてのエンジニアリング スキル (to-prd、to-issue、トリアージ、tdd など) がこの構成を自動的に読み取ります。このステップは省略され、後続のすべてのスキルで同じ質問が何度も繰り返されます。
`/grill-me` (最も軽量で純粋な生産性クラス) を試してみたいだけの場合は、問題トラッカーに依存しないため、セットアップをスキップできます。
***
## この一連のスキルと BMAD / Spec-Kit / GSD の違い
[BMAD](/ja/docs/notes/speckit/concept)、Spec-Kit、GSD などの仕様主導フレームワークをすでに使用している場合は、「なぜまだ Matt のセットが必要なのですか?」と疑問に思うかもしれません。
Matt は README に次のように直接書いています。
**主な違い**:
* BMAD/Spec-Kit/GSD は、仕様からコードまでの完全なパイプラインを指定する **フレームワーク** です。そのプロセスに従う必要があります。
* マット このセットは**コンポーネント**です。各スキルには、数行から数十行までのマークダウン ファイルがあります。いつでも分解して変更できます。
例: `grill-me` の実際の全文はこれだけです —
```markdown
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
スキル全体は7行です。しかし、クロードが決断を下す前に 20、50、さらには 100 の質問をするのは、この 7 行のせいです。 **非常に少ないテキストを使用して大きな動作の変化を活用する**という設計哲学が、この一連のスキルの人気の根本的な理由です。
***
## このシリーズの読み方
これまでに Matt のセットに出会ったことがない場合は、meta.json 内の順序で読むことをお勧めします。
1. **概要** (現在閲覧している記事) - 全体像を把握する
2. **Grill Me** - 個別にインストールして、最初に試してみます。しきい値は最も低いです
3. **Grill With Docs** - Grille-me の高度なバージョン。CONTEXT.md の導入を開始します。
4. **to-PRD + to-Issues** - 会話を実行可能なチケットに変える
5. **TDD** - マット自身、「エージェントの出力品質を向上させるためにこれまで使用した中で最も安定した方法」と述べています。
6. **コードベース アーキテクチャの改善** - AI を長期的に利用できるようにするための定期的なメンテナンス
すでにクロード コードを使用して実際のプロジェクトを作成している場合は、最も直感的な記事である [`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me) + [`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) に直接ジャンプしてください。
教えたり書いたりしている場合は、[`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me) を読むだけで十分です。これは、コードに限定されない一般的な「デザイン会話」ツールです。
***
## 私の使用方法の提案
この一連のスキルを自分でインストールした後、最大の物理的な変化は次のとおりです。
**最初に**: 急いでコードを書き始めるのはやめてください。以前は、AI が「ログインを追加してください」を受信すると、500 行のレイアウトを開始していました。ここで、`/grill-me` は最初に 20 の質問をします - 「デバイスを覚えておきたいですか?」 「セッションが期限切れになるまでどれくらい時間がかかりますか?」 「アカウントのロックに何回失敗しましたか?」 30分後に書きましょう。節約できるのは、その後の 2 時間の再作業です。
**2 番目**: CLAUDE.md は肥大化しなくなりました。以前、CLAUDE.mdには「要件を理解した上でコードを書いてください」「抽象的になりすぎないでください」など多くの禁止事項が書かれていましたが、クロードはそれでもそれを犯していました。 Matt のセットに切り替えた後、CLAUDE.md にはドメイン知識 (設計システム、コンポーネント仕様、展開) のみが組み込まれ、一般的な方法論はスキルに引き継がれます。双方の責任は明らかです。
**3 番目**: モジュールの深い思考は、スキル自体よりも価値があります。 `/improve-codebase-architecture` をインストールしなくても、SKILL.md 内の「**削除テスト**」を読むだけで (このモジュールを削除した後に複雑さが消えた場合は、パススルーであることを意味します)、PR レビュー中にすでに再確認することになります。
**コストに注意してください**:
* 5 つのスキルをインストールすると、AI はさらに質問をします。 「1文で500行を生成する」ことに慣れている人にとっては面倒に感じるでしょう
* `/tdd` が厳密に実装された後は、単純なスクリプトでも最初にテストを記述する必要がありますが、これは探索的なコードには適していません。「今回は TDD をスキップする」と指定できます。
* `/grill-with-docs` が率先して CONTEXT.md を変更します。初めて実行する前に、ドライランを行うことをお勧めします。
***
## 参考リソース
**スピーチで引用された 5 冊の本** (出現順):
* *ソフトウェア設計の哲学* — John Ousterhout (複雑さの定義、深いモジュール)
* *The Pragmatic Programmer* — David Thomas & Andrew Hunt(software entropy、outrunning headlights)
* *The Design of Design* — Frederick P. Brooks(design concept、design tree)
* *Domain-Driven Design* — Eric Evans(ubiquitous language)
* *Test-Driven Development* — Kent Beck(invest in design every day)
それぞれ20歳以上です。マットはスピーチの中で何度も「**Amazon に行って、手に入れましょう。**」という文を繰り返しました。この文自体がこのスピーチのイースターエッグです。
# その他のスキル:コミュニケーションの圧縮、引き継ぎ、指導、Skillの記述、安全ガードレール
## 個別に展開しない理由
Matt の README では、スキルは3つのカテゴリに分類されています。
* Engineering:日常的に実際のコードを書くために使用
* Productivity:一般的なワークフローツール
* Misc:彼自身が予備として保持している小さなツール
これまでの記事では、メインラインのエンジニアリングスキルをカバーしてきました。この記事では、残りの「Productivity」と「Misc」をまとめて説明します。なぜなら、それらの多くは完全な開発プロセスではなく、**特定のシナリオで非常に役立つ小さなスイッチ**だからです。
もし5つだけインストールするとしたら、私は依然として以下を優先することをお勧めします。
* [`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me)
* [`/grill-with-docs`](/ja/docs/notes/matt-pocock-skills/grill-with-docs)
* [`/to-prd` + `/to-issues`](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd)
* [`/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture)
しかし、メインラインがすでに稼働している場合は、以下のものが日常体験をよりスムーズにします。
## Productivity Skills
### caveman:コミュニケーションの極限圧縮
`/caveman` は「少トークンモード」です。エージェントに、挨拶、埋め合わせの言葉、過度な説明、曖昧なバッファリングを削除し、技術情報のみを残すように要求します。
これは以下に適しています:
* 高頻度でイテレーションしており、長い応答を読みたくない場合
* デバッグ時に事実、原因、次のステップのみが必要な場合
* 長いコンテキストがほぼ満杯で、出力を圧縮する必要がある場合
* AI にお世辞を言わせたくない場合
これは AI を失礼にするのではなく、より短い構文で完全な技術精度を維持させるものです。終了と言うまで継続的に有効であることに注意してください。
これを長期的なデフォルトではなく、一時的な設定として扱います。高リスク操作、安全警告、複雑な多段階の指示の場合、短すぎると誤読しやすくなります。
### handoff:現在のセッションを次のエージェントに引き継ぐ
`/handoff` の目標は、現在のセッションを要約した引き継ぎドキュメントに圧縮し、現在のワークスペースを汚染するのではなく、システムのテンポラリディレクトリに保存することです。
これには以下が含まれます:
* 現在の目標
* 決定済み事項
* 主要なパスとファイル
* 残りのタスク
* 次のエージェントに呼び出すべきスキルに関する推奨事項
* 機密情報のマスキング
これは、長時間のタスクの中断、コンテキストのほぼ満杯、または別のエージェントに作業を継続させたい場合に特に適しています。
重要な点:PRD、issue、ADR、commit、diff にすでに存在する内容はコピーせず、パスまたは URL のみを引用してください。引き継ぎドキュメントの価値は、会話に散らばっている状態を補完することであり、プロジェクトドキュメントを再作成することではありません。
### teach:現在のディレクトリを学習ワークスペースにする
`/teach` はこのセットの中で最も重いものです。現在のディレクトリを長期学習ワークスペースとして扱い、以下を維持します。
* `MISSION.md`:なぜこのトピックを学ぶのか
* `RESOURCES.md`:高品質なリソースリスト
* `learning-records/*.md`:学習記録、ADR のようなもの
* `lessons/*.html`:各セッションごとのインタラクティブなレッスン
* `reference/*.html`:クイックリファレンス資料
* `NOTES.md`:指導の好みと作業ノート
そのハイライトは、学習を単発の質問ではなく、長期的なシステムとして捉えることです。特に強調されているのは:
* ミッション先行:何を学ぶかよりも、なぜ学ぶかが重要
* リトリーバルプラクティス:想起練習によって長期記憶を構築する
* 間隔反復:短期的な流暢さに騙されない
* 高信頼性リソース:まず資料を探し、モデルの記憶に頼って話さない
単に「X について説明して」と質問するだけなら必要ありません。数週間かけて一つのトピックを学びたい場合は、これが非常に適しています。
### write-a-skill:新しいスキルのためのスケルトンを作成する
`/write-a-skill` は、スキル構造自体の Matt による抽象化です。
スキルは少なくとも以下を持つ必要があります。
```text
skill-name/
├── SKILL.md
├── REFERENCE.md
├── EXAMPLES.md
└── scripts/
```
もちろん、後半の3つは必須ではありません。内容が長すぎる場合、例に価値がある場合、または操作がスクリプト化可能な場合にのみ追加します。
最も重要な判断は、`description` がエージェントがスキルをロードするかどうかを決定する際に最初に目にする唯一の情報であるということです。したがって、description は「ドキュメントの処理を支援する」のような空虚な言葉で書くべきではなく、以下を説明する必要があります。
* 提供する能力
* いつトリガーされるか
* トリガーワードまたはコンテキストは何か
これは私が自分でスキルを書く経験とも一致しています。多くのスキルが機能しないのは、本文の書き方が悪いからではなく、description が広すぎるため、エージェントがそれをロードすべきかどうかを知らないからです。
## Misc Skills
### git-guardrails-claude-code:危険な git コマンドをブロックする
このスキルは、Claude Code に `PreToolUse` フックを装備し、Bash を実行する前に危険な git コマンドをインターセプトします。
デフォルトでブロックされるのは以下のコマンドです。
* `git push`
* `git reset --hard`
* `git clean -f` / `git clean -fd`
* `git branch -D`
* `git checkout .` / `git restore .`
その価値は非常に直接的です。エージェントがあなたの許可なしにプッシュしたり、ハードリセットしたり、追跡されていないファイルを削除したりするのを防ぎます。
AI に実際のレポジトリで作業させる頻繁な場合は、このスキルをインストールする価値があります。AI を信頼しないのではなく、高破壊的な操作をツールレベルでインターセプトし、プロンプトで祈るのではなく、それを行います。
### setup-pre-commit:プロジェクトにコミット前チェックを追加する
`/setup-pre-commit` は以下を設定します。
* Husky pre-commit hook
* lint-staged + Prettier
* typecheck
* test
まずパッケージマネージャーを検出し、次にプロジェクトに既存のスクリプトに基づいて pre-commit で実行すべきものを決定します。`typecheck` または `test` がない場合、強制的に作成せず、省略して通知します。
このスキルの価値は設定自体ではなく、Matt の品質観にあります。**AI にコードに問題がないと言わせるだけでなく、確定的なチェックを通過させるべきです**。
### migrate-to-shoehorn:テストで `as` を少なく書く
これは非常に Total TypeScript スタイルの小さなツールです。テスト内の TypeScript の `as` 型アサーションを `@total-typescript/shoehorn` に移行します。
典型的な置換:
| 旧記法 | 新記法 | シナリオ |
| --------------------------- | ------------------ | ---------------------------------- |
| `obj as Request` | `fromPartial(obj)` | テストで大きなオブジェクトのいくつかのフィールドのみに関心がある場合 |
| `obj as unknown as Request` | `fromAny(obj)` | 意図的に間違った型を渡してエラーパスをテストする場合 |
| 完全なオブジェクトの偽データ | `fromExact(obj)` | 完全な形状を強制する必要がある場合 |
これはテストコード専用であり、本番コードには使用されないことが明記されています。
このスキルは非常に狭いですが、Matt のエンジニアリングのセンスに合致しています。型システムのためにテストで20個の無意味なフィールドを作成したり、生の `as` を使用して型安全を完全に無効にしたりしないことです。
### scaffold-exercises:コースリポジトリ用の演習ディレクトリを生成する
このスキルは明らかに Matt 自身のコース作成ワークフローから来ています。仕様に従って以下を作成します。
```text
exercises/
└── 05-memory-skill-building/
└── 05.02-short-term-memory/
├── explainer/
├── problem/
└── solution/
```
各サブディレクトリには少なくとも空でない `readme.md` があり、必要に応じて `main.ts` があり、`pnpm ai-hero-cli internal lint` をパスする必要があります。
これはほとんどのエンジニアリングプロジェクトには役立ちませんが、コース、ブートキャンプ、練習リポジトリには非常に実用的です。さらに重要なのは、優れたスキルの特徴を示していることです。**繰り返し、機械的で、詳細を見落としやすいフォーマット作業をエージェントに任せる**ことです。
## 現時点でメインラインに含めるべきではないディレクトリ
アップストリームリポジトリには、`deprecated/`、`in-progress/`、`personal/` もあります。
これらを正式な使用ガイドとして一時的に記述しないことをお勧めします。
| ディレクトリ | メインラインに含めない理由 |
| -------------- | ----------------------------------------- |
| `deprecated/` | 廃止されており、読者が古いプロセスを採用し続ける誤解を招きやすい |
| `in-progress/` | まだ実験段階であり、動作や命名が変更される可能性がある |
| `personal/` | Matt 自身の個人的なワークスペースに近く、一般的な読者には適さない可能性がある |
将来的に記述する場合は、「Matt Pocock skills リポジトリの考古学」のような単独の記事として作成できます。安定した推奨事項に混在させるのではなく。
## このツールのセットの共通点
これらのスキルは散らばっているように見えますが、その背後には同じ原則があります。
> エージェントが漂流しやすいことを、小さく明確な作業パターンに変える。
* `caveman` はコミュニケーションの漂流を防ぐ
* `handoff` はコンテキストの損失を防ぐ
* `teach` は学習が単発の質問になるのを防ぐ
* `write-a-skill` はスキルの構造が手書きになるのを防ぐ
* `git-guardrails` は危険なコマンドを自己判断に頼るのを防ぐ
* `setup-pre-commit` は品質チェックを AI の自己申告に頼るのを防ぐ
* `migrate-to-shoehorn` はテストの型アサーションの制御不能になるのを防ぐ
* `scaffold-exercises` はコース構造の手動での項目漏れを防ぐ
これも Matt のこのセットで最も学ぶ価値のある点です。スキルは壮大である必要はありません。頻繁に発生する小さな偏差が、20行の指示で安定して修正できるなら、スキルとして記述する価値があります。
## 参考資料
# プロトタイプ:破棄可能なコードで設計上の疑問に答える
## プロトタイプは「とりあえず書いてみる」ではない
`/prototype` の最初の定義は重要です。
> プロトタイプは、質問に答えるために書かれた、破棄可能なコードです。
この一文は、プロトタイプと「手抜き実装」を切り離します。プロトタイプは本番コードの前身ではなく、後で正式版に修正していく途中段階のものでもありません。最初から破棄可能(throwaway)であるべきだと明記されています。
したがって、`/prototype` の鍵は速く書くことではなく、まず明確にすることです。
> このプロトタイプは何を明らかにしたいのか?
## 2つのブランチ
Matt はプロトタイプを 2 つのタイプに分け、出力は全く異なります。
| 明らかにしたいこと | ブランチ | 出力 |
| --------------------- | --------------- | -------------------------- |
| 論理、ステートマシン、データモデルの妥当性 | Logic prototype | 実行可能なターミナルミニプログラム |
| このインターフェースはどうあるべきか | UI prototype | ルート内に切り替え可能な複数の UI ソリューション |
これは非常に実用的です。多くのチームが「プロトタイプを作ろう」と言いますが、インタラクションの外観を検証したいのか、ステートフローを検証したいのかを明確にしません。両者は全く異なるものを必要とします。
## Logic prototype:ステートをターミナルに展開する
質問が「このステートマシンは正しいか」「このビジネスルールは実行可能か」である場合、プロトタイプは非常に小さなコマンドラインプログラムであるべきです。
その特徴:
* メモリ上で動作し、実際のデータベースに依存しない
* 1 つのコマンドで起動
* 各操作後に、関連する完全なステートを出力する
* 紙の上では推演しにくい分岐を網羅する
* テストは書かない、例外処理はしない、フレームワークに抽象化しない
例:サブスクリプションのステートフローを設計したい場合。
本番コードを直接変更しないでください。まず `subscription-prototype.ts` を書き、ユーザーがターミナルで以下を選択できるようにします。
```text
1. start trial
2. pay
3. cancel
4. expire
5. refund
6. print state
```
各ステップで、現在のエンタイトルメント、トライアルクォータ、支払い済みステート、次の更新日などを表示します。これにより、一部のステートの組み合わせを全く考えていなかったことにすぐに気づくでしょう。
このタイプのプロトタイプの価値は、「抽象的なルールを実行可能なオブジェクトにする」ことです。
## UI prototype:ルート内に複数の大胆なソリューションを配置する
質問が「インターフェースのデザインはどうあるべきか」である場合、プロトタイプは、同じソリューションを 3 回微調整するのではなく、十分に異なる複数の UI セットを生成すべきです。
`/prototype` の UI ブランチの要件:
* 1 つのルート内に複数のバリエーションを配置する
* URL の検索パラメータまたは下部のフローティング切り替えバーで切り替える
* ソリューション間には明確な違いを持たせる
* プロトタイプコードは将来の本番ページに近いものにするが、命名はプロトタイプであることを明確に示す
* 早すぎる段階で実際のデータや永続化に接続しない
これは通常の AI による UI 生成との違いです。「最適なソリューション」を一度に AI に求めるのではなく、実際のブラウザでいくつかの方向性を比較できるようにします。
例えば、ダッシュボードの空の状態について、AI に文言だけを変更させるのではなく、以下のようなものを作成させます。
* A:表形式、高密度、次の操作を強調
* B:タスク指向、左側にチェックリスト + 右側にプレビュー
* C:ガイド付き、主要な CTA と過去の例を強調
そして、チャットで 3 枚のスクリーンショットを見て想像するのではなく、同じルートで切り替えます。
## すべてのプロトタイプは削除可能でなければならない
`/prototype` の一般的なルールの中で最も重要なのは「削除可能」であることです。
* ファイル名またはパスで、これがプロトタイプであることを示す
* デフォルトで本番データベースに接続しない
* 一般的な抽象化をしない
* 過度なエラー処理をしない
* 完了したら削除するか、学んだ結論を正式なコードに吸収する
もしプロトタイプが削除できない場合、それはすでに本番コードの負債になっています。
これは AI プログラミングにおいて特に重要です。AI はプロトタイプを「動作するように見える」ように書くのが得意ですが、人間は削除するのを怠り、最終的に誰も触りたがらない一時的なコードの山がプロジェクトにできてしまいます。
## プロトタイプ完了後に残すべきもの
プロトタイプコード自体は保存する価値はありませんが、その答えは保存する価値があります。
Matt は以下のことを永続的な場所に書き留めることを推奨しています。
* プロトタイプが明らかにしようとした質問
* 観察された結論
* 選択した方向性
* 却下した方向性
* 必要であれば、ADR、issue、PRD、またはコミットメッセージに変換する
つまり、`/prototype` の成果物はコードではなく、**意思決定**です。
## 使用すべきでない場合
`/prototype` を使用すべきでないシナリオ:
* 要件が明確で、実装のみが必要な場合
* バグが再現可能で、[`/diagnose`](/ja/docs/notes/matt-pocock-skills/diagnose-and-triage) を使用すべき場合
* リファクタリングの方向性が明確で、[`/tdd`](/ja/docs/notes/matt-pocock-skills/tdd) で保護して実施すべき場合
* UI の微調整のみで、複数のソリューションを作る価値がない場合
* 削除または吸収する時間が取れない場合
プロトタイプのコストは書くことではなく、後処理にあります。後処理がなければ、始めないでください。
## 便利なプロンプト
以下のように呼び出すことができます。
```text
/prototype
このチェックアウトのステートマシンが妥当かどうか検証したいです。logic ブランチでお願いします。
破棄可能なターミナルプロトタイプのみを作成し、実際の DB には接続しないでください。
各操作後に完全なステートを出力してください。
```
または:
```text
/prototype
プロジェクト詳細ページの 3 つの情報アーキテクチャを比較したいです。UI ブランチでお願いします。
既存のルート体系内の prototype ルートに配置し、下部に切り替えバーを提供してください。
本番コンポーネントは変更しないでください。
```
ここで最も重要なのは、「明らかにしたい質問」を明確にすることです。この質問が明確であれば、プロトタイプが迷走することは少なくなります。
## Grill Me との関係
[`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me) は質問を通じて意思決定を収束させるのに適しており、`/prototype` は試用を通じて意思決定を収束させるのに適しています。
質問で解決できる問題もあります(例:「匿名コメントは承認が必要か」)。触ってみないとわからない問題もあります(例:「このドラッグ&ドロップ並べ替えステートマシンは使いにくいのではないか」)。後者はプロトタイプを使うべきです。
そのため、ワークフローの分岐点に配置します。
```text
アイデアが曖昧
↓
/grill-me
↓
それでも体験や検証が必要な場合
↓
/prototype
↓
結論を保持し、プロトタイプを削除
↓
/to-prd または /tdd
```
## 参考資料
次の記事:[Improve Codebase Architecture:shallow を deep modules にリファクタリングする](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture)。
# Setup Matt Pocock Skills:プロジェクトのルールを明確にする
## この Skill はインストールの問題を解決するものではありません
`/setup-matt-pocock-skills` は、「インストール後に実行する初期化コマンド」と誤解されがちですが、実際には**プロジェクト契約ジェネレーター**のようなものです。後続のスキルに対して、このリポジトリがどのようにタスクを追跡し、issue をどのようにタグ付けし、ドメイン言語やアーキテクチャの決定事項をどこから読み取るかを伝えます。
Matt は README の Quickstart で、インストール時に `/setup-matt-pocock-skills` を選択し、エージェントで実行するように注意喚起しています。その理由は単純です。`to-prd`、`to-issues`、`triage`、`diagnose`、`tdd`、`improve-codebase-architecture`、`zoom-out` はすべて同じプロジェクトコンテキストを必要とします。各スキルが個別に問い合わせると、プロセスが断片的になってしまいます。
これは「Claude の設定」を行うのではなく、3 つのエンジニアリング上の質問に答えるものです。
| 質問 | 何を明確にするか | 後続で誰が使用するか |
| --------------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------------- |
| Issue tracker はどこにあるか | GitHub、GitLab、ローカル Markdown、それとも他のシステムか | `to-prd`、`to-issues`、`triage` |
| Triage タグのマッピング | `needs-triage`、`needs-info`、`ready-for-agent` などの役割が、実際のどのタグに対応するか | `triage` |
| ドメインドキュメントはどこにあるか | 単一の `CONTEXT.md` か、複数の `CONTEXT-MAP.md` + 分割された ADR か | `grill-with-docs`、`diagnose`、`tdd`、`zoom-out`、`improve-codebase-architecture` |
## なぜ重要なのか
この一連のスキルの中核的な考え方は、「小さく、組み合わせ可能」ということです。小さいことの代償は、これらのスキルがプロジェクトプロセス全体を自分で管理しようとしないため、プロジェクト内の実際の規約を知る必要があるということです。
例:`/to-issues` は issue を作成する必要があります。セットアップがないと、以下のいずれを行うべきかを知りません。
* `gh issue create` を呼び出す
* `glab issue create` を呼び出す
* `.scratch//` に書き込む
* それとも、Linear/Jira でコピー可能なテキストを生成する
さらに、`/triage` は issue を `ready-for-agent` に移動する必要があります。リポジトリ内の実際のタグが `ai:ready` で、スキルが新しいタグ `ready-for-agent` を自分で作成した場合、issue tracker はすぐに汚れてしまいます。
したがって、`/setup-matt-pocock-skills` の価値は自動化ではなく、**暗黙の規約を明示化する**ことです。
## どのファイルを読むか
このスキルは、仮定するのではなく、まずリポジトリを探索することから始まります。
* `git remote -v` と `.git/config`:GitHub/GitLab プロジェクトかどうかを判断します。
* ルートディレクトリの `AGENTS.md` / `CLAUDE.md`:`## Agent skills` セクションが既に存在するかどうかを確認します。
* ルートディレクトリの `CONTEXT.md` / `CONTEXT-MAP.md`:ドメイン言語ドキュメントの形式を判断します。
* `docs/adr/` と `src/*/docs/adr/`:ADR がグローバルかモジュールレベルかを判断します。
* `docs/agents/`:セットアップが既に実行されているかを確認します。
* `.scratch/`:ローカル Markdown issue の規約が既に存在するかどうかを判断します。
これは、Matt のワークフロー全体に共通するスタイルです。**まずプロジェクトの実際の状態を確認し、それからルールを記述する**。
## 3 つの決定事項
### 1. Issue tracker
これは後続の作業単位が実装される場所です。
デフォルトでは GitHub が優先されます。なぜなら、このスキルセットは最も初期に GitHub Issues を中心に設計されたからです。しかし、GitLab とローカル Markdown も同等に扱われます。
| 選択 | どのようなシナリオに適しているか |
| -------------- | ----------------------------------------------------- |
| GitHub | オープンソースプロジェクト、GitHub issue workflow が既に存在する場合 |
| GitLab | 会社プロジェクトが GitLab にあり、`glab` に慣れている場合 |
| Local markdown | 個人プロジェクト、一時的な探索、リモートの issue tracker がない場合 |
| Other | Jira、Linear、Feishu、多次元テーブルなど、実際のプロセスをテキストで記録する必要がある場合 |
重要なのはどれを選ぶかではなく、**チームが実際に使用しているもの**を選ぶことです。間違ったものを選択すると、後続のスキルが誤ったシステムでタスクを作成します。
### 2. Triage label vocabulary
`/triage` は内部で 5 つの状態役割を使用します。
| 役割 | 意味 |
| ----------------- | ------------------------ |
| `needs-triage` | 保守担当者の判断待ち |
| `needs-info` | レポーターからの情報補完待ち |
| `ready-for-agent` | AFK エージェントに渡せるほど明確になっている |
| `ready-for-human` | 人間の判断または実装が必要 |
| `wontfix` | 対応しない |
セットアップでは、これらの役割に対応する実際のタグ名を尋ねられます。プロジェクトに既存のタグがない場合は、デフォルト名を使用すれば十分です。独自の命名体系がある場合は、スキルが新しいセットを作成するのではなく、ここでマッピングする必要があります。
### 3. Domain docs
これは Matt のスキルセットと通常のプロンプトの最大の違いです。現在の会話だけでなく、プロジェクト内の**ドメイン言語**と**アーキテクチャの決定事項**も読み取ります。
最もシンプルな形式:
```text
/
├── CONTEXT.md
└── docs/
└── adr/
```
大規模なモノレポでは、複数のコンテキストを使用できます。
```text
/
├── CONTEXT-MAP.md
├── apps/
│ └── web/
│ ├── CONTEXT.md
│ └── docs/adr/
└── services/
└── billing/
├── CONTEXT.md
└── docs/adr/
```
セットアップは、今すぐすべてのドキュメントを完成させることを強制するのではなく、後続のスキルに対して、どこを探すべきか、見つからなかった場合にどのように作成すべきかを伝えます。
## 何を記述するか
最終的に 2 種類の出力があります。
1 つ目は、`AGENTS.md` または `CLAUDE.md` 内の `## Agent skills` セクションです。
```markdown
## Agent skills
### Issue tracker
...
### Triage labels
...
### Domain docs
...
```
2 つ目は、`docs/agents/` 下の 3 つの説明です。
| ファイル | 内容 |
| ------------------------------ | ----------------------------------------- |
| `docs/agents/issue-tracker.md` | issue システム、コマンド、作成/更新の規約 |
| `docs/agents/triage-labels.md` | 標準的な役割から実際のタグへのマッピング |
| `docs/agents/domain.md` | `CONTEXT.md`、`CONTEXT-MAP.md`、ADR の読み取り規則 |
既存の `CLAUDE.md` を優先的に編集し、`CLAUDE.md` がない場合にのみ `AGENTS.md` を考慮することに注意してください。これは、**プロジェクト内に競合する 2 つのエージェントルールエントリポイントを作成しない**という重要な抑制を示しています。
## 推奨される使い方
Matt のスキルセットを初めてインストールする際は、以下の順序で行ってください。
1. インストール:`npx skills@latest add mattpocock/skills`
2. `/setup-matt-pocock-skills` を選択します。
3. `/setup-matt-pocock-skills` を実行します。
4. 実際のプロジェクトの状態に基づいて、issue tracker、タグ、ドメインドキュメントの 3 つの質問に答えます。
5. 生成された `## Agent skills` と `docs/agents/*.md` を確認します。
6. その後、[`/grill-with-docs`](/ja/docs/notes/matt-pocock-skills/grill-with-docs)、[`/to-prd`](/ja/docs/notes/matt-pocock-skills/to-prd-and-issues)、[`/triage`](/ja/docs/notes/matt-pocock-skills/diagnose-and-triage) を使用します。
単独で [`/grill-me`](/ja/docs/notes/matt-pocock-skills/grill-me) を体験したい場合は、セットアップをスキップできます。しかし、エンジニアリングプロセスに入る場合は、まずセットアップを行うことをお勧めします。
## この Skill の設計から得られるインスピレーション
`/setup-matt-pocock-skills` は非常にシンプルに見えますが、エージェントワークフローで最も一般的な問題の 1 つを解決します。それは、**ルールが個人の頭の中に散らばっている**ことです。
多くのチームは、「どのタグを使うか」「どの issue を AI に渡せるか」「CONTEXT.md はどこにあるか」といった情報を口頭での約束として扱います。人間は知っていますが、AI は知りません。AI が知らないと、繰り返し質問するか、さらに悪いことに、自分で推測します。
セットアップの役割は、これらの口頭での約束を読み取り可能なファイルにすることです。後続のスキルはより賢くなる必要はなく、同じプロジェクト契約を安定して読み取るだけで済みます。
これも、私が単独で記事を書く価値があると感じた理由です。これは派手なスキルではありませんが、ワークフロー全体が長期的に機能するための基盤となります。
## 参考資料
次の記事:[Grill Me:AI にコードを書く前に 50 の質問をさせる](/ja/docs/notes/matt-pocock-skills/grill-me)。
# TDD: レッド/グリーン リファクタリングを使用して AI に小さなステップを強制する
## 障害モード: 「AI は正しい動作をしますが、実行できません」
Matt の講演の 3 番目の失敗モード: **方向は正しいが、機能しない**。
最も直接的な解決策は、AI 用のフィードバック インフラストラクチャをインストールすることです。
* TypeScript (静的型付けはありません *クレイジーです*)
* LLM がブラウザにアクセスしてページを単独で表示できるようにする
* 自動テスト
しかし、Matt は、**このフィードバックがインストールされていても、LLM はうまく機能しない**ということを観察しました。一度に 500 行を書いてから、「ああ、タイプチェックをしなければいけない」と考える傾向があります。これは、プラグマティック プログラマーが *ヘッドライトを追い越す* と呼ぶものです。ヘッドライトが照らすよりも速く運転し、壁にぶつかるのは時間の問題です。
> "The rate of feedback is your speed limit, which means you should be testing as you go, taking small deliberate steps. **And the AI by default is really not very good at that.**"
この問題を解決するには、AI をツール レベルで段階的に強制的に停止する必要があります。 Matt の答えは、TDD - **最初にテストするとチェックポイントが強制される可能性がある**です。
## 古典的な理論: ケント・ベックの赤と緑の再構築
TDD の標準リズムは、Kent Beck の 2003 年の著書『テスト駆動開発: 例による』で次のように定義されています。
1. **赤**: 失敗したテストを作成します (何をすべきかを説明します)
2. **緑**: テストに合格するのに十分な最小のコードを作成します。
3. **REFACTOR**: テスト保護下のコード構造を改善する
各ループは非常に短く、数分程度です。すべてのステップで自動チェック (テストの合格/不合格) が行われます。
Matt はこのリズムを直接踏襲していますが、彼の SKILL.md は **アンチパターン** について話すのに多くの時間を費やしています。これが核心です。
## 主要なアンチパターン: 水平方向にスライスされた赤と緑
多くの人は、TDD とは「最初にすべてのテストを作成し、次にすべての実装を作成する」ことを意味すると考えています。 Matt は、SKILL.md のこれは間違っていると直接述べています。
```
WRONG (horizontal slicing):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical slicing via tracer bullets):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
なぜ水平が間違っているのでしょうか? SKILL.mdさんは次の3つの理由を挙げています。
> 1. Tests written in bulk test *imagined* behavior, not *actual* behavior
> 2. You end up testing the *shape* of things (data structures, function signatures) rather than user-facing behavior
> 3. Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
人間の格言: **すべてのテストを一度に書くことは、実際のコードではなく、頭の中にあるものをテストしていることになります**。 impl3 を作成すると、test1 の設計が間違っていることに気づきますが、この時点では、test2/test3/test4 はすべて間違った設計に結合されています。戻って変更を加えます。
正しいアプローチは、1 つの実装をテストし、1 つのペアを作成した後で次のペアを開くことです。各ペアが完了し、この実装から何かを学んだ後、想像力ではなく実際の経験に基づいて次のテストのペアを設計できます。
## スキルの全文構造
`engineering/tdd/SKILL.md` は、TDD 自体に多くのニュアンスがあるため、Matt が書いた最も長いスキルの 1 つです。コア構造は次のとおりです。
### 哲学
> **Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
> **Good tests** are integration-style: they exercise real code paths through public APIs. They describe *what* the system does, not *how* it does it. A good test reads like a specification.
> **Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly). The warning sign: your test breaks when you refactor, but behavior hasn't changed.
診断を思い出してください: **内部関数の名前を変更すると、テストはひざまずきます。その場合、このテストは動作ではなく実装をテストすることになり、これは悪いテストです**。
### ワークフロー (チェックリスト付き)
#### 1. Planning
コードを記述する前にユーザーと調整します。
```
[ ] Confirm with user what interface changes are needed
[ ] Confirm with user which behaviors to test (prioritize)
[ ] Identify opportunities for deep modules (small interface, deep impl)
[ ] Design interfaces for testability
[ ] List the behaviors to test (not implementation steps)
[ ] Get user approval on the plan
```
重要な質問: 「**パブリック インターフェイスはどのようなものであるべきですか?テストするのに最も重要な動作はどれですか?**」
> "**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case."
これは非常に直観に反するものです。デフォルトでは、AI はすべてのエッジ ケースを網羅しようとしますが、Matt は **優先度** を強調し、すべての動作が測定に値するわけではなく、コア パスに火力を集中させます。
#### 2. Tracer Bullet
**1つ**のことを検証する\*\*テストを作成します。
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
これは「曳光弾」です。最初に発射して照準を確認してください。 Matt 氏は、この作業は **エンドツーエンド** であるべきであると強調しました。つまり、最初にスキーマを作成し、次に API を作成し、次に UI を作成するのではなく、スタック全体を通る最も細いパスをカットする必要があると強調しました。
#### 3. Incremental Loop
後続の動作ごとに赤→緑を繰り返します。
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
ルール:
* 一度に 1 つのテスト
* 現在のテストに合格するのに十分なコードのみを記述します
* **将来のテストを予測しないでください**
* テストは観察可能な動作に焦点を当てます
「予測しない」ことが特に重要です。 AI は、「とにかくこの関数は X をサポートする必要があるので、ついでに追加しましょう」と考えずにはいられません。これにより、水平方向のスライスが開始されます。
#### 4. Refactor
すべてのテストに合格したら、リファクタリングの機会を探します。
```
[ ] Extract duplication
[ ] Deepen modules (move complexity behind simple interfaces)
[ ] Apply SOLID principles where natural
[ ] Consider what new code reveals about existing code
[ ] Run tests after each refactor step
```
> **Never refactor while RED.** Get to GREEN first.
赤のリファクタリング = テストとコードを同時に変更 = テストが間違っているのかコードが間違っているのかわかりません。 **最初にグリーン、次にリファクタリング**。
### Per-Cycle Checklist
赤と緑の各サイクルの終わりに、マットは AI に次の自己チェックを依頼します。
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
これら 5 つのポイントは、不適切なテストと過剰実装を特定するために使用されます。 AI セルフチェックにより、最も一般的な間違いを防ぐことができます。
## 実際の使用: 発行から PR まで
`/tdd` は、Matt のワークフローの `/to-issues` からの次のステップです。垂直スライスの問題を考慮すると、プロセスは次のようになります。
```
你: 实现 issue #43
↓
/tdd
↓
Claude 读 issue acceptance criteria
↓
Claude 探索代码库 → 找到 CONTEXT.md → 用项目术语
↓
Planning 阶段:
- 列出准备改的接口
- 列出准备测的行为(按优先级排序)
- 让你点头
↓
Tracer Bullet:
- RED: 写第一个测试(基于 acceptance criteria 第 1 条)
- 跑测试,确认 fail
- GREEN: 写最小实现
- 跑测试,确认 pass
↓
Incremental Loop:
- 每个 acceptance criteria 一个 RED→GREEN
↓
Refactor:
- 看 deep module 提取机会
- 每次重构后跑全套测试
↓
PR
```
赤と緑のサイクルごとに AI が停止し、「テストは失敗しました」/「テストは成功しました。差分は次のとおりです」というステータスが表示されます。 **これらの一時停止は、ヘッドライトを追い越すための解毒剤です** - AI には一度に 1,000 行をレイアウトする可能性はありません。
## モックについて: マットの強い意見
SKILL.md はモックの危険性について特に言及しており、別の `mocking.md` も提供しています。核となるアイデア:
> "Bad tests... mock internal collaborators."
内部協力者をモック化 = テストと実装が 1:1 で結合 = リファクタリング時にテスト チームがひざまずく。 Matt の好みは **統合スタイルのテスト**です。実際のデータベース (メモリ内またはテストコンテナ)、実際の HTTP (MSW)、および実際のファイル システム (tmp dir) を使用してみてください。本当に高価な境界または不安定な境界 (OpenAI API の呼び出しなど) でのみモックを実行してください。
これは多くのチームの現状とは対照的です。ほとんどのコード ライブラリは単体テストでいっぱいで、実際のコードよりもモックの方が多いのです。 Matt はスピーチで次のように判断しました: **優れたコード ベース = コード ベースのテストが簡単**。テストするために多数のモックを作成する必要がある場合は、コード構造に問題があることを意味するため、最初にアーキテクチャを変更する必要があります ([`/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) に進みます)。
## AI時代におけるTDDの新たな意義
ケント・ベックが 23 年前にその本を書いたとき、TDD の中心的な利点は「人々が間違ったコードを書かなくなること」でした。 AI 時代において、TDD には次のような追加の意味があります。
**これは、AI が理解できる唯一の「成功基準」です**。
マットはスピーチの後半でカルパシーの知恵の言葉を引用しました。
「成功基準」の最良の形式は **テスト** です。これは機械検証可能でバイナリであり、議論の余地はありません。 AI にテスト スイートを与えて「合格させてください」とすることは、AI に要件の説明を与えて「実装してください」と与えるよりも 10 倍信頼性が高くなります。
したがって、`/tdd` は単なる品質保証ツールではなく、エージェント ループへの入力インターフェイスでもあります。赤と緑の各サイクルは、完全な「インプット→アクション→フィードバック」です。 AI はサイクル内でこの実装の実際の状況を学習し、次のサイクルではより正確になります。
## インストールと使用方法
```bash
npx skills@latest add mattpocock/skills
```
`tdd` + `setup-matt-pocock-skills` を確認してください。
現在 Codex を主に使用している場合は、インストール製品を `.agents/skills/` に引き継ぎ、プロジェクト レベルのワークフロー、テスト コマンド、発行トラッカー ルールを `AGENTS.md` に書き込みます。 Matt の `/tdd` の本質は赤と緑のリファクタリング サイクルであり、クロード コードとは関係ありません。
**呼び出し方法**:
* 直接: `/tdd` - 現在の会話コンテキストから何を測定するかを推測させます
* 問題を取得: `/tdd implement #43` - 問題を取得して開きます。
* バグを修正: `/tdd reproduce this bug then fix it` - まずバグを再現できる失敗したテストを作成し、それから修正します。
## 注意事項
**すべてのタスクに適しているわけではありません**。 1 回限りのスクリプト、プレイグラウンド探索コード、UI の微調整 - TDD は使用しないでください。ペースが遅くなります。 Matt 自身も、TDD は「永続的な価値があり、保守が必要な」コードに適していると述べています。
**最初にテスト インフラストラクチャを準備します**。プロジェクトにテスト フレームワーク (Vitest / Jest / Playwright など) がインストールされていない場合は、最初にテスト フレームワークをインストールしてから `/tdd` を使用します。それ以外の場合は、最初にテスト フレームワークがインストールされますが、そのステップには多くの質問があります。
**e2e テストを自動的に追加させないでください**。 e2e はゆっくりで歯切れがよく、TDD のリズムは分単位です。 `/tdd` のデフォルトは e2e ではなく統合テストですが、「ユニット + 統合のみ、e2e ではない」と明示的に指定できます。
**再建段階は最も制御不能になりやすい段階です**。 AI が GREEN ステータスを取得すると、熱心にさまざまなものをリファクタリングします。それを見つめ、各リファクタリングの後にテストを実行します。この部分はAI逸脱のリスクが高い領域です。
## 参考リソース
次の記事: [コードベース アーキテクチャの改善: 浅いモジュールを深いモジュールに再構築する](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) - 定期的なメンテナンスにより、コード ベースで AI を長期的に実行できるようになります。
# to-PRD + to-Issues: グリルからの会話を実行可能な垂直スライスに凝縮します。
## ワークフロー内のこのセクションの位置
Matt のワークフロー図に戻ります。
```
/grill-me 或 /grill-with-docs ← 谈清楚
↓
/to-prd ← 凝固成 PRD(你在这里)
↓
/to-issues ← 切成可领取的 vertical slice(你在这里)
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture
```
`/to-prd` と `/to-issues` は、抽象的な対話による決定\*\* を実行可能な作業単位\*\* に変換するという、前と次の間のリンクです。マットの実際の使用では、これらは 2 つの連続したステップであるため、これらを組み合わせます。
## 失敗モード: グリルした後は何もしない
多くの人が `/grill-me` を使用した後に行き詰まってしまいます。多くの決定事項を伴う長い会話になりますが、**コードの作成をどのように開始すればよい**のでしょうか?
クロードに直接「やってください」と言うのは間違いです。理由は次のとおりです。
1. 完全な機能を一度に実装 = AI 出力 1000 行以上 = レビューが難しく、テストが難しく、バグを見つけるのが難しい
2. AI が記憶を失うと、次のセッションにはコンテキストがなくなります。
3. 追跡がない - どこに到達したか、どれだけ残っているかを知る方法がありません
正しいアプローチは、決定を成果物 (PRD) に凍結し、次に PRD を独立して完了できるほど小さい作業パッケージ (課題) に分割することです。これは 30 年間ソフトウェア エンジニアリングの常識でしたが、AI 時代には新たな意味を持ちます。
> AFK エージェント (不在時に実行されるエージェント) が独立して取得して完了できるように、十分に細かく切り刻みます。
## /to-prd: 会話を PRD に圧縮します
### スキルの主な制約
`/to-prd` の SKILL.md には、冒頭に非常に重要な一文が書かれています。
> "This skill takes the current conversation context and codebase understanding and produces a PRD. **Do NOT interview the user — just synthesize what you already know.**"
もう質問しないでください。これは、grill-me ステージが行うことです。to-prd は **合成** のみを行います。したがって、*コンテキストをクリアせずに to-prd を実行してください*\*。前のグリルからのすべてのダイアログに依存します。
### スキル処理の流れ
1. **コード ベースを探索します** (まだ探索していない場合) - プロジェクトの CONTEXT.md ボキャブラリーを使用し、既存の ADR を尊重します
2. **ドラフトモジュール**——インターフェイスを独立してテストできるように、深いモジュールとして抽出できる機会を積極的に探します。
3. **モジュールをユーザーと調整する** - 「これらのモジュールは正しいですか?どれをテストする必要がありますか?」
4. テンプレートに従って **PRD を生成**し、問題トラッカーに送信して、`needs-triage` でタグ付けします。
### PRD テンプレート
Matt が提供したテンプレートは次のとおりです。
```markdown
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list:
1. As a , I want a , so that
2. ...
## Implementation Decisions
- The modules that will be built/modified
- The interfaces of those modules
- Technical clarifications from the developer
- Architectural decisions
- Schema changes / API contracts / Specific interactions
(NO specific file paths or code snippets — they rot fast.)
## Testing Decisions
- What makes a good test (test external behavior, not internals)
- Which modules will be tested
- Prior art (similar tests in the codebase)
## Out of Scope
What's NOT in this PRD.
## Further Notes
```
いくつかの主要なデザイン:
* **ユーザー ストーリーが大部分を占めます**: 要件が長く、番号付きのリスト - 完全な機能ポイントを徹底的に列挙する必要があります。これにより、「わかっていると思っていた」という盲点を避けることができます。
* **実装ではファイル パスやコードは作成されません**: Matt は直接、「それらはすぐに古くなってしまう可能性がある」と言いました。これは LLM 時代特有の考慮事項です。特定のパスは再構築後すぐに廃止されますが、「モジュール境界」と「インターフェイス コントラクト」のライフ サイクルは長くなります。
* **必須範囲外**: この段落はほとんどの PRD テンプレートでは無視されますが、後で問題を切り出す際の境界線の保険となります。
## /to-issues: PRD を垂直方向のスライスに切ります
### 垂直スライスとは何ですか?
これは、Matt の方法論全体の中で最も重要な概念の 1 つです。 SKILL.mdは直接こう言いました。
> Each issue is a thin vertical slice cutting through ALL integration layers end-to-end, NOT a horizontal slice of one layer.
最もわかりやすい例は、「コメント機能」を作成することです。
**水平方向のスライス (間違った方法)**:
* 問題 1: データベース スキーマ
* 問題 2: API エンドポイント
* 問題 3: UI コンポーネント
* 問題 4: テスト
**縦方向のスライス (Tracer Bullet 法)**:
* 問題 1: 「訪問者は匿名のコメントを送信できる」 (スキーマ + API + UI + テストはすべて含まれますが、範囲が非常に狭いため匿名のみにできます)
* 問題 2: 「ログイン ユーザーからのコメントがアカウントに関連付けられている」
* 問題 3: 「コメントに返信できるようになった」
* 問題 4: 「管理者はコメントを削除できる」
水平スライスの問題: 各スライスを個別にテストすることができません。問題 1 が完了した後は、実証するものが何もなく、リンク全体を実行できるようになったのは問題 4 でした。そのときになって初めて、スキーマ設計が間違っていたことがわかりました。
完成した各垂直スライスは、**エンドツーエンドで利用可能な機能サブセット**であり、プラグマティック プログラマーの言葉では **トレーサー ブレット** (トレーサー ブレット) と呼ばれます。最初に 1 ラウンドを撃って照準を確認し、次のラウンドを調整します。
### HITL vs AFK
`/to-issues` は各スライスにもラベルを付けます。
* **HITL** (Human in the Loop) - 人々が意思決定に参加することが求められます。アーキテクチャ上の決定、設計レビューなど
* **AFK** (キーボードから離れた状態) - エージェントは独立して作業を完了でき、あなたは結果を確認するために戻ってくることができます。
> 「可能な限り HITL よりも AFK を優先します。」
これは、Matt のワークフローにおける非常に斬新なアイデアです。問題の切り出しが完了したら、不在時 (夜間や週末など) に実行されるエージェントに問題を直接送信します。翌日戻ってくると、PR はすでにそこに横になってレビューを待っています。 HITL 部分は日中は留まり、エージェントと連携して動作します。
### スライス確認リンク
`/to-issues` は、発生してもすぐには問題を生成しません。最初にタイル スキーマを番号付きリストとして表示します。
```
1. Title: 访客提交匿名评论
Type: AFK
Blocked by: None
User stories covered: #1, #2
2. Title: 评论关联到登录账户
Type: AFK
Blocked by: #1
User stories covered: #3
3. Title: 评论审核流程
Type: HITL(需要确认审核 UI 设计)
Blocked by: #1
User stories covered: #4, #5
```
次に、次のように尋ねます。
* 粒度は正しいですか?厚すぎる/薄すぎる?
* 依存関係は正しいですか?
* どれを統合/分割する必要がありますか?
* HITL/AFK マークは正しいですか?
実際に問題トラッカーに送信する前に、依存関係の順 (最初にブロッカー) でうなずくまで繰り返し、後の問題が最初の問題の実際の問題 ID を参照できるようにします。
### 問題テンプレート
```markdown
## Parent
A reference to the parent issue (if any).
## What to build
A concise description. Describe end-to-end behavior, NOT layer-by-layer
implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
## Blocked by
- A reference to the blocking ticket
(or "None - can start immediately")
```
「エンドツーエンドの動作を説明する」という行に注目してください。これは垂直スライスの精神と一致しています。合格基準は合格リストであり、クロードはこれを `/tdd` 段階で 1 つずつテストに変換します。
## 使用方法: 完全なプロセスの例
自分のブログにコメント機能を追加したいとします。完全なプロセス:
```
你: 我想给博客加评论功能
↓
/grill-me → Claude 问 30 个问题(要不要登录?匿名?嵌套?审核?……)
↓
你回答完毕,达成共识
↓
/to-prd → Claude 生成结构化 PRD,提交到 GitHub Issues #42
↓
/to-issues → Claude 提议切成 4 个 vertical slice
让你确认粒度和依赖
你点头
按依赖顺序发布到 GitHub Issues #43~#46
↓
你回家睡觉
↓
夜里 AFK agent 抓 #43(无依赖),跑 /tdd 完成 → 提 PR
你早上 review、merge
↓
agent 抓 #44 / #45 ……
```
プロセス全体で、画面の前に座って細部まで監視する必要はありません。重要な決定はグリルミーの段階で行われます。
## インストールと前提条件
```bash
npx skills@latest add mattpocock/skills
```
`to-prd`、`to-issues`、`setup-matt-pocock-skills` を確認してください。
**最初に `/setup-matt-pocock-skills`** を実行する必要があります。これにより、問題トラッカー (GitHub / GitLab / ローカル マークダウン) とトリアージ ラベルの語彙が AGENTS.md/CLAUDE.md に書き込まれます。そうしないと、to-prd および to-issue は問題の送信先を認識できなくなります。
サポートされている問題トラッカー:
* **GitHub の問題** (デフォルト、`gh` CLI を使用)
* **GitLab の問題** (`glab` CLI を使用)
* **ローカル マークダウン** (`.scratch//` の下にファイルを作成) - 個人プロジェクトまたはリモートのないプロジェクトに適しています
* **その他** (Jira、Linear など) - 散文を使用してワークフローを説明すると、その説明に従ってスキルが呼び出されます。
## よくある質問
\*\*Q: 既製の PRD はすでに存在します。 to-prd をスキップして to-issue に直接移動できますか? \*\*
A: はい。 `/to-issues` は、問題の参照をパラメーターとして受け入れ (「問題 #42 を垂直方向のスライスに分割する」)、問題のコンテンツをフェッチしてスライスします。
\*\*Q: 私のプロジェクトでは課題トラッカーを使用していませんが、使用できますか? \*\*
A: はい。セットアップ中に「ローカル マークダウン」を選択すると、すべての課題が `.scratch//001-foo.md` のようなローカル ファイルになります。
\*\*Q: PRD が長すぎて AI 自体が処理できない場合はどうすればよいですか? \*\*
A: これはスライスの粒度の兆候です。PRD は 1 つの巨大な PRD ではなく、複数の独立した PRD にスライスされる必要があります。グリルの段階でそれを感じるはずです。50 号について話しており、まだ新しい機能を導入している場合は、まず停止して 2 つの PRD に分割し、バッチで実行してください。
\*\*Q: AFK エージェントはどのようにして問題を自動的に検出しますか? \*\*
A: Matt のリポジトリにはこの部分が提供されていないため、独自のエージェントの手配と調整する必要があります (たとえば、GitHub Actions が Claude Code をトリガーして問題を実行します)。最も簡単な方法は、cron を実行して `is:open no:assignee label:agent-ready` を 1 時間ごとにチェックすることです。
## このプロセスの真価
`/to-prd` と `/to-issues` は「自動プロジェクト管理」のように見えますが、Matt がワークフローの中心に置いているのには、より深い理由があります。
**「**完全に小さなこと**とは何か**」を考えざるを得ません。機能を垂直方向のスライスに切り分けなければならないとき、ソフトウェア エンジニアリングで最も難しい作業の 1 つである継ぎ目を見つけることになります。思考そのものは、これら 2 つのスキルよりも価値があります。
そして、垂直スライスのサイズはAIが一度に処理できるサイズです。 **タスクのサイズを AI の能力の上限に合わせます** - これが LLM とのコラボレーションの基本的なリズムです。
## 参考リソース
次の記事: [TDD: 赤と緑の再構築を使用して AI にスモール ステップを実行させる](/ja/docs/notes/matt-pocock-skills/tdd)——問題を切り取った後、実際に AI にスモール ステップを実行させるにはどうすればよいでしょうか?
# Zoom Out:迷ったときにAIにまず地図を描かせる
## 最短だが、非常に有用
`/zoom-out` は、Matt の安定したエンジニアリングスキルセットの中で、おそらく最も短いものです。そのコアとなる指示は、次の一文に要約できます。
> このコードブロックには慣れていません。抽象度を一段階上げ、関連するモジュールと呼び出し元のマップを、プロジェクトのドメイン言語で描いてください。
これはコードを書くためでも、リファクタリングするためでもありません。これは、非常に一般的な状態に対処するために使用されます。**あなたも AI も特定のファイルに深く入り込んでしまったが、そのファイルが存在する理由を忘れ始めている状態**です。
## 解決する失敗パターン
AI によるコーディングは、局所的な最適解に陥りやすいです。
1. ユーザーがファイルを指定する
2. AI がそのファイルを読み込む
3. AI がローカルコードに基づいて意図を推測する
4. 変更後、初めて上位の呼び出し元、ドメインルール、または ADR がその変更をサポートしていないことに気づく
人間も同様です。デバッグに時間を費やすと、特定の関数に集中しすぎて、システム内のその位置を忘れてしまうことがあります。
`/zoom-out` の役割は、このようなトンネルビジョンを断ち切ることです。エージェントに実装を一時停止させ、まず次のように回答させます。
* このコードブロックはどのドメイン概念に属しますか?
* 誰がそれを呼び出しますか?
* それは何を呼び出しますか?
* それの背後にある不変条件は何ですか?
* `CONTEXT.md` の用語とどのように対応しますか?
* 特定の ADR によって制約されていますか?
## /improve-codebase-architecture との違い
`/zoom-out` と [`/improve-codebase-architecture`](/ja/docs/notes/matt-pocock-skills/improve-codebase-architecture) はどちらもシステム全体を考慮しますが、目的は完全に異なります。
| スキル | 目的 | 出力 |
| -------------------------------- | -------------------- | --------------------------- |
| `/zoom-out` | 見慣れないコードの理解を助ける | マップ、呼び出し関係、ドメインの説明 |
| `/improve-codebase-architecture` | アーキテクチャを深化させる機会を見つける | リファクタリング候補、削除テスト、インターフェース設計 |
`/zoom-out` は「この部分について教えてください」という感じです。リファクタリング案を急いで提案したり、直接コードを変更したりするべきではありません。そのタスクは、認知負荷を軽減することです。
## ドメイン用語を強調する理由
このスキルは、プロジェクトのドメイン用語集を使用することを明確に要求しています。理由は次のとおりです。ファイル名だけで説明すると、AI は次のようなものを出力しがちです。
```text
OrderService は OrderRepository を呼び出し、OrderRepository は db client を呼び出します。
```
これは説明しているように聞こえますが、実際には説明していません。より有用なマップは次のようになります。
```text
Checkout flow では、Order Draft はユーザーが支払う前の仮注文です。
Order Finalization は Draft を不変の Order に変換し、Inventory Reservation をトリガーします。
`OrderService.finalize()` はこの変換のシームであり、呼び出し元は主に Payment Callback と Admin Retry からです。
```
2番目の説明は、コードをビジネス言語に戻しています。「誰が誰を呼び出すか」を知るだけでなく、「なぜそれが存在するのか」も知ることができます。
## いつ使用するのが適切か
次のシナリオで `/zoom-out` を積極的に呼び出すことをお勧めします。
* 見慣れないモジュールを引き継ぐ前
* バグを修正するが、関連する呼び出しチェーンがまだ不明確な場合
* AI が生成したコードをレビューしていて、正しいレベルで変更されたかどうか不明な場合
* PRD を書く準備をしていて、モジュールの境界を確認したい場合
* すでに 3 つのファイルを見たが、システム図が形成されていない場合
* `/improve-codebase-architecture` を実行する準備をしているが、候補領域がまだ不明確な場合
これは「始める前の 5 分間」として特に適しています。一部のバグはコードが難しいからではなく、最初からレベルを間違って見ているために発生します。
## 再利用可能な出力形式
元のスキルは非常に短いですが、使用時に AI にこの形式で出力させることをお勧めします。
```markdown
## システムにおけるこのコードブロックの位置
## 主要なドメイン用語
## 主要モジュール
| モジュール | 責任 | 呼び出し元 | 呼び出される対象 |
|---|---|---|---|
## 主要なフロー
## 既知の制約 / ADR
## 確認すべきファイル
```
この形式は通常の解説よりも安定しており、後続の `/to-prd` や `/diagnose` のコンテキストに変換するのにも適しています。
## プランニングモードとして使用しない
`/zoom-out` の危険性は、AI がマップを説明した後、そのまま「このように変更できます」と提案し始めることです。コードを理解したいだけであれば、明確に制限する必要があります。
```text
実装案は提示せず、ファイルを変更せず、構造のみを説明してください。
```
その価値は、意思決定と理解を分離することにあります。理解が不十分な場合に提案を行うと、誤解をより美しく包装するだけになることがよくあります。
## 私の使用に関する推奨事項
`/zoom-out` は、他のスキルと組み合わせて使用するのに適しています。
* `/zoom-out` → `/diagnose`:まずシステムマップを確認し、次にフィードバックループを構築する
* `/zoom-out` → `/grill-with-docs`:まず既存のドメイン言語を理解し、次に新しい要件を厳しく問いただす
* `/zoom-out` → `/to-prd`:まずモジュールの位置を確認し、次に PRD を書く
* `/zoom-out` → `/improve-codebase-architecture`:まずマップを描き、次に浅い/深い問題を見つける
これは完全なプロセスではなく、単なるブレーキです。AI がローカルファイル内で変更を増やし始めたら、まず `zoom out` させることで、通常は後続のやり直し作業を 1 ラウンド節約できます。
## 参考資料
次の記事:[Prototype:破棄可能なコードで設計上の質問に答える](/ja/docs/notes/matt-pocock-skills/prototype)。
# Pi Agent とは
## はじめに
機能だけを見ると、Pi Agent は過小評価されがちです。ターミナルで動作し、ファイルを読み書きし、コマンドを実行し、セッションを保存し、モデルを切り替えることもできます。まるで別の Claude Code や Codex のように聞こえます。
しかし、Pi の真に興味深い点は、何を追加したかではなく、何を削減したかだと私は思います。AI プログラミングツールの最も核となる層、つまりモデル、コンテキスト、ツール、セッション、拡張機能を保持し、ユーザーのワークフローを事前に固定しないように最大限努めています。
したがって、私は Pi を次のように理解したいと考えています。
**Pi Agent は「より完全な」AI プログラミング製品ではなく、より薄く、より透明性の高いコーディングエージェントハーネスです。**
この判断は機能リストよりも重要です。なぜなら、Pi をどのように学ぶべきかを決定するからです。コマンドを覚えることから始めるのではなく、コーディングエージェントが実際にどのような層で構成されているかを理解することから始めるべきです。
## まず Pi を適切な位置に置く
AI プログラミングツールは通常、モデル、ハーネス、エンジニアリング環境の3つの層に大まかに分けられます。しかし、この3つの層だけではまだ抽象的すぎます。Pi が本当に注目すべきは、中間層のハーネスがさらにどのようなモジュールに分割されているかです。
この図には3つの重要なポイントがあります。
第一に、Pi は単なる「チャット UI」ではありません。CLI、インタラクティブ TUI、print/JSON、RPC、SDK は単なる入り口であり、実際にタスクを引き受けるのは `AgentSessionRuntime` と `AgentSession` です。
第二に、Pi はモデルにリクエストを送信する前にリソースの読み込みを行います。`ResourceLoader` は `AGENTS.md`、`CLAUDE.md`、スキル、拡張機能、プロンプトテンプレートなどを整理し、`SystemPrompt Builder` に渡してモデルが実際に参照するコンテキストを構築します。
第三に、モデルがツールを呼び出す際、モデルが直接ファイルシステムを制御するわけではありません。`AgentHarness` と `AgentLoop` がツールの検証、実行、結果の受け取り、次のラウンドへの継続を担当します。Extensions、Tool Registry、SessionManager は、機能の拡張と状態の保存をサポートします。
したがって、Pi は中間層に位置しますが、この「中間」は単なる空虚な言葉ではありません。具体的には、どのコンテキストがモデルに入力されるか、どのツールが呼び出されるか、ツールの結果がどのようにセッションに戻るか、どの機能が拡張機能によって補完されるかを制御します。
これが、Pi を紹介する多くの記事が minimal、transparent、extensible を強調する理由です。これらはすべて同じことを言っています。Pi はエージェントのコア実行層を小さくし、ユーザーが見て変更できるようにしようとしています。
## 作業の流れ
Pi のリクエストは「モデルに一言尋ね、モデルが一言答える」というものではありません。より正確には、分岐を伴う一連のシーケンスです。
ここで最も重要なのは、ステップ4とステップ5の間の往復です。モデルはファイルシステムに直接触れることはなく、`tool_call` を提案するだけです。Pi はこの呼び出しを受け取り、ツール名とパラメータを検証し、存在する可能性のある拡張フックをトリガーし、実際の操作を実行し、`tool_result` をコンテキストに戻します。モデルは新しいコンテキストに基づいて次のステップを判断します。
これがコーディングエージェントと通常のチャットボットの違いです。チャットボットは主にテキスト内でタスクを完了しますが、コーディングエージェントはエンジニアリングシステムにアクセスする必要があるため、ツール、コンテキスト、状態を管理するためのハーネスが必要です。
Pi のデフォルトツールは非常に少ないです。
| ツール | 意味 |
| ------- | --------------- |
| `read` | ファイルを読み込む |
| `edit` | 既存のファイルを変更する |
| `write` | ファイルを作成または上書きする |
| `bash` | シェルコマンドを実行する |
`grep`、`find`、`ls` のような読み取り専用ツールも有効化または制限できます。このツールセットは控えめに見えますが、コードの読み込み、コードの変更、テストの実行、エラーに基づく修正というプログラミングの閉ループをすでに形成しています。
この設計の背後にある問題は、「Pi がもっと多くのことをするかどうか」ではなく、「より多くのものがデフォルトでコアに含まれるべきかどうか」です。Pi の答えは明確です。必ずしもそうではありません。
## なぜ多くの機能を急いで組み込まないのか
多くの AI プログラミング製品は、計画モード、TODO、サブエージェント、MCP、権限ポップアップ、バックグラウンドタスク、ブラウザツールなどを製品に組み込みます。これにより、すぐに使い始めることができますが、代償も伴います。モデルが実際にどのようなコンテキストを受け取ったのかを知るのが難しくなり、製品のワークフローを自分のワークフローに変更することも難しくなります。
Pi のアプローチは逆です。コアを非常に小さく保ち、ワークフローを外部に配置します。
| 何を変更したいか | Pi はどこに任せるか |
| ----------- | ------------------------- |
| プロジェクトルール | `AGENTS.md` / `CLAUDE.md` |
| 特定のタスクメソッド | Skills |
| カスタムツールと UI | Extensions |
| 共有可能な機能セット | Pi Packages |
| モデル選択 | Provider / Model 設定 |
これは「機能が足りない」のではなく、製品のトレードオフです。コアはエージェントループのみを管理し、具体的なワークフローはユーザーとチーム自身が組み合わせることに任せます。
例えば、Pi にはデフォルトで DeepSearch が組み込まれていません。しかし、これは深度検索ができないという意味ではありません。Pi の考え方に沿ったより良い方法は、`deep_search` ツールを登録する拡張機能を作成し、Tavily、Exa、Brave Search、または社内検索を接続し、必要に応じてモデルに呼び出させることです。
これは「検索ボタン」を製品にハードコーディングするのとは異なります。前者はエージェントの機能を拡張しているのに対し、後者は製品がワークフローを決定しています。
## 重要な点:自分のワークフローを拡張する方法
Pi を単に「試す」だけでなく、実際に活用するには、自分のワークフローを拡張することが重要です。
ここで混同しやすいのは、Pi には「プラグイン」という拡張方法しかないわけではないということです。むしろ、プロジェクトルール、タスクメソッド、実際のツール、共有可能なパッケージという4つの層の入り口が提供されています。まず、自分が何を定着させたいのかを判断する必要があります。
| 定着させたいもの | 何を使うか | どのようなシナリオに適しているか |
| ------------- | ------------------------- | ----------------------------------------------------- |
| プロジェクトの習慣と制約 | `AGENTS.md` / `CLAUDE.md` | エージェントにコードの変更方法、実行するチェック、触れてはいけないディレクトリを指示する |
| 再利用可能なメソッドセット | Skill | コードレビュー、記事執筆、リリース、ドキュメント生成、画像処理のような「手順と経験」 |
| 実際の機能 | Extension | ツールの登録、ツール呼び出しのインターセプト、スラッシュコマンドの追加、UI の追加、外部 API の接続 |
| 配布可能な機能セット | Pi Package | 拡張機能、スキル、プロンプトテンプレート、テーマを自分またはチームで再利用できるようにパッケージ化する |
私の理解では、**Skill は作業マニュアル、Extension は実行可能なプラグイン、Package は配布コンテナです。**
例えば、現在のこのブログのワークフローは次のように分解できます。
| ワークフロー要件 | どこに配置するか |
| ------------------------------------------------------------------------ | ----------------------- |
| 「コンテンツは日本語のみで記述し、画像は `BlogImage` を使用し、MDX 変更後は `pnpm types:check` を実行する」 | `AGENTS.md` |
| 「コンセプト記事を書く際は、誤解、定義、メカニズム、例、境界線に従って構成する」 | `article-writing` Skill |
| 「Pi に `deep_search` ツールを追加し、Tavily / Exa / Brave Search を検索できるようにする」 | Extension |
| 「執筆 Skill、DeepSearch Extension、WeChat 公開コマンドを複数のプロジェクトで利用できるようにパッケージ化する」 | Pi Package |
これは単に「プラグインをインストールする」と言うよりも正確です。なぜなら、多くのワークフローはコードを書く必要がなく、良いルールやスキルがあれば十分だからです。しかし、エージェントに外部検索、データベース検索、CI の呼び出し、危険なコマンドのインターセプトなどの新しい機能を追加したい場合は、Extension を書くべきです。
### Extension:真のプラグイン層
Pi の Extension は TypeScript モジュールです。いくつかの種類のことができます。
| 機能 | 例 |
| ------------ | -------------------------------------------------- |
| ツールの登録 | `deep_search`、`query_logs`、`open_issue` |
| コマンドの登録 | `/review`、`/publish`、`/checkpoint` |
| イベントのインターセプト | `bash` で `rm -rf`、`sudo`、`.env` の書き込みを実行する前に確認を求める |
| UI の変更 | TUI でステータス、選択ボックス、確認ボックス、タスクパネルを表示する |
| 状態の保存 | TODO、接続プール、前回の検索結果、タスクフェーズを記録する |
| 外部システムとの接続 | CI、GitHub、ログシステム、社内 API |
Extension はグローバルに配置することも、プロジェクト内に配置することもできます。
```text
~/.pi/agent/extensions/ # グローバル拡張機能、すべてのプロジェクトで利用可能
.pi/extensions/ # プロジェクト拡張機能、現在のプロジェクトでのみ利用可能
```
一時的な拡張機能をテストするには、次のようにします。
```bash
pi -e ./my-extension.ts
```
自動検出ディレクトリに配置した後、Pi で次のように使用できます。
```text
/reload
```
extensions、skills、prompts、context files を再読み込みします。
これが Pi の最も価値のある点だと私は思います。単に「モデルにコードを書いてもらう」だけでなく、モデルのために制御可能な作業環境を設計しているのです。Extension はモデルが呼び出せる機能を決定し、hooks はどの動作をインターセプトするかを決定し、commands は自分のワークフローがどのようにトリガーされるかを決定します。
### Skill:すべてをプラグインとして書かない
ある機能が主に「どうやるか」であり、「実際の API を呼び出す」または「プログラムを実行する」ではない場合、それは Skill として書くのに適しています。
Skill の構造は通常次のとおりです。
```text
my-skill/
SKILL.md
scripts/
templates/
references/
```
Pi は起動時に完全なスキルをすべてコンテキストに詰め込むことはありません。まずスキルの名前と説明を読み込み、タスクが一致したときにモデルに完全な `SKILL.md` を読み込ませます。これは段階的開示と呼ばれます。利点は、複雑な方法論を保存できる一方で、毎回コンテキストを汚染する必要がないことです。
例えば、「良い記事を書く」「公式アカウントに公開する」「ブラウザ QA を行う」などは、より Skill に近いものです。それらの価値は主に手順、判断基準、参考資料にあり、LLM が呼び出し可能なツールを登録する必要があるとは限りません。
### Package:自分のワークフローをパッケージ化する
安定した機能セットができた場合は、それを Pi Package として作成することを検討できます。
Package には次のものが含まれます。
| 内容 | 役割 |
| ---------------- | ----------------------- |
| extensions | 実行可能なプラグイン、ツール、コマンド、フック |
| skills | 作業方法とタスクマニュアル |
| prompt templates | よく使うプロンプトテンプレート |
| themes | TUI テーマ |
インストール方法は次のとおりです。
```bash
pi install npm:@scope/my-pi-package
pi install git:github.com/user/repo@v1
pi install ./relative/path/to/package
pi list
pi remove npm:@scope/my-pi-package
pi update --extensions
```
デフォルトのインストールは個人設定に書き込まれます。チームプロジェクトで共有したい場合は、プロジェクトレベルの設定を使用し、パッケージを `.pi/settings.json` に記録できます。これにより、他の人がプロジェクトに入って Pi を起動したときに、不足しているパッケージが自動的に補完されます。
ただし、ここでも非常に注意が必要です。Package、Extension、Skill はすべてエージェントの動作に影響を与える可能性があります。サードパーティのパッケージはブラウザプラグインのような低権限の装飾ではなく、コードを実行したり、モデルにコマンドを実行させたりする可能性があります。インストールする前にソースコードを確認する必要があります。
したがって、私は Pi の拡張機能を次の順序で学習します。
1. まず `AGENTS.md` を使用してプロジェクトルールを明確に記述します。
2. 次に、繰り返しのメソッドを Skill にします。
3. 実際のツール機能が必要な場合は、Extension を記述します。
4. 複数のプロジェクトで再利用する場合は、最後に Package にします。
このように学習する方が安定しています。いきなりプラグインを書くのではなく、まずワークフローを「ルール、メソッド、ツール、配布」の4つのカテゴリに分解し、次に各カテゴリを Pi のどの層に配置するかを決定します。
## ソースコードから何が見えるか
Pi のソースコードを見たとき、各関数を追うことよりも、いくつかのファイルがそれぞれどの設計層を表しているかを見るのが最も役立ちました。
| ソースコードの場所 | 説明 |
| ---------------------------------------------------- | -------------------------------------------------------- |
| `packages/agent/src/agent-loop.ts` | コアループ:ユーザーメッセージ、モデル応答、ツール呼び出し、ツール結果を連結する |
| `packages/agent/src/harness/agent-harness.ts` | ハーネスの状態:セッション、システムプロンプト、ツール、フック、メッセージキューを管理する |
| `packages/coding-agent/src/core/tools/index.ts` | 組み込みツールセット:デフォルトのコーディングツールは `read`、`bash`、`edit`、`write` |
| `packages/coding-agent/src/core/resource-loader.ts` | リソースローダー:プロジェクトの指示、拡張機能、スキル、プロンプトテンプレート、テーマを読み込む |
| `packages/coding-agent/src/core/system-prompt.ts` | システムプロンプトの構築:ツール説明、プロジェクトコンテキスト、スキル、現在のディレクトリをプロンプトに含める |
| `packages/coding-agent/src/core/extensions/types.ts` | 拡張システム:拡張機能がツール、コマンド、ショートカット、UI、ライフサイクルイベントを登録できるようにする |
これらを合わせると、Pi の心臓部がほぼ完成します。まずコンテキストとツールを組み立て、次にリクエストをモデルに渡します。モデルがツールを呼び出す必要がある場合、Pi がツールを実行します。結果が返されると、ループが続行されます。
したがって、Pi の「極めてシンプル」は空虚な言葉ではありません。ソースコードの構造自体もこの考えを表現しています。エージェントループ、ハーネス、コーディングツール、リソースローダー、拡張システムを分離し、各層が比較的明確です。
## Pi の境界
Pi は非常に自由度が高いですが、自由度が高いからといって安全であるとは限りません。
Pi パッケージと拡張機能はコードを実行できます。スキルもモデルにスクリプトの実行を指示する可能性があります。`bash` は実際のシステムにアクセスできます。公式ドキュメントとセキュリティ分析は、サードパーティのパッケージ、拡張機能、スキルは自分でレビューする必要があることを警告しています。
私は Pi の境界を3つの点として理解しています。
1. **サンドボックスではない**:これを自然な隔離環境として扱わないでください。危険なプロジェクトはコンテナ、一時ディレクトリ、またはクリーンなワークツリーに配置するのが最善です。
2. **権限を判断しない**:Pi の核となる哲学は、多くのポップアップでリスクを管理することではなく、ツール、コンテキスト、拡張機能を自分で制御することです。
3. **エンジニアリングの境界を理解している人に適している**:エージェントにコマンドを実行させるタイミング、読み取り専用ツールのみを与えるタイミング、git チェックポイントを最初に作成するタイミングを知る必要があります。
これも Pi と、より製品化されたエージェントとの違いです。製品化されたツールは、より多くのセキュリティとインタラクションの詳細をパッケージ化してくれますが、Pi はより直接的な制御権を与え、同時に多くの責任をユーザーに返します。
## Pi をどのように学ぶべきか
Pi を学ぶ際、「どのようなコマンドがあるか」から始めることはお勧めしません。コマンドはすぐに調べられますが、本当に学ぶべきは次の質問です。
| 質問 | なぜ重要か |
| ------------------------- | --------------------------------- |
| Pi はどのようにコンテキストを組み立てるか | モデルが実際に何を知っているかを決定する |
| Pi のツールセットがなぜこれほど小さいのか | エージェントの動作が観察可能かどうかを決定する |
| Extension はどのようにツールを登録するか | 自分のワークフローを接続できるかどうかを決定する |
| Skill と Extension の違いは何か | いつ説明を書き、いつコードを書くかを決定する |
| Session はどのように保存され、分岐するか | エンジニアリング探索が回復、レビュー、継続できるかどうかを決定する |
Claude Code や Codex を使用したことがある場合は、Pi を「分解して見る」機会として捉えることができます。同じようにモデルにコードを書かせるのに、なぜあるツールはブラックボックス製品のように見え、あるツールは改造可能なランタイムのように見えるのでしょうか?
この質問は、「Pi が特定のツールを置き換えられるか」という質問よりも問う価値があります。
## 最後に
Pi Agent の最も価値のある点は、AI プログラミングツールの中間層を公開していることです。
それは私たちに、コーディングエージェントの能力はモデルだけでなく、ハーネスの設計からも生まれることを思い出させます。モデルは思考を担当し、ハーネスはモデルが実際のエンジニアリング環境で行動できるようにします。コンテキストがどのように入力され、ツールがどのように出力され、結果がどのように戻され、拡張機能がどのように挿入されるか、これらの詳細がエージェントの信頼性、透明性、制御可能性を共同で決定します。
したがって、私は Pi を単に「Claude Code の代替品」とは見なしません。むしろ、開発者が研究し、改造するのに適したエージェントランタイムのようなものです。これを使って直接コードを書くこともできますし、自分のエージェントワークフローを設計する方法を学ぶこともできます。
次の実践編では、この考え方に沿って続けます。通常のチュートリアルではなく、Pi Extension を使用して `deep_search` ツールを作成し、外部検索機能をエージェントループに接続する方法を見ていきます。
## 参考文献
# Pi Agent 実践ガイド
## クイックレビュー
概念編では、Pi Agent を極めてシンプルな **Agent Harness** と理解しました。これはモデル、ターミナル、ファイルシステム、シェル、セッション、拡張システムを接続しますが、重厚なワークフローを事前に設定することはありません。
そのため、実践編では一般的な「Pi にファイルを変更させる」というケースは避けたいと思います。そのケースは基本的な閉ループを説明できますが、Pi の拡張性を十分に示せません。
Pi にとってより適切な実践ケースは、デフォルトでは持っていないが多くの人が実際に必要としている機能、**DeepSearch** を追加することです。
ここでの DeepSearch は単なるインターネット検索ではなく、研究型のワークフローです。
| フェーズ | 内容 |
| ----------- | ----------------------------------------------- |
| 問題分解 | 曖昧な問題を検索可能な複数のサブ問題に分解する |
| 多段階検索 | 公式ドキュメント、コードリポジトリ、ブログ、ディスカッションフォーラム、論文をそれぞれ検索する |
| 情報源のフィルタリング | 重複排除、低品質な結果の除外、一次情報源の優先 |
| 証拠の整理 | 重要な事実、リンク、時間、バージョン、不確実性を抽出する |
| 総合的な回答 | 結論を提示し、その根拠と制約を説明する |
私の判断では、**DeepSearch は Pi 本体に組み込むべきではなく、プロンプトだけで無理やり実現すべきでもありません。Pi Extension として作成するのがより適切です。**
理由は簡単です。DeepSearch にはネットワークリクエスト、サードパーティの検索 API、情報源のフィルタリング、結果の切り詰め、引用形式、セキュリティ境界が関わってきます。これらはすべてワークフロー機能であり、コーディングエージェントの最小限の核ではありません。
## 設計目標
このケースで実現するのは完璧な研究システムではなく、動作する最小バージョンです。
目標は以下の通りです。
```text
Pi に deep_search ツールを追加する。
これは以下を受け取ります。
- query:ユーザーが調査したい問題
- depth:検索深度
- maxResults:返却する候補資料の最大数
これは以下を出力します。
- 構造化された検索結果
- 各結果のタイトル、URL、要約、関連性
- モデルが使用するための証拠プロンプト
Pi はこれらの証拠を受け取った後、現在のモデルによって最終的な結論を生成します。
```
意図的に「検索」と「総合」を分離します。
| 部分 | 担当者 | 理由 |
| ------------ | -------------------- | ----------------------------------- |
| 検索 API 呼び出し | DeepSearch extension | これは確定的な外部機能です |
| 結果の重複排除と切り詰め | DeepSearch extension | コンテキストがノイズで埋め尽くされるのを防ぎます |
| どの証拠が重要かの判断 | Pi の現在のモデル | 推論とコンテキスト理解が必要です |
| 最終的な回答の作成 | Pi の現在のモデル | ユーザーの質問とプロジェクトのコンテキストを組み合わせる必要があります |
このようにすることで、より安定します。Extension はそれ自体でモデルを呼び出す必要はなく、ネストされたエージェントになる必要もありません。高品質な証拠を提供するだけで、Pi の元のモデルが推論を続けます。
## 準備作業
Pi extension はグローバルディレクトリに置くことも、プロジェクトディレクトリに置くこともできます。ここではまずプロジェクトディレクトリに置くことをお勧めします。
```text
.pi/extensions/deepsearch/
package.json
index.ts
```
プロジェクトローカルの extension の利点は、境界が明確であることです。この DeepSearch 機能は現在のプロジェクトでのみ有効になり、すべての Pi セッションに影響を与えません。
検索サービスは Tavily、Exa、Brave Search、SerpAPI、あるいは独自の検索バックエンドを選択できます。最初のバージョンではサービスプロバイダーにこだわらず、まず `searchWeb()` 関数として抽象化します。
例えば、環境変数に API キーを保存します。
```bash
export TAVILY_API_KEY=tvly-...
```
サードパーティの検索 API を接続したくない場合は、まずローカルのモックデータを使用して extension を動作させることもできます。ツールの登録、パラメータの受け渡し、結果の形式が安定した後で、実際の検索サービスに接続します。
## Step 1: Extension ディレクトリの作成
まずディレクトリを作成します。
```bash
mkdir -p .pi/extensions/deepsearch
```
extension が依存関係を必要とする場合、`package.json` を置くことができます。
```json
{
"name": "pi-deepsearch-extension",
"private": true,
"dependencies": {
"typebox": "*",
"@earendil-works/pi-ai": "*",
"@earendil-works/pi-coding-agent": "*"
},
"pi": {
"extensions": ["./index.ts"]
}
}
```
次に依存関係をインストールします。
```bash
cd .pi/extensions/deepsearch
npm install
```
Pi の extension は TypeScript モジュールであり、手動でコンパイルする必要はありません。この体験はツールの迅速な実験に適しています。
## Step 2: deep\_search ツールの登録
コアファイルは `.pi/extensions/deepsearch/index.ts` です。
最初のバージョンは次のように書けます。
```typescript
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { StringEnum } from "@earendil-works/pi-ai";
import { Type } from "typebox";
type SearchResult = {
title: string;
url: string;
snippet: string;
score?: number;
};
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "deep_search",
label: "DeepSearch",
description: "Search the web for source-backed evidence about a question.",
promptSnippet: "Research a question with web search and return source-backed evidence.",
promptGuidelines: [
"Use deep_search when the user asks for current facts, external sources, comparison, investigation, or source-backed research.",
"After deep_search returns results, synthesize an answer with citations and clearly separate facts, inference, and uncertainty.",
"Do not treat deep_search results as final truth; inspect source quality and mention gaps."
],
parameters: Type.Object({
query: Type.String({
description: "The research question or search query."
}),
depth: Type.Optional(StringEnum(["quick", "normal", "deep"] as const)),
maxResults: Type.Optional(Type.Number({
minimum: 3,
maximum: 10,
default: 6
}))
}),
async execute(_toolCallId, params, signal) {
const depth = params.depth ?? "normal";
const maxResults = params.maxResults ?? 6;
const results = await searchWeb(params.query, depth, maxResults, signal);
return {
content: [
{
type: "text",
text: formatResultsForModel(params.query, results)
}
],
details: {
query: params.query,
depth,
results
}
};
}
});
}
async function searchWeb(
query: string,
depth: "quick" | "normal" | "deep",
maxResults: number,
signal: AbortSignal
): Promise {
const apiKey = process.env.TAVILY_API_KEY;
if (!apiKey) {
throw new Error("Missing TAVILY_API_KEY. Set it before starting pi.");
}
const response = await fetch("https://api.tavily.com/search", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
api_key: apiKey,
query,
search_depth: depth === "quick" ? "basic" : "advanced",
max_results: maxResults,
include_answer: false,
include_raw_content: depth === "deep"
}),
signal
});
if (!response.ok) {
throw new Error(`Search failed: ${response.status} ${response.statusText}`);
}
const data = await response.json() as {
results?: Array<{
title?: string;
url?: string;
content?: string;
score?: number;
}>;
};
return dedupeByUrl((data.results ?? []).map((item) => ({
title: item.title ?? "Untitled",
url: item.url ?? "",
snippet: item.content ?? "",
score: item.score
}))).filter((item) => item.url);
}
function dedupeByUrl(results: SearchResult[]): SearchResult[] {
const seen = new Set();
const deduped: SearchResult[] = [];
for (const result of results) {
const key = normalizeUrl(result.url);
if (seen.has(key)) continue;
seen.add(key);
deduped.push(result);
}
return deduped;
}
function normalizeUrl(url: string): string {
try {
const parsed = new URL(url);
parsed.hash = "";
parsed.searchParams.delete("utm_source");
parsed.searchParams.delete("utm_medium");
parsed.searchParams.delete("utm_campaign");
return parsed.toString();
} catch {
return url;
}
}
function formatResultsForModel(query: string, results: SearchResult[]): string {
if (results.length === 0) {
return `DeepSearch found no results for: ${query}`;
}
const lines = results.map((result, index) => {
return [
`## Source ${index + 1}`,
`Title: ${result.title}`,
`URL: ${result.url}`,
result.score === undefined ? undefined : `Score: ${result.score}`,
`Snippet: ${result.snippet}`
].filter(Boolean).join("\n");
});
return [
`DeepSearch query: ${query}`,
"",
"Use these sources as evidence. Cite URLs when making factual claims.",
"Separate confirmed facts from inference and uncertainty.",
"",
...lines
].join("\n\n");
}
```
このコードは最も重要なことだけを行います。
| コード位置 | 役割 |
| ------------------------- | ---------------------------- |
| `pi.registerTool()` | `deep_search` をモデル呼び出しに公開する |
| `parameters` | ツールが必要とするパラメータをモデルに伝える |
| `promptGuidelines` | いつ使用し、使用後にどのように処理するかをモデルに伝える |
| `searchWeb()` | 実際の検索サービスを呼び出す |
| `dedupeByUrl()` | 重複する URL を削除する |
| `formatResultsForModel()` | 検索結果をモデルが引用しやすい証拠ブロックに整理する |
最初のバージョンでは複雑にしすぎないでください。DeepSearch の本当の難しさは、検索リクエストを書くことではなく、情報源の品質、コンテキストの長さ、引用形式、不確実性を制御することです。
## Step 3: /deepsearch コマンドの追加
ツールはモデルが呼び出すものですが、ユーザーも直接アクセスできる入り口が必要です。
別のコマンドを登録して、ユーザー入力をより明確な研究タスクに書き換えることができます。
```typescript
export default function (pi: ExtensionAPI) {
pi.registerCommand("deepsearch", {
description: "Run a source-backed DeepSearch task",
handler: async (args, ctx) => {
const query = String(args ?? "").trim();
if (!query) {
ctx.ui.notify("Usage: /deepsearch ", "warning");
return;
}
pi.sendUserMessage(
[
"以下の質問について DeepSearch を実行してください。",
"",
`質問:${query}`,
"",
"要件:",
"1. まず deep_search を呼び出す必要があるかどうかを判断してください。",
"2. 問題が複雑な場合は、2〜4つのサブ問題に分解してそれぞれ検索してください。",
"3. 最終的な回答には情報源のリンクを含める必要があります。",
"4. 事実、推論、まだ不確実な部分を区別してください。",
"5. 検索結果をそのまま羅列するのではなく、総合的な判断を示してください。"
].join("\n"),
{ deliverAs: "followUp" }
);
}
});
pi.registerTool({
// deep_search tool definition...
});
}
```
これにより、ユーザーは直接次のように入力できます。
```text
/deepsearch Pi Coding Agent の extension メカニズムはどのような機能に適していますか?
```
`/deepsearch` は直接検索するのではなく、Pi により完全なタスク説明を送信します。モデルは説明に基づいて `deep_search` を呼び出し、その結果に基づいて総合的な回答を完成させます。
私はこの設計の方が好きです。なぜなら、エージェントの判断の余地を残しているからです。検索ツールは証拠の入り口に過ぎず、最終的な回答生成器ではありません。
## Step 4: 起動と検証
プロジェクトローカルの extension を配置したら、プロジェクトのルートディレクトリで Pi を直接起動できます。
```bash
TAVILY_API_KEY=tvly-... pi
```
一時的にテストするだけであれば、明示的に extension を指定することもできます。
```bash
TAVILY_API_KEY=tvly-... pi -e ./.pi/extensions/deepsearch/index.ts
```
Pi に入ったら、まず外部の事実を必要とする質問をします。
```text
/deepsearch Pi Coding Agent の最新バージョンの extension システムはどのような機能をサポートしていますか?
```
許容される出力は、単なる検索結果のリストではなく、以下を含むべきです。
| チェックポイント | 適切なパフォーマンス |
| ----------- | ------------------------------------- |
| ツールが呼び出されたか | `deep_search` が呼び出されたことが確認できる |
| 情報源が明確か | 各重要な事実の後に URL がある |
| 重複排除されているか | 同じページを繰り返し引用しない |
| 判断があるか | 資料を羅列するだけでなく、適用シナリオを要約できる |
| 不確実性があるか | バージョン変更、サードパーティ API、コミュニティ拡張に対して境界を保つ |
結果が単なる「検索結果リスト」である場合、`promptGuidelines` が十分に強力ではないことを示しています。ガイドラインをより明確にすることができます。
```typescript
promptGuidelines: [
"Use deep_search to gather evidence, not to produce the final answer.",
"After deep_search, write a concise research brief with citations.",
"Prefer official documentation, source code, release notes, and primary sources.",
"Mention when sources disagree or when the evidence is incomplete."
]
```
## Step 5: DeepSearch をより研究ツールらしくする
最初のバージョンが動作したら、さらに3種類の機能を追加できます。
### サブ問題の分解
DeepSearch が最も失敗しやすいのは、大きな問題を直接検索 API に投げ込むことです。
例えば:
```text
Pi Agent は Claude Code の代替になりえますか?
```
これは良い検索クエリではありません。少なくとも次のように分解できます。
| サブ問題 | 役割 |
| ---------------------------------- | -------- |
| Pi Agent のコア設計は何ですか | 位置付けを探す |
| Pi Agent はどのようなツールと拡張機能をサポートしていますか | 機能の境界を探す |
| Claude Code のデフォルト機能は何ですか | 比較対象を探す |
| 両者の権限、セキュリティ、拡張性にはどのような違いがありますか | 判断を形成する |
最初のバージョンではモデル自身に分解させることができます。2番目のバージョンでは、`/deepsearch` コマンドでモデルにまずサブ問題をリストアップさせ、次に `deep_search` を個別に呼び出すように強制できます。
### 情報源の品質階層化
DeepSearch の出力は、検索 API のスコアだけで並べ替えるべきではありません。実際に技術記事を書く際には、私は以下の優先順位で情報源を確認します。
| 優先度 | 情報源 |
| --- | ----------------------------- |
| P0 | 公式ドキュメント、ソースコード、リリースノート |
| P1 | 作者のブログ、メンテナーの説明、issue / PR |
| P2 | 高品質なチュートリアル、技術分析 |
| P3 | コミュニティディスカッション、Reddit、X、フォーラム |
Extension は `formatResultsForModel()` 内で情報源の種類を事前にマークできます。
```typescript
function classifySource(url: string): "official" | "source" | "community" | "other" {
const host = new URL(url).hostname;
if (host === "pi.dev") return "official";
if (host === "github.com") return "source";
if (host.includes("reddit.com")) return "community";
return "other";
}
```
これにより、モデルが総合する際に、コミュニティの噂と公式ドキュメントを同じ証拠レベルで扱うことがなくなります。
### コンテキストの切り詰め
検索結果はコンテキストを簡単に汚染します。DeepSearch のツール出力は少なく、かつ洗練されているべきです。
私の提案は次のとおりです。
| 内容 | ツール出力に含めるか |
| --------------- | ------------------------ |
| タイトル | 含める |
| URL | 含める |
| 200-500 字の要約 | 含める |
| ページ全文 | デフォルトでは含めない |
| 元の HTML | 含めない |
| 検索 API の元の JSON | `details` に含めるが、本文には含めない |
全文を読む必要がある場合は、別のツールを作成できます。
```text
fetch_source(url)
```
これにより、DeepSearch の最初のステップでは候補となる情報源を見つけ、2番目のステップでは最も重要な2〜3ページのみを取得します。最初から十数ページのウェブページ全体をモデルに渡すべきではありません。
## よくある質問
### なぜ直接 bash で検索スクリプトを実行しないのですか?
可能です。しかし、extension の方が安定しています。
bash を使用する際の問題は、モデルが毎回コマンド、パラメータ、出力形式、エラー処理を再決定する必要があることです。Extension はこれらの詳細を固定するため、モデルは `deep_search` を呼び出すだけで済みます。
### なぜ要約も extension に書かないのですか?
最初のバージョンではお勧めしません。
extension 自身が別のモデルを呼び出して要約を行う場合、ネストされたモデル呼び出し、コスト計算、コンテキストのずれ、引用責任の問題に直面します。より簡単な方法は、extension が証拠のみを返し、Pi の現在のセッション内のモデルが総合を担当することです。
### この DeepSearch は MCP と見なされますか?
いいえ、違います。これは Pi extension が登録したローカルツールです。
もし成熟した MCP 検索サーバーをすでに持っている場合、Pi の MCP 関連パッケージや extension を通じて接続することもできます。しかし、このケースでは Pi 自体の拡張メカニズムを理解するために、直接 extension を書くことを選択しました。
### セキュリティに関して注意すべきことは何ですか?
少なくとも4つのことに注意してください。
| リスク | 対策 |
| -------------- | ------------------------------------ |
| API キーの漏洩 | 環境変数からのみ読み込み、リポジトリに書き込まない |
| 信頼できないウェブコンテンツ | ウェブコンテンツをシステム命令として扱わず、検証すべき証拠としてのみ扱う |
| 検索結果の汚染 | 公式情報源とソースコードを優先し、コミュニティ結果の重みを下げる |
| コンテキストの爆発 | 結果の数と要約の長さを制限する |
DeepSearch は「検索強化」のように見えますが、本質的には外部ウェブページをエージェントのコンテキストに取り込むことです。外部コンテンツがコンテキストに入ると、プロンプトインジェクションを現実的なリスクとして扱う必要があります。
## まとめ
Pi Agent の最初の実践ケースを **DeepSearch Extension** としたのは、Pi の3つの重要な特徴を同時に示すことができるからです。
* Pi のコアは非常に小さく、すべてのワークフローを内蔵していません。
* 本当に役立つ機能は extension を通じて追加できます。
* Extension は単にコマンドを追加するだけでなく、モデルが外部世界にアクセスする境界を定義することでもあります。
このケースが動作すれば、Pi は単なるローカルコード編集エージェントではなく、制御可能な研究の入り口を持つことになります。外部資料が必要な問題に遭遇した場合、まず検索し、次にフィルタリングし、情報源を付けて回答することができます。
これは、モデルが記憶に基づいて回答するよりも信頼性が高く、毎回検索コマンドを手書きするよりも再利用性が高いです。
## 参考資料
# Ralph Wiggum 徹底解説
## はじめに
退勤前に AI にタスクを割り当て、翌朝起きたら使えるコードが出来上がっている——この夢は、複雑な Agent クラスタや精巧なオーケストレーションシステムが必要に思えます。しかし、2025年最もホットな AI プログラミング技術の核心は、たったこの1行です:
```bash
while :; do cat PROMPT.md | claude ; done
```
タスクを Claude に繰り返し投入する無限ループ。これが **Ralph Wiggum** です。恥ずかしいほどシンプルですが、実際にこれを使って、元々 $50,000 の見積もりだったプロジェクトを $297 で完成させた人がいます。
なぜこんなシンプルな方法が効果的なのでしょうか?Anthropic が公式プラグインをリリースした後、発明者の Geoffrey Huntley が「This isn't it」と言ったのはなぜでしょうか?
## Ralph とは
名前はアニメ『ザ・シンプソンズ』のキャラクターに由来しています。Ralph Wiggum は警察署長の息子で、劇中で最も「純真」な人物です。自分が何をしているのかよく分かっていませんが、決して止まりません。彼の名セリフ「I'm helping!」は、この技術の本質を見事に表しています:**素朴で飽くなき粘り強さ**(Naive and relentless persistence)。
ここで重要な区別があります:**Ralph は方法論であり、ツールではありません**。「アジャイル開発」が特定のソフトウェアではなく方法論であるのと同じように、Ralph は働き方を記述するものです。実装の違いによって効果は大きく異なる可能性があります——この点については後ほど詳しく説明します。
## なぜ Ralph が必要なのか:Context Rot 問題
Ralph がなぜ効果的なのかを理解するには、まずそれが解決する問題を理解する必要があります。
### AI はどのように「愚かになる」のか
Claude で複雑なタスクを処理する際、こんな経験はありませんか?最初は会話がスムーズで、Claude は正確に理解し的確に実行します。しかし会話が長くなるにつれて「鈍く」なっていきます——重要な情報を忘れ、同じ間違いを繰り返し、コード品質が低下し、不可解な「ハルシネーション」まで発生するようになります。
これは AI の頭が悪いからではありません。問題は**コンテキストウィンドウが汚染された**ことにあります。
こんなシナリオを想像してください:Claude にある機能を書かせたところ、最初は失敗しました。「これを修正して」と言うと、試みますがまた失敗します。10回やり取りした後、Claude のコンテキストには9回分の失敗したコード、9組のエラーメッセージ、大量のもはや関係のない議論が詰め込まれています。こうしたノイズの中から重要な情報を見つけ出すことは、ますます困難になります。
### Dumb Zone
Geoffrey Huntley とコミュニティの開発者たちは、「Dumb Zone」と呼ばれる現象を発見しました:
| コンテキストサイズ | パフォーマンス |
| ----------------- | ---------------- |
| 0 - 50k tokens | 最高パフォーマンス |
| 50k - 100k tokens | 良好、わずかに低下 |
| 100k+ tokens | 明らかな劣化、指示を無視し始める |
| 150k+ tokens | 深刻な劣化 |
正確な閾値はありませんが、経験則として:**コンテキストが容量の半分程度に達したら警戒すべきです**。200k tokens の Claude の場合、100k を超えると「愚かになった」AI とやり取りしている可能性があります。
### 蓄積されたコンテキストは負債である
ここに直感に反する洞察があります:蓄積されたコンテキストは資産ではなく、負債です。
私たちは記憶力は良ければ良いほど、保持する情報は多ければ多いほど良いと考えがちです。しかし大規模言語モデルの世界では、この直感は間違っています。会話が長くなるほど、コンテキストに蓄積される「ネガティブな情報」が増えていきます:失敗したコード、もはや関係のない議論、訂正された誤った理解。これらはスペースを占有するだけでなく、AI の「注意力」も分散させます。
## Ralph の動作原理
Context Rot を理解すれば、Ralph の解決策は明確になります:**コンテキストの蓄積が問題なら、蓄積しなければよい**。
Ralph は3つの柱の上に構築されています:
### 1. 新しいセッション
ループの各イテレーションで、**完全に新しい Claude インスタンス**を起動し、完全にクリーンなコンテキストウィンドウを取得します。「会話履歴をクリア」するだけではありません——その方法では蓄積された状態がまだ残っている可能性があります。現在のプロセスを完全に終了し、新しいプロセスを起動するのです。
これにより、各イテレーションの開始時に Claude は最高の状態にあります。以前のエラーに悩まされることも、古い議論に注意を散らされることもありません。
**だからこそループは Claude Code の外部で実行する必要があります**——bash ループが Claude プロセスのライフサイクルを制御できなければなりません。
### 2. ファイルが真実の情報源
毎回新しいコンテキストから始めるなら、AI は以前何をしたかをどうやって知るのでしょうか?答えは:会話履歴ではなく、ファイルシステムを通じてです。
主要なファイル:
* **PRD/spec ファイル** — 目標、機能リスト、成功基準を定義
* **IMPLEMENTATION\_PLAN.md** — タスク分解と進捗管理
* **progress.txt** — 自由形式のログ、各イテレーション終了時に学んだ内容を追記
* **Git 履歴** — コード変更の証拠
各イテレーションの開始時に、Claude はこれらのファイルを読み取って目標と進捗を把握します。混沌とした会話履歴ではなく、丁寧に整理された状態のスナップショットを見ることになります。
### 3. フィードバックループ
クリーンなコンテキストと永続的な状態だけでは十分ではありません。AI がバグのあるコードを書いてコミットしてしまえば、エラーが蓄積されてしまいます。
フィードバックループは自動化された品質ゲートとして機能します:
* **TypeScript 型チェック** — 型の正確性に対する即座のフィードバック
* **ユニットテスト** — 機能が期待通りであることを検証
* **CI/CD** — コードがビルドおよび統合できることを確認
テストが失敗した場合、コードはコミットされず、Claude は失敗メッセージを確認します。次のイテレーションの新しい Claude インスタンスが問題の修正を試みます。
> 完全な品質保証体制の構築については、[私の Claude Code 品質管理フロー](/ja/blog/claude-code-quality-control)で5層防御の実践経験を共有しています:Hooks 自動化、テスト戦略、AI Review、Pre-commit、GitHub 統合。
## Human on the Loop
Geoffrey Huntley は繰り返し、ある概念的な区別を強調しています:
| Human **in** the Loop | Human **on** the Loop |
| --------------------- | ------------------------ |
| つきっきりの世話 | 監督型マネジメント |
| AI は毎ステップであなたの確認を待つ | あなたが目標と境界を設定し、AI が自律的に実行 |
| あなたがワークフローのボトルネック | あなたは時々進捗を確認するだけ |
実際の使用には2つのモードがあります:
* **AFK モード**:退勤前に起動し、帰宅して寝て、翌朝結果を確認
* **Human-in-the-loop モード**:各イテレーション後に一時停止して確認、複雑または不確実なタスクに適しています
## Ralph に適したタスク
Ralph は万能ではありません。そのコアの強みは「成功するまで反復」であり、特定のタイプのタスクに適しています。
### 適したタスク
| シナリオ | 理由 |
| --------------- | --------------------------- |
| 明確な成功基準があるタスク | 完了を自動的に検証できる(テスト合格、型チェック合格) |
| 反復的な改善が必要なタスク | Ralph のコアの強みはまさに絶え間ない試行 |
| グリーンフィールドプロジェクト | 既存コードを壊す心配がない |
| 自動テストがあるプロジェクト | テストがバックプレッシャーとして品質を確保 |
### 適さないタスク
| シナリオ | 理由 |
| --------------- | ----------------------- |
| 人間の判断が必要なデザイン決定 | 「見た目が良いか」は自動検証できない |
| 一回限りの操作 | 反復不要なタスクに Ralph を使うのは無駄 |
| 本番環境のデバッグ | リスクが高すぎて無人運用には不向き |
| 成功基準が不明確なタスク | いつ停止すべきか判断できない |
### 3つの使用パターン
**フル実装モード**
これは Ralph の最も一般的な使い方です:完全な機能やプロジェクトをゼロから構築します。spec ファイルと実装計画を準備し、Ralph にすべてのタスクを自動実行させます。
典型的なシナリオ:
* 新しい REST API の構築
* CLI ツールの開発
* 新しい機能モジュールの実装
実例:ある開発者がこのパターンを使用して、元々 $50,000 の見積もりだったプロジェクトを完成させ、API コストの合計はわずか $297 でした。MVP 開発、テスト作成、コードレビューを含む全プロセスが完全に自動化されました。別の事例では、レガシーコードベースを React v16 から v19 にアップグレードし、Ralph は14時間稼働して、人間の介入は一切不要でした。
**探索モード**
すべてのタスクがコード出力を必要とするわけではありません。理解が必要なこともあります——新しく引き継いだコードベースの理解、複雑なシステムのアーキテクチャの理解、特定のモジュールの動作原理の理解。
典型的なシナリオ:
* 不慣れなプロジェクトを引き継ぎ、素早く全体像を把握する必要がある
* 既存のコードベースのドキュメントを生成する
* システムアーキテクチャを分析し、潜在的な問題を特定する
このモードでは、プロンプトは「機能 X を実装して」ではなく、「このコードベースを読んでアーキテクチャドキュメントを生成して」や「すべての API エンドポイントを見つけてその役割を説明して」になります。Claude は各イテレーションでより深く探索し、段階的により完全な理解を構築していきます。
**ブルートフォーステストモード**
症状は分かっている、期待する正しい動作も分かっている、でも根本原因がどうしても見つからない——そんなバグがあります。このような時は、Ralph に「力ずくで解決」させましょう。
典型的なシナリオ:
* 再現困難な間欠的なバグ
* 原因不明で時々失敗するテスト
* ボトルネックが不明なパフォーマンス問題
目標を設定します:「このバグを修正して、このテストを安定的にパスさせて」。Ralph は効果的な解決策が見つかるまで、さまざまな修正アプローチを試し続けます。この方法は「直し方は分からないけど、直ったかどうかは分かる」という問題に特に適しています。
## 実装方式の選択
Ralph の方法論を理解したら、次に直面するのは実際の問題です:このループをどのように実装するか?
コミュニティでは、エンジニアリングの成熟度が異なる2つの実装が発展しました:
**ミニマル路線**——[snarktank/ralph](/ja/docs/notes/ralph-wiggum/snarktank):数百行の bash スクリプトで、毎回新しいセッション、ループそのものに集中しています。軽量で始めやすく、素早くスタートするのに適しています。
**エンジニアリング路線**——[frankbria/ralph-claude-code](/ja/docs/notes/ralph-wiggum/frankbria):完全なツールチェーン(監視ダッシュボード、サーキットブレーカー、レート制限、セッション有効期限管理)。デフォルトでは `--continue` でセッションを再利用し、`--no-continue` で新しいセッションモードに切り替えることもできます。
| 項目 | ミニマル(snarktank) | エンジニアリング(frankbria) |
| --------- | --------------- | ----------------------- |
| セッションモード | 毎回新規 | デフォルト再利用、新規に切替可能 |
| 監視 | 手動確認 | 内蔵 tmux ダッシュボード |
| 安全機構 | max\_iterations | サーキットブレーカー+レート制限+タイムアウト |
| インストール難易度 | Skill コピー | install.sh+ウィザード |
2つの実装にはそれぞれ長所と短所があり、選択はエンジニアリングツールへのニーズ次第です。詳しい使い方と比較分析は、それぞれの実践記事をご覧ください。
## おわりに
Ralph は私たちに重要な教訓を与えてくれます:時にはシンプルな方法こそ最も効果的だということです。誰もがより複雑なアーキテクチャを追い求めている中、1つの bash ループがゲームのルールを変えました。
もちろん、Ralph はパズルの一部に過ぎません。その力を最大限に発揮するには、良いプロンプト、適切なプロジェクト、正しいフィードバック機構が必要です。原理を理解した上で、ニーズに合った実装方式を選びましょう:
* 長時間の AFK や大量のイテレーションが必要?→ 『[snarktank/ralph 実践ガイド](/ja/docs/notes/ralph-wiggum/snarktank)』
* エンジニアリング的な監視と安全機構が必要?→ 『[frankbria/ralph-claude-code 実践ガイド](/ja/docs/notes/ralph-wiggum/frankbria)』
***
**関連記事**:
* [snarktank/ralph 実践ガイド](/ja/docs/notes/ralph-wiggum/snarktank) — ミニマルな外部ループ、インストールから実践までの完全操作マニュアル
* [frankbria/ralph-claude-code 実践ガイド](/ja/docs/notes/ralph-wiggum/frankbria) — エンジニアリング的実装:監視、サーキットブレーカーと安全機構
* [Claude Subagent 完全ガイド](/ja/docs/notes/claude-subagent) — コンテキストをクリーンに保つもう一つの方法
* [Claude Skills とは](/ja/docs/notes/claude-skills/concept) — Claude の再利用可能なワークブックを探る
* [GSD 徹底解説](/ja/docs/notes/gsd/concept) — Ralph の基盤の上に構築された完全なコンテキストエンジニアリングシステム
* [Claude システムアーキテクチャ全解説](/ja/docs/notes/claude-architecture) — Hooks、Subagent などのコンポーネントの全体アーキテクチャを理解する
# frankbria/ralph-claude-code 実践ガイド
## はじめに
[前回の記事](/ja/docs/notes/ralph-wiggum/concept)では Ralph の方法論を紹介し、[snarktank/ralph](/ja/docs/notes/ralph-wiggum/snarktank) では極めてシンプルな外部ループ実装を紹介しました。今回はもう一つのアプローチ、[frankbria/ralph-claude-code](https://github.com/frankbria/ralph-claude-code) を見ていきます。
snarktank/ralph の哲学が「最小限のコードで最大限のことをする」だとすれば、frankbria の哲学は「**すべてをエンジニアリング化する**」です——インタラクティブな設定ウィザード、リアルタイム監視ダッシュボード、サーキットブレーカー、レート制限、セッション有効期限管理。簡潔さではなく、**制御性**を追求しています。
どちらの実装にも優劣はなく、異なるユースケースに適しています。本記事では frankbria の完全なツールチェーンを紹介します。
## インストールと設定
### グローバルインストール
```bash
# リポジトリをクローン
git clone https://github.com/frankbria/ralph-claude-code.git
cd ralph-claude-code
# グローバルインストール
./install.sh
```
インストール完了後、以下のグローバルコマンドが使用可能になります:
| コマンド | 説明 |
| --------------- | ---------------------- |
| `ralph` | Ralph ループを起動 |
| `ralph-enable` | 既存プロジェクトで Ralph を有効化 |
| `ralph-setup` | 新規プロジェクトを作成し Ralph を設定 |
| `ralph-import` | 既存の PRD/要件ドキュメントをインポート |
| `ralph-monitor` | リアルタイム監視ダッシュボードを起動 |
### プロジェクトの初期化
既存プロジェクトの場合、インタラクティブウィザードを使用します:
```bash
cd your-project
ralph-enable
```
ウィザードがプロジェクトタイプ(Node.js、Python、Go など)とフレームワーク(Next.js、FastAPI など)を自動検出し、対応する設定ファイルを生成します。
新規プロジェクトの場合:
```bash
ralph-setup my-new-project
```
これによりプロジェクトディレクトリの作成、Git の初期化、`.ralph/` 設定ディレクトリの生成が行われます。
### 既存要件のインポート
すでに PRD ドキュメントや要件定義がある場合:
```bash
ralph-import path/to/your-prd.md
```
Ralph がドキュメントを解析し、タスクリストを抽出して、構造化された `fix_plan.md` を生成します。
## .ralph/ ディレクトリ構造
frankbria のメモリと設定は `.ralph/` ディレクトリに集約されています:
```
.ralph/
├── PROMPT.md # プロジェクト目標とコンテキスト
├── fix_plan.md # タスクリスト(prd.json に相当)
├── AGENT.md # ビルド/テストコマンド(自動管理)
├── specs/ # 詳細要件ドキュメント
│ ├── feature-a.md
│ └── feature-b.md
└── sessions/ # セッション永続化データ
├── current.json
└── history/
```
**snarktank/ralph との比較**:
| frankbria | snarktank | 役割 |
| ------------- | ------------------------------------------ | ---------------- |
| `PROMPT.md` | `prd.json` の `projectName` + `description` | プロジェクト目標の定義 |
| `fix_plan.md` | `prd.json` の `userStories` | タスクリストと進捗 |
| `AGENT.md` | `CLAUDE.md` / `AGENTS.md` | ビルドコマンドとプロジェクト規約 |
| `specs/` | `prd.json` の `notes` フィールド | 詳細要件 |
| `sessions/` | なし(毎回新規プロセス) | セッション状態の追跡 |
`AGENT.md` は**自動管理**されることに注意してください。Ralph が実行中に発見したプロジェクト規約に基づいてこのファイルを自動更新します。snarktank/ralph の `progress.txt` に似ていますが、より構造化されています。
## コアコマンド
### 基本実行
```bash
# Ralph ループを起動
ralph
# リアルタイム監視付き
ralph --monitor
# tmux で起動(長時間実行におすすめ)
ralph --live
```
### 監視ダッシュボード
```bash
# 監視を単独で起動
ralph-monitor
```
`ralph-monitor` は tmux ダッシュボードを開き、以下をリアルタイム表示します:
* 現在実行中のタスク
* 完了/未完了のタスク数
* API 呼び出し回数とコスト見積もり
* サーキットブレーカーの状態
* 最近のエラーログ
### よく使うパラメータ
| パラメータ | 説明 | デフォルト値 |
| ----------------- | ------------- | ------ |
| `--resume` | 前回の中断箇所から再開 | - |
| `--calls ` | 最大 API 呼び出し回数 | 100 |
| `--timeout ` | タイムアウト時間(分) | 300 |
| `--monitor` | リアルタイム監視を有効化 | false |
| `--live` | tmux で実行 | false |
```bash
# API 呼び出し50回制限、2時間タイムアウト
ralph --calls 50 --timeout 120
# 前回の中断箇所から再開
ralph --resume
```
## セーフティ機構
frankbria の最大の差別化ポイントは、多層のセーフティ機構です。
### サーキットブレーカー(Circuit Breaker)
サーキットブレーカーは「進捗なし」を検知すると自動的にループを停止し、無意味な API 消費を防ぎます:
**連続無進捗検知**:連続 N 回のイテレーションで新たなタスクが完了しない場合、サーキットブレーカーが発動します。
**同一エラー検知**:同じエラーメッセージが連続で発生する場合、AI が無限ループに陥っていることを示しており、サーキットブレーカーが発動します。
### レート制限
デフォルトで 100 calls/hour に制限されており、予期しない API 請求額の急増を防ぎます。パラメータで調整可能です:
```bash
ralph --calls 200 # 200 calls に引き上げ
```
### 5時間 API 制限の三層検知
Anthropic API には 5 時間スライディングウィンドウの使用制限があります。frankbria には三層の検知機能が組み込まれています:
1. **事前検知**:各 API 呼び出し前に残り枠を推定
2. **レスポンス検知**:API レスポンスの rate limit headers を解析
3. **フォールバック戦略**:制限に近づくと自動的に呼び出し頻度を低下
### セッション有効期限管理
デフォルトのセッション有効期限は 24 時間です。期限超過後はセッションデータを自動クリーンアップし、期限切れのコンテキストが後続の実行に影響するのを防ぎます。
## スマート終了検知
frankbria は単純にすべてのタスク完了後に終了するわけではありません。**ダブルコンディション終了ゲート**を使用しています:
```
終了条件 = completion_indicators >= 2 AND EXIT_SIGNAL: true
```
**completion\_indicators** は AI の出力から検知された完了シグナルの数で、以下を含みます:
* 「すべてのタスクが完了」
* 「これ以上の TODO はなし」
* テストがすべてパス
* fix\_plan.md のすべての項目が done とマークされている
**EXIT\_SIGNAL** は AI が出力で明示的に宣言する終了意図です。
なぜ 2 つの条件が必要なのでしょうか?**早期終了を防ぐ**ためです。単一のシグナルは誤判定の可能性があります。例えば AI が「タスク完了」と言っても、実際には現在の story しか完了していないかもしれません。ダブルコンディションにより、複数の独立したシグナルがすべて完了を確認した場合にのみ本当に終了します。
## snarktank/ralph との比較
| 観点 | snarktank/ralph | frankbria/ralph-claude-code |
| ------------ | ---------------------- | ----------------------------------- |
| **実装方式** | 外部 bash ループ(毎回新規セッション) | 外部 bash ループ(`--continue` でセッション再利用) |
| **セッションモード** | 毎回新規 | デフォルトで再利用(`--no-continue` で新規に切替可能) |
| **コンテキスト** | 毎回新規 | `--continue` でイテレーション間を跨いで蓄積 |
| **インストール** | Skill コピー | install.sh + インタラクティブウィザード |
| **タスク形式** | prd.json | PROMPT.md + fix\_plan.md |
| **監視** | 手動 `cat`/`jq` | 組み込み tmux ダッシュボード |
| **セーフティ機構** | max\_iterations | サーキットブレーカー + レート制限 + タイムアウト |
| **タスクソース** | PRD のみ | beads / GitHub Issues / PRD |
| **適用シーン** | 長期 AFK、大量イテレーション | 短中期イテレーション、監視が必要な場合 |
### コアの違い:エンジニアリング化の程度
どちらも外部 bash ループで新しい Claude プロセスを起動します。コアの違いはセッション管理方式(frankbria は `--no-continue` で新規セッションモードに切替可能)ではなく、**エンジニアリング化の程度**にあります:
* **snarktank**:極めてシンプルなスクリプト、数百行の bash、ループそのものに特化
* **frankbria**:完全にエンジニアリング化されたツールチェーン——監視ダッシュボード、サーキットブレーカー、レート制限、セッション有効期限管理
frankbria はデフォルトで `--continue` を有効にしてセッションを再利用するため、短いタスクに適しています。長いタスクの場合は `--no-continue` に切り替えて新規セッションモードにすることで、snarktank と同じ Context Rot 防護を得られると同時に、frankbria のエンジニアリング化された利点を保持できます。
### セッション再利用を無効にする方法
frankbria では `--continue` を無効にする 3 つの方法があります:
```bash
# 方法1:コマンドラインパラメータ
ralph --no-continue
# 方法2:環境変数
export CLAUDE_USE_CONTINUE=false
# 方法3:.ralphrc 設定
SESSION_CONTINUITY=false
```
無効にすると、frankbria の動作は snarktank と同等(毎回新規セッション)になりますが、すべてのエンジニアリング化ツール(監視、サーキットブレーカー、レート制限など)は保持されます。
## Context Rot の現実的なトレードオフ
セッション再利用方式の選択は、本質的に Context Rot と起動オーバーヘッドのトレードオフです:
**短いタスク(\< 50k tokens)**:セッション再利用の方が有利です。コンテキストがまだ劣化する前であり、最初の数回のイテレーションの記憶が後続で活用できます。毎回新規セッションを作成する起動オーバーヘッドはむしろ無駄になります。
**長いタスク(100k+ tokens)**:新規セッションの方が信頼性が高いです。100k tokens を超えると Context Rot が明らかに加速し、蓄積されたコンテキストは資産から負債に変わります。新規セッションには起動オーバーヘッドがありますが、毎回最良の状態で開始できます。
**実践的なアドバイス**:
| シーン | 推奨 | 理由 |
| ----------------- | ----------------------------------------- | ----------------------------- |
| 5 つ以下の小タスク | frankbria(デフォルトモード) | 起動が速い、コンテキスト再利用可能 |
| 10 以上のタスク、AFK が必要 | snarktank または frankbria + `--no-continue` | Context Rot を回避、より信頼性が高い |
| リアルタイム監視が必要 | frankbria | 組み込みダッシュボード |
| タスク量が不明 | frankbria + `--no-continue` | エンジニアリング化ツール + Context Rot 防護 |
frankbria ユーザーはタスクの規模に応じて柔軟に選択できます:短いタスクにはデフォルトの `--continue` モード、長いタスクには `--no-continue` モードに切り替えます。snarktank と比較した frankbria の利点は、どちらのモードでも完全なエンジニアリング化ツールチェーンが保持されることです。
## まとめ
frankbria/ralph-claude-code は Ralph 方法論のエンジニアリング化された実装路線を代表しています。snarktank の簡潔さを一部犠牲にして、より充実した監視、セーフティ、設定機能を実現しています。
どちらの実装を選ぶかは具体的なニーズ次第です——「より正しい」答えはなく、「より適した」選択があるだけです。
***
**関連記事**:
* [Ralph Wiggum 詳細解説](/ja/docs/notes/ralph-wiggum/concept) — コア原理と方法論
* [snarktank/ralph 実践ガイド](/ja/docs/notes/ralph-wiggum/snarktank) — 極めてシンプルな外部ループ実装
* [GSD 詳細解説](/ja/docs/notes/gsd/concept) — Ralph をベースに構築された完全なコンテキストエンジニアリングシステム
* [Claude システムアーキテクチャ完全解説](/ja/docs/notes/claude-architecture) — Hooks、Subagent などのコンポーネントの全体アーキテクチャを理解する
# ラルフの実践ガイド
## はじめに
[前の記事](/ja/docs/notes/ralph-wiggum/concept) では、無限ループ + 毎回新しいコンテキスト + 真実の唯一の情報源としてのファイルという Ralph の核となる原則を理解しました。これら 3 つの柱は単純に聞こえますが、コンセプトの理解から実際に実行するまで、多くの詳細を検討する必要があります。
この記事では、始めましょう。 [snarktank/ralph](https://github.com/snarktank/ralph) を使用して、インストールから実行までの完全なプロセスを完了する方法を学びます。 snarktank/ralph は、コミュニティ内で最も成熟した Ralph 実装の 1 つです (10,000 つ星以上)。 Claude Code と Amp の 2 つのツールと、PRD 生成、JSON 変換、自動実行のための完全なツール チェーンをサポートしています。
## 前提条件
始める前に、環境が次の要件を満たしていることを確認してください。
| 依存関係 | 説明 |
| ------------------ | ----------------------------------------------------------------- |
| **AI プログラミング ツール** | クロード コード (`npm install -g @anthropic-ai/claude-code`) または Amp CLI |
| **jq** | JSON処理ツール(macOS:`brew install jq`) |
| **Git** | プロジェクトは Git リポジトリである必要があります。 |
```bash
# 检查依赖
claude --version # Claude Code CLI
jq --version # JSON 处理
git --version # Git
```
## インストールと構成
snarktank/ralphは、利用シーンに合わせて選べる多彩な設置方法をご用意しています。
### 方法 1: クロード コードに直接インストールする (推奨)
最も簡単な方法 - クロード コードの会話に GitHub リンクを貼り付け、クロードが自動的にインストールを完了できるようにします。
```
Install this skill for me: https://github.com/snarktank/ralph
```
Claude Code はリポジトリのクローンを自動的に作成し、スキル ファイルを正しい場所にコピーします。 `/prd` および `/ralph` コマンドは、インストール後に使用できるようになります。
### 方法 2: クロード コード マーケットのインストール
マーケットコマンド経由でインストールします。
```bash
# 添加并安装插件
/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace
```
インストール後、`/prd` (PRD の生成) と `/ralph` (JSON への変換) の 2 つのスキルを使用できます。
### 方法 3: 手動スキルのインストール (クロード コード/アンプ)
スキル ファイルを対応するツールのグローバル構成ディレクトリに手動でコピーします。
```bash
# 先克隆仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# Claude Code 用户
cp -r /tmp/ralph/skills/prd ~/.claude/skills/
cp -r /tmp/ralph/skills/ralph ~/.claude/skills/
# Amp 用户
cp -r /tmp/ralph/skills/prd ~/.config/amp/skills/
cp -r /tmp/ralph/skills/ralph ~/.config/amp/skills/
```
`/prd` および `/ralph` コマンドは、インストール後に使用できるようになります。
### 方法 4: プロジェクト レベルのインストール
Ralph スクリプトをプロジェクトに直接コピーします。チーム共有またはカスタム スクリプトが必要なシナリオに最適です。
```bash
# 克隆 Ralph 仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# 复制核心文件到项目
mkdir -p scripts/ralph
cp /tmp/ralph/ralph.sh scripts/ralph/
cp /tmp/ralph/CLAUDE.md scripts/ralph/ # Claude Code 用户
# 或
cp /tmp/ralph/prompt.md scripts/ralph/ # Amp 用户
# 赋予执行权限
chmod +x scripts/ralph/ralph.sh
```
インストールが完了すると、プロジェクトの構造は次のようになります。
```
your-project/
├── scripts/ralph/
│ ├── ralph.sh # 核心循环脚本
│ └── CLAUDE.md # Claude Code 的 Prompt 模板
├── tasks/ # PRD 文件目录(执行时自动创建)
│ └── prd.json # 你的任务定义
└── ...
```
> **提案**: 最初の方法が最も簡単です。GitHub リンクを Claude Code に投げるだけです。インストール プロセスを手動で制御する場合は、方法 2 (マーケット コマンド) または方法 3 (手動コピー) を選択します。スクリプトをチームと共有する必要がある場合、またはスクリプトをカスタマイズする必要がある場合は、方法 4 を選択してください。
***
## コアファイル構造
ラルフの記憶は完全にファイル システムに依存しています。各ファイルの役割を理解することは、Ralph を上手に使用するための前提条件です。
### ralph.sh - ループエンジン
これは Ralph の核心であり、新しい AI インスタンスを常に生成する bash スクリプトです。
```bash
# 基本用法
./scripts/ralph/ralph.sh [max_iterations] # 默认:Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # 使用 Claude Code
```
反復ごとに、ralph.sh は次の手順を実行します。
1. 機能ブランチを作成します (prd.json の `branchName` から)
2. 最も優先度の高い未完のストーリーを選択します (`passes: false`)
3. このストーリーを実装するための **新しい** AI インスタンスを生成します
4. 品質チェックの実行 (タイプチェック、テスト)
5. チェックに合格しました → git commit;チェックに失敗しました → 次の反復に残します
6. prd.json を更新し、ストーリーを `passes: true` としてマークします。
7. 学んだ教訓を progress.txt に追加します。
8. すべてのストーリーが完了するか、最大反復回数に達するまで繰り返します。
デフォルトの反復制限は 10 です。プロジェクトの複雑さに応じて調整します。
```bash
# 简单项目
./scripts/ralph/ralph.sh --tool claude 10
# 复杂项目
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json - タスク定義
これはラルフの「脳」であり、すべてのタスクが定義されます。これはフラットな JSON ファイルです。
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists with valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve MDX component props format"
}
]
}
```
**フィールドの説明**:
| フィールド | 説明 |
| -------------------- | -------------------------------- |
| `projectName` | ログとブランチの命名に使用されるプロジェクト名 |
| `branchName` | Git ブランチ名 - Ralph が自動的に作成します |
| `id` | ストーリーの一意の識別子、推奨される `US-001` 形式 |
| `title` | 短いタイトル |
| `description` | 詳細な説明 - 具体的であればあるほど良い |
| `acceptanceCriteria` | 受け入れ基準のリスト - **これは最も重要なフィールドです** |
| `priority` | 優先順位番号 - 番号が小さいほど、最初に実行されます。 |
| `passes` | 完了したかどうか - Ralph は自動的に更新します |
| `dependsOn` | 依存ストーリー ID リスト |
| `notes` | 追加のヒントとコンテキスト |
### progress.txt - 体験ログ
これがラルフの「長期記憶」です。各反復の後、AI は追加の学習エクスペリエンスを追加します。
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
次の反復のためのクロードの新しいインスタンスは、このファイルを読み取り、すぐに以前のすべてのエクスペリエンスを取得します。これが、Ralph の実行がますます良くなっている理由です。**繰り返しの間に知識は蓄積されますが、コンテキストはクリーンなままです**。
### AGENTS.md - 永続的な知識ベース
progress.txt に加えて、Ralph はプロジェクト内の `AGENTS.md` (または `CLAUDE.md`) ファイルも更新します。 Claude Code と Amp は両方とも、起動時にこれらのファイルを自動的に読み取ります。
progress.txt とは異なり、AGENTS.md はプロジェクト全体にわたる安定した知識を記録します。
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## PRD を書き込む
PRD (製品要件ドキュメント) の品質は、Ralph の実行結果を直接決定します。よく書かれています、ずっと順風満帆です、ラルフ。下手に書くと、ラルフは同じストーリーで何度も失敗することになります。
### スキルを使用して PRD を生成する
snarktank/ralph スキルがインストールされている場合は、対話的に PRD を生成できます。
```bash
# 在 Claude Code 或 Amp 中
/prd I want to add i18n support to the blog, translating all Chinese content to English
```
AI は、いくつかの明確な質問 (どのドキュメントが関係しているか、テクノロジー スタックの制限、品質基準など) を尋ね、構造化された PRD ドキュメントを生成します。
生成後、`/ralph` コマンドを使用して PRD を `prd.json` 形式に変換します。
```bash
/ralph # 转换 PRD 为 prd.json
```
### PRD を手動で書き込む
prd.json を直接記述することもできます。以下は主要な設計原則です。
**原則 1: ストーリーの粒度は中程度である**
各ストーリーは 1 回の反復で完了できるほど小さく、独立して価値を提供できるほど大きい必要があります。
```json
// ❌ 太大:一次迭代完不成
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, forgot password, OAuth, permission management..."
}
// ❌ 太小:没有独立价值
{
"id": "US-001",
"title": "Create email field on User table",
"description": "Add email field to User model"
}
// ✅ 刚好:一次迭代能完成,有独立价值
{
"id": "US-001",
"title": "Implement email/password login",
"description": "Create login API and login page with email/password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form submits successfully",
"All tests pass"
]
}
```
**経験則**: ストーリーには 1 ~ 3 つのファイル変更が含まれ、3 ~ 5 つの受け入れ基準があります。
**原則 2: 受け入れ基準は自動的に検証可能でなければなりません**
ラルフはストーリーが完了したかどうかを判断する必要があるため、受け入れ基準は客観的に評価可能である必要があります。
```json
// ❌ 模糊的标准
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ 可验证的标准
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**原則 3: dependOn を使用して実行順序を制御します**
一部のストーリーには依存関係があります。 `dependsOn` フィールドは、Ralph が正しい順序で実行されることを保証します。
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**原則 4: メモにコンテキストを提供する**
メモフィールドは AI に追加のヒントを提供します。 AI が知らない可能性がある、あなたが知っている情報をここに書きます。
```json
{
"notes": "Project uses fumadocs framework, i18n files follow .en.mdx suffix naming. Reference content/docs/notes/speckit/concept.en.mdx for translation style."
}
```
***
## ラルフループを実行する
PRD の準備ができたら、ループを実行します。
### 実行開始
```bash
# 使用 Claude Code,默认 10 次迭代
./scripts/ralph/ralph.sh --tool claude
# 指定迭代次数
./scripts/ralph/ralph.sh --tool claude 30
# 使用 Amp(默认)
./scripts/ralph/ralph.sh 20
```
### 実行プロセス
開始すると、次のような出力が表示されます。
```
=== Ralph Loop - Iteration 1 ===
Branch: ralph/i18n-translation
Selected story: US-001 - Translate homepage metadata
Spawning fresh Claude instance...
[Claude Code executing...]
Quality check: pnpm types:check ... PASSED
Committing: feat: [US-001] - Translate homepage metadata
Updating prd.json: US-001 passes: true
Appending to progress.txt
=== Ralph Loop - Iteration 2 ===
Selected story: US-002 - Translate blog post hello-world
Spawning fresh Claude instance...
```
各反復はクロードの完全に新しいインスタンスです。 prd.json を読み取ることで何をすべきかを認識し、progress.txt を読み取ることで以前に学習した内容を認識します。
\###完全なシグナル
すべてのストーリーが `passes: true` とマークされると、Ralph は完了シグナルを出力して終了します。
```
All stories completed!
COMPLETE
```
### 監視とデバッグ
Ralph の実行中に、次のコマンドを使用して進行状況を表示できます。
```bash
# 查看每个 story 的完成状态
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 查看经验日志
cat progress.txt
# 查看最近的 git 提交
git log --oneline -10
# 实时跟踪 Ralph 输出
tail -f progress.txt
```
### 自動アーカイブ
新しい機能を開始すると (別の `branchName` を使用して)、Ralph は最後に実行したファイルを `archive/YYYY-MM-DD-feature-name/` ディレクトリに自動的にアーカイブし、作業ディレクトリをクリーンな状態に保ちます。
***
## フィードバック ループと品質ゲート コントロール
ラルフの「自己修正」能力は、フィードバック ループの品質に完全に依存します。フィードバック ループがなければ、Ralph は盲目的にループする単なるスクリプトになります。コードが正しいかどうかを判断できないまま、コードを大量に出力し続けます。
### 品質チェックを構成する
CLAUDE.md(またはprompt.md)でQAコマンドを定義します。
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 品質アクセス制御レベル
| 階層 | ツール | 捕捉された問題 |
| --------- | ---------------- | ----------------------------- |
| 即時フィードバック | TypeScript コンパイラ | 型エラー、構文エラー |
| 機能検証 | 単体テスト | 論理エラー、特殊なケース |
| 統合検証 | ビルドコマンド | 依存関係の問題、構成エラー |
| 実行時検証 | 開発ブラウザスキル | UI レンダリングの問題 (フロントエンド プロジェクト) |
> フロントエンド ストーリーの場合、Ralph は次の受け入れ基準を追加することを推奨しています。「開発ブラウザ スキルを使用してブラウザで検証する」 - AI に実際にブラウザを開いて、ページが正しくレンダリングされていることを確認させます。
### 品質チェックに失敗した場合
ストーリーが繰り返し QA に失敗した場合、Ralph は同じストーリーを無期限に再試行しません。反復上限に達すると停止し、現在の状態が保持されます。次のことができます。
1. progress.txt を表示して、スタックの理由を理解します。
2. 問題を手動で修正して、再度実行します。
3. ストーリーの粒度を調整します (大きすぎる可能性があります)
4. メモフィールドにコンテキストを追加します。
***
## 迅速なカスタマイズ
Ralph のプロンプト テンプレート (CLAUDE.md または prompt.md) は、AI の動作を制御する主な手段です。インストール後、プロジェクトに応じてカスタマイズする必要があります。
### 主要なカスタマイズ項目
**1.プロジェクト固有の品質コマンド**
```markdown
## Project-Specific Commands
- Typecheck: `pnpm types:check` (not `tsc` or `pnpm typecheck`)
- Test: `pnpm vitest run`
- Build: `pnpm build`
- Lint: `pnpm lint`
```
**2.コードスタイルの制約**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**3.既知の落とし穴**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**4.スタック時の対処**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 実際のケース: Ralph によるブログの翻訳
Ralph が実際にどのように機能するかを示すために、実際の例を示します。Ralph スタイルの自律エージェントを使用して、ブログ全体を中国語から英語に翻訳します。
### プロジェクト設定
このプロジェクトでは、22 以上のコンテンツ ファイル (ブログ投稿、ドキュメント、ナビゲーション メタデータ) を中国語から英語に翻訳する必要があります。このプロジェクトは fumadocs の Next.js ブログに基づいており、i18n をサポートしています。このタスクは `prd.json` ファイルで定義されており、それぞれに明確な受け入れ基準を持つ 16 のユーザー ストーリーが含まれています。
```
scripts/ralph/
├── prd.json # 16 个 user story,带验收标准
└── progress.txt # 经验日志,每个 story 完成后更新
```
すべてのユーザー ストーリーは一貫したパターンに従います。
* **明示的な成果物**: "コンテンツ/blog/xxx.en.mdx を作成"
* **検証可能な基準**: 「タイプチェックに合格」、「内部リンクは /en/ プレフィックスを使用」
* **技術的制約**: 「コード ブロックを翻訳しないようにする」、「QuoteCard にdefaultLang='en' を設定する」
### 実行モード
エージェントは、ラルフの方法論の中核原則に従います。
1. **文書は真実の情報源です**: `prd.json` はどのストーリーが通過したかを追跡します (`passes: true/false`)。 `progress.txt` 反復の間に経験を積みます - 例: 「Typecheck コマンドは `pnpm types:check` であり、`pnpm typecheck` ではありません」
2. **自動品質ゲート**: 各翻訳の後、`pnpm types:check` が実行され、MDX ファイルが正しくコンパイルされたかどうかが検証されます。型チェックが失敗した場合は、送信する前に修正してください。
3. **段階的な進歩**: 各ストーリーは、説明的な提出情報 (`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`) とともに個別に提出され、必要に応じて簡単にロールバックできます。
4. **並列実行**: 長い記事の場合、複数のサブエージェントが同時に翻訳されます。たとえば、US-010 (claude-skills コンセプト + プラクティス)、US-011 (speckit コンセプト + プラクティス)、および US-012 (claude-architecture + claude-subagent) が並行して実行されます。
### 重要な教訓
| 経験 | 詳細 |
| ------------------------ | ------------------------------------------------------------------------------------------- |
| **知識の蓄積は重要です** | 初期のストーリーで発見されたパターン (QuoteCard `defaultLang`、リンク プレフィックス ルール) により、後続のストーリーをより速く完了できるようになります。 |
| **フィードバック ループとしての型チェック** | 問題が積み重なる前に、欠落しているインポートや不正な MDX を検出します。 |
| **並列化されスケーラブル** | 6 つの翻訳エージェントが同時に実行され、完了時間は 1 つのエージェントとほぼ同じです。 |
| **PRD の粒度は重要です** | スコープはストーリーごとに 1 ~ 2 ファイル - 確実に完了するのに十分な大きさ、意味のあるのに十分な大きさ |
| **進捗状況ログにより、ミスの繰り返しを防止** | progress.txt の「コードベース パターン」部分は、同じ問題の再発見を防ぐためのナレッジ ベースになります。 |
### 結果
16 のユーザー ストーリーすべてが 1 回のセッションで完了しました。8 つの meta.en.json ナビゲーション ファイルが作成され、3 つのブログ投稿が翻訳され、12 のドキュメント ページが翻訳され、完全なサイトの構築が検証されました。受け入れ基準が明確であり、フィードバック ループ (タイプチェック) が問題を即座に検出するため、各翻訳は一貫した品質を維持します。
このプロジェクトは、Ralph の **完全な実装モデル** - 明確に定義されたタスク + 明確な成功基準 + 自動検証 + ファイル システムを介した増分配信を実証します。
***
## コミュニティの実装と代替案
スナークタンク/ラルフが唯一の選択肢ではありません。これらの実装にはそれぞれ、ニーズに応じて独自の長所と短所があります。
| リソース | リンク | 説明書 |
| -------------- | ------------------------------------------------------------------------------- | ------------------------------------------------- |
| スナークタンク/ラルフ | [スナークタンク/ラルフ](https://github.com/snarktank/ralph) | この記事で使用される、最も完全な関数 |
| ラルフ・オーケストレーター | [マイクヨブリエン/ラルフオーケストレーター](https://github.com/mikeyobrien/ralph-orchestrator) | Mickey O'Brien によって開発され、より多くのカスタマイズ オプションを備えています。 |
| ラルフ・ループ・エージェント | [vercel-labs/ralph-loop-agent](https://github.com/vercel-labs/ralph-loop-agent) | AI SDK に基づく Vercel の実装 |
| ラルフィー | [マイケルシメレス/ラルフィー](https://github.com/michaelshimeles/ralphy) | Michael Shimeles による軽量実装 |
### 代替: GSD
GSD は厳密には Ralph の「コミュニティ実装」ではなく、**代替**です。これは、Ralph の中心原則 (コンテキスト管理、アトミック タスク) を適用していますが、より完全なワークフロー (議論→計画→実行→検証) を提供します。
| リソース | リンク | 説明書 |
| ----------------- | ----------------------------------------------------------------- | -------------------------- |
| GSD (ゲット・スタッフ・ダン) | [キラキラカウボーイ/くそったれ](https://github.com/glittercowboy/get-shit-done) | アイデアから PRD、実行までの完全なフレームワーク |
Ralph が「生」すぎて、より多くのプロセス サポートが必要だと感じる場合は、GSD の方が適している可能性があります。詳細については、[GSD の詳細な分析](/ja/docs/notes/gsd/concept) を参照してください。
***
## 推奨リソース
**公式情報源**:
| リソース | リンク | 説明書 |
| --------------- | ---------------------------------------------------------------------- | ------------- |
| ジェフリー・ハントリーのブログ | [ghuntley.com/ralph](https://ghuntley.com/ralph/) | 発明者によるオリジナル記事 |
| ラルフ・ウィガムのハウツー | [ガントリー/ラルフ・ウィガムのハウツー](https://github.com/ghuntley/how-to-ralph-wiggum) | 公式ユーザーガイド |
**ビデオチュートリアル**:
| リソース | リンク | 説明書 |
| --------------------- | ----------------------------------------------------------------------- | --------------------------------- |
| ラルフ・ウィガムの徹底したディスカッション | [なぜクロードコードの実装ではないのか](https://www.youtube.com/watch?v=O2bBWDoxO4s) | Geoffrey Huntley が正式な実装の問題点を説明します |
| Ralph の正しい使い方 | [ラルフ ウィガム ループの使い方は間違っています](https://www.youtube.com/watch?v=I7azCAgoUHc) | Roman (Mentat) の使用デモ |
| ラルフについて話さなければなりません | [ラルフについて話さなければなりません](https://www.youtube.com/watch?v=Yr9O6KFwbW4) | 論争に対するテオの分析 |
***
## ベスト プラクティスとよくある質問
### コスト管理
Ralph の自動実行は、API 料金が継続的に発生することを意味します。いくつかの管理措置:
* **常に `max_iterations`** を設定します: これは最も基本的なセーフティ ネットです
* **ストーリーの粒度を適度に保ちます**: ストーリーが大きすぎると、複数の反復が必要になります。ストーリーが詳細すぎると、起動時のオーバーヘッドが増加します。
* **最初は小規模なテスト**: 新しいプロジェクトは最初に 3 ~ 5 回繰り返し実行され、プロンプトと品質アクセス制御が適切に機能していることを確認した後に拡張されます。
### よくある落とし穴
**罠 1: ストーリーが大きすぎる**
症状: ストーリーが繰り返し失敗し、反復回数がすぐに使い果たされてしまいます。
解決策: 2 ~ 3 つの小さなストーリーに分割します。 「完全な認証システムの構築」は、「ログイン API の実装」+「ログイン ページの作成」+「JWT ミドルウェアの追加」に分かれています。
**トラップ 2: フィードバック ループがない**
症状: ラルフはストーリーが完了したと主張していますが、実際のコードには何か問題があります。
解決策: 実行可能チェック コマンドを受け入れ基準に追加します。 「コードが書かれている」は合格基準ではなく、「pnpm テストにすべて合格する」が合格基準です。
**トラップ 3: progress.txt は使用されません**
症状: 同じエラーが異なる反復で繰り返し表示されます。
解決策: プロンプト テンプレートで「progress.txt を読み、そこに含まれるルールに従う」ことを明示的に指示していることを確認してください。 AIが自動的に経験値を追加しない場合は、プロンプトに「各ストーリーが完了したら、progress.txtに経験値を追加します」と追加してください。
**トラップ 4: 依存関係の順序が間違っている**
症状: ストーリーがまだ存在しないコードに依存しているため、実装が失敗します。
解決策: `dependsOn` フィールドを正しく設定します。インフラストラクチャの話を必ず最初に置いてください。
### FAQ
\*\*Q: Ralph と公式プラグインの違いは何ですか? \*\*
主な違い: snarktank/ralph は反復ごとに新しいプロセス (実際には完全に新しいコンテキスト) を生成しますが、公式プラグインは同じセッション内でループします (コンテキストは蓄積し続けます)。詳細は[前回の記事の分析](/ja/docs/notes/ralph-wiggum/concept#the-problem-with-the-official-plugin)を参照してください。
\*\*Q: prd.json は実行中に手動で変更できますか? \*\*
はい。 Ralph は、各反復の開始時に prd.json を再読み込みします。反復間でストーリーの説明を変更したり、新しいストーリーを追加したり、ストーリーを手動で `passes: true` としてマークしたり (スキップ) することができます。
\*\*Q: 何度も失敗するストーリーに行き詰まった場合、ラルフはどうすればよいですか? \*\*
1. progress.txt を確認して、失敗の理由を理解します。
2. メモにコンテキストを追加する
3. ストーリーを分割します (大きすぎるかもしれません)
4. ブロックの問題を手動で修正して、再度実行します。
\*\*Q: ラルフの実行中に他のことをすることはできますか? \*\*
はい。ラルフは「人間としての役割を果たす」ように設計されており、じっと見つめる必要はありません。 AFK モードでは、仕事を降りる前に開始し、翌朝結果を確認します。 Ralph が作業しているファイルを変更しないでください。
\*\*Q: コストを管理するにはどうすればよいですか? \*\*
3 つの方法: 適切な `max_iterations` を設定する、ストーリーの粒度を適切に保つ (無駄な反復を減らすため)、最初に小規模な試行を実行してプロセスが正しいことを確認します。一般的に、10 ~ 20 ストーリーのプロジェクトの場合、API 料金は 50 ~ 100 ドル以内です。
***
## 概要
ラルフのワークフローは 5 つのステップに要約できます。
```
安装 → 编写 PRD → 配置质量门禁 → 运行循环 → 检查结果
```
中心となる概念は変わりません。**ドキュメントを唯一の真実の情報源とし、各反復をゼロから開始し、品質ゲートに制御してもらいます**。
次に、プロジェクトに戻り、prd.json を準備し、`./scripts/ralph/ralph.sh --tool claude` を実行して、コーヒーを作りに行きます。
### さらに読む
* [ラルフ ウィガムの詳細な分析](/ja/docs/notes/ralph-wiggum/concept) - ラルフの核となる原則を再考する
* [GSD の詳細な分析](/ja/docs/notes/gsd/concept) - Ralph に基づいて構築された完全なコンテキスト エンジニアリング システム
* [クロードスキルとは](/ja/docs/notes/claude-skills/concept) ——ラルフのPRDスキルはクロードスキルです
* [Speckit 実践ガイド](/ja/docs/notes/speckit/practice) - もう 1 つの構造化された AI プログラミング ワークフロー
# snarktank/ralph 実践ガイド
## はじめに
[前回の記事](/ja/docs/notes/ralph-wiggum/concept)では、Ralph のコア原理——無限ループ + 毎回新しいコンテキスト + ファイルを唯一の真実の情報源とすること——を学びました。3つの柱はシンプルに聞こえますが、理解から実際に動かすまでには多くの細かなポイントがあります。
今回は実際に手を動かしていきます。[snarktank/ralph](https://github.com/snarktank/ralph) は Ralph メソドロジーの**外部ループ実装**です。各イテレーションで新しい Claude プロセスを起動し、Context Rot 問題を根本的に解決します。コミュニティで最も完成度の高い Ralph 実装の一つ(10k+ stars)で、Claude Code と Amp の両プラットフォームに対応し、PRD 生成、JSON 変換、自動実行のフルツールチェーンを提供しています。
> もう一つの実装は [frankbria/ralph-claude-code](/ja/docs/notes/ralph-wiggum/frankbria) で、完全なエンジニアリングツールチェーン(モニタリングダッシュボード、サーキットブレーカー、レート制限)を提供し、制御性と安全メカニズムに重点を置いています。両者の比較はそちらの記事をご覧ください。
## 前提条件
始める前に、以下の環境要件を満たしていることを確認してください:
| 依存関係 | 説明 |
| ----------------- | -------------------------------------------------------------------- |
| **AI プログラミングツール** | Claude Code (`npm install -g @anthropic-ai/claude-code`) または Amp CLI |
| **jq** | JSON 処理ツール(macOS: `brew install jq`) |
| **Git** | プロジェクトが Git リポジトリである必要があります |
```bash
# 依存関係の確認
claude --version # Claude Code CLI
jq --version # JSON 処理
git --version # Git
```
## インストールと設定
最も簡単な方法は、Claude Code の会話で GitHub リンクを直接貼り付けることです:
```
帮我安装这个 skill:https://github.com/snarktank/ralph
```
Claude Code が自動的にリポジトリをクローンし、skill ファイルを正しい場所にコピーします。インストール完了後、`/prd` と `/ralph` コマンドが使用可能になります。
> snarktank/ralph は Marketplace インストール、手動での skill ファイルコピー、プロジェクトレベルのインストールなど他の方法にも対応しています。詳しくは [GitHub リポジトリの説明](https://github.com/snarktank/ralph) をご覧ください。
***
## コアファイル構成
Ralph の記憶は完全にファイルシステムに依存しています。各ファイルの役割を理解することが、Ralph を使いこなすための前提です。
### ralph.sh — ループエンジン
これが Ralph の中核です。新しい AI インスタンスを繰り返し起動する bash スクリプトです。
```bash
# 基本的な使い方
./scripts/ralph/ralph.sh [max_iterations] # デフォルトは Amp を使用
./scripts/ralph/ralph.sh --tool claude [iterations] # Claude Code を使用
```
各イテレーションで ralph.sh は以下のことを行います:
1. 機能ブランチを作成(prd.json の `branchName` に基づく)
2. 最も優先度の高い未完了の story を選択(`passes: false`)
3. その story を実装するために**完全に新しい** AI インスタンスを起動
4. 品質チェックを実行(型チェック、テスト)
5. チェック通過 → git commit、失敗 → 次のイテレーションに持ち越し
6. prd.json を更新し、story を `passes: true` としてマーク
7. progress.txt に今回学んだ経験を追記
8. すべての story が完了するかイテレーション上限に達するまで繰り返し
デフォルトのイテレーション上限は10回です。プロジェクトの複雑さに応じて調整してください:
```bash
# シンプルなプロジェクト
./scripts/ralph/ralph.sh --tool claude 10
# 複雑なプロジェクト
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — タスク定義
これが Ralph の「頭脳」です。すべてのタスクがここで定義されます。フォーマットはフラットな JSON ファイルです:
```json
{
"projectName": "ブログ i18n 翻訳",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "トップページのメタデータを翻訳",
"description": "content/docs/meta.en.json を作成し、すべてのナビゲーション項目の英語翻訳を含める",
"acceptanceCriteria": [
"meta.en.json ファイルが存在し JSON 形式が正しい",
"すべてのナビゲーションタイトルが英語に翻訳されている",
"pnpm types:check が通過する"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "既存の meta.json の構造を参考にする"
},
{
"id": "US-002",
"title": "ブログ記事 hello-world を翻訳",
"description": "content/blog/hello-world.en.mdx を作成し、中国語から英語に翻訳する",
"acceptanceCriteria": [
"hello-world.en.mdx ファイルが存在する",
"すべての QuoteCard コンポーネントに defaultLang='en' が設定されている",
"内部リンクに /en/ プレフィックスが使用されている",
"コードブロックは翻訳されていない",
"pnpm types:check が通過する"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "MDX コンポーネントの props フォーマットを維持すること"
}
]
}
```
**フィールド説明**:
| フィールド | 説明 |
| -------------------- | --------------------------- |
| `projectName` | プロジェクト名、ログとブランチ命名に使用 |
| `branchName` | Git ブランチ名、Ralph が自動作成 |
| `id` | Story の一意識別子、`US-001` 形式を推奨 |
| `title` | 簡潔なタイトル |
| `description` | 詳細な説明、具体的であるほど良い |
| `acceptanceCriteria` | 受け入れ基準リスト——**最も重要なフィールド** |
| `priority` | 優先度の数値、小さいほど先に実行 |
| `passes` | 完了しているかどうか、Ralph が自動更新 |
| `dependsOn` | 依存する story ID のリスト |
| `notes` | 追加の備考やヒント |
### progress.txt — 経験ログ
これが Ralph の「長期記憶」です。各イテレーション終了後、AI がここに今回学んだ内容を追記します:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
次のイテレーションの新しい Claude インスタンスがこのファイルを読み取り、これまでのすべての経験を即座に取得します。これが Ralph がイテレーションを重ねるごとにスムーズに動くようになる理由です——**知識はイテレーション間で蓄積されますが、コンテキストはクリーンに保たれます**。
### AGENTS.md — 永続化ナレッジベース
progress.txt に加えて、Ralph はプロジェクト内の `AGENTS.md` ファイル(または `CLAUDE.md`)も更新します。Claude Code と Amp はどちらも起動時にこれらのファイルを自動的に読み取ります。
progress.txt とは異なり、AGENTS.md には**安定した、プロジェクト横断で汎用的な知識**が記録されます:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## PRD の作成
PRD(Product Requirements Document)の品質が Ralph の実行効果を直接左右します。うまく書ければ Ralph はスムーズに進みますが、書き方が悪いと同じ story で繰り返し失敗することになります。
### Skill を使った PRD 生成
snarktank/ralph の skill をインストールしていれば、インタラクティブに PRD を生成できます:
```bash
# Claude Code または Amp で
/prd 我想为博客系统添加 i18n 支持,需要将所有中文内容翻译成英文
```
AI がいくつかの確認質問(対象ファイル、技術スタックの制約、品質基準など)を行い、その後構造化された PRD ドキュメントを生成します。
生成後、`/ralph` コマンドで PRD を `prd.json` 形式に変換します:
```bash
/ralph # PRD を prd.json に変換
```
### 手動での PRD 作成
prd.json を直接作成することもできます。以下が重要な設計原則です。
**原則1:Story の粒度を適切にする**
各 story は1回のイテレーションで完了できる程度に小さく、独立した成果物として意味がある程度に大きくすべきです。
```json
// ❌ 大きすぎる:1回のイテレーションで完了できない
{
"id": "US-001",
"title": "完全なユーザー認証システムを構築",
"description": "登録、ログイン、パスワード忘れ、OAuth、権限管理を実装..."
}
// ❌ 小さすぎる:独立した価値がない
{
"id": "US-001",
"title": "User テーブルの email フィールドを作成",
"description": "User モデルに email フィールドを追加"
}
// ✅ 適切:1回で完了でき、独立した価値がある
{
"id": "US-001",
"title": "メール・パスワードログインを実装",
"description": "ログイン API とログインページを作成し、メール・パスワード認証に対応",
"acceptanceCriteria": [
"POST /api/auth/login が email + password を受け付ける",
"JWT token を返す",
"ログインページのフォームが送信可能",
"すべてのテストが通過する"
]
}
```
**経験則**:1つの story は1〜3ファイルの修正を含み、3〜5個の受け入れ基準を持ちます。
**原則2:受け入れ基準は自動検証可能であること**
Ralph は story の完了を判断する必要があるため、受け入れ基準は客観的に判定できるものでなければなりません:
```json
// ❌ 曖昧な基準
"acceptanceCriteria": [
"コード品質が良い",
"パフォーマンスが良い",
"ユーザー体験がスムーズ"
]
// ✅ 検証可能な基準
"acceptanceCriteria": [
"pnpm types:check が通過する",
"pnpm test が通過する",
"API レスポンス時間 < 200ms",
"ファイル src/auth/login.ts が存在し loginHandler 関数をエクスポートしている"
]
```
**原則3:dependsOn で順序を制御する**
story 間に依存関係がある場合、`dependsOn` フィールドで Ralph が正しい順序で実行することを保証します:
```json
{
"userStories": [
{
"id": "US-001",
"title": "データベーススキーマを作成",
"dependsOn": []
},
{
"id": "US-002",
"title": "ユーザー登録 API を実装",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "ログインページを実装",
"dependsOn": ["US-002"]
}
]
}
```
**原則4:notes でコンテキストを提供する**
notes フィールドは AI への追加ヒントです。あなたが知っていて AI が知らない可能性のある情報をここに書きましょう:
```json
{
"notes": "プロジェクトは fumadocs フレームワークを使用。i18n ファイルの命名規則は .en.mdx サフィックス。content/docs/notes/speckit/concept.en.mdx の翻訳スタイルを参考にすること。"
}
```
***
## Ralph Loop の実行
PRD の準備ができたら、ループを開始しましょう。
### 実行の開始
```bash
# Claude Code を使用、デフォルト10回のイテレーション
./scripts/ralph/ralph.sh --tool claude
# イテレーション回数を指定
./scripts/ralph/ralph.sh --tool claude 30
# Amp を使用(デフォルト)
./scripts/ralph/ralph.sh 20
```
### 実行プロセス
起動すると、以下のような出力が表示されます:
```
Starting Ralph - Tool: claude - Max iterations: 35
===============================================================
Ralph Iteration 1 of 35 (claude)
===============================================================
## US-001 Complete
**Summary of what was done:**
1. Created meta.en.json with all navigation items translated
2. Ran pnpm types:check — PASSED
3. Committed: feat: [US-001] - Translate homepage metadata
There are still **15 user stories with `passes: false`** remaining.
The next story is **US-002: 翻译博客文章 hello-world**.
Iteration 1 complete. Continuing...
===============================================================
Ralph Iteration 2 of 35 (claude)
===============================================================
```
各イテレーションは完全に新しい Claude インスタンスです。prd.json を読み取って現在何をすべきかを把握し、progress.txt を読み取ってこれまでに学んだことを把握します。
### 完了シグナル
すべての story が `passes: true` としてマークされると、Ralph は完了シグナルを出力して終了します:
```
All stories completed!
COMPLETE
```
### モニタリングとデバッグ
Ralph の実行中、以下のコマンドで進捗を確認できます:
```bash
# 各 story の完了状態を確認(アイコン付きで直感的)
cat tasks/prd.json | python3 -c "
import json,sys
for s in json.load(sys.stdin)['userStories']:
print(f'{\"✅\" if s[\"passes\"] else \"⬜\"} {s[\"id\"]}: {s[\"title\"]}')"
# または jq で確認
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 経験ログを確認
cat progress.txt
# 最近の git コミットを確認
git log --oneline -10
# Ralph の出力をリアルタイムで確認
tail -f progress.txt
# 完了後、メインブランチとの差分全体を確認
git diff main...ralph/your-branch-name --stat
```
### 中断と再開
Ralph の実行は長時間になることがありますが、途中での中断は完全に安全です:
* **中断**:`Ctrl+C` で直接中断できます。完了した story(`passes: true`)は失われません。すでに commit され prd.json に書き込まれています
* **再開**:同じコマンドを再度実行するだけで、Ralph は最初の `passes: false` の story から自動的に続行します
```bash
# 中断後の再開は、同じコマンドを再度実行するだけ
./scripts/ralph/ralph.sh --tool claude 35
```
特定の story が繰り返し失敗してブロックされている場合は、手動でスキップできます。`prd.json` を編集して該当 story の `passes` フィールドを `true` に変更し、再度実行すると、Ralph はそれをスキップして後続の story の処理を続行します。
### 自動アーカイブ
新しい `branchName` で別の機能を開始すると、Ralph は前回の実行ファイルを `archive/YYYY-MM-DD-feature-name/` ディレクトリに自動的にアーカイブし、作業ディレクトリを整理された状態に保ちます。
***
## フィードバックループと品質ゲート
Ralph の「自己修正」能力は完全にフィードバックループの品質に依存しています。フィードバックループのない Ralph は盲目的にループするスクリプトに過ぎません。コードを生成し続けますが、そのコードが正しいかどうかを判断できないのです。
### 品質チェックの設定
CLAUDE.md(または prompt.md)で品質チェックコマンドを定義します:
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 品質ゲートの階層
| 階層 | ツール | 検出される問題 |
| --------- | ------------------- | --------------------------- |
| 即時フィードバック | TypeScript compiler | 型エラー、構文エラー |
| 機能検証 | ユニットテスト | ロジックエラー、エッジケース |
| 統合検証 | Build コマンド | 依存関係の問題、設定エラー |
| ランタイム検証 | dev-browser skill | UI レンダリングの問題(フロントエンドプロジェクト) |
> フロントエンドの story では、Ralph は受け入れ基準に「Verify in browser using dev-browser skill」を追加することを推奨しています——AI に実際にブラウザを開いてページのレンダリングが正しいことを確認させます。
### 品質チェックが失敗した場合
特定の story の品質チェックが繰り返し失敗しても、Ralph は同じ story を無限にリトライしません。イテレーション上限に達すると停止し、現在の状態を残します。その場合は以下の対応ができます:
1. progress.txt を確認して AI がどこで詰まっているかを確認
2. 手動で問題を修正してから再実行
3. story の粒度を調整(大きすぎる可能性がある)
4. notes に追加のコンテキストを補足
### Prompt のカスタマイズ
Ralph の prompt テンプレート(CLAUDE.md または prompt.md)が AI の動作を制御する主要な手段です。インストール後、自分のプロジェクトに合わせてカスタマイズしましょう。主なカスタマイズの方向性:
**コーディングスタイルの制約**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**よくある落とし穴**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**詰まった場合の対処方法**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 実践事例:Ralph でブログの i18n 翻訳を完了する
Ralph が実際のプロジェクトでどのように動作するかを示すために、ここで実際の事例を紹介します:Ralph スタイルの自律 agent を使って、ブログ全体を中国語から英語に翻訳したケースです。
### プロジェクトの設定
プロジェクトでは22以上のコンテンツファイル(ブログ記事、ドキュメント、ナビゲーションメタデータ)を中国語から英語に翻訳する必要がありました。目標は fumadocs ベースの Next.js ブログでの i18n 対応です。タスクは `prd.json` ファイルに定義されており、16個の user story が含まれ、それぞれに明確な受け入れ基準がありました:
```
scripts/ralph/
├── prd.json # 16個の user story、受け入れ基準付き
└── progress.txt # 経験ログ、各 story 完了後に更新
```
各 user story は一貫したパターンに従っています:
* **明確な成果物**:"Create content/blog/xxx.en.mdx"
* **検証可能な基準**:"Typecheck passes"、"Internal links use /en/ prefix"
* **技術的な制約**:"Keep code blocks untranslated"、"Set defaultLang='en' on QuoteCard"
### 実行パターン
Agent は Ralph メソドロジーのコア原則に従いました:
1. **ファイルが唯一の真実の情報源**:`prd.json` が各 story のステータス(`passes: true/false`)を追跡。`progress.txt` がイテレーション間で経験を蓄積——例えば「Typecheck コマンドは `pnpm types:check` であり、`pnpm typecheck` ではない」など
2. **自動化された品質ゲート**:翻訳完了ごとに `pnpm types:check` を実行して MDX ファイルが正しくコンパイルされることを検証。typecheck が失敗した場合は、問題を修正してからコミット
3. **インクリメンタルな前進**:各 story を独立してコミットし、説明的なコミットメッセージを使用(`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`)。必要な場合にロールバックしやすくなります
4. **並列実行**:長い記事の場合、複数の subagent が同時に翻訳を実行——例えば US-010(claude-skills concept + practice)、US-011(speckit concept + practice)、US-012(claude-architecture + claude-subagent)が同時並行で実行
### 重要な学び
| 学び | 詳細 |
| ---------------------------- | ---------------------------------------------------------------------------------- |
| **知識の蓄積が重要** | 初期の story で発見したパターン(QuoteCard の `defaultLang`、リンクプレフィックスのルール)が後続の story をより速く完了させた |
| **フィードバックループとしての Typecheck** | 問題が積み重なる前に、欠落した import やフォーマット不正の MDX をキャッチ |
| **並列化はスケールする** | 6つの翻訳 agent が同時に実行され、完了時間は1つの場合とほぼ同じ |
| **PRD の粒度が極めて重要** | 各 story を1〜2ファイルに限定——確実に完了できる程度に小さく、意味がある程度に大きく |
| **進捗ログが同じミスの繰り返しを防ぐ** | `progress.txt` の "Codebase Patterns" セクションがナレッジベースとなり、同じ落とし穴を再び踏むことを防止 |
### 成果
16個の user story がすべて単一セッションで完了しました:8個の meta.en.json ナビゲーションファイルの作成、3つのブログ記事の翻訳、12個のドキュメントページの翻訳、完全なサイトビルドの検証通過。受け入れ基準が明確で、フィードバックループ(typecheck)が問題を即座に検出できたため、各翻訳は一貫した品質を維持しました。
このプロジェクトは Ralph の\*\*完全実装モード(Full Implementation Mode)\*\*を示しています——明確なタスク定義、明確な成功基準、自動化された検証、そしてファイルシステムを通じたインクリメンタルなデリバリーです。
***
## ベストプラクティスとよくある質問
### コスト管理
Ralph の自動実行は API コストが継続的に発生することを意味します。いくつかの制御手段があります:
* **常に `max_iterations` を設定する**:これが最も基本的なセーフティネットです
* **Story の粒度を合理的にする**:大きすぎる story は複数回のイテレーションを消費し、細かすぎる story は起動オーバーヘッドを増加させます
* **まず小規模でテストする**:新しいプロジェクトではまず3〜5回のイテレーションで試行し、prompt と品質ゲートが正常に機能することを確認してから本格的に実行しましょう
### よくある落とし穴
**落とし穴1:Story が大きすぎる**
症状:1つの story が繰り返し失敗し、イテレーション回数がすぐに尽きる。
解決策:2〜3個のより小さな story に分割する。「完全な認証システムを構築」を「ログイン API の実装」+「ログインページの作成」+「JWT ミドルウェアの追加」に分割する。
**落とし穴2:フィードバックループがない**
症状:Ralph が story 完了を報告するが、実際にはコードに問題がある。
解決策:受け入れ基準に実行可能なチェックコマンドを含める。「コードが書けた」は受け入れ基準ではなく、「pnpm test がすべて通過する」が受け入れ基準です。
**落とし穴3:progress.txt が活用されていない**
症状:同じエラーが異なるイテレーションで繰り返し発生する。
解決策:prompt テンプレートに「progress.txt を読み取り、その中の経験に従う」という明確な指示があることを確認する。AI が自動的に学習内容を追記しない場合は、prompt に「After each story, append learnings to progress.txt」を追加する。
**落とし穴4:依存関係の順序が間違っている**
症状:ある story が依存するコードがまだ存在せず、実装が失敗する。
解決策:`dependsOn` フィールドを正しく設定し、インフラストラクチャの story が先に来るようにする。
### よくある質問
**Q: 実行中に prd.json を手動で修正して介入できますか?**
できます。Ralph は各イテレーション開始時に prd.json を再読み込みします。イテレーションの合間に story の説明を修正したり、新しい story を追加したり、特定の story を手動で `passes: true` にマーク(スキップ)したりできます。
**Q: Ralph が1つの story で繰り返し失敗する場合はどうすればいいですか?**
1. progress.txt で失敗の原因を確認する
2. notes に追加のコンテキストを補足する
3. story を分割する(粒度が大きすぎる可能性がある)
4. ブロックしている問題を手動で修正してから再実行する
**Q: Ralph の実行中に他のことをしても大丈夫ですか?**
大丈夫です。Ralph は「Human on the Loop」として設計されています——ずっと監視している必要はありません。AFK モードで、退社前に起動して翌日結果を確認するだけで OK です。実行中は Ralph が操作しているファイルを変更しないようにしましょう。
**Q: コストをどう管理すればいいですか?**
3つの方法があります:合理的な `max_iterations` の設定、適切な story の粒度の維持(無駄なイテレーションの削減)、そしてまず小規模で試行してフローが正しいことを確認すること。一般的に、10〜20個の story のプロジェクトは $50〜100 の API コスト範囲内に収まります。
***
## まとめ
Ralph の使用フローは5つのステップに要約できます:
```
インストール → PRD 作成 → 品質ゲートの設定 → ループ実行 → 成果の確認
```
コアの考え方は常に変わりません:**ファイルを唯一の真実の情報源とし、各イテレーションを完全に新しいスタートとし、品質ゲートにチェックを任せる**。
さあ、自分のプロジェクトに戻って、prd.json を準備し、`./scripts/ralph/ralph.sh --tool claude` を実行して、コーヒーでも飲みに行きましょう。
### 関連記事
* 《[Ralph Wiggum 詳細解説](/ja/docs/notes/ralph-wiggum/concept)》— Ralph のコア原理を振り返る
* 《[frankbria/ralph-claude-code 実践ガイド](/ja/docs/notes/ralph-wiggum/frankbria)》— エンジニアリング指向の Ralph 実装:モニタリング、サーキットブレーカーと安全メカニズム
* 《[GSD 詳細解説](/ja/docs/notes/gsd/concept)》— Ralph の上に構築された完全なコンテキストエンジニアリングシステム
* 《[Claude Skills とは](/ja/docs/notes/claude-skills/concept)》— Ralph の PRD skill は Claude Skill の一つ
* 《[Speckit 実践ガイド](/ja/docs/notes/speckit/practice)》— もう一つの構造化 AI プログラミングワークフロー
# コンセプト紹介
## はじめに
2025年10月、GitHubはSpec Kitというツールキットをオープンソース化し、「スペック駆動開発」(Spec-Driven Development)という概念をAIプログラミングの世界に正式に持ち込みました。一見レトロに思えるこのアイデア――スペックを先に書いてからコードを書く――は、AIプログラミングツールを使いこなすための新しいパラダイムになりつつあります。
Claude Code、Cursor、GitHub Copilotなどの AIコーディングアシスタントを日常的に使っている方であれば、こんなフラストレーションを経験したことがあるのではないでしょうか。「ユーザーログイン機能を追加して」とお願いすると、AIは張り切って大量のコードを生成しますが、よく見てみると――使い慣れていないフレームワークを使っていたり、セキュリティ方針が期待と違っていたり、UIスタイルが合っていなかったり……そこから何度も修正を繰り返し、疲弊してしまいます。
問題はどこにあるのでしょうか。AIが賢くないのではなく、提供した情報が足りないのです。「ユーザーログイン機能を追加して」は一見明確に見えますが、実際には何百もの未指定の判断が隠れています。どの認証方式を使うのか?パスワードの要件は?ログイン失敗時はどう処理するのか?ログイン状態を記憶する必要はあるのか?サードパーティログインに対応するのか?……AIは推測するしかなく、推測は必ずズレを生みます。
スペック駆動開発は、まさにこの問題を解決するために生まれました。
## Vibe Coding:スピードの代償
2025年初頭、元Tesla AI ディレクターの Andrej Karpathy が「Vibe Coding」という言葉を生み出しました。これは「AIの提案を深くレビューせずに受け入れる」開発スタイルを表す言葉です。この用語は瞬く間に広まり、Collins辞書の2025年のWord of the Yearにも選ばれました。
Vibe Codingの魅力は明白です。アイデアを説明すれば、AIがコードを生成し、動いているように見えればそれでよし。ラピッドプロトタイピング、ハッカソン、使い捨てスクリプトであれば、この方法は確かに効率的です。しかし、本番システムに適用すると問題が生じます。
このような話は業界では珍しくありません。AI生成のデータベースクエリが小規模テストでは問題なく動作するものの、実際のトラフィック下ではシステムが這うように遅くなったり、つぎはぎの認証モジュールがQAを通過したのに、2週間後に無効化されたアカウントがまだ管理ツールにアクセスできることが判明したり。Final Round AIの2025年の調査によると、**18人のCTOのうち16人がAI生成コードによる本番障害を経験しています**。
これはVibe Codingに価値がないということではありません。重要なのは**境界を見極める**ことです。
| シナリオ | Vibe Coding | スペック駆動開発 |
| --------- | ----------- | -------- |
| プロトタイプ/デモ | 適している | 過剰 |
| 使い捨てスクリプト | 適している | 過剰 |
| 本番機能 | リスクが高い | 推奨 |
| セキュリティ関連 | 危険 | 必須 |
| チーム開発 | 保守が困難 | 推奨 |
スペック駆動開発は、AIの効率性を維持しつつ、Vibe Codingの落とし穴を避けることを目指しています。
## スペック駆動開発とは
スペック駆動開発のコアとなる考え方は一言で要約できます。**「何を作るか」を先に定義し、「どう作るか」はその後に考える**。
これはソフトウェアエンジニアリングでは当たり前のことのように聞こえますが、AIプログラミングの時代においては新たな意味を持ちます。従来の要件定義書は人間向けに書かれており、往々にして冗長で曖昧で、専門用語に満ちていました。スペック駆動開発における「スペック」はAI向けに書かれるものであり、簡潔で構造化され、実行可能なものです。
家を建てる場面を想像してみてください。従来のAIプログラミング方式は、施工チームに「快適な3LDKの家を建てて」と伝えて、あとはお任せするようなものです。結果は良いかもしれませんが、想像とは大きくかけ離れたものになる可能性が高いでしょう。スペック駆動開発は、まず建築設計図を描くことに相当します。何階建てか、各階の面積はどれくらいか、窓の向きは、使用する資材の規格は……施工チームは設計図通りに施工するため、結果は自然と期待通りになります。
AIプログラミングにおいて、この設計図が\*\*スペック(仕様書)\*\*です。どのプログラミング言語やフレームワークを使うかには触れず、機能が何を達成すべきか、ユーザーがどんなタスクを完了する必要があるか、成功の基準は何かだけに焦点を当てます。
従来の開発フローと比較すると、スペック駆動開発には根本的な違いがあります。
| 従来のAIプログラミング | スペック駆動開発 |
| ----------------------- | -------------------------- |
| 要件を直接記述 → AIがコード生成 | 要件 → スペック → 計画 → タスク → コード |
| AIが多くの詳細を推測する必要がある | 各ステップが明確で、AIは実行するだけ |
| 手戻りが頻繁で、コミュニケーションコストが高い | 前工程に投資し、後工程がスムーズ |
| シンプルなタスク向き | 複雑な機能向き |
この「段階的な精緻化」のプロセスこそが、スペック駆動開発の真髄です。一足飛びにコードに到達するのではなく、複数のステージを通じて要件を段階的に明確化していきます。各ステージでレビューと調整が可能です。
## Speckit ワークフロー概要
GitHubのSpec KitとClaude Codeのspeckitコマンドは、どちらも同様のワークフローに従っており、大まかに6つのステージに分けられます。
```
Constitution → Specify → Clarify → Plan → Tasks → Implement
↓ ↓ ↓ ↓ ↓ ↓
プロジェクト 機能スペック 曖昧さの 技術計画 タスク分解 実行・実装
憲法 解消
```
**1. Constitution(プロジェクト憲法)**
プロジェクト憲法は、プロジェクト全体の基本原則と制約を定義します。例えば「テストファースト」「シンプル至上」「APIファースト」などです。これらの原則は後続のすべてのステージに適用され、AIが生成するソリューションが技術的な好みに合致することを保証します。
**2. Specify(機能スペック)**
これが最も重要な最初のステップです。自然言語で望む機能を記述すると、AIがそれを構造化されたスペックドキュメントに整理してくれます。内容は以下を含みます。
* ユーザーストーリー:誰が何をしたいのか、なぜか
* 機能要件:システムが備えるべき能力
* 成功基準:機能が基準を満たしているかをどう判断するか
重要なのは、スペックドキュメントは\*\*「何を作るか」にのみ焦点を当て、「どう作るか」には触れない\*\*ことです。具体的な技術スタックやコード構造は含みません。
**3. Clarify(曖昧さの解消)**
AIがスペック内の曖昧な点をチェックし、最大5つの重要な質問を投げかけます。これらの質問は通常、機能の境界、ユーザータイプ、セキュリティ要件などに関するものです。この質疑応答を通じて、スペックはより明確になります。
**4. Plan(技術計画)**
明確なスペックが揃って初めて、技術的なアプローチの検討に入ります。このステップでは以下が成果物として生まれます。
* 技術選定(言語、フレームワーク、データベース)
* データモデル設計
* APIコントラクト定義
* リサーチレポート(技術的な判断の解決)
**5. Tasks(タスク分解)**
技術計画を実行可能なタスクリストに分解します。各タスクには明確なID、説明、ファイルパスがあり、AIに直接渡して実行させることができます。タスクはユーザーストーリーごとにグループ化され、並行開発に対応しています。
**6. Implement(実行)**
タスクリストに従って一つずつ実行していきます。各タスクの完了時にマークをつけ、トレーサビリティを確保します。
これら6つのステージの成果物は、明確なチェーンを形成します。
| ステージ | 成果物 | 目的 |
| ------------ | -------------------- | ----------- |
| Constitution | constitution.md | プロジェクト原則の定義 |
| Specify | spec.md | 機能要件の記述 |
| Clarify | 更新されたspec.md | 曖昧さの解消 |
| Plan | plan.md, research.md | 技術ソリューション設計 |
| Tasks | tasks.md | 実行可能なタスクリスト |
| Implement | 実際のコード | 最終成果物 |
## なぜこのアプローチが有効なのか
スペック駆動開発が「AIに直接コードを書かせる」方法で失敗する場面で成功するのは、AIプログラミングの根本的な矛盾である**情報の非対称性**を解決するからです。
「写真共有機能を追加して」と言ったとき、あなたの頭の中には完全なイメージがあるかもしれませんが、AIにはその数文字しか見えません。AIは推測するしかありません。どこに共有するのか?誰が見られるのか?圧縮は必要か?ウォーターマークは?一括処理は対応するのか?……一つ一つの推測が間違っている可能性があります。
スペック駆動開発は、「まず考え抜くことを強制する」ことでこの問題を解決します。ユーザーストーリー、機能要件、成功基準を書き出すことを求められると、「当たり前」だと思っていた詳細が浮かび上がってきます。このプロセス自体に価値があります。AIを使わなくても、要件を明確に書き出すだけでコミュニケーションコストを削減できるのです。
さらに、段階的な精緻化により、エラーがより早く表面化します。Specifyステージで要件のズレを発見すれば、修正コストはほぼゼロです。コードを書き終わってから発見すると、やり直しになるかもしれません。
もちろん、スペック駆動開発は万能ではありません。明確な適用場面があります。
**適している場合**:
* 複雑な機能開発(複数のモジュールやインタラクションが関わる)
* チーム開発プロジェクト(スペックドキュメントがコミュニケーションの媒体になる)
* 品質要件が高い場面(トレーサビリティと検証可能性が求められる)
**適していない場合**:
* 単純なバグ修正や軽微な変更
* 探索的プログラミング(何を作るかまだ決まっていない)
* 時間が極度に限られている場合(スペックを書く余裕がない)
重要なのは、タスクの複雑さを見極めることです。1時間で完了できるものに1時間かけてスペックを書く必要はありません。しかし、1週間かかる機能開発であれば、2時間かけてスペックを書く価値は十分にあります。
## ただし、スペックは銀の弾丸ではない
よくある誤解を一つ明確にしておきます。**スペック駆動開発は推測を減らしますが、レビューの必要性をなくすわけではありません**。
完全なスペックがあっても、AIは以下のような問題を起こす可能性があります。
* エッジケースの見落とし(スペックがカバーしていない極端なシナリオ)
* パフォーマンス要件を満たさないコードの生成
* 潜在的なセキュリティ脆弱性の混入
* スタイルが一貫しない実装の生成
これは建築と同じです。詳細な設計図があっても、検査は必要です。施工チームが設計図通りに家を建て終えたからといって、そのまま引っ越したりはしないでしょう。配線が安全か、配管が正常に機能しているか、ドアや窓がしっかりしているかを確認するはずです。
スペック駆動開発の価値は、**エラーを発見しやすくする**ことにあります。エラーそのものをなくすことではありません。
## まとめ
スペック駆動開発の核心はシンプルな真理です。**タスクが複雑であればあるほど、着手する前にしっかり考え抜く必要がある**。AIプログラミングツールはこの真理の重要性を増幅させます。なぜなら、AIはあなたの指示を忠実に実行しますが、あなたの意図を本当に理解することはできないからです。
3つのキーワードを覚えておきましょう。
| キーワード | 意味 |
| ---------------- | ----------------------------- |
| **スペックが先、コードが後** | 「何を作るか」を先に定義し、「どう作るか」はその後に考える |
| **段階的な精緻化** | 曖昧から明確へ、各ステップでレビューと調整が可能 |
| **推測を減らす** | 明確なスペック = AIの推測の余地が少なくなる |
理念を理解したところで、次の記事『[Speckit 実践ガイド](/ja/docs/notes/speckit/practice)』では、実際に手を動かしていきます。speckitコマンドを使って、一つの機能のスペック駆動開発フローを完成させる方法を紹介します。
『[Claude Skills](/ja/docs/notes/claude-skills/concept)』と組み合わせることで、スペックの実行をさらに自動化・標準化することができます。
# 実践ガイド
## はじめに
[前回の記事](/ja/docs/notes/speckit/concept)では、仕様駆動開発の理念を紹介しました。「何を作るか」を先に定義してから「どう作るか」を考えるというアプローチです。一見遠回りに思えるこのプロセスですが、実際には AI プログラミングにおける手戻りやコミュニケーションコストを大幅に削減できます。
今回は実際に手を動かしていきます。speckit コマンドスイートを使って、要件定義からコード実装までの完全なワークフローを学びましょう。
## インストールと設定
Speckit コマンドは GitHub 公式の [Spec Kit](https://github.com/github/spec-kit) プロジェクトに由来しています。使用シナリオに応じて、いくつかの導入方法があります。
### 新規プロジェクトの初期化
新規プロジェクトの場合は、公式の specify-cli ツールを使った初期化を推奨します。
```bash
# uv で specify-cli をインストール
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git
# 新規プロジェクトを初期化し、AI アシスタントとして Claude を指定
specify init my-project --ai claude
```
これにより、`.specify/` 設定ディレクトリや関連テンプレートファイルを含むプロジェクトディレクトリ構造が自動的に作成されます。
### 既存プロジェクトへの導入
Speckit コマンドを使うには設定ファイルが必要です。既存プロジェクトに speckit を導入するには、specify-cli を使います。
```bash
cd your-existing-project
specify init . --ai claude # 注意:. はカレントディレクトリを意味します
```
これにより、プロジェクトに以下が作成されます。
```
your-project/
├── .specify/
│ ├── templates/ # 仕様書、計画書などのテンプレート
│ ├── scripts/ # ヘルパースクリプト
│ └── memory/ # constitution.md
├── .claude/
│ └── commands/ # Claude Code コマンド設定
│ ├── speckit.specify.md
│ ├── speckit.plan.md
│ └── ...
└── specs/ # 機能仕様書の保存ディレクトリ
```
初期化では既存ファイルは上書きされません。完了後、Claude Code で `/speckit.*` コマンドスイートが使用可能になります。
> **注意**:speckit コマンドは Claude Code に組み込まれていません。上記の初期化手順を先に完了する必要があります。初期化せずに `/speckit.specify` を実行すると「コマンドが見つかりません」というエラーが表示されます。
***
## コマンドリファレンス
Speckit は仕様駆動開発の各フェーズをサポートするコマンドセットを提供しています。各コマンドには明確な入力と出力があり、トレーサブルなチェーンを形成します。
### /speckit.specify — 機能仕様書の作成
これはワークフロー全体の出発点です。実現したい機能を自然言語で記述すると、AI が構造化された仕様書に整理してくれます。
**機能**:自然言語の記述から機能仕様書を作成します
**入力**:機能の説明(自然言語)
**出力**:
* `specs/[番号]-[機能名]/spec.md` — 機能仕様書
* 新しい git ブランチ(例:`001-user-auth`)
**使用例**:
```
/speckit.specify ユーザーログイン機能を追加したい。メール・パスワード認証と「ログイン状態を保持する」オプションが必要
```
実行後、AI は以下を行います。
1. 短い機能名を生成(例:`user-auth`)
2. 新しい機能ブランチを作成
3. ユーザーストーリー、機能要件、成功基準を含む仕様書を生成
4. 不明確な箇所に `[NEEDS CLARIFICATION]` マークを付与
**仕様書の基本構造**:
```markdown
# Feature Specification: ユーザーログイン
## User Scenarios & Testing
### User Story 1 - ユーザーログイン (Priority: P1)
ユーザーがメールアドレスとパスワードでシステムにログインする...
**Acceptance Scenarios**:
1. Given 正しいメールアドレスとパスワードを入力, When ログインをクリック, Then システムに正常にログインできる
## Requirements
### Functional Requirements
- FR-001: システムはメール・パスワードによるログインをサポートすること
- FR-002: システムは「ログイン状態を保持する」オプションを提供すること
## Success Criteria
- SC-001: ユーザーが 30 秒以内にログインフローを完了できること
```
仕様書には**技術的な詳細は一切含まれない**ことに注意してください。フレームワークの指定も、データベーススキーマも、API 定義もありません。それらは後のフェーズで扱います。
***
### /speckit.clarify — 曖昧点の解消
仕様書の作成後、まだ曖昧な部分が残っている場合があります。このコマンドは仕様書をレビューし、重要な質問をして曖昧点の解消を支援します。
**機能**:仕様書の曖昧点を特定し、Q\&A を通じて仕様を洗練させます
**入力**:既存の spec.md ドキュメント
**出力**:更新された spec.md(明確化記録付き)
**使用例**:
```
/speckit.clarify
```
実行後、AI は以下を行います。
1. 仕様書の曖昧点をスキャン
2. 優先度順に並び替え(スコープ > セキュリティ > UX > 技術詳細)
3. 一度に一つずつ質問
4. 回答に基づいて仕様書を更新
**Q\&A の例**:
```markdown
## Question 1: ログイン失敗の処理
**Context**: 仕様書にはユーザーログインについて記載されていますが、ログイン失敗時の処理方法が定義されていません。
**Recommended:** Option B - 5 回連続失敗後のアカウントロックはセキュリティのベストプラクティスです
| Option | Description |
|--------|-------------|
| A | エラーメッセージのみ表示、制限なし |
| B | 5 回連続失敗後、15 分間アカウントをロック |
| C | ブルートフォース攻撃防止のため CAPTCHA を使用 |
オプション文字(例:"B")で回答するか、"yes" で推奨案を受け入れるか、独自の回答を提供してください。
```
各明確化の後、仕様書は自動的に更新され、明確化記録が追加されます。
```markdown
## Clarifications
### Session 2025-12-20
- Q: ログイン失敗時の処理は? → A: 5 回連続失敗後、15 分間アカウントをロック
```
***
### /speckit.plan — 技術計画の生成
仕様が明確になったら、技術設計フェーズに進みます。このステップでは技術計画とリサーチレポートが作成されます。
**機能**:仕様書から技術実装計画を生成します
**入力**:spec.md ドキュメント
**出力**:
* `plan.md` — 技術計画(アーキテクチャ、データモデル、API 設計)
* `research.md` — リサーチレポート(技術選定の判断根拠)
* `data-model.md` — データモデル(該当する場合)
* `contracts/` — API コントラクト(該当する場合)
**使用例**:
```
/speckit.plan Next.js + Prisma + PostgreSQL を使用
```
コマンドの後に技術スタックの希望を追加できます。実行後、AI は以下を行います。
1. 仕様書の機能要件を分析
2. 関連技術のベストプラクティスを調査
3. データモデルと API 構造を設計
4. 完全な技術計画を作成
**技術計画の主な内容**:
```markdown
# Implementation Plan: ユーザーログイン
## Technical Context
**Language/Version**: TypeScript 5.x
**Primary Dependencies**: Next.js 15, Prisma, PostgreSQL
**Authentication**: NextAuth.js with credentials provider
## Project Structure
src/
├── app/
│ └── (auth)/
│ ├── login/
│ └── api/auth/
├── lib/
│ └── auth/
└── prisma/
└── schema.prisma
## Data Model
- User: id, email, passwordHash, createdAt, updatedAt
- Session: id, userId, expiresAt
```
***
### /speckit.tasks — タスクの分解
技術計画ができたら、次はそれを実行可能なタスクリストに分解します。
**機能**:技術計画を実行可能なタスクリストに分割します
**入力**:plan.md ドキュメント
**出力**:`tasks.md` — 依存関係順にソートされたタスクリスト
**使用例**:
```
/speckit.tasks
```
実行後、AI は以下を行います。
1. plan.md から技術アプローチを抽出
2. spec.md からユーザーストーリーの優先度を抽出
3. ユーザーストーリーごとにグループ化されたタスクを生成
4. 並列実行可能なタスクに `[P]` マークを付与
5. 各タスクに具体的なファイルパスを指定
**タスクリストの形式**:
```markdown
## Phase 1: Setup
- [ ] T001 プロジェクト構造の作成
- [ ] T002 [P] Prisma schema の設定
- [ ] T003 [P] NextAuth の設定
## Phase 2: User Story 1 - ユーザーログイン (P1)
- [ ] T004 [US1] User モデルの作成 in prisma/schema.prisma
- [ ] T005 [US1] ログイン API の実装 in src/app/api/auth/[...nextauth]/route.ts
- [ ] T006 [US1] ログインページの作成 in src/app/(auth)/login/page.tsx
```
各タスクには以下が含まれます。
* **タスク ID**(T001, T002...)— トラッキング用
* **\[P] マーカー** — 他の \[P] タスクと並列実行可能であることを示す
* **\[US] タグ** — 所属するユーザーストーリーを示す
* **ファイルパス** — 作業対象のファイルを明示
***
### /speckit.implement — 実装の実行
すべての準備が整いました。タスクリストの実行に移りましょう。
**機能**:タスクリストに従って順番にタスクを実行します
**入力**:tasks.md ドキュメント
**出力**:実際のコード
**使用例**:
```
/speckit.implement
```
実行前に、AI はチェックリスト(存在する場合)を確認します。実行時は以下のように進みます。
1. フェーズ順にタスクを実行
2. 完了したタスクを `[X]` でマーク
3. タスクの依存関係を遵守
4. 並列タスクは同時に実行可能
**実行の例**:
```
Phase 1: Setup
✓ T001 プロジェクト構造の作成
✓ T002 Prisma schema の設定
✓ T003 NextAuth の設定
Phase 2: User Story 1
✓ T004 User モデルの作成
T005 実行中...
```
### 実装後のレビュー
`/speckit.implement` の完了後、**コードを直接マージしないでください**。AI が生成したコードには人間によるレビューが必要です。
**必須の検証ステップ**:
1. **テストスイートの実行**
```bash
npm test # またはプロジェクトのテストコマンド
```
AI が既存の機能を壊していないことを確認します。
2. **コードレビューのチェックポイント**
* コードが仕様の意図に沿っているか(spec.md と照合)
* プロジェクトのコーディングスタイルに従っているか
* セキュリティ上の潜在的な問題がないか
3. **境界テスト**
AI が見落とした可能性のあるエッジケースを手動でテストします。
* null 値の処理
* 極端な入力
* 並行処理のシナリオ
* エラーパス
4. **パフォーマンスチェック**
データベース操作や API コールが関係する場合、N+1 クエリなどのパフォーマンス問題を確認します。
> **ヒント**:仕様書が十分に詳細であっても、AI は実装の細部で逸脱する可能性があります。レビューは仕様駆動開発への不信ではなく、エンジニアリング規律の一環です。
***
### /speckit.analyze — 一貫性分析
これはオプションの品質チェックステップで、仕様書・計画書・タスク間の一貫性を検証します。
**機能**:ドキュメント間の一貫性と品質の分析
**入力**:spec.md, plan.md, tasks.md
**出力**:分析レポート(ファイルは変更されません)
**使用例**:
```
/speckit.analyze
```
実行後、以下を確認します。
* すべての要件に対応するタスクがあるか
* タスクがすべてのユーザーストーリーをカバーしているか
* 用語が統一されているか
* 抜け漏れや重複がないか
***
### その他のコマンド(オプション)
上記のコアコマンドに加えて、speckit にはいくつかの補助コマンドがあります。メインワークフローには含まれませんが、特定のシナリオで役立ちます。
**`/speckit.constitution`** — プロジェクト憲章の作成
プロジェクトの開発原則や規約を定義するために使用します。チーム全員が統一された開発標準に従うことを保証するのに最適です。
* **入力**:インタラクティブな Q\&A または直接提供される原則
* **出力**:`.specify/constitution.md` プロジェクト憲章ファイル
* **ユースケース**:新規チームプロジェクトの初期化、コーディングスタイルやアーキテクチャ判断の統一
**`/speckit.checklist`** — 品質チェックリストの生成
機能仕様書に基づいてカスタマイズされた品質チェックリストを生成し、実装前の品質保証に使用します。
* **入力**:spec.md ドキュメント
* **出力**:`checklists/` ディレクトリ内のチェックリスト
* **ユースケース**:重要な機能リリース前の品質ゲート、コードレビューの参考
**`/speckit.taskstoissues`** — タスクを GitHub Issues に変換
tasks.md のタスクを自動的に GitHub Issues に変換し、チームコラボレーションとタスク割り当てを容易にします。
* **入力**:tasks.md ドキュメント
* **出力**:GitHub Issues(gh CLI 経由で作成)
* **ユースケース**:チーム開発、スプリント計画、タスクトラッキング
***
## ツールエコシステム
本記事で紹介した speckit コマンドは [GitHub Spec Kit](https://github.com/github/spec-kit) プロジェクトに由来しています。これに加えて、2025 年には複数の主要 AI コーディングツールが同様の仕様駆動ワークフローをサポートし始めました。
| ツール | 特徴 | 適した用途 |
| --------------------------------------------------------- | -------------------------------------------------------------- | -------------------- |
| **[GitHub Spec Kit](https://github.com/github/spec-kit)** | 本記事で使用したツール、MIT ライセンス、Claude Code / Copilot / Gemini CLI をサポート | コマンドライン派、クロスツール連携 |
| **[AWS Kiro](https://kiro.dev/)** | VS Code フォーク、ビジュアルワークフロー、EARS 表記法 | GUI 派、AWS エコシステムユーザー |
| **[JetBrains Junie](https://blog.jetbrains.com/junie/)** | IntelliJ エコシステム統合、Think More 推論モード | JetBrains IDE ユーザー |
| **Cursor Plan Mode** | 組み込みの計画フェーズ、実行計画の自動生成 | Cursor を既に使用している開発者 |
**選び方**:
* Claude Code、GitHub Copilot、または Gemini CLI を使用している場合は、GitHub Spec Kit を推奨します
* グラフィカルなインターフェースやビジュアルワークフローを好む場合は、AWS Kiro をお試しください
* JetBrains ユーザーの場合は、Junie の方が IDE との統合が自然です
* 既に Cursor を使用している場合は、Plan Mode が同様の計画機能を提供しています
核心的な思想は共通です。ツールはあくまで手段であり、重要なのは**仕様ファースト、コードセカンド**という考え方です。
***
## エンドツーエンドのウォークスルー
実際の事例を使って、speckit の完全なワークフローを一通り見ていきましょう。
**シナリオ**:ブログシステムに記事ブックマーク機能を追加する
### Step 1: 仕様書の作成
```
/speckit.specify ブログに記事ブックマーク機能を追加したい。ユーザーがお気に入りの記事を保存して、プロフィールで閲覧できるようにしたい
```
AI が仕様書 `specs/002-article-bookmark/spec.md` を生成します。
```markdown
# Feature Specification: 記事ブックマーク機能
## User Scenarios & Testing
### User Story 1 - 記事をブックマーク (Priority: P1)
記事を閲覧中に、ユーザーがブックマークボタンをクリックしてブックマークリストに追加する。
**Acceptance Scenarios**:
1. Given ユーザーがログイン済み, When ブックマークボタンをクリック, Then 記事がブックマークリストに追加される
2. Given ユーザーが未ログイン, When ブックマークボタンをクリック, Then ログインプロンプトが表示される
### User Story 2 - ブックマークリストの閲覧 (Priority: P2)
ユーザーがプロフィールで自分のブックマークした記事をすべて閲覧する。
## Requirements
- FR-001: ユーザーは記事をブックマーク/ブックマーク解除できること
- FR-002: ブックマークボタンは現在のブックマーク状態を反映すること
- FR-003: プロフィールにブックマークリストを表示すること
## Success Criteria
- SC-001: ブックマーク操作が 500ms 以内に完了すること
- SC-002: ブックマークリストはページネーションをサポートし、1 ページあたり 10 件表示すること
```
### Step 2: 要件の明確化
```
/speckit.clarify
```
AI が質問します:「ブックマーク数に上限はありますか?」
回答:「ユーザーあたり最大 100 件のブックマーク」
仕様書が更新されます。
* FR-004: 各ユーザーのブックマークは最大 100 件まで
* 上限に達した場合に通知を表示
### Step 3: 計画の生成
```
/speckit.plan Next.js + Prisma を使用
```
生成された技術計画には以下が含まれます。
* Bookmark モデル(userId, articleId, createdAt)
* API ルート設計(POST/DELETE /api/bookmarks)
* コンポーネント設計(BookmarkButton, BookmarkList)
### Step 4: タスクの分解
```
/speckit.tasks
```
生成されたタスクリスト:
```markdown
## Phase 1: Setup
- [ ] T001 Prisma schema に Bookmark モデルを追加
## Phase 2: US1 - 記事のブックマーク
- [ ] T002 [US1] ブックマーク API の作成 in src/app/api/bookmarks/route.ts
- [ ] T003 [US1] BookmarkButton コンポーネントの作成 in src/components/BookmarkButton.tsx
- [ ] T004 [US1] 記事ページへの統合
## Phase 3: US2 - ブックマークリスト
- [ ] T005 [US2] ブックマークリストページの作成 in src/app/profile/bookmarks/page.tsx
- [ ] T006 [US2] ページネーションロジックの実装
```
### Step 5: 実装の実行
```
/speckit.implement
```
タスクが順番に実行され、完了するごとに `[X]` でマークされます。
***
## ベストプラクティスと注意事項
### Speckit を使うべきタイミング
**適しているケース**:
* 新機能の開発(3 つ以上のファイルに関わる場合)
* 要件が完全には明確でない場合(clarify で解消)
* 複数人での共同プロジェクト(仕様書が共通認識として機能)
* 重要な機能(トレーサビリティが必要な場合)
**適していないケース**:
* 単純なバグ修正
* 1 行のコード変更
* 緊急のホットフィックス
* 純粋な探索的実験
### よくある落とし穴
speckit を使う際に注意すべき一般的な落とし穴がいくつかあります。
**落とし穴 1:仕様が曖昧すぎる**
症状:AI が生成したコードが期待と大きくかけ離れ、大幅な手戻りが必要になります。
```markdown
# ❌ 曖昧な仕様
ユーザーが記事を検索できる
# ✓ 明確な仕様
- FR-001: ユーザーはタイトルのキーワードで記事を検索できること
- FR-002: 検索結果は関連度順にソートされ、1 ページあたり 10 件表示すること
- FR-003: 検索結果内で検索キーワードがハイライト表示されること
- FR-004: 検索キーワードが空の場合、トレンド記事を表示すること
```
解決策:`/speckit.clarify` を実行するか、手動で機能要件と成功基準を追加してください。
**落とし穴 2:仕様が詳細すぎる**
症状:AI が過度に制約されて力を発揮できず、硬直したコードを生成するか、指示の一部を無視してしまいます。
```markdown
# ❌ 過度に詳細(実装の詳細を指定)
lodash の debounce 関数を使用し、遅延 300ms、
useCallback でラップし、依存配列は [searchTerm]...
# ✓ 適切な詳細レベル(何をするかだけを記述)
検索入力はデバウンスし、過剰なリクエストを避けること
```
解決策:仕様は「何をするか」のレベルに留め、「どうするか」は Plan フェーズに委ねてください。
**落とし穴 3:Plan フェーズのスキップ**
症状:タスクが粗すぎるか細かすぎ、実装中に頻繁な手戻りが発生し、タスク間の依存関係が混乱します。
解決策:複雑な機能では必ず Plan フェーズを完了してください。Plan は技術アプローチを策定するだけでなく、潜在的なアーキテクチャ上の問題を特定するのにも役立ちます。
**落とし穴 4:レビューなしでマージ**
症状:デプロイ後にエッジケースの未処理、セキュリティ脆弱性、パフォーマンス問題が発見されます。
解決策:上記の「実装後のレビュー」セクションを参照してください。マージ前に必ずテストとコードレビューを実施してください。
### よくある質問
**Q: すべての機能でワークフロー全体を実行する必要がありますか?**
いいえ、必ずしもそうではありません。シンプルな変更はそのままコーディングに進んで構いません。複雑な機能の場合は、少なくとも specify + plan を完了することを推奨します。
**Q: 仕様書はとても詳細なのに、AI が期待どおりのコードを生成しない場合は?**
仕様が本当に「詳細」かどうかを確認してください。明確に書いたつもりでも、実際には曖昧な点が残っていることが多いです。`/speckit.clarify` を実行して、見落としがないか確認してみてください。
**Q: 一部のステップをスキップできますか?**
はい、可能です。最小ワークフローは specify → tasks → implement です。ただし、clarify と plan をスキップすると、後から手戻りが発生するリスクが高まります。
**Q: 既に生成された仕様書を修正するには?**
spec.md ファイルを直接編集するだけで大丈夫です。変更後は、一貫性を保つために plan と tasks を再実行することを推奨します。
**Q: AI が生成したコードが全く間違っている場合、どうデバッグしますか?**
段階的にトラブルシューティングしてください。
1. **仕様を確認**:仕様が本当に明確ですか?`/speckit.clarify` を実行して見落としがないか確認します
2. **計画を確認**:plan.md の技術アプローチは妥当ですか?妥当でなければ、直接編集してタスクを再生成します
3. **スコープを絞る**:AI にタスクを一つだけ実行させ、出力が期待に沿っているか観察します
4. **制約を追加**:constitution.md により明確な技術プリファレンスを追加します
**Q: Plan と Tasks の間に不整合がある場合は?**
`/speckit.analyze` を実行すると不整合を検出できます。よくある原因は以下の通りです。
* Plan を更新した後、Tasks を再生成していない
* Tasks を手動で編集したが、Plan を更新していない
* 仕様変更後、一部のドキュメントしか更新していない
解決策:spec.md を信頼できる唯一の情報源として、plan.md と tasks.md を順番に再生成してください。
**Q: 機能間の依存関係はどう扱いますか?**
機能 B が機能 A に依存している場合、2 つのアプローチがあります。
1. **仕様を統合**:A と B を同じ spec.md に記述し、AI に一括で計画させる
2. **段階的に開発**:まず機能 A のワークフローを完了してから、機能 B の specify を開始する
依存関係のある複数の機能を同時に開発することは推奨しません。インテグレーションの問題が発生しやすくなります。
***
## まとめ
Speckit の核心的な価値は、プロセスを増やすことではなく、**暗黙知を形式知に変える**ことにあります。ユーザーストーリー、機能要件、成功基準を書き出すことを求められると、「言わなくても分かる」と思っていた細部が自然と浮かび上がってきます。
このワークフローを覚えておいてください。
```
Specify → Clarify → Plan → Tasks → Implement
定義 明確化 設計 分解 実行
```
各ステップが次のステップの曖昧さを減らしていきます。最終的に AI が受け取るのは、曖昧な意図の記述ではなく、明確なタスクリストです。
さあ、あなたのプロジェクトに戻って、`/speckit.specify` で最初の仕様駆動開発ワークフローを始めてみましょう。
### 参考資料
* 《[仕様駆動開発とは](/ja/docs/notes/speckit/concept)》— 核心理念を振り返る
* 《[GSD 詳細解説](/ja/docs/notes/gsd/concept)》— 仕様駆動の考え方を採用したもう一つのコンテキストエンジニアリングシステム
* 《[Claude Skills とは](/ja/docs/notes/claude-skills/concept)》— Speckit 自体が Claude Skill の一つ
# AI 時代の TDD: モデルに最初に赤信号を当てさせます
## まず結論から話しましょう
AI がコードを作成した後、TDD は時代遅れではありませんが、その立場は変わりました。
以前、TDD について話すとき、最初にテストを作成し、次に実装を作成し、小さなステップでリファクタリングするというプログラマーの自制心についてよく話していました。 AI プログラミングに関して言えば、ブレーキ システムに似ています。なぜなら、モデルが最も得意とすることは、同時に最も危険なことでもあるからです。それは、完全に見える大きなコードを素早く書くことができるからです。
関数の実装を依頼すると、次のような結果が得られます。
* 実装ファイル
* 一連のテスト
* 説明
* 「完了」という言葉
問題は、「完全に見える」ということはエンジニアリングの意味では行われないということです。工学的な意味で完了するには、少なくとも次のように答えてください。
> この動作は、明示的に失敗したテストによって定義されているのでしょうか?
> この失敗は実装によって緑色になったのでしょうか?
> 緑色になった後、動作を変えずにコードを整理しましたか?
AI時代にTDDが再び話題になっているのはこのためだ。
それはプロセスを高度に見せることではなく、「モデルへの信頼」を「フィードバックへの信頼」に置き換えることです。
## 1. よくある誤解: TDD は「最初にテストを書く」ものではありません
多くの人が TDD を嫌うのは、TDD を儀式だと理解しているからです。
```text
先写测试。
再写代码。
最后跑一下。
```
これは確かに退屈で、形式主義になってしまいがちです。
本当に役立つ TDD は、「テスト ファイルが早く現れる」ことではなく、「障害が十分早く現れる」ことです。
### 重要なのはテストではなく赤です
TDD の最初のステップは、TEST ではなく RED と呼ばれます。
赤は、システムが明らかに失敗するように最初にテストを作成することを意味します。これが失敗するには、次の 3 つのことが当てはまる必要があります。
1. 失敗します。
2. ターゲットの動作が存在しないため、失敗します。
3. 期待通りに失敗する。
最初に赤を見なければ、その後に続く緑は意味がありません。
たとえば、`slugify("Hello World") -> "hello-world"` を実装するとします。貴重な RED は、「テスト ファイルを作成した」ことではなく、次のことです。
```text
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
Failure: NameError: name 'slugify' is not defined
Reason: 目标函数还不存在,符合预期
```
このとき、テストが仕様になります。次のステップに進むには、この動作を true にするだけで済みます。
### まずグリーンにしてからテストを補い、通常はストーリーを作り上げます
AI が逆の方向に進むことは簡単です。最初に実装を記述し、後でテストを追加します。
これはスムーズな体験です。コードが実行され、テストが実施されたのを見ると、心の中で「もうすぐ」と感じるでしょう。しかし、これには致命的な問題があります。テストは現在の実装を遡及的に実装する可能性が高いのです。
「要件がどうあるべきか」を問うのではなく、「現在のコードを簡単に通過できるようにどのように記述するか」を問うのです。
AI がテストを作成する際に次のような匂いがするのはこのためです。
* アサーションが現在の実装に固有すぎる
* モックが多すぎて、実際の境界が測定されていません。
* ハッピーパスのみをテストします
* 既存のコードを通過させるには、アサーションを非常に広範囲に作成します
* 古いコードがもともと間違っていたことを証明できるテストはありません
TDD ではその逆が必要です。まず要件を失敗させてから、コードが要件に追いつくようにします。
## 2. AI 時代に TDD がより必要とされる理由
AI プログラミングの核心的な矛盾は、「コードの記述が遅い」ことではなく、「フィードバックが遅い」ことです。
TDD を使用しない場合、通常は次のように作業します。
```text
描述需求 -> AI 写一堆代码 -> 人肉看 diff -> 跑一下 -> 发现问题 -> 回头修
```
最後まで疑問は山積するだろう。それが間違っていると気づく頃には、次の 3 つのカテゴリのものが混在している可能性があります。
* 要件の誤解
* 実装パスが間違っています
* リファクタリングにより古い動作が破壊される
TDD の役割は、この長いチェーンを短縮することです。
### モデルに決定可能な目標を与えます
「エレガントに書く」ことが目的ではありません。
「ユーザーのログインステータスの有効期限が切れた後、自動的にログインページに戻る」という表現は十分に具体的ではありません。
より良い目標は次のとおりです。
```text
当 access token 过期时:
1. 请求返回 401。
2. 客户端清理本地 session。
3. 用户被重定向到 /login。
4. 原始目标地址被保存在 redirect 参数里。
```
さらに一歩進んで、そのうちの 1 つを失敗したテストに変えます。
```text
given expired session
when user opens /settings
then app redirects to /login?redirect=/settings
```
現時点では、AI は「ログイン状態の期限切れをどのように処理するか」を推測するのではなく、明確な動作を完了します。
### 大きなタスクを小さな閉じたループに分割します
AI が制御を失いやすいのは、すべてを一度に実行することです。
ログイン、権限、リフレッシュ トークン、エラー プロンプト、ルート ジャンプを一度に実装すると、最終的には大きな差が得られます。うまくいくかもしれないが、レビューコストは高い。ビジネス、ステータス、ルーティング、境界、テスト、リファクタリングを同時に判断する必要があります。
TDD のリズムは次のようになります。
```text
一个行为 -> 一个失败测试 -> 最小实现 -> 变绿 -> 再下一个行为
```
一度に少量だけ進めてください。それは理解できるほど小さく、AI がストーリーを作るのが難しいほど小さく、失敗したときにすぐに特定できるほど小さいです。
### モデルが「スムーズにプレイできる」ように制限されます。
AI によくある問題は、過度の熱意です。
境界のバグを修正するように要求すると、ヘルパーが抽出されます。テストを追加するように要求すると、実装が変更されます。リファクタリングを要求すると、動作が変わります。
TDD はフェーズを使用してこれらのアクションを分離します。
| ステージ | 何をすべきか | してはいけないこと |
| ------ | ---------- | ------------------ |
| レッド | 失敗するテストを書く | 本番環境の実装を作成する |
| グリーン | 最小限の実装を書く | 緑色になるようにテストを変更します。 |
| リファクター | クリーンアップ構造 | 新しい動作を導入する |
「ご注意ください」よりもこの表の方が役に立ちます。これにより、モデルがどの段階にあるかを認識できるようになり、人間が境界違反を検出しやすくなります。
## 3. 赤と緑の再構築: 3 つのスローガンではなく 3 つのドア
「レッド、グリーン、リファクタリング」は簡単にスローガンになる可能性があります。実際に使用すると、3 つのドアのように見えるはずです。ドアを通過するたびに証拠を残さなければなりません。
### 最初の扉: 赤、ニーズが満たされていないことを証明
RED フェーズ中の最も重要な質問は次のとおりです。
> このテストが失敗した場合、目標とする行動がまだ不足していることが証明されるのでしょうか?
悪い赤:
```py
assert True
```
REDもあまり良くありません:
```py
assert "hello" in format_title("Hello World")
```
広すぎます。多くの欠陥のある実装も合格します。
より良い赤:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
このテストは小規模ですが、明らかです。入力、出力、および動作を指定します。
### 2 番目のドア: 緑、現在のテストのみを通過させます
GREEN フェーズは、最終的なアーキテクチャを記述することではありません。
ミッションは 1 つだけです。それは、現在失敗しているテストに最小限のコードで合格することです。
この発言は直感に反して聞こえます。多くの人は、「最小限の実装」があまりにも醜いのではないかと心配しています。はい、時には醜いこともあります。しかし、その価値は設計圧力を維持することにあります。
最初のテストが次の場合:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
許容される GREEN は次のとおりです。
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
中国語、アクセント、連続句読点、絵文字、SEO の特殊なケースをすぐにサポートする必要はありません。これらは後のテストによって駆動される必要があります。
### 3 番目のドア: REFACTOR、構造のみを変更し、動作は変更しません
REFACTOR ステージは AI によって最も混乱されやすいです。
「コードを整理する」を「ついでに強化する」と解釈します。これではうまくいきません。リファクタリングの定義は非常に狭いです。外部の動作は変わりませんが、内部構造は改善されます。
適切なリファクタリングは次のようになります。
* 変数名をより正確なものに変更します
* 繰り返される表現を抽出する
* 深すぎる条件分岐を削除する
* モジュールの責任を明確にするために関数の場所を移動します。
悪いリファクタリングは次のようになります。
* 新しい入力がサポートされるようになりました
* エラーメッセージを簡単に変更しました
* 依存関係を簡単に変更
* テストアサーションを便利に変更しました
判断基準はシンプルです。
> このコミットが `refactor:` と呼ばれただけであれば、テストの前後で同じ緑色になるはずで、ユーザーの動作も同じになるはずです。
## 4. 良いテストの味
TDD は、テストが多ければ多いほど良いというものではありません。 AI は、ほとんど価値のないテストを大量に生成することにも優れています。
さらに重要なのは味を試すことです。
### 優れたテストは仕様書のようなものです
優れたテストは、ビジネス仕様のように読み取る必要があります。
```text
当用户没有权限时,保存按钮不可点击。
当标题为空时,表单显示错误信息。
当重复提交同一个请求时,只创建一条记录。
```
それは内部で行われることではなく、外部の動作に関係します。
悪いテストは実装ノートのように見えます。
```text
应该调用 validateInput 三次。
应该读取 state.user.flags。
应该触发 handleClick 内部函数。
```
実装の詳細が決まると、リファクタリングは困難になります。内部構造を変更しただけですが、大規模なテストは失敗しました。このようなテストでは、コードを保護するのではなく、コードをフリーズします。
### 優れたテストには限界がある
テストは 1 つの質問にのみ答えるように設計するのが最適です。
テストで次のこともアサートされる場合:
* 正しい形式
* 権限は正しいです
* ネットワークリクエストは正しいです
* トースト コピーは正しいです
* データベースのステータスが正しい
失敗すると、何が問題なのかを知るのは困難です。
AI は一度に多くのことを証明したいため、このような「大規模で包括的な」テストを作成する傾向があります。 TDD はその逆で、アクション、失敗、実装です。
### 優れたテストにより、実装での不正行為が困難になります
テストがあまりにも具体的な入力のみをカバーする場合、AI はちょうど一致する偽の実装を作成する可能性があります。
たとえば:
```py
def slugify(text: str) -> str:
return "hello-world"
```
最初のテストでは合格しますが、2 番目のテストでは実際のロジックが強制的に実行されます。
```py
def test_slugify_handles_another_title():
assert slugify("Test Driven Development") == "test-driven-development"
```
したがって、TDD は常に 1 つのテストだけを作成するわけではなく、各ラウンドで 1 つの動作プレッシャーのみを追加します。徐々に圧力が増し、模様が徐々に伸びていきます。
## 5. AI はどのように TDD を回避するのでしょうか?
AI は本来テストを尊重しないため、この部分を明確にする必要があります。
最適化の目標はシンプルです。先ほど述べたタスクを完了することです。 「テストをパスさせてください」と言うと、人間が望まないアクションを実行する可能性があります。
### 最初のタイプ: テストを変更して緑色にします
最も典型的なもの:
* `assert slugify("Hello World") == "hello-world"` を現在の出力に変更します
* 失敗したアサーションを削除する
* `skip` をテストに追加します
* 厳密なアサーションを緩やかなアサーションに変更する
これはTDDではなく、赤い光を消しているのです。
### 2 番目のタイプ: 書き込みオーバーフィッティング実装
たとえば、テストには入力が 1 つだけあります。
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
モデルは次のように記述できます。
```py
def slugify(text: str) -> str:
if text == "Hello World":
return "hello-world"
return text
```
このとき叱る必要はありません。実装がハードコーディングされ続けないように、次の動作を追加し続ける必要があります。
### 3 番目の方法: モックを使用して実際の境界をカバーする
AIはモックが大好きです。モックを使用すると、テストの作成が容易になり、実際の問題の多くが解消されます。
嘲笑できないわけではありませんが、次のことを尋ねる必要があります。
> 私が今嘲笑しているのは遅い依存関係ですか、それとも本当に検証したい境界線ですか?
支払いコールバック解析を検証したいが、解析層を模擬する場合、テストは無意味になります。
## 6. TDD を使用しない場合
TDD には価値がありますが、すべてが価値があるわけではありません。
### 不適切なシーン
* 純粋な視覚的な微調整
* ワンタイムスクリプト
* 技術探索デモ
* 要件自体が明確に考えられていないプロトタイプ
* テスト フレームワークはまだウェアハウスをセットアップしていません
このようなシナリオでは、まず探索速度を追求し、プロセスに足を引っ張られないようにしてください。
### 適したシーン
* バグ修正
* 権限、課金、ステートマシン
* データ変換と境界処理
* 長期間保守されるコアモジュール
* AIが繰り返し変更するコードパス
判断基準は「この機能が素晴らしいかどうか」ではなく、
> 間違っていた場合、コストは明らかですか?
コストは明らかなので、最初にテストを作成する価値があります。
## 7. 実行可能なメンタルメソッド
AI に一言だけ与えるとしたら、次のようには言いません。
```text
请高质量实现这个功能。
```
私ならこう言います:
```text
先写一个失败测试,运行它,确认失败原因符合预期。不要写实现,直到我说 go。
```
この文は、モデルに「適切に動作する」ことを求めているのではなく、チェック可能なプロセスに入るように求めているため、より高品質です。
もう少し完全なもの:
```text
每轮只处理一个行为。
RED:写一个失败测试并运行。
GREEN:写最小实现,不改测试。
REFACTOR:只在绿色状态下整理结构。
每轮报告测试文件、命令、失败原因、通过结果。
```
これがAI時代のTDDの核心です。
テストやプロセスについて迷信深いのではなく、すべてのステップについて証拠を持つことについて考えています。
## 終わりに
AI プログラミングに最も必要なのは、より多くのコードではなく、より短いフィードバックです。
TDD の価値は次のとおりです。TDD は、「正しいはずだと思っていた」ことを、「ここで失敗したが、その後青信号になった」に変えるのです。変化は小さいですが、十分に現実的です。
一文しか覚えていない場合は、次のことを思い出してください。
> AI にコードを直接配信させないでください。最初に赤色の光を照射し、次に赤色の光を緑色に変えます。
次の記事 [実践ガイド](/ja/docs/notes/tdd-with-ai/practice) では、このリズムを直接コピーできるワークフローに変換します。
## 推奨リソース
# AI時代のTDD:Codex実践マニュアル
## まず地図をください
[コンセプト](/ja/docs/notes/tdd-with-ai/concept) 内容: AI プログラミングにおける TDD の価値は、「最初にテストを書く」という儀式ではなく、最初に検証可能な赤信号を作成し、その後実装で青信号に変えることです。
この記事では、Codex を実装する方法について説明します。
すぐに「どの書類を照合する必要があるか」と尋ねないでください。もっと良い質問は次のとおりです。
> Codex を毎回同じ TDD ワークフローで動作させるにはどうすればよいですか?
このワークフローは 4 つのレベルに分割できます。
| 階層 | 何をしている | どこ | いつが適切ですか | |
| -- | ------------ | ----------------------------------- | -------------------------- | ------ |
| L1 | プロジェクトの規律を書く | `AGENTS.md` | すべてのプロジェクトには | が必要です。 |
| L2 | 固化プロセス | `.agents/skills/tdd-codex/SKILL.md` | 要件を満たすために TDD を繰り返し使用する | |
| L3 | 隔離段階 | `.codex/agents/*.toml` | 複雑なタスク、テストと実装の間の相互汚染の恐れ | |
| L4 | 自動リマインダー | `.codex/hooks.json` | 重要な倉庫、AI が密かにテストを変更するのを恐れる | |
利用可能な最小のバージョンは L1 + L2 です。
完全な防御ラインは L1 + L2 + L3 + L4 です。
## 1. まず「完了」とは何かを定義します。
完了基準がない場合、Codex は簡単に「書かれたコード」を「完了」として扱うことができます。
TDD シナリオでは、完了基準をより具体的にする必要があります。
### 証拠はすべてのラウンドで提出されなければなりません
Codex にラウンドごとに次の 6 つの項目を報告させます。
```text
Behavior: 这一轮实现哪个行为
Test: 测试文件和测试名
Command: 跑了什么命令
RED: 失败原因是否符合预期
GREEN: 通过结果是什么
REFACTOR: 是否重构,为什么
```
これは「完了」よりもはるかに便利です。
これにより、最初に実装を記述してから合理的と思われるテストを追加するのではなく、モデルが実際に赤と緑のサイクルを通過したことがわかります。
### ラウンドで処理される動作は 1 つだけです
これは重要です。
Codex にテスト マトリックス全体を一度に生成させないでください。これは「水平敷設テスト」になります。
```text
RED: test1, test2, test3, test4, test5
GREEN: 一次写一个大实现
```
あなたが望むのは、それを縦にスライスすることです。
```text
RED test1 -> GREEN impl1 -> REFACTOR
RED test2 -> GREEN impl2 -> REFACTOR
RED test3 -> GREEN impl3 -> REFACTOR
```
最初の実装ラウンドでは、問題の理解が変わります。すべてのテストを一度に作成しないでください。
## 2. L1: AGENTS.md に規律を書き込む
`AGENTS.md` は、プロジェクトに入るときに Codex が読み取る記述ファイルです。
OpenAI の公式ドキュメントには、Codex が最初にグローバル記述を読み取り、次にプロジェクトのルート ディレクトリから現在のディレクトリまでそれを読み取ると記載されています。各層は最初に `AGENTS.override.md` を読み取り、それ以外の場合は `AGENTS.md` を読み取ります。現在のディレクトリに近い記述ほど後から表示されるため、優先度が高くなります。デフォルトのマージ制限は `32 KiB` であるため、ここに長いチュートリアルを書くことはできません。
### プロジェクトの交通ルールのようなものであるべきです
`AGENTS.md` は、Codex に TDD とは何かを教える責任はありません。このプロジェクトではどのような行為が許可されていないのかを明確に記述することのみを担当します。
この段落を直接配置できます。
```markdown
# TDD Rules
- For new behavior and bug fixes, use red/green TDD.
- RED: write exactly one failing behavior test first.
- Run the smallest relevant test command and confirm the failure is expected.
- Do not edit production implementation during RED.
- GREEN: write the minimum production code required to pass the current failing test.
- Never modify, delete, skip, or weaken tests to make implementation pass.
- REFACTOR only after tests are green.
- Keep structural changes and behavior changes separate.
- Report Behavior, Test, Command, RED, GREEN, and REFACTOR for each cycle.
```
追加のプロジェクト コマンド:
```markdown
# Verification
- Use `pytest` or the smallest relevant pytest command for Python behavior tests.
- Use `npm run types:check` only when this blog site's MDX or TypeScript changes.
- Use the smallest targeted command during RED/GREEN loops.
- If a command is slow, explain what targeted command was used first and what full command remains.
```
### 百科事典として書いてはいけません
不正な `AGENTS.md` には次の内容が入ります。
* TDDの歴史
* すべてのテスト哲学
* 大量のフレームワークチュートリアル
* 複雑なプロンプトテンプレート
* さまざまな言語での完全な仕様
これらのことは、本当に重要なルールを薄めます。
私の提案は次のとおりです: `AGENTS.md` 常駐の規律のみを配置します。長いプロセスにはスキルを使用してください。
## 3. L2: プロセスをコーデックススキルに変換する
`AGENTS.md` は「デフォルトの規律」を解決し、スキルは「完全なプロセス」を解決します。
Codex に対して「TDD で実行してください」とよく言う場合、この文をプロジェクト レベルのスキルにアップグレードする必要があります。
### ディレクトリ構造
ここに入れてください:
```text
.agents/
skills/
tdd-codex/
SKILL.md
```
Codex は、現在のディレクトリから `.agents/skills` までをスキャンします。ウェアハウスのルート ディレクトリにあるスキルは、チームが使用するワークフローに適しています。
### 利用可能な最小限の SKILL.md
```markdown
---
name: tdd-codex
description: Implementing or fixing maintainable code with Codex using strict red-green-refactor TDD. Use for new behavior, bug reproduction, behavior tests, or safe AI coding.
---
# TDD Codex Workflow
Use one behavior slice per cycle.
## Phase 0: Scope
Identify one observable behavior.
Name the public API, user flow, or integration boundary under test.
Do not edit production code.
## Phase 1: RED
Write exactly one failing behavior test.
Prefer public behavior over implementation details.
Run the smallest relevant test command.
Confirm the failure is expected.
Stop and report:
- Behavior
- Test file
- Command
- Failure reason
## Phase 2: GREEN
Write the minimum production code to pass the current failing test.
Never modify, delete, skip, or weaken tests to pass.
Do not add speculative features or abstractions.
Run the same test command.
Report the passing result.
## Phase 3: REFACTOR
Only refactor after tests are green.
If the code is already simple, skip.
If refactoring, make one structural change at a time.
Run tests after each refactor.
Do not change behavior.
## Cycle Report
Return:
- Behavior:
- Test:
- Command:
- RED:
- GREEN:
- REFACTOR:
- Next slice:
```
### 呼び出しメソッド
これからは次のように言えます。
```text
用 tdd-codex skill 做这个需求。
每轮只处理一个行为。
先 RED,确认失败后停下来,不要直接写实现。
```
またはもっと短くすると:
```text
用 tdd-codex。先写红灯,等我说 go。
```
重要なのは、プロンプトがどれほど美しいかということではなく、Codex を毎回同じ軌道に戻すということです。
## 4. L3: サブエージェントを使用して赤と緑の再構築を分離する
すべてのタスクにサブエージェントが必要なわけではありません。
ただし、タスクが複雑で、テストが実装によって汚染されやすく、リファクタリングが制御不能になりやすい場合は、RED、GREEN、および REFACTOR を異なるエージェントに分割する方が安定します。
### 解体する価値があるのはいつですか?
分解に適しています:
* 権限、課金、ステートマシン
-マルチモジュール機能
* バグは非常に隠されているため、最初に再発テストを作成する必要があります
* モデルは常にグリーンを取得するためのテストに変更されます
* テスト品質のレビューのみを担当する人が欲しい
分解には適していません:
* ガジェット機能
* コピーライティングの変更
* 純粋な視覚的な微調整
* ワンタイムスクリプト
エージェントを解体するコストは現実的です。分離によるメリットが通信のコストを上回る場合にのみ使用してください。
### RED agent
```toml
# .codex/agents/tdd-test-writer.toml
name = "tdd_test_writer"
description = "RED phase agent. Writes one failing behavior test and stops before implementation."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the RED phase.
Write exactly one behavior-focused test for the requested slice.
Prefer public APIs and user-visible behavior over implementation details.
Run the smallest relevant test command.
Confirm the test fails for the expected reason.
Do not edit production implementation.
Do not add multiple tests at once.
Return Behavior, Test, Command, and RED failure reason.
"""
```
### GREEN agent
```toml
# .codex/agents/tdd-implementer.toml
name = "tdd_implementer"
description = "GREEN phase agent. Implements the minimum production code to pass the current failing test."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the GREEN phase.
Read the failing test and relevant production code.
Write the minimum implementation required to pass the current test.
Never modify, delete, skip, or weaken tests to make them pass.
Do not add speculative features, helpers, configuration, or abstractions.
Run the relevant tests and return the command plus passing output.
"""
```
### REFACTOR agent
```toml
# .codex/agents/tdd-refactorer.toml
name = "tdd_refactorer"
description = "REFACTOR phase agent. Improves structure only after tests are green."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the REFACTOR phase.
Start by running the relevant tests to confirm the code is green.
Look for duplication, unclear names, needless branching, or misplaced responsibility.
Skip refactoring when the code is already simple.
If you refactor, make one structural change at a time.
Run tests after each refactor.
Never change behavior in this phase.
"""
```
### メインセッションをコマンドする方法
```text
按三阶段 TDD 做这个 slice:
1. tdd_test_writer 只写一个失败测试,并确认 RED。
2. 等我确认后,tdd_implementer 写最小实现,并确认 GREEN。
3. tdd_refactorer 判断是否需要结构重构。
不要批量铺测试。
不要在 GREEN 阶段修改测试。
```
ここでのポイントは、コンテキストを分離することです。テストを作成する人は、実装の詳細に影響されないようにする必要があります。実装を作成する人は手動でテストできません。リファクタリングを行う人は新しい動作を導入できません。
## 5. L4: フックを使用してテストの差分に焦点を当てる
ルールのみに依存している場合でも、モデルが範囲外になる可能性があります。
最も一般的なクロスボーダーは、テストが赤であり、モデルがテストを変更して緑に変えることです。
フックの価値は「絶対に安全」であることではなく、このアクションを即座に公開することです。
### Codex フックを有効にする
まず、設定内の機能フラグを開きます。
```toml
# ~/.codex/config.toml 或 /.codex/config.toml
[features]
codex_hooks = true
```
Codex は、アクティブな構成レイヤーの隣にあるフックを探します。一般的な場所:
* `~/.codex/hooks.json`
* `~/.codex/config.toml`
* `/.codex/hooks.json`
* `/.codex/config.toml`
プロジェクト レベルでは、ウェアハウスに従うことができるため、最初に `/.codex/hooks.json` を使用することをお勧めします。
### hooks.json
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/watch-test-edits.sh\"",
"timeout": 10,
"statusMessage": "Checking test file edits"
}
]
},
{
"matcher": "Bash|apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/run-fast-check.sh\"",
"timeout": 120,
"statusMessage": "Running fast checks"
}
]
}
]
}
}
```
### テストファイルが変更されているかどうかを確認します
```bash
# .codex/hooks/watch-test-edits.sh
#!/usr/bin/env bash
set -euo pipefail
changed_tests="$(
git diff --name-only |
grep -E '(^|/)(__tests__|tests?)/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.go$|test_.*\.py$' || true
)"
if [ -n "$changed_tests" ]; then
cat < "hello-world"
- trim leading/trailing spaces
- collapse repeated spaces into one hyphen
- remove punctuation
- normalize "Café" -> "cafe"
- empty input returns empty string
```
### ステップ 1: 最初の赤信号
```text
读取 SPEC.md。
只实现第一条行为:"Hello World" -> "hello-world"。
先 RED:只写一个失败测试,运行它,确认失败。
不要写生产实现。
```
理想的な出力:
```text
Behavior: basic title becomes lowercase hyphenated slug
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
RED: failed because slugify is not defined
```
そうして初めて続行できるのです。
### ステップ 2: 最小限の緑
```text
go
```
Codex は最小限の実装を記述します。
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
次に、次のように報告します。
```text
GREEN: pytest tests/test_slugify.py -q passed
REFACTOR: skipped, implementation is still simple
Next slice: trim leading/trailing spaces
```
### ステップ 3: 2 番目の赤信号
```text
继续下一条:去掉首尾空格。
先 RED,只写一个测试。
```
テスト:
```py
def test_slugify_trims_spaces():
assert slugify(" Hello World ") == "hello-world"
```
現在の実装が `-hello-world-` を出力する場合、赤信号は true です。
次に、緑:
```py
def slugify(text: str) -> str:
return text.strip().lower().replace(" ", "-")
```
### ステップ 4: 抽象化を急がない
この時点で、多くの AI は `normalizeInput`、`removePunctuation`、`toAscii` を描画したいと考えます。
まだ急がないでください。
TDD の設計は、想像力ではなく、テストのプレッシャーによって押し出される必要があります。 Unicode、句読点、空の文字列を追加し、構造上の圧力が実際に発生するまで待ってから、リファクタリングを行ってください。
## 8. 毎日の使用状況のクイックチェック
### 新機能
```text
用 TDD 实现这个需求。
每轮只处理一个行为。
先写一个失败测试并运行确认 RED。
不要写生产实现,直到我说 go。
```
### バグを修正
```text
先写一个能复现这个 bug 的失败测试。
确认它因为这个 bug 失败后,再写最小修复。
不要改测试来适配当前实现。
```
### 複雑な関数
```text
先不要写代码。
请给出 TDD 分解计划:
- 外圈集成测试是什么
- 内圈每个行为 slice 是什么
- 每轮用什么命令验证
- 哪些地方不能 mock
等我确认后再开始 RED。
```
### Review
```text
Review 这次改动,重点看:
- 是否先有失败测试
- 测试是否测行为而不是实现
- 是否存在为了通过而弱化测试
- 结构改动和行为改动是否混在一起
- 是否缺少外圈集成测试
```
## 9. 最終チェックリスト
Codex に TDD を依頼するたびに、この表を使用して最後に確認してください。
| 質問 | 資格基準 |
| ---------------------------- | ------------------------- |
| まず本当に人気があるのか | 失敗したコマンドと失敗の理由はありますか。 |
| 赤い色は正しいですか | 失敗の理由は、ターゲット動作の欠如に対応します。 |
| ラウンドごとに 1 つの動作だけを行う | バッチテストなし |
| 緑 テストは変更されましたか? | 緑色になるように変更されたテストはありません。 |
| テストは動作をテストしますか?内部実装の詳細に依存しない | |
| 混合動作はリファクタリングされましたか | 構造的変化と行動的変化を分けて考える |
| 完全な検証は行われていますか | 対象のテストと必要な完全な検査が実行されています。 |
このテーブルが通過できない場合は、急いでマージしないでください。
## 推奨リソース
# AI가 "공부 없이도 시험을 통과"하게 해준다면, 대학에는 무엇이 남는가?
Anthropic은 최근 프린스턴, 버클리, 런던정치경제대학(LSE), 애리조나주립대학교 출신 학생 네 명을 초청해 캠퍼스 내 AI의 실제 현황에 대한 인터뷰를 진행했습니다. 약 40분간 이어진 이 인터뷰에는 홍보성 문구가 없었습니다. 오직 진솔한 혼란과 불안, 그리고 깊은 사유만이 담겨 있었습니다.
이 인터뷰는 보다 근본적인 문제를 드러냅니다. **AI는 단순히 학습 방식을 바꾸는 것이 아니라, 교육 체계 전체의 기저 논리를 해체하고 있습니다.**
***
## 외면할 수 없는 현실
인터뷰 시작 부분에서 진행자는 직접적인 질문을 던집니다. 현재 캠퍼스에서 AI에 대한 분위기는 어떠한가?
답은 명확했습니다. **학생의 90%가 AI를 사용하고 있습니다.** 가끔 사용하는 것이 아니라, 일상적인 워크플로의 일부로서 말입니다. 강의 노트 요약, 문제지 풀기, 과제 피드백 받기, 비즈니스 케이스 분석, 시장 조사, 재무 연구 수행 등 다양한 용도로 활용되고 있습니다. 심지어 일부 학생들은 퀴즈를 완성하는 데도 AI를 사용하는데, 그 이유는 매우 현실적입니다. 대학원생으로서 여러 개의 아르바이트를 병행할 때는 항상 충분한 시간이 있지 않기 때문입니다.
더욱 흥미로운 점은, 거의 모든 학생이 AI를 사용하고 있음에도 불구하고 아무도 규칙이 무엇인지 명확히 알지 못한다는 사실입니다. 일부 수업은 AI 사용을 명시적으로 금지하고, 일부는 적극적으로 장려하며, 대부분은 모호한 회색 지대에 있습니다. 학생들은 경계가 어디에 있는지 모르고, 교수들도 어떻게 관리해야 할지 모릅니다. 이러한 상태를 "회색 지대"라고 합니다. 사용하고 싶지만 규정 위반이 두렵고, 사용하지 않으면 뒤처지는 것 같은 느낌이 드는 상황입니다.
이 회색 지대에서 가장 위험한 점은 학생들이 규정을 위반할 수 있다는 것이 아닙니다. **진정으로 가치 있는 논의가 이루어지지 못하도록 막는다는 것입니다.** 학생들은 AI 활용 모범 사례를 공개적으로 공유할 수 없고, 교수들은 학생들에게 책임감 있는 도구 사용법을 지도할 수 없습니다. 학술 공동체 전체가 표면적으로는 금지하면서 이면에서는 사용하는 어색한 상태에 빠져 있습니다.
규칙이 효과적으로 집행될 수 없을 때, 그것은 "잘 위장하는 사람"과 "그렇지 못한 사람"을 가르는 도구로 전락합니다. 명시적으로 금지하면서도 이면에서 광범위하게 사용되는 상황은 **학생들의 AI 사용을 막는 것이 아니라, 더 나은 사용 방법에 대한 공개적인 논의를 막을 뿐입니다.**
***
## AI는 거울이다
인터뷰 중 한 가지 통찰이 특히 날카롭게 와닿았습니다. 인공지능, 특히 학생들이 인공지능을 어떻게 사용하는가는 그 동기를 매우 잘 드러낸다는 것입니다.
이 발언의 핵심 통찰은 이것입니다. **AI는 거울이 되어, 당신이 대학에 다니는 진짜 목적을 비춥니다.**
인터뷰에서는 대학의 목표를 세 가지로 분류했습니다. 첫째, 전공 지식을 깊이 학습하고 특정 분야에 대한 심층적인 이해를 습득하는 것. 둘째, 취업을 준비하고 좋은 직장을 얻으며 직업 네트워크를 구축하는 것. 셋째, 인맥을 넓히고 사교 생활을 즐기며 대학 문화를 경험하는 것. 각 학생마다 이 세 가지 목표의 비중이 다르며, AI 사용 방식은 그 비중을 정확하게 드러냅니다.
"시험 통과"와 "학위 취득"만을 목표로 한다면, AI의 출력을 그대로 과제로 제출할 것입니다. 이것은 도덕적 판단이 아니라 현실입니다. 기술이 최소한의 비용으로 목표를 달성할 수 있게 해준다면, 굳이 먼 길을 돌아갈 이유가 있을까요? 학위를 취득하고 취업하는 것이 목표라면, AI로 과제를 완성하는 것은 완전히 합리적인 선택입니다.
하지만 진정으로 배우고 싶고 특정 분야를 깊이 이해하고 싶다면, AI를 대화 파트너로 활용할 것입니다. AI에게 질문을 던지고, 개념을 설명해 달라고 요청하고, 자신의 말로 다시 표현합니다. AI가 초안 코드를 작성하게 한 뒤 직접 리팩토링하고 최적화합니다. 그리고 모든 단계에서 자신이 무슨 일이 일어나고 있는지 진정으로 이해하고 있는지 확인합니다.
이러한 분화는 학생들 간에만 존재하는 것이 아니라, 전공 간에도 나타납니다. 인문계 학생들은 종종 AI를 멀리하는 경향이 있습니다. 그들의 학습은 정독, 즉 원문을 꼼꼼히 읽고, 언어의 세부 사항을 음미하며, 저자의 의도를 이해하는 과정을 필요로 하기 때문입니다. AI는 이 과정을 방해합니다. AI가 제공하는 것은 요약과 번역이지, 원문을 직접 경험하는 것이 아니기 때문입니다. 반면 공학과 경영학 학생들은 AI를 대량으로 활용합니다. AI가 기술적 장벽을 낮춰 컴퓨터 배경지식이 없는 사람도 코드를 작성하고, 웹사이트를 구축하고, 데이터를 분석할 수 있게 해주기 때문입니다.
**이러한 양극화는 본질적으로 "무엇이 배울 가치가 있는가"에 대한 서로 다른 이해를 반영합니다.** 인문계 학생들에게는 셰익스피어 원문을 읽는 경험 자체가 학습입니다. 공학 학생들에게는 문제를 해결할 수 있는가 여부가 중요하지, 모든 코드를 직접 작성하는 것이 중요한 것이 아닙니다. AI는 이러한 차이를 더욱 뚜렷하게 만들고 있습니다.
***
## 도구인가 목발인가? 간단한 판단 기준
인터뷰에서 핵심적인 질문이 제기되었습니다. AI가 도구인지 목발인지를 어떻게 구분할 수 있는가?
학생들의 답변은 놀라울 정도로 일치했습니다. **설명할 수 있는가의 여부입니다.**
자신이 만든 것을 설명할 수 없고, AI가 그 과정에서 어떤 역할을 했는지 말할 수 없다면, 그것은 목발입니다. 초등학생에게 설명하듯 명확하게 설명할 수 있고, 저수준 및 고수준 설명을 모두 제시할 수 있다면, 그것은 도구입니다.
이 기준은 단순해 보이지만, 학습의 본질을 건드립니다. 파인만 학습법의 핵심 논리는 이것입니다. 어떤 개념을 쉬운 언어로 설명할 수 없다면, 아직 그것을 진정으로 이해하지 못한 것입니다. AI 시대의 학습도 마찬가지입니다. AI가 해준 일을 설명할 수 없다면, "사고를 외주화"하는 것이지 "사고를 강화"하는 것이 아닙니다.
인터뷰에서 흥미로운 사례가 소개되었습니다. 어떤 학생은 강의 슬라이드를 입력하면 AI가 각 슬라이드 옆에 교수의 주석과 유사한 설명을 생성해 주는 도구를 개발했습니다. 이 학생은 이렇게 말했습니다. "이것이 유용한 이유는 제가 AI에게 제가 알고 싶은 것이 무엇인지 이미 프롬프트로 알려줬기 때문입니다. 슬라이드의 특정 내용에 대한 정의 같은 것들 말이죠. 슬라이드는 때로 추상적이고 맥락이 부족해서 옆에 컨텍스트를 추가해야 합니다."
핵심은 "제가 AI에게 제가 알고 싶은 것을 이미 알려줬다"는 부분입니다. 이 학생은 자신의 지식 공백이 어디에 있는지 알고, 어떤 도움이 필요한지 알며, 그 도움을 제공하도록 AI를 능동적으로 이끌었습니다. 이것이 도구입니다. 반면 슬라이드를 그냥 AI에게 던져 "이 수업을 요약해줘"라고 한 뒤 그대로 암기한다면, 그것은 목발입니다.
**차이는 능동성과 이해에 있습니다.** 도구를 사용하는 사람은 자신이 무엇을 하고 있는지 알고, 전체 과정을 주도합니다. 목발을 사용하는 사람은 기술에 통제권을 넘기고 스스로는 수동적인 수신자가 됩니다.
***
## 학교의 지체는 대응이 느린 것이 아니라, 본질적으로 대응할 수 없는 것이다
인터뷰에서는 몇몇 학교의 시도도 언급되었습니다. LSE의 한 필수 과목은 학생들에게 Claude 사용법을 지도하기 시작했습니다. Claude와 대화하고, 다양한 역할을 부여하고, 학생들이 AI와 어떻게 상호작용하는지 확인하기 위해 대화 기록을 제출하도록 요구합니다. 애리조나주립대학교의 커리어 관리 센터는 다양한 시나리오에 대한 프롬프트 템플릿을 제공하는 프롬프트 라이브러리를 구축했습니다. 이러한 시도들은 좋은 방향입니다. 핵심 아이디어는 AI를 금지하는 것이 아니라, 학생들에게 책임감 있는 사용법을 가르치는 것입니다.
하지만 이것은 소수의 사례에 불과합니다. 대부분의 학교는 여전히 "학생들이 AI를 사용해도 되는가"를 논의 중입니다. 일부 교수는 사용해도 되지만 과제에 사용 방식을 명기해야 한다고 하고, 일부 수업은 직접 금지하며, 또 일부는 아예 언급하지 않고 학생들이 사용하지 않을 것이라고 가정합니다. 통합적인 프레임워크도 없고, 통일된 기준도 없으며, 전체 체계가 혼란 상태에 있습니다.
더 깊은 문제는 이것입니다. **이 문제는 본질적으로 규정과 제도만으로는 해결할 수 없습니다.**
전통적인 교육의 관리 논리는 이러합니다. 학교가 규칙을 설정하고, 학생은 규칙을 따르며, 위반자는 처벌을 받습니다. 이 논리가 성립하는 전제는 "위반 행위가 탐지될 수 있다"는 것입니다. 그러나 AI는 이 전제를 무너뜨렸습니다.
학생들이 과제를 제출할 때 AI를 사용하는 것을 금지할 수 있지만, 학생들이 사고 과정에서 AI를 사용했는지 감시할 수는 없습니다. AI 탐지 도구를 사용할 수 있지만, 이러한 도구의 정확도는 처벌 근거로 사용할 수 있는 수준에 훨씬 못 미칩니다. 오판율이 너무 높고, 학생들은 금방 탐지를 우회하는 방법을 배웁니다. 더 중요한 것은, **근본적으로 "AI 도움으로 완성된 고품질 과제"와 "학생이 독립적으로 완성한 고품질 과제"를 구별할 수 없다는 점입니다.** 좋은 AI 활용은 원래 원활하게 이루어져야 하기 때문입니다.
인터뷰에서 이런 말이 직접적으로 나왔습니다. 근본적으로, 어떤 규정도 학생들이 AI를 사용하는 방식을 바꿀 수 없습니다. 책임은 학생의 손에 있습니다. 이것은 책임을 회피하는 것이 아니라 현실입니다.
기술이 "공부하지 않고도 시험을 통과"하는 것을 가능하게 만들 때, 학교가 직면하는 것은 관리 문제가 아니라 존재론적 문제입니다. **시험이 학습을 증명할 수 없다면, 학교의 존재 의미는 무엇인가?**
***
## 교육 체계의 기저 논리가 무너졌다
이 문제는 교육 체계의 근본적인 모순을 건드립니다.
전통적인 교육은 몇 가지 핵심 가정 위에 세워져 있습니다. 첫째, 지식은 희소하며 전달하기 위해 전문 기관(학교)과 전문 인력(교수)이 필요합니다. 둘째, 학습 성과는 시험으로 측정할 수 있습니다. 셋째, 학위는 특정 분야의 지식을 습득했음을 증명하므로, 관련 직업에 종사할 자격이 있습니다.
AI는 이러한 가정들을 하나씩 무너뜨리고 있습니다.
**지식은 더 이상 희소하지 않습니다.** YouTube에는 무료 Stanford 강의가 있고, Claude는 언제든지 질문에 답해주며, GitHub에는 학습할 수 있는 무수한 오픈소스 프로젝트가 있습니다. 학교에 가지 않아도 지식에 접근할 수 있고, 심지어 유료 구독 없이도 기본적인 AI 튜터링을 받을 수 있습니다.
**시험은 학습을 측정할 수 없습니다.** AI가 대부분의 시험 문제를 풀 수 있을 때, 시험은 "이해도를 측정하는 도구"에서 "AI를 얼마나 잘 사용하는지 측정하는 도구"로 변합니다. 이것은 시험이 완전히 무용하다는 말이 아니라, 더 이상 "진정으로 이해한 사람"과 "도구를 잘 활용하는 사람"을 정확히 구별할 수 없다는 말입니다.
**학위의 가치가 하락하고 있습니다.** 고용주들이 학위가 지원자가 관련 지식을 실제로 갖추고 있음을 보장하지 못한다는 것을 깨달을수록, 실질적인 역량의 증거인 포트폴리오, 프로젝트 경험, 인턴십 성과를 더욱 중시하게 됩니다. 학위는 "역량 증명"에서 "기본 조건"으로 격하됩니다.
**이러한 가정들이 무너질 때, 교육 체계의 가치 제안은 새롭게 정의되어야 합니다.**
인터뷰는 하나의 답을 제시했습니다. 대학의 가치는 "**지식 전달**"에서 "**환경 제공**"으로 전환됩니다. 이 환경은 실수를 허용하고, 탐색을 허용하며, 다른 사람들과 아이디어를 충돌시킬 수 있는 공간입니다. 룸메이트와 주말을 보내며 "졸업 전 버킷리스트"를 만들고, hackathon에서 "어쩌면 어리석을 수도 있는" 아이디어를 테스트하고, 교수와 논쟁하고, 동료와 토론하며, 커리어 리스크 없이 실패에서 배울 수 있는 환경 말입니다.
인터뷰에서 언급된 학생 프로젝트들이 이를 잘 보여줍니다. "자동 수강신청 알림", "빈 강의실 찾기", "졸업 전 소원 순위표" 같은 것들입니다. 기술적으로는 모두 복잡하지 않으며, 많은 창작자는 컴퓨터 배경지식조차 없습니다. 하지만 이것들은 진실된 인간적 감정에서 비롯됩니다. 놓치는 것에 대한 두려움, 편의에 대한 욕구, 대학 생활에 대한 소중함.
**기술 장벽이 낮아진 후, 중요한 것은 더 이상 "코드를 쓸 줄 아는가"가 아니라 "어떤 문제를 해결하고 싶은가"입니다.** 대학이 제공하는 것은 이러한 문제들을 자유롭게 탐색할 수 있는 공간, 아이디어를 현실로 만들고 실패에서 배울 수 있는 환경입니다.
AI는 과제를 완성하는 데 도움을 줄 수 있지만, "실수하고 탐색할 수 있는" 이 시간을 대신 보내줄 수는 없습니다.
***
## 취업 시장의 역설
인터뷰 후반부에서는 취업에 대해 이야기하며 또 다른 역설을 드러냈습니다.
학생들은 AI로 이력서를 작성하고, 기업은 AI로 이력서를 걸러냅니다. 전체 채용 과정이 화면을 향해 이야기하는 것으로 변했습니다. AI를 향해 자기소개서를 쓰고, 카메라를 향해 질문에 답하고, 마지막에는 AI가 생성한 거절 통보를 받습니다. 이력서를 제출하고 거절 통보를 받기까지 15분밖에 걸리지 않을 수도 있습니다. 효율은 매우 높지만, 인간적인 면은 거의 없습니다.
이것은 "AI 대 AI"의 취업 시장을 만들어냈습니다. 학생들은 AI가 "좋은" 이력서를 작성하도록 훈련하고, 기업들은 AI가 "좋은" 지원자를 선별하도록 훈련합니다. 이 과정에서 진짜 인간의 역할은 점점 줄어듭니다. 화면을 향해 이야기하는 것에는 화학 반응이 없으며, 수치화하기 어렵지만 중요한 자질들, 즉 유머 감각, 임기응변 능력, 팀 협업의 미묘한 점들을 보여줄 수 없습니다.
그러나 역설의 다른 면은 이것입니다. AI 능숙도 자체가 새로운 경쟁력이 되었습니다. 빅4 컨설팅 회사들은 예전에는 generalist MBA를 채용했지만, 이제는 AI 역량을 갖춘 MBA를 전문적으로 찾습니다. AI를 다양한 산업에 적용하는 방법을 안다면, 당신이 그들의 최우선 후보입니다.
**역설은 이것입니다. AI는 취업 시장을 더 차갑게 만들면서도, 동시에 AI 역량을 더욱 중시하게 만들었습니다. 피할 수 없으니, 효과적으로 사용하는 방법을 배워야 합니다.**
이것은 다시 그 핵심 질문으로 돌아옵니다. "효과적인 사용"이란 무엇인가? ChatGPT로 이메일을 작성할 줄 아는 것이 아닙니다. 어떤 문제에 AI를 적용하는 것이 적합한지 파악하고, 필요한 결과를 얻기 위해 AI를 이끄는 프롬프트를 설계하고, AI 출력의 품질을 평가하고 필요한 수정을 가할 수 있는 능력입니다.
이러한 능력은 AI를 금지함으로써 키워지는 것이 아니라, 대량의 실천과 시행착오를 통해 길러집니다. 그렇기에 AI를 적극적으로 수용하고, Claude Builder Club을 만들고, hackathon을 조직하는 학교들이 학생들에게 더 가치 있는 교육을 제공하고 있습니다. 그들은 상대적으로 안전한 환경에서 학생들이 AI와 협력하는 방법을 배우게 합니다.
***
## 책임의 전이
인터뷰에서 가장 핵심적인 통찰은 아마도 이것일 것입니다. 기술이 "공부하지 않고도 시험을 통과"하게 해줄 때, 배움 자체의 의미는 각자가 스스로 대답해야 하는 문제가 됩니다.
이것은 책임의 전이입니다. **학교에서 학생으로, 규칙에서 자각으로, 외부 동기에서 내부 동기로.**
전통적인 교육의 동기는 외부적입니다. 학위를 얻기 위해 시험을 통과해야 하고, 취업을 위해 학위가 필요합니다. 이 외부 인센티브 체계가 학생들의 학습을 추동합니다. 하지만 AI가 공부하지 않고도 시험을 통과할 수 있게 해줄 때, 이 인센티브 체계는 효력을 잃습니다.
남는 것은 내부 동기뿐입니다. 진정으로 배우고 싶습니까? 이 분야에 진정으로 관심이 있습니까? 진정으로 깊이 이해하고 싶습니까, 아니면 그냥 학위만 원합니까?
인터뷰에서 흥미로운 세부 사항이 있었습니다. 어떤 학생은 대학원 생활의 좋지 않은 면을 언급했습니다. 여러 개의 아르바이트를 병행하다 보니 시간이 없어 AI로 빠르게 퀴즈를 완성할 때가 있다는 것입니다. 하지만 그는 이렇게 덧붙였습니다. 대학원은 비판적 사고를 넓히고 더 결단력 있는 모습을 보여야 하는 시기여야 한다고요. **그는 모순을 인식했지만, 효율을 선택했습니다.**
이것은 도덕적 판단이 아닙니다. 현실적인 압박 하에서 효율은 종종 이상보다 더 중요합니다. 하지만 이 선택은 하나의 사실을 드러냅니다. **외부 압박(퀴즈 완성)과 내부 동기(깊이 있는 학습)가 충돌할 때, 많은 사람들은 전자를 선택합니다.**
AI는 이러한 충돌을 더욱 첨예하게 만들었습니다. AI가 "시험 준비"를 극도로 쉽게 만들었기 때문입니다. AI가 없던 시대에는, 시험을 그냥 통과하고 싶어도 최소한 무언가를 배워야 했습니다. AI는 이 중간 단계를 제거했습니다. 전혀 학습하지 않고도 통과할 수 있게 된 것입니다.
**이것은 모든 사람이 그 질문에 직면하게 만듭니다. 당신은 왜 대학에 다니는가?**
답이 "학위를 얻고 취업하기 위해서"라면, AI로 과제를 완성하는 것은 완전히 합리적입니다. 답이 "이 분야를 진정으로 배우고 싶어서"라면, AI가 제공하는 지름길의 유혹에 능동적으로 저항해야 합니다.
학교는 이 선택을 대신 해줄 수 없습니다. 규칙은 내부 동기를 강제할 수 없습니다. **이것은 당신 자신의 책임입니다.**
***
## 기술은 당신이 준비될 때까지 기다리지 않는다
인터뷰 마지막 부분에는 하나의 태도가 일관되게 흐릅니다. "우리가 방법을 찾아낼 것이다."
학교 규칙이 따라오지 못하고 있습니까? 일단 사용하고, 무엇이 효과적인지 학교에 알려주면 됩니다. AI가 부정행위에 사용될 수 있습니까? 천천히 책임감 있는 사용법을 배우면 됩니다. 취업 시장이 변했습니까? 새로운 게임 규칙에 적응하면 됩니다.
이것은 맹목적인 낙관주의가 아니라 현실주의입니다. 기술은 이미 여기 있으며, 당신이 준비될 때까지 기다렸다가 세상을 바꾸지 않습니다. 저항을 선택할 수도 있고, 적응을 선택할 수도 있습니다. 하지만 시간을 멈추는 것은 선택할 수 없습니다.
**이 세대 학생들과 AI의 관계는 공포도 아니고 맹목적인 수용도 아닙니다. 혼란 속에서 더듬어 나아가고, 시행착오 속에서 배우는 것입니다.**
그들이 Claude Builder Club에서 만드는 프로젝트들은 기술적으로는 복잡하지 않지만, 진짜 문제를 해결합니다. hackathon에서 시도하는 아이디어들은 어리석을 수도 있지만, 적어도 도전하고 있습니다. 수업에서의 혼란, 규칙이 불분명하지만, 적어도 생각하고 있습니다.
이러한 "**행동하면서 배우는**" 자세는, 어떤 규정보다도 AI 시대에 적응하는 데 더 큰 도움이 될 것입니다.
Kevin Kelly는 《기술의 충동(What Technology Wants)》에서 "technium" 개념을 제시했습니다. 기술은 하나의 전체로서 마치 자신의 의지가 있는 것처럼, 점점 더 강력해지려는 것 같다는 것입니다. 인터뷰의 한 관찰이 이와 잘 호응합니다. 지난 2년 동안, AI가 계속 발전하기 위해 필요한 것이라면 무엇이든 얻을 수 있었습니다. 원자력에 대한 태도 변화, 우주 데이터 센터에 대한 논의 등, AI의 발전을 늦출 수 있는 어떤 병목이 있을 때마다 그 장애물은 제거되었습니다.
하지만 더 정확한 말은 이것일지도 모릅니다. **학생들에게 필요한 것은 창조될 것입니다.**
AI는 단지 도구입니다. 미래를 결정하는 것은, 이 세대 학생들이 이 도구를 어떻게 사용하기로 선택하는가입니다. 사고를 회피하는 데 쓸 것인가, 사고를 강화하는 데 쓸 것인가. 시험을 통과하는 데 쓸 것인가, 세계를 탐색하는 데 쓸 것인가. 목발로 쓸 것인가, 도구로 쓸 것인가.
이 선택은 학교가 관리할 수 없습니다. 오직 그들 스스로 결정할 수 있을 뿐입니다.
그리고 이 인터뷰를 보건대, 적어도 일부 학생들은 이 문제를 진지하게 고민하고 있습니다. 그것으로 충분할지도 모릅니다.
# Agent처럼 생각하기:Claude Code 팀의 도구 설계 철학
Thariq은 Anthropic 엔지니어이자 Claude Code의 핵심 개발자 중 한 명입니다. 2월 말 그는 X에 장문의 글을 게시하며, Claude Code를 구축하는 과정에서 겪은 agent 도구 설계에 관한 다섯 가지 실제 사례를 공유했습니다——이론적 프레임워크가 아니라, 현장에서 직접 부딪히며 얻은 경험입니다. 이 글은 351만 회 조회, 209개의 답글, 9,691개의 좋아요를 기록하며 수많은 고품질 논의를 이끌어냈습니다.
글 전체의 핵심 질문을 이끌어내기 위해, Thariq은 매우 적절한 비유를 사용합니다. 어려운 수학 문제를 풀 때, 어떤 도구가 필요할까요?
**종이와 펜**은 최소한의 구성——하지만 수작업 계산에 묶이게 됩니다. **계산기**는 더 낫습니다——하지만 고급 기능 사용법을 알아야 합니다. **컴퓨터**가 가장 강력합니다——하지만 코드를 작성할 줄 알아야 합니다.
도구의 선택은 사용자의 능력에 달려 있습니다. 프로그래밍을 모르는 사람에게 컴퓨터를 주는 것보다는 계산기를 주는 편이 낫습니다. 프로그래머에게 계산기를 주면 오히려 그의 능력을 제한하게 됩니다.
agent의 경우도 마찬가지입니다. 문제는 "어떤 도구가 가장 강력한가"가 아니라, "어떤 도구가 모델의 현재 능력에 가장 잘 맞는가"입니다. 이 글이 공유하는 것은, Claude Code 팀이 그 일치점을 찾아가는 과정에서 겪었던 시행착오입니다.
다음은 원문과 커뮤니티 토론에서 제가 정리한 네 가지 점진적 주제입니다.
***
## 적을수록 많다——도구 수에 관한 역발상
직관적으로는 도구가 많을수록 agent의 능력이 강해진다고 생각하기 쉽습니다. 하지만 Claude Code 팀의 경험은 정반대였습니다.
Claude Code가 현재 보유한 도구는 약 20개에 불과하며, 팀은 이 모든 도구가 실제로 필요한지 지속적으로 검토하고 있습니다. 새 도구를 추가하는 기준은 높습니다. 이는 모델이 고려해야 할 선택지를 하나 더 늘리는 것을 의미하기 때문입니다——도구가 늘어날수록, 모델이 의사결정을 내릴 때 "인지적 부담"이 커집니다.
이 발견은 커뮤니티에서 강한 공감을 불러일으켰습니다. leon.M은 자신의 실무 경험을 공유했습니다.
Apple의 CodeAct 연구는 이 관점에 정량적 근거를 제공합니다. **단일 코드 실행 프리미티브(code execution primitive)는 복잡한 태스크에서 방대한 전용 도구 세트보다 최대 20% 높은 성능을 보입니다.** 적을수록, 실제로 더 많을 수 있습니다.
AskUserQuestion 도구의 세 차례 반복은 이 원칙을 가장 잘 보여주는 사례입니다. Claude Code 팀은 Claude가 사용자에게 질문하는 능력을 향상시키고 싶었습니다——Claude가 일반 텍스트로 질문할 수는 있었지만, 그 질문에 답하는 것이 번거롭게 느껴졌습니다. 마찰을 어떻게 줄일 수 있을까요?
**첫 번째 시도**: ExitPlanTool에 파라미터를 추가해, 계획을 출력하는 동시에 관련 질문들을 함께 내보내도록 했습니다. 결과——Claude가 혼란스러워했습니다. 계획과 그 계획에 대한 질문을 동시에 출력하도록 요구하면, 사용자의 답변이 계획과 충돌할 경우 어떻게 처리해야 할까요?
**두 번째 시도**: 출력 지침을 수정해 Claude가 특정 markdown 형식으로 질문하고, 프론트엔드에서 이를 파싱해 표시하도록 했습니다. 결과——불안정했습니다. Claude는 불필요한 문장을 추가하거나, 선택지를 빠뜨리거나, 전혀 다른 형식을 사용하는 경우가 있었습니다.
**세 번째 시도**: 독립적인 AskUserQuestion 도구를 만들었습니다. Claude는 언제든 이를 호출할 수 있으며, 호출하면 팝업 창에 질문이 표시되고 사용자가 답변할 때까지 agent 루프가 블로킹됩니다. **성공했습니다.**
Thariq은 원문에서 흥미로운 문장을 남겼습니다.
> Claude는 이 도구를 기꺼이 호출하는 것처럼 보였으며, 출력 결과도 좋다는 것을 확인했습니다. 아무리 잘 설계된 도구라도 Claude가 어떻게 호출하는지 이해하지 못한다면 작동하지 않습니다.
phuong은 예리하게 물었습니다. "'Claude가 이 도구를 기꺼이 호출하는 것처럼 보였다'는 문장이 여기서 가장 흥미롭고도 설명이 부족한 부분입니다. 모델이 도구에 대해 갖는 '선호도'를 어떻게 감지하셨나요——transcript를 읽는 방식으로요, 아니면 내부 호출 빈도 지표로요? 이 휴리스틱이 좀 더 구체화된다면, agent 도구 설계 방식이 완전히 달라질 것 같습니다."
이 질문은 끝내 답을 얻지 못했지만, 중요한 직관을 가리키고 있습니다. **도구 설계의 성공 기준은 "사람이 합리적이라고 느끼는가"가 아니라, "모델이 사용법을 이해하고, 실제로 사용하려 하는가"입니다.**
Emeka는 엔터프라이즈 도구 개발의 관점에서 매우 직관적인 교훈을 덧붙였습니다. "예전에 agent용 도구를 만들 때 모든 입력과 출력을 통제하려 했는데, 모델이 그냥……우회해버렸습니다. 엔지니어링 공수를 아끼세요. 모델이 모호함을 처리할 수 있다고 신뢰하세요."
***
## 도구는 낡아진다——능력 향상 후의 족쇄 효과
"적을수록 많다"가 공간적 차원의 인식에 관한 것이라면, "도구는 낡아진다"는 시간적 차원의 교훈입니다.
**한때 모델을 도왔던 도구가, 모델이 발전하면서 오히려 제약이 될 수 있습니다.**
Claude Code 초기 출시 당시, 팀은 모델이 방향성을 유지하려면 할 일 목록이 필요하다는 것을 알았습니다——시작 시에 할 일을 기록하고, 작업이 완료되면 하나씩 체크해 나가는 방식이었습니다. 이를 위해 Claude에게 TodoWrite 도구를 제공했습니다. 그럼에도 Claude는 자신이 무엇을 해야 하는지 자주 잊었습니다.
팀의 대응은 5턴마다 시스템 알림을 삽입해 Claude에게 목표를 상기시키는 것이었습니다.
하지만 모델이 발전하면서 문제는 역전되었습니다. 모델은 할 일 목록 알림이 더 이상 필요 없었을 뿐만 아니라, 오히려 그 알림을 제약으로 느끼기 시작했습니다. **반복적으로 할 일 목록을 상기시키자, Claude는 목록을 유연하게 조정하는 것이 아니라 엄격하게 따라야 한다고 느끼게 되었습니다.** 동시에, Opus 4.5는 서브 agent 활용 면에서 크게 향상되었지만, 서브 agent 간에 공유되는 할 일 목록을 어떻게 조율할 것인가라는 문제가 생겼습니다.
그래서 팀은 TodoWrite를 Task Tool로 교체했습니다. 두 도구의 차이는 본질적입니다. Todos의 역할은 모델이 궤도를 유지하도록 하는 것——마치 상사가 직원의 업무 목록을 감시하는 것과 같습니다. Tasks는 agent 간 소통을 지원하는 데 중점을 둡니다——팀의 협업 칸반 보드와 같습니다. Tasks는 의존 관계를 지원하고, 서브 agent 간에 업데이트를 공유하며, 모델이 이를 수정하거나 삭제할 수도 있습니다.
**TodoWrite → 5턴마다 알림 → Task Tool, 이것은 세 번의 재설계입니다.** 이전 설계가 "잘못되었기" 때문이 아니라, 모델이 이미 성장했기 때문입니다.
modi의 댓글이 이 현상을 정확하게 요약합니다.
이것은 흥미로운 논쟁도 불러일으켰습니다. Danny Cosson은 AskUserQuestion이 "실은 나쁜 설계"라고 생각합니다——Claude Code의 매우 우아한 텍스트 입력/텍스트 출력 모델을 특정 인터랙션 패턴에 억지로 끼워 맞춘 것이며, 실질적인 이점도 거의 없다는 주장입니다.
하지만 @Toong의 반론은 설득력이 있습니다. AskUserQuestion의 가치는 두 가지에 있습니다——**주도권**(모델에게 질문할 권한이 있다는 것을 명확히 전달하는 것)과 **상태 관리**(일반적인 텍스트 출력과 "사용자 개입을 기다리는" 블로킹 상태를 명확하게 구분하는 것).
같은 도구에 대해 완전히 다른 두 가지 평가가 존재합니다. 이것은 Thariq이 글 말미에서 말한 것을 정확히 증명합니다. **도구 설계는 과학이 아니라 예술입니다.** 사용하는 모델, agent의 목표, 그리고 그것이 놓인 환경에 따라 달라집니다.
***
## Agent 스스로 답을 찾게 하라——"제공"에서 "자율 탐색"으로
이 부분이 전체 글에서 제가 가장 실용적 가치가 높다고 생각하는 내용입니다.
Claude Code는 처음에 RAG 벡터 데이터베이스를 사용해 Claude에게 컨텍스트를 제공했습니다. RAG는 강력하고 빠르지만 두 가지 문제가 있었습니다. 하나는 인덱싱과 설정이 필요하며 환경에 따라 취약할 수 있다는 것, 다른 하나는 더 근본적인 문제——**이 방식은 컨텍스트를 Claude에게 제공하는 것이지, Claude가 스스로 찾아가는 것이 아니라는 점**입니다.
팀은 중요한 전환을 했습니다. Claude가 웹을 검색할 수 있다면, 코드베이스도 검색할 수 있지 않을까요? Claude에게 Grep 도구를 제공함으로써, Claude가 스스로 파일을 검색하고 컨텍스트를 구성하도록 했습니다.
Beacon의 댓글이 핵심을 찌릅니다.
**1년이라는 시간 동안, Claude는 거의 자율적으로 컨텍스트를 구성하지 못하던 상태에서, 여러 레이어에 걸쳐 중첩 검색을 수행하며 필요한 컨텍스트를 정확히 찾아내는 수준으로 진화했습니다.** 이 진화의 핵심은 Claude에게 더 많은 정보를 주는 것이 아니라, 더 나은 검색 능력을 주는 것이었습니다.
Claude Code에 Agent Skills가 도입되면서, 팀은 공식적으로 **점진적 공개(Progressive Disclosure)** 개념을 제시했습니다. agent가 탐색을 통해 관련 컨텍스트를 단계적으로 발견할 수 있도록 하는 방식입니다.
구체적인 구현은 우아합니다. Claude는 skill 파일을 읽을 수 있고, 그 파일들은 또 다른 파일을 참조할 수 있으며, 모델은 이를 재귀적으로 읽어나갈 수 있습니다. skill의 일반적인 용도는 Claude에게 더 많은 검색 능력을 추가하는 것입니다——예를 들어 API 사용법이나 데이터베이스 쿼리 방법에 관한 지침을 제공하는 것입니다.
Brian Wagner는 Claude Code의 사고방식과 완전히 일치하는 3계층 실천 사례를 공유했습니다.
> SKILL.md는 간결하게 유지하고(약 100줄), 무거운 컨텍스트는 Claude가 필요할 때 발견하는 파일에 둡니다. 저는 이것을 3번째 레이어라고 부릅니다. 여러분은 점진적 공개라고 부르시는군요. 본질은 같습니다.
PrimeLine은 정량화된 3계층 시스템을 구축하기도 했습니다. JSON 설정(약 500 token) > 스키마 요약(약 300 token) > 전체 markdown(약 3K token). 컨텍스트 라우터가 태스크 키워드에 따라 어느 계층을 로드할지 결정합니다.
이 계층화 전략의 핵심 논리는 이것입니다. **컨텍스트는 유한한 자원이며, 한계 수익은 체감합니다.** 모든 정보를 한꺼번에 agent에게 쏟아 넣으면 token을 낭비할 뿐만 아니라, 진정으로 중요한 정보가 희석됩니다. 필요에 따라 제공하고, agent 스스로 더 깊은 세부 사항이 필요한 시점을 판단하게 하는 것——이것이 확장 가능한 방식입니다.
Claude Code Guide 서브 agent는 점진적 공개의 또 다른 영리한 응용 사례입니다. 팀은 Claude가 Claude Code 자체를 사용하는 방법에 대해 충분히 알지 못한다는 것을 발견했습니다——MCP를 추가하는 방법이나 특정 슬래시 명령의 역할을 물어봐도 답하지 못했습니다.
모든 정보를 시스템 프롬프트에 집어넣을 수도 있었지만, 사용자가 이런 질문을 하는 경우는 드물기 때문에 컨텍스트 오염이 생기고, Claude Code의 주된 업무인 코드 작성에 방해가 될 것입니다.
먼저 Claude에게 문서 링크를 주고 스스로 검색하게 했습니다——가능하기는 했지만, Claude가 올바른 답을 찾기 위해 대량의 결과를 컨텍스트에 불러들이는 문제가 있었습니다. 최종 해결책은 전용 서브 agent를 구축하는 것이었습니다. Claude Code Guide입니다. 이 서브 agent는 상세한 검색 지침을 가지고 있으며, 문서를 효율적으로 검색하는 방법과 반환해야 할 내용을 파악하고 있습니다. **새로운 도구를 추가하지 않고도, Claude의 action space를 확장했습니다.**
Lance Martin은 agent 설계 패턴에 관한 글에서 보완적인 시각을 제시했습니다. **agent에게 수십 개의 도구를 정의하는 것보다, 컴퓨터를 주고 코드로 도구를 편성하게 하는 것을 고려하세요.** Claude Code의 핵심 추상화는 CLI입니다——agent는 여러분의 컴퓨터 위에서 살며, bash와 파일 시스템이라는 기초 프리미티브를 통해 복잡한 태스크를 수행합니다. 소수의 원자적 도구(bash 도구 등)는 방대한 도구 세트보다 더 유연하고 token 효율이 높습니다.
***
## 모델 공감——Agent처럼 생각하기
앞서 살펴본 세 가지 주제——적을수록 많다, 도구는 낡아진다, agent 스스로 답을 찾게 하라——의 배경에는 공통된 메타 방법론이 있습니다. Thariq은 글 서두에서 이미 그것을 가리켰습니다. **agent처럼 세상을 보라.**
이것은 규칙의 집합이 아니라 하나의 사고방식입니다. David Zhang은 이것에 이름을 붙였습니다.
**모델 공감(Model Empathy)**——인간의 관점에서 "합리적인" 도구를 설계하는 것이 아니라, 모델의 관점에서 그것이 실제로 무엇을 보고 있는지, 어떻게 이해하는지, 어떻게 사용하는지를 생각하는 것입니다.
Terminally Drifting은 모든 agent 팀이 경험하는 동일한 교훈을 세 단계로 압축했습니다.
> 1. 인간을 위해 도구를 설계한다
> 2. 모델이 관리자 권한을 가진 라쿤처럼 도구를 사용한다
> 3. token 경제와 예측 가능한 부작용을 위해 재설계한다
"agent처럼 생각하기"가 바로 그 결정적인 돌파구입니다.
Vish는 자신의 전환점 경험을 공유했습니다. "우리는 계속해서 인간에게 합리적으로 느껴지는 도구 인터페이스를 구축해왔는데, agent가 왜 이상한 선택을 하는지 이해할 수 없었습니다. **한번 마음의 모델을 뒤집어서, 모델이 도구 정의 안에서 실제로 무엇을 보고 있는지를 생각하자, 모든 것이 달라졌습니다.**"
Emeka는 엔터프라이즈 도구 개발의 관점에서 같은 이야기를 했습니다. "기본값은 모두 인간의 멘탈 모델입니다——스키마, 필드명, 플로우, 모두 인간이 읽기 위해 최적화되어 있습니다. 하지만 agent에게 이것은 잘못된 프레임워크입니다."
"마음의 모델을 뒤집는다"는 것은 말로는 쉽지만, 실제로는 끊임없는 연습이 필요합니다. agent의 출력을 꼼꼼히 읽어야 합니다——무엇을 올바르게 했는지를 보는 것이 아니라, 왜 어떤 선택을 했는지, 어디서 망설였는지, 어디서 돌아갔는지를 보는 것입니다. 이러한 "이상한 행동"은 종종 모델의 버그가 아니라, 도구 설계의 버그입니다.
커뮤니티 토론에서는 몇 가지 생각해볼 만한 추가 시각도 등장했습니다.
OAIR은 맹점을 제기했습니다. **이 모든 도구 반복은 agent가 상태가 없다는 것을 전제로 합니다——각 세션은 처음부터 시작됩니다. 만약 가장 중요한 "도구"가 action space 안에 있는 것이 아니라, 이 코드베이스가 어떻게 작동하는지에 대한 영속적인 기억 속에 있다면 어떨까요?** Claude Code는 이후 CLAUDE.md 파일과 memory 시스템으로 이 문제에 부분적으로 답했지만, 영속적 상태 관리는 여전히 agent 설계의 열린 과제로 남아 있습니다.
Clinker는 또 다른 설계 원칙을 제시했습니다. **원시 능력이 아니라 복구 가능성(recoverability)을 최적화하세요——명시적인 도구 사전 조건, 관찰 가능한 상태, 저비용 재시도는 빠르게 더 많은 도구를 추가하는 것보다 훨씬 효과적인 경우가 많습니다.** 이것은 소프트웨어 엔지니어링에서 "오류 없는 시스템보다 결함 허용 시스템을 만들라"는 이념과 일맥상통합니다.
范式折叠의 정리가 가장 날카롭습니다.
> Action space 설계는 본질적으로 권한 설계입니다——AI에게 어떤 권한을 주는지에 따라, 그것이 어떤 역할이 되는지가 결정됩니다. 팀 관리와 똑같습니다. 병목이 사람의 능력에 있다고 생각했는데, 실제로는 자신이 그은 권한 경계가 병목이었던 것입니다.
***
## 마치며
Thariq의 결론으로 돌아가겠습니다.
> 많이 실험하고, 출력을 읽고, 새로운 방법을 시도하세요. Agent처럼 관찰하세요.
Claude Code의 헤비 유저로서, 이 글을 읽고 가장 크게 느낀 것은 이것입니다. **겉으로 "자연스러워 보이는" 기능들 뒤에는, 수없이 많은 "이렇게는 안 되겠다, 다른 방법을 써보자"는 반복이 있었다는 것입니다.** AskUserQuestion은 세 번을 시도했고, TodoWrite는 세 번을 재설계했으며, RAG는 Grep으로 대체되었습니다. 매번의 개선은 더 영리한 방안을 생각해냈기 때문이 아니라, 모델의 실제 행동을 꼼꼼히 관찰했기 때문입니다.
이 글의 부제는 "Seeing like an Agent"——agent처럼 세상을 보는 것입니다. 하지만 각도를 바꿔 생각해보면, 이것은 사실 모든 좋은 엔지니어링 실천의 본질이기도 합니다. **자신의 시각에서 시스템을 설계하는 것이 아니라, 사용자의 시각에서 설계하는 것.** 다만 이번에는, 그 사용자가 AI 모델이라는 것이 다를 뿐입니다.
앞으로 agent를 구축하는 모든 개발자는, David Zhang이 말한 "모델 공감"을 익혀야 할 것입니다. 이것은 어떤 신비로운 능력이 아닙니다. 핵심은 세 가지입니다. **모델의 실제 행동을 관찰하고, 그 출력을 읽고, 본 것을 바탕으로 설계를 조정하는 것.**
Agent처럼 관찰하세요.
***
**더 읽어보기:**
# 칩 전쟁부터 우주 데이터센터까지:AI 산업의 다음 10년
> 투자는 진실을 추구하는 행위입니다. 진실을 먼저 발견하고 올바르게 판단한다면, 그것이 바로 알파를 창출하는 방법입니다. 그리고 그것은 반드시 다른 사람들이 아직 보지 못한 진실이어야 합니다.
이 말은 Atreides Management 창립자 Gavin Baker가 Patrick O'Shaughnessy의 팟캐스트 "Invest Like the Best"에서 한 인터뷰에서 나온 것입니다. Gavin은 테크 투자 분야에서 가장 열정적이고 통찰력 있는 투자자 중 한 명으로 꼽힙니다. 거의 두 시간에 걸친 이 대화는 GPU, TPU, AI 경제학, 우주 데이터센터, SaaS의 미래, 그리고 스키 강사에서 투자자로의 그의 인생 전환점까지 폭넓게 다루고 있습니다.
이 인터뷰는 정보 밀도가 매우 높아서 "잠깐, 한번 생각해봐야겠다"는 순간이 너무나 많았습니다. 다음은 가장 깊이 생각해볼 만한 몇 가지 관점입니다.
***
## AI 발전을 어떻게 추적할까요?먼저 200달러를 쓰세요
인터뷰 초반에 Patrick은 매우 실질적인 질문을 던졌습니다. Gemini 3 같은 새 모델이 출시되면 그 정보를 어떻게 처리하느냐는 것이었습니다.
Gavin의 대답은 직접적이었습니다. **직접 써봐야 한다**는 것이었습니다.
하지만 핵심은 "사용"이 아니라 어떤 버전을 사용하느냐에 있습니다. 그는 무료 버전 AI를 써보고 "AI가 별것 아니다"라는 결론을 내리는 투자자들에게 놀라움을 표했습니다.
> 무료 버전은 10살짜리 아이를 상대하는 것과 같습니다. 그 10살짜리의 모습을 보고 35살이 됐을 때 어떨지를 예측하는 거죠. 유료로 사용할 수 있습니다——실제로 최고 등급 멤버십을 이용하려면 반드시 유료여야 하는데, 월 200달러입니다. 그쪽이야말로 진짜 30\~35세 성인입니다.
이 비유는 매우 정확합니다. 국내외 대형 모델의 격차도 비슷한 이치입니다. 많은 사람들이 무료 모델을 잠깐 써보고 "AI가 별것 아니다"라고 생각하지만, Claude 4.5 Opus, Gemini 3 Pro, GPT-5.2 Reasoning 같은 최고급 모델을 사용해본다면 느낌이 완전히 달라집니다.
유료 구독에 관해 말하자면, 저는 매달 다양한 AI 제품 구독에 약 2,000위안 정도를 씁니다. 그 중 큰 부분은 250달러짜리 Claude Code Max 플랜입니다. 개발자로서 코딩 수요가 많다면, 각종 미러 사이트를 이용하지 말고 공식 Claude Code(125달러부터)를 직접 구독하시길 강력히 권합니다. 미러 사이트 뒤에서 실제로 어떤 모델을 사용하는지 알 수 없을 뿐만 아니라, Max 플랜의 가성비는 실제로 매우 높아서 종량제보다 훨씬 저렴하게 계산됩니다.
정보를 얻는 채널로는 Gavin의 답변이 많은 분들에게 의외일 수 있습니다. 바로 \*\*X (Twitter)\*\*입니다.
그는 AI의 발전이 상당 부분 "X 플랫폼에서 실시간으로 일어나고 있다"고 말했습니다. 지구상에서 AI 최전선을 진정으로 이해하는 사람은 약 500\~1,000명 정도이며, 상당수가 중국에 있고, 그 사람들을 면밀히 주시해야 한다고 했습니다. 특히 Andrej Karpathy를 언급했습니다.
> 안드레이 카르파시가 쓴 모든 글은 최소 세 번씩 읽어야 합니다. 최소한.
AI 동향을 주시하는 사람으로서, 저도 이에 깊이 공감합니다. Twitter의 AI 토론은 어떤 뉴스 미디어보다 더 실시간으로, 더 깊이 있게 이루어집니다. 연구소 연구원들이 최신 진전을 직접 포스팅해서 논의하고, 심지어 서로 "논쟁"을 벌이기도 합니다. Gavin은 Meta의 PyTorch 팀과 Google의 Jax 팀이 X에서 공개적인 논쟁을 벌인 적이 있었고, 결국 두 연구소의 책임자들이 나서서 "우리 연구소 사람들은 상대 연구소 험담을 하면 안 된다"고 밝혀야 했다고 언급했습니다.
***
## Scaling Laws:우리의 "고대 이집트의 순간"
Gemini 3가 출시된 후, 많은 사람들이 주목한 것은 이것이 scaling laws(규모의 법칙)에 대해 무엇을 말해주는가였습니다. Gavin은 제가 이전에 들어본 적 없는 시각을 제시했습니다.
> 사전 학습 scaling laws에 대한 우리의 이해는 고대 이집트인들이 태양을 이해하는 방식과 비슷할지도 모릅니다. 그들은 매우 정밀하게 측정할 수 있었습니다. 대피라미드의 동서 축이 춘분과 추분에 완벽하게 정렬될 만큼, 스톤헨지도 마찬가지로. 완벽한 측정이었습니다. 하지만 그들은 궤도 역학을 몰랐습니다. 왜 태양이 동쪽에서 떠서 서쪽으로 지는지 몰랐습니다.
이 비유는 저를 한참 멈춰서 생각하게 만들었습니다. 우리는 실제로 매우 정밀하게 예측할 수 있습니다. 모델에 컴퓨팅 파워를 10배 늘리면 성능이 얼마나 향상될지를요. 하지만 우리는 왜 그런지 모릅니다. 이것은 "법칙"이 아니라 "경험적 관찰"입니다. 극히 정밀하게 측정하지만 원리는 이해하지 못하는 경험적 관찰인 것입니다.
그렇다면 Gemini 3가 왜 중요할까요?그것이 이 "경험적 관찰"이 여전히 유효함을 증명했기 때문입니다. Blackwell 칩의 지연과 모든 사람들이 "scaling laws가 이미 실패한 것이 아닌가"를 걱정하던 시기에, Gemini 3는 명확한 답을 주었습니다. **실패하지 않았다**는 것입니다.
하지만 더 흥미로운 것은 Gavin이 이어서 한 말입니다. "추론 모델(reasoning models)"의 등장이 없었다면, 2024년부터 2025년까지의 AI 발전은 원래 정체되었을 것이라고 했습니다.
왜일까요?XAI가 20만 장의 Hopper GPU 협업 작동을 해결한 이후, 다음 단계는 Blackwell 칩을 기다려야 했기 때문입니다. 20만 장 이상의 Hopper를 "일관성(coherent)"있게——쉽게 말해 하나의 통합체처럼 작동하게——유지할 수 없었습니다. 그런데 Blackwell이 지연된 것입니다.
> 추론 모델이 없었다면, 2024년 중반부터 지금까지 AI에는 아무런 진전도 없었을 것입니다. 모든 것이 정체되었을 것입니다. 이것이 시장에 무엇을 의미하는지 상상할 수 있나요?우리는 완전히 다른 환경에서 살고 있었을 것입니다. 추론 모델은 어떤 의미에서 AI를 구했습니다. Blackwell 없이도 AI가 계속 발전할 수 있게 해줬으니까요.
이것은 제가 이전에 인식하지 못했던 시각입니다. 추론 모델(예: o1)은 단순히 새로운 기능이 아니라, 실제로 AI 산업 전체의 발전 리듬을 "구했"습니다.
***
## 칩 전쟁:Google이 "산소를 빨아들이고 있다"
GPU와 TPU의 경쟁에 관해 이야기하면서, Gavin은 인상 깊은 말을 했습니다.
> Google은 현재 토큰의 최저 비용 생산자입니다. 그들이 계속해온 일은 제가 "AI 생태계의 경제적 산소를 빨아들이는 것"이라고 표현하겠습니다——이는 그들에게 극도로 합리적인 전략입니다.
최저 비용 생산자로서, Google은 저가(심지어 손실을 감수하면서)로 AI 서비스를 제공하며 경쟁자들을 힘들게 만들고 있습니다. 이는 테크 업계의 전형적인 전략이지만, Gavin은 흥미로운 변화를 지적했습니다.
> AI는 제 커리어에서 처음으로 "최저 비용 생산자"라는 것이 테크 분야에서 진정으로 중요해진 사례입니다. Apple은 저비용으로 스마트폰을 만들어서 수조 달러 시가총액이 된 것이 아닙니다. Microsoft는 저비용으로 소프트웨어를 만들어서 수조 달러 시가총액이 된 것이 아닙니다. NVIDIA도 저비용으로 AI 가속기를 만들어서 수조 달러 시가총액이 된 것이 아닙니다. 이것은 한 번도 중요한 적이 없었습니다.
하지만 AI 시대에 전력이 제한 요소가 되면서, **와트당 얼마나 많은 토큰을 생산할 수 있는가**가 결정적으로 중요해졌습니다. 와트당 3~~5배의 토큰을 생산할 수 있다면, 그것은 3~~5배의 수익입니다. 컴퓨팅 가격은 무관해지는데, 병목은 전력이기 때문입니다.
이 구도는 곧 바뀔 것입니다. Blackwell 칩이 드디어 배포되기 시작했고, Gavin은 첫 번째 Blackwell 모델이 XAI에서 나올 것이라고 예측했습니다.
> Jensen에 따르면, Elon보다 데이터센터를 더 빠르게 짓는 사람은 없다고 합니다. Jensen이 공개적으로 한 말입니다.
Blackwell과 이후의 Ruben 칩이 대규모로 배포되면, 최저 비용 생산자로서 Google의 우위는 사라질 것입니다. 그때가 되면 그들이 여전히 -30%의 마진으로 AI 비즈니스를 계속 운영하려 할까요?이 셈법은 완전히 달라질 것입니다.
***
## 우주 데이터센터:미치광이 같지만 제1원칙으로 보면 맞는 말
Patrick이 "잘 논의되지 않는 황당한 아이디어가 있느냐"고 물었을 때, Gavin은 우주 데이터센터 이야기를 시작했습니다. 처음에는 농담인 줄 알았지만, 그의 분석을 다 듣고 나서 이것이 인터뷰 전체에서 가장 선견지명 있는 부분일 수 있다는 것을 깨달았습니다.
> 제1원칙의 관점에서 보면, 우주 데이터센터는 모든 차원에서 지구상의 데이터센터보다 우수합니다.
그의 논거는 다음과 같습니다.
**1. 에너지**:우주에서 위성은 24시간 햇빛에 노출될 수 있으며, 태양 복사 강도는 지상보다 6배 높습니다. 게다가 항상 햇빛이 있으므로 배터리가 필요 없습니다. 배터리는 비용의 상당 부분을 차지합니다. 따라서 태양계에서 가장 저렴한 에너지원은 "우주 태양광"입니다.
**2. 냉각**:지구상의 데이터센터는 비용과 무게의 상당 부분이 냉각에 사용됩니다. 하지만 우주에서는 어떨까요?냉각이 무료입니다. 방열판을 위성의 음영 면에 설치하면, 그곳은 절대영도에 가깝습니다.
**3. 네트워크**:데이터센터 내부에서 랙 사이는 광섬유로 연결됩니다. 본질적으로 레이저가 케이블을 통해 전달되는 것입니다. 이보다 더 빠른 것은 무엇일까요?레이저가 진공을 통과하는 것입니다. 따라서 우주의 위성들을 레이저로 연결하면, 네트워크가 실제로 지상 데이터센터보다 더 빠릅니다.
**4. 사용자 경험**:지금 AI에게 질문하면, 신호가 휴대폰에서 기지국을 거쳐, 광섬유를 통해, 어떤 데이터센터로 가서 계산한 후 다시 돌아옵니다. 하지만 위성이 휴대폰과 직접 통신할 수 있다면(Starlink는 이미 휴대폰 직접 연결 기능을 증명했습니다), 전체 경로가 훨씬 짧아집니다.
물론 이를 위해서는 Starship의 대규모 발사가 필요하고, 아마 5\~6년은 더 걸릴 것입니다. 하지만 Gavin은 흥미로운 수렴을 지적했습니다. Tesla, SpaceX, XAI가 합류하고 있습니다. XAI는 Optimus 로봇의 "지능 모듈"이 되고, SpaceX는 우주에 데이터센터를 구축해 AI에 컴퓨팅 파워를 제공하며, 이 세 회사는 서로 경쟁 우위를 강화하는 선순환 구조를 형성하고 있습니다.
***
## SaaS의 "불타는 플랫폼"
앞의 내용이 AI의 미래에 대해 흥분감을 느끼게 했다면, 이 부분은 많은 기존 기업의 미래에 대해 우려를 갖게 할 수 있습니다.
Gavin은 직접적으로 말했습니다. **애플리케이션 SaaS 기업들은 오프라인 소매업체들이 전자상거래에 맞섰을 때 저질렀던 것과 완전히 동일한 실수를 범하고 있다**고 말입니다.
오프라인 소매업체들은 당시 아마존을 보며 "전자상거래는 저마진 사업인데, 어떻게 우리보다 효율적일 수 있겠느냐?지금 고객들이 스스로 돈 내고 매장에 와서 물건을 직접 들고 간다"고 생각했습니다. 그들은 고객의 니즈를 분명히 보면서도, 전자상거래의 마진 구조가 마음에 들지 않아 투자를 거부했습니다. 결과는 어떻게 됐을까요?아마존의 북미 소매 비즈니스 마진은 현재 많은 전통 소매업체보다 높습니다.
SaaS 기업들은 지금 같은 상황에 처해 있습니다. 전통적인 소프트웨어는 한 번 만들면 무한정 복사 배포할 수 있어 매출총이익률이 80\~90%에 달할 수 있습니다. 하지만 AI는 다릅니다. 매번 사용할 때마다 새로 계산해야 하므로, 좋은 AI 기업의 매출총이익률은 40%에 불과할 수 있습니다.
> AI 에이전트를 만들고 싶은데 35% 이하의 마진으로 운영하기 싫다면, 절대 성공하지 못할 것입니다. AI 네이티브 기업들은 바로 그 마진으로 운영하고 있으니까요. 80%의 마진을 지키고 싶다면, AI 분야에서 성공하지 못한다는 것을 스스로 보장하는 겁니다. 절대적인 보장입니다.
Gavin은 이것이 "생사를 가르는 결정"이라고 말하며, **Microsoft를 제외하고 거의 모든 기업이 실패하고 있다**고 했습니다.
그는 Nokia의 유명한 "불타는 플랫폼" 메모를 인용했습니다. 당신의 플랫폼이 불타고 있습니다. 하지만 옆에 아주 좋은 새 플랫폼이 있고, 그쪽으로 뛰어넘어간 다음 돌아와서 원래 플랫폼의 불을 끌 수 있습니다. 그러면 이제 두 개의 플랫폼이 생기는 것입니다.
Salesforce, ServiceNow, HubSpot, GitLab, Atlassian——그는 이 모든 기업들이 다음 전략을 실행할 수 있고 또 해야 한다고 봅니다. AI 수익을 공개하고, AI 매출총이익률을 공개하며(낮은 마진이야말로 "진정한 AI"임을 증명합니다), 그런 다음 아직 적자인 벤처캐피털 지원 경쟁자들을 가리키며 "나는 그들이 없는 것을 갖고 있다: 현금흐름을 창출하는 사업"이라고 말하는 것입니다.
***
## 한 투자자의 성장 이야기
인터뷰 마지막에 Patrick은 좀 더 개인적인 질문을 했습니다. 젊은 사람들에게 자신이 하는 일을 어떻게 소개하겠느냐는 것이었습니다.
Gavin의 대답은 "투자는 진실을 추구하는 행위"에서 시작했지만, 진정으로 흥미로운 것은 그의 인생 이야기였습니다.
원래 그의 계획은 이랬습니다. 겨울에는 스키 강사, 여름에는 래프팅 가이드, 비수기에는 암벽등반을 하며 소설 쓰기와 야생동물 사진을 시도하는 것이었습니다. 이것이 대학 시절 그의 "인생 계획"이었고, 부모님도 많이 지지해주셨습니다.
하지만 부모님은 한 가지 작은 부탁을 했습니다. 전문직 인턴십을 딱 하나만, 무엇이든 상관없으니 해볼 수 없겠느냐고요.
그가 찾을 수 있는 유일한 인턴십은 한 증권사의 프라이빗 웰스 매니지먼트 부서였습니다. 일은 간단했습니다. 회사가 리서치 보고서를 발행할 때마다, 어떤 고객들이 해당 주식을 보유하고 있는지 확인한 다음 보고서를 발송하는 것이었습니다.
그러다가 그는 그 보고서들을 읽기 시작했습니다.
> 제 생각에 "세상에, 이게 내가 상상할 수 있는 가장 흥미로운 일이구나"라는 생각이 들었습니다.
그는 투자를 "기술과 운이 공존하는 게임"으로 이해했습니다. 포커와 비슷합니다. 운이 나빠서 질 수도 있습니다——예를 들어 투자한 회사 본사에 운석이 떨어진다든지——하지만 대부분의 경우 기술이 중요합니다. 우위를 얻는 방법은 가장 철저한 역사적 지식과 현재 세계에 대한 가장 정확한 이해를 결합하여, "다음에 무슨 일이 일어날 것인가"에 대한 차별화된 판단을 형성하는 것입니다.
그것은 인턴십 3일째 되던 날이었습니다. 그는 서점에 가서 Peter Lynch의 책을 사서 이틀 만에 다 읽었습니다. 그런 다음 Warren Buffett을 읽고, 『Market Wizards』를 읽고, 버핏이 주주들에게 보낸 편지를 읽었습니다——두 번씩이나요. 그리고 회계를 독학했습니다. 학교로 돌아간 후 전공을 영어와 역사에서 역사와 경제학으로 바꿨습니다.
그는 또한 청소부로 일했던 경험을 언급했습니다. Alta 스키장에서 아르바이트를 할 때 객실 청소를 했습니다. 한번은 방을 청소하던 중 투숙객이 읽고 있는 책이 자신이 읽고 있는 책과 같은 것을 보고, "좋은 책이네요, 저도 마침 비슷한 부분을 읽고 있었어요"라고 말했습니다. 상대방은 그를 외계인 보듯이 바라보다가 더욱 놀란 표정으로 물었습니다. "청소하는 사람이 책도 읽어요?"
> 이 일은 다른 사람들을 대하는 방식에 영구적인 영향을 미쳤습니다.
***
## 에필로그:AI가 무엇을 필요로 하든 얻게 됩니다
인터뷰가 거의 끝날 무렵, Gavin은 제가 가장 흥미롭다고 느낀 말을 했습니다.
> 지난 2년간, AI가 계속 발전하기 위해 무엇이 필요하든 그것을 얻었습니다. 미국 대중의 여론이 원자력 문제에 대해 이렇게 빠르게 바뀐 것을 본 적 있나요?그냥 그렇게 됐습니다. 그것도 정확히 AI가 그것을 필요로 할 때. 이제 지구상의 전력 한계에 부딪히자, 갑자기 우주 데이터센터에 관한 논의가 등장했습니다. AI 발전을 늦출 수 있는 병목이 생길 때마다, 모든 것이 오히려 가속화되었습니다.
이것은 Kevin Kelly가 『기술은 무엇을 원하는가(What Technology Wants)』에서 제시한 "technium"(기술권) 개념을 떠올리게 합니다. 기술은 하나의 전체로서 점점 더 강력해지려는 나름의 의지를 가진 것처럼 보입니다.
이것이 우연일 수도 있습니다. 단지 많은 똑똑한 사람들이 문제를 해결하는 것일 수도 있습니다. 하지만 Gavin이 관찰한 이 패턴——AI가 장애물에 부딪힐 때마다 그 장애물이 어떤 방식으로든 제거된다——은 분명 깊이 생각해볼 만한 가치가 있습니다.
# 큐레이션 해설
# Curations 큐레이션 해설
여기에는 기술 전문가의 영상을 시청하고, 양질의 블로그를 읽은 후의 사고와 정리를 담았습니다.
단순한 노트 발췌가 아니라, 저 자신의 이해와 실천 경험을 녹여낸 내용입니다.
## 콘텐츠 출처
* 기술 영상 해설
* 블로그 아티클 선별
* 팟캐스트 / 인터뷰 정리
## 최신 콘텐츠
### Agent처럼 생각하기: Claude Code 팀의 도구 설계 철학
Anthropic 엔지니어 Thariq이 Claude Code를 구축하는 과정에서 경험한 에이전트 도구 설계 노하우를 공유합니다. AskUserQuestion의 세 번에 걸친 반복 개선, TodoWrite의 세 차례 리팩터링, RAG에서 점진적 공개(Progressive Disclosure)에 이르기까지, 모든 사례는 하나의 핵심 방법론을 가리킵니다. 바로 Agent의 시선으로 세상을 바라보는 것입니다.
[전문 읽기 →](./claude-code-seeing-like-an-agent)
### 대학생의 90%가 AI를 사용하지만, 규칙을 아는 사람은 아무도 없다
프린스턴, UC 버클리, LSE 출신의 학생 네 명이 캠퍼스 내 AI의 실제 현주소에 대해 이야기합니다. 부정행위, 혼란, 양극화, 그리고 아무도 감히 묻지 못하는 질문, 즉 대학은 여전히 어떤 의미를 지니는가에 대해 솔직하게 다룹니다. 홍보 영상이 아닌, 진솔한 고민과 성찰의 기록입니다.
[전문 읽기 →](./ai-on-campus-student-perspectives)
### 반도체 전쟁에서 우주 데이터센터까지: AI 산업의 다음 10년
반도체 전쟁과 우주 데이터센터, SaaS의 생존 기로와 투자의 본질까지. Gavin Baker는 "Invest Like the Best" 팟캐스트에서 AI 산업에 대한 가장 깊이 있는 통찰을 나눕니다. 왜 반드시 유료 AI를 써야 하는지, 스케일링 법칙(Scaling Laws)이 왜 "고대 이집트인이 태양을 이해하는 방식"과 같은지, 그리고 추론 모델이 어떻게 AI 산업 전체의 발전 리듬을 '구했는지'에 대해 이야기합니다.
[전문 읽기 →](./gavin-baker-ai-economics)
# Karpathy의 트윗은 별 62,000개로 폭발했습니다. andrej-karpathy-skills는 정확히 무엇을 했나요?
2026년 1월 27일, Andrej Karpathy는 11개 섹션, 약 1,400 단어로 구성된 프로그래밍 에세이인 X에 매우 긴 트윗을 게시하여 11월 "80% 필기 + 20% 에이전트"에서 12월 "80% 에이전트 + 20% 연마"로 전환하는 동안 직면한 함정을 기록했습니다. 해당 트윗은 결국 **769만 조회 수, 좋아요 39,000개, 북마크 36,000개**에 도달했습니다.
3개월 후 Karpathy의 트윗을 설치 가능한 4개의 규칙으로 패키징하는 `forrestchang/andrej-karpathy-skills`이라는 GitHub 저장소가 출시되었습니다. 2주 만에 **62.7,000개의 별과 5,500개의 포크**에 도달하여 2026년 4월 GitHub 주간 목록에서 1위를 차지했습니다.
웨어하우스 온톨로지: **Markdown 파일**.
***
## 1. Karpathy는 무엇에 대해 불평하고 있나요?
Karpathy의 긴 기사에는 기본적으로 LLM으로 코딩된 네 가지 "만성 질환"이 나열되어 있습니다.
**첫 번째 질병: 몰래 가정을 하는 것**
> "가장 일반적인 유형의 실수는 모델이 잘못된 가정을 한 다음 이를 검증하지 않고 적용한다는 것입니다. 그들은 자신의 혼란을 관리하지 않고, 설명을 구하지도 않고, 불일치를 보여주지도 않고, 절충안을 제시하지도 않으며, 반박할 때가 되면 반박하지도 않고, 약간 너무 아첨합니다."
이것은 "집단 오류"입니다. "로그인 추가"라고 말하면 어떤 인증을 사용할지, 장치를 기억할지, 세션을 관리하는 방법을 묻지 않고 단지 합리적이라고 생각하는 솔루션을 제시할 뿐입니다. 검토를 마치고 원하는 내용과 다르다는 것을 발견할 때쯤에는 이미 500줄이 작성되어 있을 것입니다.
**두 번째 질병: 과잉 엔지니어링**
> "그들은 특히 코드와 API를 지나치게 복잡하게 만들고, 추상화 레이어를 부풀리고, 죽은 코드를 정리하지 않습니다. 그들은 1000줄의 코드를 사용하여 비효율적이고, 부풀리고, 깨지기 쉬운 구조를 구현할 것이며, 그들이 '물론이죠!'라고 말하기 전에 어린아이처럼 속여서 '글쎄, 그냥 이걸 하는 게 어때?'라고 말해야 합니다. 그런 다음 즉시 100줄로 줄이세요."
이는 LLM 코딩의 가장 일반적인 관찰자 효과입니다. 맥락이 "관대하게" 주어지는 동시에 복잡성도 "관대하게" 보상됩니다\*\*. 전략 모드, 팩토리 모드, 종속성 주입 등이 모두 제공됩니다.
**세 번째 질병: 바꾸라고 하지 않은 것을 바꾸세요**
> "그들은 때때로 일부 주석과 코드가 마음에 들지 않거나 완전히 이해하지 못하기 때문에 일부 주석과 코드를 수정하거나 삭제합니다. 이러한 변경 사항이 현재 작업과 아무 관련이 없더라도 마찬가지입니다."
버그 수정을 요청하면 "더 이상 필요하지 않은 것 같다"는 이유로 옆에 있는 완료되지 않은 TODO 주석을 편리하게 삭제합니다.
**질병 4: CLAUDE.md에 규칙을 작성해도 여전히 깨집니다**
> "CLAUDE.md에서 간단한 복구 시도를 했는데도 위의 문제가 여전히 존재합니다."
트윗 전체에서 가장 가슴 아픈 문장입니다. 전 OpenAI 창립 팀원이자 Tesla AI 이사였던 Karpathy는 Claude를 완벽하게 유지하는 CLAUDE.md를 작성할 수 없었습니다.
***
## 2. `andrej-karpathy-skills`에 대한 해결책
`forrestchang` 이 네 가지 질병에 대한 해결책을 네 가지 원칙으로 체계화하고 이를 CLAUDE.md 파일에 패키지합니다.
| 원칙 | 해당 질병 | 핵심활동 |
| ------------------ | --------- | ------------------------------------------------------------- |
| **코딩하기 전에 생각해보세요** | 비밀리에 가정 | 가정을 명확하게 기술하고, 다양한 해석을 나열하고, 혼란스러우면 잠시 멈추고 질문하고, 필요할 때는 반박하세요 |
| **단순성 우선** | 과도한 엔지니어링 | 필요한 최소 코드만 작성하세요. 추측성 유연성, 오류 처리, 추상화를 작성하지 마십시오 |
| **수술적 변화** | 무단 변경 | 필요한 것만 만지십시오. 스타일을 리팩토링하거나 변경하지 마세요. 다른 데드 코드만 보고하고 삭제하지 마세요 |
| **목표 중심 실행** | 방법 정렬 불량 | 검증 가능한 성공 기준 + 테스트를 제공하고 모델 루프가 자동으로 통과하도록 허용 |
네 번째 원칙은 Karpathy의 트윗에 있는 또 다른 유명한 대사인 "Leverage"를 직접 참조하는 것입니다.
이것이 전체 프로젝트 방법론의 발판입니다. \*\*처음 세 가지 원칙은 LLM이 어지러워지는 것을 방지합니다. 네 번째 원칙은 그 강점을 진정으로 활용하는 방법을 알려줍니다. \*\*
***
## 3. 설치 및 사용방법
**방법 A: Claude Code 플러그인으로** (이것을 먼저 사용하는 것이 좋습니다)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
설치 후 모든 프로젝트의 Claude Code 대화 상자는 자동으로 이 네 가지 원칙을 준수합니다. 전역적으로 적용되며 언제든지 `/plugin`에 의해 해제될 수 있습니다.
**방법 B: CLAUDE.md를 수동으로 복사**
저장소를 입력하고 → `CLAUDE.md`을 열고 → 복사한 후 프로젝트 루트 디렉터리의 `CLAUDE.md`에 붙여넣습니다. 이 프로젝트에만 유효합니다.
저장소는 또한 주류 AI IDE를 다루는 콘텐츠 세트인 추가 `CURSOR.md` 및 `.cursor/rules/` 적응을 제공합니다.
***
## 4. 왜 62,000개의 별에 도달할 수 있나요?
이는 풀어볼 가치가 있는 현상입니다. 별 62.7,000개는 "단일 파일 저장소"에 대한 과장된 숫자입니다. 비교를 위해 같은 기간 동안 Microsoft markitdown(9k)과 Addy Osmani의 에이전트 기술(4.6k)을 합친 것은 그만큼 많지 않습니다.
충격 무게별로 분류:
**1. Karpathy IP 보증** - `forrestchang-skills`이라고 불리는 경우 동일한 콘텐츠가 10,000개를 초과할 수 없습니다. Karpathy에는 "전 OpenAI 창립팀 + Tesla AI 디렉터 + CS231n 강사"라는 문화적 자본이 함께 제공되며 그의 트윗에는 "필독" 라벨이 붙어 있습니다.
**2. 완벽한 타이밍** — Opus 4.7은 4월 16일에 출시되었으며 과도한 엔지니어링에 대한 불만이 최고조에 달했습니다. 이 레포는 모두가 "클로드가 너무 미치게 만드는 것을 막기 위해" 해독제를 찾고 있을 때 나타났습니다.
**3. 문제점은 보편적입니다** - 모든 클로드 코드/커서 사용자는 이 네 가지 함정을 밟았으며 공감률은 100%에 가깝습니다.
**4. 임계값이 매우 낮습니다** - 파일 1개 또는 명령 2줄. 스타의 가격은 무시할 수 있을 정도로 저렴합니다. "설치하지 않으면 손해를 보게 됩니다."
**5. 강력한 검증 가능성** - 4가지 원칙은 명확하고 기억하기 쉬우며 스크린샷을 찍어 전달하기도 쉽습니다. 금지된 1000라인 프롬프트 프로젝트 가이드와는 다릅니다.
**6. 이중 언어 README** ——`README.zh.md`는 중국 AI 서클 트래픽을 직접 소모하며, V2EX는/즉시/Weibo에서 동시에 폭발합니다.
**7. 저자에 의한 교차 프로모션** - 상단 열의 문장 \*"내 새 프로젝트 Multica를 확인하세요"\*는 저자의 자체 상업 에이전트 플랫폼 `multica-ai/multica`으로 트래픽을 유도합니다. \*\*이 저장소는 본질적으로 Multica 고객 확보 퍼널의 최상위입니다. \*\*
**8. 메타 적합** - 여기서 논의하는 "LLM 코딩 오류"는 LLM으로 코딩할 때 모든 독자가 경험하는 것과 정확히 같습니다. 읽기와 사용이 통합되어 전환율이 매우 높습니다.
한마디로, 판매하는 것은 코드나 도구가 아니라 **Karpathy의 감정을 설치 가능한 규칙으로 포장**하는 것입니다. 이는 2026년 AI 프로그래밍계에서 가장 일반적인 "콘텐츠가 제품" 사례입니다.
***
## 5. 나의 사용 제안
**먼저 방법 A를 사용하여 전역적으로 설치합니다**. 도구와 스크립트를 작성할 때 경험이 향상되는지 확인하십시오. 특히 Claude가 다른 사람의 코드를 변경하도록 허용한 경우 무작위 변경 문제가 줄어드는지 확인하십시오.
**CLAUDE.md 프로젝트에 병합할지 여부는 1\~2주 후에 결정됩니다**. 각 프로젝트의 CLAUDE.md는 이미 도메인 지식(설계 시스템, 구성 요소 사양, 배포 프로세스)으로 채워져 있는 반면 Karpathy의 세트는 일반적인 방법론입니다. 두 가지가 충돌하지 않으며 겹쳐질 수 있습니다. 그러나 그것이 유용하다고 정말로 확신할 때까지 타이밍을 기다려야 합니다.
**비용에 주의하세요**: Claude가 더 많은 질문을 하게 될 것이며, 이는 "한 문장으로 생성"하는 데 익숙한 사람들에게는 짜증스러울 것입니다. 수행되어야 하는 약간의 청소를 수행하지 못할 수도 있습니다(너무 엄격함). 매우 모호한 탐사 작업으로 인해 제약을 받게 됩니다.
**더 깊은 가치**: 요구 사항을 명확하게 명시해야 하며 이는 모든 고품질 소프트웨어 엔지니어링의 전제 조건입니다.
**주목할 만한 후속 조치**: forrestchang 자신도 기술 메커니즘을 제품화하기 위해 "오픈 소스 관리형 에이전트 플랫폼"인 Multica를 홍보하고 있습니다. 이 네 가지 원칙이 결국 사실상의 표준이 된다면 Multica는 상용차가 될 것입니다. 이 줄에 주목하세요.
***
## 참조 리소스
# Indie Dev
준비부터 출시까지 인디 개발의 전체 과정을 기록합니다.
# Bark
# Bark
간단한 HTTP 요청으로 iPhone에 맞춤 푸시 알림을 전송할 수 있는 도구입니다. 무료, 오픈소스, 자체 배포를 지원합니다.
## 추천하는 이유
* **극도로 간단한 API** - curl 명령어 하나로 알림 전송이 가능하며, 복잡한 설정이 필요하지 않습니다
* **오픈소스 무료** - MIT 라이선스로 완전 공개되어 있으며 비용이 전혀 없습니다
* **개인 정보 보호** - 자체 서버 구축을 지원하여 푸시 데이터를 직접 관리할 수 있습니다
## 활용 사례
* **스크립트 알림** - 데이터 백업 완료 등 오래 걸리는 작업이 끝났을 때 알림을 받을 수 있습니다
* **서비스 모니터링** - 서버 이상 발생 시 즉시 경고를 받을 수 있습니다
* **자동화 연동** - CI/CD 빌드 결과, 예약 작업 완료 알림 등에 활용할 수 있습니다
## 빠른 시작
1. [App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865)에서 Bark를 다운로드합니다
2. 앱을 열고 푸시 주소를 복사합니다
3. 첫 번째 알림을 전송합니다:
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
제목이 포함된 알림:
```bash
curl https://api.day.app/YOUR_KEY/제목/내용
```
## 프로젝트 정보
* GitHub: [Finb/Bark](https://github.com/Finb/Bark)
* Stars: 7.2k+
* 라이선스: MIT
# Toolkit
# Toolkit
제가 발견한 우수한 GitHub 프로젝트와 실용적인 소프트웨어를 모아두었습니다.
# Claude Agent Teams 완전 가이드
## 서론
Claude Code의 Subagent를 사용해본 적이 있다면, 병렬 개발이 이미 충분히 강력하다고 느꼈을 것입니다. 하지만 Subagent에는 한계가 있습니다. 메인 Agent에게 결과를 보고하는 것만 가능하며, 서로 소통할 수 없습니다.
Agent Teams는 이 상황을 완전히 바꿉니다. 상상해보십시오. 한 Agent가 보안 감사를 담당하고, 다른 Agent가 성능 최적화를 담당하며, 세 번째 Agent가 테스트 커버리지를 담당합니다. 이들은 병렬로 작업할 수 있을 뿐만 아니라, 직접 대화하고, 서로 도전하며, 합의에 도달할 수 있습니다. 이것이 Agent Teams의 핵심 가치입니다.
## Agent Teams 이해하기
Agent Teams의 아키텍처는 실제 개발 팀과 매우 유사합니다.
```
┌─────────────────────────────────────────────────────────┐
│ 여러분 (사용자) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (메인 Claude 인스턴스, 조율 담당) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│ 보안 감사 │◄─►│ 성능 최적화 │◄─►│ 테스트 커버리지│
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└──────────────┼──────────────┘
▼
┌─────────────────┐
│ 공유 작업 목록 │
└─────────────────┘
```
### Subagent와의 차이점
| 특성 | Subagent | Agent Teams |
| ------------ | ------------------------- | ----------------------- |
| **컨텍스트** | 독립 컨텍스트, 결과를 메인 Agent에 반환 | 독립 컨텍스트, 완전히 독립적으로 실행 |
| **통신 방식** | 메인 Agent에게만 보고 가능 | Teammates 간 직접 통신 가능 |
| **작업 조율** | 메인 Agent가 모든 작업 관리 | 공유 작업 목록으로 자체 조율 |
| **적용 시나리오** | 결과만 필요한 집중 작업 | 토론과 협업이 필요한 복잡한 작업 |
| **Token 비용** | 낮음: 결과 요약이 메인 컨텍스트로 반환 | 높음: 각 Teammate가 독립 인스턴스 |
간단히 말하면: **Subagent는 작업을 수행하러 보내는 하청업체이고, Agent Teams는 같은 방에서 협업하는 프로젝트 팀**입니다.
### Agent Teams가 효과적인 이유
핵심 통찰: **전문화가 집중을 가져옵니다**.
단일 Agent가 복잡한 다단계 작업을 처리할 때, 컨텍스트가 계속 팽창하여 자주 `/clear`로 초기화해야 합니다. Agent Teams는 각 Teammate가 좁은 전문 영역을 유지하도록 하여 컨텍스트를 깨끗하게 유지하고, 성능이 더 안정적입니다.
## Agent Teams 활성화
Agent Teams는 현재 실험적 기능으로, 기본적으로 비활성화되어 있습니다. 수동으로 활성화해야 합니다.
**방법 1: 환경 변수**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**방법 2: settings.json (권장, 영구 적용)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## 핵심 사용법
### 첫 번째 Agent Team 생성
활성화 후, 자연어로 Claude에게 팀 생성을 요청하면 됩니다.
```
agent team을 생성하여 PR #142를 리뷰해줘.
세 명의 리뷰어를 생성해줘:
- 보안 문제에 집중하는 리뷰어
- 성능 영향을 확인하는 리뷰어
- 테스트 커버리지를 검증하는 리뷰어
각자 리뷰하고 발견 사항을 보고하게 해줘.
```
**키워드 팁**: "create an agent team" 또는 "spawn an agent team"을 사용하십시오. "spawn agents"만 말하면 Subagent와 Agent Teams가 혼동될 수 있습니다.
### 표시 모드
Agent Teams는 두 가지 표시 모드를 지원합니다.
| 모드 | 설명 | 요구 사항 |
| --------------- | ------------------------- | ----------------- |
| **In-process** | 모든 Teammates가 메인 터미널에서 실행 | 특별한 요구 사항 없음 |
| **Split panes** | 각 Teammate가 독립 창에서 실행 | tmux 또는 iTerm2 필요 |
기본값은 `auto`입니다. tmux에서 실행 중이면 split panes를 사용하고, 그렇지 않으면 in-process를 사용합니다.
**표시 모드 설정**:
```json
{
"teammateMode": "in-process"
}
```
**단일 세션에서 지정**:
```bash
claude --teammate-mode in-process
```
### 자주 사용하는 단축키
| 동작 | 단축키 |
| --------------------- | ------------ |
| Teammates 간 전환 | `Shift+Down` |
| 이전 Teammate로 돌아가기 | `Shift+Up` |
| 작업 목록 표시 전환 | `Ctrl+T` |
| 현재 Teammate 중단 | `Escape` |
| **Delegate Mode 활성화** | `Shift+Tab` |
| Teammate 세션 진입 | `Enter` |
### Delegate Mode (중요)
Delegate Mode는 Agent Teams의 가장 중요한 기능 중 하나입니다.
| 모드 | Lead 동작 |
| ----------------- | --------------------------------- |
| **일반 모드** | Lead가 직접 작업을 구현하고 코드를 작성할 수 있음 |
| **Delegate Mode** | Lead는 조율만 가능하며, 코드 작성이나 테스트 실행 불가 |
**Delegate Mode가 필요한 이유**:
이 제한이 없으면, Lead가 자주 "일을 빼앗아" 합니다. 세 명의 Teammates가 작업을 기다리고 있는데, Lead가 직접 코드를 작성하기 시작합니다. Delegate Mode를 활성화하면, Lead는 순수한 프로젝트 매니저로 강제됩니다. 작업 관리, Teammates와의 소통, 출력 검토만 가능합니다.
```
# 팀 시작 후 즉시 Shift+Tab을 눌러 활성화
```
## 실전 사례
### 사례 1: 병렬 코드 리뷰
단일 리뷰어는 특정 유형의 문제에 깊이 빠지기 쉽습니다. 리뷰 차원을 독립된 영역으로 분리하면 보안, 성능, 테스트 커버리지 모두 동등한 관심을 받을 수 있습니다.
```
agent team을 생성하여 이 PR을 리뷰해줘. 세 명의 리뷰어를 생성해줘:
- 보안 리뷰어: 인증, 권한, 인젝션 취약점 검사
- 성능 리뷰어: 알고리즘 복잡도, 데이터베이스 쿼리, 캐싱 전략 분석
- 테스트 리뷰어: 테스트 커버리지, 경계 조건, 오류 처리 검증
각자 리뷰한 후, 발견한 문제에 대해 서로 토론하게 해줘.
```
### 사례 2: 경쟁적 가설 디버깅
근본 원인이 불명확할 때, 단일 Agent는 그럴듯한 설명을 하나 찾으면 멈추는 경향이 있습니다. Teammates가 서로 도전하게 하면 이 문제를 피할 수 있습니다.
```
사용자가 앱에서 메시지를 하나 보내면 연결을 유지하지 않고 종료된다고 보고했습니다.
5개의 agent teammates를 생성하여 다른 가설을 조사하게 해줘. 서로 토론하며
서로의 이론을 반박하도록 하고, 과학적 토론처럼 진행해줘. 합의된 발견 사항을
조사 보고서에 업데이트해줘.
```
**핵심 메커니즘**: 토론 구조입니다. 여러 독립적인 조사자가 적극적으로 서로의 이론을 뒤집으려 시도하며, 최종적으로 살아남는 가설이 진정한 근본 원인일 가능성이 더 높습니다.
### 사례 3: 콘텐츠 대량 생산
비기술적 작업의 전형적인 활용입니다. 하나의 입력을 여러 출력으로 변환합니다.
```
agent team을 생성하여 이 영상 스크립트를 네 개 플랫폼의 콘텐츠로 변환해줘:
- LinkedIn 기사 작성자
- Twitter 스레드 작성자
- Newsletter 작성자
- 블로그 글 작성자
스크립트 위치: /content/scripts/video-20.md
```
각 Teammate가 독립적으로 작성하면서도 콘텐츠 일관성을 유지합니다.
### 사례 4: QA 품질 검사 클러스터
블로그 웹사이트의 품질 검사에 5개의 Agent를 배포하여 서로 다른 측면을 병렬로 테스트합니다.
```
agent team을 생성하여 블로그에 대한 전면적인 품질 검사를 수행해줘:
- Agent 1: 핵심 페이지 테스트 (홈페이지, 소개 페이지, 연락처 페이지)
- Agent 2: 글 페이지 테스트 (렌더링, 네비게이션, SEO 메타데이터)
- Agent 3: 링크 검사 (내부 링크, 외부 링크, 깨진 링크)
- Agent 4: SEO 검증 (제목, 설명, 구조화된 데이터)
- Agent 5: 접근성 테스트 (ARIA 라벨, 대비도, 키보드 네비게이션)
우선순위별로 정렬된 문제 보고서를 생성해줘.
```
**효과**: 원래 사람이 순차적으로 실행해야 했던 전면 검사를 몇 분 만에 완료합니다. 각 Agent가 자신의 영역에 집중하고, 마지막에 우선순위별로 정렬된 문제 목록으로 종합합니다.
### 사례 5: 다중 라운드 토론 모드
유용한 프롬프트 패턴으로, Teammates가 회의하듯 토론하게 합니다.
```
Agent Teams를 사용하여 4명의 teammates를 생성하고 [기술 결정]에 대해 토론하게 해줘.
3라운드 토론을 진행해줘. 각 라운드에서 teammates가 서로 교류하게 해줘.
한 teammate는 전담으로 Red Team 관점을 맡아 비판적 의견을 제시하게 해줘.
```
이 패턴은 아키텍처 결정, 기술 선택 등 다각도 비교가 필요한 시나리오에 특히 적합합니다.
### 사례 6: C 컴파일러 프로젝트
Anthropic은 16개의 Agent를 사용하여 Linux 커널을 컴파일할 수 있는 C 컴파일러를 처음부터 구축했습니다.
| 지표 | 데이터 |
| --------- | ----------------------- |
| Agent 수 | 16개 병렬 인스턴스 |
| 세션 수 | \~2,000개 Claude Code 세션 |
| 비용 | \~$20,000 |
| 코드 라인 수 | 100,000줄 |
| Token 사용량 | 20억 입력 + 1.4억 출력 |
최종 결과물: x86, ARM, RISC-V에서 부팅 가능한 Linux 6.9를 빌드할 수 있는 Rust 컴파일러입니다.
## 팀 관리
### Teammates 및 모델 지정
Claude가 작업에 따라 자동으로 생성할 Teammates 수를 결정하지만, 명시적으로 지정할 수도 있습니다.
```
4명의 teammates를 생성하여 이 모듈들을 병렬로 리팩토링해줘.
각 teammate는 Sonnet 모델을 사용해줘.
```
### 계획 승인 요청
복잡하거나 위험도가 높은 작업의 경우, Teammates에게 실행 전에 먼저 계획을 수립하도록 요청할 수 있습니다.
```
아키텍트 teammate를 생성하여 인증 모듈을 리팩토링해줘.
변경을 수행하기 전에 계획 승인을 요청해줘.
```
Teammate가 계획을 완료하면 Lead에게 승인 요청을 보냅니다. Lead가 검토 후 승인하거나 수정 의견을 반려할 수 있습니다.
### Teammates에게 직접 대화하기
각 Teammate는 완전한 Claude Code 세션입니다. 어떤 Teammate에게든 직접 메시지를 보낼 수 있습니다.
* **In-process 모드**: `Shift+Down`으로 전환한 후 메시지 입력
* **Split-pane 모드**: 해당 창을 직접 클릭
### Teammates 종료
```
보안 감사 teammate를 종료해줘
```
Lead가 종료 요청을 보내며, Teammate는 승인하거나 거부(이유 설명 포함)할 수 있습니다.
### 팀 정리
완료 후, Lead에게 리소스 정리를 요청합니다.
```
팀을 정리해줘
```
**중요**: 항상 Lead를 통해 정리하십시오. Teammates에게 정리를 맡기면 리소스 상태가 불일치할 수 있습니다.
## 모범 사례
### 팀 규모 관리
규모 권장 사항:
| 팀 규모 | 적합한 시나리오 |
| ----- | ------------------------- |
| 3명 | 간단한 다각도 리뷰 |
| 4-5명 | 표준적인 기능 개발 또는 리팩토링 |
| 6명 이상 | 대규모 마이그레이션 또는 복잡한 아키텍처 작업 |
**경험 법칙**: 각 Teammate에게 5-6개의 작업을 할당하는 것이 적절합니다. 15개의 독립적인 작업이 있다면, 3명의 Teammates가 좋은 출발점입니다.
### 작업 세분화
* **너무 작음**: 조율 오버헤드가 이점을 초과
* **너무 큼**: Teammates가 너무 오래 작업하여 체크포인트가 없고, 낭비 위험 증가
* **적절함**: 독립적이고 완결되며, 산출물이 명확한 작업 단위 (함수 하나, 테스트 파일 하나, 리뷰 보고서 하나)
### 파일 충돌 방지
두 Teammates가 같은 파일을 편집하면 덮어쓰기가 발생합니다. 작업을 분배할 때 각 Teammate가 다른 파일 집합을 담당하도록 하십시오.
```
Teammate 1: src/auth/ 디렉토리 담당
Teammate 2: src/api/ 디렉토리 담당
Teammate 3: src/utils/ 디렉토리 담당
```
### 모니터링 및 안내
정기적으로 Teammates의 진행 상황을 확인하고, 부적절한 방향을 적시에 수정하십시오. 팀을 너무 오래 무감독으로 실행하면 낭비 위험이 증가합니다.
Lead가 Teammates를 기다리지 않고 직접 작업을 구현하기 시작하면:
```
teammates가 작업을 완료한 후에 진행해줘
```
### 충분한 컨텍스트 제공
Teammates는 프로젝트 컨텍스트(CLAUDE.md, MCP servers, skills)를 자동으로 로드하지만, Lead의 대화 히스토리는 상속하지 않습니다. 생성 시 충분한 작업 세부 정보를 제공하십시오.
```
보안 감사 teammate를 생성하되, 다음과 같은 프롬프트를 제공해줘:
"src/auth/ 디렉토리의 인증 모듈 보안 취약점을 감사해줘.
토큰 처리, 세션 관리, 입력 검증에 중점을 둬줘.
앱은 httpOnly 쿠키에 저장된 JWT 토큰을 사용합니다.
발견 사항 보고 시 심각도 등급을 첨부해줘."
```
### 자기 보고 검증 패턴
작업 설명에 명확한 검증 기준을 포함하여, Teammates가 완료 후 자체 점검하도록 합니다.
```
작업 완료 시 Lead에게 보고해줘:
1. 어떤 파일을 검사했는지
2. 어떤 문제를 발견했는지
3. 어떤 수정을 했는지
4. 검증 기준(테스트 통과, lint 경고 없음 등)이 충족되었는지
```
이 패턴은 Lead의 검증 작업량을 줄이면서, 작업이 "완료된 것처럼 보이는" 것이 아닌 진정으로 완료되었는지 확인합니다.
## 고급 기법
### Hooks를 사용한 강제 품질 게이트
Hooks를 통해 Teammates가 작업을 완료할 때 규칙을 강제 적용합니다.
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle`: Teammate가 유휴 상태에 진입하기 직전에 실행됩니다. exit code 2를 반환하면 피드백을 보내고 Teammate가 작업을 계속하게 할 수 있습니다
* `TaskCompleted`: 작업이 완료로 표시될 때 실행됩니다. exit code 2를 반환하면 완료를 차단하고 피드백을 보낼 수 있습니다
### 사전 승인 권한
Teammate의 권한 요청이 Lead에게 전달되어 빈번한 중단을 야기할 수 있습니다. 생성 전에 일반적인 작업을 사전 승인하십시오.
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### Worktree와 결합
Agent Teams는 Worktree와 함께 사용할 수 있으며, 각 Teammate가 자신의 worktree에서 작업합니다.
```
agent team을 생성하되, 각 teammate가 독립된 worktree에서 작업하여
파일 충돌을 피하게 해줘.
```
### 서드파티 오케스트레이션 도구
네이티브 Agent Teams 외에도, 커뮤니티에서 개발한 오케스트레이션 도구들이 있습니다.
| 도구 | 설명 |
| --------------- | -------------------------- |
| **Gas Town** | 여러 병렬 Claude 세션을 관리하는 도구 |
| **Multiclaude** | 여러 터미널 창에서 Claude 인스턴스를 실행 |
이러한 도구들은 Agent Teams 실험적 기능 외에 대안을 제공하지만, 더 많은 수동 설정이 필요합니다. 네이티브 Agent Teams가 요구 사항을 충족한다면, 공식 기능을 우선 사용하는 것을 권장합니다.
## 현재 제한사항
Agent Teams는 아직 실험적 기능이므로, 제한사항을 이해하는 것이 중요합니다.
| 제한사항 | 설명 |
| --------------------------- | ---------------------------------------------------- |
| In-process teammates 복원 불가 | `/resume` 및 `/rewind`가 in-process teammates를 복원하지 않음 |
| 작업 상태 지연 가능 | Teammates가 때때로 작업 완료 표시를 잊음 |
| 종료가 느릴 수 있음 | Teammates가 현재 요청을 완료한 후에야 종료됨 |
| 세션당 하나의 팀 | Lead가 한 번에 하나의 팀만 관리 가능 |
| 중첩 팀 불가 | Teammates가 자체적으로 팀을 생성할 수 없음 |
| Lead 고정 | 팀을 생성한 세션이 Lead이며, 이전 불가 |
| Split panes는 tmux/iTerm2 필요 | VS Code 터미널, Windows Terminal, Ghostty는 미지원 |
| Plan mode는 세션 레벨 | Teammate의 Plan mode 상태는 생성 시 고정되며, 세션 중 변경 불가 |
## 비용 고려사항
Agent Teams의 Token 소모는 단일 세션보다 현저히 높습니다.
| 시나리오 | Token 소모 | 비용 배수 |
| -------------------------- | ------------- | ------------ |
| 단일 Agent 세션 | \~200k tokens | 1x |
| 3명의 Teammates | \~800k tokens | \~4x |
| 5명의 Teammates | \~1.2M tokens | \~6x |
| 16명의 Teammates (C 컴파일러 사례) | 20억 tokens | $20,000 / 2주 |
**비용 분석**:
* 각 Teammate는 완전히 독립된 Claude 인스턴스로, 자체 컨텍스트를 가짐
* Teammates 간의 통신도 Token을 소모
* Lead가 모든 Teammates를 조율해야 하므로 추가 오버헤드 발생
**언제 가치가 있는가**:
* 병렬 탐색이 필요한 연구 작업
* 다각도 리뷰 (보안, 성능, 테스트)
* 토론을 통해 합의에 도달해야 하는 결정
* 순차적으로 완료할 수 있는 일반 작업에는 부적합
* 상호 통신이 필요 없는 병렬 작업에는 부적합 (Subagent가 더 경제적)
## 사용 경험
### 언제 Agent Teams를 사용하는가
판단 기준은 다음과 같습니다.
1. **작업에 다양한 관점이 필요**: 서로 다른 영역의 전문 지식 (보안 + 성능 + 테스트)
2. **토론과 합의가 필요**: 경쟁적 가설, 아키텍처 결정
3. **병렬 탐색에 가치가 있음**: 여러 구현 방안 비교
단순히 병렬 실행만 필요하고 상호 통신이 불필요하다면, Subagent나 Worktree가 더 적합합니다.
### 연구와 리뷰부터 시작하기
Agent Teams를 처음 사용한다면, 코드 작성이 필요 없는 작업부터 시작하십시오. PR 리뷰, 기술 방안 연구, 버그 조사 등입니다. 이런 작업은 명확한 경계를 가지며, 병렬 탐색의 가치를 보여주면서 병렬 구현에 따른 조율 과제를 피할 수 있습니다.
### 다른 기능과의 조합
| 조합 | 효과 |
| ---------------------- | ----------------------- |
| Agent Teams + Worktree | 각 Teammate가 격리된 환경에서 작업 |
| Agent Teams + Hooks | 자동화된 품질 검사 및 피드백 |
| Agent Teams + Skills | 각 Teammate가 전문 역량 보유 |
## 마치며
Agent Teams는 AI 보조 개발의 새로운 패러다임을 대표합니다. "하나의 AI 어시스턴트"에서 "하나의 AI 팀"으로의 전환입니다.
Anthropic은 16개의 Agent를 사용하여 2주간 $20,000을 들여 10만 줄의 코드로 된 C 컴파일러를 만들었습니다. 이 프로젝트의 핵심 교훈은: **테스트 품질이 무엇보다 중요하다**는 것입니다. Agent는 주어진 문제를 자율적으로 해결하므로, 작업 검증기가 거의 완벽해야 합니다. 그렇지 않으면 Agent가 잘못된 문제를 해결하게 됩니다.
세 가지 핵심 요점을 기억하십시오.
| 요점 | 설명 |
| ------ | ---------------------------------- |
| **협업** | Teammates가 직접 소통하며, 결과만 보고하는 것이 아님 |
| **공유** | 공유 작업 목록을 통해 작업 조율 |
| **감독** | 정기적으로 진행 상황을 확인하고, 적시에 방향 수정 |
시작은 간단합니다.
```
agent team을 생성하여 [당신의 작업]을 해줘
```
***
**관련 읽을거리**:
* [Claude Worktree 완전 가이드](/ko/docs/notes/claude-worktree) — Worktree와 Agent Teams의 조합 이해하기
* [Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent) — Subagent와 Agent Teams의 사용 시나리오 비교
* [Tmux 빠른 시작 가이드](/ko/docs/notes/tmux-tutorial) — Tmux를 사용한 다중 Agent 세션 관리
**참고 자료**:
* [Claude Code 공식 문서 - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**영상 튜토리얼**:
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) — 7가지 비기술적 활용 사례 데모
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) — 다중 에이전트 워크플로 상세 해설
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) — 팀 오케스트레이션 심층 분석
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) — 완전한 설정 튜토리얼
# Claude 시스템 아키텍처 완전 해부
## 서론
2025년 9월, Anthropic은 **$183B 기업가치**로 $13B 자금을 조달하며 세계 4위 비상장 기업이 되었습니다. 핵심 제품인 Claude Code는 2월 출시 이후 **11.5만** 명의 활성 개발자를 유치했으며, 매주 **1.95억 줄**의 코드를 처리하고, 사용자 성장률은 \*\*300%\*\*에 달합니다.
더 흥미로운 것은 Anthropic CEO Dario Amodei가 밝힌 사실입니다: **Claude Code 코드의 90%가 자기 자신이 작성했습니다**.
**어떻게 가능할까요?**
하나의 AI 프로그래밍 어시스턴트가 어떻게 "스스로를 작성"할 수 있을까요? 어떤 아키텍처 설계가 이처럼 효율적으로 인간 개발자를 보조하고, 심지어 대체할 수 있게 만드는 것일까요?
정답은 Claude의 **모듈식 아키텍처**에 있습니다: MCP는 도구를 제공하고, Skills는 사용법을 가르치며, Subagents는 병렬로 실행하고, Hooks는 제어 가능성을 보장합니다. 이러한 구성 요소들이 협력하여 Claude에게 "프로그래머처럼 일하는" 능력을 부여합니다.
이 문서에서는 이 아키텍처의 **전체 그림을 조감**하겠습니다. 각 구성 요소의 역할, 구성 요소 간의 협업 관계, 그리고 빠르게 시작할 수 있는 설정 예시를 다룹니다. 후속 글에서는 각 구성 요소의 세부 사항을 깊이 분석할 예정입니다.
## 전체 아키텍처 개요
Claude 시스템은 **모듈식 아키텍처** 설계를 채택하고 있으며, 각 구성 요소는 기능별로 분류되어 **상호 보완적으로 협력**하며 계층적 의존이 아닙니다:
**핵심 이해**: 이러한 구성 요소들은 **동급 보완적** 확장 기능이며 계층적 의존이 아닙니다. 필요에 따라 자유롭게 조합할 수 있습니다:
| 원하는 것... | 사용... | 한 줄 정의 |
| ------------------- | ------------- | ------------------------------------------- |
| 외부 데이터 소스 및 서비스 연결 | **MCP** | Claude에게 "손과 발"을 달아 데이터베이스, API, 파일 시스템에 접근 |
| Claude에게 특정 워크플로 교육 | **Skills** | Claude가 특정 분야에서 "어떻게 해야 하는지 알게" 합니다 |
| 복잡한 작업 병렬 처리 | **Subagents** | 큰 작업을 작은 작업으로 분해하여 여러 Agent가 동시에 수행 |
| 반복 작업 빠르게 트리거 | **Commands** | 원클릭으로 자주 쓰는 워크플로를 시작하여 반복 지시를 절약 |
| 특정 작업의 필수 실행 보장 | **Hooks** | Claude의 판단과 무관하게 반드시 실행되는 단계 |
***
## 핵심 런타임
### Agent SDK — 런타임 엔진
Agent SDK는 전체 Claude Agent 시스템의 **런타임 핵심 엔진**으로, 다음을 제공합니다:
* **메인 루프 (Main Loop)**: Agent의 핵심 작업 루프
* **컨텍스트 관리**: Token 예산, 자동 압축 (92% 사용률에서 트리거)
* **도구 디스패치**: 어떤 도구를 사용할지, 어떻게 실행할지 결정
* **권한 시스템**: 도구 접근 권한 제어
Agent의 핵심 작업 모드는 간단한 **피드백 루프**입니다:
```
컨텍스트 수집 → 작업 실행 → 작업 검증 → 반복
```
***
### Built-in Tools — 핵심 도구
Claude Agent에는 20개 이상의 핵심 도구가 내장되어 있으며, 세 가지 범주로 나뉩니다:
| 범주 | 도구 | 설명 |
| -------- | ------------------- | ------------------- |
| **읽기** | Read, Glob, Grep | 파일 읽기, 패턴 매칭, 내용 검색 |
| **작업** | Write, Edit, Bash | 파일 쓰기, 편집, 명령 실행 |
| **네트워크** | WebSearch, WebFetch | 웹 검색, 웹 페이지 크롤링 |
이러한 도구는 **기본적으로 사용 가능**하며 추가 설정이 필요 없습니다. Claude는 이러한 도구를 통해 컴퓨터와 상호작용하며, 프로그래머가 IDE를 사용하는 것과 같습니다.
***
## 설정 및 컨텍스트
### CLAUDE.md — 영구 컨텍스트
새로운 대화를 시작할 때마다 프로젝트 배경, 코딩 규칙, 아키텍처 약속을 반복해서 설명해야 합니다... CLAUDE.md를 사용하면 이러한 정보를 **한 번 설정하고 자동으로 로드**할 수 있습니다.
CLAUDE.md는 프로젝트의 **README for AI**와 같습니다. Claude에게 이 프로젝트의 배경 지식, 작업 방식, 약속을 알려줍니다.
#### 계층적 오버라이드
Claude는 다음 순서로 CLAUDE.md를 로드하며, **더 구체적인 것이 더 높은 우선순위**를 가집니다:
```
Enterprise (최저)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md) (최고)
```
#### 내용 제안
CLAUDE.md에는 다음과 같은 핵심 정보를 포함해야 합니다:
| 범주 | 내용 예시 |
| ----------- | ---------------------------------------------- |
| **기술 스택** | Next.js 14 + TypeScript, Tailwind CSS |
| **빌드 명령** | `npm run dev`, `npm run build`, `npm run test` |
| **코드 규칙** | 명명 규칙, Lint 도구 설정 |
| **프로젝트 구조** | 핵심 디렉토리의 용도 설명 |
**핵심 원칙**: 간결하게 유지하세요. CLAUDE.md는 **매 대화마다 로드**되므로, 너무 길면 소중한 Token을 낭비하게 됩니다.
***
## 패키징 및 배포
### Plugins — 설치 가능한 단위
팀 설정이 분산되어 있고, 공유 및 표준화가 어렵습니다. 각자 자신만의 Skills, Commands, Hooks 세트를 가지고 있는데... 어떻게 통합 관리할 수 있을까요?
Plugins는 **Skills + Commands + Subagents + Hooks + MCP**를 **설치 가능한 단위**로 패키징하여, 원클릭 배포와 팀 표준화를 실현합니다.
#### 디렉토리 구조
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 플러그인 매니페스트 (필수)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Hooks 설정
├── .mcp.json # MCP Server 설정
└── README.md # 설명 문서
```
#### 설정 예시
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "프론트엔드 개발 툴킷",
"author": "Your Team"
}
```
```bash
# 설치 방법
claude plugin install github:your-org/your-plugin # GitHub에서
claude plugin install /path/to/plugin # 로컬에서
```
**관련 리소스**
| 리소스 | 설명 |
| -------------------------------------------------------------------------------- | ------------------------------- |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Anthropic 공식 플러그인 저장소 |
| [wshobson/agents](https://github.com/wshobson/agents) | ⭐ 24.3k, 고품질 Agent 템플릿 모음 |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | 커뮤니티 플러그인 마켓, 다양한 플러그인 검색 가능 |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | ⭐ 19.3k, 엄선된 Claude Code 리소스 목록 |
***
## 확장 기능 (상호 보완 모듈)
Claude 시스템의 확장 기능은 여러 **상호 보완 모듈**로 구성되어 있으며, 각자의 역할을 수행하면서 협력합니다:
| 모듈 | 기능 정의 | 활성화 방식 |
| ------------- | ----------------------------- | ---------- |
| **MCP** | 외부 데이터 및 서비스 연결 (WHAT) | 설정 후 사용 가능 |
| **Skills** | 절차적 지식, Claude에게 방법을 교육 (HOW) | 자동 매칭 |
| **Subagents** | 독립 컨텍스트, 병렬 작업 위임 | 명시적 호출 |
| **Commands** | 반복적인 워크플로 | 수동 `/cmd` |
| **Hooks** | 결정론적 제어, 이벤트 기반 | 자동 트리거 |
***
### MCP — 외부 연결
#### 설계 철학
기존 방식에서는 각 외부 데이터 소스마다 **커스텀 통합**이 필요했으며, 이로 인해 N×M의 통합 지옥이 발생했습니다. MCP는 표준화된 프로토콜을 제공하여 **한 번 연결하면 어디서나 사용** 가능하게 합니다.
MCP (Model Context Protocol)는 **AI 애플리케이션의 USB-C 인터페이스**로 설계되었습니다:
| 특성 | 설명 |
| ---------- | ------------------------------------- |
| **개방 표준** | 2024년 11월 발표, 2025년 12월 Linux 재단에 기부 |
| **업계 채택** | OpenAI, Microsoft, Google, AWS 등이 채택 |
| **생태계 규모** | 월 9,700만+ SDK 다운로드, 수천 개의 커뮤니티 Server |
#### 아키텍처 패턴
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, AI 도구)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (프로토콜 클라이언트)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (특정 기능 노출)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**사용 시나리오**: 데이터베이스 연결, 서드파티 서비스 통합 (GitHub, Slack, Notion), 프라이빗 API 접근, 실시간 데이터 스트림 처리.
#### 설정 예시
프로젝트 루트 디렉토리에 `.mcp.json`을 생성합니다:
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — 수동 워크플로
Slash Commands는 **수동으로 트리거**하는 반복적인 워크플로를 제공합니다.
| 특성 | 설명 |
| ---------- | ----------------------- |
| **트리거 방식** | 수동으로 `/command-name` 입력 |
| **저장 위치** | `.claude/commands/` |
| **용도** | 반복적인 워크플로, 표준화된 작업 |
**예시**: `.claude/commands/review.md` 생성
```markdown
현재 변경 사항에 대해 코드 리뷰를 수행하세요. 다음에 중점을 두세요:
1. 코드 스타일과 일관성
2. 잠재적인 성능 문제
3. 보안 취약점
4. 테스트 커버리지
```
그런 다음 `/review`를 입력하면 트리거됩니다.
***
### Hooks — 결정론적 제어
Hooks는 **결정론적 제어**의 핵심입니다. 특정 작업은 반드시 실행되어야 하며 LLM의 판단에 의존할 수 없습니다.
| 범주 | 이벤트 | 트리거 시점 |
| ----------- | -------------------- | ---------------- |
| **도구** | `PreToolUse` | 도구 실행 전 |
| | `PostToolUse` | 도구 성공적 실행 후 |
| | `PostToolUseFailure` | 도구 실행 실패 후 |
| | `PermissionRequest` | 권한 요청 시 |
| **세션** | `SessionStart` | 세션 시작 시 |
| | `SessionEnd` | 세션 종료 시 |
| | `Stop` | Claude가 응답 완료 시 |
| **서브 에이전트** | `SubagentStart` | 서브 에이전트 시작 시 |
| | `SubagentStop` | 서브 에이전트 중지 시 |
| **기타** | `UserPromptSubmit` | 사용자가 prompt 제출 후 |
| | `Notification` | 알림 이벤트 |
| | `PreCompact` | 컨텍스트 압축 전 |
#### 설정 예시
TypeScript 파일 자동 포맷팅:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### 실전: 자율 루프
**Ralph Wiggum**은 Anthropic 공식 플러그인으로, Stop hook을 활용하여 자율 반복 루프를 구현합니다:
```bash
/ralph-loop "TODO API 구현, CRUD 및 테스트 포함" --max-iterations 20
```
**작동 원리**: Stop hook이 Claude의 종료를 가로채서 원래 prompt를 다시 주입하고, 작업이 완료되거나 최대 반복 횟수에 도달할 때까지 계속 반복합니다.
**적용 시나리오**: 여러 차례 반복이 필요한 작업 (테스트 통과, 코드 리팩토링), 자동 검증 수단이 있는 작업.
***
### Subagents — 작업 위임 및 병렬 실행
#### 설계 철학
단일 Agent가 직면하는 과제: 제한된 컨텍스트 윈도우, 병렬 처리 불가, 불명확한 책임 분담. Subagents는 **Orchestrator-Worker** 아키텍처 패턴을 채택하여 이러한 문제를 해결합니다:
```
Main Agent (Orchestrator)
├── 사용자 요청 분석
├── 계획 수립
├── 작업 분해
└── 전문화된 서브 에이전트 생성
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - 병렬 실행)
│ 코드 │ 테스트 │ 문서 │
└────────┴────────┴────────┘
↓
결과 통합 → 메인 에이전트 종합 출력
```
#### 핵심 특성
| 특성 | 설명 |
| ------------ | ----------------------------------- |
| **컨텍스트 격리** | 각 Subagent는 독립된 컨텍스트를 가지며 오염을 방지합니다 |
| **작업 전문화** | 커스텀 시스템 프롬프트로 전용 역할을 정의합니다 |
| **도구 권한 제어** | Subagent가 특정 도구만 사용하도록 제한할 수 있습니다 |
| **병렬 실행** | 여러 Subagent가 동시에 작업합니다 |
**성능 데이터**: 다중 에이전트 시스템은 단일 에이전트보다 90.2% 높은 성능을 보이며, 병렬화는 연구 시간을 90% 단축할 수 있습니다 (Token 소비는 약 15배이지만, 복잡한 작업에는 가치가 있습니다).
#### 설정 예시
`.claude/agents/`에 Markdown 파일을 생성합니다:
```markdown
---
name: Code Reviewer
description: 코드 리뷰 전문 서브 에이전트
tools:
- Read
- Grep
- Glob
---
당신은 숙련된 코드 리뷰 전문가입니다. 다음에 중점을 두세요:
1. 코드 품질과 유지보수성
2. 잠재적인 버그와 경계 조건
3. 성능 최적화 기회
4. 보안 취약점
```
***
### Skills — 절차적 지식
#### 설계 철학
Skills는 AI를 위한 **재사용 가능한 작업 매뉴얼**입니다. 모듈화된 지식 패키지로, Claude가 필요에 따라 동적으로 로드할 수 있습니다. 핵심 설계 원칙은 \*\*점진적 공개 (Progressive Disclosure)\*\*입니다:
```
📚 Skills 작업 매뉴얼
│
├─ 📋 목차 ────────────── 【메타데이터 레이어】시작 시 사전 로드 (~30-50 tokens)
│ name: "weekly-report"
│ description: "표준화된 주간 보고서 생성"
│
├─ 📖 본문 섹션 ─────────── 【핵심 문서 레이어】관련 시 로드 (~수백-수천 tokens)
│ # Weekly Report Generator
│ ## Instructions
│ 다음 구조에 따라 주간 보고서 생성...
│
└─ 📎 부록 ────────────── 【참조 리소스 레이어】필요 시 로드
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills**: MCP는 Claude에게 도구에 접근하는 능력을 제공하고 (WHAT), Skills는 Claude에게 이러한 도구를 효과적으로 사용하는 방법을 가르칩니다 (HOW).
#### 핵심 장점
| 장점 | 설명 |
| ------------- | -------------------------------------------------- |
| **Token 효율적** | 메타데이터는 30-50 tokens만 차지하여 수십 개의 Skills를 동시에 활성화 가능 |
| **자동 활성화** | 작업 컨텍스트에 따라 자동 매칭, 수동 트리거 불필요 |
| **조합 가능** | 여러 Skills가 자동으로 협력 |
| **이식 가능** | Claude.ai, Claude Code, API 전반에서 일관된 경험 |
#### 설정 예시
`.claude/skills/`에 디렉토리를 생성합니다:
```
my-skill/
├── SKILL.md # 핵심 지시 사항 (필수)
├── scripts/ # 실행 가능한 스크립트 (선택)
└── references/ # 참고 자료 (선택)
```
SKILL.md 핵심 구조:
```yaml
---
name: code-review # Skill 이름
description: 코드 리뷰, 품질 및 보안 검사 # 간단한 설명 (자동 매칭에 사용)
---
# Code Review Skill
## Instructions
[구체적인 단계 설명...]
## Output Format
[출력 형식 요구 사항...]
```
**핵심 사항**: frontmatter의 `description`은 자동 매칭에 사용되므로, 간결하고 정확하게 유지하세요.
***
## 공식 참조 링크
**설계 철학**
| 리소스 | 설명 |
| --------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Agent 아키텍처의 핵심 문서 |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Agent SDK 엔지니어링 실습 |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Skills 설계 철학 |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | MCP 발표 공지 |
**공식 문서**
| 리소스 | 설명 |
| ----------------------------------------------------------------------------- | ------------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Skills와 다른 구성 요소 비교 |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | CLAUDE.md 사용 가이드 |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Subagents 공식 문서 |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Hooks 공식 문서 |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | MCP 공식 문서 |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | MCP 프로토콜 사양 |
**심층 분석**
| 리소스 | 설명 |
| --------------------------------------------------------------------------------------------------------- | ------------------- |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | 아키텍처 다이어그램 포함 |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Skills 원리 심층 분석 |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Claude Code 구축 비하인드 |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Skills와 MCP 기술 비교 |
***
## 더 읽을거리
Skills의 개념과 실습을 더 깊이 알고 싶다면 다음을 참고하세요:
* [Claude Skills란 무엇인가](/ko/docs/notes/claude-skills/concept) — Skills 핵심 원리 상세 해설
* [Claude Skills 실전 가이드](/ko/docs/notes/claude-skills/practice) — 첫 번째 Skill 직접 만들기
* [Claude Subagent 완벽 가이드](/ko/docs/notes/claude-subagent) — 서브 에이전트 사용 및 커스터마이징
* [GSD 심층 해부](/ko/docs/notes/gsd/concept) — 컨텍스트 엔지니어링 기반 AI 프로그래밍 시스템
* [나의 Claude Code 모범 사례](/ko/blog/claude-code-best-practices) — Claude Code 일상 사용 팁
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# Claude Worktree 완전 가이드
## 서론
Claude Code로 복잡한 작업을 처리할 때, 이런 곤란한 상황을 겪어본 적이 있을 것입니다. 세 가지 독립적인 작업을 처리해야 하는데, 같은 디렉토리에서 여러 Claude 인스턴스를 실행하면 코드가 충돌합니다. 한 Agent가 파일을 수정하는 동안 다른 Agent도 같은 파일을 건드리고, 마지막에 병합할 때 엉망이 됩니다.
2025년 2월, Anthropic은 `--worktree` 명령을 출시하여 이 상황을 완전히 바꾸었습니다. 이제 세 개의 터미널에서 각각 `claude -w feature-1`, `claude -w feature-2`, `claude -w bugfix-1`을 실행할 수 있으며, 세 Agent가 각자 격리된 환경에서 작업하여 서로 간섭하지 않습니다.
## Worktree 이해하기
여러분이 건축가이고, 동시에 세 개의 다른 방을 설계하고 있다고 상상해보십시오. 전통적인 방식은 같은 도면에 그리는 것인데, 수정을 반복하다 보면 쉽게 엉망이 됩니다. Worktree의 방식은 세 장의 독립된 도면을 제공하여 각 도면을 하나의 방 설계에 전용하고, 마지막에 주 도면에 통합하는 것입니다.
기술적으로 말하면, Worktree는 Git의 네이티브 기능입니다. Claude Code의 `--worktree` 명령은 이 기능을 더 간단하고 사용하기 쉽게 래핑했습니다. 하나의 명령으로 격리 환경 생성, Claude 인스턴스 시작, 완료 후 자동 정리까지 수행합니다.
### 왜 직접 여러 번 Clone하지 않는가
이런 의문이 들 수 있습니다. 왜 코드를 여러 번 clone하지 않는가?
| 방안 | 디스크 사용량 | 동기화 난이도 | 정리 복잡도 |
| ------------------- | ------------------ | --------------- | --------------------- |
| 다중 Clone | 각각 완전한 저장소 | 수동 pull/push 필요 | 수동으로 디렉토리 삭제 필요 |
| Git Worktree | 작업 파일만 복사, .git 공유 | 히스토리 자동 공유 | `git worktree remove` |
| Claude `--worktree` | 작업 파일만 복사, .git 공유 | 히스토리 자동 공유 | 종료 시 자동 정리 |
Worktree는 동일한 `.git` 데이터베이스를 공유하며, 모든 커밋 히스토리와 브랜치 정보가 공유됩니다. 이는 하나의 worktree에서 생성한 commit이 다른 worktree에서 즉시 보인다는 의미입니다.
### 언제 Worktree를 사용해야 하는가
작업을 시작하기 전에, 먼저 해당 작업이 worktree에 적합한지 판단해보십시오.
경험 법칙: **작업에 30분 이상이 필요하면 worktree 사용을 고려하십시오**. 짧은 작업에 worktree를 사용하면 오히려 시간 낭비입니다. 환경 생성, 의존성 설치, 최종 병합까지 합치면 작업 자체보다 오래 걸릴 수 있습니다. 하지만 깊은 작업이 필요한 작업에서는 worktree의 격리성이 매우 유용합니다.
| Worktree 사용에 적합 | 적합하지 않음 |
| ------------------- | ------------------- |
| 독립적인 기능 개발 | 10분이면 끝나는 작은 수정 |
| 서로 다른 모듈의 병렬 리팩토링 | 빈번한 상호작용이 필요한 작업 |
| 장시간 실행되는 작업 | 진행 중인 다른 수정에 강하게 의존 |
| 격리된 테스트가 필요한 실험적 변경 | 간단한 bug fix |
### 전제 조건
Worktree를 사용하기 전에 다음 조건을 충족하는지 확인하십시오.
| 조건 | 설명 |
| ------------- | ----------------------------------------- |
| Git 초기화 완료 | Git 저장소 디렉토리 내에 있어야 합니다 (`.git` 디렉토리 존재) |
| 최소 하나의 commit | 빈 저장소에서는 worktree를 생성할 수 없습니다 |
| 원격 브랜치 사용 가능 | 기본적으로 원격 브랜치에서 체크아웃합니다 (예: `origin/main`) |
## 완전한 워크플로: 생성부터 정리까지
아래에서 실제 개발 순서에 따라, Worktree 생성부터 최종 정리까지의 완전한 흐름을 살펴보겠습니다.
### 1단계: Worktree 생성
#### 원격 기본 브랜치에서 생성
`-w` 또는 `--worktree` 매개변수로 Claude를 시작합니다.
```bash
# "feature-auth"라는 이름의 worktree를 생성하고 Claude 시작
claude -w feature-auth
# 랜덤 이름 자동 생성 (예: "bright-running-fox")
claude -w
```
이 명령은 실제로 네 가지 작업을 수행합니다.
1. `/.claude/worktrees/feature-auth/`에 새 작업 디렉토리 생성
2. `worktree-feature-auth`라는 새 브랜치 생성
3. 원격 기본 브랜치(예: `origin/main` 또는 `origin/master`)에서 코드 체크아웃 — **주의: 현재 있는 브랜치가 아닙니다**
4. 새 디렉토리에서 Claude Code 시작
모든 worktree는 `.claude/worktrees/` 디렉토리 아래에 있습니다.
```
your-project/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← 첫 번째 worktree
│ ├── bugfix-123/ ← 두 번째 worktree
│ └── refactor-api/ ← 세 번째 worktree
├── src/
└── package.json
```
이 경로를 `.gitignore`에 추가하는 것을 권장합니다.
```bash
# .gitignore
.claude/worktrees/
```
#### 현재/특정 브랜치에서 생성
`-w`는 항상 원격 기본 브랜치에서 체크아웃하며, 현재 기본 브랜치 지정은 지원하지 않습니다. 현재 브랜치(또는 특정 브랜치)를 기반으로 worktree를 생성하려면 세 가지 방법이 있습니다.
**방법 1: 수동 Git 생성**
```bash
# 현재 HEAD 기반으로 worktree 생성
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# 또는 특정 브랜치 기반
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# 해당 디렉토리에서 Claude 시작
cd .claude/worktrees/my-feature && claude
```
이 방법은 완전한 제어권을 제공합니다. 임의의 브랜치, 임의의 commit을 기반으로 worktree를 생성할 수 있으며, 작업 디렉토리가 처음부터 올바른 브랜치에 있습니다. 공식 문서에서도 "더 많은 브랜치 및 위치 제어가 필요한 경우, Git으로 직접 worktree를 생성한 다음 해당 디렉토리에서 Claude를 실행하세요"라고 권장합니다.
**방법 2: 대화 중 생성 (권장)**
기존 Claude 세션에서 직접 Claude에게 worktree 생성을 요청합니다.
```
> 从当前分支开启一个worktree
> start a worktree
```
`-w` 명령과 달리, 대화 중에 생성하는 worktree는 **현재 브랜치를 자동으로 기반으로** 하며, 원격 기본 브랜치가 아닙니다. Claude가 자동으로 worktree 생성을 완료하고 전환하며, 전체 과정에서 수동으로 Git 명령을 조작할 필요가 없습니다. 이미 특정 feature 브랜치에서 작업 중이라면, 한 마디로 현재 브랜치 기반의 격리 환경을 만들 수 있는 가장 편리한 방법입니다.
**방법 3: `-w`로 먼저 생성 후, 세션에서 브랜치 전환**
먼저 `claude -w`로 worktree를 생성한 후, 세션에 진입하여 Claude에게 대상 브랜치로 전환을 요청합니다. 이 방법의 단점은 먼저 원격 기본 브랜치를 가져온 후 전환한다는 것입니다. 한 단계가 더 필요하며, 앞의 두 방법보다 깔끔하지 않습니다. 또한 대상 브랜치가 다른 worktree에서 이미 사용 중이면 브랜치 충돌이 발생합니다.
스크린샷에서 보듯이, Claude가 브랜치 충돌을 감지하고 두 가지 선택지를 제공합니다. 메인 디렉토리로 돌아가서 작업하거나, 대상 브랜치 기반으로 새 작업 브랜치를 생성합니다. 최종적으로는 동작하지만, 전체 과정이 방법 1과 방법 2보다 직접적이지 않습니다.
**고급: Makefile로 원클릭 명령 래핑**
현재 브랜치에서 worktree를 자주 생성해야 한다면, 프로젝트 루트 디렉토리의 `Makefile`에 단축 명령을 추가하여 생성 + 에디터 열기 + Claude 시작을 하나의 확정적인 파이프라인으로 연결할 수 있습니다.
```makefile
# 현재 브랜치에서 worktree를 생성하고 개발 환경 시작
# 사용법: make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "사용법: make worktree name="; \
echo "예시: make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ $$(git branch --show-current)에서 worktree 생성: $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ 환경 초기화"
cd .claude/worktrees/$(name) && npm install
@echo "→ Zed에서 열기"
zed .claude/worktrees/$(name)
@echo "→ Claude 시작"
cd .claude/worktrees/$(name) && claude
```
사용법은 매우 간결합니다.
```bash
# 현재 브랜치 기반으로 worktree 생성, Zed로 열기, Claude 시작
make worktree name=fix-login-bug
# 여러 병렬 작업 추가 생성
make worktree name=feature-search
make worktree name=refactor-api
```
여러 개의 Git/cd/claude 명령을 수동으로 입력하는 것에 비해, `make worktree name=xxx` 한 줄이면 되며, 매번 실행되는 과정이 완전히 동일합니다. 특정 단계를 잊거나 경로를 잘못 입력하는 일이 없습니다. 주의할 점은, Makefile이 네이티브 `git worktree add`를 사용하기 때문에 Claude Code의 `WorktreeCreate` Hook이 트리거되지 않습니다(이 Hook은 `claude -w` 또는 대화 중 worktree 생성 시에만 동작합니다). 따라서 환경 초기화 단계(의존성 설치, `.env` 복사 등)는 위 예시의 `npm install`처럼 Makefile에 직접 작성해야 합니다.
### 2단계: 환경 초기화
Worktree 생성이 완료되면, 첫 번째로 할 일은 개발 환경을 초기화하는 것입니다. 각 새 worktree는 독립된 디렉토리이므로, `node_modules`, 가상 환경, `.env` 파일 등은 자동으로 가져오지 않습니다.
Claude Code는 `WorktreeCreate` Hook을 제공하여 환경 설정을 자동화합니다.
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
이렇게 하면 worktree를 생성할 때마다 의존성이 자동으로 설치되고, 환경 변수 파일이 자동으로 복사됩니다. 일반적인 초기화 단계는 다음과 같습니다.
| 프로젝트 유형 | 초기화 명령 |
| ------- | ---------------------------------------------- |
| Node.js | `npm install` 또는 `yarn` |
| Python | `pip install -r requirements.txt` 또는 가상 환경 활성화 |
| Go | `go mod download` |
| 공통 | `.env` 파일 복사, 환경 변수 설정 |
Hook을 설정하지 않았다면, 각 worktree 세션을 시작할 때 `/init`을 실행하여 Claude가 현재 작업 디렉토리의 컨텍스트를 올바르게 이해하고, 프로젝트 구조와 CLAUDE.md 설정을 다시 읽도록 하는 것이 좋습니다.
### 3단계: 커밋과 병합
환경이 준비되고 개발이 완료되면, 다음 단계는 변경 사항을 대상 브랜치에 병합하는 것입니다.
**main 브랜치로 병합**
가장 일반적인 경우로, worktree가 `origin/main`에서 분기되었고 변경 사항도 `main`에 병합해야 합니다. Worktree의 Claude 세션에서 직접 말하면 됩니다.
```
> 모든 변경 사항을 커밋하고, 원격에 푸시한 다음, main으로 PR을 생성해줘
```
Claude가 자동으로 commit → push → `gh pr create`의 전체 과정을 완료합니다.
**feature 브랜치로 병합**
`feature-x` 브랜치에서 개발하고 있다면, worktree의 변경 사항은 `main`이 아닌 `feature-x`로 병합해야 합니다.
```
> 변경 사항을 커밋하고 푸시한 다음, feature-x 브랜치로 PR을 생성해줘
```
Claude가 `gh pr create --base feature-x`를 실행하여 feature 브랜치를 대상으로 하는 PR을 직접 생성합니다.
또한 worktree 세션을 종료한 후(worktree 유지를 선택), 메인 디렉토리로 돌아가서 Claude를 시작할 수도 있습니다.
```
> worktree-my-task 브랜치의 변경 사항을 현재 브랜치에 병합해줘
```
worktree의 일부 commit이 필요 없다면, 선택적으로 cherry-pick할 수 있습니다.
```
> worktree-my-task 브랜치의 커밋 히스토리를 확인하고, 인증 모듈 관련 커밋만 현재 브랜치에 cherry-pick해줘
```
> **팁**: 모든 worktree는 동일한 `.git` 데이터베이스를 공유하므로, worktree에서 생성한 commit은 메인 디렉토리에서 즉시 볼 수 있으며, 추가적인 push/pull 작업이 필요 없습니다.
### 4단계: 종료와 정리
코드 병합이 완료되면, worktree 세션을 종료할 수 있습니다.
worktree 세션을 종료할 때, Claude는 상황에 따라 자동으로 처리합니다.
| 상태 | 처리 방식 |
| --------------- | --------------------- |
| **변경 없음** | 자동으로 worktree와 브랜치 삭제 |
| **변경 또는 커밋 있음** | 유지 또는 삭제를 선택하라고 안내 |
유지된 worktree는 계속 존재하므로, 나중에 작업을 이어갈 수 있습니다.
`WorktreeRemove` Hook을 설정하여 정리를 자동화할 수도 있습니다.
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**수동 관리 명령**
수동으로 worktree를 관리해야 한다면, 표준 Git 명령을 사용할 수 있습니다.
```bash
# 모든 worktree 나열
git worktree list
# 특정 worktree 수동 삭제
git worktree remove .claude/worktrees/feature-auth
# 오래된 worktree 참조 정리
git worktree prune
```
> **주의**: worktree 디렉토리를 직접 `rm -rf`로 삭제하지 마십시오. 올바른 방법은 `git worktree remove`를 사용하는 것이며, 이미 잘못 삭제했다면 `git worktree prune`을 실행하여 남은 참조를 정리하십시오.
## 병렬 개발 패턴
기본 워크플로를 숙지한 후, worktree를 활용하여 병렬 개발을 실현하는 방법을 살펴보겠습니다.
### 다중 터미널 병렬
가장 일반적인 사용법은 여러 터미널 탭에서 동시에 실행하는 것입니다.
```bash
# 터미널 1: 사용자 인증 기능 처리
claude -w feature-auth
# 터미널 2: 결제 버그 수정
claude -w bugfix-payment
# 터미널 3: API 모듈 리팩토링
claude -w refactor-api
```
각 Claude 인스턴스는 자신의 worktree에서 작업하며, 수정 사항이 서로 영향을 미치지 않습니다. 다음과 같이 할 수 있습니다.
* 한 터미널에서 Claude에게 새 기능 개발을 맡기기
* 다른 터미널에서 Claude에게 버그 수정을 맡기기
* 세 번째 터미널에서 직접 코드 리뷰 진행하기
### 경쟁적 구현
효율적인 사용법 중 하나는 여러 Agent가 동일한 기능을 독립적으로 구현하게 하는 것입니다.
```bash
# 세 개의 터미널에서 각각 실행
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
동일한 요구사항 설명을 제공하고, 각자 구현하게 합니다. 마지막에 세 가지 방안을 비교하고 가장 좋은 것을 병합합니다. 이는 LLM의 비결정성을 활용하는 것입니다. 동일한 입력이 다른 출력을 생성할 수 있으며, 때로는 두 번째 버전이 더 나을 수 있습니다.
UI 디자인 탐색도 이 패턴에 매우 적합합니다. 애플리케이션의 인터페이스를 재설계하고 싶지만 어떤 스타일이 더 좋을지 확실하지 않다면:
```bash
# 세 Agent에게 각각 다른 스타일을 구현하게 하기
claude -w ui-minimal # 미니멀 스타일
claude -w ui-colorful # 선명한 배색
claude -w ui-glassmorphism # 글래스모피즘 스타일
```
완료 후 세 버전의 개발 서버를 동시에 실행하고(다른 포트에서), 나란히 비교하여 가장 만족스러운 방안을 메인 브랜치에 병합합니다. 전통적인 "하나 수정하고, 결과 확인하고, 불만족이면 다시 수정하는" 프로세스보다 훨씬 효율적입니다.
### Subagent 격리
Worktree는 메인 Claude 인스턴스뿐만 아니라 Subagent에도 적용할 수 있습니다. 커스텀 Subagent의 frontmatter에 `isolation: worktree`를 추가합니다.
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
대화 중에 Claude에게 직접 말할 수도 있습니다.
```
> 使用 worktree 来隔离你的 agents
> use worktrees for your agents
```
Subagent가 worktree 격리로 설정되면:
```
메인 Agent (메인 디렉토리)
│
├── Migration Agent 1 시작 ──→ worktree-migration-1/
│ └── src/auth/ 디렉토리 처리
│
├── Migration Agent 2 시작 ──→ worktree-migration-2/
│ └── src/api/ 디렉토리 처리
│
└── Migration Agent 3 시작 ──→ worktree-migration-3/
└── src/utils/ 디렉토리 처리
```
각 Subagent가 자신의 worktree에서 독립적으로 작업하며, 서로 간섭하지 않습니다. 완료 후 커밋되지 않은 변경 사항이 없으면 worktree가 자동으로 정리됩니다.
### Tmux 및 IDE와 결합
`--tmux` 매개변수와 함께 사용하면 새 Tmux 세션에서 자동으로 시작되며, 터미널을 닫아도 Claude가 백그라운드에서 계속 실행됩니다.
```bash
claude -w feature-auth --tmux
```
VS Code 또는 Cursor를 사용한다면, 소스 컨트롤 패널이 모든 worktree를 자동으로 인식합니다. 메인 저장소는 하나의 repo로 표시되고, 각 worktree는 독립된 repo로 표시되어 IDE에서 직접 전환, 커밋, 푸시할 수 있습니다. Worktree는 [Ralph 루프](/ko/docs/notes/ralph-wiggum/concept)와도 결합할 수 있으며, 각 Ralph 루프가 자신의 worktree에서 실행되어 루프가 실패하더라도 메인 브랜치에 영향을 미치지 않습니다.
## 주의사항 및 모범 사례
### 흔한 함정
1. **브랜치 출처 혼동**: `-w`로 생성한 worktree는 **원격 기본 브랜치**에서 체크아웃되며, 현재 있는 브랜치가 아닙니다. `feature-x` 브랜치에서 `claude -w my-task`를 실행하면, 새 worktree의 코드는 `origin/main`에서 오며 `feature-x`의 변경 사항을 포함하지 않습니다. 현재 브랜치 기반으로 작업하려면 [현재/특정 브랜치에서 생성](#현재특정-브랜치에서-생성)을 참고하십시오.
2. **커밋되지 않은 변경은 가져오지 않음**: worktree 생성 시, 메인 디렉토리에서 스테이징되지 않았거나 커밋되지 않은 수정 사항은 새 worktree에 나타나지 않습니다. Worktree는 commit 히스토리를 기반으로만 생성되므로, 중요한 변경 사항이 커밋되었는지 확인하십시오.
3. **같은 브랜치를 여러 Worktree에서 사용할 수 없음**: Git은 두 worktree가 동시에 같은 브랜치를 체크아웃하는 것을 허용하지 않습니다. 메인 디렉토리에서 이미 `feature-x` 브랜치에 있다면, worktree에서도 `feature-x`를 체크아웃하려 하면 오류가 발생합니다. 각 worktree는 서로 다른 브랜치에 있어야 합니다.
4. **환경 재초기화 필요**: 각 새 worktree에는 `node_modules` 등 런타임 의존성이 포함되지 않으므로, `WorktreeCreate` Hook을 설정하여 자동화하는 것을 권장합니다([2단계: 환경 초기화](#2단계-환경-초기화) 참고).
### 사용 권장 사항
욕심내지 마십시오. 기술적으로는 많은 worktree를 열 수 있지만, 각 Claude 인스턴스는 API 할당량을 소모하고, 너무 많은 병렬 작업은 추적하기 어려우며, 최종 병합 시 충돌도 더 복잡해집니다.
**명명 규칙**: 좋은 명명 습관을 들여 후속 관리를 편리하게 하십시오.
```bash
# 좋은 명명
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# 좋지 않은 명명
claude -w test
claude -w temp
claude -w 1
```
## 비 Git 버전 관리
SVN, Perforce 또는 Mercurial을 사용한다면, `WorktreeCreate`와 `WorktreeRemove` Hook을 설정하여 유사한 격리 효과를 구현할 수 있습니다. 이러한 Hook을 설정한 후, `--worktree`를 사용할 때 기본 Git 동작 대신 사용자 정의 명령이 호출됩니다.
## 마치며
Worktree는 Claude Code 팀이 매일 사용하는 기능으로, Boris Cherny는 이를 "1순위 생산성 팁"이라고 부릅니다. 핵심 가치는 간단합니다. **여러 Agent가 서로 간섭하지 않고 병렬로 작업할 수 있게 하는 것**입니다.
시작은 간단합니다. 다음 명령만 실행하면 됩니다.
```bash
claude -w your-task-name
```
***
**관련 읽을거리**:
* [Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent) — Subagent와 Worktree의 조합 이해하기
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) — AI 프로그래밍 효율을 높이는 또 다른 방법
* [Claude 시스템 아키텍처 완전 해설](/ko/docs/notes/claude-architecture) — 전체 아키텍처에서 Worktree의 위치 이해하기
**참고 자료**:
* [Claude Code 공식 문서 - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Boris Cherny의 Worktree 발표 공지](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Git Worktree 공식 문서](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**영상 튜토리얼**:
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — worktree 전체 워크플로 실전 데모
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — 수동 worktree 생성 방법
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — 네이티브 worktree 기능 상세 해설
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — 빠른 시작 튜토리얼, Subagent 사용법 포함
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — 다중 worktree 병렬 개발의 장점
# 문서
문서 센터에 오신 것을 환영합니다. 여기에는 제가 정리한 Claude Code 관련 기술 문서와 튜토리얼이 수록되어 있습니다.
# Tmux 빠른 시작 가이드
## 서론
Claude Code의 Agent Teams를 사용해 본 적이 있거나 여러 Claude 인스턴스를 동시에 실행하고 싶다면, Tmux는 거의 필수 도구입니다. Tmux를 사용하면 하나의 터미널 창에서 여러 세션을 실행할 수 있고, 터미널을 닫아도 세션이 백그라운드에서 계속 실행되며, Claude가 Tmux 내에서 자동으로 여러 Agent를 생성하고 관리할 수 있습니다.
이 튜토리얼은 Claude Code 사용자를 위해 설계되었으며, Tmux 기초 지식과 Claude Code와의 통합 기법을 모두 다룹니다.
## Tmux 이해하기
Tmux의 세 가지 핵심 개념:
```
┌─────────────────────────────────────────────────────────┐
│ Session(세션) │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window(윈도우) ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| 개념 | 비유 | 설명 |
| ----------- | ------ | ---------------------------- |
| **Session** | 워크스페이스 | 최상위 컨테이너로, 연결이 끊어져도 계속 실행됩니다 |
| **Window** | 브라우저 탭 | 하나의 세션에 여러 윈도우를 포함할 수 있습니다 |
| **Pane** | 화면 분할 | 하나의 윈도우를 여러 패널로 분할할 수 있습니다 |
## 설치 및 기본 사항
### Tmux 설치
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
설치 확인:
```bash
tmux -V
# 출력 예시: tmux 3.6a
```
### 접두사 키
Tmux의 모든 명령은 **접두사 키**로 시작하며, 기본값은 `Ctrl+B`입니다.
명령 입력 방법:
1. `Ctrl+B`를 누릅니다 (놓지 마세요)
2. 놓은 후 명령 키를 누릅니다
예를 들어, 윈도우 분할: `Ctrl+B` 후 `%`를 누릅니다
## 자주 쓰는 명령어 빠른 참조
### 세션 관리
| 명령어 | 설명 |
| --------------------------- | ------------------- |
| `tmux` | 새 세션 생성 |
| `tmux new -s name` | 이름이 지정된 세션 생성 |
| `tmux ls` | 모든 세션 목록 표시 |
| `tmux attach -t name` | 세션에 연결 |
| `tmux kill-session -t name` | 세션 종료 |
| `Ctrl+B d` | 현재 세션 분리 (백그라운드 실행) |
### 윈도우 관리
| 단축키 | 설명 |
| ------------ | ------------ |
| `Ctrl+B c` | 새 윈도우 생성 |
| `Ctrl+B n` | 다음 윈도우 |
| `Ctrl+B p` | 이전 윈도우 |
| `Ctrl+B 0-9` | 지정된 윈도우로 전환 |
| `Ctrl+B ,` | 현재 윈도우 이름 변경 |
| `Ctrl+B &` | 현재 윈도우 닫기 |
### 패널 관리
| 단축키 | 설명 |
| ------------ | ----------- |
| `Ctrl+B %` | 수직 분할 (좌우) |
| `Ctrl+B "` | 수평 분할 (상하) |
| `Ctrl+B 방향키` | 패널 간 이동 |
| `Ctrl+B x` | 현재 패널 닫기 |
| `Ctrl+B z` | 패널 최대화/복원 |
| `Ctrl+B {` | 패널 왼쪽으로 이동 |
| `Ctrl+B }` | 패널 오른쪽으로 이동 |
### 기타 자주 사용하는 기능
| 단축키 | 설명 |
| ---------- | ----------------- |
| `Ctrl+B [` | 복사 모드 진입 (스크롤 가능) |
| `q` | 복사 모드 종료 |
| `Ctrl+B ?` | 모든 단축키 표시 |
## Claude Code와의 통합
### Claude Code에 Tmux가 필요한 이유
1. **Agent Teams의 Split-pane 모드**: 각 Teammate가 독립된 패널에 표시됩니다
2. **백그라운드 실행**: 터미널을 닫아도 작업이 계속 실행됩니다
3. **세션 지속성**: 연결이 끊어져도 다시 연결하면 완전한 컨텍스트가 복원됩니다
4. **다중 인스턴스 관리**: 여러 Claude 세션을 동시에 실행할 수 있습니다
### 기본 사용법: Claude 백그라운드 실행
```bash
# tmux에서 Claude 시작
tmux new -s claude-work
claude
# 세션 분리 (Claude 계속 실행)
# Ctrl+B d
# 나중에 다시 연결
tmux attach -t claude-work
```
### --tmux 매개변수 사용
Claude Code는 Tmux 통합을 기본 지원합니다:
```bash
# 새 tmux 세션에서 Claude 시작
claude --tmux
# worktree와 함께 사용
claude -w feature-auth --tmux
```
이 명령은 자동으로:
1. 새 tmux 세션을 생성합니다
2. 그 안에서 Claude Code를 시작합니다
3. 세션 이름을 `claude-{랜덤ID}`로 지정합니다
### Agent Teams의 Tmux 모드
Agent Teams는 split-pane 표시 모드를 사용할 수 있으며, 각 Teammate가 독립된 패널에서 실행됩니다:
```json
// settings.json
{
"teammateMode": "tmux"
}
```
또는 명령줄을 통해:
```bash
claude --teammate-mode tmux
```
효과:
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## 실용적인 설정
### 권장 \~/.tmux.conf
`~/.tmux.conf`를 생성하거나 편집합니다:
```bash
# Ctrl+A를 접두사 키로 사용 (누르기 더 편함)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# 마우스 지원 활성화
set -g mouse on
# 히스토리 버퍼 증가 (Claude 출력이 많음)
set -g history-limit 50000
# vim 스타일 패널 탐색
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# 더 직관적인 분할 단축키
bind | split-window -h
bind - split-window -v
# 설정 빠른 리로드
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# 256색 지원
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# 윈도우 번호를 1부터 시작 (0은 너무 멀리 있음)
set -g base-index 1
setw -g pane-base-index 1
# 상태 표시줄 최적화
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
설정 리로드:
```bash
tmux source-file ~/.tmux.conf
```
### Claude Code 전용 설정
Claude Code에 최적화된 설정:
```bash
# Claude 세션 팝업 단축키
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
이 설정의 효과:
1. `Ctrl+A y`를 눌러 Claude 팝업을 엽니다
2. 각 디렉토리마다 독립된 Claude 세션이 있습니다
3. 팝업을 닫아도 세션은 계속 실행됩니다
4. 다시 열면 이전 대화가 복원됩니다
## 일반적인 워크플로
### 워크플로 1: 다중 프로젝트 병렬 작업
```bash
# 각 프로젝트에 독립 세션 생성
tmux new -s project-a
# 안에서 Claude 시작
claude -w feature-x
# 분리 후 다른 세션 생성
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# 세션 간 전환
tmux switch -t project-a
tmux switch -t project-b
# 또는 모든 세션을 나열하여 선택
# Ctrl+B s
```
### 워크플로 2: 개발 대시보드
다중 패널 개발 환경 생성:
```bash
# 세션 생성
tmux new -s dev
# 세 개의 패널로 분할
# Ctrl+B % (수직 분할)
# Ctrl+B " (오른쪽 수평 분할)
# 패널 레이아웃:
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# 첫 번째 패널에서 Claude 실행
claude
# 두 번째 패널로 전환 (Ctrl+B 오른쪽 화살표)
tail -f logs/app.log
# 세 번째 패널로 전환
npm test -- --watch
```
### 워크플로 3: 원격 개발
Tmux의 가장 강력한 기능은 세션 지속성으로, 특히 SSH 원격 개발에 적합합니다:
```bash
# 원격 서버에 연결
ssh user@server
# tmux 세션 생성
tmux new -s remote-claude
# Claude 시작
claude
# SSH 연결 끊기 (Claude 계속 실행)
# Ctrl+B d
exit
# 나중에 다시 연결
ssh user@server
tmux attach -t remote-claude
# Claude 세션 완전 복원
```
### 워크플로 4: Agent Teams 모니터링
tmux를 사용하여 Agent Teams의 모든 Teammate를 모니터링합니다:
```bash
# Claude를 tmux 모드로 시작
claude --teammate-mode tmux
# Agent Team 생성
# "agent team을 만들어서 코드를 리뷰..."
# 화면이 자동으로 분할되며, 각 Teammate마다 하나의 패널이 생성됩니다
# 다른 패널을 클릭하여 해당 Teammate와 직접 소통할 수 있습니다
```
## 문제 해결
### 자주 발생하는 문제
| 문제 | 해결 방법 |
| ---------------- | --------------------------------- |
| 색상이 올바르게 표시되지 않음 | `TERM=xterm-256color`를 확인하세요 |
| 마우스가 작동하지 않음 | 설정에 `set -g mouse on`을 추가하세요 |
| 복사/붙여넣기 문제 | 복사 모드에서 `Enter`를 사용하여 복사하세요 |
| 세션이 사라짐 | `tmux ls`를 확인하세요, 시스템 재시작일 수 있습니다 |
### 고아 세션 정리
Claude Code는 때때로 정리되지 않은 tmux 세션을 남길 수 있습니다:
```bash
# 모든 세션 나열
tmux ls
# 특정 세션 종료
tmux kill-session -t session-name
# 모든 세션 종료 (주의!)
tmux kill-server
```
### iTerm2 사용자
macOS의 iTerm2를 사용하는 경우, 네이티브 통합을 활용할 수 있습니다:
```bash
# iTerm2의 tmux 통합 모드 사용
tmux -CC
# 또는 Claude Code에서
claude --teammate-mode tmux
```
iTerm2는 자동으로 tmux 패널을 네이티브 탭과 화면 분할로 변환합니다.
## 사용 후기
### Tmux가 필요한 경우
| 시나리오 | Tmux 필요 여부 |
| ----------------- | ---------- |
| 간단한 일회성 Claude 대화 | 불필요 |
| 장시간 실행되는 작업 | 필요 |
| Agent Teams | 강력 추천 |
| 원격 개발 | 필수 |
| 다중 프로젝트 병렬 작업 | 추천 |
### 최소 설정
설정을 복잡하게 하고 싶지 않다면, 다음 몇 가지 명령만 기억하면 됩니다:
```bash
# 세션 생성
tmux new -s work
# 분리 (백그라운드 실행)
Ctrl+B d
# 다시 연결
tmux attach -t work
# 패널 분할
Ctrl+B % # 좌우 분할
Ctrl+B " # 상하 분할
# 패널 전환
Ctrl+B 방향키
```
### Claude Code와의 최적 조합
1. **Worktree + Tmux**: 각 worktree를 독립된 tmux 세션에서 실행합니다
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux**: 모든 Teammate를 시각적으로 관리합니다
```bash
claude --teammate-mode tmux
```
3. **장기 작업 + 분리**: 시작 후 분리하고, 나중에 돌아와서 확인합니다
```bash
# 시작
tmux new -s migration
claude
# "데이터베이스 마이그레이션 실행..."
# Ctrl+B d
# 몇 시간 후
tmux attach -t migration
```
## 마무리
Tmux는 Claude Code를 효율적으로 사용하기 위한 핵심 도구이며, 특히 다음과 같은 시나리오에서 중요합니다:
| 핵심 | 설명 |
| --------- | --------------------------------- |
| **지속성** | 연결이 끊어져도 세션이 유실되지 않습니다 |
| **병렬 처리** | 여러 Claude 인스턴스를 동시에 관리합니다 |
| **시각화** | Agent Teams의 split-pane 표시를 지원합니다 |
세 가지 핵심 명령만으로 시작할 수 있습니다:
* `tmux new -s name` 세션 생성
* `Ctrl+B d` 세션 분리
* `tmux attach -t name` 다시 연결
***
**관련 읽을거리**:
* [Claude Agent Teams 완벽 가이드](/ko/docs/notes/claude-agent-teams) — Agent Teams는 split-pane 모드를 위해 Tmux가 필요합니다
* [Claude Worktree 완벽 가이드](/ko/docs/notes/claude-worktree) — Worktree는 Tmux와 함께 백그라운드에서 실행할 수 있습니다
**참고 자료**:
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**동영상 튜토리얼**:
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) — Tmux 기초 입문
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) — Claude Code 통합 워크플로
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) — 고급 설정 기법
# MVP 스프린트: 2주 만에 핵심 기능 완성하기
이것은 테스트 글입니다.
# Apple 개발자 계정 등록하기
직접 개발한 앱을 App Store에 출시하려면, 첫 번째 단계는 Apple Developer Program(애플 개발자 프로그램)에 등록하는 것입니다. 연간 ¥688(99달러)로, 모든 iOS 인디 개발자가 반드시 지출해야 하는 비용입니다.
이 글에서는 계정 유형의 차이점, 등록 전에 준비해야 할 사항, 그리고 전체 등록 절차를 안내합니다.
## 계정 유형 비교
Apple Developer Program에는 세 가지 계정 유형이 있으며, 각각 다른 개발 시나리오에 적합합니다.
| 특성 | 개인 계정 | 조직 계정 | 기업 계정 |
| ------------ | ------------------ | ------------------- | ------------ |
| 연회비 | ¥688($99) | ¥688($99) | ¥1,988($299) |
| App Store 출시 | ✅ | ✅ | ❌(내부 배포만 가능) |
| 개발자 이름 표시 | 개인 이름 | 조직/회사명 | 조직명 |
| 팀 멤버 관리 | ❌ | ✅ | ✅ |
| D-U-N-S 번호 | 불필요 | 필요 | 필요 |
| 심사 기간 | 비교적 빠름(보통 48시간 이내) | 비교적 느림(조직 정보 확인 필요) | 비교적 느림 |
| 적합한 대상 | 인디 개발자, 개인 | 회사, 스튜디오 | 대기업 내부 앱 |
**인디 개발자의 선택**: 개인 개발자라면 **개인 계정**을 선택하면 됩니다. 절차가 가장 간단하고 심사도 가장 빠르며, 기능도 충분합니다. App Store에 표시되는 개발자 이름은 실명이 됩니다.
## 등록 전 준비
### 필수 조건
등록을 시작하기 전에 다음 항목을 준비했는지 확인하세요.
* **Apple ID**: 아직 없다면 [appleid.apple.com](https://appleid.apple.com)에서 생성하세요. 자주 사용하는 이메일로 등록하는 것을 권장합니다. 이후 모든 개발 관련 알림이 이 이메일로 발송됩니다.
* **이중 인증**: Apple ID에 이중 인증(Two-Factor Authentication)이 반드시 활성화되어 있어야 합니다. iPhone에서 「설정 → Apple ID → 로그인 및 보안 → 이중 인증」으로 이동하여 활성화하세요.
* **Apple 기기**: 등록 과정에서 iPhone 또는 iPad를 통한 본인 확인이 필요하며, Apple Developer 앱을 다운로드해야 합니다.
### 조직 계정 추가 요구사항
조직 계정을 등록하는 경우, 다음 사항도 필요합니다.
* **D-U-N-S 번호**: Dun & Bradstreet 공식 웹사이트에서 미리 신청하세요. 심사에 5\~14영업일이 소요됩니다.
* **법인 신분**: 등록자는 조직의 법인 또는 위임된 대리인이어야 합니다.
* **조직 정보**: 등록 주소, 법인 이름, 연락처 등이 포함됩니다.
## 개발 기기 준비
개발자 계정 등록은 첫 번째 단계에 불과합니다. iOS 개발에는 몇 가지 하드웨어와 소프트웨어 도구도 필요합니다.
### 필수 기기
* **Mac 컴퓨터** — Xcode는 macOS에서만 실행되므로 이것은 필수 조건입니다. Apple Silicon(M 시리즈 칩) Mac을 권장합니다. 컴파일 속도가 빠르고 iOS 시뮬레이터를 직접 실행할 수 있습니다. MacBook Air M 시리즈로도 인디 개발에 충분하며, 예산이 한정적이라면 Mac mini를 고려할 수 있습니다.
* **iPhone / iPad(권장하지만 필수는 아님)** — 시뮬레이터가 대부분의 디버깅 시나리오를 커버하지만, 실제 기기 테스트는 성능, 센서(카메라/GPS/NFC), 푸시 알림 등의 측면에서 대체할 수 없습니다. 유료 개발자 계정이 없어도 무료 Apple ID로 실제 기기에서 디버깅할 수 있습니다(단, 7일 재서명 등의 제한이 있으며, 자세한 내용은 하단 Q\&A를 참조하세요).
### 개발 도구
* **Xcode** — Apple 공식 IDE로, Mac App Store에서 무료로 다운로드할 수 있습니다. 용량이 큰 편(약 12GB 이상)이므로 첫 설치 시 인내심이 필요합니다.
* **Apple Developer App** — 계정 등록, WWDC 영상 및 문서 확인에 사용합니다.
* **TestFlight** — 내부 테스트 배포 도구로, 사용자를 초대하여 앱을 테스트하는 공식 채널입니다.
### 주의사항
* macOS와 Xcode 버전은 최신 상태를 유지해야 합니다. Apple은 매년 WWDC 이후 새 버전의 Xcode를 출시하며, 보통 최근 1\~2개의 macOS 메이저 버전을 요구합니다.
* Xcode 업데이트가 빈번하고 용량이 크므로, 충분한 디스크 공간을 확보해 두세요(최소 50GB 이상).
* 개발하는 앱이 하드웨어 기능(카메라, 블루투스, NFC 등)을 사용하는 경우, 실제 기기 테스트가 필수입니다.
* Mac이 없는 경우, 클라우드 Mac 서비스(MacStadium, AWS EC2 Mac 등)가 대안이 될 수 있지만, 네이티브 기기만큼의 경험을 제공하지는 않습니다.
## 등록 절차
### 1단계: Apple Developer 앱 다운로드
iPhone 또는 iPad에서 App Store를 열고 「Apple Developer」를 검색하여 다운로드 및 설치합니다.
### 2단계: 로그인 및 등록 시작
Apple Developer 앱을 열고 Apple ID로 로그인합니다. 「계정」 탭을 누른 다음 「지금 Apple Developer Program 등록하기」를 누릅니다.
### 3단계: 정보 입력 및 본인 확인
안내에 따라 개인 정보를 입력합니다.
1. **신원 정보 확인**: 이름, 주소 등 기본 정보를 확인합니다.
2. **본인 확인**: 거주 지역에 따라 앱에서 정부 발급 유효 신분증(여권, 운전면허 등) 촬영이나 셀카를 요청할 수 있습니다.
3. **약관 동의**: Apple Developer Program 사용권 계약을 읽고 동의합니다.
> 본인 확인 단계에서는 조명이 밝은 환경에서 진행하여 사진이 선명하게 촬영되도록 하세요. 전체 등록 절차는 동일한 기기에서 완료해야 합니다.
### 4단계: 연회비 결제
정보가 정확한지 확인한 후 연회비 ¥688($99)를 결제합니다. Apple ID에 연결된 결제 수단으로 결제할 수 있습니다. 결제가 완료되면 확인 이메일을 받게 됩니다.
### 5단계: 심사 대기
* **개인 계정**: 보통 48시간 이내에 심사가 완료됩니다. 저는 3월 14일에 결제했고, 3월 15일 오전에 환영 이메일을 받았습니다. 24시간도 걸리지 않았습니다.
* **조직 계정**: Apple에서 조직 정보와 D-U-N-S 번호를 확인하므로, 더 오래 걸릴 수 있습니다.
심사가 완료되면 [developer.apple.com](https://developer.apple.com)에서 개발자 대시보드에 로그인하여 모든 개발 리소스에 접근할 수 있습니다.
## 구독 관리 및 갱신
Apple Developer Program은 연간 구독 방식으로, 매년 ¥688가 자동 갱신됩니다.
### 자동 갱신
기본적으로 자동 갱신이 활성화되어 있으며, 만료 전에 Apple ID에 연결된 결제 수단에서 자동으로 결제됩니다. 자동 갱신을 유지하여 계정 만료로 인해 출시된 앱에 영향이 가지 않도록 하는 것을 권장합니다(자세한 내용은 아래 자주 묻는 질문 참조).
### 구독 취소 또는 관리
갱신 설정을 변경해야 하는 경우, iPhone에서 「설정 → Apple ID → 구독」을 열고 Apple Developer Program을 찾아 관리하세요.
## 자주 묻는 질문
**Q: 갱신을 잊어버리면 어떻게 되나요?**
계정이 만료되면 앱이 App Store에서 내려가지만, 삭제되지는 않습니다. 다시 결제하면 앱이 복구됩니다. 다만 그 사이에 사용자 다운로드와 업데이트가 영향을 받으므로, 자동 갱신을 활성화하는 것을 권장합니다.
**Q: D-U-N-S 번호는 어떻게 신청하나요?**
아래 링크를 방문하여 회사 정보를 입력하고 신청서를 제출하세요. 심사에는 보통 5\~14영업일이 소요됩니다. 개인 계정에는 이 번호가 필요하지 않습니다.
**Q: 문제가 발생하면 Apple에 어떻게 연락하나요?**
[Apple Developer 지원](https://developer.apple.com/contact/)을 방문하면 온라인 채팅이나 전화로 연락할 수 있습니다. 한국어 지원도 가능하며, 응답 속도가 꽤 좋습니다.
**Q: 먼저 개발하고 나중에 등록해도 되나요?**
먼저 체험해 볼 수는 있지만, 무료 계정의 제한 사항을 알아둬야 합니다. 무료 Apple ID만으로도 Xcode에서 코드를 작성하고 시뮬레이터로 디버깅할 수 있으며, 자신의 기기에 설치하여 실행할 수도 있습니다. Swift를 배우고 기본 UI 아이디어를 검증하는 데는 충분합니다.
하지만 무료 계정에는 여러 제한이 있습니다. 실제 기기에 설치한 앱은 7일마다 다시 컴파일하여 설치해야 하며, 플랫폼당 최대 3대의 기기만 사용할 수 있습니다. 또한 푸시 알림, iCloud, TestFlight, 앱 내 구매 등의 기능을 사용할 수 없습니다. 앱에 이러한 기능이 필요하다면 개발 단계에서부터 유료 멤버십이 필요합니다. App Store 출시만을 위한 것이 아닙니다.
권장 사항: Swift 입문을 배우고 데모를 실행하는 정도라면 무료 계정을 먼저 사용해도 됩니다. 하지만 정식 프로젝트를 시작하면 가능한 한 빨리 유료 멤버십에 등록하여, 이후 기능 제한으로 인한 진행 지연을 피하세요.
# 아이디어 검증: 막연한 영감에서 실행 가능한 방향으로
이것은 테스트 글입니다.
# 기술 스택 선택: Next.js + Supabase를 선택한 이유
이것은 테스트 글입니다.
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# 개념 소개
## 서론
Claude Code에서 완벽한 워크플로를 구성했을 때 — 커스텀 명령어, 코드 리뷰 훅, 전용 Skills — 이런 생각이 들 수 있습니다: 이것들을 패키징해서 팀이나 커뮤니티에 공유할 수는 없을까?
바로 이것이 Plugin이 해결하는 문제입니다.
Skills가 AI를 위한 "작업 매뉴얼"이라면, Plugin은 "도구 상자"입니다. Skills, Commands, Hooks, MCP 서버 등 모든 설정을 하나로 묶어 패키징하여 단일 명령으로 설치하고 배포할 수 있게 해줍니다.
## Plugin 이해하기
오랜 세월 동안 신뢰할 수 있는 도구들을 축적해 온 숙련된 장인을 상상해 보세요: 망치, 톱, 자, 각종 드라이버. 작업대를 옮길 때마다 도구를 하나하나 옮기고 다시 정리해야 합니다. Plugin은 잘 설계된 도구 상자와 같습니다. 모든 도구를 담을 수 있을 뿐만 아니라 카테고리별로 깔끔하게 정리되어 있어, 어디로 가져가든 바로 작업을 시작할 수 있습니다.
기술적 관점에서 Plugin은 Claude Code의 확장 패키지 메커니즘입니다. Plugin에는 다음을 포함할 수 있습니다:
| 컴포넌트 | 역할 | 파일 위치 |
| ------------- | ---------------- | ----------- |
| **슬래시 명령어** | 빠른 작업 진입점 | `commands/` |
| **Subagents** | 전문화된 서브 에이전트 | `agents/` |
| **Skills** | AI 지식 패키지 | `skills/` |
| **Hooks** | 이벤트 트리거 자동화 스크립트 | `hooks/` |
| **MCP 서버** | 외부 시스템 연결 | `.mcp.json` |
| **LSP 서버** | 언어 서버 설정 | `.lsp.json` |
이러한 컴포넌트들이 함께 작동하여 완전한 워크플로 솔루션을 형성합니다.
## Plugin vs 독립 설정
Claude Code에서는 설정을 프로젝트의 `.claude/` 디렉토리에 배치할 수도 있고, Plugin으로 패키징할 수도 있습니다. 두 가지의 핵심적인 차이는 **배포 방식**과 **네임스페이스**에 있습니다:
| 측면 | 독립 설정 (`.claude/`) | Plugin |
| ------- | ------------------- | ---------------------- |
| 명령어 이름 | `/hello` | `/plugin-name:hello` |
| 사용 사례 | 개인 워크플로, 프로젝트 특정 설정 | 팀 공유, 커뮤니티 배포 |
| 버전 관리 | 프로젝트 코드와 함께 관리 | Semantic Versioning 지원 |
| 업데이트 방식 | 수동 동기화 | 자동 업데이트 지원 |
| 충돌 처리 | 다른 설정과 충돌 가능성 있음 | 네임스페이스 격리 |
**Plugin을 선택해야 할 때**:
* 팀 멤버와 워크플로 설정을 공유해야 할 때
* 여러 프로젝트에서 같은 도구 세트를 재사용하고 싶을 때
* 커뮤니티에 설정을 배포할 계획이 있을 때
* 버전 관리와 자동 업데이트가 필요할 때
**독립 설정을 사용해야 할 때**:
* 개인적인 빠른 실험
* 재사용이 필요 없는 프로젝트 특정 설정
* 간단한 일회성 명령어
## Plugin 디렉토리 구조
표준 Plugin 구조는 다음과 같습니다:
```
my-plugin/
├── .claude-plugin/ # 元数据目录
│ └── plugin.json # 必需:插件清单
├── commands/ # 斜杠命令
│ ├── review.md
│ └── deploy.md
├── agents/ # 子代理
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # 事件钩子
│ └── hooks.json
├── scripts/ # 辅助脚本
│ └── format-code.sh
├── .mcp.json # MCP 服务器配置
└── .lsp.json # LSP 服务器配置
```
**핵심 주의사항**:
* `plugin.json`은 반드시 `.claude-plugin/` 디렉토리 안에 배치해야 합니다
* 나머지 디렉토리(commands, agents, skills 등)는 플러그인 루트에 배치합니다
* 기능 디렉토리를 `.claude-plugin/` 안에 넣지 마세요
## 핵심 설정 파일
Plugin의 핵심은 `.claude-plugin/plugin.json`으로, 플러그인의 메타데이터와 컴포넌트 경로를 정의합니다:
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "一个示例插件",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| 필드 | 필수 | 설명 |
| ------------- | --- | ------------------------ |
| `name` | 예 | 플러그인 고유 식별자, 소문자와 하이픈 사용 |
| `version` | 아니오 | 시맨틱 버전 번호 |
| `description` | 아니오 | 플러그인 간단한 설명 |
| `author` | 아니오 | 저자 정보 |
| `keywords` | 아니오 | 검색용 태그 |
| `commands` | 아니오 | 명령어 파일 또는 디렉토리 경로 |
| `agents` | 아니오 | 에이전트 파일 또는 디렉토리 경로 |
| `skills` | 아니오 | Skills 디렉토리 경로 |
| `hooks` | 아니오 | 훅 설정 경로 |
| `mcpServers` | 아니오 | MCP 설정 경로 |
## 설치 범위
Plugin은 다양한 사용 사례에 대응하기 위해 네 가지 설치 범위를 지원합니다:
| 범위 | 설정 파일 | 용도 |
| --------- | ----------------------------- | ------------------------ |
| `user` | `~/.claude/settings.json` | 개인 플러그인, 모든 프로젝트에서 사용 가능 |
| `project` | `.claude/settings.json` | 팀 플러그인, 버전 관리를 통해 공유 |
| `local` | `.claude/settings.local.json` | 프로젝트 특정, gitignore 대상 |
| `managed` | `managed-settings.json` | 기업 관리 (읽기 전용) |
기본 설치 범위는 `user`입니다. 플러그인 설정을 Git에 커밋하여 팀에서 사용하려면 `project` 범위를 선택하세요.
## 핵심 장점
### 네임스페이스 격리
Plugin의 명령어에는 네임스페이스 접두사가 붙습니다(예: `/my-plugin:review`). 이를 통해 다른 플러그인이나 프로젝트 설정과의 이름 충돌을 방지합니다. 팀 협업에서 특히 중요합니다 — 서로 다른 팀이 개발한 플러그인이 문제없이 공존할 수 있습니다.
### 버전 관리
Plugin은 Semantic Versioning을 지원하여 다음이 가능합니다:
* 플러그인의 변경 이력 추적
* 필요시 이전 버전으로 롤백
* 호환 가능한 업데이트 자동 수신
### 간편한 배포
Plugin Marketplace를 통해 다음이 가능합니다:
* GitHub에서 플러그인 호스팅
* 사용자가 간단한 명령으로 설치 가능
* 의존성과 업데이트 자동 처리
### 팀 협업
Plugin은 팀 시나리오에 특히 적합합니다:
* 팀의 개발 도구 체인 통일
* 새 멤버가 단일 명령으로 모든 도구 확보
* 중앙 집중식 설정 관리로 중복 작업 감소
## Plugin 생태계
Claude Code의 Plugin 생태계는 빠르게 성장하고 있습니다. 2025년 초 기준으로, 생태계는 상당한 규모에 도달했습니다:
* **229개 이상의 플러그인**이 생태계에서 활발하게 사용 중
* **239개의 Agent Skills**가 마켓플레이스 전반에 분포
* **200개 이상의 MCP 서버**가 Docker 툴킷에 사전 구축
**공식 리소스**:
| 리소스 | 링크 | 설명 |
| ---------------------------------- | -------------------------------------------------------------------------------------- | --------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 공식 Skills 리포지토리 |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 공식 플러그인 디렉토리 |
| Docker MCP Toolkit | [공식 사이트](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200개 이상의 사전 구축 MCP 서버 |
**커뮤니티 셀렉션**:
| 리소스 | 링크 | 설명 |
| ---------------------- | ----------------------------------------------------------------- | ---------------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243개 플러그인 자동 수집 |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 모범 사례 모음 |
| claude-plugins.dev | [공식 사이트](https://claude-plugins.dev/) | 커뮤니티 레지스트리 및 CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99개 에이전트 + 15개 오케스트레이터 |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17개 전문 에이전트 |
## 다른 기능과의 관계
Plugin은 "컨테이너" 개념으로, Claude Code 생태계의 다른 기능을 포함할 수 있습니다:
```
Plugin(컨테이너)
├── Skills(지식 패키지)
├── Commands(빠른 명령어)
├── Agents(서브 에이전트)
├── Hooks(이벤트 훅)
└── MCP/LSP(외부 연결)
```
이 계층 관계를 이해하는 것이 중요합니다:
* **Skills**는 Claude에게 무언가를 하는 방법을 가르칩니다
* **Commands**는 빠른 트리거 진입점을 제공합니다
* **Agents**는 독립적인 전문 작업을 처리합니다
* **Hooks**는 이벤트 기반 자동화를 구현합니다
* **Plugin**은 이 모든 것을 묶어 배포와 관리를 용이하게 합니다
### Skills vs Plugins
처음 접하는 분들은 Skills와 Plugins의 차이에 혼란을 느낄 수 있습니다. [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins)의 분석에 따르면:
| 특성 | Skills | Plugins |
| --------- | ----------------------------- | ------------------------------------ |
| **범위** | 모든 Claude 제품 (Web, API, Code) | Claude Code 전용 |
| **포함 내용** | Markdown 가이드 + 선택적 스크립트 | Commands, Agents, Hooks, MCP, Skills |
| **활성화** | 자동 (모델이 사용 시점 결정) | 가변적 (컴포넌트 유형에 따라 다름) |
| **최적 용도** | Claude에게 도메인 전문 지식 교육 | Claude Code 환경 확장 |
| **배포 방식** | GitHub 리포지토리, 파일 시스템 | 탈중앙화 Marketplace |
**핵심 인사이트**: Skills는 모델이 자동으로 트리거하며 수동 호출이 필요 없습니다. Plugins는 패키징 메커니즘으로 분산 공유의 과제를 해결합니다. 두 가지는 함께 사용할 수 있습니다 — Plugin은 Skills를 포함할 수 있습니다.
## 요약
Claude Code Plugin은 본질적으로 **워크플로 패키징 및 배포 메커니즘**입니다. 설정 재사용과 팀 협업의 어려움을 해결하여, 정성껏 구축한 도구 체인을 더 많은 사람과 공유할 수 있게 해줍니다.
세 가지 키워드를 기억하세요:
| 키워드 | 의미 |
| ------- | ---------------------- |
| **패키징** | 여러 설정 컴포넌트를 하나의 단위로 통합 |
| **격리** | 네임스페이스로 충돌 방지 |
| **배포** | Marketplace를 통한 손쉬운 공유 |
개념을 이해했으니, 다음 글 [Claude Code Plugin 실전 가이드](/ko/docs/notes/claude-plugin/practice)에서 직접 실습합니다: Plugin을 처음부터 만들고, Marketplace에 게시하고, 팀 협업의 모범 사례를 알아봅니다.
Plugin에 포함할 수 있는 컴포넌트가 아직 익숙하지 않다면, 먼저 [Claude Skills란 무엇인가](/ko/docs/notes/claude-skills/concept)를 읽어 Skills의 핵심 개념을 이해하시기 바랍니다.
# 실전 가이드
## 빠른 복습
이전 글에서 Plugin의 핵심 개념을 알아보았습니다. Plugin은 Claude Code의 워크플로를 패키징하고 배포하는 메커니즘으로, Commands, Skills, Agents, Hooks 등의 컴포넌트를 하나로 통합하여 팀 공유 및 커뮤니티 배포를 가능하게 합니다. 이번 글에서는 실전적인 관점에서 생성부터 배포까지의 전체 프로세스를 함께 진행하겠습니다.
## 첫 번째 Plugin 만들기
### 1단계: 디렉토리 구조 생성
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### 2단계: 플러그인 매니페스트 생성
`.claude-plugin/plugin.json`에 플러그인의 메타데이터를 정의합니다:
```json
{
"name": "my-first-plugin",
"description": "我的第一个 Claude Code 插件",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### 3단계: 슬래시 명령어 추가
`commands/` 디렉토리에 Markdown 파일을 생성합니다. 각 파일이 하나의 명령어에 해당합니다:
`commands/hello.md`:
```markdown
---
description: 向用户发送友好的问候
---
# Hello 命令
请热情地问候用户,并询问今天可以帮助他们做什么。
```
### 4단계: 플러그인 테스트
`--plugin-dir` 플래그를 사용하여 로컬 플러그인을 로드하고 테스트합니다:
```bash
claude --plugin-dir ./my-first-plugin
```
Claude Code에서 명령어를 실행합니다:
```
/my-first-plugin:hello
```
### 5단계: 명령어 인자 추가
명령어는 사용자가 입력한 인자를 받을 수 있습니다. `hello.md`를 업데이트합니다:
```markdown
---
description: 向指定用户发送个性化问候
---
# Hello 命令
请热情地问候名为 "$ARGUMENTS" 的用户,并询问今天可以帮助他们做什么。
如果用户没有提供名字,就使用"朋友"作为称呼。
```
인자를 포함하여 명령어를 테스트합니다:
```
/my-first-plugin:hello 小明
```
**지원되는 인자 플레이스홀더**:
* `$ARGUMENTS` - 모든 사용자 입력
* `$1`, `$2`, `$3` - 개별 인자
## 컴포넌트 추가하기
### Skills 추가
`skills/` 디렉토리를 생성합니다. 각 Skill은 `SKILL.md` 파일을 포함하는 폴더입니다:
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md`:
```yaml
---
name: code-review
description: 审查代码质量、安全性和可维护性
---
当审查代码时,请检查以下方面:
1. **代码组织**:结构是否清晰
2. **错误处理**:异常是否被妥善处理
3. **安全隐患**:是否存在安全漏洞
4. **测试覆盖**:关键逻辑是否有测试
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Subagents 추가
`agents/` 디렉토리를 생성합니다:
`agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家。
当被调用时:
1. 运行 git diff 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(如注入、敏感信息泄露)
- 性能优化机会
```
### Hooks 추가
Hooks를 사용하면 특정 이벤트가 발생할 때 스크립트를 자동으로 실행할 수 있습니다. `hooks/hooks.json`을 생성합니다:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**중요**: `${CLAUDE_PLUGIN_ROOT}` 환경 변수를 사용하여 플러그인 디렉토리 내의 파일을 참조하세요. 이렇게 하면 플러그인이 어디에 설치되더라도 경로가 올바르게 해석됩니다.
대응하는 스크립트 `scripts/format-code.sh`를 생성합니다:
```bash
#!/bin/bash
# 格式化代码
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
실행 권한을 부여하는 것을 잊지 마세요:
```bash
chmod +x scripts/format-code.sh
```
### MCP 서버 추가
플러그인이 외부 시스템에 연결해야 하는 경우, `.mcp.json`을 생성합니다:
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## 전체 Plugin 구조
모든 기능을 갖춘 Plugin은 다음과 같은 구조를 가집니다:
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单
├── commands/
│ ├── review.md # 代码审查命令
│ ├── deploy.md # 部署命令
│ └── test.md # 测试命令
├── agents/
│ ├── code-reviewer.md # 代码审查代理
│ └── debugger.md # 调试代理
├── skills/
│ └── code-standards/
│ └── SKILL.md # 代码规范知识
├── hooks/
│ └── hooks.json # 事件钩子配置
├── scripts/
│ ├── format-code.sh # 格式化脚本
│ └── run-tests.sh # 测试脚本
├── .mcp.json # MCP 配置
├── LICENSE
├── README.md
└── CHANGELOG.md
```
대응하는 `plugin.json`:
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "开发者工具箱:代码审查、测试、部署一站式解决方案",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## Marketplace에 배포하기
### Marketplace란
Marketplace는 Plugin의 배포 센터입니다. "플러그인 스토어"라고 생각하시면 됩니다. 사용자는 간단한 명령어 하나로 여러분이 배포한 플러그인을 설치할 수 있습니다.
### Marketplace 설정 생성
GitHub 저장소에 `.claude-plugin/marketplace.json`을 생성합니다:
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "开发者工具箱",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "文档生成工具"
}
]
}
```
### 플러그인 소스 유형
Marketplace는 여러 가지 소스 유형을 지원합니다:
**상대 경로** (동일 저장소 내 플러그인):
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**GitHub 저장소**:
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**임의의 Git 저장소**:
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### 배포 절차
1. **GitHub 저장소를 생성합니다**
2. **코드를 푸시합니다**:
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **사용자가 여러분의 Marketplace를 추가합니다**:
```bash
/plugin marketplace add your-username/your-repo
```
4. **사용자가 플러그인을 설치합니다**:
```bash
/plugin install dev-toolkit@your-marketplace
```
## Plugin 설치 및 관리
### 대화형 메뉴로
```bash
/plugin
```
이 명령을 실행하면 플러그인을 탐색, 설치, 활성화, 비활성화할 수 있는 대화형 인터페이스가 열립니다.
### 명령줄로
**Marketplace 추가**:
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # 本地
```
**플러그인 설치**:
```bash
# 安装到用户范围(默认)
/plugin install formatter@my-marketplace
# 安装到项目范围(团队共享)
/plugin install formatter@my-marketplace --scope project
# 安装到本地范围(gitignored)
/plugin install formatter@my-marketplace --scope local
```
**기타 관리 명령어**:
```bash
/plugin enable # 启用插件
/plugin disable # 禁用插件
/plugin uninstall # 卸载插件
/plugin update # 更新插件
```
### 플러그인 검증
배포 전에 플러그인 설정이 올바른지 검증합니다:
```bash
claude plugin validate .
```
또는 Claude Code 내에서:
```
/plugin validate .
```
## 팀 협업 설정
### 프로젝트 내 플러그인 설정 공유
플러그인 설정을 버전 관리에 커밋하면 팀원이 자동으로 해당 설정을 사용할 수 있습니다:
`.claude/settings.json`:
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
팀원이 프로젝트를 클론하면 이 플러그인들이 자동으로 사용 가능해집니다.
### 엔터프라이즈 Marketplace 제한
엄격한 관리가 필요한 엔터프라이즈 환경에서는 managed settings에서 허용되는 Marketplace를 제한할 수 있습니다:
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
빈 배열 `[]`로 설정하면 외부 플러그인을 완전히 차단할 수 있습니다.
## CLI 명령어 레퍼런스
| 명령어 | 설명 |
| -------------------------------------- | --------------------- |
| `/plugin` | 대화형 관리 인터페이스 열기 |
| `/plugin install @` | 플러그인 설치 |
| `/plugin uninstall ` | 플러그인 제거 |
| `/plugin enable ` | 플러그인 활성화 |
| `/plugin disable ` | 플러그인 비활성화 |
| `/plugin update ` | 플러그인 업데이트 |
| `/plugin validate .` | 현재 디렉토리의 플러그인 설정 검증 |
| `/plugin marketplace add ` | Marketplace 추가 |
| `/plugin marketplace list` | 추가된 Marketplace 목록 표시 |
| `/plugin marketplace update` | Marketplace 캐시 업데이트 |
| `/plugin marketplace remove ` | Marketplace 제거 |
## 모범 사례
### 개발 모범 사례
1. **Skills는 하나에 집중하기**: 각 Skill은 하나의 일을 잘 수행하도록 하고, 모든 것을 다 담으려 하지 마세요
2. **명확한 설명 작성하기**: Claude가 컴포넌트를 언제 사용해야 하는지 이해할 수 있도록 하세요
3. **팀 내에서 먼저 테스트하기**: 커뮤니티에 배포하기 전에 팀 내부에서 먼저 검증하세요
4. **버전 변경 사항 기록하기**: CHANGELOG.md에 각 버전의 변경 내용을 기록하세요
### 디렉토리 구조 모범 사례
* `commands/`, `agents/`, `skills/`는 플러그인 루트 디렉토리에 배치합니다
* `.claude-plugin/` 디렉토리에는 `plugin.json`만 배치합니다
* 플러그인 내 파일을 참조할 때는 `${CLAUDE_PLUGIN_ROOT}`를 사용합니다
* `../`를 사용하여 플러그인 외부의 파일에 접근하지 마세요
### Hooks 모범 사례
1. 스크립트에 실행 권한 부여: `chmod +x script.sh`
2. shebang으로 인터프리터 선언: `#!/bin/bash`
3. `${CLAUDE_PLUGIN_ROOT}` 변수를 사용하여 경로 정확성 확보
4. Hook에 통합하기 전에 스크립트를 단독으로 테스트
### 버전 관리 모범 사례
Semantic Versioning을 따릅니다:
* **MAJOR** (1.0.0 → 2.0.0): 호환성을 깨는 변경
* **MINOR** (1.0.0 → 1.1.0): 새 기능 추가 (하위 호환)
* **PATCH** (1.0.0 → 1.0.1): 버그 수정 (하위 호환)
## 자주 발생하는 문제 해결
| 문제 | 가능한 원인 | 해결 방법 |
| -------------- | ----------------- | ----------------------------------------------------- |
| 플러그인이 로드되지 않음 | plugin.json 형식 오류 | `claude plugin validate`로 검증 |
| 명령어가 표시되지 않음 | 디렉토리 구조 오류 | `commands/`가 루트 디렉토리에 있고 `.claude-plugin/` 내부가 아닌지 확인 |
| Hooks가 실행되지 않음 | 스크립트에 실행 권한 없음 | `chmod +x script.sh` 실행 |
| 경로를 찾을 수 없음 | 상대 경로 사용 | `${CLAUDE_PLUGIN_ROOT}`로 변경 |
| MCP 서버 실패 | 환경 변수 미설정 | `.mcp.json`의 경로 설정 확인 |
## 기존 설정에서 마이그레이션
이미 `.claude/` 디렉토리에 설정이 있는 경우, 다음 단계를 따라 Plugin으로 마이그레이션할 수 있습니다:
1. **Plugin 구조 생성**:
```bash
mkdir my-plugin/.claude-plugin
```
2. **plugin.json 생성**:
```json
{
"name": "my-plugin",
"description": "从现有配置迁移的插件",
"version": "1.0.0"
}
```
3. **기존 파일 복사**:
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **Hooks 마이그레이션**:
`.claude/settings.json`에서 `hooks` 설정을 `hooks/hooks.json`으로 복사합니다
5. **테스트**:
```bash
claude --plugin-dir ./my-plugin
```
## 학습 리소스
### 공식 문서
| 리소스 | 링크 | 설명 |
| ----------------- | ----------------------------------------------------------------------------------------- | ------------ |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | 플러그인 레퍼런스 문서 |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | 플러그인 생성 가이드 |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | 공식 모범 사례 |
| Agent Skills 표준 | [agentskills.io](https://agentskills.io) | 오픈 표준 사양 |
### 공식 저장소
| 리소스 | 링크 | 설명 |
| ---------------------------------- | -------------------------------------------------------------------------------------- | ----------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | 공식 Skills 저장소 |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | 공식 플러그인 카탈로그 |
| Docker MCP Toolkit | [공식 사이트](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200개 이상의 프리빌트 MCP |
### 커뮤니티 리소스
| 리소스 | 링크 | 설명 |
| ---------------------- | ---------------------------------------------------------------------------- | ------------------------ |
| claude-plugins.dev | [공식 사이트](https://claude-plugins.dev/) | 커뮤니티 레지스트리 및 CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243개 플러그인 컬렉션 |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 모범 사례 모음 |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99개 에이전트 + 15개 오케스트레이터 |
| jeremylongshore 튜토리얼 | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | 수백 개 플러그인 + Jupyter 튜토리얼 |
### 추천 읽을거리
| 글 | 출처 |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## 전망
Plugin 메커니즘으로 Claude Code의 확장 능력이 비약적으로 향상되었습니다. 커뮤니티의 발전과 함께 다음과 같은 변화를 기대할 수 있습니다:
* **더 풍부한 플러그인 생태계**: 다양한 개발 시나리오와 워크플로를 포괄
* **엔터프라이즈급 기능**: 더 완성도 높은 권한 관리 및 감사 기능
* **크로스 플랫폼 호환성**: Skills 오픈 표준이 이미 여러 벤더에 채택되고 있습니다
지금이야말로 시작하기에 최적의 시기입니다. 간단한 명령어부터 시작하여 점차 Skills와 Hooks를 추가하고, 궁극적으로 완전한 워크플로 솔루션을 구축해 보세요.
Plugin에 포함할 수 있는 Subagent 컴포넌트에 대해 자세히 알고 싶다면, 《[Claude Code Subagent란 무엇인가](/ko/docs/notes/claude-subagent/concept)》를 참조하세요.
# Claude Code Subagent의 개념 소개
## 소개
Claude Code를 사용하여 복잡한 작업을 처리할 때 다음과 같은 딜레마에 직면했을 수 있습니다. 주요 대화의 맥락이 점점 길어지고 AI가 이전의 중요한 정보를 "잊기" 시작하며 응답 품질이 점차 저하됩니다.
이 문제를 해결하기 위해 Subagent가 탄생했습니다.
Skills가 Claude의 "작업 매뉴얼"인 경우 Subagent는 귀하가 고용하는 "정규직 직원"입니다. 이들은 자체 독립 워크스테이션(컨텍스트)을 갖고 특정 유형의 작업에 집중하며 완료 후 결과를 보고합니다.
## 하위 에이전트 이해
당신이 회사의 CEO라고 상상해보십시오. 회사 규모가 작으면 모든 일을 직접 처리합니다. 그러나 비즈니스가 확장됨에 따라 재무 회계사, 채용 HR, 개발 엔지니어 등 정규 직원을 고용하기 시작합니다. 각 직원은 자신의 스테이션에서 근무하며 작업을 완료한 후 귀하에게 보고합니다.
Subagent는 Claude Code에서 바로 이 역할을 수행합니다.
기술적 관점에서 서브에이전트는 다음과 같은 특성을 지닌 전문 AI 보조자입니다.
| 특징 | 설명 |
| ------------- | ----------------------------- |
| **독립적인 컨텍스트** | 각 하위 에이전트는 자체 컨텍스트 창에서 실행됩니다. |
| **특화 역량** | 특정 작업 유형에 최적화 |
| **구성 가능한 도구** | 지정된 도구 세트에만 액세스할 수 있습니다 |
| **맞춤형 프롬프트** | 행동을 안내하는 특별한 시스템 프롬프트가 있습니다 |
## 왜 독립적인 컨텍스트가 필요한가요?
이는 Subagent의 핵심 설계 개념이며 깊은 이해가 필요합니다.
일반적인 대화에서는 모든 정보가 동일한 맥락에 쌓입니다. Claude가 코드 베이스를 검색하고, 파일을 분석하고, 수정하도록 하면 이 모든 중간 처리가 컨텍스트 공간을 차지합니다. 대화가 진행되고 상황이 더욱 혼잡해짐에 따라 Claude는 이전의 중요한 정보를 "잊기" 시작할 수 있습니다.
하위 에이전트는 다음을 변경합니다.
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
하위 에이전트의 분석 프로세스는 기본 대화를 오염시키지 않습니다. 주요 대화는 명확성과 초점을 유지하면서 정제된 결과만 얻었습니다.
## 내장 하위 에이전트 유형
Claude Code는 가장 일반적인 사용 시나리오를 다루는 세 가지 강력한 내장 하위 에이전트를 제공합니다.
### 하위 에이전트 탐색
**타겟팅**: 코드베이스를 빠르게 읽기 전용으로 탐색합니다.
**특징**:
* Haiku 모델 사용(빠르고 낮은 대기 시간)
* 엄격한 읽기 전용 - 파일을 생성, 수정 또는 삭제할 수 없습니다.
* 사용 가능한 도구: Glob, Grep, Read, Bash(읽기 전용 작업)
**사용 시기**:
"이 기능은 어디에 구현되어 있나요?"와 같은 탐구적인 질문을 하면 그리고 "오류는 어떻게 처리되나요?" Claude는 자동으로 Explore Subagent를 호출합니다.
**세부정보 수준**:
| 레벨 | 설명 | 적용 가능한 시나리오 |
| ------ | --------------- | ------------------- |
| 빠른 | 최소한의 탐색으로 빠른 검색 | 간단한 타겟 쿼리 |
| 중간 | 보통 탐색 | 속도와 완성도의 균형 |
| 매우 철저함 | 종합분석 | 심층적인 이해가 필요한 복잡한 문제 |
### 하위 에이전트 계획
**포지셔닝**: 코드 베이스를 연구하고 구현 계획을 준비합니다.
**특징**:
* Sonnet 모델 사용(강력한 추론 기능)
* 전용 탐색 도구: Read, Glob, Grep, Bash
* 계획 모드에서 자동으로 호출됩니다.
**사용 시기**:
계획 모드에 들어가서 Claude가 계획을 제안하기 전에 연구를 수행해야 하는 경우 Plan Subagent는 자동으로 정보를 수집한 다음 연구 결과에 따라 계획 제안을 제공합니다.
### 범용 하위 에이전트
**포지셔닝**: 복잡한 다단계 작업을 처리합니다.
**특징**:
* Sonnet 모델 사용
* 모든 도구에 대한 액세스(읽기 및 쓰기 포함)
* 탐구와 수정이 필요한 복잡한 작업에 적합
**사용 시기**:
작업에 여러 단계가 포함되어 수정하기 전에 검색이 필요하거나 초기 검색이 실패하고 여러 전략을 시도해야 하는 경우.
## 나의 이해와 실천
세 가지 공식 내장 하위 에이전트를 주의 깊게 관찰하면 한 가지 공통점을 발견할 수 있습니다. **모두 연구 및 계획 작업입니다**. Explore는 코드 베이스 탐색을 담당하고, Plan은 계획 수립을 담당하며, General-Purpose도 주로 연구 및 분석에 사용됩니다. 이들 중 어느 것도 코드 작성을 위해 특별히 설계된 것은 없습니다.
이는 Subagent에 대한 나의 이해를 확증해 줍니다. **Subagent의 핵심 가치는 "클린 컨텍스트"가 아니라 주 에이전트가 작업에 집중할 수 있도록 하는 것입니다**.
### 노동 형태의 분업
저의 사용방법은 매우 간단합니다. Subagent는 조사, 기획, 검토 등의 '정보수집' 업무를 담당하고, Main Agent는 실제 실행을 담당합니다.
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### Subagent에서 코드를 작성하도록 하면 어떨까요?
어떤 사람들은 주 에이전트가 여러 하위 에이전트를 예약하여 코드를 작성하는 것을 좋아합니다. 나는 이것이 신뢰할 수 없다고 생각합니다. 이유는 간단합니다. **컨텍스트가 심각하게 누락되었습니다**.
하위 에이전트의 컨텍스트는 독립적입니다. 주요 대화에서 어떤 내용이 논의되었는지, 어떤 결정이 내려졌는지, 어떤 제약이 부과되었는지 알 수 없습니다. 코드 작성을 요청하는 것은 신입 사원에게 배경 정보 없이 작업을 완료하도록 요청하는 것과 같습니다. 생성된 코드가 기대와 일치하지 않을 가능성이 높습니다.
반대로 Subagent를 "연구원"으로 지정하는 것이 훨씬 더 합리적입니다.
* 연구 과제 자체에는 많은 맥락이 필요하지 않습니다.
* 코드 대신 정보가 반환되며, 이는 전체 컨텍스트를 기반으로 주체가 사용할 수 있습니다.
* 설문조사 결과가 편향되더라도 주체가 이를 수정할 수 있습니다.
### 나의 일일 사용량
1. **새 작업을 시작하기 전**: Explore 에이전트가 관련 코드의 구조를 빠르게 이해할 수 있도록 하세요.
2. **복잡한 작업 계획**: 계획 에이전트가 요구 사항을 분석하고 구현 단계를 공식화하도록 합니다.
3. **코드 검토**: 검토 담당자가 코드 품질 및 보안 문제를 확인하도록 합니다.
4. **실제 코딩**: 주체는 수집된 컨텍스트를 기반으로 코드를 작성합니다.
현재 제가 사용하고 있는 에이전트 목록은 다음과 같습니다.
이것의 장점은 "하위 에이전트 검색 프로세스 중에 생성된 중간 결과 묶음" 대신 "내가 알아야 할 정보"만 포함하여 주 에이전트의 컨텍스트 창이 깨끗하게 유지된다는 것입니다.
## 다른 기능과의 비교
### Subagent vs Skills
이것이 가장 흔한 혼란입니다. 핵심 차이점: **기술은 클로드에게 지식을 주입합니다. 하위 에이전트는 독립 작업자**를 생성합니다.
| 치수 | 기술 | 하위 에이전트 |
| --------------- | ------------------------- | ------------------- |
| **핵심 기능** | 전문지식과 지침 제공 | 독립적으로 작업을 수행하는 에이전트 |
| **컨텍스트** | 주요 대화 내용 공유 | 독립적인 맥락을 갖는다 |
| **트리거 방법** | 설명에 따른 자동 매칭 | 자동 위임 또는 수동 호출 |
| **적용 가능한 시나리오** | 특정 유형의 작업에서 Claude의 능력 향상 | 복잡하고 다단계 독립 작업 |
비유적으로 말하면, 기술은 훈련 자료와 같아서 클로드가 어떤 일을 하는 방법을 배울 수 있게 해줍니다. Subagent는 자신의 작업장에서 독립적으로 작업을 완료하고 결과를 보고하는 정규 직원과 같습니다.
두 가지를 결합할 수 있습니다. 코드 검토 하위 에이전트는 코드 사양 스킬을 로드하여 "전문가 + 전문 지식"의 결합된 효과를 얻을 수 있습니다.
### 하위 에이전트와 슬래시 명령 비교
| 치수 | 하위 에이전트 | 슬래시 명령 |
| ---------- | --------------- | ---------- |
| **활성화 방법** | 자동 위임 또는 명시적 호출 | 사용자 수동 입력 |
| **컨텍스트** | 독립형 컨텍스트 | 공유된 주요 대화 |
| **복잡성** | 복잡한 작업에 적합 | 간단한 작업에 적합 |
슬래시 명령은 바로 가기 키이며 `/review`을 입력하면 미리 정의된 작업을 트리거할 수 있습니다. Subagent는 복잡한 다단계 작업을 자율적으로 완료할 수 있는 독립적인 작업자입니다.
### Subagent vs Plugin
플러그인은 하위 에이전트를 포함할 수 있는 "컨테이너" 개념입니다.
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
플러그인의 `agents/` 디렉터리에 하위 에이전트를 정의하고 플러그인과 함께 배포할 수 있습니다.
## 에이전트틱 디자인 패턴
Anthropic은 공식 문서에 6가지 핵심 Agentic 디자인 패턴을 요약합니다. 이러한 패턴을 이해하면 하위 에이전트 시스템을 더 효과적으로 설계하는 데 도움이 될 수 있습니다.
| 패턴 | 핵심 아이디어 | 하위 에이전트 신청 |
| --------------- | ----------------------- | ---------------------------- |
| **프롬프트 연결** | 복잡한 작업을 여러 순차적 단계로 분해 | 여러 하위 에이전트 연쇄 호출 |
| **라우팅** | 입력 유형에 따라 특수 프로세서에 분산 | 다양한 유형의 작업이 전문화된 하위 에이전트 |
| **병렬화** | 동시에 여러 개의 독립적인 하위 작업 실행 | 여러 하위 에이전트를 병렬로 시작 |
| **오케스트레이터-작업자** | 중앙 코디네이터가 작업자에게 작업을 할당 | 코디네이터인 Claude, 작업자인 Subagent |
| **평가자-최적화자** | 생성기 출력, 평가기 최적화 | 하위 에이전트 생성 + 하위 에이전트 검토 |
| **에이전트** | 자율적인 결정을 내리는 독립 에이전트 | 각 하위 에이전트는 독립적으로 실행됩니다 |
이러한 모드는 조합하여 사용할 수 있습니다. 예를 들어, 코드 품질 시스템은 다음 두 가지를 모두 사용할 수 있습니다.
* **병렬화**: 보안 검사와 성능 분석을 동시에 실행
* **오케스트레이터-작업자**: Master Claude는 여러 전문 하위 에이전트를 조정합니다.
* **Evaluator-Optimizer**: 생성 후 즉시 코드 검토
## 핵심 장점
### 컨텍스트 보호
Subagent의 가장 큰 가치는 주요 대화의 맥락을 보호하는 것입니다. 코드 검색, 파일 분석 등의 중간 프로세스가 메인 다이얼로그에 누적되지 않아 메인 다이얼로그가 항상 높은 수준의 목표에 집중할 수 있습니다.
### 전문화 기능
자세한 지침과 적절한 도구로 구성된 특정 도메인에 대한 특수 하위 에이전트를 생성할 수 있습니다. 특수 하위 에이전트는 범용 Claude보다 특정 작업에서 더 나은 성능을 발휘합니다.
### 유연한 권한 제어
각 하위 에이전트는 서로 다른 도구 액세스 권한을 가질 수 있습니다. 예를 들어 탐색 클래스 Subagent는 읽기 전용 권한만 부여하고 수정 클래스 Subagent는 쓰기 권한만 부여합니다. 이렇게 세밀하게 제어하면 보안이 향상됩니다.
### 재사용성
Subagent는 일단 생성되면 프로젝트 전체에서 재사용하거나 플러그인을 통해 팀과 공유할 수 있습니다.
## 서브에이전트를 사용해야 하는 경우
**하위 에이전트 사용에 적합한 시나리오**:
* 작업을 실행하려면 독립적인 컨텍스트가 필요합니다.
* 작업은 복잡한 다단계 작업 흐름입니다.
* 기본 대화와 다른 도구 세트가 필요합니다.
* 작업을 실행하는 데 시간이 오래 걸릴 수 있습니다.
**하위 에이전트 사용에 적합하지 않은 시나리오**:
* 간단한 일회성 쿼리
* 기본 대화와 긴밀한 상호 작용이 필요합니다.
* 미션을 빠르게 완료할 수 있습니다.
## 일반적인 애플리케이션 시나리오
### 코드 검토
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
코드 조각을 완성하면 Code Review Subagent가 주요 개발 작업을 방해하지 않고 독립적인 컨텍스트에서 코드를 검토하도록 할 수 있습니다.
### 디버깅 분석
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
오류가 발생하면 디버깅 하위 에이전트는 오류의 원인을 심층적으로 분석하고 다양한 가설을 시도한 후 최종적으로 복구 제안을 제공할 수 있습니다.
### 코드베이스 탐색
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
새 프로젝트를 처음 접하는 경우 Explore Subagent를 사용하면 수많은 검색 결과로 인해 주요 대화가 복잡해지지 않고 코드 베이스를 신속하게 매핑할 수 있습니다.
## 학습 리소스
### 공식 리소스
| 자원 | 링크 | 지침 |
| -------------- | --------------------------------------------------------------------------------------------------------- | ------------------------ |
| 클로드 코드 문서 | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 공식 문서 항목 |
| 하위 에이전트 가이드 | [클로드 코드 문서](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | 하위 에이전트 공식 문서 |
| 다중 에이전트 시스템 연구 | [인류공학](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2% 성능향상 연구내용 |
| 에이전트 디자인 패턴 | [인류학 문서](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 6가지 핵심 디자인 패턴에 대한 자세한 설명 |
### 커뮤니티 리소스
| 자원 | 링크 | 지침 |
| ------------- | ----------------------------------------------------------------- | ---------------------- |
| wshobson/에이전트 | [GitHub](https://github.com/wshobson/agents) | 에이전트 99명 + 오케스트레이터 15명 |
| 복합공학 | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17개 전문 에이전트용 플러그인 |
| 멋진 클로드 코드 | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | 모범 사례 요약 |
## 요약
Claude Code Subagent는 기본적으로 상황에 독립적인 전문 AI 도우미입니다. 컨텍스트 격리를 통해 복잡한 작업에서 정보 과부하 문제를 해결하고, 주요 대화를 항상 명확하고 집중적으로 유지합니다.
세 가지 핵심 단어를 기억하세요.
| 키워드 | 의미 |
| ------ | ------------------------------------------- |
| **독립** | 각 하위 에이전트에는 자체 컨텍스트 창이 있습니다. |
| **전문** | 특정 작업 유형에 최적화 |
| **위임** | Claude는 작업을 Subagent에 자동 또는 수동으로 위임할 수 있습니다 |
개념을 이해한 후 다음 글 "[Claude Code Subagent 실용 가이드](/ko/docs/notes/claude-subagent/practice)"에서는 사용자 정의 서브에이전트 생성, 도구 권한 구성, 실제 프로젝트에서의 모범 사례 등을 실습해 보겠습니다.
Subagent가 불러올 수 있는 스킬에 대해 알고 싶으시면 "[클로드 스킬이란?](/ko/docs/notes/claude-skills/concept)"을 읽어보세요. 배포용 Subagent를 패키징하려면 "[Claude Code 플러그인이란 무엇입니까](/ko/docs/notes/claude-plugin/concept)"를 읽어보시기 바랍니다.
# Claude Code Subagent 실용 가이드
## 빠른 검토
이전 글에서 Subagent의 핵심 개념에 대해 알아보았습니다. Subagent는 컨텍스트 격리를 통해 복잡한 작업에서 정보 과부하 문제를 해결하는 컨텍스트 독립적 전문 AI 도우미입니다. Claude Code에는 탐색, 계획 및 범용이라는 세 가지 기본 하위 에이전트가 있습니다. 이 문서에서는 실용적인 관점에서 사용자 정의 하위 에이전트를 만들고 고급 사용법을 익힐 수 있도록 안내합니다.
## 하위 에이전트 관리
### /agents 명령을 통해
가장 쉬운 방법은 대화형 인터페이스를 사용하는 것입니다.
```bash
/agents
```
그러면 다음을 수행할 수 있는 메뉴가 열립니다.
* 모든 하위 에이전트 보기(내장 + 사용자 정의)
-새 하위 에이전트 만들기
* 기존 하위 에이전트에 대한 구성 및 도구 권한 편집
* 불필요한 Subagent 삭제
* 이름 충돌이 있을 때 어떤 하위 에이전트가 활성화되어 있는지 확인
### 파일관리를 통해
하위 에이전트는 Markdown 파일로 저장됩니다. 파일을 직접 생성하고 편집할 수도 있습니다.
**저장 위치**:
| 위치 | 경로 | 범위 |
| ------- | -------------------- | ---------------- |
| 프로젝트 수준 | `.claude/agents/` | 현재 프로젝트 전용이며 Git |
| 사용자 수준 | `~/.claude/agents/` | 모든 프로젝트에서 사용 가능 |
| 플러그인 | 플러그인의 `agents/` 디렉토리 | 플러그인과 함께 설치됨 |
**우선순위**: 프로젝트 수준 > 사용자 수준 > 플러그인 수준
동일한 이름을 가진 하위 에이전트가 여러 위치에 존재하는 경우 우선 순위가 높은 하위 에이전트가 우선 순위가 낮은 하위 에이전트를 덮어씁니다.
## 첫 번째 하위 에이전트 만들기
### 1단계: 디렉터리 만들기
```bash
mkdir -p .claude/agents
```
### 2단계: 마크다운 파일 만들기
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### 3단계: 하위 에이전트 테스트
클로드 코드에서:
```
> 用 code-reviewer 代理审查我最近的修改
```
아니면 Claude가 자동으로 선택하도록 하세요.
```
> 帮我审查一下代码质量
```
`description`이 충분히 명확하게 작성되면 Claude가 자동으로 하위 에이전트를 인식하고 호출합니다.
## 구성 필드에 대한 자세한 설명
하위 에이전트 구성 파일은 YAML 머리말과 Markdown 본문의 두 부분으로 구성됩니다.
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| 필드 | 필수 | 설명 |
| ---------------- | --- | --------------------------------------------- |
| `name` | 이다 | 고유 식별자, 소문자 및 하이픈 사용 |
| `description` | 예 | 자연어 설명(Claude는 이를 사용하여 언제 전화할지 결정합니다) |
| `tools` | 아니요 | 쉼표로 구분된 도구 목록입니다. 생략하면 모든 도구가 상속됩니다 |
| `model` | 아니요 | 모델 선택: `sonnet`, `opus`, `haiku` 또는 `inherit` |
| `permissionMode` | 아니요 | 권한 모드(아래 참조) |
| `skills` | 아니요 | 자동 로드된 기술(하위 에이전트는 상위 세션의 기술을 상속하지 않음) |
### 권한 모드
| 모드 | 설명 |
| ------------------- | -------------------- |
| `default` | 일반 권한 확인 |
| `acceptEdits` | 편집 작업을 자동으로 수락 |
| `bypassPermissions` | 모든 권한 확인 건너뛰기 |
| `plan` | 계획을 제안만 하고 실행은 하지 않음 |
| `ignore` | 이 하위 에이전트 무시 |
### 마크다운 텍스트
텍스트는 Subagent의 시스템 프롬프트입니다. 더 자세히 작성할수록 Subagent의 성능이 향상됩니다.
좋은 시스템 프롬프트에는 다음이 포함되어야 합니다.
* 명확한 역할 정의
* 특정 작업 단계
* 주요 체크리스트
* 출력 형식 요구 사항
## 트리거 메커니즘
### 자동 위임
Claude는 작업 내용과 하위 에이전트의 `description`을 기반으로 위임 여부를 자동으로 결정합니다.
**자동 사용을 장려하는 팁**: `description`에서 유발 단어를 사용하세요.
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
또는:
```yaml
description: MUST BE USED when encountering errors or test failures
```
### 명시적 호출
사용할 하위 에이전트를 Claude에게 직접 알려주십시오.
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
## 도구 구성
### 일반적으로 사용되는 도구 목록
| 도구 | 지침 |
| ----------- | -------- |
| `Read` | 파일 내용 읽기 |
| `Write` | 파일에 쓰기 |
| `Edit` | 파일 편집 |
| `Glob` | 파일 패턴 일치 |
| `Grep` | 정규식 검색 |
| `Bash` | 쉘 명령 실행 |
| `WebFetch` | 웹 콘텐츠 받기 |
| `WebSearch` | 웹 검색 |
### 도구 구성 전략
**읽기 전용 하위 에이전트**(탐색, 분석):
```yaml
tools: Read, Grep, Glob, Bash
```
참고: Bash가 포함되어 있더라도 Subagent는 읽기 전용 명령(ls, git status, git log 등)에만 사용해야 합니다.
**하위 에이전트 읽기 및 쓰기**(수리, 리팩터링):
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**최소 권한의 원칙**: 실수로 인한 작업을 방지하기 위해 필요한 도구만 부여합니다.
## 실용적인 하위 에이전트 템플릿
### 코드 검토자
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### 디버깅 전문가
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### 테스트 러너
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### 문서 생성기
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### 보안 스캐너
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## 고급 사용법
### 생산 수준 디자인 패턴
프로덕션 환경에는 다중 에이전트 협업을 위한 몇 가지 입증된 패턴이 있습니다.
#### 3 아미고스 모드
제품, 아키텍처, 구현의 세 가지 역할로 구성된 협업 모델:
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| 역할 | 책임 | 도구 구성 |
| -------- | --------------- | -------------- |
| PM 에이전트 | 기능 정의, 요구 사항 정렬 | 읽기, 웹 검색 |
| 건축가 에이전트 | 기술 솔루션 설계 | 읽기, Glob, Grep |
| 클로드 코드 | 코드 구현 | 모든 도구 |
#### 3단계 파이프라인
복잡한 작업을 세 가지 명확한 단계로 나누세요.
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
각 단계는 전용 하위 에이전트를 담당하며 출력은 다음 단계의 입력 역할을 합니다.
#### 모델 조정 전략
비용과 효과를 최적화하기 위해 다양한 모델이 다양한 단계에서 사용됩니다.
| 무대 | 추천 모델 | 이유 |
| ----- | ----- | ------------ |
| 기획단계 | 소네트 | 깊은 추론이 필요합니다 |
| 실행 단계 | 하이쿠 | 빠르고 저렴한 비용 |
| 검토 단계 | 소네트 | 종합적인 판단이 필요함 |
구성 예:
```yaml
---
name: quick-executor
model: haiku
---
```
### 하위 에이전트 링크
복잡한 작업 흐름의 경우 여러 하위 에이전트를 연결할 수 있습니다.
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### 재개 가능한 실행
이전 컨텍스트 전체를 유지하면서 하위 에이전트 실행을 일시 중지하고 재개할 수 있습니다.
**최초 통화**:
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**복구 에이전트**:
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**사용 시나리오**:
* 여러 세션을 거쳐 완성된 장기 연구
* 반복적인 개선, 컨텍스트 유지
* 관련 작업을 순차적으로 처리하는 다단계 워크플로우
### 하위 에이전트에 대한 기술 구성
하위 에이전트는 상위 세션의 기술을 자동으로 상속하지 않습니다. 필요한 경우 명시적으로 선언합니다.
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### CLI 동적 정의
파일을 저장할 필요가 없으며 명령줄에서 직접 임시 하위 에이전트를 정의합니다.
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
빠른 테스트 또는 일회용 사용에 적합합니다.
## 모범 사례
### 1. 집중하세요
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
한 가지 일을 잘 수행하는 하위 에이전트가 많은 일을 수행하는 하위 에이전트보다 낫습니다.
### 2. 명확한 설명을 작성하세요.
Claude는 `description`을 사용하여 Subagent를 사용할 시기를 결정합니다. 좋은 설명은 다음과 같이 답해야 합니다.
1. \*\*이 하위 에이전트의 기능은 무엇입니까? \*\* 특정 능력을 나열하십시오.
2. \*\*언제 사용해야 하나요? \*\* 유발 단어가 포함되어 있습니다.
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. 도구 액세스 제한
필요한 도구만 부여하십시오.
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
이렇게 하면 Subagent가 실수로 파일을 수정하는 것을 방지하고 검토 작업에 더 집중할 수 있습니다.
### 4. 자세한 시스템 프롬프트 작성
시스템 프롬프트가 더 자세할수록 Subagent가 더 나은 성능을 발휘합니다.
* 명확한 역할 정의
* 특정 작업 단계
* 주요 체크리스트
* 출력 형식 요구 사항
### 5. 버전 관리
Git에 프로젝트 수준 하위 에이전트를 커밋합니다.
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
팀 구성원은 프로젝트를 복제한 후 자동으로 동일한 하위 에이전트를 얻습니다.
## 일반적인 문제 해결
| 문제 | 가능한 원인 | 솔루션 |
| ---------------- | ------------------------- | ---------------------------------------------------- |
| 하위 에이전트가 호출되지 않음 | 설명이 충분히 명확하지 않습니다 | 더욱 구체적으로 만들기 위해 유발 단어를 추가하세요 |
| 하위 에이전트가 호출되지 않음 | 잘못된 파일 위치 | 파일이 `.claude/agents/` 또는 `~/.claude/agents/`에 있는지 확인 |
| 도구를 사용할 수 없습니다 | 도구 필드 구성 오류 | 도구 이름의 철자를 확인하고 쉼표로 구분되어 있는지 확인하세요. |
| 출력이 불안정합니다 | 시스템 프롬프트가 너무 모호함 | 특정 단계 및 출력 형식 요구 사항 추가 |
| 컨텍스트 손실 | 세션이 종료되었습니다 | 재개 가능한 실행 사용 |
| 이름 충돌 | 여러 위치에 동일한 이름을 가진 하위 에이전트 | `/agents`을 사용하여 어느 것이 활성화되어 있는지 확인하세요 |
## 팀과 공유
### 방법 1: Git을 통해
Subagent를 `.claude/agents/` 디렉터리에 배치하고 프로젝트 저장소에 제출합니다. 복제 후 팀원은 자동으로 획득됩니다.
### 방법 2: 플러그인을 통해
플러그인의 `agents/` 디렉터리에 Subagent를 배치하고 플러그인 메커니즘을 통해 배포합니다.
### 방법 3: 사용자 수준 공유
일반적으로 사용되는 하위 에이전트를 `~/.claude/agents/`에 배치하여 모든 프로젝트에서 사용할 수 있도록 합니다. 여러 시스템 간의 동기화는 도트파일을 사용하여 관리할 수 있습니다.
## 학습 리소스
### 공식 문서
| 자원 | 링크 | 지침 |
| ----------- | --------------------------------------------------------------------------------------------------------- | ----------------- |
| 클로드 코드 문서 | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | 공식 문서 항목 |
| 하위 에이전트 가이드 | [클로드 코드 문서](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | 하위 에이전트 구성 세부사항 |
| 에이전트 디자인 패턴 | [인류학 문서](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | 6가지 핵심 디자인 패턴 |
| 다중 에이전트 연구 | [인류공학](https://www.anthropic.com/engineering/multi-agent-research-system) | 90.2% 성능 향상 연구 내용 |
### 커뮤니티 리소스
| 자원 | 링크 | 지침 |
| ------------- | ----------------------------------------------------------------- | --------------------------------- |
| wshobson/에이전트 | [GitHub](https://github.com/wshobson/agents) | 99명의 에이전트 + 15개의 Orchestrator 템플릿 |
| 복합공학 | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17개 전문 에이전트용 플러그인 |
| 멋진 클로드 코드 | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Claude Code 모범 사례 요약 |
### 추천 도서
| 기사 | 소스 | 주제 |
| ----------------------------- | ------------------------------------------------------------------------------------- | --------------- |
| 효과적인 에이전트 구축 | [인류](https://www.anthropic.com/research/building-effective-agents) | 에이전트 설계 원칙 |
| 다중 에이전트 연구 시스템을 구축한 방법 | [인류](https://www.anthropic.com/engineering/multi-agent-research-system) | 다중 에이전트 아키텍처 실습 |
| Claude 기술, 명령, 하위 에이전트 및 플러그인 | [젊은 리더스 테크](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | 기능 비교 분석 |
## 요약
Claude Code Subagent는 AI 프로그래밍 효율성을 향상시키는 강력한 도구입니다. 독립적인 컨텍스트와 전문화된 구성을 통해 복잡한 작업을 관리 가능하게 만듭니다.
빠른 시작:
1. `/agents`을 실행하여 관리 인터페이스를 엽니다.
2. 간단한 하위 에이전트(예: 코드 검토자) 만들기
3. 자동 위임 및 명시적 호출 테스트
4. 필요에 따라 구성을 조정합니다.
사용이 심화됨에 따라 점차적으로 다음을 수행할 수 있습니다.
* 팀을 위한 전용 하위 에이전트 만들기
* 복잡한 작업 흐름을 처리하도록 하위 에이전트 링크 구성
* 재개 가능한 실행으로 장기 작업 처리
Subagent를 다른 구성으로 패키징하여 배포하고자 하는 경우에는 "[Claude Code Plugin Practical Guide](/ko/docs/notes/claude-plugin/practice)"를 읽어보시기 바랍니다.
# 고급편
## 터미널 알림: 작업 완료 알림
Claude가 작업을 완료했을 때 알림을 받고 싶으신가요?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
iTerm2의 알림 기능과 함께 사용하거나, `terminal-notifier`를 활용하여 맞춤 알림을 설정할 수 있습니다([모범 사례](/ko/blog/claude-code-best-practices)의 Hooks 설정을 참조하세요).
## Hooks 심화 활용
Hooks는 단순히 셸 명령을 실행하는 것만이 아닙니다. 실제로 네 가지 유형이 있습니다:
1. **command**:Shell 命令(最常见)
2. **http**:POST JSON 到 URL(支持自定义 headers 和环境变量展开)
3. **prompt**:发给 Claude 评估(比如「所有任务都完成了吗?」)
4. **agent**:启动一个有工具访问权限的子代理来验证
알아두면 유용한 고급 Hook 이벤트:
* `PostCompact`:压缩完成后触发,适合注入提醒让 Claude 重新读取关键文件
* `SessionStart`:写入 `$CLAUDE_ENV_FILE` 可以给整个会话持久化环境变量
* `PreToolUse`:可以修改工具输入(`updatedInput`),甚至自动批准或拒绝操作
## 플러그인 생태계
`/plugin`을 사용하여 커뮤니티 플러그인을 탐색하고 설치할 수 있습니다. 주목할 만한 플러그인:
* **dx**(by ykdojo):提供 `/handoff`(自动写交接文档)、`/clone`(克隆对话)、`/half-clone`(只克隆最近的对话减少上下文)
* **mine**(by anipotts):把所有 Claude Code 会话数据导入 SQLite,支持成本追踪、缓存分析、错误记忆等查询
## Agent Teams: 멀티 에이전트 협업
환경 변수를 설정하여 실험적 Agent Teams 기능을 활성화합니다:
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
활성화하면 세션이 Team Lead 역할을 하며, git worktree를 통해 여러 에이전트가 동시에 작업하도록 조율할 수 있습니다. 각 에이전트는 자체 컨텍스트 윈도우에서 독립적으로 실행되므로, 대규모 프로젝트의 병렬 개발에 적합합니다.
다만 토큰 소비가 4\~15배 증가하므로, 상황에 맞게 사용하시기 바랍니다.
## 프롬프트 전략
다음 팁들은 Boris Cherny가 Twitter에서 공유한 팀 실천 방법에서 가져온 것입니다. Claude Code에 적용된 "프롬프트 엔지니어링"의 모범 사례라고 할 수 있습니다.
### Claude를 코드 리뷰어로 활용하기
Claude에게 코드를 작성하게만 하지 말고, 여러분의 코드를 리뷰하게 하세요:
```
Grill me on these changes and don't make a PR until I pass your test.
```
또는 코드가 작동하는지 증명하게 하세요:
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### 불만족스러운 답변을 다시 묻지 않기
Boris의 팁 #6: Claude가 평범한 답변을 주었다면, 다른 표현으로 다시 질문하지 마세요. 대신 "이 솔루션은 충분하지 않다, 구체적으로 어디를 개선할 수 있는지 알려줘"라고 말하세요. 기존 답변을 개선하는 것이 처음부터 다시 시작하는 것보다 효과적입니다.
### Claude 스스로 CLAUDE.md를 업데이트하게 하기
실수를 수정한 후, 다음 한마디를 추가하세요:
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris에 따르면, Claude는 자기 자신을 위한 규칙을 작성하는 데 놀라울 정도로 뛰어나다고 합니다. 시간이 지남에 따라 CLAUDE.md는 점점 더 정확해지고, 대화 품질도 지속적으로 향상됩니다.
### "fix"라고만 말하기
Slack MCP를 활성화한 상태에서 Slack의 버그 리포트를 붙여넣고 한 단어만 말하세요: **fix**. 컨텍스트 스위칭 제로입니다.
또는 CI가 실패했을 때, 간단히:
```
Go fix the failing CI tests.
```
수동으로 로그를 분석하거나 문제를 설명할 필요가 없습니다. Claude가 직접 로그를 확인하고, 문제를 진단하고, 수정하게 하세요.
## 마치며
Claude Code는 매우 빠르게 발전하고 있으며, 이러한 기법들도 끊임없이 개선되고 있습니다. 공식 Changelog를 팔로우하여 최신 정보를 확인하시기 바랍니다.
이전 글을 아직 읽지 않으셨다면, 기본 워크플로우부터 시작하시는 것을 추천합니다:
### 추가 읽기
* [나의 Claude Code 모범 사례](/ko/blog/claude-code-best-practices) — 워크플로우 핵심 기법과 슬래시 명령어 가이드
* [AI 프로그래밍 품질 관리: 코드 품질을 보장하는 5가지 방어선](/ko/blog/claude-code-quality-control) — Claude Code 프로그래밍의 품질 보증 체계
* [Claude 시스템 아키텍처 완전 해설](/ko/docs/notes/claude-architecture) — MCP, Skills, Subagents, Hooks 등 컴포넌트 이해하기
# 실용 명령어와 자동화
## `/diff`: 인터랙티브 Diff 뷰어
`/diff`를 입력하면 인터랙티브한 diff 뷰가 열립니다:
* **좌우 화살표 키**: git diff(전체 변경 사항)와 Claude의 각 턴별 변경 사항 간 전환
* **상하 화살표 키**: 다른 파일 탐색
터미널에서 `git diff`를 실행하는 것보다 훨씬 편리하며, 특히 여러 파일에 걸친 변경 사항이 있을 때 유용합니다.
## `/simplify`: 멀티 에이전트 코드 리뷰
`/simplify`를 실행하면 3개의 병렬 리뷰 에이전트가 동시에 시작됩니다:
* **코드 재사용** 에이전트: 중복 패턴을 찾습니다
* **코드 품질** 에이전트: 가독성과 구조를 검사합니다
* **효율성** 에이전트: 불필요한 성능 오버헤드를 분석합니다
세 에이전트가 독립적으로 작업한 후 결과를 취합하여, 유효한 문제는 자동으로 수정하고 오탐은 건너뜁니다.
## `/security-review`: 보안 스캔
현재 브랜치의 변경 사항에 대해 보안 감사를 수행하여 SQL 인젝션, XSS, 인증 결함, 데이터 처리 문제 및 의존성 취약점을 검사합니다. 각 발견 사항은 적대적 검증을 거쳐 오탐을 줄입니다.
## `/copy`의 숨겨진 기능
`/copy`는 단순히 마지막 응답을 복사하는 것이 아닙니다. 응답에 코드 블록이 포함되어 있으면, 전체 응답을 복사하는 대신 특정 코드 블록을 선택할 수 있는 인터랙티브 선택기가 나타납니다. 숫자를 전달하여 이전 응답을 복사할 수도 있습니다: `/copy 2`는 마지막에서 두 번째, `/copy 3`은 마지막에서 세 번째 응답을 복사하므로 스크롤하여 수동으로 선택할 필요가 없습니다.
## `/batch`: 대규모 병렬 리팩토링
```
/batch 把 src/ 下所有组件从 Class 组件迁移到函数组件
```
이것은 강력한 기능입니다. `/batch`는 코드베이스를 분석하고, 작업을 5\~30개의 독립적인 단위로 분해한 후, 각 단위마다 격리된 git worktree에서 작업하는 독립 에이전트를 시작하고, 마지막으로 각 에이전트가 커밋하고 PR을 생성합니다.
대규모 마이그레이션, 일괄 타입 어노테이션 추가, 전역 리네이밍 등의 시나리오에 적합합니다.
## `/loop`: 예약 작업
```
/loop 5m 检查部署是否完成
/loop 1h /review-pr 1234
```
세션 내에서 지정된 간격으로 반복 실행되는 예약 작업을 생성합니다. 배포 상태 폴링, PR 정기 확인 등에 유용합니다. 세션 수준(종료하면 사라짐)이며, 최대 50개 작업, 3일 후 자동 만료됩니다.
## 파이프 입력: 무엇이든 Claude에게 전달하기
```bash
# 让 Claude 分析错误日志
cat error.log | claude -p "分析这个错误日志,找出根本原因"
# 让 Claude 总结最近的改动
git diff HEAD~3 | claude -p "总结这三次提交的改动"
# 让 Claude 解读命令输出
kubectl get pods | claude -p "哪些 pod 状态异常?"
```
`-p`는 Headless 모드(비대화형)로, 스크립트 및 CI/CD 파이프라인에서 사용하기에 적합합니다.
## Headless 모드의 숨겨진 매개변수
`-p` 모드에는 매우 강력하지만 잘 알려지지 않은 매개변수들이 있습니다:
```bash
# 设置花费上限(超过就停)
claude -p --max-budget-usd 5.00 "重构认证模块"
# 限制对话轮数
claude -p --max-turns 3 "修复这个测试"
# 输出 JSON 格式(方便程序解析)
claude -p --output-format json "分析这个项目"
# 要求输出符合特定 JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "总结项目"
# 多轮 headless 对话(用 session-id 保持上下文)
claude -p --session-id my-task "第一步:分析代码"
claude -p --session-id my-task "第二步:生成测试"
# 指定备用模型(主模型过载时自动切换)
claude -p --fallback-model sonnet "复杂分析"
# 限制可用工具
claude -p --tools "Read,Grep,Glob" "只读分析,不要改代码"
# 完全替换系统提示词
claude -p --system-prompt "你是一个 Python 专家" "优化这段代码"
```
# 설정과 진단
## `/statusline`: 커스텀 상태 표시줄
`/statusline`을 사용하면 자연어 설명으로 하단 상태 표시줄에 표시되는 정보를 커스터마이즈할 수 있습니다. 또는 `~/.claude/statusline.sh` 스크립트를 직접 생성할 수도 있습니다.
표시할 수 있는 정보에는 현재 모델, git 브랜치, 커밋되지 않은 파일 수, 컨텍스트 사용 진행률, 세션 비용 등이 있습니다. 여러 Claude 창을 열어 서로 다른 작업을 처리할 때, 상태 표시줄을 통해 각 창이 무엇을 하고 있는지 빠르게 구분할 수 있습니다.
## settings.json 자동 완성
settings.json 시작 부분에 `$schema`를 추가하면 VS Code / Cursor가 설정 항목의 자동 완성과 유효성 검사를 제공합니다:
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## 유용한 숨겨진 설정
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration`: 각 대화 턴의 소요 시간 표시
* `DISABLE_AUTOUPDATER`: 자동 업데이트 확인 비활성화로 컨텍스트 오버헤드 감소
## `/stats`와 `/insights`: 사용 분석
* `/stats`: 일별 사용량, 세션 기록, 사용 스트릭, 모델 선호도를 시각화하며 날짜 범위 필터링 지원
* `/insights`: 모든 Claude Code 기록을 분석하여 어떤 워크플로가 효과적인지, 어디서 병목이 발생하는지 알려주고 최적화 제안도 생성
## history.jsonl: 프롬프트 기록
Claude는 보낸 모든 프롬프트를 `~/.claude/history.jsonl`에 저장합니다. Claude에게 이 파일을 분석하도록 요청하면 프롬프트 패턴과 최적화 기회를 찾을 수 있습니다.
## `/doctor`: 상태 점검
이상한 문제가 발생하면 `/doctor`(또는 터미널에서 `claude doctor`)를 실행하세요. 설치 상태, 버전, 인증 상태, 시스템 종속성을 점검하여 문제를 빠르게 파악할 수 있도록 도와줍니다.
## 커뮤니티 도구
커뮤니티 도구 `ccusage`로 토큰 사용량을 추적할 수 있습니다:
```bash
npx ccusage daily
```
`--dangerously-skip-permissions`를 사용했거나 많은 명령을 승인한 경우, `cc-safe`로 위험을 스캔할 수 있습니다:
```bash
npx cc-safe .
```
`.claude/settings.json`에서 `sudo`, `rm -rf`, `chmod 777`, `git reset --hard` 등의 고위험 명령을 검사합니다.
## 알아두면 좋은 슬래시 명령
| 명령 | 기능 |
| --------------------- | ------------------------------------------------------- |
| `/export [filename]` | 대화를 일반 텍스트로 내보내기 |
| `/pr-comments [PR]` | PR 댓글 가져오기 (현재 브랜치 자동 감지) |
| `/release-notes` | 현재 버전 변경 로그 보기 |
| `/plugin` | 커뮤니티 플러그인 탐색 및 설치 |
| `/fast` | 빠른 모드 전환 |
| `claude --debug` | 시작 시 디버그 로그 활성화 (카테고리 필터링 지원, 예: `--debug "api,hooks"`) |
| `/install-github-app` | GitHub App 설치로 자동 PR 리뷰 구현 |
# 컨텍스트 관리
## `/compact`는 인수를 받을 수 있다
많은 사람이 `/compact`로 컨텍스트를 압축할 수 있다는 것은 알지만, 인수를 지정해서 무엇을 보존할지 정할 수 있다는 것은 잘 모릅니다:
```
/compact 保留所有关于数据库 schema 的讨论,以及当前的重构方案
```
이렇게 하면 압축 시 지정한 내용이 우선적으로 보존되어 핵심 컨텍스트가 손실되는 것을 방지할 수 있습니다.
## CLAUDE.md에 압축 생존 지침 작성하기
CLAUDE.md에 `## Compact Instructions` 섹션을 추가하여 압축 시 반드시 보존해야 할 내용을 Claude에게 알려주세요:
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
이렇게 하면 자동 압축이 되더라도 중요한 정보가 손실되지 않습니다.
## 토큰 예산 때문에 Claude가 조기 중단하는 것 방지하기
CLAUDE.md에 다음을 추가하세요:
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
가끔 Claude는 컨텍스트가 거의 가득 차면 "컨텍스트가 거의 가득 찼습니다"라고 말하며 스스로 멈추는 경우가 있습니다. 이 내용을 추가하면 조기 중단을 방지할 수 있습니다.
## Handoff 프로토콜: 세션 인수인계
컨텍스트가 거의 가득 찼지만 작업이 아직 완료되지 않았을 때, Claude에게 인수인계 문서를 작성하게 하세요:
```
把剩余的计划写到 HANDOFF.md 里,说明你尝试了什么、什么有效、什么没效。
```
그런 다음 새 세션을 열고 `@HANDOFF.md`만 입력하면 전체 컨텍스트를 복원할 수 있습니다. 10K+ 토큰의 컨텍스트를 2K 미만으로 압축하며, `/compact`보다 훨씬 정확합니다.
## 70-80%에서 선제적으로 압축하기
놓치기 쉬운 포인트: 컨텍스트가 한계에 가까워지면 Claude가 자동으로 압축을 트리거합니다. 하지만 작업 도중에 자동 압축이 발생하면 핵심 정보가 손실되어 이후 응답 품질이 저하될 수 있습니다.
더 좋은 방법은 **선제적 관리**입니다: 컨텍스트가 70-80%에 도달하면 수동으로 `/compact`를 실행하세요. 자동 압축을 기다리는 것보다 훨씬 효과적입니다. 작업을 완료한 후에는 즉시 `/clear`를 실행하여 컨텍스트가 무한히 커지지 않도록 하세요.
환경 변수를 통해 자동 압축을 더 일찍 트리거할 수도 있습니다:
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context`: 컨텍스트 진단
컨텍스트 창에 얼마나 공간이 남았는지 확실하지 않다면 `/context`가 알려줍니다:
* 어떤 도구나 MCP 서비스가 가장 많은 컨텍스트를 소비하는지
* 현재 용량 사용 비율
* 맞춤형 최적화 제안
일부 MCP 서비스는 등록만 해두고 사용하지 않아도 컨텍스트 창의 30% 이상을 차지할 수 있다는 것을 발견했습니다. `/context`로 확인하고 사용하지 않는 MCP를 정리하면 상당한 공간을 확보할 수 있습니다.
## MCP 도구 자동 지연 로딩
MCP 도구 정의가 컨텍스트의 10%를 초과하면, Claude Code는 자동으로 Tool Search를 활성화합니다. 전체 도구 정의 대신 경량 검색 인덱스를 로드합니다. 이를 통해 MCP 컨텍스트 소비를 85% 이상 줄일 수 있습니다(예: 77K 토큰에서 8.7K으로). 이 기능은 **기본적으로 활성화**되어 있으며 수동 설정이 필요하지 않습니다.
주의: Tool Search는 Sonnet 4+ 및 Opus 4+ 모델만 지원하며, Haiku는 지원하지 않습니다. `ANTHROPIC_BASE_URL`이 비공식 프록시를 가리키는 경우, Tool Search는 자동으로 비활성화됩니다(대부분의 프록시가 `tool_reference` 블록을 전달하지 않기 때문).
동작을 사용자 정의하려면 settings.json에서 설정할 수 있습니다:
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
지원되는 설정 값:
* **미설정**: 기본적으로 활성화
* **`true`**: 강제 활성화(비공식 프록시 시나리오 포함)
* **`auto`**: 컨텍스트가 10%를 초과할 때 활성화(기본 동작과 동일)
* **`auto:`**: 사용자 정의 임계값, 예: `auto:5`는 5% 초과 시 활성화
* **`false`**: 비활성화, 모든 MCP 도구가 사전 로드됨
# Claude Code 숨겨진 꿀팁 모음
창립자 Boris Cherny의 트윗, 커뮤니티, changelog에서 정리한 Claude Code 실용 팁입니다. 대부분 '이런 기능이 있었어?' 싶은 숨겨진 기능들입니다 -- 단축키, 숨은 기능, 커맨드라인 트릭 등 한번 쓰기 시작하면 다시는 돌아갈 수 없습니다.
## 목차
* [단축키 편](./shortcuts) -- Shift+Tab 모드 전환, Esc+Esc 이력 호출, Ctrl+S 임시 저장 등
* [입력과 인터랙션 편](./input-interaction) -- `!` 터미널 명령어, `@` 파일 주입, URL 붙여넣기, /btw 끼어들기, Vim 모드
* [사고 및 모델 제어 편](./thinking-model) -- think/ultrathink 키워드, /effort, subagents, opusplan
* [세션 관리 편](./session-management) -- /rename, /branch, /color, 원격 제어
* [컨텍스트 관리 편](./context-management) -- /compact 파라미터, 압축 지시어, Handoff 프로토콜, MCP 지연 로딩
* [명령어와 자동화 편](./commands-automation) -- /diff, /simplify, /batch, /loop, Headless 모드
* [설정과 진단 편](./config-diagnostics) -- statusline, settings.json, /stats, /doctor
* [고급 편](./advanced) -- Hooks 심화, 플러그인 생태계, Agent Teams, 프롬프트 철학
### 더 읽어보기
* [Claude 시스템 아키텍처 완전 해설](/ko/docs/notes/claude-architecture) -- MCP, Skills, Subagents, Hooks 등 구성 요소 이해하기
* [Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent) -- 서브에이전트의 개념과 실전 활용
# 입력과 상호작용
## `!`:直接运行终端命令
在输入框以 `!` 开头,可以直接在 Claude Code 内执行终端命令,不需要切换到另一个终端窗口:
```
! git status
! npm run build
! docker ps
```
输入 `!` 加命令前缀后按 Tab 还能自动补全历史命令。
## `@` + 文件路径:注入文件上下文
在输入时用 `@` 加文件路径,可以把文件内容直接注入到上下文中:
```
帮我看看 @src/auth/login.ts 和 @src/auth/middleware.ts 之间的逻辑有没有问题
```
支持 Tab 键自动补全路径,不需要手动输入完整路径。比让 Claude 自己去读文件更快,因为省去了工具调用的开销。
## 直接粘贴 URL
直接把 URL 粘贴到输入中,Claude 会自动抓取网页内容作为上下文:
```
参考这个 API 文档 https://docs.example.com/api/v2 来写客户端代码
```
## 喂 `/llms-full.txt` 让 Claude 自己查文档
很多开源项目的文档站点会提供 `/llms-full.txt` 文件(LLM 友好的完整文档)。遇到某个库的问题时,把这个文件的 URL 粘贴给 Claude,它能自己查文档解决绝大部分问题:
```
参考 https://docs.astro.build/llms-full.txt 帮我解决这个路由问题
```
## `/btw`:在 Claude 工作时插嘴
这是 2026 年 3 月刚加的新功能。当 Claude 正在执行任务时,你可以用 `/btw` 发起一个旁路对话——问问它在想什么、给它补充信息,而不需要打断当前任务。
正如 Anthropic 工程师 @trq212 在推特上说的:「没人会用 Ctrl+C 打断同事,你只需要说一句 'btw',他们就会抬头看你。」
## Vim 模式
输入 `/vim` 开启 Vim 模式,支持:
* 模式切换(Normal/Insert)
* 导航(h/j/k/l, w/b/e, 0/$)
* 编辑操作(d, c, y, p)
* 文本对象(iw, aw, i", a())
如果你是 Vim 用户,这比默认的输入体验好太多。用 `/config` 可以设置为永久开启。
## 语音模式
输入 `/voice` 激活语音模式,长按空格键说话,松开发送。适合不想打字但又需要给 Claude 交代任务的时候。按键可以在 `keybindings.json` 中自定义。
# 세션 관리
## `/rename`:세션에 이름 붙이기
```
/rename my-auth-refactor
```
현재 세션에 이름을 지정합니다. 이름을 지정하면 대화형 세션 선택기(`claude --resume`)에서 이름이 있는 세션을 바로 선택하여 복원할 수 있으며, Enter 키를 눌러 확인할 필요가 없습니다. 터미널에서 `claude --resume my-auth-refactor`로 바로 실행할 수도 있습니다.
선택기에서 텍스트를 입력하면 바로 검색 및 필터링이 가능하며, 다음 단축키도 지원합니다: `Ctrl+V`로 세션 미리보기, `Ctrl+R`로 이름 변경, `Ctrl+A`로 전체 프로젝트 표시 전환, `Ctrl+B`로 브랜치별 필터링.
## `/branch`:대화 분기 만들기
git의 브랜치처럼 `/branch`는 현재 대화 지점에서 포크를 생성합니다. 포크에서 다양한 접근법을 시도할 수 있으며, 원래 대화에는 영향을 주지 않습니다. 결과가 만족스럽지 않으면 원래 브랜치로 돌아가서 계속하면 됩니다.
## `/color`:창에 색상 입히기
현재 세션의 프롬프트 바에 색상을 설정합니다. red, blue, green, yellow, purple, orange, pink, cyan을 지원합니다.
## 명령줄 세션 관리 도구 모음
```bash
# 恢复当前目录最近的会话
claude --continue
# 打开会话选择器,或按名称恢复
claude --resume
claude --resume my-auth-refactor
# 启动时直接命名会话
claude -n "auth-refactor"
# Fork 上一次会话(保留上下文,创建新分支)
claude -c --fork-session
# 恢复与特定 PR 关联的会话
claude --from-pr 123
# 在隔离的 git worktree 中启动
claude -w
```
Claude Code의 세션은 자동 저장됩니다(Ctrl+S가 필요 없습니다). 터미널을 열 때마다 `--continue`로 지난 작업을 이어서 할 수 있습니다.
## `claude --remote`:다른 기기에서 이어서 작업하기
```bash
claude --remote "your task description"
```
웹 세션을 시작하여 claude.ai나 모바일 앱에서 계속 작업할 수 있습니다.
## `/remote-control`:스마트폰으로 로컬 Claude 원격 제어
컴퓨터의 Claude Code에서 `/remote-control`을 입력하면 연결 코드가 생성됩니다. 그런 다음 스마트폰의 Claude 앱에서 이 연결 코드를 입력하면 로컬 Claude Code 세션을 원격으로 제어할 수 있습니다. 스마트폰에서 컴퓨터의 Claude에게 명령을 내릴 수 있는 것입니다.
# 키보드 단축키
Claude Code의 단축키 체계는 대부분의 사용자가 생각하는 것보다 훨씬 풍부합니다. `?`를 누르면 현재 컨텍스트에서 사용 가능한 모든 단축키를 확인할 수 있습니다.
## Shift+Tab: 모드 순환 전환
아마도 가장 중요한 단축키입니다. `Shift+Tab`을 누르면 세 가지 모드를 순환하며 전환할 수 있습니다:
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
`/plan`이나 `/auto-accept`를 직접 입력할 필요 없이 키 하나로 해결됩니다. 저의 사용 방식은 새로운 작업을 받으면 두 번 눌러 Plan Mode로 전환하고, 방향을 확인한 후 한 번 더 눌러 Auto-Accept로 전환하여 Claude가 자율적으로 실행하도록 하는 것입니다.
## Esc + Esc: 타임머신
`Esc`를 두 번 연속 누르면 되감기 메뉴(Rewind)가 나타납니다:
* **코드와 대화 복원**: 이전 체크포인트로 돌아가 파일과 대화 기록 모두 롤백됩니다
* **대화만 복원**: 메시지를 롤백하되 현재 코드 변경 사항은 유지합니다
* **코드만 복원**: 파일 수정을 되돌리되 대화 기록은 유지합니다
Claude는 파일을 편집할 때마다 자동으로 체크포인트를 기록합니다. 이는 `git checkout .`보다 훨씬 세밀한 제어가 가능합니다. 마지막 커밋으로만 돌아가는 것이 아니라 원하는 편집 단계로 돌아갈 수 있기 때문입니다.
단, 주의할 점이 있습니다. Claude가 도구를 통해 직접 편집한 파일만 추적됩니다. 직접 수정한 파일이나 `git push` 등 외부 작업은 되감기 대상이 아닙니다.
## Ctrl+S: 프롬프트 임시저장 (Prompt Stash)
프롬프트를 작성하던 중 다른 작업을 먼저 처리해야 할 때, `Ctrl+S`를 누르면 현재 입력 내용이 임시저장됩니다:
그런 다음 다른 명령이나 질문을 입력할 수 있습니다. 해당 메시지를 전송하면 임시저장된 내용이 입력란에 **자동으로 복원**되어 중단한 곳에서 이어서 작성할 수 있습니다.
`git stash`의 프롬프트 버전이라고 생각하면 됩니다. 사용 예시: 긴 리팩토링 요구사항을 작성하다가 Claude에게 먼저 특정 파일을 확인해달라고 하고 싶을 때, `Ctrl+S`로 요구사항 설명을 임시저장하고 파일 관련 질문을 한 뒤 답변을 받으면 요구사항 설명이 자동으로 돌아옵니다.
## Ctrl+B: 작업을 백그라운드로 보내기
Claude가 시간이 오래 걸리는 작업(대규모 리팩토링 등)을 처리하고 있는데 다른 작업을 하고 싶다면, `Ctrl+B`를 눌러 현재 작업을 백그라운드로 보낼 수 있습니다. 터미널이 즉시 새로운 입력을 받을 수 있게 됩니다.
`Ctrl+T`로 백그라운드 작업 목록을 확인하고, `Ctrl+F`를 두 번 눌러 모든 백그라운드 에이전트를 종료할 수 있습니다.
> tmux 사용자 주의: tmux의 기본 프리픽스 키도 `Ctrl+B`이므로, Claude의 백그라운드 기능을 사용하려면 두 번 눌러야 합니다.
## Ctrl+G: 에디터에서 긴 프롬프트 작성하기
Claude에게 상세한 지시를 내려야 하는데 터미널에서 타이핑하기 불편할 때가 있습니다. `Ctrl+G`를 누르면 시스템 기본 `$EDITOR`(VS Code, Vim 등)가 열리고, 거기서 프롬프트를 작성한 후 저장하고 닫으면 자동으로 Claude에 전송됩니다.
기본 에디터를 변경하려면 셸 설정 파일(`~/.zshrc` 또는 `~/.bashrc`)에서 설정합니다:
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
`--wait` 매개변수가 중요합니다. 파일을 닫을 때까지 에디터가 대기하도록 지시하는 것으로, 이것이 없으면 Claude가 빈 내용을 즉시 받게 됩니다. Vim 같은 터미널 에디터는 기본적으로 블로킹 동작이므로 이 매개변수가 필요 없습니다.
여러 단락에 걸친 요구사항 설명이나 대량의 참고 자료를 붙여넣을 때 특히 유용합니다. Plan Mode에서는 `Ctrl+G`를 사용하여 Claude가 생성한 계획을 에디터에서 직접 편집할 수도 있습니다.
## Cmd+T: 확장 사고 전환
기본 단축키는 `Cmd+T`(Windows/Linux에서는 `Meta+T`)이며, 확장 사고(Extended Thinking) 모드를 켜고 끕니다. 활성화하면 Claude가 응답하기 전에 더 깊이 추론하므로 복잡한 아키텍처 결정이나 까다로운 버그 추적에 적합합니다.
주의할 점은 대부분의 터미널(iTerm2, Terminal.app, Warp 등)이 `Cmd+T`를 '새 탭'으로 가로채기 때문에 실제로 이 단축키가 작동하지 않는 경우가 많다는 것입니다. 해결 방법은 두 가지입니다: `/keybindings`로 충돌하지 않는 키에 재할당하거나, `/effort` 명령으로 사고 깊이를 전환하는 것입니다(동일한 효과이며 레벨을 세밀하게 제어할 수도 있습니다).
## Readline 단축키
Claude Code의 입력란은 표준 Readline 단축키를 지원하며, 터미널에 익숙한 분들이라면 바로 사용할 수 있습니다:
| 단축키 | 기능 |
| ----------------- | ------------------ |
| Ctrl+A | 줄 시작으로 이동 |
| Ctrl+E | 줄 끝으로 이동 |
| Ctrl+W | 이전 단어 삭제 |
| Ctrl+U | 줄 시작까지 삭제 |
| Ctrl+K | 줄 끝까지 삭제 |
| Ctrl+Y | 마지막으로 삭제한 텍스트 붙여넣기 |
| Alt+Y | 삭제 기록 순환 |
| Option+Left/Right | 단어 단위 이동 (Mac) |
## 승인 단축키: `y/n/d/e`
Claude가 파일 변경을 제안하고 확인을 기다릴 때, 네 개의 단일 키 단축키로 흐름을 제어합니다:
* `y`: 수락
* `n`: 거부
* `d`: 전체 diff 보기
* **`e`: 편집 후 수락**
`e`는 가장 간과되지만 가장 유용한 단축키입니다. Claude의 변경 사항을 미세 조정한 후 적용할 수 있습니다. 몇 줄의 코드가 마음에 들지 않더라도 거부하고 다시 시작할 필요 없이 `e`를 눌러 수정하면 됩니다.
## 빠른 참조
| 단축키 | 기능 |
| ----------- | ------------------------------------------------- |
| Shift+Tab | 모드 전환: Normal → Auto-Accept → Plan |
| Esc+Esc | 되감기 메뉴 열기 |
| Ctrl+S | 현재 입력 임시저장, 다음 전송 후 자동 복원 |
| Ctrl+B | 현재 작업을 백그라운드로 보내기 |
| Ctrl+T | 백그라운드 작업 목록 보기 |
| Ctrl+F (x2) | 모든 백그라운드 에이전트 종료 |
| Ctrl+G | 외부 에디터에서 프롬프트 작성 |
| Ctrl+O | 상세 도구 출력 보기 전환 |
| Cmd+T | 확장 사고 전환 (터미널에 가로채일 수 있음. 재할당 또는 `/effort` 사용 권장) |
| `\` + Enter | 여러 줄 입력 (설정 불필요) |
| Shift+Enter | 여러 줄 입력 (`/terminal-setup` 먼저 실행 필요) |
| Up / Down | 입력 기록 탐색 |
| Ctrl+R | 명령 기록 검색 |
| Ctrl+L | 화면 지우기 (기록 유지) |
| Ctrl+C | 현재 생성 취소 |
| Ctrl+D | Claude Code 종료 |
| `?` | 사용 가능한 모든 단축키 표시 |
## 사용자 정의 키바인딩
기본 단축키가 맞지 않으면 `/keybindings`로 `~/.claude/keybindings.json`을 열어 사용자 정의할 수 있습니다. 변경 사항은 즉시 적용되며 재시작할 필요가 없습니다.
조합 키 구문(예: `ctrl+shift+c`)과 코드 모드(예: `ctrl+k ctrl+s` -- Ctrl+K를 누르고 놓은 후 Ctrl+S를 누름)를 지원합니다. 16가지 바인딩 컨텍스트(Chat, Autocomplete, Confirmation, DiffDialog 등)가 있으며, 각 컨텍스트에 서로 다른 동작을 할당할 수 있습니다.
# 사고 및 모델 제어
## 키워드로 사고 깊이 제어하기
프롬프트에 특정 키워드를 추가하면 서로 다른 수준의 사고 예산을 트리거할 수 있습니다. 이것은 Claude Code만의 고유 기능입니다(claude.ai 웹 버전에는 없습니다):
| 키워드 | 사고 예산 | 적용 시나리오 |
| ----------------------------- | ----------- | ---------------- |
| `think` | \~4,000 토큰 | 일상적인 코딩 질문 |
| `think hard` / `megathink` | \~10,000 토큰 | 복잡한 로직, 다중 파일 연관 |
| `think harder` / `ultrathink` | \~31,999 토큰 | 아키텍처 설계, 까다로운 버그 |
실제 사용에서는 Claude가 얕은 답변을 할 때 `think hard`를 추가하여 다시 질문하는 경우가 많습니다. 특히 복잡한 문제(예: 여러 서비스에 걸친 버그 추적)의 경우에는 바로 `ultrathink`를 사용합니다.
## `/effort`: 사고 깊이 제어하기
키워드(think / ultrathink) 외에도 `/effort`로 사고 깊이를 직접 설정할 수 있습니다:
```
/effort low # 简单任务,跳过深度思考,更快更省
/effort high # 复杂任务,深度推理
/effort max # 最大思考预算(仅 Opus)
/effort auto # 让 Claude 自己判断
```
설정은 전체 세션 동안 유지됩니다. 간단한 파일 수정에는 `low`를, 복잡한 아키텍처 설계에는 `max`를 사용하면 품질을 희생하지 않으면서 비용을 절약할 수 있습니다.
## `use subagents` 키워드
요청 뒤에 `use subagents`를 추가하면 Claude가 작업을 여러 서브 에이전트로 분해하여 병렬로 처리합니다. 이렇게 하면 속도가 빨라질 뿐만 아니라 메인 에이전트의 컨텍스트 윈도우를 깔끔하게 유지할 수 있습니다.
Boris는 트위터에서 이 점을 특별히 언급했습니다: 개별 작업을 서브 에이전트에 위임하여 메인 에이전트의 컨텍스트를 집중시킨다는 것입니다.
## `opusplan`: 최고의 가성비 모델 전략
한 줄 요약: **Opus가 생각하고, Sonnet이 실행한다**.
## `/model`: 모델 전환하기
`/model`을 사용하면 세션 중 언제든지 모델을 전환할 수 있습니다. 예를 들어 평소에는 Sonnet을 사용하다가 복잡한 문제를 만나면 임시로 Opus로 전환하고, 해결 후 다시 돌아오는 방식입니다.
## 출력 스타일 제어
`/config`에서 "Output style"을 선택하면 잘 알려지지 않았지만 매우 유용한 두 가지 모드가 있습니다:
* **Explanatory 모드**: Claude가 작업 사이에 "지식 포인트"를 삽입하여 관련 프레임워크와 코드 패턴을 설명합니다 — 새로운 프로젝트를 학습하는 데 적합합니다
* **Learning 모드**: 협업 학습 모드로, Claude가 직접 답을 주는 대신 코드에 `TODO(human)` 마커를 추가하여 직접 구현하도록 합니다
`~/.claude/output-styles/`에 커스텀 출력 스타일 파일(Markdown 형식)을 생성하여 시스템 프롬프트를 직접 수정할 수도 있습니다. 참고: 커스텀 출력 스타일은 `keep-coding-instructions: true`를 설정하지 않으면 기본 코딩 시스템 프롬프트를 **완전히 대체합니다**.
# 개념 소개
## 서론
2025년 10월, Anthropic은 Claude Skills라는 새로운 기능을 조용히 출시했습니다. 겉보기에 소박한 이 업데이트를 저명한 기술 블로거 Simon Willison은 "MCP보다 더 중요할 수 있다"고 평가하며, AI 도구 분야에서 "캄브리아기 대폭발"을 일으킬 것이라고 예측했습니다.
이러한 평가는 근거 없는 것이 아닙니다. AI 어시스턴트를 자주 사용하신다면 이런 어려움을 겪어본 적이 있을 것입니다: 새로운 대화를 시작할 때마다 같은 워크플로우 설명을 반복 입력해야 하고, 힘들게 AI를 만족스러운 상태로 조정해놓아도 대화 창을 바꾸면 처음부터 다시 시작해야 합니다. Skills는 바로 이 문제를 해결하기 위해 탄생했습니다.
## Claude Skills 이해하기
여러분이 한 회사의 사장이라고 상상해 보십시오. 신입 사원이 입사하면 회사의 업무 프로세스, 브랜드 규정, 자주 발생하는 문제의 처리 방법이 상세히 기록된 업무 매뉴얼을 줍니다. Claude Skills는 AI 어시스턴트에게 주는 바로 이 "업무 매뉴얼"입니다. AI가 반복 가능하고 표준화된 방식으로 특정 작업을 수행할 수 있게 해줍니다.
기술적 관점에서 보면, Skills는 지시, 스크립트, 리소스가 포함된 폴더로, Claude가 필요할 때 동적으로 로드할 수 있습니다. 각 Skill은 Claude에게 특정 유형의 작업을 일관된 방식으로 수행하는 방법을 가르치며, 이러한 지식은 대화 간에 영구적으로 저장됩니다. 즉, 한 번만 "교육"하면 이후 언제 사용하든 Claude가 어떻게 해야 하는지 기억합니다.
### 세 가지 구성 요소
완전한 Skill은 다음 세 부분으로 구성됩니다:
| 구성 요소 | 역할 | 필수 여부 |
| ------------ | --------------------------------------- | ----- |
| **SKILL.md** | 핵심 지시 문서, 메타데이터와 상세 지시 포함 | 필수 |
| **참고 자료** | 브랜드 가이드, 정책 문서, 템플릿 등 보충 정보 | 선택 |
| **스크립트** | Python/JavaScript 코드, 복잡한 계산이나 파일 작업 처리 | 선택 |
이 중 SKILL.md는 전체 Skill의 "핵심"이며, 기본 구조는 다음과 같습니다:
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
파일 시작 부분의 YAML frontmatter에는 두 가지 핵심 필드가 있습니다: `name`은 기술의 식별 이름으로 최대 64자이며, `description`은 Claude에게 이 기술이 무엇을 하는지, 언제 사용해야 하는지를 알려주며 최대 200자입니다. Claude는 바로 이 설명을 기반으로 특정 Skill을 호출할 시점을 판단하므로, 명확하고 정확하게 작성할수록 Skill이 올바르게 트리거될 확률이 높아집니다.
### 활용 시나리오
Skills의 활용 시나리오는 매우 다양하며, 일상 업무에서의 각종 반복 작업을 폭넓게 다룹니다:
**문서 처리**: Excel 스프레드시트, PPT 프레젠테이션, Word 문서, PDF 보고서를 일괄 생성합니다. Anthropic 공식에서 바로 사용할 수 있는 문서 기술 세트를 제공하고 있습니다.
**브랜드 준수**: 회사의 브랜드 색상, 로고 사용 규칙, 간격 규정, 어조 스타일을 Skill로 패키징하여 AI가 생성하는 모든 콘텐츠가 브랜드 표준을 준수하도록 합니다.
**회의록**: 회의 기록을 자동 요약하고, 액션 아이템을 추출하며, 담당자를 배정하고, 후속 이메일을 생성합니다.
**데이터 분석**: 표준화된 분석 프로세스를 실행합니다. 예를 들어 경쟁 인텔리전스 스캔(제품 업데이트, 가격 변동, 애널리스트 코멘트의 구조화된 추출), 재무 분석(재무제표 분석, 재무 모델 구축) 등이 있습니다.
**프로젝트 관리**: 목표로부터 프로젝트 계획을 수립하고, 마일스톤을 제안하며, 주간 보고서/투자자 브리핑을 생성합니다.
## 점진적 공개 아키텍처
Skills의 가장 정교한 설계는 정보 로딩 방식에 있습니다. 기존 MCP 도구 설명은 수천에서 수만 개의 token을 소비할 수 있지만, Skills의 메타데이터는 수십 개의 token만 차지합니다. 이는 많은 수의 Skills를 동시에 활성화해도 컨텍스트 윈도우가 도구 설명으로 가득 찰 걱정이 전혀 없다는 것을 의미합니다.
이러한 효율성은 **점진적 공개**(Progressive Disclosure)라는 아키텍처 설계에서 비롯됩니다. Skills는 3단계 정보 구조를 채택하여 필요에 따라 단계별로 로드합니다. 마치 목차가 있는 매뉴얼과 같습니다:
```
📚 Skills 작업 매뉴얼
│
├─ 📋 목차 ─────────────────────────── 【메타데이터 층】시작 시 사전 로드
│ │
│ │ name: "weekly-report"
│ │ description: "업무 내용을 기반으로 표준화된 주간 보고서 생성"
│ │
│ │ ✓ 30-50 tokens만 차지
│ │ ✓ 모든 Skills의 목차가 동시에 표시
│ │
│
├─ 📖 본문 섹션 ─────────────────────── 【핵심 문서 층】관련 시 로드
│ │
│ │ # Weekly Report Generator
│ │
│ │ ## Instructions
│ │ 다음 구조에 따라 주간 보고서 생성...
│ │
│ │ ## Examples
│ │ 입력: 이번 주에 로그인 기능을 완료...
│ │ 출력: ### 이번 주 완료 사항 ...
│ │
│ │ ⚡ Claude가 필요하다고 판단할 때만 펼침
│ │ 📊 수백에서 수천 tokens 소비
│ │
│
└─ 📎 부록 ─────────────────────────── 【참조 리소스 층】필요 시 로드
│
│ references/
│ ├── brand-guide.md 브랜드 규정
│ ├── template.xlsx 보고서 템플릿
│ └── examples/ 과거 주간 보고서
│
│ 🔍 명확히 필요할 때만 로드
│ 📦 대량의 참고 자료 포함 가능
```
먼저 목차를 보고 어떤 섹션이 있는지 파악하고(메타데이터 층), 필요한 섹션을 찾아 펼쳐 읽은 뒤(핵심 문서 층), 더 많은 세부 정보가 필요하면 부록을 참조합니다(참조 리소스 층).
| 계층 | 내용 | 로드 시점 | Token 소비 |
| ------------ | ------------------ | ---------- | -------- |
| **메타데이터 층** | name + description | 시작 시 사전 로드 | 30-50 |
| **핵심 문서 층** | SKILL.md 전체 내용 | 관련 시 로드 | 수백에서 수천 |
| **참조 리소스 층** | 참고 파일, 템플릿 등 | 필요 시 로드 | 필요에 따라 |
이는 대규모 언어 모델의 본질과 완전히 일치합니다. "텍스트를 입력하여 모델이 이해하게 하는 것"입니다. Skills는 복잡한 프로토콜이나 API 호출을 도입하지 않고, 정교하게 구성된 텍스트 구조를 통해 AI가 효율적으로 지식을 획득하고 활용할 수 있게 합니다. Simon Willison은 이 설계를 "경이로울 정도로 간결하다"고 평가했는데, 이는 복잡한 문제를 가장 소박한 방식으로 해결했기 때문입니다.
## 핵심 장점
### Token 효율성
Skills의 점진적 공개 아키텍처는 매우 높은 token 효율성을 제공합니다. 간단한 비교를 통해 이해할 수 있습니다:
| 방안 | 시작 시 Token 소비 | 100개 기술의 총 소비 |
| --------------- | ------------- | -------------- |
| 기존 방안 (전체 로드) | 수천에서 수만 | 컨텍스트 윈도우 초과 가능 |
| Skills (점진적 공개) | 30-50 | 3000-5000 |
각 기술의 메타데이터가 수십 개의 token만 차지하므로, 수십에서 수백 개의 Skills를 동시에 활성화할 수 있으며, 전체 내용은 필요에 따라 로드되어 소중한 컨텍스트 공간을 낭비하지 않습니다.
### 조합 가능성
여러 Skills가 자동으로 협업할 수 있습니다. 복잡한 작업을 요청하면 Claude가 어떤 Skills를 호출해야 하는지 지능적으로 식별하고, 이를 조율하여 함께 작업을 완료합니다.
예를 들어 "이 매출 데이터를 기반으로 분기 보고서를 생성해주세요"라고 말하면 Claude는 다음과 같이 할 수 있습니다:
1. 데이터 분석 Skill을 호출하여 원시 데이터 처리
2. 차트 생성 Skill을 호출하여 시각화 생성
3. 문서 Skill을 호출하여 최종 보고서 생성
전체 과정에서 어떤 Skill을 사용할지 수동으로 지정할 필요가 없으며, Claude가 작업 요구에 따라 자동으로 선택하고 조합합니다.
### 이식 가능성
같은 Skill을 Anthropic 생태계의 모든 플랫폼에서 사용할 수 있습니다:
| 플랫폼 | 설명 |
| ----------- | -------------------- |
| Claude.ai | 웹 버전, 일반 사용자에게 적합 |
| Claude Code | 명령줄 도구, 개발자에게 적합 |
| API | 프로그래밍 통합, 시스템 개발에 적합 |
팀을 위해 만든 브랜드 라이팅 Skill은 이 모든 플랫폼에서 일관된 동작을 유지하며, 진정한 의미의 **한 번 구축, 어디서나 사용**을 실현합니다.
> **다른 AI 플랫폼의 전략**: 현재 Skills는 Anthropic 고유의 기능입니다. OpenAI는 Custom GPTs + Assistants API의 이원 전략(두 시스템이 통합되지 않음)을 채택하고 있으며, Microsoft Copilot과 Google Gemini는 재사용 가능한 기술 모듈보다는 각자의 생태계 깊은 통합에 집중하고 있습니다. Claude Skills는 의미 있는 차별화 특성으로 평가받고 있습니다.
### 효율성 데이터
Anthropic 내부 벤치마크 테스트에 따르면, Skills를 사용하는 팀은 **반복 프롬프트 엔지니어링 시간을 73% 절감**했습니다. 이는 효율성 향상뿐 아니라, 더 중요하게는 워크플로우의 표준화와 재사용성을 의미합니다. 팀원들이 각자 프롬프트 세트를 유지할 필요 없이 검증된 동일한 Skills를 공유할 수 있습니다.
## 요약
Claude Skills는 본질적으로 AI 어시스턴트를 위한 **재사용 가능한 작업 매뉴얼**입니다. 점진적 공개 아키텍처를 통해 매우 높은 token 효율성을 달성하여, AI가 소중한 컨텍스트 공간을 차지하지 않으면서 대량의 전문 지식을 보유할 수 있게 합니다.
세 가지 키워드만 기억하면 Skills의 정수를 파악한 것입니다:
| 키워드 | 의미 |
| --------- | ------------------------------- |
| **효율적** | 메타데이터는 수십 개의 token만 차지, 필요 시 로드 |
| **조합 가능** | 여러 Skills가 자동으로 협업 |
| **이식 가능** | 크로스 플랫폼 일관된 경험 |
개념을 이해하셨다면, 다음 글 《[Claude Skills 실전 가이드](/ko/docs/notes/claude-skills/practice)》에서 직접 실습해 보겠습니다: Skills를 활성화하고 설치하는 방법, 첫 번째 커스텀 Skill을 만드는 방법, 그리고 흔한 함정을 피하는 방법을 다룹니다.
워크플로우를 더욱 체계적으로 정리하고 싶으시다면, 《[스펙 주도 개발이란 무엇인가](/ko/docs/notes/speckit/concept)》를 참고하여 AI 프로그래밍을 "직관"에서 "엔지니어링"으로 업그레이드하는 방법을 알아보십시오.
# 실전 가이드
## 간단한 복습
[이전 글](/ko/docs/notes/claude-skills/concept)에서 Skills의 핵심 개념을 알아보았습니다: AI 어시스턴트를 위한 재사용 가능한 작업 매뉴얼로, 점진적 공개 아키텍처를 통해 매우 높은 token 효율성을 달성하며, 효율성, 조합 가능성, 이식 가능성이라는 세 가지 특징을 갖고 있습니다. 이 글에서는 실전적 관점에서 출발하여, Skills와 다른 기능의 차이를 이해하고, Skills의 활성화, 설치 및 생성을 배우며, 모범 사례를 익히고 흔한 함정을 피하는 방법을 알려드리겠습니다.
## 기능 비교
Claude 생태계에는 다양한 기능이 있어 처음 접하면 그 차이가 헷갈릴 수 있습니다. 아래 표를 통해 빠르게 구분할 수 있습니다:
| 기능 | 무엇인가 | 가장 적합한 용도 | 지속성 |
| ------------- | --------- | --------------- | ------------- |
| **Skills** | 전문 지식 패키지 | 반복 작업, 표준화 프로세스 | 대화 간 영구 |
| **Prompts** | 즉시 지시 | 일회성 요청 | 현재 대화에만 |
| **Projects** | 지식 베이스 | 배경 정보, 프로젝트 문서 | 프로젝트 워크스페이스 내 |
| **MCP** | 커넥터 | 외부 데이터, API 호출 | 지속 연결 |
| **Subagents** | 하위 에이전트 | 작업 위임, 병렬 처리 | 세션 간 |
### Skills vs MCP
이것이 가장 흔한 혼동입니다. 핵심 차이: **MCP는 Claude를 데이터에 연결하고, Skills는 Claude에게 데이터를 처리하는 방법을 가르칩니다**. 둘은 대체 관계가 아니라 보완 관계입니다.
| 차원 | Skills | MCP |
| ------------ | -------------------------- | --------------------------- |
| **핵심 기능** | Claude에게 작업 수행 방법을 가르침 | Claude를 외부 시스템에 연결 |
| **Token 소비** | 매우 낮음 (수십 개 token) | 비교적 높음 (수천에서 수만 token) |
| **기술 복잡도** | 간단 (Markdown + YAML) | 복잡 (완전한 프로토콜 규격) |
| **대표 시나리오** | 브랜드 라이팅, 보고서 생성, 워크플로우 | 데이터베이스 쿼리, API 호출, 클라우드 서비스 |
| **이식 가능성** | Claude.ai/Code/API 크로스 플랫폼 | 여러 모델 회사에서 채택 |
이 차이를 이해하면 언제 무엇을 사용해야 하는지 알 수 있습니다. 데이터베이스 쿼리, API 호출, 클라우드 서비스 접근이 필요할 때는 MCP를, 특정 라이팅 스타일 준수, 표준화된 프로세스 실행, 전문 지식 재사용이 필요할 때는 Skills를 사용합니다.
모범 사례는 두 가지를 결합하여 사용하는 것입니다: MCP로 CRM 시스템에 연결하여 고객 데이터를 가져오고, Skills로 해당 데이터를 분석하고 보고서를 생성하는 방법을 정의합니다.
### Skills vs Subagents
핵심 차이: **Skills는 Claude가 특정 유형의 작업에 더 능숙해지게 하고, Subagents는 Claude가 독립적인 "전문 직원"에게 작업을 위임하게 합니다**.
| 차원 | Skills | Subagents |
| ----------- | ---------------------------- | ---------------------------- |
| **핵심 기능** | 전문 지식과 지시 제공 | 독립적으로 작업을 수행하는 하위 에이전트 |
| **컨텍스트** | 메인 대화 컨텍스트에 주입 | 독립적인 컨텍스트 윈도우 보유 |
| **적용 시나리오** | Claude가 특정 유형의 작업에 더 능숙해지게 함 | 복잡하고 다단계의 독립적 작업 |
| **활성화 방식** | 설명에 기반하여 자동 매칭 | 수동 호출 또는 Claude가 자동 위임 |
| **이식 가능성** | Claude.ai/Code/API 크로스 플랫폼 | Claude Code 및 Agent SDK에만 해당 |
비유하자면, Skills는 교육 자료와 같습니다. Claude가 특정 작업을 수행하는 방법을 배우게 합니다. [Subagents](/ko/docs/notes/claude-subagent)는 전담 직원과 같습니다. 자신만의 자리(컨텍스트)와 권한(도구)을 갖고, 독립적으로 작업을 완료한 후 결과를 보고합니다.
두 가지를 조합하여 사용할 수 있습니다: 예를 들어 코드 리뷰 하위 에이전트가 언어별 모범 사례 Skill을 로드하여 "전문가 + 전문 지식"의 조합 효과를 실현할 수 있습니다. [Anthropic 연구](https://www.anthropic.com/engineering/multi-agent-research-system)에 따르면, 멀티 에이전트 시스템(Claude Opus 4 메인 에이전트 + Claude Sonnet 4 하위 에이전트)은 내부 평가에서 단일 에이전트보다 90.2% 높은 성과를 보였습니다.
### Skills vs 슬래시 명령어
Claude Code를 사용해 보셨다면 `/commit`, `/review`와 같은 [슬래시 명령어](/ko/blog/claude-code-best-practices)에 익숙하실 것입니다. 핵심 차이: **Skills는 컨텍스트에 따라 자동 활성화되고, 슬래시 명령어는 수동으로 입력하여 트리거해야 합니다**.
| 차원 | Skills | 슬래시 명령어 (Slash Commands) |
| ----------- | ---------------------------------- | ------------------------ |
| **활성화 방식** | 자동 활성화 (컨텍스트 매칭 기반) | 수동 입력 (예: `/commit`) |
| **트리거 조건** | Claude가 description을 기반으로 관련 여부 판단 | 사용자가 명확히 명령어 입력 |
| **적용 시나리오** | "항시 켜짐" 능력 강화 | 명확하고 반복 가능한 작업 |
| **사용자 인식** | 인식 없이 자동 적용 | 명령어 이름을 기억해야 함 |
예를 들어 설명하겠습니다: `/commit`을 입력하면 Claude가 사전 정의된 커밋 프로세스를 실행합니다. 이것이 슬래시 명령어입니다. "주간 보고서를 작성해주세요"라고 말하면 Claude가 자동으로 주간 보고서 생성 Skill을 식별하여 로드하며, 어떤 명령어도 입력할 필요가 없습니다. 이것이 Skills입니다.
간단히 기억하면: 슬래시 명령어는 단축키로 여러분이 직접 트리거해야 합니다. Skills는 배경 지식으로 Claude가 자동으로 언제 사용할지 판단합니다.
### Skills vs Plugins
Plugins는 Claude Code의 확장 패키지 메커니즘입니다. 핵심 차이: **Skills는 자동 활성화되는 능력 확장이고, Plugins는 패키징하여 배포하는 완전한 워크플로우 구성입니다**.
| 차원 | Skills | Plugins |
| ----------- | ---------------------------- | ------------------------ |
| **핵심 기능** | 전문 능력 확장 | 워크플로우 패키징 배포 |
| **활성화 방식** | 컨텍스트에 따라 자동 활성화 | 설치 후 구성 요소 병합 |
| **적용 범위** | 크로스 플랫폼 (Claude.ai/Code/API) | Claude Code에만 해당 |
| **포함 내용** | 지시 + 스크립트 + 리소스 | 슬래시 명령어 + hooks + skills |
| **배포 메커니즘** | 개별 폴더 | marketplace를 통해 설치 |
핵심 이해: Plugins는 Skills를 포함할 수 있으며(`skills/` 디렉토리 내), 더 큰 패키징 단위입니다. Plugin을 설치하면 그 안의 Skills가 자동으로 활성화되고, 슬래시 명령어가 자동 완성에 표시되며, hooks가 기존 구성과 병합됩니다.
간단히 말해: Skills로 Claude의 능력을 확장하고, Plugins로 팀 간에 표준화된 워크플로우 구성을 배포합니다.
## 실전 튜토리얼
### 방법 1: 내장 Skills 활성화
가장 간단한 입문 방법입니다. Anthropic 공식에서 실용적인 문서 기술 세트를 제공하고 있습니다:
| 기술 | 기능 |
| --------------------- | --------------------------------- |
| **Excel (xlsx)** | 스프레드시트 생성, 데이터 분석, 차트가 포함된 보고서 생성 |
| **PowerPoint (pptx)** | 프레젠테이션 생성, 슬라이드 편집, 프레젠테이션 내용 분석 |
| **Word (docx)** | 문서 생성, 내용 편집, 텍스트 서식 지정 |
| **PDF (pdf)** | 서식이 적용된 PDF 문서 및 보고서 생성 |
**활성화 단계**:
1. [Claude.ai](https://claude.ai)에 로그인합니다
2. 오른쪽 상단의 프로필을 클릭하여 **Settings**에 진입합니다
3. **Capabilities** 옵션을 찾습니다
4. 필요한 기술을 활성화합니다
활성화 후 바로 테스트할 수 있습니다: "Q3 매출 예산 Excel 스프레드시트를 만들어 주세요. 월별 명세와 합계를 포함해 주세요."
> **참고**: Pro, Max, Team 또는 Enterprise 플랜이 필요하며, 코드 실행 기능을 활성화해야 합니다.
### 방법 2: 커뮤니티 Skills 설치
Claude Code를 사용하고 계시다면, 명령어로 커뮤니티가 기여한 Skills를 설치할 수 있습니다.
**플러그인 마켓플레이스를 통한 설치**:
```bash
# 添加官方 Skills 仓库
/plugin marketplace add anthropics/skills
# 安装文档技能包
/plugin install document-skills@anthropic-agent-skills
# 安装示例技能包
/plugin install example-skills@anthropic-agent-skills
```
**Skills 저장 위치**:
| 위치 | 경로 | 설명 |
| ----------- | ------------------- | ------------------- |
| 개인 Skills | `~/.claude/skills/` | 본인만 사용 가능 |
| 프로젝트 Skills | `.claude/skills/` | git 버전 관리와 함께, 팀 공유 |
### 방법 3: 커스텀 Skill 생성
이것이 Skills의 진정한 위력이 있는 부분입니다. 자신만의 워크플로우를 만들 수 있습니다.
**1단계: 폴더 구조 생성**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
완전한 Skill 폴더의 예시는 다음과 같습니다:
```
weekly-report/
├── SKILL.md # 핵심 지시 (필수)
├── template.md # 주간 보고서 템플릿 (선택)
└── examples/ # 예시 주간 보고서 (선택)
├── good-example.md
└── bad-example.md
```
**2단계: SKILL.md 작성**
SKILL.md는 전체 Skill의 핵심입니다. YAML frontmatter(메타데이터)와 Markdown 본문(상세 지시) 두 부분으로 구성됩니다.
**필수 메타데이터**:
| 필드 | 요구 사항 | 설명 |
| ------------- | ------- | ----------------------------------- |
| `name` | 최대 64자 | 기술의 고유 식별 이름 |
| `description` | 최대 200자 | Claude에게 이 기술을 언제 사용할지 알려줌 (매우 중요!) |
**선택 메타데이터**:
| 필드 | 설명 |
| --------------- | ---------------------------------------------- |
| `dependencies` | 필요한 소프트웨어 패키지, 예: `python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | 허용된 도구 목록 |
| `model` | 선택적 모델 오버라이드 |
완전한 주간 보고서 생성 Skill 예시:
```yaml
---
name: weekly-report
description: 根据本周工作内容生成标准化的周报,包含进展、问题和下周计划
---
# 周报生成助手
## 使用场景
当用户需要生成周报、工作总结或进度汇报时,使用此技能。
## 输出格式
请按以下结构生成周报:
### 本周完成
- 列出已完成的主要工作项
- 每项包含简短说明和成果
### 进行中
- 列出正在进行的工作
- 标注当前进度和预期完成时间
### 遇到的问题
- 列出阻碍进展的问题
- 如果有,说明需要的支持
### 下周计划
- 列出下周的主要任务
- 按优先级排序
## 风格要求
- 使用简洁的表达
- 避免过于技术化的术语
- 突出成果和影响
## 示例
**输入**:这周完成了用户登录功能,修复了 3 个 bug,参加了产品评审。
**输出**:
### 本周完成
- 用户登录功能开发:完成前后端联调,支持邮箱和手机号登录
- Bug 修复:解决了 3 个高优先级问题,提升系统稳定性
### 进行中
- (无)
### 遇到的问题
- (无)
### 下周计划
- 开始用户注册功能开发
- 编写单元测试用例
```
**3단계: 테스트**
Claude에서 테스트합니다: "이번 주 주간 보고서를 작성해 주세요. 이번 주에 사용자 로그인 기능 개발을 완료하고, 3개의 버그를 수정하고, 제품 리뷰 회의에 2번 참석했습니다."
### Skill Creator 사용하기
SKILL.md를 처음부터 작성하고 싶지 않다면, Claude에 내장된 skill-creator 기술을 사용하여 대화형으로 안내받을 수 있습니다:
```
Help me create a skill for [your workflow]
```
Claude가 일련의 질문을 통해 요구 사항을 정리한 다음, SKILL.md 초안을 생성해 줍니다.
## 기술 원리
### Skills의 메타 도구 시스템
Skills는 본질적으로 **메타 도구 시스템**입니다. 코드를 직접 실행하지 않고, 전문화된 지시를 대화 컨텍스트에 주입하여 Claude의 추론 방식을 변경합니다.
Skill을 트리거하면 두 가지 일이 발생합니다:
1. **메타데이터 메시지**: 어떤 Skill이 로드되고 있는지 표시하는 가시적 상태 인디케이터
2. **기술 프롬프트**: 완전한 SKILL.md 지시가 Claude에게 전송되지만, 사용자에게는 숨겨짐
### 발견 및 선택 메커니즘
Claude는 어떤 Skill을 호출해야 하는지 어떻게 알까요? 답은: **완전히 언어 이해에 의존합니다**.
활성화된 모든 Skills의 name과 description이 동적 목록으로 형식화되어 시스템 프롬프트에 기록됩니다. 메시지를 보내면 Claude가 네이티브 언어 이해 능력을 사용하여 의도를 매칭하고, 특정 Skill을 호출할지 여부를 결정합니다.
이것이 바로 `description` 필드가 중요한 이유입니다. Claude가 판단하는 유일한 근거입니다. 복잡한 알고리즘 라우팅은 없으며, 결정은 전적으로 Claude의 추론 과정에서 이루어집니다.
## 모범 사례
대량의 실전 경험을 통해, 커뮤니티에서 Skills 생성을 위한 네 가지 황금 법칙을 정리했습니다:
**1. 집중 유지**
하나의 Skill은 한 가지만 해야 합니다. 집중된 여러 Skills가 하나의 크고 포괄적인 Skill보다 훨씬 유용하며, 이렇게 하면 유지보수가 쉬울 뿐만 아니라 조합하여 사용하기도 더 쉽습니다.
**2. 명확한 설명**
description 필드는 Claude가 언제 Skill을 호출할지 결정하므로, 적용 시나리오를 반드시 명확하게 작성해야 합니다. "매출 데이터를 기반으로 분기 분석 보고서 생성"은 좋은 설명이고, "데이터 처리"는 너무 광범위합니다.
**3. 예시 제공**
SKILL.md에 입력/출력 예시를 포함하면 출력의 안정성이 크게 향상됩니다. 특히 특정 형식 요구 사항이 있는 작업에서 효과적입니다.
**4. 간단하게 시작**
먼저 순수 Markdown으로 기본 지시를 작성하고, 효과를 검증한 후 스크립트 추가를 고려하며, 점진적으로 복잡도를 높입니다.
### 자주 발생하는 문제 해결
| 문제 | 가능한 원인 | 해결 방법 |
| --------------- | ------------------------ | ---------------------------- |
| Skill이 트리거되지 않음 | description이 충분히 정확하지 않음 | 더 구체적인 사용 시나리오 설명으로 재작성 |
| Skill이 트리거되지 않음 | Skill이 올바르게 설치되지 않음 | 파일 경로와 이름 확인 |
| 출력이 불안정함 | 예시 부족 | 더 많은 입력/출력 예시 추가 |
| 출력이 불안정함 | 지시가 너무 모호함 | 제약 조건과 형식 요구 사항 추가 |
| 로딩이 너무 느림 | 파일이 너무 큼 | 큰 파일을 references 하위 디렉토리로 이동 |
### 보안 주의 사항
Skills는 코드를 실행할 수 있으므로 보안이 매우 중요합니다:
* **신뢰할 수 있는 출처**: 신뢰할 수 있는 채널의 Skills만 사용하십시오
* **스크립트 검토**: 설치 전에 Skills 내의 스크립트 코드를 확인하십시오
* **민감 정보 보호**: Skills에 API 키나 비밀번호를 하드코딩하지 마십시오
* **권한 관리**: 팀에서 사용할 때 Skills의 공유 범위에 주의하십시오
## 현재 제한 사항
새로운 기능으로서 Skills에는 현재 몇 가지 제한 사항이 있습니다:
| 제한 사항 | 설명 |
| -------------------------- | --------------------------------- |
| ~~**Anthropic 생태계에만 한정**~~ | ✅ **해결됨** - 아래 설명 참조 |
| **감사 메커니즘 부재** | 내장된 검토 또는 감사 워크플로우 없음 |
| **학습 곡선** | 팀이 워크플로우를 조정하고 버전 관리 프로세스를 구축해야 함 |
| **초기 단계** | 생태계가 아직 발전 중 |
> **중대 업데이트 (2025년 12월 18일)**: Anthropic은 Agent Skills를 [개방형 표준](https://agentskills.io)으로 공식 발표했습니다. 규격과 참조 SDK가 [agentskills.io](https://agentskills.io)에 공개되었습니다.
>
> **채택한 회사/제품**:
> 
>
> * **Microsoft**: VS Code, GitHub이 통합
> * **OpenAI**: ChatGPT, Codex CLI가 동일한 아키텍처 채택
> * **프로그래밍 도구**: Cursor, Goose, Amp, OpenCode
> * **파트너 Skills**: Atlassian, Figma, Canva, Stripe, Notion, Zapier
>
> 동시에 Anthropic, OpenAI, Block이 공동으로 [Agentic AI Foundation](https://www.linuxfoundation.org/)(Linux Foundation에서 호스팅)을 설립했으며, Google, Microsoft, AWS도 참여했습니다. 이는 Skills가 단일 벤더 기능에서 업계 표준으로 발전하고 있음을 의미하며, Claude Code용으로 작성된 Skills가 OpenAI Codex CLI와 상호 운용이 가능해집니다.
>
> 참고 출처:
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## 학습 리소스
### 공식 리소스
| 리소스 | 링크 | 설명 |
| ------------------- | -------------------------------------------------------------------------------------------------------------------- | ----------------- |
| Skills GitHub 리포지토리 | [anthropics/skills](https://github.com/anthropics/skills) | 공식 예시, 22k+ Stars |
| Claude Code 문서 | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Skills 사용 가이드 |
| 도움말 센터 | [support.claude.com](https://support.claude.com) | 자주 묻는 질문 |
| 기술 블로그 | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | 기술 원리 심층 분석 |
| API 빠른 시작 | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | 개발자 통합 가이드 |
| Agent Skills 개방형 표준 | [agentskills.io](https://agentskills.io) | 공식 규격 및 SDK |
### 커뮤니티 추천
| 리소스 | 링크 | 설명 |
| --------------------- | ------------------------------------------------------------------------------------- | ------------------------- |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Skills 엄선 컬렉션 |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148+ 슬래시 명령어, 54개 AI 에이전트 |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | 사무 문서 생성 및 편집 기술 |
### 추천 읽기
| 문서 | 저자 |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic 공식 |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## 전망
Skills의 등장은 AI 도구 발전의 중요한 방향을 대표합니다. AI가 단순히 작업을 수행하는 것을 넘어, 특정 작업 방식을 학습하고 기억할 수 있게 하는 것입니다. Simon Willison은 Skills가 AI 도구 분야에서 "캄브리아기 대폭발"을 가져올 것이라고 예측했으며, 이 판단은 과장이 아닙니다.
점점 더 많은 개발자와 팀이 Skills를 구축하고 공유하기 시작하면서, 다음과 같은 변화를 볼 수 있을 것입니다:
* **전문화된 Skills 마켓**: 각 산업의 전문가들이 지식을 재사용 가능한 Skills로 패키징
* **Skills와 MCP의 깊은 융합**: 완전한 엔드투엔드 워크플로우 형성
* **엔터프라이즈급 Skills 플랫폼**: 팀 협업, 버전 관리, 권한 제어
지금이 바로 시작하기 좋은 시점입니다. 즉시 할 수 있는 것은 Claude.ai에 로그인하여 문서 기술을 활성화하는 것입니다. 이번 주에는 커뮤니티 Skill을 하나 설치하고 첫 번째 간단한 Skill을 만들어 볼 수 있습니다. 장기적으로는 팀 내의 반복 작업을 파악하고, 점진적으로 전용 기술 라이브러리를 구축하는 것이 효율성을 향상시키는 효과적인 방법이 될 것입니다.
### 추가 읽기
* 《[Claude 시스템 아키텍처 전체 분석](/ko/docs/notes/claude-architecture)》 - Skills의 전체 Claude 시스템에서의 위치
* 《[Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent)》 - Subagent 메커니즘 심층 이해
* 《[나의 Claude Code 모범 사례](/ko/blog/claude-code-best-practices)》 - Claude Code 일상 사용 팁
# Skill-Creator 심층 분석: 데이터를 사용하여 기술 개발을 촉진합니다.
## 소개
이 문서는 2026년 3월 정보를 기반으로 하며 Claude Code v2.1+에 해당합니다.
[개념](/ko/docs/notes/claude-skills/concept)과 [연습](/ko/docs/notes/claude-skills/practice)을 읽었다면 SKILL.md 파일을 수동으로 작성하는 방법을 이미 알고 있어야 합니다. 머리말을 정의하고 명령을 작성하고 `.claude/skills/` 디렉터리에 저장하면 끝입니다.
하지만 여기에 근본적인 질문이 있습니다. \*\*당신의 기술이 정말 유용하다는 것을 어떻게 알 수 있나요? \*\*
단락의 문구를 변경하여 더 잘 작동한다고 느낄 수도 있지만 이는 주관적인 느낌일 뿐입니다. 어쩌면 다른 프롬프트 단어를 사용하면 새 버전이 더 나쁠 수도 있습니다. 어쩌면 당신의 기술은 미숙련에 비해 전혀 향상되지 않을 수도 있습니다. Claude는 혼자서도 잘 할 수 있습니다.
개념장과 실무장에서는 기술 개발 과정을 **작성 → 시도 → 느낌 있음 → 온라인** 순으로 진행합니다. 모든 과정은 직관에 의존하고, 정량화도 없고, "이 기술이 전혀 기술이 없는 것보다 얼마나 더 나은가?"라고 답할 방법이 없습니다. 그리고 Skill-Creator는 이것을 엔지니어링으로 전환했습니다. **작성 → 스킬 유무 병렬 테스트 → 블라인드 테스트 A/B 비교 → 정량 채점 → 피드백 반복 → 데이터 검증**.
이것이 Skill-Creator가 존재하는 이유입니다. "SKILL.md 생성"에 도움이 될 뿐만 아니라 생성 → 테스트 → 평가 → 최적화 루프의 전체 세트를 제공하여 데이터가 스스로 말할 수 있도록 해줍니다.
## 스킬크리에이터란?
Skill-Creator 자체도 스킬입니다. 33KB SKILL.md 파일과 하위 에이전트 지침 파일, Python 스크립트 및 HTML 뷰어를 지원합니다. 디렉토리 구조는 다음과 같습니다:
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
설치도 매우 간단합니다.
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## 이것을 다시 따르십시오: 기존 기술을 평가하고 최적화하십시오.
내가 실제로 사용하는 스킬을 활용하여 Skill-Creator의 전체 과정을 살펴보겠습니다. 저는 Claude Code 플러그인 마켓 [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub)을 운영하고 있습니다. 여기서는 `yux-video-summary` 스킬을 사용하여 비디오 자막을 구조화된 요약으로 변환합니다. 이는 중국어 및 영어 언어 감지, DUAL\_FILE/SINGLE\_FILE 두 가지 출력 모드, 필러 단어 정리 등을 지원합니다. 스킬의 SKILL.md는 다음과 같습니다.
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
스킬은 작성되어 있지만 실제로 유용한지 어떻게 알 수 있나요? \*\* Skill-Creator가 등장하는 곳입니다.
> Skill-Creator 소스 코드에는 중요한 작성 원칙이 있습니다. *"모든 것 뒤에 있는 **이유**를 설명하기 위해 열심히 노력하십시오. ALWAYS 또는 NEVER를 모두 대문자로 쓴다면 이는 노란색 플래그입니다. 추론을 재구성하고 설명하십시오."* 의미: 좋은 스킬은 엄격한 규칙을 쌓기보다는 **이유**를 설명해야 합니다\*\*.
### 1단계: 테스트 케이스 생성 및 평가 실행
핵심 질문: \*\*이 기술이 전혀 기술이 없는 것보다 정말 나은가요? \*\*
Claude Code를 열고 다음을 직접 입력하세요.
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator는 먼저 스킬 정의와 스키마를 읽은 다음 자동으로 테스트 케이스와 정량적 주장을 생성합니다. 내 실행에서는 3개의 테스트 사례와 39개의 어설션이 생성되었습니다.
테스트 케이스를 임의로 컴파일하지 않는다는 점에 유의하십시오. 스킬에 정의된 DUAL\_FILE 및 SINGLE\_FILE의 두 가지 출력 모드를 이해하고 특히 다양한 비디오 유형(자습서, 팟캐스트 인터뷰, 기술 공유) 및 언어 조합을 다루는 시나리오를 설계합니다. Assertions의 디자인도 매우 특별합니다. 언어 감지, 출력 모드 선택, 콘텐츠 품질, 중국어 및 영어 필러 단어 정리에 이르기까지 제가 직접 테스트하려는 것보다 훨씬 더 포괄적입니다.
그런 다음 시스템은 각 테스트 사례에 대해 with\_skill(스킬 로딩) 및 **without\_skill**(기준, 스킬이 로드되지 않음)에 대해 두 개의 독립적인 하위 에이전트를 동시에 시작합니다. **6개의 병렬 에이전트**(3개의 테스트 사례 × 2개의 버전)가 동시에 시작되었으며, 각각은 서로 간섭하지 않고 **독립적인 작업 트리**에서 실행되었습니다.
> Anthropic의 PDF 기술은 이전에 채울 수 없는 양식을 처리하는 데 문제가 있었습니다. Claude는 필드를 정의하지 않고 정확한 좌표에 텍스트를 배치해야 했습니다. 오류 지점은 Eval을 통해 격리되었으며, 이후 팀에서는 위치 지정 로직을 수정했습니다. 이것이 Eval의 가치입니다. "무언가 옳지 않다고 느껴지는 것"을 "여기서 정확히 잘못된 것"으로 바꾸는 것입니다.
### 2단계: 하위 에이전트 3명 릴레이 점수 매기기
모든 작업이 완료되면 세 명의 전문 하위 에이전트가 **자동으로** 순서대로 나타납니다.
**Grader** 어설션을 하나씩 확인합니다. with\_skill 버전의 요약에 개요 테이블이 포함되어 있는지, DUAL\_FILE 모드가 올바르게 선택되었는지, 필러 단어가 정리되었는지 여부를 확인한 다음 각 항목의 통과/실패 및 증거를 기록하여 `grading.json`을 생성합니다.
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**비교기**는 블라인드 A/B 비교를 수행합니다. 두 개의 요약을 수신하지만 **어느 것이 스킬 버전이고 어느 것이 기준 버전인지 알 수 없습니다**. 'Output A'와 'Output B'만 보고 자체 품질 기준에 따라 독립적으로 판단하여 승자를 결정합니다.
**Analyzer**는 위의 결과를 결합하여 진단을 내립니다. 기술 여부에 관계없이 어떤 주장이 통과되었는지(이 주장은 차별화가 없으며 더 나은 주장으로 대체되어야 함을 나타냄), 어떤 결과가 높은 분산을 가지고 있는지(테스트가 불안정함), 시간과 토큰 간의 균형이 무엇인지 진단합니다. 마지막으로 개선을 위한 제안을 제시합니다.
### 3단계: 평가 뷰어에서 결과 검토
채점이 완료되면 Skill-Creator가 자동으로 브라우저에서 HTML 뷰어를 엽니다.
**출력 탭** 각 테스트 사례의 출력을 하나씩 볼 수 있습니다. 하단에 피드백 텍스트 상자가 있습니다. "요약에 타임라인이 부족합니다", "필러 단어가 정리되지 않았습니다" 등 충분하지 않다고 생각되는 내용을 적어보세요. 모든 사용 사례를 읽은 후 **모든 리뷰 제출**을 클릭하면 피드백이 `feedback.json`에 저장됩니다.
**벤치마크 결과 탭** with\_skill과 Without\_skill의 합격률, 시간 소모량, 토큰 소모량, 각 Assertion의 항목별 비교 등 정량적 비교를 확인할 수 있습니다.
### 4단계: 만족할 때까지 반복 및 개선
Claude Code로 돌아가서 피드백 제공을 마쳤다고 알려주세요. Skill-Creator는 `feedback.json`을 읽고 벤치마크 데이터를 기반으로 분석 및 개선 제안을 제공합니다.
제 실력은 97%의 합격률로 좋은 성적을 거두었습니다. Skill-Creator는 작은 문제를 정확하게 식별했습니다. 인터뷰 영상에 주목할만한 인용문 단락이 부족하여 수리를 제안했습니다.
핵심은 개별 테스트 사례를 패치하지 않는다는 것입니다. 피드백을 일반화하고, 그 뒤에 있는 요구 사항을 이해하고, 기술의 전체 구조를 조정한 다음 SKILL.md를 다시 작성하고, 모든 테스트를 `iteration-2/` 디렉터리에 다시 실행하고, 두 라운드의 출력을 비교할 수 있도록 새 Eval Viewer를 엽니다. 이 주기는 귀하가 만족할 때까지 계속됩니다.
> Skill-Creator 소스 코드의 주목할 만한 개선 철학: *"우리는 다양한 프롬프트에서 백만 번 사용할 수 있는 기술을 만들려고 노력하고 있습니다. 까다로운 문제가 있는 경우에는 지나치게 과적합된 변경 사항이나 억압적으로 제한하는 MUST를 적용하는 대신 분기하고 다른 비유를 사용해 보십시오."* 핵심 아이디어: **과적합을 방지**하여 테스트 사례를 피하고 일반화 기능을 추구합니다.
### 5단계(선택 사항): 스킬이 적시에 발동되도록 설명을 최적화합니다.
스킬의 품질은 검증되었지만 간과하기 쉬운 또 다른 문제가 있습니다. 스킬의 `description` 필드에 따라 Claude가 언제 호출할지 결정됩니다.
입력:
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator는 약 20개의 평가 쿼리를 자동으로 생성하고(절반은 트리거되어야 하고 나머지 절반은 트리거되지 않아야 함) 브라우저에서 검토 인터페이스가 열립니다.
이러한 쿼리는 중국어와 영어로 제공되며 다양한 실제 표현을 포괄합니다. "트리거하면 안 됩니다"라는 쿼리는 너무 터무니없어서는 안 됩니다. 좋은 반례로는 "이 회의록을 요약하는 데 도움을 주세요."가 있는데, 이는 비디오 요약과 "요약"이라는 키워드를 공유하지만 실제로는 비디오 요약보다는 문서 처리 기술이 필요합니다.
페이지에서 직접 쿼리 텍스트를 편집할 수 있고, **+ 쿼리 추가**를 클릭하여 새 쿼리를 추가하고, 삭제 버튼을 사용하여 부적절한 텍스트를 삭제하고, 각 쿼리에 대해 트리거해야 함 스위치를 전환할 수도 있습니다. 올바른지 확인한 후 **Export Eval Set**을 클릭하여 JSON 파일을 내보냅니다. Claude Code로 돌아가서 내보냈다고 알려주세요. 시스템은 백그라운드에서 최적화 루프를 자동으로 실행합니다.
전체 프로세스는 완전히 자동화되어 있습니다. 쿼리를 훈련 세트와 테스트 세트 60/40으로 분할하고, 훈련 세트에 대한 설명을 반복적으로 최적화하고(최대 5라운드), 테스트 세트 결과를 사용하여 과적합을 방지하는 최상의 버전을 선택합니다. 실행 후 최적화 전후의 설명 비교가 출력됩니다.
최적화된 설명은 더욱 구체적입니다. 즉, 지원되는 파일 형식(.vtt/.srt)을 명확하게 하고 파이프라인 기능(필러 정리, DUAL/SINGLE\_FILE 논리)을 강조하고 MUST USE를 사용하여 트리거되어서는 안 되는 시나리오를 제외합니다. Anthropic은 내부적으로 이 최적화 도구 세트를 사용하여 자체 문서 생성 기술을 실행했습니다. 그 결과, 퍼블릭 스킬 6개 중 5개 발동 정확도가 향상되었습니다.
### 고급 사용법: 동적 컨텍스트 주입
스킬을 로드할 때 자동으로 컨텍스트를 삽입하려면 Skills 2.0의 `!` 구문을 사용하여 SKILL.md에 셸 명령을 포함할 수 있습니다.
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
이러한 명령은 Claude가 스킬을 보기 전에 실행되며 데이터는 프롬프트에 직접 포함됩니다. Claude가 파일을 하나씩 탐색하도록 하는 것과 비교하면 많은 시간과 토큰이 절약됩니다.
## 두 가지 유형의 스킬: 어떤 것을 만들어야 할까요?
Skill-Creator를 사용하기 전에 Anthropic이 정의한 두 가지 스킬 유형을 이해하는 것이 필요합니다.
**능력향상형** - 이전에는 모델이 할 수 없거나 잘 할 수 없었던 일을 모델이 하게 하세요. 예를 들면:
* 이미지 생성 스킬 : 클로드는 기본적으로 이미지를 생성할 수 없지만, 스킬을 통해 나노배너 등의 도구를 호출하면 달성할 수 있습니다.
* 프론트 엔드 디자인 기술: 기본 AI 디자인은 종종 매우 "AI 취향"이며, 좋은 디자인 기술은 품질을 크게 향상시킬 수 있습니다.
**코딩 기본 설정** - 특정 작업 흐름을 강화하세요. 모델에는 이미 개별 기능이 있지만 정확한 실행 순서가 필요합니다. 예를 들면:
* PR 리뷰 스킬: 정해진 절차에 따라 코드 보안을 점검하고 위험도 보고서 출력
* 영상 요약 스킬 : 특정 템플릿 구조에 따른 출력, 자동 언어 감지 및 필러 단어 정리
이 두 가지 유형의 기술을 테스트해야 하는 이유는 다릅니다. **역량 개선 유형**은 모델이 발전함에 따라 불필요해질 수 있습니다. - 기준(without\_skill)도 모든 어설션을 통과할 수 있다면 모델이 충분히 기본적이며 이 기술이 폐기될 수 있음을 의미합니다. **코딩 유형**은 내구성이 더 좋지만 실제로 작업 흐름에 충실한지 확인해야 합니다.
Skill-Creator의 평가 기능을 사용하면 오래되었을 수 있는 기술을 맹목적으로 사용하는 대신 기술이 여전히 가치가 있는지 지속적으로 확인할 수 있습니다.
## 커뮤니티의 의견
Skill-Creator 업데이트는 X/Twitter에서 Reddit, 독립 블로그에 이르기까지 많은 논의를 촉발시켰으며 실제 피드백은 공식 문서보다 더 가치가 있습니다.
### 정말 유용할까요? 데이터가 말한다
가장 직접적인 질문: 기술을 추가하는 것이 기술을 추가하지 않는 것보다 정말 나은가요? \*\* 몇 가지 실제 측정을 통해 명확한 답을 얻을 수 있습니다.
Reddit u/hashpanak은 타이틀 생성 기술에 대한 평가를 실행했으며 \_skill이 있는 경우 100% 합격률을 보였고 60%가 없는 경우에만 60%의 합격률을 보였습니다. 토큰 비용이 그만한 가치가 있는지 묻는 질문에 그는 "물론입니다. 최적화 후에 반복되는 작업을 스크립트로 변환하여 토큰을 절약할 수 있습니다."라고 대답했습니다. u/spences10은 훨씬 더 극단적입니다. 그는 250개의 샌드박스 평가를 실행하고 스킬 활성화 비율을 84%에서 100%로 높였습니다. u/Manfluencer10kultra의 댓글 섹션에서는 "**이것이 표준 관행이 되어야 합니다.**"라고 말했습니다.
Blogger [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/)이 WordPress 보안 기술을 벤치마킹했습니다. 21개의 주장이 모두 통과되었으며(기준은 90.5%에 불과) 속도가 9.9% 더 빨랐습니다. 그의 요약: **"기술은 예술이었지만 지금은 공학입니다."**
@0zhuxiaofeng은 실제 워크플로우의 관점에서 보다 구체적인 수치를 제공했습니다. "한 달 동안 사용한 후 가장 큰 변화는 run\_eval을 통해 스킬이 스스로 점수를 매길 수 있다는 것입니다. 이제 콘텐츠 작업을 실행하는 에이전트가 각 릴리스 후 자동으로 효과를 평가하고, 불량한 스킬은 직접 제거하고 다시 작성합니다. **수동 개입이 하루 3시간에서 30분으로 단축되었습니다**."
### 간과된 맹점: 트리거 ≠ 품질
Blogger [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/)는 아무도 언급하지 않은 사각지대를 지적했습니다. **스킬은 품질 평가를 통과하지만 트리거 평가에서 실패할 수 있습니다** - 출력 품질은 매우 좋지만 결코 호출되지 않습니다. `run_loop.py` 최적화 3라운드 후에 그는 13/13에 대한 평가를 트리거했습니다. 핵심 통찰력: "기술 설명은 메타데이터가 아니라 학습 가능한 매개변수입니다. 실제 라우팅 동작을 최적화해야 합니다."
이는 @DrWang5257의 제안과 일치합니다. "모든 것을 한 번에 다시 작성하지 마십시오. 먼저 이를 트리거 조건, 입력 템플릿 및 실패 폴백의 세 섹션으로 나누고 단계별로 반복합니다. 이렇게 하면 업데이트 속도가 빠르고 롤오버 비율이 낮습니다."
### 실제 문제점
효과는 좋지만 함정도 많습니다.
* **토큰 소비량이 엄청납니다**. @konghao10은 "토큰 소비가 엄청나다"고 솔직하게 말했습니다. 동시에 6개의 병렬 에이전트를 실행하는 것은 실제로 저렴하지 않습니다. Reddit u/munkymead도 "심각한 테스트를 받는 데 비용이 많이 든다"고 말했습니다.
* **스킬이 너무 많으면 싸우게 됩니다**. [RoboRhythms 블로거 Noah Albert](https://www.roborhythms.com/best-claude-code-skills-2026/)는 **기술이 8-10에 도달하면 문제가 발생하기 시작합니다**: Claude가 출력에 대해 스스로 질문하고 더 자세한 서문을 생성하며 때때로 기술 간에 명령 충돌이 발생한다는 사실을 발견했습니다. 그러나 Reddit u/Specialist\_Solid523은 다음과 같이 반박했습니다. "잘 작성되지 않은 기술은 컨텍스트만 잠식합니다. **잘 작성된 기술은 거의 항상 토큰 사용을 더 효율적으로 만듭니다.**"
* **SKILL.md는 반복 횟수가 많아질수록 길어집니다**. Reddit u/IulianHI는 반복적인 개선을 통해 스킬 파일이 계속 확장되지만 \*\* 실제로 작업을 수행하기 위한 컨텍스트 창을 밀어낸다\*\*는 모순을 지적했습니다. Happy Path만 다루는 테스트 케이스는 Critical 5%를 놓치게 됩니다.
* **버전 관리가 누락되었습니다**. @fengqve 님은 "스킬은 왜 **버전 개념이 없나요**? 너무 여러번 업데이트가 되어서 어떤 업데이트인지 설명하기 어렵습니다. "라고 불평합니다. 이는 여러 라운드의 반복 후에 특히 고통스럽습니다.
* **헤드리스 모드에는 버그가 있습니다**. GitHub에는 주요 문제가 있습니다. 기술이 `claude -p` 모드에서 트리거되지 않아 최적화 루프를 설명하는 회상이 항상 0%가 됩니다([#36570](https://github.com/anthropics/claude-code/issues/36570)).
### 더 깊이 생각하기: 재귀적 자기 개선
@vista8은 관련 논문 [Memento-Skills: Let Agents Design Agents](https://github.com/Memento-Teams/Memento-Skills)를 공유했는데 댓글 영역의 누군가가 이를 정확하게 요약했습니다. "Skill의 핵심 병목 현상은 반복입니다. 첫 번째 버전을 작성하기는 쉽지만 실제 시나리오에서 더 좋게 사용하기는 어렵습니다. 이 '사용 → 평가 → 개선' 주기를 자동화할 수 있다면 Agent에 자체 진화 엔진을 설치하는 것과 같습니다."
Reddit r/ClaudeAI의 104와 유사한 스레드에서도 이 방향에 대해 논의합니다. 그러나 최고 댓글은 이에 대해 찬물을 끼얹었습니다. u/Tatrions는 다음과 같이 말했습니다. "재귀 루프는 작동하지만 어려운 부분은 언제 개선 사항을 신뢰할지 아는 것입니다. 우리는 증거 게이팅을 수행해야 한다는 것을 발견했습니다. 실패가 적어도 두 번 발생하지 않는 한 변경 사항을 커밋하지 마십시오. 그렇지 않으면 각 주기는 처음에 깨지지 않은 것을 '수정'하고 결국 더 악화됩니다."
## 설치 및 생태학
Skill-Creator는 Anthropic이 공식적으로 유지 관리하는 스킬 중 하나로 [anthropics/skills](https://github.com/anthropics/skills) 창고에 포함되어 있으며, 17개 이상의 생산 수준 스킬이 포함되어 있습니다.
더 넓은 Skills 생태계도 빠르게 성장하고 있습니다. [skills.sh](https://skills.sh) 시장은 편리한 검색 및 설치 경험을 제공하며 커뮤니티는 1,234개 이상의 에이전트 기술을 유지해 왔습니다.
## 마지막에 쓰세요
Skill-Creator가 해결하는 핵심 문제는 다음과 같습니다. \*\*당신의 기술이 실제로 효과적인지 어떻게 알 수 있습니까? \*\*
그것이 없으면 기술 개발은 "쓰기 → 노력 → 괜찮은 느낌"에 의존합니다. Skill-Creator를 사용하면 다음을 수행할 수 있습니다.
* **병렬 에이전트**를 사용하여 숙련된 효과와 비숙련된 효과를 모두 테스트합니다.
* **블라인드 A/B 비교**를 통해 평가 편향 제거
* **Eval Viewer**를 통해 결과를 시각화하고 피드백을 남깁니다.
* **설명 최적화**를 사용하여 스킬 발동 타이밍을 정확하게 제어
* 만족할 때까지 **반복 루프**를 사용하여 지속적으로 개선하세요.
이는 소프트웨어 엔지니어링의 테스트 중심 개발 개념과 일치합니다. 즉, "단순히 코드를 작성하고 실행할 수 있다고 생각하는 것"이 아니라 "테스트를 사용하여 실제로 예상대로 작동하는지 입증하는 것"입니다.
Anthropic은 공식 블로그에서 흥미로운 전망을 제시했습니다. 모델의 기능이 향상됨에 따라 SKILL.md는 "구현 계획"(Claude **방법** 알려주기)에서 '사양 설명'(Claude **무엇**을 알려주고 모델이 스스로 알아내도록 허용)으로 발전할 수 있습니다. Eval 프레임워크는 이 방향의 첫 번째 단계입니다. Eval은 "무엇을 해야 할지"를 설명합니다. 언젠가 이 기술 자체가 스킬이 되기에 충분하다면 Skill-Creator가 구축한 테스트 시스템은 더욱 중요해질 것입니다.
이미 기술을 사용하고 있다면 `/skill-creator`을(를) 사용하여 가장 많이 사용하는 기술을 평가해 보세요. 일부 기술은 실제로 전혀 기술이 없는 것보다 낫지 않다는 사실에 놀랄 수도 있습니다. 바로 여기에서 최적화가 시작됩니다.
관련 자료:
* [클로드 스킬이란](/ko/docs/notes/claude-skills/concept) — 스킬의 핵심 원리를 이해합니다.
* [연습 가이드](/ko/docs/notes/claude-skills/practice) — 첫 번째 스킬 만들기
# 개념 소개
## 서론
[Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept)에서 우리는 핵심 문제 하나를 알아보았습니다: **Context Rot** — 대화가 길어질수록 Claude의 컨텍스트 윈도우가 실패한 코드, 오래된 논의, 관련 없는 정보로 가득 차면서 출력 품질이 지속적으로 하락하는 현상입니다.
Ralph의 해결책은 "모든 것을 재시작"하는 것이었습니다: bash 무한 루프로 매번 완전히 새로운 Claude 인스턴스를 실행하고, 파일 시스템을 통해 상태를 전달합니다. 단순하고 효과적이지만, 명확한 한계도 있습니다 — 이것은 단지 하나의 방법론일 뿐, 프로젝트 이해도 없고, 단계별 계획도 없으며, 품질 검증도 없습니다. 스펙을 직접 작성하고, 작업을 직접 편성하고, "완료되었는지"를 직접 판단해야 합니다.
Chase AI가 영상에서 정확하게 요약한 것처럼: **Ralph Loop은 극히 강력한 무기이지만, 대부분의 사람들에게 필요한 것은 하나의 무기가 아니라 전체 무기고입니다.** Ralph 루프는 전적으로 사전 준비에 의존합니다: PRD가 충분히 좋은가? 기능 정의가 충분히 긴밀한가? "완료"가 어떤 모습인지 알고 있는가? 이 질문들에 대한 답이 정확하지 않다면, 루프를 아무리 많이 돌려도 garbage in, garbage out일 뿐입니다.
만약 **단순히 Claude를 반복 실행하는 것이 아니라, 프로젝트를 진정으로 이해하고 안정적으로 코드를 전달하는** 시스템을 원한다면 어떨까요?
이것이 바로 \*\*GSD (Get Shit Done)\*\*가 하는 일입니다.
## GSD란 무엇인가
GSD의 창시자는 **TÂCHES** (GitHub: glittercowboy)이며, 독립 개발자입니다. 그의 동기는 매우 직접적입니다:
> "나는 50명 규모의 소프트웨어 회사가 아닙니다. 기업 드라마를 하고 싶지 않습니다. 나는 그저 좋은 것을 만들고 싶은 크리에이티브한 사람일 뿐입니다."
그의 라이브 스트리밍에서 TÂCHES는 충격적인 사실을 보여주었습니다: 그는 **코드를 직접 작성하지 않습니다**. 그는 GSD를 사용하여 4시간 만에 완전한 macOS 네이티브 음악 생성 앱(Sample Digger)을 처음부터 구축했으며, 전 과정에서 코드를 직접 작성하지 않았습니다. 그의 자기 포지셔닝은 프로그래머가 아니라 "고위 프로젝트 매니저"입니다 — 비전을 설명하고, 핵심 결정을 내리고, 결과를 검증합니다. GSD는 이러한 작업 방식을 가능하게 합니다.
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> —— TÂCHES
다른 스펙 주도 개발 도구들 — BMAD, SpecKit — 은 각각의 가치가 있지만, 복잡한 기업 워크플로우를 도입하는 경향이 있습니다: 스프린트 세레모니, 스토리 포인트, 이해관계자 싱크. 독립 개발자나 소규모 팀에게 이러한 프로세스 자체가 부담입니다. Chase AI가 평가한 것처럼: "It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way."
GSD의 설계 철학은 **복잡성을 시스템 안에 숨기는 것**입니다. 사용자는 몇 가지 간단한 명령만 필요하고, 시스템이 배후에서 모든 컨텍스트 관리, 작업 편성, 품질 검증을 처리합니다. 프로젝트 출시 한 달 만에 약 3,000개의 GitHub 스타와 14,000건의 npm 설치를 달성했으며, TÂCHES는 거의 매일 15-20회 업데이트합니다.
### GSD의 도구 생태계 내 위치
| 차원 | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| -------------- | ----------------- | -------------------- | ----------------- | ---------------------- |
| 핵심 포지셔닝 | 실행 기술 (bash loop) | 스펙 생성 도구 키트 | 기업급 프레임워크 | **컨텍스트 엔지니어링 + 스펙 주도** |
| 계획 능력 | 없음 (스펙 자체 준비 필요) | 강함 (spec→plan→tasks) | 강함 (완전한 애자일 프로세스) | **강함 (연구→논의→계획)** |
| 실행 자율성 | 최고 (AFK 모드) | 각 단계 수동 트리거 필요 | 각 단계 수동 트리거 필요 | **각 단계 수동 트리거 필요** |
| 인간 참여 모드 | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Context Rot 처리 | 새 session 재시작 | 내장 방안 없음 | 내장 방안 없음 | **서브에이전트 신선한 컨텍스트** |
| 품질 검증 | 외부 테스트 의존 | 빌드 체크 | 내장 QA 프로세스 | **자동 검증 + UAT** |
| 사용자 복잡도 | 최저 | 중간 | 높음 | **낮음** |
| 시스템 복잡도 | 최저 | 중간 | 높음 | **높음** |
이 표는 핵심적인 트레이드오프를 보여줍니다: **Ralph는 최저의 시스템 복잡도로 최고의 실행 자율성을 교환했습니다** — 시작한 후 잠자러 갈 수 있습니다; 반면 **GSD는 높은 시스템 복잡도로 계획 품질과 인간 검증을 교환했습니다** — 각 단계마다 개입할 기회가 있습니다. SpecKit과 BMAD는 중간 지대에 위치하며, 계획 능력을 제공하지만 GSD의 컨텍스트 엔지니어링과 Ralph의 자율 실행이 부족합니다.
GSD와 Ralph는 모순되지 않습니다. GSD는 Ralph의 핵심 원칙 — 신선한 컨텍스트, 파일을 진실의 원천으로 사용 — 을 계승하되, 그 위에 완전한 프로젝트 이해 및 실행 체계를 구축했습니다. Ralph가 "AI에게 작업을 주고 반복 시도하게 하는 것"이라면, GSD는 "당신이 무엇을 원하는지 이해하고, 어떻게 할지 연구하고, 어떻게 단계별로 나눌지 계획하고, 실행하고 검증하는 것"입니다.
Chase AI의 요약이 매우 정확합니다: **Ralph 루프는 당신이 완전한 청사진을 가지고 온다고 가정합니다 — GSD는 그 청사진을 구축하는 것을 도와줍니다.** GSD는 당신의 반쯤 형성된 아이디어를 받아, 깊이 있게 질문하고, 대신 조사하고, 완전한 PRD를 생성하고, 이를 원자적 작업으로 분해한 다음, 프로젝트를 처음부터 끝까지 전달합니다. 그리고 코드를 실행할 때는 Ralph 루프를 강력하게 만드는 바로 그 기본 원칙을 사용합니다: 서브에이전트의 신선한 컨텍스트와 가능한 한 작고 정확한 작업.
## 핵심 워크플로우
GSD의 워크플로우는 **논의 → 계획 → 실행 → 검증**의 순환이며, 각 단계마다 명확한 입력과 출력이 있습니다.
### 1. 프로젝트 초기화
```text
/gsd:new-project
```
하나의 명령으로 전체 프로세스가 시작됩니다. 시스템은 다음을 수행합니다:
1. **질문** — 당신의 아이디어를 완전히 이해할 때까지 계속 질문합니다 (목표, 제약, 기술 선호, 엣지 케이스)
2. **조사** — 병렬 에이전트를 파견하여 관련 분야를 조사합니다 (선택 사항이지만 권장)
3. **요구사항 추출** — v1, v2, 범위 밖의 내용을 구분합니다
4. **로드맵** — 요구사항에 대응하는 단계별 계획을 생성합니다
로드맵을 승인하면 구축이 시작됩니다. TÂCHES의 경험에 따르면: 초기 설명이 상세할수록 시스템의 추가 질문이 줄어들고, 모호할수록 질문이 늘어납니다. 그는 시작 전에 대략적인 비전 문서를 준비할 것을 권장합니다 — 기술 스택이나 구현 세부사항을 알 필요는 없고, 원하는 것을 설명하기만 하면 됩니다.
**산출물 파일**: `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md`
> 기존 코드베이스가 있나요? 먼저 `/gsd:map-codebase`를 실행하면, 시스템이 병렬 에이전트를 파견하여 기술 스택, 아키텍처, 관례, 잠재적 문제를 분석합니다. 이후 `/gsd:new-project`가 기존 코드베이스를 기반으로 계획을 수립할 수 있습니다.
### 2. 논의 단계
```text
/gsd:discuss-phase 1
```
로드맵의 각 단계에는 한두 줄의 설명만 있으며, 이것만으로는 원하는 것을 구축하기에 충분하지 않습니다. 논의 단계의 역할은 조사와 계획 전에 **구현 선호도를 포착**하는 것입니다.
시스템은 현재 단계를 분석하고, "회색 지대" — 여러 가지 합리적인 구현 방식이 있는 의사결정 포인트 — 를 식별합니다:
* 시각적 기능 → 레이아웃, 인터랙션, 빈 상태 처리
* API/CLI → 응답 형식, 오류 처리, 상세 수준
* 콘텐츠 시스템 → 구조, 톤, 깊이, 플로우
여기서 내리는 모든 결정이 이후의 조사와 계획 품질에 직접 영향을 미칩니다. 이 단계를 건너뛸 수도 있지만 (시스템이 합리적인 기본값을 사용합니다), 심도 있는 논의를 하면 시스템이 기대에 더 부합하는 결과물을 구축할 수 있습니다.
**산출물 파일**: `{phase}-CONTEXT.md`
### 3. 계획 단계
```text
/gsd:plan-phase 1
```
시스템은 다음을 수행합니다:
1. **조사** — 논의 단계의 결정을 지침으로 삼아 현재 단계를 어떻게 구현할지 조사합니다
2. **계획** — 2-3개의 원자적 작업 계획을 생성하며, XML 구조화 형식을 사용합니다
3. **검증** — 계획이 요구사항을 충족하는지 확인하고, 통과할 때까지 반복 수정합니다
중요한 설계 이념은 **Goal-Backward Planning** (목표 역추적 계획)입니다. "무엇을 구축해야 하는가"에서 출발하는 것이 아니라, "목표를 달성하기 위해 어떤 조건이 성립해야 하는가?"를 묻고 — 거기서 역으로 계획과 작업을 도출합니다. TÂCHES는 이 방식이 "출력 품질을 극적으로 향상시켰다"고 말합니다. 각 작업이 자신과 다른 작업의 관계를 이해하며, 단순한 할 일 목록이 아니기 때문입니다.
각 계획은 충분히 작아서 완전히 새로운 컨텍스트 윈도우에서 실행할 수 있습니다. 이것이 핵심입니다 — **품질 저하가 없습니다**.
**산출물 파일**: `{phase}-RESEARCH.md`, `{phase}-{N}-PLAN.md`
### 4. 실행 단계
```text
/gsd:execute-phase 1
```
시스템은 다음을 수행합니다:
1. **웨이브 실행** — 독립적인 작업은 병렬로 실행하고, 의존성이 있는 것은 순차적으로 실행합니다
2. **신선한 컨텍스트** — 각 계획은 완전히 새로운 200k 토큰 컨텍스트에서 실행되며, 누적된 쓰레기가 전혀 없습니다
3. **원자적 커밋** — 각 작업은 독립적으로 git commit합니다
4. **목표 검증** — 코드베이스가 단계에서 약속한 기능을 구현했는지 확인합니다
TÂCHES의 라이브 데모에서 그는 3개의 완전한 단계 개발을 완료했으며, **메인 컨텍스트 윈도우는 시종일관 24%를 유지했습니다**. GSD Executor 서브에이전트는 1,000줄 미만의 컨텍스트만 로드하면 하나의 완전한 단계를 완수할 수 있습니다 — 10개의 계획을 연속으로 실행해도 컨텍스트는 여전히 50% 미만입니다. 이것은 Claude Code에서 직접 작업하는 경험과 완전히 다릅니다: 더 이상 "러시안 룰렛을 하면서, 언제 컨텍스트 윈도우의 벽에 부딪힐지 도박하는 것"이 아닙니다.
**산출물 파일**: `{phase}-{N}-SUMMARY.md`, `{phase}-VERIFICATION.md`
### 5. 검증 단계
```text
/gsd:verify-work 1
```
자동화된 검증은 코드가 존재하는지, 테스트가 통과하는지 확인할 수 있습니다. 하지만 기능이 **기대한 대로 작동하는지**는 당신이 직접 확인해야 합니다.
시스템은 다음을 수행합니다:
1. **테스트 가능한 산출물 추출** — 지금 할 수 있어야 하는 일들의 목록을 나열합니다
2. **하나씩 검증 안내** — "이메일로 로그인할 수 있나요?" 예/아니오, 또는 문제 설명
3. **실패 자동 진단** — 디버그 에이전트를 파견하여 근본 원인을 찾습니다
4. **수정 계획 생성** — 바로 실행할 수 있는 수정 방안을 만듭니다
모두 통과하면 다음 단계로 진행합니다. 문제가 있으면 `/gsd:execute-phase`를 다시 실행하여 수정 계획을 실행합니다.
이것이 GSD와 Ralph 루프의 가장 큰 이념적 차이입니다: **Ralph는 hands-off입니다 — 시작한 후 놓아두고 실행하게 합니다; GSD는 각 단계가 끝난 후 인간 검증 단계가 있습니다.** Chase AI는 Ralph 루프가 "정복하러 가는 타입"이라고 지적했습니다 — 스스로 작동하며 뒤돌아보지 않습니다; GSD는 각 핵심 노드에서 개입하여 방향을 바로잡을 수 있도록 보장하며, 오류가 무인 감독 하에서 겹겹이 쌓이는 것을 방지합니다.
또한 GSD는 전용 디버그 프로세스도 제공합니다. 검증에서 문제가 발견되면, `/gsd:debug`가 **격리된 디버그 서브에이전트**를 시작합니다. 이 에이전트는 자체적인 가설-증거-해결 워크플로우를 가지고 있으며, 독립적인 디버그 문서를 생성하여 전체 조사 과정을 추적하고, 메인 컨텍스트를 오염시키지 않습니다.
**산출물 파일**: `{phase}-UAT.md`
### 순환 반복
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
각 단계는 완전한 **논의 → 계획 → 실행 → 검증** 순환을 거칩니다. 컨텍스트는 신선하게 유지되고, 품질은 일관되게 유지됩니다.
모든 단계가 완료되면 `/gsd:complete-milestone`이 마일스톤을 아카이브하고 버전을 표시합니다. 그런 다음 `/gsd:new-milestone`이 다음 버전의 구축을 시작합니다.
## 왜 효과적인가: 기술 원리
GSD의 안정성은 우연이 아닙니다. 그 뒤에는 네 가지 핵심 기술적 지원이 있습니다.
### Context Engineering
Claude Code는 올바른 컨텍스트를 제공받을 때 매우 강력합니다. 대부분의 사람들은 올바른 컨텍스트를 어떻게 제공해야 하는지 모릅니다. GSD가 이 문제를 대신 처리합니다.
| 파일 | 역할 |
| ----------------- | ---------------------------- |
| `PROJECT.md` | 프로젝트 비전, 항상 로드됨 |
| `research/` | 생태계 지식 (기술 스택, 기능, 아키텍처, 함정) |
| `REQUIREMENTS.md` | 버전별 요구사항, 단계 추적 포함 |
| `ROADMAP.md` | 방향과 진행 상황 |
| `STATE.md` | 결정, 장애물, 위치 — 세션 간 기억 |
| `PLAN.md` | 원자적 작업 + XML 구조 + 검증 단계 |
| `SUMMARY.md` | 실행 기록, 히스토리에 커밋 |
각 파일에는 Claude의 품질 저하 임계값에 기반한 **크기 제한**이 있습니다. 제한 이내로 유지하면 일관된 고품질 출력을 얻을 수 있습니다. 메인 컨텍스트 윈도우는 30-40%로 유지되며, 실제 작업은 서브에이전트의 완전히 새로운 200k 컨텍스트에서 수행됩니다.
Chase AI는 Context Rot에 대해 직관적인 설명을 했습니다: **컨텍스트 윈도우가 아무리 크더라도 — Sonnet, Opus, 심지어 백만 토큰 윈도우라도 — 앞부분의 토큰이 뒷부분보다 더 효과적입니다.** 이것은 버그가 아니라 LLM의 고유한 특성입니다. Claude Code에 내장된 autocompact은 부분적으로만 완화할 수 있습니다. GSD의 방안은 더 철저합니다: 각 원자적 작업이 완전히 새로운 서브에이전트에서 실행되어, 모든 작업이 Claude의 최고 성능을 발휘할 수 있도록 보장합니다.
TÂCHES 자신의 데이터가 이를 뒷받침합니다: 그는 월 $200의 Max 플랜에서 매월 약 $30,000의 Opus 토큰을 소비합니다. 이것은 많아 보이지만, 각 작업이 신선한 컨텍스트에서 실행되기 때문에 재작업이 극히 드물고, 실제 효율은 저하된 컨텍스트에서 반복적으로 수정하는 것보다 훨씬 높습니다.
### XML Prompt Formatting
각 계획은 Claude에 최적화된 구조화 XML입니다:
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
정확한 지시, 추측 불필요, 검증이 각 작업에 내장되어 있습니다.
### Multi-Agent Orchestration
각 단계는 동일한 패턴을 사용합니다: 가벼운 오케스트레이터가 전문화된 에이전트를 파견하고, 결과를 수집하며, 다음 단계로 라우팅합니다.
| 단계 | 오케스트레이터의 역할 | 에이전트의 역할 |
| -- | ---------------- | ------------------------------------ |
| 조사 | 조정, 발견 사항 제시 | 4개의 병렬 연구원이 기술 스택, 기능, 아키텍처, 함정 조사 |
| 계획 | 검증, 반복 관리 | 플래너가 계획 생성, 체커가 검증, 통과할 때까지 반복 |
| 실행 | 웨이브 그룹화, 진행 추적 | 실행자가 병렬로 구현, 각각 완전히 새로운 200k 컨텍스트 보유 |
| 검증 | 결과 제시, 다음 단계 라우팅 | 검증자가 코드베이스 확인, 디버거가 실패 진단 |
오케스트레이터는 절대로 무거운 작업을 하지 않습니다. 에이전트를 파견하고, 기다리고, 결과를 통합합니다. 결과적으로: 전체 단계를 실행할 수 있습니다 — 심층 조사, 여러 계획 생성 및 검증, 수천 줄의 코드 병렬 작성, 자동 검증 — **그러면서도 메인 컨텍스트 윈도우는 30-40%를 유지합니다**.
### Atomic Git Commits
각 작업이 완료되면 즉시 독립적으로 커밋합니다:
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
장점: `git bisect`로 특정 실패 작업을 정확히 찾을 수 있고, 각 작업을 독립적으로 롤백할 수 있으며, 명확한 히스토리 기록이 향후 세션에서 Claude가 코드 진화를 이해하는 데 도움이 됩니다.
## GSD의 한계
GSD는 강력하지만, **할 수 없는 것**을 이해하는 것도 똑같이 중요합니다.
### GSD는 인간이 가이드하는 워크플로우이며, 자율 에이전트가 아닙니다
GSD는 지속적으로 실행될 수 없습니다. 각 단계 경계 — `discuss`에서 `plan`으로, `execute`로, `verify`로 — 에서 수동으로 명령을 입력해야 합니다. "앱 하나 만들어줘"라고 말하고 잠자러 갈 수 없습니다.
이것은 Ralph의 AFK 모드와 뚜렷한 대조를 이룹니다. Ralph는 "시작한 후 잠자러 가도록" 설계되었습니다 — bash 무한 루프가 작업이 완료되거나 실패할 때까지 계속 실행됩니다. GSD는 각 핵심 노드에서 당신이 현장에 있을 것을 요구합니다: 로드맵 승인, 논의 질문 답변, 계획 트리거, 실행 시작, 검증 결과 확인.
TÂCHES는 라이브 스트리밍에서 4시간 동안 계속 명령을 입력했습니다: `new-project`, `discuss-phase 1`, `plan-phase 1`, `execute-phase 1`, `verify-work 1`, `discuss-phase 2`...... 매번 전환할 때마다 엔터 키를 눌러야 했습니다. 이것은 우연이 아닙니다 — 의도적인 설계 선택입니다.
### 의도적인 설계 트레이드오프
Ralph는 계획 능력을 희생하여 실행 자율성을 얻었고, GSD는 실행 자율성을 희생하여 계획 품질과 인간 검증을 얻었습니다. **이것은 설계 트레이드오프이지 결함이 아닙니다.**
* **Ralph의 장점**: 잠자는 동안 전체 기능을 완성하게 할 수 있습니다. 하지만 스펙이 충분히 좋지 않으면, 잘못된 방향으로 전속력으로 질주합니다.
* **GSD의 장점**: 각 단계가 끝난 후 방향을 교정할 수 있습니다. 하지만 전 과정에 현장에 있어야 하며, 자리를 비울 수 없습니다.
이상적인 상태는 무엇일까요? GSD의 논의, 계획, 실행, 검증이 자동 순환으로 연결된다면 — Ralph의 bash loop과 비슷하지만, GSD의 구조화된 계획과 품질 검증을 갖춘 — 그것이 두 세계의 최상의 조합이 될 것입니다. 하지만 현재 그런 도구는 아직 없습니다. 이것이 아마도 다음으로 탐구할 가치가 있는 방향일 것입니다.
## 영상 자료
아래 영상은 GSD의 사용 방법과 효과를 보다 직관적으로 이해하는 데 도움이 됩니다.
## 마치며
GSD는 AI 프로그래밍 도구 진화의 한 방향을 대표합니다: "AI가 코드를 작성하게 하는 것"에서 "AI가 안정적으로 프로젝트를 전달하게 하는 것"으로.
Ralph Wiggum은 핵심적인 통찰을 증명했습니다 — 신선한 컨텍스트가 누적된 컨텍스트보다 더 가치 있다는 것. GSD는 이 기반 위에 프로젝트 이해(new-project), 의사결정 포착(discuss), 구조화된 계획(plan), 병렬 실행(execute), 품질 검증(verify)을 추가하여 완전한 폐쇄 루프를 형성했습니다.
독립 개발자와 소규모 팀에게 GSD의 가치는 복잡한 엔지니어링 실천을 몇 가지 간단한 명령으로 캡슐화한 것에 있습니다. 서브에이전트 편성이나 XML 프롬프트 엔지니어링을 이해할 필요가 없습니다 — 원하는 것을 설명하고, 시스템이 처리하도록 하면 됩니다.
Chase AI의 말이 적절합니다: GSD는 "기술 배경이 아닌 사람이지만, 여전히 Claude Code에서 지속 가능하고 반복 가능한 방식으로 프로젝트를 처음부터 끝까지 구축하고 싶은" 사람에게 적합합니다. 그리고 TÂCHES의 라이브 스트리밍이 이를 증명했습니다 — "아마 Hello World HTML 페이지 정도밖에 직접 작성할 수 없을 것"이라고 자칭하는 사람이 GSD로 완전한 네이티브 데스크톱 애플리케이션을 구축했습니다.
이것은 마법이 아닙니다. **올바른 복잡성을 올바른 곳에 배치하는 것**입니다 — 시스템이 편성의 복잡성을 담당하고, 인간은 창의성과 의사결정에 집중합니다. 그리고 그 한계도 존중할 가치가 있습니다: GSD는 인간이 항상 현장에 있는 것을 선택했으며, 이것은 제약이면서 동시에 안정성의 원천입니다.
실습을 시작하고 싶으신가요? [GSD 실전 가이드](/ko/docs/notes/gsd/practice)를 계속 읽어보세요 — 완전한 명령 참조, 구성 상세 설명, 실전 워크플로우 데모와 자주 묻는 질문을 다룹니다.
***
**관련 읽을거리**:
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) — Context Rot 문제와 Ralph 방법론의 완전한 해설
* [스펙 주도 개발이란 무엇인가](/ko/docs/notes/speckit/concept) — Vibe Coding에서 스펙 주도 개발로의 패러다임 전환
* [Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent) — 컨텍스트를 깨끗하게 유지하는 또 다른 방법
* [Claude 시스템 아키텍처 전체 해설](/ko/docs/notes/claude-architecture) — Hooks, Subagent 등 컴포넌트의 전체 아키텍처
* [나의 Claude Code 모범 사례](/ko/blog/claude-code-best-practices) — Claude Code 일상 사용 팁
# 실전 가이드
## 서론
[이전 글](/ko/docs/notes/gsd/concept)에서 GSD의 핵심 원리인 컨텍스트 엔지니어링, 서브 에이전트 오케스트레이션, 목표 역추적 계획, 원자적 커밋에 대해 깊이 알아보았습니다. 이런 개념들은 우아하게 들리지만, "원리를 이해하는 것"에서 "실제로 프로젝트를 완성하는 것" 사이에는 적지 않은 운영 세부 사항이 있습니다.
이번 글에서는 직접 실습해 보겠습니다. GSD의 완전한 명령어 체계, 설정 옵션, 산출물 파일 구조, 그리고 처음부터 완전한 기능을 납품하는 방법을 배우게 됩니다.
## 설치 및 설정
### 설치
```bash
npx get-shit-done-cc@latest
```
설치 프로그램이 다음 사항을 선택하도록 안내합니다:
1. **런타임** — Claude Code, OpenCode, Gemini CLI 또는 전부
2. **위치** — 전역(모든 프로젝트) 또는 로컬(현재 프로젝트)
설치 후 런타임에서 `/gsd:help`를 입력하여 설치가 성공했는지 확인합니다.
### 권장: 권한 건너뛰기 모드
GSD는 무마찰 자동화를 위해 설계되었습니다. 다음과 같은 방법으로 Claude Code를 실행하는 것을 권장합니다:
```bash
claude --dangerously-skip-permissions
```
이 플래그를 사용하고 싶지 않다면 `.claude/settings.json`에서 세밀한 권한을 설정할 수 있습니다.
### 업데이트
```text
/gsd:update
```
GSD는 매우 빈번하게 업데이트됩니다(TÂCHES는 거의 매일 15-20회 업데이트를 푸시합니다). 정기적으로 이 명령어를 실행하여 최신 버전을 유지하는 것을 권장합니다.
## 완전 명령어 참조
GSD의 모든 상호작용은 `/gsd:` 접두사가 붙은 슬래시 명령어로 이루어집니다. 아래에 기능별로 분류하여 전체 명령어를 나열합니다.
### 핵심 워크플로우 명령어
이 다섯 가지 명령어가 GSD의 메인 루프를 구성하며, 순서대로 사용합니다.
| 명령어 | 설명 |
| ------------------------ | ---------------------------------------------------------------------------- |
| `/gsd:new-project` | 프로젝트를 초기화합니다. 시스템이 여러분의 아이디어를 이해할 때까지 계속 질문한 후, 조사하고, 요구사항을 추출하고, 로드맵을 생성합니다 |
| `/gsd:discuss-phase [N]` | N번째 단계의 모호한 부분을 논의합니다. 구현 선호도를 파악하여 계획 수립 방향을 제시합니다 |
| `/gsd:plan-phase [N]` | N번째 단계의 원자적 작업 계획을 생성합니다. 조사, 계획, 검증의 세 가지 하위 단계를 포함합니다 |
| `/gsd:execute-phase ` | N번째 단계를 실행합니다. 서브 에이전트가 작업을 병렬로 구현하며, 각 작업은 독립적으로 커밋됩니다 |
| `/gsd:verify-work [N]` | N번째 단계의 산출물을 검증합니다. 하나씩 확인하도록 안내하며, 문제를 자동으로 진단합니다 |
> `[N]`은 선택적 매개변수를 나타냅니다. 생략하면 시스템이 자동으로 현재 단계를 감지합니다. ``은 필수 매개변수를 나타냅니다.
### 마일스톤 관리
| 명령어 | 설명 |
| --------------------------- | ------------------------------------------------------------- |
| `/gsd:audit-milestone` | 현재 마일스톤 진행 상황을 감사합니다. 모든 단계의 상태를 확인하고 미완료 항목을 식별합니다 |
| `/gsd:complete-milestone` | 현재 마일스톤을 아카이브하고, 버전을 표시하며, 다음 주기로 진입할 준비를 합니다 |
| `/gsd:new-milestone [name]` | 새 마일스톤을 생성합니다. 선택적으로 이름을 제공하면, 시스템이 완료된 작업을 기반으로 다음 단계를 계획합니다 |
### 단계 관리
| 명령어 | 설명 |
| --------------------------------- | ------------------------------------------- |
| `/gsd:add-phase` | 로드맵 끝에 새 단계를 추가합니다 |
| `/gsd:insert-phase [N]` | 지정한 위치에 긴급 단계를 삽입하며, 이후 단계는 자동으로 번호가 재할당됩니다 |
| `/gsd:remove-phase [N]` | 지정한 단계를 제거하고, 관련된 모든 산출물 파일을 연쇄적으로 삭제합니다 |
| `/gsd:list-phase-assumptions [N]` | 지정한 단계의 모든 가정과 의존성을 나열하여 잠재적 위험을 식별합니다 |
### Quick Mode 및 도구
| 명령어 | 설명 |
| ---------------------- | ------------------------------------------------------------------------ |
| `/gsd:quick [--full]` | 빠른 모드 -- 조사, 계획 검증, 검증을 건너뛰며 작은 작업에 적합합니다. `--full`은 완전한 보장을 활성화합니다 |
| `/gsd:debug [desc]` | 격리된 디버그 서브 에이전트를 시작합니다. 선택적으로 문제를 설명하면, 시스템이 가설 수립 -> 증거 수집 -> 해결을 수행합니다 |
| `/gsd:add-todo [desc]` | 아이디어를 할 일 목록에 기록합니다. 로드맵은 수정하지 않습니다 |
| `/gsd:check-todos` | 현재 할 일 목록을 확인합니다 |
| `/gsd:map-codebase` | 기존 코드베이스를 분석합니다 -- 기술 스택, 아키텍처, 관례, 잠재적 문제 |
### 세션 및 설정 관리
| 명령어 | 설명 |
| ------------------ | ----------------------------------------------------- |
| `/gsd:pause-work` | 작업을 일시 중지합니다. 현재 상태를 STATE.md에 저장하여 다음에 쉽게 복원할 수 있습니다 |
| `/gsd:resume-work` | 작업을 재개합니다. STATE.md에서 이전 상태를 읽어 중단된 지점부터 계속합니다 |
| `/gsd:progress` | 프로젝트 전체 진행 상황을 확인합니다 -- 완료된 단계 수, 현재 위치, 대기 중인 항목 |
| `/gsd:help` | 사용 가능한 모든 명령어와 간단한 설명을 표시합니다 |
| `/gsd:settings` | GSD 설정을 확인하고 수정합니다 |
| `/gsd:set-profile` | 모델 설정을 전환합니다 (quality / balanced / budget) |
| `/gsd:update` | GSD를 최신 버전으로 업데이트합니다 |
## 설정 상세
### 모델 설정
GSD는 세 가지 모델 설정을 지원하며, `/gsd:set-profile`로 전환합니다:
| 설정 | 계획 | 실행 | 검증 | 적합한 시나리오 |
| -------------- | ------ | ------ | ------ | -------------------------- |
| quality | Opus | Opus | Sonnet | 복잡한 프로젝트, 핵심 기능, 처음 사용 |
| balanced (기본값) | Opus | Sonnet | Sonnet | 일상적인 개발, 대부분의 시나리오에 최적의 균형 |
| budget | Sonnet | Sonnet | Haiku | 간단한 기능, 예산에 민감한 경우, 빠른 반복 |
### 핵심 설정
`/gsd:settings`를 통해 다음 설정을 확인하고 수정할 수 있습니다:
| 설정 | 기본값 | 설명 |
| ------------------------ | ---------- | ---------------------------------------------------------- |
| `mode` | `balanced` | 모델 설정 선택 |
| `depth` | `standard` | 조사 깊이: `quick`(빠른) / `standard`(표준) / `deep`(심층) |
| `git.branching_strategy` | `feature` | Git 브랜치 전략: `feature`(기능별 브랜치) / `phase`(단계별 브랜치) / `none` |
### 워크플로우 스위치
다음 에이전트는 개별적으로 켜고 끌 수 있으며, 속도와 품질 사이에서 균형을 맞출 수 있습니다:
| 스위치 | 기본값 | 설명 |
| -------------- | --- | --------------------------- |
| `research` | 켜짐 | 계획 수립 전에 자동 조사를 수행할지 여부 |
| `plan_check` | 켜짐 | 계획 생성 후 자동 검증을 수행할지 여부 |
| `verifier` | 켜짐 | 실행 후 자동 검증을 수행할지 여부 |
| `auto_advance` | 꺼짐 | 단계 완료 후 자동으로 다음 단계로 진행할지 여부 |
> `research`와 `plan_check`를 끄면 속도를 크게 높일 수 있지만, 계획 품질이 낮아질 수 있습니다. 프로젝트에 익숙해진 후에 끄는 것을 고려하는 것이 좋습니다.
## 산출물 파일 구조
GSD의 모든 상태와 산출물은 `.planning/` 디렉토리에 저장됩니다. 이 구조를 이해하면 디버깅과 수동 개입에 도움이 됩니다.
### 프로젝트 레벨 파일
| 파일 | 역할 | 생성 시점 |
| ----------------- | ------------------------- | ------------------------- |
| `PROJECT.md` | 프로젝트 비전과 범위 | `new-project` |
| `REQUIREMENTS.md` | 버전별 요구사항 문서, 단계 추적 포함 | `new-project` |
| `ROADMAP.md` | 단계 계획과 진행 상황 | `new-project` |
| `STATE.md` | 현재 상태 -- 결정 사항, 차단 요소, 위치 | `new-project`, 지속적으로 업데이트 |
### 단계 레벨 파일
각 단계는 다음 파일을 생성합니다 (단계 1을 예로 들면):
| 파일 | 역할 | 생성 시점 |
| -------------------- | -------------- | ----------------- |
| `01-CONTEXT.md` | 논의 단계의 의사결정 기록 | `discuss-phase 1` |
| `01-RESEARCH.md` | 조사 결과와 기술 탐구 | `plan-phase 1` |
| `01-01-PLAN.md` | 첫 번째 원자적 작업 계획 | `plan-phase 1` |
| `01-02-PLAN.md` | 두 번째 원자적 작업 계획 | `plan-phase 1` |
| `01-01-SUMMARY.md` | 첫 번째 계획의 실행 기록 | `execute-phase 1` |
| `01-02-SUMMARY.md` | 두 번째 계획의 실행 기록 | `execute-phase 1` |
| `01-VERIFICATION.md` | 자동 검증 결과 | `execute-phase 1` |
| `01-UAT.md` | 사용자 수락 테스트 기록 | `verify-work 1` |
### 디렉토리 구조 예시
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## 실전 워크플로우 시연
아래에서는 "블로그 시스템에 댓글 기능 추가"를 예로 들어 초기화부터 납품까지의 전체 흐름을 시연합니다.
### Step 1: 프로젝트 초기화
```text
/gsd:new-project
```
시스템이 계속 질문을 시작합니다:
```
> 무엇을 구축하고 싶으신가요?
"Next.js 블로그에 댓글 기능을 추가하고 싶습니다. 익명 및 로그인 댓글,
Markdown 렌더링, 관리자 대시보드를 지원합니다. 기술 스택은 Prisma + PostgreSQL입니다."
```
설명이 상세할수록 시스템의 추가 질문이 줄어듭니다. TÂCHES의 조언은 다음과 같습니다: 원하는 것을 설명하는 대략적인 비전 문서를 준비하세요 -- 기술 세부 사항을 알 필요는 없습니다.
시스템이 완료되면 네 개의 파일을 산출하고 로드맵 승인을 요청합니다. 승인 후 구축 단계에 진입합니다.
> **기존 코드베이스가 있으신가요?** 먼저 `/gsd:map-codebase`를 실행하세요. 시스템이 기존 아키텍처와 관례를 분석하며, 이후 `new-project`가 기존 코드를 기반으로 계획을 수립할 수 있습니다.
### Step 2: 논의 단계
```text
/gsd:discuss-phase 1
```
시스템이 모호한 부분을 식별하고 하나씩 질문합니다:
```
> 댓글 중첩 수준: 다단계 중첩을 지원할까요, 아니면 1단계 답글만 지원할까요?
> 익명 댓글: 인증 코드가 필요한가요, 아니면 바로 제출할 수 있나요?
> 관리자 대시보드: 일괄 작업이 필요한가요, 아니면 건별 심사인가요?
```
여기서의 모든 결정이 후속 계획 품질에 직접적으로 영향을 미칩니다. 확실하지 않으면 시스템에게 기본값을 사용하도록 할 수 있지만, 심층적인 논의를 통해 실행 단계에서의 재작업을 크게 줄일 수 있습니다.
### Step 3: 계획 단계
```text
/gsd:plan-phase 1
```
시스템이 수행하는 작업:
1. Prisma + PostgreSQL로 댓글 시스템을 구현하는 방법을 조사합니다
2. 2-3개의 원자적 작업 계획을 생성합니다 (예: 데이터 모델, API 라우트, 프론트엔드 컴포넌트)
3. 계획이 모든 요구사항을 충족하는지 자동 검증합니다
각 계획은 충분히 작아서 완전히 새로운 컨텍스트 윈도우에서 완료할 수 있습니다.
### Step 4: 실행 단계
```text
/gsd:execute-phase 1
```
시스템이 웨이브 단위로 실행을 시작합니다:
* **Wave 1** (의존성 없음): 데이터베이스 스키마, Prisma 모델 -- 병렬 실행
* **Wave 2** (Wave 1에 의존): API 라우트, 댓글 CRUD -- 병렬 실행
* **Wave 3** (Wave 2에 의존): 프론트엔드 댓글 컴포넌트 -- 독립 실행
각 작업은 완전히 새로운 200k tokens 컨텍스트에서 실행되며, 완료 후 독립적으로 git commit됩니다.
### Step 5: 검증 단계
```text
/gsd:verify-work 1
```
시스템이 하나씩 확인하도록 안내합니다:
```
> ✅ 데이터베이스 테이블이 생성되었습니다
> ✅ API 라우트가 올바른 상태 코드를 반환합니다
> ❓ 블로그 게시물 하단에 댓글 입력란이 보이시나요? [예/아니오/문제 설명]
> ❓ 댓글 제출 후 페이지가 실시간으로 업데이트되나요? [예/아니오/문제 설명]
```
실패한 항목이 있으면 시스템이 자동으로 진단하고 수정 계획을 생성합니다. `/gsd:execute-phase 1`을 다시 실행하면 수정이 실행됩니다.
### 일반적인 운영 시나리오
**긴급 단계 삽입**: 요구사항이 변경되어 현재 단계 앞에 새 작업을 삽입해야 합니다.
```text
/gsd:insert-phase 2
```
이후 단계는 자동으로 번호가 재할당됩니다 (기존 Phase 2가 Phase 3이 되고, 이후도 마찬가지입니다).
**일시 중지 및 재개**: 다른 일을 처리하기 위해 작업을 중단해야 합니다.
```text
/gsd:pause-work # 현재 상태를 저장합니다
# ... 다른 일을 처리합니다 ...
/gsd:resume-work # 마지막으로 중단된 위치로 복원합니다
```
**불만족스러운 결과 롤백**:
```bash
git reset --hard HEAD~3 # 실행 전 상태로 되돌아갑니다
```
```text
/gsd:remove-phase 2 # 해당 단계의 모든 산출물 파일을 연쇄적으로 삭제합니다
```
TÂCHES는 라이브 스트리밍에서 이 작업을 여러 번 시연했습니다 -- 마음에 들지 않으면 롤백하고, 깔끔하게 처리합니다.
## 디버그 워크플로우
검증에서 문제가 발견되거나 개발 과정에서 버그를 만났을 때, GSD는 전용 디버그 프로세스를 제공합니다.
```text
/gsd:debug 댓글 제출 후 페이지가 실시간으로 업데이트되지 않습니다
```
시스템이 **격리된 디버그 서브 에이전트**를 시작하며, 다음과 같은 워크플로우를 따릅니다:
1. **가설 수립** — 문제 설명을 기반으로 여러 가능한 근본 원인 가설을 생성합니다
2. **증거 수집** — 가설을 하나씩 검증하고, 코드, 로그, 네트워크 요청을 확인합니다
3. **해결** — 근본 원인을 파악한 후 수정 방안을 생성합니다
핵심 특성:
* **컨텍스트 격리**: 디버그 에이전트는 자체 컨텍스트 윈도우를 가지며, 메인 개발 컨텍스트를 오염시키지 않습니다
* **문서 추적**: 전체 조사 과정을 기록하는 독립적인 디버그 문서를 생성합니다
* **수정 계획**: 진단이 완료되면 바로 실행할 수 있는 수정 계획을 산출합니다
이것은 메인 컨텍스트에서 직접 디버깅하는 것보다 훨씬 효율적입니다 -- 디버그 정보가 메인 윈도우에 축적되지 않습니다.
## 실전 경험
TÂCHES의 라이브 스트리밍과 Chase AI의 사용 경험을 종합하여 다음과 같은 실전 조언을 드립니다.
### 천천히 해야 빨라집니다
TÂCHES는 초기에 GSD를 사용할 때 "빨리빨리빨리"라는 마인드셋이었다고 솔직히 말했지만, 나중에 **조사와 논의 단계에 더 많은 시간을 투자하면 실행 단계에서의 재작업이 오히려 줄어든다**는 것을 발견했습니다. 새 버전의 GSD에 `research-project`와 `define-requirements` 단계가 추가된 것은 바로 실행에 들어가기 전에 방향을 올바르게 잡기 위해서입니다.
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### 단계 사이에 컨텍스트를 정리합니다
TÂCHES의 습관은 **각 단계 사이에 `clear`를 실행**하여 메인 컨텍스트를 간결하게 유지하는 것입니다. 그는 Warp 터미널을 사용하며, 각 윈도우를 전체 화면(Command+Shift+Enter)으로 설정하고, 한 윈도우에서 현재 단계를 실행하는 동시에 다른 윈도우에서 다음 단계를 조사합니다.
### Token 비용의 절충
GSD의 서브 에이전트 방식은 확실히 Claude Code를 직접 사용하는 것보다 더 많은 token을 소비합니다. 하지만 Chase AI는 강력한 논거를 제시했습니다: **"plan twice, prompt once"(두 번 계획하고 한 번 프롬프트)가 "한 번 프롬프트하고 계속 수정"하는 것보다 장기적으로 더 token을 절약합니다.** 새로운 컨텍스트에서 한 번에 올바르게 수행하는 것이, 퇴화된 컨텍스트에서 반복적으로 수정하는 것보다 훨씬 효율적이기 때문입니다.
### 불만족스러울 때의 처리
어떤 단계의 결과가 만족스럽지 않다면, `git reset --hard`를 실행한 후 `/gsd:remove-phase`로 해당 단계의 모든 산출물 파일을 연쇄적으로 삭제할 수 있습니다. TÂCHES는 라이브 스트리밍에서 이 작업을 실제로 시연했습니다 -- 어떤 시각적 효과가 마음에 들지 않으면 만족스러운 이전 상태로 바로 롤백하고, 깔끔하게 처리합니다.
### To-Do 시스템
`/gsd:add-todo`로 언제든지 아이디어를 할 일 목록에 기록할 수 있으며, 로드맵을 수정할 필요가 없습니다. 이러한 아이디어는 `/gsd:discuss-milestone` 시에 꺼내어 다음 마일스톤의 입력으로 활용할 수 있습니다. TÂCHES의 전략은 "먼저 기능을 구현하고, milestone 2에서 UI를 다듬는 것"입니다.
## 자주 묻는 질문과 모범 사례
### 모범 사례
**상세한 초기 설명을 제공하세요**. `/gsd:new-project`의 품질은 여러분의 입력 품질에 달려 있습니다. 대략적인 비전 문서를 준비하세요 -- 목표, 사용자, 핵심 기능, 알려진 제약을 설명합니다. 설명이 정확할수록 시스템의 추가 질문이 줄어들고 계획이 더 정확해집니다.
**단계 사이에 컨텍스트를 정리하세요**. 각 단계를 완료한 후 `clear` 또는 `/compact`를 실행하여 메인 컨텍스트 윈도우를 간결하게 유지합니다. TÂCHES의 습관은 메인 컨텍스트를 30-40% 수준으로 유지하는 것입니다.
**먼저 Quick Mode로 테스트하세요**. 확실하지 않은 작은 기능의 경우, 먼저 `/gsd:quick`으로 시험해 보세요. 효과가 좋으면 정식 로드맵에 포함시킵니다.
**기존 프로젝트에는 먼저 map-codebase를 실행하세요**. 기존 코드베이스에서 GSD를 사용하기 전에 먼저 `/gsd:map-codebase`를 실행하세요. 시스템이 기술 스택, 아키텍처, 관례를 분석하며, 이후의 계획이 기존 코드에 더 적합하게 됩니다.
### FAQ
**Q: GSD는 어떤 런타임을 지원하나요?**
A: Claude Code, OpenCode, Gemini CLI입니다. 설치 시 개별 또는 전부를 선택할 수 있습니다.
**Q: Quick Mode와 전체 모드의 차이점은 무엇인가요?**
A: Quick Mode는 GSD의 기본 보장(원자적 커밋, 상태 추적)을 제공하지만, 조사, 계획 검증, 검증 단계를 건너뜁니다. 버그 수정, 작은 기능, 설정 변경 등 완전한 계획이 필요하지 않은 작업에 적합합니다.
**Q: 실행 중에 일시 중지할 수 있나요?**
A: 가능합니다. `/gsd:pause-work`가 현재 상태를 STATE.md에 저장합니다. 다음에 `/gsd:resume-work`를 실행하면 시스템이 마지막으로 중단된 위치에서 계속합니다.
**Q: token 비용을 어떻게 제어하나요?**
A: 세 가지 방법이 있습니다 -- (1) `budget` 설정으로 전환: `/gsd:set-profile budget`; (2) `research` 또는 `plan_check` 에이전트 끄기; (3) 간단한 작업에 `/gsd:quick` 사용.
**Q: GSD를 Ralph과 함께 사용할 수 있나요?**
A: 가능합니다. GSD와 Ralph은 서로 다른 문제를 해결합니다 -- GSD는 계획과 구조화된 실행을 담당하고, Ralph은 자율 순환 실행을 담당합니다. GSD의 `new-project`와 `plan-phase`로 완전한 계획을 생성한 후, 인간의 개입이 필요 없는 단계에 대해 Ralph 순환을 사용하여 실행할 수 있습니다.
**Q: 다중 협업은 어떻게 하나요?**
A: `.planning/` 디렉토리를 Git에 커밋할 수 있습니다. 여러 사람이 각자 다른 단계를 실행하고 Git으로 결과를 병합할 수 있습니다. 다만 동일한 단계를 동시에 실행하는 것은 피하는 것이 좋습니다.
## 요약
GSD의 핵심 가치는 **복잡성을 시스템 안에 숨기고, 사용자에게는 단순함을 남기는 것**입니다. 몇 가지 명령어만 있으면 됩니다 -- `new-project`, `discuss-phase`, `plan-phase`, `execute-phase`, `verify-work` -- 시스템이 뒤에서 모든 컨텍스트 관리, 서브 에이전트 오케스트레이션, 품질 검증을 처리합니다.
설치부터 납품까지, GSD는 명확한 경로를 제공합니다: 원하는 것을 설명 -> 구현 세부 사항 논의 -> 원자적 계획 생성 -> 병렬 실행 -> 산출물 검증. 매 단계마다 여러분이 개입할 기회가 있고, 매 단계마다 파일로 기록됩니다.
이것은 "버튼 하나를 누르면 끝"이라는 마법이 아닙니다. 이것은 여러분의 참여가 필요하지만 대부분의 인지 부담을 대신 져주는 시스템입니다. TÂCHES가 말했듯이: 여러분은 고위 프로젝트 매니저이고, GSD는 여러분의 실행 팀입니다.
***
**추가 읽을거리**:
* [GSD 심층 분석](/ko/docs/notes/gsd/concept) — 핵심 원리, 워크플로우와 기술 아키텍처
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) — Context Rot과 Ralph 방법론
* [snarktank/ralph 실전 가이드](/ko/docs/notes/ralph-wiggum/snarktank) — Ralph의 설치, PRD 작성과 실전
* [규격 주도 개발이란](/ko/docs/notes/speckit/concept) — Vibe Coding에서 규격 주도 개발로의 패러다임 전환
* [Speckit 실전 가이드](/ko/docs/notes/speckit/practice) — Speckit 명령어 상세와 완전한 사례
# gstack: YC CEO가 Claude Code에 기업가적 경험을 담을 때
## 소개
이전 노트에서는 [Ralph Wiggum](/ko/docs/notes/ralph-wiggum/concept)의 무한 루프부터 [GSD](/ko/docs/notes/gsd/concept)의 사양 중심 개발까지 Claude Code 생태계의 다양한 "향상 솔루션"을 살펴보았습니다. 그들은 모두 같은 질문에 대답하려고 노력하고 있습니다. \*\*AI 프로그래밍을 "적응"에서 "신뢰할 수 있는 전달"로 어떻게 바꾸나요? \*\*
Ralph의 대답은 "모든 것을 다시 시작"입니다. 컨텍스트 부패를 피하기 위해 매번 새로운 프로세스를 사용하십시오. GSD의 대답은 "사양 중심"입니다. 즉, 구조화된 단계 계획 및 검증 주기를 통해 품질을 보장합니다. 하지만 실행 시스템뿐만 아니라 완전한 가상 엔지니어링 팀을 원한다면 어떻게 될까요? CEO는 제품 결정을 내리고, 엔지니어링 관리자는 아키텍처를 검토하고, 디자이너는 경험을 제어하고, QA는 실제 브라우저 테스트를 실행하고, 릴리스 엔지니어는 출시를 관리합니다. 이 모든 작업은 AI가 수행하고 사용자가 명령합니다.
이것이 gstack의 핵심 아이디어입니다.
## gstack이란 무엇입니까?
**Garry Tan**은 gstack의 창시자로서 풍부한 기술 및 기업가적 배경을 가지고 있습니다. 그는 14세에 코드 작성을 시작했고, 스탠포드 컴퓨터 공학을 졸업했으며, Posterous(나중에 Twitter에 인수됨)를 공동 창립한 Palantir의 10번째 직원이며, 2023년부터 Y Combinator의 사장 겸 CEO로 재직하고 있습니다.
그는 gstack을 사용하여 60일 만에 600,000라인 이상의 프로덕션 코드(35% 테스트)를 릴리스했습니다. 이는 여전히 YC를 풀타임으로 운영하면서 하루 평균 10,000라인 이상입니다. 프로젝트 중 하나인 garylist.org는 150,000줄의 코드와 35%의 테스트 범위로 21일 만에 시작되었습니다. 그 자신의 말에 따르면, 코드 품질은 그가 2년 동안 500만 달러와 10명의 엔지니어를 투자한 이전 기업가 프로젝트보다 뛰어났습니다.
프로젝트는 2026년 3월 11일에 오픈 소스로 공개된 이후 3주 이내에 v0에서 v0.15.1.0으로 반복되었으며 GitHub는 60,500개 이상의 별을 받았습니다. MIT 라이센스, 완전 오픈 소스.
## 도구 생태계에서 gstack의 위치
| 치수 | 네이티브 클로드 코드 | 랄프 위검 | GSD | 스펙킷 | 초능력 | **지스택** |
| -------- | ---------------- | ---------------- | ----------------- | ------------------ | ------------- | ----------------------- |
| 핵심 포지셔닝 | 유니버설 AI 코딩 어시스턴트 | 무한 루프 반복 | 상황별 엔지니어링 + 사양 기반 | 요구 사항 → 사양 → 작업 | 프로세스 규율 + TDD | **역할 기반 가상 팀** |
| 코어 패턴 | 대화형 프로그래밍 | Bash 루프 + 새 프로세스 | 단계별 로드맵 | 사양 → 계획 → 작업 | 엄격한 개발 파이프라인 | **스프린트 7단계 프로세스** |
| 인간의 참여 | 실시간 대화 | AFK(핸즈오프) | 단계별 검증 | 사양 승인 | 단계별 검증 | **단계별 역할 검토** |
| 고유한 기능 | 기본 코딩 | 무제한 반복 | 컨텍스트 부패 관리 | 요구사항 추적 | 강제 TDD | **브라우저 자동화 + 다중 역할 검토** |
| 시나리오에 적합 | 간단한 작업 | 지속적인 반복 | 대규모 프로젝트 관리 | 엄격한 요구 사항이 있는 프로젝트 | 엔지니어링 품질 보증 | **풀 프로세스 제품 개발** |
표에서 주요 패턴을 볼 수 있습니다. \*\*이러한 도구는 서로 경쟁하지 않지만 다양한 차원에서 AI 프로그래밍 문제를 해결합니다. \*\*
Superpowers는 **프로세스 규율**을 사용하여 코드 품질(필수 TDD, 구조화된 대화, 구현 계획)을 보장합니다. GSD는 **컨텍스트 엔지니어링**을 사용하여 복잡한 프로젝트(단계 계획, 하위 에이전트의 새로운 컨텍스트, 파일 시스템 상태)를 관리합니다. gstack은 **역할 분해**를 사용하여 의사결정 품질을 향상합니다(CEO 관점에서 제품 검토, 엔지니어링 관리자 검토 아키텍처, QA에서 실제 브라우저 실행).
간단히 말하면 Superpowers는 프로세스 가드레일을 기반으로 하고 gstack은 역할 설계를 기반으로 합니다. 전자는 1에서 N까지의 프로젝트 구현에 적합하고 후자는 0에서 1까지의 제품 구성에 적합합니다. \*\*둘은 경쟁 제품이 아닌 상호 보완적인 제품입니다. \*\*
## 핵심 작업 흐름: Sprint 7단계
gstack은 전체 개발 프로세스를 **Think → Plan → Build → Review → Test → Ship → Reflect**의 주기로 구성합니다. 이를 "The Sprint"라고 합니다. Agile Sprint가 아니라 "역할이 순차적으로 나타나는" 개발 리듬입니다.
### 1. 생각하기 — 제품 클리닉
```text
/office-hours
```
gstack의 가장 특징적인 스킬입니다. 영감은 YC의 근무 시간에서 직접 나옵니다. 기업가는 YC 파트너를 만나 자기 성찰을 하러 갑니다. AI가 **6가지 강제 질문**을 묻습니다.
1. 특별히 누가 이것을 필요로 합니까?
2. 오늘 없으면 어떻게 되나요?
3. 이 문제가 지금 긴급한 이유는 무엇입니까?
4. 그것이 작동하는지 어떻게 알 수 있나요?
5. 아무것도 하지 않으면 어떻게 되나요?
6. 출시할 수 있는 가장 작은 버전은 무엇입니까?
목적은 코드 작성을 돕는 것이 아니라 코드를 작성하기 전에 **문제 자체를 재검토**하는 것입니다.
### 2. 계획 — 다중 역할 검토
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
CEO Review는 본질적으로 "창립자 모드"입니다. 요구 사항을 문자 그대로 실행하는 대신 한 걸음 물러나 "이 제품의 진정한 목적은 무엇입니까?"라고 묻습니다. 범위 확장, 선택적 확장, 범위 유지 및 범위 축소의 네 가지 모드를 지원합니다.
### 3. 빌드 — 코딩 구현
승인된 계획에 따라 코딩을 시작하세요. 이 단계에서는 표준 Claude Code 기능을 사용합니다.
### 4. 검토 - 병행 전문가 검토
```text
/review
```
이 기술은 **7개의 병렬 하위 에이전트**를 한 번에 파견하여 테스트, 유지 관리 가능성, 보안, 성능, 데이터 마이그레이션, API 계약, 레드팀 공격 등 7가지 관점에서 코드를 검토합니다. 명백한 문제는 자동으로 수정됩니다.
### 5. 테스트 - 실제 브라우저 QA
```text
/qa
```
연습 시험이 아닙니다. QA 기술은 실제 테스터가 하는 것처럼 **실제 헤드리스 Chromium 브라우저**를 실행하고, 앱을 열고, 버튼을 클릭하고, 양식을 작성하고, 스크린샷을 찍습니다. 자동으로 버그를 수정하고, 회귀 테스트를 생성하고, 버그 발견 후 다시 검증합니다.
### 6. 배송 — 원클릭 게시
```text
/ship
```
마스터 브랜치를 자동으로 동기화하고, 테스트를 실행하고, 차이점을 검토하고, 버전 번호 및 CHANGELOG를 업데이트하고, 커밋, 푸시, PR 생성을 수행합니다. 프로젝트에 테스트 프레임워크가 없으면 먼저 프레임워크를 빌드합니다.
### 7. 성찰 - 검토 및 학습
```text
/retro
```
엔지니어링 관리자 스타일 주간 보고서: 커밋 기록, 테스트 비율, 코드 품질 추세를 분석합니다. 여러 사람으로 구성된 팀 분석 및 "연속 출시 일수"와 같은 추적 지표를 지원합니다.
## 작동 이유: 기술 원리
### 데몬 찾아보기: AI에 주목
gstack의 가장 독특한 기술적 기여는 Browse Daemon입니다. 이는 localhost HTTP를 통해 통신하는 지속적인 헤드리스 Chromium 인스턴스입니다. 첫 번째 호출은 브라우저를 시작하고(~~3초) 각 후속 명령에는 100~~200ms만 소요됩니다. 이는 AI가 DOM 구조를 추측하는 것이 아니라 실제로 앱을 볼 수 있다는 것을 의미합니다.
또한 CSS 선택기를 작성하지 않고도 접근성 트리를 통해 요소를 찾을 수 있는 **Ref System**(요소 참조 `@e1`, `@e2`)을 도입했습니다. 이는 커뮤니티(비평가 포함)에서 일반적으로 인정받는 "진정한 기술적 기여"입니다.
### 역할 분류: 상담원이 아니라 팀
gstack이 하는 일은 모든 역할을 독립적인 프롬프트 파일로 분해하여 Claude Code가 다양한 단계에서 다양한 역할의 관점으로 전환하여 코드를 검토할 수 있도록 하는 것입니다. 이는 본질적으로 세련된 프롬프트 엔지니어링입니다.
핵심 통찰력은 다음과 같습니다. \*\*계획은 검토와 같지 않고, 검토는 출시와 같지 않으며, 창업자의 취향과 엔지니어링의 엄격함은 완전히 다른 사고 방식입니다. \*\* 일반 에이전트가 모든 작업을 수행하도록 하는 대신 필요할 때 "브레인 모드"를 전환합니다(창업자적 사고, 엔지니어링 엄격함, 편집증적 검토, 빠른 실행).
### 세 가지 주요 철학
gstack의 ETHOS.md에는 세 가지 핵심 개념이 기록되어 있습니다.
1. **호수 끓이기**: AI가 완전성의 한계 비용을 0으로 만들 때 항상 완전한 구현(100% 테스트 적용 범위, 모든 엣지 케이스, 모든 오류 경로)을 선택하십시오. "릴리스 단축키"는 옛날 생각입니다.
2. **구축 전 검색**: 세 가지 지식 계층 - 오랜 시간 테스트를 거친 패턴, 새롭고 인기 있는 솔루션, 첫 번째 원칙. 모두가 무엇을 하고 있는지 이해하고, 그들의 가정에 의문을 제기하고, 일반적인 솔루션이 왜 잘못된지 알아내는 것부터 시작하세요.
3. **사용자 주권**: AI 추천, 인간의 의사결정. 두 AI 모델이 합의에 도달하더라도 사용자의 판단이 여전히 우선합니다. 사용자에게는 도메인 지식, 전략적 관점 및 취향이 있기 때문입니다.
## gstack의 경계와 논란
gstack에 대한 커뮤니티의 반응은 아마도 모든 AI 프로그래밍 도구 중에서 가장 양극화되어 있을 것입니다.
**밝은 면**: 창립자와 비기술 개발자는 일반적으로 동의합니다. 특히 `/office-hours` 및 `/plan-ceo-review`과 같은 "제품 사고" 기술은 많은 독립 개발자가 코딩을 시작하기 전에 제품 방향을 재검토하는 데 도움이 되었습니다. 엔지니어링 검토(`/review`)를 통해 실제로 숨겨진 보안 취약점을 발견할 수 있습니다. 이 다각도 병렬 검토 모델은 실용적인 가치를 가지고 있습니다.
**질문하는 측면**도 매우 직접적입니다.
* **LOC 표시는 그다지 중요하지 않습니다**: 60일 동안 600,000줄의 코드. 코드 줄 수는 결코 품질 지표가 아닙니다. 많은 양의 코드는 단지 스캐폴딩과 상용구일 수도 있습니다.
* **기본적으로 프롬프트 템플릿**: 각 스킬은 SKILL.md 파일이며 기술 임계값은 높지 않습니다. 실제 가치는 파일 자체에 있는 것이 아니라 프롬프트 디자인의 품질에 있습니다.
* **AI 자체 검토 코드의 한계**: `/review` AI가 작성한 코드를 AI가 검토하도록 하는 것은 자신의 숙제를 수정하는 것과 같습니다. 다중 역할 병렬 처리는 이 문제를 완화할 수 있지만 여전히 동일한 모델입니다.
* **연예인 효과 보너스**: 창업자가 YC CEO가 아닐 경우, 이 프로젝트가 그다지 큰 주목을 받지 못할 확률이 높습니다.
**내 의견**: 논란은 제쳐두고, gstack의 정말 가치 있는 부분은 두 가지입니다. 바로 Browse Daemon의 브라우저 자동화 기술과 역할 분해의 디자인 패턴입니다. 이 중 어느 것도 Garry Tan이 누구인지에 달려 있지 않습니다. 역할화의 핵심 의미는 기술 수준이 아니라 행동 수준에 있습니다. 이는 일반 에이전트에게 모든 것을 맡기는 대신 AI 워크플로를 보다 의식적으로 구성하는 데 도움이 됩니다.
gstack은 분기 및 사용자 정의에 적합합니다. 필요한 기술을 습득하고 프롬프트를 모두 복사하는 대신 원하는 대로 변경할 수 있습니다.
## 비디오 리소스
## 마지막에 쓰세요
gstack은 AI 프로그래밍 도구의 흥미로운 방향을 나타냅니다. 즉, AI를 보다 자율적으로 만들거나(Ralph의 경로) 프로세스를 보다 엄격하게 만드는 것이 아니라(Superpowers의 경로) AI가 다양한 역할을 수행하여 의사 결정의 품질을 향상시키는 것입니다. 그 논란은 AI 프로그래밍 생태계의 풍부함을 보여줄 뿐입니다. 모든 사람에게 적합한 솔루션은 없습니다.
gstack에 관심이 있다면 다음 단계는 [실용 장](/ko/docs/notes/gstack/practice)을 읽어보는 것입니다. 설치부터 전체 워크플로우를 통한 실행까지 단계별 튜토리얼입니다.
***
**관련 자료**:
* [GSD 개념 소개](/ko/docs/notes/gsd/concept) — 또 다른 구조화된 AI 프로그래밍 솔루션
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) — 무한 루프 반복의 시작점 이해
* [Claude Skills Concept](/ko/docs/notes/claude-skills/concept) — 스킬의 기본 메커니즘을 이해합니다.
# gstack 프런트엔드 기술 파노라마: 디자인부터 출시까지 AI 워크플로
## 소개
이전 노트에서는 [gstack이란 무엇인가](/ko/docs/notes/gstack/concept), [워크플로 실행 방법](/ko/docs/notes/gstack/practice), [Skill의 엔지니어링 구조](/ko/docs/notes/gstack/skill-architecture)에 대해 이야기했습니다. 하지만 아직 논의되지 않은 질문이 하나 있습니다. gstack 설치 후 가져온 60개 이상의 기술 중 프런트엔드/UI 디자인과 관련된 기술은 무엇입니까? 어떤 순서로? \*\*
이 노트는 두 가지 작업을 수행합니다. 먼저 프런트 엔드와 관련된 \~27개의 기술을 기능별로 분류한 다음 흥미로운 작은 프로젝트(카운트다운 기념 페이지)를 사용하여 실제 효과를 볼 수 있도록 각 단계의 스크린샷과 함께 전체 워크플로를 처음부터 끝까지 단계별로 안내합니다.
## 프런트엔드 스킬 키트 파노라마
gstack의 프런트엔드 기술은 기초부터 지붕까지 6개의 기능 계층으로 나눌 수 있으며, 각 계층은 서로 다른 단계에서 문제를 해결합니다.
### 설계 인프라
디자인 언어를 결정하기 위한 프로젝트 수준의 일회성 설정과 모든 후속 기술은 이러한 벤치마크를 참조합니다.
| 스킬 | 해야 할 일 | 사용 시기 |
| ---------------------- | --------------------------------------- | -------------------------------------- |
| `/design-consultation` | 완벽한 디자인 시스템 상담, 출력 색상 매칭, 글꼴, 간격, 질감 방향 | 새 프로젝트를 시작하거나 시각적 스타일을 재정의하려는 경우 |
| `/teach-impeccable` | 디자인 기본 설정을 한 번에 수집하여 AI 구성 파일에 기록 | AI가 미학을 기억할 수 있도록 gstack을 설치한 후 한 번 실행 |
| `/brand-guidelines` | 기존 브랜드 컬러 매칭 및 글꼴 사양 적용 | 기존 브랜드 매뉴얼이 있는 경우 바로 적용 |
> 프로젝트에 이미 `DESIGN.md`이 있는 경우 이 레벨을 건너뛸 수 있습니다.
### 디자인 탐색
방향이 확실하지 않은 경우 여러 옵션을 빠르게 비교하세요.
| 스킬 | 해야 할 일 | 사용 시기 |
| ------------------ | -------------------------------- | --------------------------------- |
| `/design-shotgun` | 3\~5개의 시각적 솔루션 생성, 비교 패널 열기 | 어떤 스타일을 원하는지 잘 모르겠으나 가능성을 보고 싶습니다 |
| `/frontend-design` | 인식 가능한 프로덕션 수준 프런트엔드 인터페이스 코드 생성 | 방향이 명확한 후 바로 작업 |
| `/canvas-design` | 포스터, 시각 예술 생성(PNG/PDF) | 웹 구성요소가 아닌 정적인 시각적 디자인이 필요함 |
### 디자인 구현
계획을 실제로 실행 가능한 코드로 전환하고 조판, 레이아웃 및 응답성을 처리합니다.
| 스킬 | 해야 할 일 | 사용 시기 |
| ------------------------ | ------------------------------ | -------------------------------- |
| `/design-html` | 확정된 디자인 초안을 프로덕션급 HTML/CSS로 변환 | 직접 구현하고 싶은 모형이 있습니다 |
| `/mobile-responsiveness` | 모바일 우선 반응형 레이아웃 및 터치 상호작용 | 처음부터 모바일 적응 |
| `/adapt` | 장치 및 화면 크기에 따른 중단점 적응 | 데스크톱 버전이 있으며 휴대폰/태블릿에 적용해야 함 |
| `/typeset` | 글꼴 선택, 수준, 크기, 두께 및 가독성 최적화 | 텍스트 레이아웃은 "거의 의미가 없어 보입니다" |
| `/arrange` | 레이아웃 간격, 시각적 리듬, 정렬 복구 | 일관성 없는 간격, 레이아웃이 혼잡하거나 흩어져 있는 느낌 |
### 디자인 개선
기능적 완성도를 바탕으로 역동적인 효과와 개성, 감성적인 디테일을 주입합니다.
| 스킬 | 해야 할 일 | 사용 시기 |
| ------------ | ------------------------------------------- | ------------------------------- |
| `/animate` | 의도적인 마이크로 상호 작용 및 애니메이션 추가 | 페이지 기능은 양호하지만 "단단한" 느낌 |
| `/delight` | 놀라운 세부 사항과 개인화된 터치를 추가하세요 | 사용자가 이 페이지를 기억하도록 하기 |
| `/bolder` | 시각적 효과 증폭 | 디자인이 너무 평범하고 너무 안전하다 |
| `/colorize` | 단조로운 인터페이스에 전략적인 색상 추가 | 페이지가 너무 회색이고 너무 평범하며 따뜻함이 부족합니다 |
| `/overdrive` | 기술적 폭발 수준 효과 - 셰이더, 스프링 물리학, 스크롤 드라이브 애니메이션 | 특정 지역은 와우 효과를 원합니다 |
| `/onboard` | 새로운 사용자 안내 프로세스, 빈 상태 디자인 | 최초 사용자 경험 |
이 네 가지 강화 기술은 **진행적 관계**에 있습니다. `animate`은 기본 동적 효과, `delight`은 감정, `bolder`는 증폭, `overdrive`은 폭발입니다. 프로젝트 요구 사항에 따라 단계별로 쌓으므로 모두 사용할 필요는 없습니다.
### 디자인 최적화
융합 및 개선 – 과잉 제거, 편차 정렬, 거친 가장자리 연마.
| 스킬 | 해야 할 일 | 사용 시기 |
| ------------ | -------------------------- | --------------------------- |
| `/polish` | 최종 품질 개선: 정렬, 간격, 일관성 | 출시 전 최종 패스 |
| `/quieter` | 시각적 자극의 강도 감소 | 디자인이 너무 화려하고 시끄럽다 |
| `/distill` | 불필요한 복잡성 최소화 및 제거 | 페이지에 요소가 너무 많아 요소를 줄이고 싶습니다 |
| `/normalize` | 디자인 시스템 표준 정렬(토큰, 간격, 색상) | 스타일이 DESIGN.md 사양에서 벗어남 |
| `/clarify` | UX 카피라이팅, 오류 메시지, 라벨 문구 개선 | 카피라이팅은 혼란스럽고 오류 메시지는 불친절합니다 |
### 설계 검토 및 검증
온라인에 접속하기 전에 체계적인 검사를 통해 문제를 찾아내고 점수를 매기고 수정합니다.
| 스킬 | 해야 할 일 | 사용 시기 |
| --------------------- | --------------------------------- | ----------------------------- |
| `/plan-design-review` | 구현 전 설계 계획 검토(0\~10점) | AI가 디자이너의 관점에서 계획을 검토하기를 원합니다 |
| `/design-review` | 구현 후 Visual QA, 자동 스크린샷 비교 및 수정 | 코드 작성 후 시각적인 복원 정도 확인 |
| `/critique` | UX 평가: 시각적 계층, 인지 부하, 정서적 공명 | 구조화된 디자인 검토 보고서를 원하십니까 |
| `/audit` | 기술 검토: 접근성, 성능, 테마, 반응성 | 라이브 전 체계적인 점검 |
| `/benchmark` | 성능 기준 테스트, 비교 전/후 | 변경이 성능에 미치는 영향을 정량화하고 싶습니까? |
## 실제 시연: 카운트다운 기념일 페이지를 사용하여 전체 과정을 진행하세요.
분류표만 보면 너무 추상적이다. 우리는 위의 기술을 결합하기 위해 작은 프로젝트를 사용합니다. **카운트다운/기념일 단일 페이지** 만들기: 의미 있는 날짜를 선택하고 디지털 애니메이션과 배경 효과가 포함된 카운트다운 디스플레이를 만듭니다.
이 프로젝트는 작지만 완벽하여 6가지 기술 수준의 대부분을 다룰 수 있을 만큼 충분합니다. 전체 프로세스는 다음과 같습니다.
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> 매번 7단계를 모두 실행할 필요는 없습니다. 익숙해지면 일반적으로 사용되는 링크는 `/frontend-design → /animate → /polish → /ship` 네 단계에 불과합니다. 완전한 능력을 보여주기 위해 여기에서 모든 단계를 밟습니다.
### 1단계: 인프라 - 설계 언어 결정
**Skill**:`/design-consultation` + `/teach-impeccable`
프로젝트 초기에 한 번만 수행하면 됩니다. AI가 사용자의 디자인 선호도를 기억할 수 있도록 `DESIGN.md`을 출력합니다. 프로젝트에 이미 `DESIGN.md`이 있는 경우 직접 건너뜁니다.
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* TODO: 스크린샷 — 디자인 컨설팅으로 생성된 DESIGN.md 조각 */}
### 2단계: 탐색 - 여러 옵션 비교
**Skill**:`/design-shotgun`
방향이 확실하지 않은 경우 AI가 3\~5개의 시각적 솔루션을 생성하고 비교 패널을 열어 선택할 수 있습니다.
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* TODO: 스크린샷 — design-shotgun으로 생성된 솔루션 비교 패널 3개 */}
3가지 옵션 중에서 방향을 선택하세요. 원하는 것이 무엇인지 정확히 알고 있다면 이 단계를 건너뛰고 바로 3단계로 이동하세요.
### 3단계: 빌드 - 프로덕션 수준 코드 생성
**Skill**:`/frontend-design` + `/adapt`
핵심 링크. 처음부터 응답성을 보장하면서 코드를 작성합니다.
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* TODO: 스크린샷 - 구성이 완료된 후 데스크톱 페이지 효과 */}
{/* TODO: 스크린샷 - 모바일 버전 효과(/adapt 적응 후) */}
### 4단계: 강화 – 움직임과 개성 주입
**스킬**: `/animate` → `/delight` (요청 시 `/overdrive`)
이 세 가지는 점진적인 관계에 있습니다. `animate`은 기본 동적 효과, `delight`은 감정적 세부 사항, `overdrive`은 폭발 효과입니다. 필요에 따라 레이어를 추가합니다.
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
`DESIGN.md`에 참조된 애니메이션 제약 조건에 유의하세요. 디자인 시스템이 150ms 호버 전환만 허용하는 경우 `/overdrive`은 적용되지 않습니다. 이것은 좋은 판단 연습입니다.
{/* TODO: 스크린샷 또는 GIF - 모션 향상 전과 후 */}
### 5단계: 튜닝 - 융합 및 연마
**스킬**: `/typeset` + `/polish` (요청 시 `/distill`, `/normalize`)
간격 정렬, 글꼴 계층, 시각적 리듬. 너무 많이 추가했다고 생각되면 `/distill`을(를) 사용하여 빼세요.
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* 할 일: 스크린샷 — 광택 전과 후의 자세한 비교 */}
### 6단계: 검토 – 체계적인 점검
**Skill**:`/design-review` + `/audit`
시각적 QA + 기술 검토. `/design-review`은 자동으로 스크린샷을 찍어 문제를 비교하고 수정하며, `/audit`은 접근성과 성능을 확인합니다.
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* TODO: 스크린샷 — 감사를 통해 생성된 채점 보고서 */}
### 7단계: 출시
**Skill**:`/ship`
표준 gstack 릴리스 프로세스 - 테스트, 차이점 검토, PR 생성.
```text
> /ship
```
***
**예상 결과**: 프로세스에서 8\~10개의 프런트엔드 기술을 사용하는 시각적으로 아름다운 카운트다운 페이지. 무엇보다 중요한 것은 '어떤 단계에서 어떤 스킬을 사용하는가'에 대한 직관을 확립하는 것입니다.
## 일일 치트 시트
이상이 전체 과정입니다. 일상적인 개발 중에 특정 문제가 발생하는 경우 다음 표를 확인하세요.
| 내 현재 질문 | 무엇을 사용할 것인가 |
| -------------------------- | ---------------------------------- |
| 어떤 스타일을 원하는지 모르겠어요 | `/design-shotgun` |
| 페이지 기능은 좋지만 "거의 쓸모없다"는 느낌 | `/polish` |
| 뭔가 문제가 있는 것 같지만 설명할 수 없습니다 | `/design-review` |
| 글꼴/레이아웃이 이상해 보입니다 | `/typeset` |
| 지저분한 간격과 혼잡한 레이아웃 | `/arrange` |
| 스타일이 디자인 시스템에서 벗어남 | `/normalize` |
| 공용 구성 요소를 추출하고 싶습니까 | `/extract` |
| 페이지가 너무 복잡해서 빼고 싶어요 | `/distill` |
| 디자인이 너무 평범하고 너무 안전하다 | `/bolder` 또는 `/colorize` |
| 디자인이 너무 화려하고 시끄럽다 | `/quieter` |
| 오류 메시지 텍스트가 친절하지 않습니다 | `/clarify` |
| 휴대폰 디스플레이에 문제가 있습니다 | `/adapt` |
| 애니메이션 효과를 추가하고 싶습니까 | `/animate`(기본) 또는 `/overdrive`(폭발) |
| 온라인 접속 전 체계적인 점검 | `/audit` |
| debug | `/investigate` |
## 요약
이 메모는 두 가지 작업을 수행합니다.
1. **파노라마** - gstack의 27개 프론트엔드 스킬은 6개 레이어(인프라→탐색→구현→개선→최적화→검토)로 분류됩니다.
2. **실용적 시연** - 카운트다운 기념일 페이지를 사용하여 전체 워크플로를 살펴보고 각 단계에서 사용되는 스킬과 그 이유를 보여줍니다.
주요 내용: 이러한 스킬의 가장 강력한 사용법은 개별적으로 호출하는 것이 아니라 파이프라인에서 이를 결합하는 것입니다. 즉, 각 단계에서 명확한 스킬 선택을 통해 방향 탐색, 구현 구축, 개선 강화, 릴리스 검토 등이 가능합니다.
하지만 프로세스에 얽매이지 마세요. 일단 능숙해지면 대부분의 경우 `/frontend-design → /animate → /polish → /ship` 4단계이면 충분합니다.
***
**관련 자료**:
* [gstack 개념](/ko/docs/notes/gstack/concept) — gstack은 무엇이고 어떤 문제를 해결하나요?
* [gstack 실무 장](/ko/docs/notes/gstack/practice) — 설치부터 실행까지 전체 워크플로
* [gstack Skill Architecture Teardown](/ko/docs/notes/gstack/skill-architecture) — 스킬 개발자는 무엇을 배울 수 있나요?
* [Claude Skills Concept](/ko/docs/notes/claude-skills/concept) — 스킬의 기본 메커니즘을 이해합니다.
# gstack 실습: 설치부터 실행까지 전체 워크플로
## 소개
[콘셉트](/ko/docs/notes/gstack/concept)에서는 Claude Code를 가상 엔지니어링 팀으로 바꾸는 역할 기반 기술 세트인 gstack의 핵심 포지셔닝과 GSD, Superpowers, Ralph 및 기타 솔루션과 비교하여 AI 프로그래밍 도구 생태계에서의 차별화된 포지셔닝에 대해 알아봤습니다.
이 실용적인 기사는 설치 및 구성부터 전체 워크플로 실행까지 **사용 방법**에 중점을 두어 30분 안에 gstack을 시작할 수 있도록 도와줍니다.
## 설치 및 구성
### 전제조건
* **클로드 코드**가 설치되어 사용 가능합니다.
* **Git** 설치됨
* **Bun v1.0+** 설치됨(gstack은 Bun에 구축됨)
* Windows 사용자에게는 Node.js도 필요합니다.
### 전역 설치(권장, 30초 안에 완료)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
설치 스크립트는 다음 세 가지 작업을 수행합니다.
1. `CLAUDE.md` 파일에 gstack의 스킬 정보를 추가하세요.
2. 모든 스킬 파일을 스킬 디렉터리에 넣습니다.
3. Playwright 및 해당 Chromium 브라우저(`/browse` 및 `/qa`용)를 설치합니다.
### 프로젝트 수준 설치(팀 공유)
저장소를 복제한 후 팀 구성원이 자동으로 gstack을 얻도록 하려면 다음을 수행하세요.
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
\###다중 에이전트 지원
gstack은 Claude Code에만 국한되지 않고 현재 **10개의 AI 프로그래밍 에이전트**를 지원합니다. `./setup`은(는) 기본적으로 설치된 호스트를 자동으로 감지합니다.
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
각 호스트의 스킬 설치 경로는 `~/./skills/gstack-*/` 형태로 되어 있어 서로 간섭하지 않습니다.
> 💡 **OpenClaw 사용자를 위한 추가 옵션**: OpenClaw는 ACP를 통한 호출 외에도 ClawHub를 통해 4가지 기본 방법론 기술(`gstack-openclaw-office-hours`, `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`, `gstack-openclaw-retro`)을 직접 설치할 수 있으며, 이는 Claude Code 세션 없이 대화식으로 사용할 수 있습니다.
### 팀 모드(팀 공유 + 자동 업데이트, 권장)
v1.x에는 팀 모드가 도입되었습니다. 각 개발자는 gstack을 전역적으로 설치하고 웨어하우스는 "우리는 gstack을 사용합니다"만 기록하고 업데이트는 자동으로 발생합니다.
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
`required`을 `optional`로 바꾸는 것은 필수가 아닌 "부드러운 알림"입니다. Claude Code를 시작할 때마다 자동으로 업데이트 확인이 실행됩니다(시간당 한 번씩 조절되며, 네트워크에 장애가 발생하면 안전하고 조용합니다). 웨어하우스에는 판매된 파일이 없으며 버전 드리프트도 없습니다.
### 업데이트
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
또는 Claude Code에서 직접 `/gstack-upgrade`을 사용하세요.
## 전체 명령 참조
### 스프린트 프로세스
| 명령 | 역할 | 설명 |
| --------------------- | --------- | ----------------------------------------------------------------------------------------------- |
| `/office-hours` | YC 근무시간 | 제품 방향을 재구성하고 디자인 문서를 생성하기 위한 6가지 강제 질문 |
| `/plan-ceo-review` | CEO/설립자 | 4개 제품군 모델로 제공되는 10성급 제품을 찾고 있습니다 |
| `/plan-eng-review` | 엔지니어링 관리자 | 잠금 아키텍처, 데이터 흐름, 엣지 케이스, 테스트 매트릭스 |
| `/plan-design-review` | 수석 디자이너 | 디자인 차원 0-10 점수, 10점 달성 방법 설명 |
| `/plan-devex-review` | 개발자 경험 리더 | 개발자 초상화, 벤치마크 TTHW 및 디자인 마법의 순간을 살펴보세요. 세 가지 모드(DX EXPANSION / POLISH / TRIAGE), 20\~45개의 강제 질문 |
| `/autoplan` | 파이프라인 검토 | CEO → 디자인 → 엔지니어링 → DX 검토를 순차적으로 자동 실행하여 코딩 의사결정 원칙에 따라 자동 결정하고 "취향 결정"만 던집니다 |
### 디자인
| 명령 | 설명 |
| ---------------------- | ----------------------------------------------- |
| `/design-consultation` | 완전한 디자인 시스템을 처음부터 구축하고 DESIGN.md 생성 |
| `/design-shotgun` | 여러 AI 디자인 변형을 생성하고 브라우저에서 선택 항목 비교 |
| `/design-html` | 프로덕션급 HTML/CSS 생성, React/Svelte/Vue 프레임워크 감지 지원 |
### 검토 및 보안
| 명령 | 역할 | 설명 |
| ---------------- | ----------------- | -------------------------------------------------------------------------------------- |
| `/review` | 직원 엔지니어 | CI를 통과할 수 있지만 프로덕션 환경에서 폭발할 버그를 찾고, 명백한 문제를 자동으로 수정하고, 무결성 격차를 표시합니다 |
| `/investigate` | 디버깅 전문가 | 체계적인 근본 원인 디버깅. 철칙: 근본 원인을 찾을 때까지 버그를 수정하지 마세요. 3번의 수정 실패 후 중지 |
| `/design-review` | 코드를 작성할 수 있는 디자이너 | 시각적 감사 + 자동 복구, 원자 제출, 전후 비교 스크린샷 |
| `/devex-review` | DX 테스터 | 실제 온보딩 실행: 문서 찾아보기, 입력 프로세스 실행, TTHW 타이밍, 스크린샷 오류, `/plan-devex-review` 점수와 비교 |
| `/cso` | 보안 담당자 | OWASP 상위 10개 + STRIDE 위협 모델링, 17개의 거짓 긍정 제외 규칙, 8/10 신뢰도 임계값, 각 결과에는 특정 활용 시나리오가 수반됩니다 |
### 테스트 및 QA
| 명령 | 설명 |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/qa` | 실제 브라우저 테스트를 열고 버그 찾기 → Atomic 커밋 수정 → 회귀 테스트 생성 → 재검증 |
| `/qa-only` | 위와 동일하지만 보고만 가능하고 코드 수정은 없습니다 |
| `/benchmark` | 기준 성능 테스트: 페이지 로딩, 코어 웹 바이탈, 리소스 크기, 지원 전후 비교 |
| `/browse` | \~100ms 수준의 브라우저 명령, 실제 Chromium, 스크린샷, 양식 채우기, 요소 클릭 |
| `/open-gstack-browser` | GStack 브라우저 시작: 가시적 AI 제어 Chromium, 사이드바 확장, 크롤링 방지 스텔스, 자동 모델 라우팅(Sonnet 작업/Opus 분석), 원클릭 쿠키 가져오기 지원 |
| `/setup-browser-cookies` | 로그인이 필요한 페이지를 테스트하기 위해 실제 브라우저(Chrome/Arc/Brave/Edge)에서 헤드리스 세션으로 쿠키를 가져옵니다. |
| `/pair-agent` | 교차 AI 에이전트 브라우저 페어링: 동일한 GStack 브라우저를 OpenClaw/Hermes/Codex/Cursor 등에 공유합니다. 각 에이전트에는 독립적인 탭이 있으며 원격 에이전트를 지원하기 위한 ngrok 터널, 범위 토큰 + 탭 격리 + 속도 제한 + 동작 속성 |
### 출시 및 운영 및 유지보수
| 명령 | 설명 |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------- |
| `/ship` | 메인 브랜치 동기화 → 테스트 실행 → 적용 범위 감사 → 버전 업데이트 → 푸시 제출 → PR 생성; 프로젝트에 테스트 프레임워크가 없을 때 자동 부트스트랩 |
| `/land-and-deploy` | PR 병합 → CI 대기 → 배포 → 프로덕션 환경 상태 확인 |
| `/canary` | 배포 후 카나리아 모니터링: 콘솔 오류, 성능 회귀, 페이지 오류 |
| `/setup-deploy` | `/land-and-deploy` 일회성 구성: 자동 감지 플랫폼(Fly.io/Render/Vercel/Netlify/Heroku/GitHub Actions/custom) + 프로덕션 URL + 배포 명령 |
| `/setup-gbrain` | 한 번의 클릭으로(5분 이내) GBrain 데이터베이스 시작: PGLite 로컬, Supabase 기존 URL 또는 관리 API를 통해 자동으로 새 Supabase 프로젝트 생성 MCP 등록 + 웨어하우스 수준 읽기-쓰기/읽기 전용/거부 권한 |
### 검토 및 학습
| 명령 | 설명 |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| `/retro` | 팀 인식 주간 보고서: 1인당 분석, 연승 통계, 테스트 상태 추세, 성장 기회; 모든 프로젝트에 걸쳐 `/retro global` + AI 도구(Claude Code / Codex / Gemini) |
| `/document-release` | 게시된 코드(README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS)와 일치하도록 프로젝트 문서를 자동으로 업데이트합니다. `/ship`은 이제 자동으로 호출됩니다 |
| `/learn` | 세션 간 학습 추억 관리: 프로젝트별 보기, 검색, 정리, 내보내기, 누적 |
| `/context-save` `/context-restore` | 연속 체크포인트 모드 패키지: 컨텍스트를 저장하기 위한 자동 WIP 커밋, `/context-restore`를 사용하여 충돌/전환 후 세션을 다시 작성 |
### 보안 보호
| 명령 | 설명 |
| ----------------------- | ------------------------------------------- |
| `/careful` | 위험한 작업 경고: rm -rf, DROP TABLE, force-push 등 |
| `/freeze` / `/unfreeze` | 특정 디렉토리에 대한 편집 범위 잠금/잠금 해제 |
| `/guard` | `/careful` + `/freeze` 조합, 최고 보안 모드 |
| `/checkpoint` | 작업상태 스냅샷 저장/복원 |
### 도구 통합
| 명령 | 설명 |
| -------------------------------------------------- | ------------------------------------------------------------------------------ |
| `/codex` | OpenAI Codex CLI 통합: 독립적인 코드 검토(통과/실패 게이트), 대결 모드, 상담 모드 모델 간 중복 분석은 `/review` |
| `/health` | 코드 품질 대시보드: tsc + biome + knip + shellcheck + 테스트 → 0-10 전체 점수 |
| `/skillify` | 현재 워크플로우를 재사용 가능한 기술로 통합 |
| `/scrape` | 웹 스크래핑 작업흐름 |
| `/landing-report` | 랜딩페이지 실적 및 경험 보고서 |
| `/make-pdf` | PDF 문서 생성 |
| `/benchmark-models` `/model-overlays` `/plan-tune` | 모델 간 비교, 적용 범위 오버레이, 계획 최적화 |
### Standalone CLI(v0.19+)
슬래시 명령 외에도 gstack에는 독립형 CLI 세트도 함께 제공됩니다(Claude Code 세션 내에서는 실행되지 않음).
| 명령 | 설명 |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| `gstack-model-benchmark` | 모델 간 평가: 동일한 프롬프트에서 Claude / GPT(Codex CLI를 통해) / Gemini를 실행하고 지연, 토큰, 비용 및 (선택 사항) LLM 심사 품질 점수를 비교합니다. 사용할 수 없는 공급자는 자동으로 건너뜁니다. |
| `gstack-taste-update` | 디자인 취향 학습: `/design-shotgun`의 승인/비승인을 프로젝트 수준 취향 파일에 기록하고 매주 5%씩 감소하며 후속 변형 생성에 피드백 |
## 구성 세부정보
### CLAUDE.md 콘텐츠 추가
설치 후 gstack은 `CLAUDE.md`에 사용 가능한 모든 기술의 목록과 간단한 설명을 추가합니다. 이를 통해 Claude Code는 어떤 명령을 사용할 수 있는지 알 수 있습니다.
### 스킬 디렉토리 구조
주요 입구는 최상위 `~/.claude/skills/gstack/SKILL.md`이고, 각 하위 명령은 플랫 디렉터리 형태로 존재하며, 핵심은 `SKILL.md` 파일입니다.
```text
~/.claude/skills/gstack/
├── SKILL.md # 主入口 skill
├── browse/ # 浏览器 daemon
├── qa/ # QA 测试
├── review/ # 代码审查
├── ship/ # 发布流程
├── plan-ceo-review/ # CEO 审查
├── office-hours/ # 产品门诊
├── pair-agent/ # 跨 Agent 浏览器配对
├── open-gstack-browser/ # GStack Browser 启动器
├── setup-gbrain/ # GBrain 数据库一键上手
├── hosts/ # 10 个 host 配置(claude/codex/cursor/...)
├── bin/ # standalone CLI(gstack-model-benchmark 等)
└── ... # 当前 v1.x 共 50 个 skill 目录
```
동작을 사용자 정의하기 위해 `SKILL.md`을 자유롭게 수정할 수 있습니다. 이것이 "포크 및 사용자 정의"의 장점입니다.
### Browse Daemon
Browse Daemon은 영구 Chromium 인스턴스입니다. 주요 구성:
* **포트**: 무작위로 선택되는 10000-60000, 10개 이상의 병렬 작업 공간 지원
* **보안**: 로컬 호스트만 바인딩하고 각 세션마다 베어러 토큰 인증을 사용합니다.
* **쿠키**: Chrome/Arc/Brave/Edge에서 가져오려면 `/setup-browser-cookies`을 사용하세요.
## 실제 작업 흐름 시연
다음은 일반적인 gstack 작업 흐름을 보여줍니다. 명령과 출력은 문서와 비디오의 실제 사례를 기반으로 합니다.
> 💡 **참고**: 다음 출력은 연구를 기반으로 편집된 일반적인 예입니다. 향후 실제 사례를 바탕으로 특정 프로젝트의 스크린샷이 추가될 예정입니다.
### 1단계: 제품 클리닉
```text
> /office-hours
[YC Office Hours] 6 forcing questions:
1. Who specifically needs this?
2. What do they do today without it?
3. Why is this urgent right now?
4. How will you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can ship?
→ Design doc generated
```
서두르지 말고 먼저 YC Office Hours의 관점에서 AI가 아이디어를 고문하도록 하세요.
### 2단계: 다중 역할 검토 계획
```text
> /autoplan
[CEO Review] Finding the 10-star product...
[Design Review] Rating dimensions 0-10...
[Eng Review] Locking architecture + edge cases...
→ Fully reviewed plan ready
```
`/autoplan`은 CEO → 설계 → 엔지니어링 검토의 3단계를 자동으로 실행하여 완전한 사후 검토 계획을 생성합니다.
### 3단계: 코딩 구현
일반적으로 승인된 계획에 따라 코딩합니다. 표준 Claude Code 대화를 사용할 수 있습니다.
### 4단계: 여러 전문가의 코드 검토
```text
> /review
Dispatching 7 specialist reviewers...
- Testing coverage ✓
- Maintainability ✓
- Security: Found 1 issue (auto-fixing)
- Performance ✓
- Data migration ✓
- API contract ✓
- Red team: No vulnerabilities found
→ Review complete, 1 auto-fix applied
```
### 5단계: 브라우저 QA
```text
> /qa
Opening headless browser...
Testing user flows:
- Login flow ✓
- Dashboard load ✓
- Form submission: Bug found → fixing → re-testing ✓
- Image upload ✓
→ 4 flows tested, 1 bug fixed, regression test generated
```
### 6단계: 게시
```text
> /ship
Syncing with main...
Running tests: 42 passed, 0 failed
Reviewing diff: 3 files changed
Updating VERSION: 1.2.0 → 1.3.0
Creating PR: "Add screenshot feature"
→ PR #47 created, ready for merge
```
## 실용적인 팁과 커뮤니티 경험
### Garry Tan의 제안
gstack의 ETHOS.md, 세 가지 핵심 원칙:
1. **Boil the Lake**: AI는 완전성을 거의 무료로 만듭니다. 항상 완전한 작업을 수행하고 지름길을 택하지 않습니다.
2. **구축 전 검색**: 먼저 검색하고, 먼저 이해하고, 3단계 지식 검증 후 시작
3. **사용자 주권**: AI 추천은 귀하가 결정합니다. 두 AI 모델이 모두 동의하더라도 여전히 귀하의 판단이 우선합니다.
gstack의 README는 Karpathy의 인용문으로 시작됩니다. 이는 Garry Tan이 gstack을 구축하려는 이유를 설명하는 출발점이기도 합니다.
### 긍정적인 커뮤니티 경험
* **`/office-hours`(YC 지원)**: Reddit r/ycombinator의 여러 S26 지원자는 gstack의 업무 시간을 사용하여 지원 자료를 스트레스 테스트하는 것이 매우 효과적이라고 보고했습니다.
* **보안 감사를 통해 실제 취약점이 발견되었습니다**: CTO 피드백이 있었습니다. `/review`은(는) 팀이 인식하지 못한 XSS 취약점을 발견했습니다.
* **`/browse` 실제 브라우저 테스트**: 커뮤니티(비평가 포함)에서 "진정한 기술적 기여"로 인정
### 일반적인 함정
* **잦은 권한 메시지**: 일부 사용자는 "권한 메시지는 30초마다 승인을 받아야 하기 때문에 잠을 잘 수 없다"고 보고했습니다. Claude Code 설정에서 적절한 자동 승인 규칙을 구성하는 것이 좋습니다.
* **높은 토큰 소비**: 특성화된 프롬프트는 컨텍스트 소비를 증가시킵니다. 비용에 민감한 경우 가장 필요한 기술을 선택적으로 사용할 수 있습니다.
* **에이전트 루프**: HN에는 에이전트가 70분 루프에 갇혔다고 사용자가 보고한 사례가 있습니다. 합리적인 시간 초과 및 체크포인트를 설정하는 것이 좋습니다.
* **모든 사람에게 해당되는 것은 아닙니다**: 숙련된 개발자는 대부분의 기술이 불필요한 래퍼라고 느낄 수 있습니다. gstack은 성숙한 엔지니어링 프로세스를 갖춘 팀보다는 **독립적인 창립자 및 소규모 팀**에 더 적합합니다.
## 자주 묻는 질문(FAQ) 및 모범 사례
\*\*Q: gstack과 Superpowers를 동시에 사용할 수 있나요? \*\*
그렇습니다. 둘은 서로를 보완합니다. Superpowers는 프로세스 규율과 TDD 보증에 능숙하고 gstack은 제품 사고와 다중 역할 검토에 능숙합니다. 많은 팀이 일일 코딩 훈련에 Superpowers를 사용하고 제품 계획 및 QA에 gstack을 사용합니다.
\*\*Q: 토큰은 비싼가요? \*\*
기본 Claude Code보다 높습니다. 각 기술의 역할 프롬프트는 상황에 맞는 창을 차지합니다. 그러나 귀하의 시간이 토큰 수수료보다 더 가치가 있다면 이는 일반적으로 좋은 거래입니다.
\*\*Q: 어떤 유형의 프로젝트에 적합합니까? \*\*
아이디어부터 출시까지 **전체 프로세스 제품 개발**에 가장 적합합니다. 단순히 버그를 수정하거나 작은 기능을 만드는 정도라면 네이티브 클로드 코드만으로도 충분합니다. gstack의 가치는 "완전한 프로세스"에서 극대화됩니다.
\*\*Q: 스킬을 어떻게 맞춤설정하나요? \*\*
각 스킬은 `SKILL.md` 파일입니다. 직접 편집하세요.
1. 스킬 디렉토리 찾기: `~/.claude/skills/gstack//`
2. `SKILL.md` 편집
3. `./setup` 다시 실행
커뮤니티에서는 전역 설치를 직접 수정하는 대신 저장소를 포크하고 사용자 정의하는 것을 권장합니다.
### 모범 사례
1. **먼저 `/office-hours` 후 코드**: 코드를 작성하기 전에 제품 클리닉을 하는 습관을 들입니다.
2. **`/browse` 확인을 잘 활용하세요**: 코드만 보지 말고 AI가 애플리케이션을 실제로 "볼" 수 있도록 하세요.
3. **주기적 `/retro`**: 코드 품질 및 작업 속도에 대한 가시성을 유지합니다.
4. **점진적 채택**: 모든 기술을 한 번에 사용할 필요가 없습니다. `/office-hours` + `/review` + `/ship`부터 시작
5. **포크 사용자 정의**: 부적절한 프롬프트가 표시되면 직접 변경하세요. 이것이 오픈소스의 장점이다
## 요약
gstack의 핵심 가치는 특정 기술이 얼마나 강력한지에 있는 것이 아니라 **구조화된 AI 협업 모드**를 제공한다는 점입니다. 역할 전환을 통해 다양한 단계에서 다양한 유형의 AI 지원을 받을 수 있습니다. 먼저 CEO 관점에서 제품 방향을 검토하고, 엔지니어링 관리자의 엄밀한 아키텍처 검토를 거쳐 최종적으로 QA의 실제 브라우저를 통해 결과를 검증합니다.
다음으로 직접 설치해 보고 `/office-hours`에서 첫 번째 gstack 프로젝트를 시작할 수 있습니다.
***
**자세한 내용**:
* [gstack 개념](/ko/docs/notes/gstack/concept) — gstack의 핵심 개념과 도구 생태학적 포지셔닝을 이해합니다.
* [GSD 실무장](/ko/docs/notes/gsd/practice) — 또 다른 구조화된 AI 프로그래밍 솔루션에 대한 실무 가이드
* [클로드 스킬실습편](/ko/docs/notes/claude-skills/skill-creator) — 스킬 생성 메커니즘을 이해한다
# Teardown gstack: 개발자가 배울 수 있는 기술
## 소개
[개념](/ko/docs/notes/gstack/concept)과 [실용](/ko/docs/notes/gstack/practice)에서는 gstack이 무엇인지, 사용자 관점에서 어떻게 사용하는지 알아보았습니다. 이 노트는 다른 관점에서 작성되었습니다 - **기술 개발자로서**, gstack Warehouse 파일을 파일별로 읽은 후 어떤 엔지니어링 설계에서 배우고 배울 가치가 있는지 알아보십시오.
gstack은 단순한 23개의 프롬프트 파일 모음 그 이상입니다. 그 뒤에는 템플릿 생성, 자동 업그레이드, 학습 및 메모리, 점진적인 지침, 다중 플랫폼 적응, 계층화된 테스트 등 완전한 엔지니어링 시스템이 있습니다. 이는 기술 프로젝트를 "사용 가능"에서 "사용하기 쉬운" 것으로 바꾸는 열쇠입니다.
***
## 1. SKILL.md는 손으로 쓰지 않습니다 - 템플릿 생성 시스템
gstack의 가장 반직관적인 디자인: \*\*각 SKILL.md는 자동으로 생성되며 직접 편집할 수 없습니다. \*\*
```text
SKILL.md.tmpl (人写) → gen-skill-docs → SKILL.md (机生)
```
사람이 작성한 `.tmpl` 템플릿에는 워크플로 논리와 모범 사례, `{{PLACEHOLDER}}` 자리 표시자가 포함되어 있습니다. 빌드 스크립트는 소스 코드에서 명령 참조, 브라우저 플래그 목록, 서문 시작 코드 등을 추출하고 이를 자리 표시자에 채워 최종 SKILL.md를 생성합니다.
```text
{{PREAMBLE}} ← 从 resolvers/preamble.ts 生成的启动代码
{{BROWSE_SETUP}} ← 浏览器初始化指令
{{COMMAND_REFERENCE}} ← 从 commands.ts 提取的命令文档
{{SNAPSHOT_FLAGS}} ← 从源代码常量提取的快照选项
```
\*\*왜 이러는 걸까요? \*\*
* 문서와 코드는 절대로 동기화되지 않습니다. 명령 참조는 소스 코드에서 생성되며 소스 코드가 변경되면 문서가 자동으로 업데이트됩니다.
* 23개의 스킬이 동일한 서문(약 220라인)을 공유하며, 모든 스킬이 동시에 업데이트됩니다.
* CI는 재생성하는 것을 잊지 않도록 생성된 파일이 만료되었는지 여부를 `--dry-run` 확인할 수 있습니다.
**요점**: 여러 기술을 유지 관리하는 경우 기술 간에 공유되는 모든 콘텐츠를 템플릿으로 추출하고 빌드 단계에서 사용하여 최종 파일을 생성해야 합니다. 동일한 콘텐츠의 여러 복사본을 수동으로 동기화하면 조만간 문제가 발생할 수 있습니다.
***
## 2. 업그레이드 메커니즘 - 탐지부터 실행까지 완전한 연결
gstack의 업그레이드 시스템은 매우 정교하게 설계되었으며 세 가지 계층으로 나뉩니다.
### 첫 번째 레이어: 버전 감지
`bin/gstack-update-check`은 다음을 수행하는 독립형 bash 스크립트입니다.
1. 로컬 `VERSION` 파일을 읽습니다.
2. 캐시 `~/.gstack/last-update-check` 확인(60분 동안 UP\_TO\_DATE 캐시, 720분 동안 UPGRADE\_AVAILABLE 캐시)
3. 캐시가 만료되면 HTTP는 GitHub의 `raw.githubusercontent.com/.../VERSION`을 요청합니다.
4. 버전 번호를 비교하고 `UPGRADE_AVAILABLE <旧> <新>`을 출력합니다.
### 두 번째 레이어: 프리앰블 통합
**각 스킬의 SKILL.md 시작 코드의 첫 번째 줄은 버전 감지입니다**:
```bash
_UPD=$(~/.claude/skills/gstack/bin/gstack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
```
즉, 사용자가 스킬을 호출하면 업데이트가 자동으로 감지됩니다. 특별히 업그레이드 명령을 실행할 필요가 없으며 존재감은 없지만 적용 범위는 100%입니다.
### 세 번째 수준: 점진적 알림 + 자동 업그레이드
새 버전을 감지한 후 사용자를 즉시 방해하지 않지만 점진적인 백오프를 위해 다시 알림 메커니즘(Snooze)을 사용합니다.
* 1차 알림: 24시간 후에 다시 언급해 주세요.
* 2차 알림: 48시간 후에 다시 언급해 주세요.
* 3회 이상 : 7일 이후 다시 언급해주세요.
* 새 버전 릴리스에서는 스누즈 카운터가 재설정됩니다.
사용자는 `gstack-config set auto_upgrade true` 자동 업그레이드를 활성화하고 확인을 건너뛰어 직접 실행할 수 있습니다.
업그레이드를 수행하면 5가지 설치 유형(글로벌 git, 로컬 git, 공급업체 등)이 구분됩니다. git 설치는 `git fetch + reset`을 사용하고, 공급업체 설치는 먼저 백업한 다음 교체하고, 실패할 경우 `.bak`에서 복원합니다. 업그레이드 후에는 프로젝트의 로컬 공급업체 복사본도 자동으로 동기화됩니다.
**배울만한 가치가 있는 점**:
* "각 통화 감지" 모드는 적용 범위가 매우 높으며 사용자가 감지할 수 없습니다.
* 점진적인 백오프로 빈번한 중단 방지
* 설치 유형을 차별화하고 단일 크기 대신 다양한 업그레이드 전략을 구현합니다.
* 백업 및 복원은 업그레이드 실패로 인해 전체 스킬이 중단되지 않도록 보장합니다.
***
## 3. 학습 시스템 - 사용할수록 스킬이 스마트해집니다.
gstack은 가볍지만 효과적인 **교차 세션 메모리 시스템**을 구현합니다.
### 스토리지
각 프로젝트에는 추가로 작성된 독립적인 학습 로그(`~/.gstack/projects/$SLUG/learnings.jsonl`)가 있습니다.
```json
{
"skill": "review",
"type": "pitfall",
"key": "n-plus-one",
"insight": "这个项目的 User model 有 N+1 查询问题,findAll 要加 include",
"confidence": 8,
"source": "observed",
"files": ["src/models/user.ts"],
"ts": "2026-04-01T14:30:00Z"
}
```
### 자동수집
각 기술이 완료되기 전에 "운영 자체 개선" 링크가 있습니다. 실행 중에 발견된 예상치 못한 실패, 우회 또는 프로젝트 문제가 있는지 여부를 반영하고 모든 내용이 learnings.jsonl에 자동으로 기록됩니다. 사용자가 수동으로 트리거할 필요가 없습니다.
### 자동 로딩
새 세션이 시작될 때마다 프리앰블은 처음 3개의 신뢰도가 높은 학습 항목을 로드하여 컨텍스트를 주입하여 새 세션이 역사적 지식을 상속할 수 있도록 합니다.
### 신뢰도 하락
`observed` 및 `inferred` 소스의 항목은 30일마다 1포인트씩 감소합니다. 지식 기반을 수동으로 정리할 필요가 없습니다. 오래된 지식은 자연스럽게 사라지고 새로운 관찰이 자연스럽게 그 자리를 차지합니다.
### 관리 인터페이스
```text
/learn # 显示最近 20 条
/learn search # 搜索
/learn prune # 检测过期条目(引用的文件已删除)
/learn export # 导出为 markdown 可加入 CLAUDE.md
```
**배울만한 가치가 있는 점**:
* 추가 쓰기 전용 설계는 간단하고 안정적이며 동시성이 안전합니다.
* 신뢰도 하락은 유지 관리가 적은 지식 노화 관리입니다. 수동 청소보다 훨씬 효율적입니다.
* 경로 대신 git 원격 URL을 사용하여 프로젝트를 식별하고(`gstack-slug`을 통해) 다른 위치에 복제하고 재사용할 수 있습니다.
* 프로젝트 간 쿼리를 지원하지만 기본적으로 격리됨
***
## 4. 프리앰블 주입 - 스킬의 "미들웨어 레이어"
이것은 gstack의 가장 똑똑한 아키텍처 디자인 중 하나입니다. 각 SKILL.md는 웹 프레임워크의 미들웨어처럼 기능하는 약 220줄의 프리앰블 코드를 공유합니다.
```text
┌─ 更新检测 ──────────────────────────────────┐
│ 会话追踪 (sessions/$PPID) │
│ 配置读取 (proactive, skill_prefix, telemetry)│
│ 学习历史加载 (前 3 条高置信度) │
│ 上下文恢复 (最近的 checkpoint + timeline) │
│ 路由规则检测 │
│ 首次使用引导流程 │
└──────────────────────────────────────────────┘
↓
Skill 特有逻辑
```
Preamble의 bash 스크립트는 키-값 쌍(`BRANCH: main`, `PROACTIVE: true`)을 출력한 다음 템플릿은 자연어 조건을 사용하여 Claude가 그에 따라 행동을 조정할 수 있도록 합니다.
```text
If PROACTIVE is false, do not invoke skills automatically.
Instead suggest: "I think /skillname might help here -- want me to run it?"
```
이는 본질적으로 bash 출력을 Claude의 "환경 변수"로 처리하는 것입니다. 즉, 런타임 감지에는 bash를 사용하고 동작 라우팅에는 자연어를 사용합니다.
**배울 가치가 있는 점**: 여러 기술이 있는 경우 각 기술에 대한 복사본을 작성하는 대신 공유 논리(구성 로딩, 상태 복구, 버전 감지)를 통합된 서문으로 추출해야 합니다.
***
## 5. 프로그레시브 부팅 - Sentinel 파일 모드
gstack의 최초 사용자 경험은 세심하게 디자인되었습니다. 각 부팅 단계가 터치 파일(센티넬 파일)을 통해 한 번만 발생하는지 확인하세요.
```text
~/.gstack/.completeness-intro-seen ← "Boil the Lake" 理念介绍
~/.gstack/.telemetry-prompted ← 遥测选择(community/anonymous/off)
~/.gstack/.proactive-prompted ← 主动触发开关
~/.gstack/.routing-prompted ← CLAUDE.md 路由规则写入
~/.gstack/.welcome-seen ← 安装欢迎消息
```
스킬이 시작될 때마다 이러한 파일이 존재하는지 확인하십시오. 그렇지 않은 경우 해당 부팅 및 터치 파일을 표시합니다. 이미 본 단계는 다시 표시되지 않습니다.
**배울 가치가 있는 점**: 구성에서 `"onboarding_step": 3` 상태를 유지하는 것과 비교할 때 sentinel 파일은 더 간단하고 안정적입니다. 구성 파일 손상의 영향을 받지 않으며 각 단계는 독립적으로 제어됩니다.
***
## 6. SKILL.md 구조 설계 - 3계층 아키텍처
각 SKILL.md는 표준 3계층 구조를 따릅니다.
### 첫 번째 레이어: YAML Frontmatter
```yaml
---
name: qa
preamble-tier: 3
version: 0.15.1.0
description: |
Systematically QA test a web application...
Use when asked to "qa", "test this site", "find bugs"...
benefits-from: [office-hours]
allowed-tools:
- Bash
- Read
- Write
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "bash ${CLAUDE_SKILL_DIR}/bin/check-careful.sh"
---
```
주요 분야:
* `allowed-tools`: 도구 수준 권한 허용 목록, 각 기술은 필요한 도구만 선언합니다.
* `benefits-from`: 사전 종속성 기술을 명시적으로 선언합니다.
* `hooks`: 도구가 호출되기 전에 가로챌 수 있는 PreToolUse 후크(예: 신중한 가로채기 `rm -rf`)
* `description`: 모든 자연어 트리거 단어를 포함합니다.
### 레이어 2: 공유 프리앰블 + 일반 규칙
프리앰블 시작 코드 + 음성 정의 + 컨텍스트 복구 + 무결성 원칙 + 검색 우선 순위 + 완료 상태 프로토콜 + 업그레이드 규칙 등 모든 기술은 동일하며 템플릿에서 생성됩니다.
### 세 번째 레이어: 스킬별 로직
이는 워크플로 정의, 역할 설정, 인지 모델 주입, 상호 작용 게이팅 등 각 기술의 "영혼"입니다.
**배울 가치가 있는 점**: 3계층 분리를 통해 각 스킬은 고유한 로직에만 집중할 수 있으며, 공유된 부분은 일관성을 위해 프레임워크로 보장됩니다.
***
## 7. 신속한 엔지니어링 팁 컬렉션
SKILL.md를 모두 읽은 후 학습할 가치가 있는 즉각적인 디자인 기술은 다음과 같습니다.
### 아첨 금지 규칙
근무 시간의 시작 모드는 AI의 일반적인 "및 진흙탕" 동작을 명시적으로 금지합니다.
```text
Never say:
- "That's an interesting approach" → take a position instead
- "There are many ways to think about this" → pick one
- "You might want to consider..." → say "This is wrong because..."
- "That could work" → say whether it WILL work
```
### 금지어 목록
음성 섹션에는 명확한 금지 단어와 문구가 있습니다.
* 금지된 단어: 탐구, 결정적, 강력함, 포괄적, 미묘함, 중추적, 풍경...
* 금지된 문구: "여기에 핵심이 있습니다", "플롯 반전", "이것을 분석해 보겠습니다"...
* 비활성화된 형식: em 대시(쉼표/마침표로 대체)
이는 LLM에서 흔히 사용되는 "AI 스타일의" 단어이며, 비활성화된 후에 출력이 확실히 더 자연스러워집니다.
### 인지 모델 주입
각 검토 기술은 서로 다른 사고 체계를 주입합니다.
* **CEO 리뷰**: 18가지 인지 모델(베조스의 일방향/양방향 문 의사결정, 멍거의 역사고, 잡스의 집중과 뺄셈...)
* **Eng Review**: 15가지 엔지니어링 관리 패턴("기본적으로 지루함", 폭발 반경 직관, 콘웨이 법칙...)
* **디자인 검토**: 12가지 디자인 인식 패턴(서비스로서의 계층 구조, 제약 조건 숭배, "내가 눈치챌까?" 테스트...)
이러한 모드는 AI가 기계적으로 수행하도록 허용하지 않지만 똑똑한 신인에게 이전 AI의 경험 목록을 제공하는 것과 마찬가지로 AI에 **사고 프레임워크**를 제공합니다.
### 사양기준
```text
Not "you should test this"
but `bun test test/billing.test.ts`
Not "this might be slow"
but "this queries N+1, ~200ms per page load with 50 items"
Not "there's an issue in the auth flow"
but "auth.ts:47, the token check returns undefined"
```
### 신뢰도 교정
검토 기술을 사용하려면 각 발견에 신뢰도 점수가 수반되어야 하며 신뢰도가 낮은 결과는 자동으로 다운그레이드되거나 숨겨집니다.
| 점수 | 의미 | 처리 |
| ---- | ------------- | ------------ |
| 9-10 | 특정 코드를 읽고 확인 | 일반 표시 |
| 7-8 | 높은 신뢰도 패턴 매칭 | 일반 표시 |
| 5-6 | 보통, 허위 경보 가능성 | 지침과 함께 표시 |
| 3-4 | 낮은 신뢰도 | 보고서에서 숨기기 |
| 1-2 | 순수한 추측 | P0 수준에서만 표시됨 |
### 대화형 게이팅
선박 스킬은 사용자를 멈추고 언제 기다려야 하는지, 언제 자동으로 계속해야 하는지를 정확하게 정의합니다.
```text
Only stop for:
- Tests failing with no obvious fix
- Merge conflicts requiring human judgment
- Unclear which changes to include
Never stop for:
- Normal git operations
- CHANGELOG/VERSION updates
- PR creation
```
**배울 가치가 있는 점**: 좋은 기술은 "AI가 모든 것을 수행"하는 것이 아니라 인간-기계 경계를 정확하게 정의하는 것입니다.
***
## 8. 상태 관리 - 파일 시스템은 데이터베이스입니다.
gstack의 모든 지속성은 `~/.gstack/`에 저장된 파일 시스템을 통해 수행됩니다.
| 경로 | 목적 | 형식 |
| ------------------------------------- | -------- | ------ |
| `config.yaml` | 글로벌 구성 | YAML |
| `sessions/$PPID` | 활성 세션 | 터치 파일 |
| `projects/$SLUG/learnings.jsonl` | 학습기록 | JSONL |
| `projects/$SLUG/timeline.jsonl` | 스킬 타임라인 | JSONL |
| `projects/$SLUG/checkpoints/*.md` | 체크포인트 | 마크다운 |
| `projects/$SLUG/health-history.jsonl` | 건강검진 내역 | JSONL |
| `analytics/skill-usage.jsonl` | 원격 측정 사용 | JSONL |
| `last-update-check` | 버전 캐시 | 일반 텍스트 |
거의 모든 시계열 데이터는 **JSONL**(행당 하나의 JSON 개체)을 사용하여 추가 방식으로 작성됩니다. 이 선택은 현명합니다.
* 쓰기 자연 동시성 안전성 추가
* 데이터베이스 종속성이 필요하지 않습니다.
* `grep` / `jq`을 사용하여 직접 쿼리할 수 있습니다.
* 마지막 줄이 누락될 때까지 손상됨
***
## 9. 크로스 스킬 통합 모드
### 파일전송상품
파일 시스템을 통해 기술 간에 작업 산출물을 전송합니다.
```text
/office-hours → design doc → /plan-ceo-review 读取
/plan-ceo-review → ceo-plans/*.md → /autoplan 读取
/review → reviews.jsonl → /ship 读取并展示 Dashboard
/qa → qa-reports/ → /retro 读取
```
### Review Readiness Dashboard
배송 스킬은 `reviews.jsonl`을 읽고 게시 전 교차 스킬 검토 상태를 표시합니다.
```text
| Review | Runs | Last Run | Status | Required |
| Eng Review | 1 | 2026-03-16 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
```
### 종속성 사전 제안
plan-ceo-review가 디자인 문서가 없음을 감지하면 먼저 `/office-hours` 실행을 적극적으로 권장합니다.
```text
"No design doc found. /office-hours produces a structured problem statement...
Takes about 10 minutes."
Options: A) Run /office-hours now B) Skip
```
### 시퀀스 예측 사용
Context Recovery는 최근 스킬 사용 순서를 분석하고 다음 단계를 예측합니다.
```text
If pattern repeats (e.g., review → ship → review),
suggest: "Based on your recent pattern, you probably want /ship."
```
***
## 10. 그 외 주목할만한 디자인
### 후크 시스템
주의, 동결 및 보호의 세 가지 기술은 `PreToolUse` 후크를 사용합니다. 이는 도구가 호출되기 전에 가로챌 수 있는 유일한 메커니즘입니다.
* **주의**: Bash를 가로채고 `rm -rf`, `DROP TABLE`, `git push --force`를 확인하세요.
* **동결**: 편집/쓰기를 차단하고 경로가 허용 범위 내에 있는지 확인합니다.
* **가드**: 위의 두 가지를 결합합니다.
### 다중 플랫폼 적응
동일한 템플릿 세트는 `--host` 매개변수를 통해 다양한 플랫폼에 대한 기술 파일을 생성합니다.
```bash
bun run gen:skill-docs --host claude # Claude Code 格式
bun run gen:skill-docs --host codex # OpenAI Codex 格式
bun run gen:skill-docs --host kiro # AWS Kiro 格式
bun run gen:skill-docs --host factory # Factory Droid 格式
```
경로와 머리말은 자동으로 조정되며, 스킬 로직은 변경되지 않습니다.
### 완료 상태 프로토콜
각 기술이 끝나면 표준화된 완료 상태가 출력되어야 합니다.
```text
DONE — 全部完成,提供证据
DONE_WITH_CONCERNS — 完成但有顾虑
BLOCKED — 无法继续
NEEDS_CONTEXT — 需要更多信息
```
### 세 가지 업그레이드 규칙 실패
```text
If you have attempted a task 3 times without success, STOP and escalate.
```
AI가 무한 재시도 루프에 빠지는 것을 방지합니다.
### Diff 기반 테스트 선택
E2E 테스트 비용은 각각 약 $4입니다(Claude 에이전트 시작 필요). 따라서 gstack은 `touchfiles.ts`을 통해 각 테스트가 의존하는 소스 파일을 선언하고 `git diff`에 따라 영향을 받는 테스트만 실행합니다.
```typescript
// test/helpers/touchfiles.ts
{
"qa-workflow": ["qa/SKILL.md.tmpl", "browse/src/server.ts"],
"ship-flow": ["ship/SKILL.md.tmpl", "scripts/resolvers/preamble.ts"]
}
```
***
## 요약: 가져갈 수 있는 디자인 원칙
gstack의 엔지니어링 실무에서 Skill 개발자에게 가장 중요한 다음과 같은 설계 원칙을 추출했습니다.
1. **템플릿 생성 > 수동 동기화**: 기술 전반에 걸쳐 공유되는 콘텐츠는 템플릿 + 빌드 단계를 사용하여 자동으로 생성되며 복사 및 붙여넣기는 불가능합니다.
2. **패시브 감지 > 액티브 감지**: 모든 스킬 호출에 업그레이드 감지가 포함되어 있으며 사용자는 인식하지 못하지만 적용률은 100%입니다.
3. **로그 추가 > 복잡한 데이터베이스**: JSONL + 파일 시스템은 간단하고 안정적으로 대부분의 지속성 요구 사항을 충족할 수 있습니다.
4. **프로그레시브 부팅 > 단일 구성**: 센티넬 파일을 사용하여 부팅 단계를 제어하며 각 단계는 한 번만 나타납니다.
5. **정확한 게이팅 > 완전 자동**: "중지 및 사용자 대기"와 "자동 계속" 사이의 경계를 명확하게 정의합니다.
6. **신뢰도 정량화 > 퍼지 판단**: 각 AI 판단에는 신뢰도 점수가 부여되며 낮은 신뢰도는 자동으로 하향 조정됩니다.
7. **Time Decay > 수동 정리**: 학습 기록에 대한 신뢰도는 시간이 지나면서 쇠퇴하고, 오래된 지식은 자연스럽게 사라집니다.
8. **금지어 목록 > 스타일 가이드**: 금지어를 직접 나열하는 것이 "자연스러운 어조를 사용하세요"보다 훨씬 효과적입니다.
***
**관련 자료**:
* [gstack 개념](/ko/docs/notes/gstack/concept) — gstack은 무엇이고 어떤 문제를 해결하나요?
* [gstack 실무 장](/ko/docs/notes/gstack/practice) — 설치부터 실행까지 전체 워크플로
* [gstack Front-end Skill](/ko/docs/notes/gstack/frontend-skills) — 프런트엔드/UI 디자인 스킬 파노라마 및 권장 워크플로
* [Claude Skills Concept](/ko/docs/notes/claude-skills/concept) — 스킬의 기본 메커니즘을 이해합니다.
# Diagnose와 Triage: 먼저 피드백 루프를 구축하고, 누구에게 맡길지 결정하기
## 두 가지 스킬을 함께 설명하는 이유
`/diagnose`와 `/triage`는 README에서 두 개의 독립적인 스킬이지만, 동일한 엔지니어링 문제의 두 가지 측면을 해결합니다.
* `/diagnose`는 **이 버그가 무엇인지, 어떻게 재현하고, 어떻게 수정되었음을 증명할 수 있는지**에 관심을 가집니다.
* `/triage`는 **이 이슈가 현재 정보를 기다려야 하는지, 에이전트에게 맡겨야 하는지, 사람에게 맡겨야 하는지, 아니면 처리하지 않아야 하는지**에 관심을 가집니다.
하나는 사실을 담당하고, 다른 하나는 프로세스를 담당합니다. 실제 프로젝트에서는 이 두 가지가 종종 연결됩니다. 먼저 버그 이슈를 분류하고, 정보가 부족하면 `needs-info`로 전환합니다. 정보가 충분하면 diagnose를 사용하여 피드백 루프를 구축합니다. 재현이 명확해지면 `ready-for-agent`로 할지 `ready-for-human`으로 할지 결정합니다.
## /diagnose의 핵심: 피드백 루프가 전부입니다.
`/diagnose`에서 가장 기억해야 할 문장은 **먼저 에이전트가 실행할 수 있는 pass/fail 신호를 구축하라**는 것입니다.
Matt는 진단을 6단계로 나눕니다.
| 단계 | 목표 |
| --------------------- | ------------------------------------- |
| Build a feedback loop | 빠르고, 확실하며, 반복적으로 실행 가능한 실패 신호를 구축합니다. |
| Reproduce | 이 신호가 사용자가 설명한 동일한 버그를 재현하도록 합니다. |
| Hypothesise | 3-5개의 반증 가능한 가설을 나열합니다. |
| Instrument | 최소한의 프로브로 가설을 검증합니다. |
| Fix + regression test | 올바른 테스트 표면에서 회귀 테스트를 작성한 다음 수정합니다. |
| Cleanup + post-mortem | 임시 프로브를 정리하고 실제 근본 원인을 기록합니다. |
이는 많은 사람들이 버그를 디버깅하는 순서와 반대입니다. 일반적인 디버깅 프로세스는 코드를 보고, 원인을 추측하고, 수정하고, 페이지를 새로 고치는 것입니다. Matt는 반대로, 먼저 버그를 반복 가능한 기계 신호로 만든 다음 가설을 논의합니다.
## 좋은 피드백 루프란 무엇인가?
`/diagnose`는 우선순위가 높은 순서대로 피드백 루프를 제공합니다.
| 루프 | 적합한 시나리오 |
| ----------------------------- | ------------------------------------------------ |
| 실패 테스트 | 적절한 테스트 표면이 있고 버그를 직접 표현할 수 있는 경우 |
| curl / HTTP script | API 버그, 서버 측 동작은 요청으로 재현 가능 |
| CLI + fixture | 명령줄 도구, 파서, 변환기 |
| Headless browser | UI 버그, 콘솔 오류, 네트워크 동작 |
| Replay captured trace | 온라인 실제 페이로드, 이벤트 스트림, 로그 추적 |
| Throwaway harness | 시스템의 작은 부분만 시작하고 복잡한 종속성을 격리 |
| Property / fuzz loop | 간헐적인 오류 출력, 트리거율을 높여야 하는 경우 |
| Bisection / differential loop | 특정 버전 이후에 문제가 발생하여 이분법 또는 이전 버전과 비교가 필요한 경우 |
| HITL script | 수동으로만 클릭할 수 있는 경우에도 스크립트에 따라 안정적인 출력을 제공하도록 합니다. |
여기에는 매우 엄격한 판단이 있습니다. **루프가 없으면 가설 단계로 넘어가지 마십시오.** 신호가 없으면 모든 분석은 "그렇게 보이는" 것으로 변질됩니다.
## 비결정적 버그는 어떻게 처리해야 하는가?
`/diagnose`는 간헐적인 버그에 대해서도 실용적인 태도를 취합니다. 목표는 처음부터 100% 재현하는 것이 아니라, 재현율을 디버깅 가능한 수준으로 높이는 것입니다.
예를 들어:
* 100번 반복하여 트리거
* 동시 트리거
* sleep을 주입하여 경쟁 조건 창 확대
* 무작위 시드 또는 시간 고정
* 환경 변수 및 외부 종속성 축소
1%의 간헐적인 버그는 디버깅하기 어렵지만, 50%의 간헐적인 버그는 이미 디버깅 가능한 대상입니다. 이 접근 방식은 프런트엔드 비동기, 메시지 큐, 결제 콜백, 스트리밍 출력에 매우 유용합니다.
## 가설은 반증 가능해야 한다.
Matt는 검증을 시작하기 전에 3-5개의 가설을 나열하고, 각 가설에 대한 예측을 작성하도록 요구합니다.
```text
만약 X가 원인이라면, Y를 변경한 후 버그가 사라져야 합니다.
또는 Z를 관찰할 때 특정 특징이 나타나야 합니다.
```
이는 에이전트가 첫 번째 그럴듯한 설명에 갇히는 것을 방지합니다. 더 중요한 것은, 특정 실험이 정보 가치가 있는지 판단할 수 있게 해줍니다.
나쁜 가설:
```text
캐시 문제일 수 있습니다.
```
반증 가능한 가설:
```text
만약 브라우저 캐시로 인해 오래된 스크립트가 실행된다면, 캐시를 비활성화하고 강제 새로 고침을 한 후 콘솔의 오래된 번들 해시가 사라지고 버튼 클릭 이벤트도 복구되어야 합니다.
```
후자만이 검증할 가치가 있습니다.
## 수정 단계에서 가장 흔히 저지르는 실수
`/diagnose`는 올바른 테스트 표면이 있다면, 최소 재현을 실패 테스트로 전환한 다음 코드를 수정하도록 요구합니다.
핵심은 "올바른 테스트 표면"입니다. 단순히 유닛 테스트를 추가하는 것이 회귀 테스트가 아닙니다. 올바른 테스트 표면은 실제 버그 패턴을 커버해야 합니다.
* 버그가 여러 호출자의 조합으로 트리거되는 경우, 단일 함수만 테스트해서는 안 됩니다.
* 버그가 실제 페이로드 구조로 트리거되는 경우, 수동으로 작성된 장난감 객체만 테스트해서는 안 됩니다.
* 버그가 브라우저 이벤트 순서로 트리거되는 경우, 순수 함수만 테스트해서는 안 됩니다.
올바른 테스트 표면을 찾을 수 없다면, 그 자체가 결론입니다. 코드 구조가 버그를 고정할 수 있는 여지를 남기지 않았다는 것입니다. 수정 후에는 이 정보를 [`/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)에 전달해야 합니다.
## /triage의 핵심: 이슈는 상태 머신이다.
`/triage`는 AI가 단순히 "이슈를 봐주는" 것이 아닙니다. 이슈를 작은 상태 머신으로 간주합니다.
각 이슈는 동시에 다음을 가져야 합니다.
* 카테고리: `bug` 또는 `enhancement`
* 상태: `needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`
이 상태 세트의 가치는 유지보수자가 다음 질문에 빠르게 답할 수 있도록 하는 데 있습니다.
* 아직 아무도 보지 않은 것은 무엇인가?
* 보고자가 정보를 보충해야 하는 것은 무엇인가?
* AFK 에이전트에게 맡길 수 있을 만큼 명확한 것은 무엇인가?
* 사람이 직접 해야 하는 것은 무엇인가?
* 닫아야 하고 그 이유를 기록해야 하는 것은 무엇인가?
## ready-for-agent의 기준
`ready-for-agent`는 이 프로세스에서 가장 중요한 상태입니다. 이는 "이 작업을 AI가 시도해 볼 수 있다"는 의미가 아니라,
> 작업이 부재중인 에이전트가 독립적으로 수령, 구현, 검증할 수 있을 만큼 명확하다.
이는 일반적으로 이슈에 다음이 포함되어 있음을 의미합니다.
* 배경 및 문제 진술
* 관련 코드 경로 또는 모듈
* 명확한 수락 기준
* 알려진 제약 사항
* 버그인 경우, 재현 방법이 있는 것이 좋습니다.
* 추가적인 제품/디자인 판단이 필요하지 않습니다.
이러한 것들이 부족하다면 `needs-info` 또는 `ready-for-human`이어야 하며, 에이전트에게 억지로 넘겨서는 안 됩니다.
## needs-info는 구체적인 질문을 해야 한다.
`/triage`가 `needs-info`에 제공하는 템플릿은 매우 단순하지만, 핵심은 질문이 구체적이어야 한다는 것입니다.
```markdown
## Triage Notes
**지금까지 확인된 사항:**
- ...
**(@reporter)님께 필요한 추가 정보:**
- ...
```
나쁜 질문:
```text
더 많은 정보를 제공해 주십시오.
```
좋은 질문:
```text
문제를 트리거한 브라우저 버전, 오류 페이지 URL, 클릭 순서, 그리고 Network 패널의 `/api/orders/:id` 응답 본문을 제공해 주십시오.
```
AI는 예의 바른 헛소리를 쉽게 쓸 수 있지만, 이 스킬은 "우리가 이미 알고 있는 것"과 "아직 부족한 것"을 분리하도록 강요합니다.
## wontfix도 기록해야 한다.
`/triage`는 개선 사항에 대한 `wontfix`에 대해 흥미로운 디자인을 가지고 있습니다. 단순히 이슈를 닫는 것이 아니라, 거부 이유를 `.out-of-scope/` 지식 기반에 작성하고, 댓글에 링크를 추가합니다.
이렇게 하면 다음에 유사한 요구 사항이 발생할 때 AI가 동일한 논의를 다시 시작하지 않습니다. AI는 먼저 `.out-of-scope/`를 읽고 유지보수자에게 "이 방향은 이전에 X라는 이유로 거부되었습니다"라고 상기시킬 수 있습니다.
이는 ADR의 정신과 매우 유사합니다. 모든 결정을 기록하는 것이 아니라, 미래에 혼란을 야기하고 반복적으로 나타날 결정만 기록합니다.
## 두 가지 스킬의 협력 방법
일반적인 버그 이슈는 다음과 같이 진행될 수 있습니다.
1. `/triage`가 이슈, 댓글, 태그 및 관련 코드를 읽습니다.
2. `bug + needs-triage`로 판단합니다.
3. 먼저 재현을 시도합니다. 단계가 부족하면 `needs-info`로 전환합니다.
4. 정보가 충분하면 `/diagnose`를 시작합니다.
5. `/diagnose`는 재현 루프를 구축하고, 가설을 나열하고, 근본 원인을 찾습니다.
6. 수정 경로가 명확하고 테스트 표면이 명확하면 이슈는 `ready-for-agent`로 변경됩니다.
7. 제품 판단, 외부 권한, 수동 검증이 필요한 경우 이슈는 `ready-for-human`으로 변경됩니다.
8. 수정 후 근본 원인과 회귀 테스트를 이슈 또는 PR에 다시 작성합니다.
이 프로세스의 핵심은 "AI가 자동으로 버그를 수정하는 것"이 아니라, 이슈를 모호한 설명에서 실행 가능한 작업 패키지로 전환하는 것입니다.
## 나의 사용 제안
한 가지만 기억한다면:
> `/diagnose`는 먼저 "어떻게 망가졌는지 증명할 수 있는가"를 묻고, `/triage`는 먼저 "지금 어떤 상태에 있어야 하는가"를 묻습니다.
이 두 가지 질문은 많은 저품질 AI 프로그래밍을 막을 수 있습니다.
* 재현 없이 수정
* 수락 없이 작업 시작
* 근본 원인 없이 리팩토링
* 정보 없이 에이전트에게 떠넘기기
Matt의 이 두 가지 스킬은 화려하지 않지만, 실제 팀에서 숙련된 엔지니어가 하는 일과 매우 유사합니다. 먼저 사실을 수렴하고, 그 다음 프로세스를 진행합니다.
## 참고 자료
다음 글: [TDD: Red-Green-Refactor를 사용하여 AI가 작은 단계를 밟도록 강제하기](/ko/docs/notes/matt-pocock-skills/tdd).
# Grill Me: 코드를 작성하기 전에 AI가 50가지 질문을 하도록 합니다.
## 실패 모드: “AI가 내가 원하는 대로 하지 않았습니다”
Matt가 연설에서 언급한 첫 번째 실패 모드는 마음 속의 요구 사항이 매우 명확하다고 생각하고 AI가 이를 작성하도록 하는 것입니다. 전혀 그렇지 않습니다.
> "I would run it, and I would try not to look at the code, but I would look at the code, and I realized I would get worse code. I did it again, I got even worse code... I did it again, kept running the compiler, and I would just end up with garbage."
많은 사람들이 이 느낌에 익숙합니다. "로그인을 추가해 주세요"라고 말하면 AI는 "장치를 기억하시겠습니까?"라고 묻지 않습니다. "계정 잠금에 몇 번이나 실패하셨나요?" "세션이 만료되는 데 얼마나 걸리나요?" 합리적이라고 생각되는 계획을 직접 제시합니다. 검토할 때쯤에는 500줄이 작성되었습니다. 재작업하는 데 2시간이 걸렸습니다.
## 왜 이런 일이 일어나는가: 디자인 컨셉이 벗어남
Matt는 “The Design of Design”에서 Frederick P. Brooks의 **디자인 컨셉**(디자인 컨셉)을 인용합니다.
> 여러 사람이 협력하여 무언가를 디자인하면 여러분 사이에 무언가가 창조될 것입니다. 그것은 여러분의 마음 속에 떠다니는 보이지 않는 "이것에 대한 이론"입니다. 그것은 자산도 아니고, 마크다운 파일에 담긴 자산도 아니고, 눈에 보이지 않는 합의입니다.
AI는 코드가 나오자마자 코드를 작성합니다. 즉, 동일한 디자인 컨셉을 사용자와 전혀 공유하지 않는다는 의미입니다. 코드를 작성할 때 잘못된 점은 구문이 아니라 전제입니다.
이 문제를 해결하려면 시작하기 전에 먼저 디자인 컨셉 조정을 해야 합니다. Brooks가 제공한 도구는 **디자인 트리**라고 합니다. 즉, 결정을 여러 분기로 분할한 다음 각 분기를 분할합니다. 업스트림 결정을 건너뛰고 다운스트림 결정을 직접 내릴 수는 없습니다. 그렇지 않으면 업스트림이 다운스트림으로 변경되면 모든 것을 다시 실행해야 합니다.
## 맷의 스킬 전문
[`mattpocock/skills`](https://github.com/mattpocock/skills)에서 Matt의 이 이론 구현은 `productivity/grill-me/SKILL.md`이며, 전체 파일과 머리말을 합한 내용은 15줄 미만입니다.
```markdown
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
문장별로 해석하기:
* **"가차없이 인터뷰하세요"** - 핵심 단어는 *가차없이*(놓지 않음)입니다. 기본적으로 LLM은 1\~2개의 질문을 한 다음 "거의"라고 느끼고 조치를 취하는 경향이 있습니다. 이 단어는 이러한 경향을 강제로 억제합니다.
* **"디자인 트리의 각 가지를 따라 내려가다"** —— Brooks의 디자인 트리 개념. Claude가 요구 사항을 트리로 처리하여 업스트림을 먼저 해결한 다음 다운스트림을 해결하도록 합니다. "로그인"이라고 말하면 먼저 "인증 방법"(트리 루트)을 묻고, 답변에 따라 "세션 관리 방법"/"토큰 저장 방법"(하위 노드)을 확장합니다.
* **"결정 간의 종속성을 하나씩 해결"** - 패키지 질문을 명시적으로 금지합니다. 결정 간에 종속성이 있는 경우가 많습니다(SSO를 선택하는 경우 다운스트림에는 비밀번호 정책 문제가 필요하지 않음). 먼저 업스트림에서 많은 다운스트림 문제를 제거할 수 있는지 확인하세요.
* **"각 질문에 대해 권장 답변을 제공하세요"** - 주요 보너스 포인트. AI는 질문만 하는 것이 아니라 답변도 추천해준다. 당신은 고개를 끄덕이거나 아니오로 타이핑 시간을 80% 절약할 수 있습니다.
* **"한 번에 하나씩 질문하세요"** - AI가 한 번에 10개의 질문을 제공하는 것을 방지합니다.
* **"코드베이스를 탐색하여 질문에 대한 답을 얻을 수 있으면 대신 코드베이스를 탐색하십시오."** - 프로젝트에 이미 존재하는 사실인 경우(예: "프로젝트에서 어떤 테스트 프레임워크가 사용되는지") Claude가 직접 확인하도록 하세요. 묻지 마세요.
7줄이지만 각 문장은 특정 LLM 행동 편향에 해당합니다.
## 설치 및 사용방법
**설치**:
```bash
npx skills@latest add mattpocock/skills
```
`grill-me` 및 `setup-matt-pocock-skills`을 확인하세요(grill-me는 후자에 의존하지 않지만 다른 스킬은 이에 의존하므로 함께 설치하는 것이 좋습니다).
**통화**: Claude Code 대화 상자에 `/grill-me`을 입력합니다.
**일반적인 프로세스**:
1. 하고 싶은 일을 설명하는데, 내용이 매우 모호할 수 있습니다. ("내 블로그에 댓글 기능을 추가하고 싶어요")
2. `/grill-me`을(를) 입력하세요
3. 클로드는 하나씩 질문을 하기 시작했고, 각 질문에 추천 답변을 주었습니다.
4. 각 질문에 하나씩 대답합니다(끄덕끄덕/아니요/맞습니다)
5. 일반적으로 20\~50개의 질문 후에 합의에 도달하며, 클로드가 요약을 제공합니다.
6. 요약은 [`/to-prd`](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues)에 직접 공급되어 PRD가 되거나 [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)에 직접 전달되어 작성을 시작할 수 있습니다.
## 실제 사례: 영상 편집 기능의 가격은 얼마인가요?
Matt는 ["내가 매일 사용하는 5가지 상담원 기술"](https://www.aihero.dev/5-agent-skills-i-use-every-day)에서 몇 가지 구체적인 수치를 제시했습니다.
* **새로운 동영상 편집기 기능** - 합의에 도달하기 위한 16가지 질문
* **복잡한 기능** - 30\~50문항
* **매우 복잡함** - 100개 질문, 세션 최대 45분
샘플 질문(Matt의 비디오/블로그 게시물에서 복원됨):
* "Should video clips be reorderable, or only added/removed in sequence?"
* "When a clip is deleted, do we keep its source file, or delete the file too?"
* "Does the editor need undo/redo? How many steps deep?"
* "Should we render previews in the browser, or rely on a backend service?"
이러한 문제는 기술적 문제가 아니었고 모두 제품 결정이었습니다. 그러나 **모든 결정은 수백 줄의 코드 모양을 결정합니다**. 이러한 질문을 건너뛰고 AI가 직접 작성하게 하면 AI가 스스로 답을 제시하고 작성 후 다시 돌아와서 하나씩 거부하게 됩니다.
## 플랜 모드와 클로드코드 내장 플랜 모드의 차이점
Claude Code는 `plan mode`과 함께 제공됩니다(입력하려면 Shift+Tab을 누르세요). 표면적으로는 grill-me와 유사해 보입니다. 조치를 취하기 전에 먼저 논의해야 합니다. 그러나 Matt는 연설에서 grill-me를 선호한다고 직접 말했습니다.
구체적인 차이점:
| 치수 | 계획 모드 | /그릴나 |
| ------------ | ------------------------------ | ------------------------------------- |
| 기본 목표 | 가능한 한 빨리 실행 가능한 계획 수립 | 먼저 합의에 도달, 계획은 부산물 |
| 질문 수 | 0\~5 | 20\~100 |
| 질문 형식 | 한 번에 한 문단씩 질문하기 | 한 번에 하나씩 질문하기 |
| 추천 답변을 원하시나요 | 아니요 | 예 |
| 코드 베이스 탐색 여부 | 가끔 | 적극적으로(명시적 지시) |
| 시나리오에 적합 | 이미 명확하게 생각했으며 구현 계획을 확인하고 싶습니다 | 아직 명확하게 생각하지 않았으므로 명확하게 생각하도록 강요받아야 함 |
가장 큰 실질적인 차이점은 "**긴급 여부**"입니다. 플랜 모드는 서둘러 시작하고, 그릴미는 서두르지 않고, 프롤로그보다는 '명확하게 생각하기'를 주요 과제로 여긴다.
## 고급 사용법
### 1. 비프로그래밍 시나리오
`grill-me`은 코드를 바인딩하지 않으며 순수한 제품 의사결정 대화에도 사용할 수 있습니다. Matt가 직접 사용합니다.
* 코스 강의 계획서 디자인
* 기사 작성
* 내부 의사소통 서류
마음 속에 막연한 생각이 있고, 그것을 충분히 생각해보고 싶은 한, 그것을 사용할 수 있습니다.
### 2. [`/to-prd`](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues)과 협력하세요
grill-me 세션이 끝난 후 `/to-prd`이라고 말하면 Claude가 전체 대화를 구조화된 PRD(사용자 스토리, 모듈 분할 및 테스트 전략 포함)로 압축하여 이슈 트래커에 제출합니다. **핵심사항: 중간에 맥락을 지우지 마세요** - to-prd는 대화 맥락에서 직접 추출되며 다시 묻지 않습니다.
### 3. [`/grill-with-docs`](/ko/docs/notes/matt-pocock-skills/grill-with-docs)과 협력하세요
프로젝트에 이미 `CONTEXT.md`(도메인 언어) 및 `docs/adr/`(아키텍처 결정)이 있는 경우 `/grill-me` 대신 `/grill-with-docs`을 사용하세요. 고문을 받는 동안 **동기적으로 CONTEXT.md**를 업데이트합니다. 문서가 업데이트되는 동안 결정이 내려지며 더 이상 "문서가 영원히 오래된 것"이라는 문제가 없습니다.
### 4. 질문 깊이 맞춤설정
시간이 촉박한 경우 `/grill-me` 바로 뒤에 "질문을 10개로 제한하고 아키텍처 결정에만 집중하세요."라는 문장을 추가할 수 있습니다. 그것은 당신의 한계에 따라 수렴할 것입니다. 하지만 Matt는 이를 권장하지 않습니다. 그는 "더 많은 질문을 하는 것"이 바로 이 기술의 가치라고 믿으며 이를 중단하는 것은 거의 계획 모드와 같습니다.
## 메모
**처음 실행하면 짜증날 것입니다**. "한 문장에 500줄을 생성하는 것"에 익숙한 사람들은 처음으로 AI가 30번이나 질문하는 것이 시간 낭비라고 느낄 것이다. Matt의 조언은 처음 5개의 질문을 붙잡으라는 것입니다. 처음 5개의 질문은 종종 당신이 생각지도 못한 것들을 드러내줍니다. 그 기준점을 넘기면 중독이 됩니다.
**매우 작은 작업에는 적합하지 않습니다**. 오타를 수정하고 console.log를 추가하세요. grill-me를 사용하지 마세요. "새로운 것을 만든다"거나 "부작용이 있는 오래된 것을 바꾸는 것"에 적합합니다.
**때때로 AI가 기술적인 세부 사항을 묻는 경우가 있습니다**. 신경 쓰지 않고 판단하게 하고 싶다면 "당신의 전화"와 "당신이 결정합니다"라고 대답하면 수락하고 계속됩니다.
## 이 스킬이 왜 인기가 있나요?
`/grill-me`은 Matt의 스킬 중 가장 자주 스크린샷되고 전달되는 스킬입니다. 그 이유는 복잡하지 않습니다.
1. **매우 미니멀함**: 7줄의 마크다운, 복사하여 붙여넣기만 하면 됩니다.
2. **즉각적 효과**: 첫 실행 시 AI의 '문제 밀도' 변화를 느낄 수 있습니다.
3. **휴대용**: Claude Code에 의존하지 않고 Codex, Cursor, Aider 모두 사용 가능
4. **LLM 방지 기본 동작 제공**: 각 단어는 높은 엔지니어링 미학을 갖춘 LLM 방지 편차입니다.
그 성공은 "**실력이 꼭 길 필요는 없다**"는 최고의 논거가 되기도 했습니다.
## 참조 리소스
다음 기사: [Grill With Docs: 프로젝트 언어 및 ADR 유지 관리](/ko/docs/notes/matt-pocock-skills/grill-with-docs) - 도메인이 복잡한 프로젝트를 위한 grill-me의 고급 버전입니다.
# Grill With Docs: 도메인 언어와 ADR을 사용하여 AI에 프로젝트 메모리 장착
## 실패 모드: "AI가 너무 장황합니다."
Matt 연설의 두 번째 실패 모드는 다음과 같습니다.
> AI는 간단한 말을 여러 장의 말로 표현합니다. 그것은 당신에게 두 가지 언어를 말하는 것과 같습니다.
이는 코드의 양과는 아무런 관련이 없으며 **어휘의 잘못된 배치**입니다. AI는 기본적으로 일반 용어("아이템", "데이터", "핸들러")를 사용하며, 여러분이 생각하는 프로젝트의 실제 용어는 "코스", "초안 버전", "유령 레슨"일 수 있습니다. AI는 이러한 단어가 프로젝트에서 특정 의미를 가지고 있다는 것을 모르기 때문에 주변에 동의어인 새로운 단어를 많이 생성합니다. 결과는 다음과 같습니다.
* 장황한 사고 과정(독점적인 단어는 피하세요)
* 구현이 마음 속 디자인과 잘못 정렬되었습니다(동일한 의미 공간에 있지 않기 때문).
* 세션 간에 재사용할 수 없습니다(각 대화마다 컨텍스트를 다시 설정해야 함).
## 고전 이론: DDD의 유비쿼터스 언어
Matt는 Eric Evans의 "Domain-Driven Design"을 인용했습니다. 이 책은 2003년에 출판되었으며 **유비쿼터스 언어**라는 개념을 제안했습니다.
> 동일한 용어 집합을 사용하여 도메인 전문가, 개발자 및 코드를 연결합니다. 제품 토론, 코드 주석, 변수 이름, 문서에 나오는 단어는 동일한 의미를 가져야 합니다.
DDD의 목표는 코드를 도메인 전문가의 두뇌처럼 보이게 만드는 것입니다. AI 시대에는 새로운 역할이 있습니다. **LLM도 이 언어로 되어 있어야 합니다**. LLM은 스탠드업 회의에 참석하지 않고 제품 요구 사항 회의를 볼 수 없으며 그룹의 속어를 이해할 수 없습니다. LLM은 귀하가 제공한 문서를 통해서만 배울 수 있습니다.
Matt는 이것을 기술로 전환했습니다. 코드 베이스를 스캔하여 용어를 추출하고 마크다운 파일 `CONTEXT.md`을 생성한 다음 이를 인간과 AI 모두에 맞게 조정했습니다.
## 기술의 진화: 유비쿼터스 언어에서 문서를 이용한 그릴로
가장 초기의 기술은 `ubiquitous-language`이라고 불렸습니다. 이 기술은 단 한 가지 작업만 수행했습니다. 코드 베이스를 스캔하여 용어집을 생성하는 것이었습니다. 그러나 Matt는 나중에 단순히 문서를 생성하는 것만으로는 충분하지 않다는 사실을 발견했습니다.
* **문서는 최신이 아닙니다**: 오늘 생성되었으며 내일 코드가 변경되고 용어집이 유지되지 않습니다.
* **사람들은 그것을 보려고 하지 않을 것입니다**: 거기에 놓으면 죽을 것입니다.
그는 이를 세 가지를 결합한 `grill-with-docs`로 리팩토링했습니다.
1. **고문 요구 사항**(grill-me에서 상속된 모든 능력)
2. **기존 용어집에 도전**: 귀하가 말한 단어가 CONTEXT.md에 작성된 내용과 일치하지 않습니까? 당장 지적해
3. **결정을 내릴 때 문서를 동시에 업데이트**: 고문 과정에서 도달한 새로운 결론은 CONTEXT.md에 인라인으로 기록되거나 새 ADR을 생성합니다.
이는 "정적 문서 생성"에서 "대화는 문서 유지"로의 패러다임 전환입니다.
## 스킬 전문
`engineering/grill-with-docs/SKILL.md`의 핵심 구조:
```markdown
---
name: grill-with-docs
description: Grilling session that challenges your plan against the
existing domain model, sharpens terminology, and updates
documentation (CONTEXT.md, ADRs) inline as decisions crystallise.
---
Interview me relentlessly about every aspect of this plan until we
reach a shared understanding. Walk down each branch of the design
tree, resolving dependencies between decisions one-by-one. For each
question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question
before continuing.
If a question can be answered by exploring the codebase, explore the
codebase instead.
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
If a CONTEXT-MAP.md exists at the root, the repo has multiple contexts.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in
CONTEXT.md, call it out immediately.
"Your glossary defines 'cancellation' as X, but you seem to mean Y —
which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise
canonical term.
"You're saying 'account' — do you mean the Customer or the User?
Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with
specific scenarios.
### Cross-reference with code
When the user states how something works, check whether the code
agrees. If you find a contradiction, surface it.
### Update CONTEXT.md inline
When a term is resolved, update CONTEXT.md right there. Don't batch
these up — capture them as they happen.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. Hard to reverse
2. Surprising without context
3. The result of a real trade-off
```
## 실제 CONTEXT.md는 어떤 모습인가요?
Matt의 [`course-video-manager`](https://github.com/mattpocock/course-video-manager/blob/main/CONTEXT.md) 저장소는 완전한 CONTEXT.md 예제를 제공합니다. 이에 대한 느낌을 얻기 위해 몇 가지 용어를 선택해 보겠습니다.
| 용어 | 정의 |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| **Course** | The primary domain entity: a structured collection of versions, sections, lessons, and videos |
| **Draft Version** | The single mutable CourseVersion that is currently being edited; always the latest by `createdAt` |
| **Published Version** | An immutable CourseVersion with a name and description, created by the Publish flow |
| **Ghost Lesson** | A lesson that exists in the database but not yet on the file system (`fsStatus = "ghost"`) |
| **Export Hash** | A SHA256 hash derived from a video's clip filenames, timestamps, clip order |
| **Unexported Video** | A video whose current Export Hash does not match any file on disk; blocks publishing |
| **Materialization Cascade** | The chain reaction when materializing a lesson inside a ghost course |
| **Clip** | A timestamped segment of source footage within a video |
| **Fractional Index** | A string-based ordering value that allows inserting items between existing items |
| **Purge** | The deliberate deletion of an Exported Video's `.mp4` file from disk |
몇 가지 사항에 유의하세요.
1. **각 용어는 동명사 또는 고유명사입니다** – “주문 상태”와 같은 설명 문구가 아닙니다.
2. **각 정의는 온톨로지 네트워크를 형성하는 다른 용어**(강좌 → 버전 → 강의 → 비디오)를 참조합니다.
3. **코드 필드가 직접 나타납니다** (`fsStatus = "ghost"`) - 문서와 코드의 1:1 매핑
4. **결정 설명 포함**('게시 차단', '연쇄 반응') - 명사뿐 아니라 규칙도 포함
코드 작성 시 AI가 이 문서를 보면 '파일이 없는 강의' 대신 '유령 강의'를 사용하게 됩니다. 코드, 대화, 커밋 메시지가 모두 통합됩니다.
## ADR: 생성 시
스킬에는 중요한 제한 사항이 있습니다.
> Only offer to create an ADR when all three are true:
>
> 1. **Hard to reverse** — the cost of changing your mind later is meaningful
> 2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
> 3. **The result of a real trade-off** — there were genuine alternatives
ADR(Architecture Decision Record) 문화는 2011년 Michael Nygard의 블로그에서 시작되었지만 많은 팀에서 이를 사용하여 모든 결정에 대한 ADR을 작성합니다. ADR 20개 중 18개가 계정을 실행하고 있습니다. Matt 이 삼각측량은 훌륭한 도구입니다. **ADR은 세 가지 조건이 동시에 충족되는 경우에만 가치가 있습니다**. 그렇지 않으면 CONTEXT.md에서 요약하고 코드에서 요약하도록 하세요.
## 설치 및 사용방법
**전제 조건**: `/setup-matt-pocock-skills`을 먼저 실행하십시오(CONTEXT.md를 어디에 넣을지, ADR 디렉토리를 어디에 둘 것인지 묻습니다).
**전화**: `/grill-with-docs`
**일반적인 프로세스**:
1. 하고 싶은 일을 설명해주세요
2. `/grill-with-docs`
3. Claude는 먼저 CONTEXT.md 및 docs/adr/을 **스캔**하여 기존 용어와 결정을 컨텍스트에 로드합니다.
4. 그 과정에서 고문을 시작하세요:
* 사용하신 단어가 CONTEXT.md와 충돌합니다 → 즉석에서 지적해주세요
* 모호한 단어를 사용했습니다(예: "계정"은 고객 또는 사용자일 수 있음) → 둘 중 하나를 선택하여 문서에 놓도록 합니다.
* 말씀하신 동작이 기존 코드와 일치하지 않습니다. → 충돌되는 부분을 지적해 주세요.
5. 결정에 도달하면 **CONTEXT.md**를 동기식으로 업데이트합니다(백로그 없음, 일괄 처리 없음).
6. 되돌릴 수 없는 주요 결정 → ADR 생성 여부를 묻습니다.
아직 프로젝트에 CONTEXT.md 및 docs/adr/이 없으면 느리게 생성됩니다. 첫 번째 용어를 작성하고 첫 번째 ADR을 빌드해야 할 때까지 파일이 생성되지 않습니다. 우리는 즉시 빈 템플릿을 제공하지 않습니다.
\##다중 컨텍스트 프로젝트(CONTEXT-MAP.md)
프로젝트가 너무 커서 하나의 CONTEXT.md에 들어갈 수 없는 경우(예를 들어 주문과 청구가 두 개의 독립적인 경계 컨텍스트인 경우) 루트 디렉터리에 일반 디렉터리로 `CONTEXT-MAP.md`을 넣을 수 있습니다.
```
/
├── CONTEXT-MAP.md ← 总目录
├── docs/adr/ ← 系统级决策
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← 模块级决策
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
`/grill-with-docs`은 CONTEXT-MAP.md의 존재를 자동으로 인식하고 해당 하위 디렉터리로 이동합니다. 이는 DDD의 **제한된 컨텍스트** 개념을 직접 구현한 것입니다. 각 컨텍스트의 "순서"는 서로 다른 의미를 가질 수 있으며 오염을 방지하기 위해 별도로 유지 관리됩니다.
\##과 grill-me의 차이점
| 치수 | /그릴나 | /grill-with-docs |
| ------------------- | ---------------- | ------------------------------------ |
| 질문 능력 | ✅ | ✅ (모두 상속) |
| 프로젝트 용어 검증 | ❌ | ✅ |
| 실시간 업데이트 CONTEXT.md | ❌ | ✅ |
| ADR 트리거 판단 | ❌ | ✅ |
| 적용단계 | 초기 아이디어, 개인 프로젝트 | 도메인 복잡성이 있는 실제 프로젝트 |
| 창업비용 | 0 | 설정 필요 + 프로젝트에 CONTEXT.md를 빌드할 의향이 있음 |
단순하고 조악한 판단:
* **개인 대본 작성, 기사 작성, 교육 과정** → `/grill-me`
* **장기적인 유지관리가 필요한 실제 프로젝트** → `/grill-with-docs`
## 직관에 반하는 이점: AI가 "닥치는" 법을 배우게 하세요.
CONTEXT.md는 AI만을 위한 것이 아니라 미래의 AI 세션을 위한 것입니다. 새로운 대화가 시작될 때마다 Claude는 CONTEXT.md를 읽고 즉시 프로젝트 컨텍스트에 들어갈 수 있어 긴 온보딩 시간을 절약할 수 있습니다.
더 미묘한 점은 Matt가 연설에서 CONTEXT.md를 추가한 후 AI의 *사고 추적*을 볼 수 있다고 말했습니다.
> "AI가 덜 장황한 방식으로 생각할 수 있게 해줍니다."
왜요? CONTEXT.md가 없으면 AI는 "사용자, 즉 항목을 주문한 사람, 이하..."라고 생각할 때 자체 용어를 지속적으로 정의해야 하기 때문입니다. CONTEXT.md를 사용하면 "고객"이라고 직접 표시되므로 사고 체인이 훨씬 짧고 응답이 더 빠릅니다.
**LLM의 토큰 경제는 사고 경로 단축 = 더 빠르고, 더 정확하고, 더 저렴한 결과를 결정합니다**. CONTEXT.md는 이러한 효율성을 위한 숨겨진 레버입니다.
## 메모
**CONTEXT.md는 첫 번째 실행에서 크게 변경됩니다**. 프로젝트에 손으로 작성한 CONTEXT.md가 이미 있는 경우 먼저 git stash를 실행하거나 실행하기 전에 시험 실행하도록 합니다("변경할 내용을 먼저 나열하고 파일을 직접 쓰지 마십시오"라는 프롬프트에 문장을 추가할 수 있습니다).
**ADR 조정은 실제 조정입니다**. 흥분하지 말고 모든 결정을 내리면 ADR이 생성됩니다. 귀하의 문서/adr/은 6개월 안에 쓰레기로 가득 차게 될 것입니다. Matt의 세 가지 기준을 엄격히 준수해야 합니다.
**CONTEXT.md 구현 세부 정보를 입력하지 마세요**. Skill에는 "CONTEXT.md를 구현 세부 사항에 연결하지 마십시오. 도메인 전문가에게 의미 있는 용어만 포함하십시오."라는 말이 있습니다. CONTEXT.md에 "PostgreSQL"을 쓰는 것은 잘못된 것입니다. 도메인 전문가는 데이터베이스 선택에 관심이 없으며 이는 ADR의 문제입니다.
## 참조 리소스
다음 글: [to-PRD + to-이슈: 다이얼로그에서 실행 가능한 티켓으로](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues)——고문 이후, 다이얼로그를 실행 가능한 작업 단위로 굳히는 방법.
# 코드베이스 아키텍처 개선: 얕은 모듈을 깊은 모듈로 재구성
## 실패 모드: “잘못된 코드 베이스에서 떠도는 AI”
Matt의 강연에서 네 번째 실패 모드는 그림으로 비유한 것입니다.
> "코드베이스의 얕은 모듈은 다음과 같습니다. 작은 덩어리가 많이 있고 AI는 이를 수정하기 전에 여러 모듈을 거쳐 모든 종속성을 이해해야 합니다."
> "AI is really good at creating codebases like this. So you'll have a situation where AI doesn't understand what your code is doing. It will attempt to explore the code, but because it's poorly laid out, filled with shallow modules, it doesn't get to the right module in time, or doesn't understand all the dependencies."
이는 AI 프로그래밍 특유의 악순환입니다.
```
AI 写代码倾向于产生 shallow 模块(小、多、互相依赖)
↓
代码库变得 shallow
↓
下次 AI 进来探索更难,更容易写错
↓
更多 shallow 模块被加进去
↓
代码库越来越烂,AI 越来越无能
```
이 주기를 깨기 위해서는 수동 역리팩토링을 주기적으로 수행하여 얕은 모듈을 깊은 모듈로 병합해야 합니다. 이것이 바로 `/improve-codebase-architecture`이 하는 일입니다.
## 고전 이론: Ousterhout의 심층 모듈
John Ousterhout는 스탠포드의 CS 교수이자 Tcl 언어 및 Raft 논문의 저자입니다. 그의 2018년 저서 "A Philosophy of Software Design"은 간단하지만 강력한 통치자를 제안합니다.
**모듈의 "깊이" = 인터페이스의 숨겨진 복잡성**
| 유형 | 인터페이스 | 구현 | 이미지 |
| ------ | ----- | -- | ----------- |
| **깊은** | 단순 | 부자 | 직사각형: 좁고 깊은 |
| **얕음** | 복잡한 | 단순 | 직사각형: 넓고 얕은 |
이상적인 모듈은 심층적입니다. 사용자는 짧은 인터페이스만 보면 되고 복잡성은 내부에 숨겨져 있습니다. 극단적인 반례는 얕은 모듈입니다. 인터페이스는 구현만큼 복잡하며, 이는 캡슐화가 없음을 의미합니다. 사용자는 구현을 직접 볼 수도 있습니다.
Ousterhout의 판단: **좋은 코드 기반은 소수의 심층 모듈로 구성됩니다. 나쁜 코드 베이스는 수많은 얕은 모듈**로 구성됩니다. 이는 "기능을 가능한 한 작게, 파일을 가능한 한 짧게, 모듈을 최대한 많이 유지"하는 전통적인 신조와 완전히 반대입니다. 그는 이러한 종류의 신조가 정확히 얕은 모듈을 생성한다고 믿습니다.
## Matt의 확장 프로그램: 삭제 테스트
Matt는 Ousterhout의 이론을 **삭제 테스트**라고 부르는 운영 엔지니어링 테스트로 번역했습니다.
> **Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.**
인간의 말:
* **삭제하면 복잡함이 사라집니다** → 이 모듈은 원래 pass-through(transit) 모듈이라 작동하지 않으니 잘라주세요.
* **삭제하면 복잡성이 N명의 호출자에게 전파됩니다** → 원래는 복잡성을 숨기는 데 도움이 되므로 정말 깊으므로 그대로 두세요.
이 테스트의 장점은 양방향이라는 것입니다. 즉, "삭제해야 하는 얇은 패키지"와 "추출해야 하는 공통 논리"를 모두 식별할 수 있습니다. 코드 조각을 삭제한 후 복잡성이 5곳으로 퍼지는 것을 발견하면 이 코드를 심층 모듈로 추출할 가치가 있음을 의미합니다.
## 핵심 용어(Matt의 정확한 정의)
`improve-codebase-architecture/SKILL.md`에는 **이 단어의 엄격한 사용**을 요구하는 용어집이 있습니다. "구성 요소", "서비스", "API" 및 "경계"로 이동하지 마십시오.
| 용어 | 정의 |
| --------- | ------------------------------------------------------------- |
| **모듈** | 인터페이스와 구현(함수/클래스/패키지/슬라이스)이 있는 모든 것 |
| **인터페이스** | 호출자가 알아야 하는 모든 것 - 유형, 불변, 오류 모드, 순서, 구성(함수 서명뿐만 아니라) |
| **구현** | 모듈 내부의 코드 |
| **깊이** | 인터페이스의 레버. 깊음 = 높은 레버리지, 얕음 = 인터페이스가 구현만큼 복잡함 |
| **심**(심) | 인터페이스의 위치 - 내부 수정 없이 동작을 변경할 수 있는 곳입니다. **"경계"가 아닌 "이음새" 사용** |
| **어댑터** | 솔기에서 인터페이스의 특정 구현 구현 |
| **레버리지** | 호출자가 "깊은"에서 얻는 이점 |
| **지역** | 관리자가 "깊이"를 통해 얻을 수 있는 이점 - 변경 사항, 버그 및 지식이 모두 한 곳에 집중되어 있습니다 |
몇 가지 핵심 원칙:
* **삭제 테스트**: 위 참조
* **인터페이스는 테스트 표면입니다**: 테스트는 인터페이스를 통해서만 실행될 수 있습니다. 이는 심층 모듈 테스트 가능성의 기초입니다.
* **어댑터 1개 = 가상 솔기. 두 개의 어댑터 = 실제 이음새.**: 구현이 하나만 있는 인터페이스는 거짓 이음새입니다. **실제 조인트에는 최소 2개의 어댑터가 필요합니다**
마지막은 특히 반직관적입니다. 많은 팀이 "향후 확장을 위해" 미리 인터페이스를 추상화하지만 실제로는 하나의 구현만 있습니다. Matt의 판단: **쓸모 없음, 삭제**. 두 번째가 이루어질 때까지 기다리십시오. 이는 YAGNI와 동일한 유래를 갖는다.
## 스킬 작업흐름
### 1. 탐색
Skill은 먼저 AI가 `CONTEXT.md` 및 `docs/adr/`을 읽은 다음 `subagent_type=Explore`를 사용하여 하위 에이전트를 코드 베이스로 보냅니다.
경직된 영감 대신 **마찰**을 신호로 사용하세요.
> * Where does understanding one concept require bouncing between many small modules?
> * Where are modules **shallow** — interface nearly as complex as the implementation?
> * Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
> * Where do tightly-coupled modules leak across their seams?
> * Which parts of the codebase are untested, or hard to test through their current interface?
의심스러운 점을 발견할 때마다 삭제 테스트를 적용하세요. 삭제하면 복잡성이 사라지나요, 아니면 퍼질까요? '해산'이라는 답은 심화할 가치가 있는 후보다.
### 2. 현 후보자(성명후보자)
번호가 매겨진 후보자 목록을 제시하십시오.
```
1. Files: src/orders/parser.ts, src/orders/validator.ts, src/orders/normalizer.ts
Problem: 三个文件互相调用,理解 Order 入站需要在三处跳转
Solution: 合并为单一 OrderIntake 模块,对外只暴露 parse(raw) → ValidatedOrder
Benefits:
- Locality: Order 入站的所有逻辑、错误处理、bug 修复集中一处
- Leverage: 调用方从理解 3 个接口降为 1 个
- Tests: 只需测 parse() 的输入输出,不再需要 mock 内部协作
```
요구사항:
* **CONTEXT.md 어휘**를 사용하여 도메인("FooBarHandler"가 아닌 "주문 접수 모듈")에 대해 이야기합니다.
* **용어집**("심", "깊이", "지역성")을 사용하여 건축에 대해 이야기합니다.
* **인터페이스 디자인을 즉시 제안하지 마세요** - 사용자가 먼저 흥미로운 후보를 선택하도록 하세요.
후보자가 기존 ADR과 충돌하는 경우 충돌로 인해 ADR을 다시 방문해야 하는 경우에만 이를 언급하고 다음과 같이 명확하게 표시하십시오.
> "contradicts ADR-0007 — but worth reopening because…"
ADR에서 금지하는 모든 리팩토링을 파헤치지 마세요.
### 3. 그릴링 루프
사용자가 후보를 선택한 후 굽기 모드로 전환합니다([`/grill-with-docs`](/ko/docs/notes/matt-pocock-skills/grill-with-docs)에서 상속됨).
* 디자인 트리를 살펴보세요 - 제약 조건, 종속성, 심화 후 모듈의 모양, 이음매 뒤에 숨겨진 것, 테스트가 살아남을 수 있음
* **부작용은 즉시 발생합니다**:
* 심화 모듈에 CONTEXT.md에 없는 이름을 부여 → 즉시 CONTEXT.md에 추가
* 고문에서 모호한 용어가 날카롭게 표현되었습니다. → CONTEXT.md를 즉시 업데이트하세요.
* 사용자가 합리적인 부하 부담(중요, 미래 탐험가가 알아야 할 사항)으로 후보를 거부 → ADR 생성 제안
* 심화 모듈의 다양한 인터페이스 디자인을 탐색하고 싶음 → `INTERFACE-DESIGN.md` 별도 프로세스로 이동
문서 유지 관리와 아키텍처 변환은 동일한 대화에서 이루어지며 두 번 반복되지 않습니다.
## 실제 사례: 메즈바 아흐메드(Mejba Ahmed)의 사례
타사 개발자 Mejba Ahmed는 이 기술을 사용한 경험을 자세히 기록하기 위해 ["Deep Modules: The Claude Code Skill Saving My Codebase"](https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules)라는 기사를 작성했습니다. 핵심 사항:
* 그는 원래 한 프로젝트에 50개 이상의 파일을 갖고 있었고, 각 파일의 길이는 100줄 미만이었습니다. - **일반적인 얕은 라이브러리**
* `/improve-codebase-architecture`에서 심화 후보 8명을 뽑았습니다.
* 그는 3개의 심화 항목을 선택했습니다(2개의 데이터 처리 모듈 병합, 1개의 도구 세트 병합).
* 결과: 파일 수가 50개 이상에서 30개 이상으로 줄었지만 **총 코드 크기는 기본적으로 동일하게 유지됩니다** - 소수의 심층 모듈로 복잡성이 압축되었습니다.
* 이 라이브러리에서 Claude의 후속 코드 변경 적중률이 크게 향상되었습니다. ("60%에서 90%로"라고 말했는데, 이는 엄격하게 측정되지는 않았지만 강하게 느껴졌습니다)
Mejba에는 **한 번에 8단계씩 심화하지 마세요**라는 알림도 있습니다. 한 번에 하나만 선택하고 테스트 + 커밋 + 관찰을 실행한 후 다음 항목을 선택하세요. 그렇지 않으면 완료한 후 롤백할 수 있는 방법이 없습니다.
## 설치 및 사용방법
```bash
npx skills@latest add mattpocock/skills
```
`improve-codebase-architecture` + `setup-matt-pocock-skills`을 확인하세요.
**전화**: `/improve-codebase-architecture`
**권장 리듬**:
* **일주일에 한 번 또는 각 스프린트가 끝날 때 실행**
* 또는 \*\*집약적 개발의 물결을 마친 후 한 번 실행합니다. (특히 AI 코드를 자주 작성한 후 얕은 모듈을 쌓기 쉽습니다.)
* **바쁠 때 뛰지 마세요** - 바쁠 때 소화할 시간이 없을 정도로 큰 변화를 제안합니다.
**일반적인 프로세스**:
1. `/improve-codebase-architecture`
2. AI 탐색 + N개 후보 나열(삭제 테스트 인수 포함)
3. 가장 기분이 좋은 것을 고르세요
4. 그릴 루프 정렬 디자인에 드롭
5. AI 구현 리팩토링([`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)과 함께 실행하는 것이 좋습니다. 리팩토링에는 테스트 보호 기능이 있어야 합니다)
6. 커밋 + 관찰
7. 일주일 뒤에 또 오세요
## 이 스킬이 Matt 작업 흐름의 "폐쇄 루프"인 이유는 무엇입니까?
Matt의 작업 흐름 다이어그램으로 돌아가서:
```
/grill-me → /to-prd → /to-issues → /tdd → /improve-codebase-architecture → 回到 /grill-me
```
시작점으로 되돌아가는 것에 주목하세요. `/improve-codebase-architecture`은(는) 일회성 도구가 아니며 **정기적인 유지 관리**입니다. 그 이유는 다음과 같습니다.
1. AI는 계속해서 코드 베이스에 얕은 모듈을 추가합니다. (이것이 기본 경향이며, 너무 많이 쓰면 쌓이게 됩니다.)
2. 비즈니스가 계속 발전함에 따라 오래된 이음새는 쓸모 없게 될 것입니다.
3. CONTEXT.md의 용어는 계속해서 날카로워지고 있으며 이전 이름은 유지되지 않습니다.
**이 스킬을 실행할 때마다 코드 베이스의 AI 친화성이 새로워집니다**. 이것이 LLM을 사용하여 장기간 코드베이스를 건강하게 유지하는 **유일한** 방법입니다. 새로 고치지 않으면 3개월 후에 코드베이스에서 AI가 종료됩니다.
## 스킬 자체보다 이런 사고방식이 더 가치있습니다
`/improve-codebase-architecture`을 전혀 설치하지 않더라도 다음 세 가지만 기억하시면 PR 리뷰의 질이 한 단계 향상됩니다.
1. **삭제 테스트**: 새 모듈을 볼 때마다 "삭제하면 복잡성이 사라지는가, 아니면 퍼질 것인가?"라고 자문해 보세요.
2. **최소 두 개의 어댑터가 실제 연결됨**: 단일 구현 인터페이스 = 거짓 추상, 삭제
3. **인터페이스는 테스트 표면입니다**: 측정할 수 없음 = 인터페이스 디자인에 문제가 있음
이 세 가지 항목에는 AI나 기술이 필요하지 않습니다. 이는 엔지니어링 미학의 경화입니다. Matt는 이를 일괄 실행을 위한 기술로 묶었지만 실제 활용은 세 가지 원칙 자체입니다.
## 메모
**너무 깊이 들어가지 마세요**. Ousterhout 자신은 딥 모듈이 교리라기보다는 목표라고 말했습니다. 모든 기능을 그 안에 집어넣는 대규모 Util 클래스는 딥 모듈이 아니라 신 모듈입니다. 판단 기준은 '간단한 인터페이스 + 응집력 있는 구현'이며 이 두 가지를 모두 충족해야 합니다.
**심화에는 테스트 보호 기능이 있어야 합니다**. 구조적 변경은 위험도가 높은 작업이며, 테스트 없이 대담하게 리팩토링하는 것 = 책임을 지기를 기다리는 것입니다. 현재 테스트가 없으면 [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)으로 이동하여 중요 경로에 테스트를 추가한 다음 다시 돌아옵니다.
**ADR 결정은 즉흥적으로 이루어져서는 안 됩니다**. 굽는 동안 후보자를 거부하면 AI는 쉽게 ADR 생성을 제안합니다. 이유가 실제로 "미래 사람들이 알아야 할 사항"인 경우에만 이를 수락합니다. 그렇지 않으면 docs/adr/이 저널 항목으로 채워집니다.
**코드 중 일부가 얕아도 상관없습니다**. 로거 래퍼, 상수 파일, 일회용 스크립트 등은 얕아서 문제가 되지 않습니다. 이 기술은 추상화에 도움이 되는 척하지만 실제로는 혼란을 더하는 얕은 모듈을 찾습니다.
## 참조 리소스
***
## 시리즈 결론
지금까지 6개의 글을 모두 읽었습니다. 전체 워크플로를 검토합니다.
```
/grill-me 或 /grill-with-docs ← 谈清楚要做什么
↓
/to-prd ← 凝固成 PRD
↓
/to-issues ← 切成 vertical slice
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture ← 周期性深化
↓
回到 /grill-me
```
이 프로세스의 정신은 한 문장으로 요약될 수 있습니다.
> \*\*AI는 지상의 전술적 군인이고, 당신은 전략적 계층입니다. '문제 정의', '문제 해체', '문제 테스트' 세 가지를 뒤로하고 스스로 해보고, '코드 작성'은 AI에게 맡기는 것이 AI 시대 엔지니어의 진정한 입장이다. \*\*
Matt의 일련의 기술은 궁극적인 답이 아니라 현 단계의 모범 사례입니다. 3개월 후에는 더 나은 것이 있을 수 있지만 영적인 측면은 변하지 않습니다. 좋은 코드 기반은 항상 나쁜 코드 기반보다 더 중요하며 기본 소프트웨어 기술은 항상 가치가 있습니다.
[개요](/ko/docs/notes/matt-pocock-skills/overview)로 돌아가거나, 가장 유용한 스킬을 선택하여 설치하여 사용해 보세요.
# 소프트웨어 기본 사항이 그 어느 때보다 중요합니다. Matt Pocock의 Claude Code 기술 세트
## 스펙 대 코드의 물결에 압도당했지만 여전히 차분한 사람
2026년은 AI 프로그래밍의 "사양에서 코드까지" 서술이 가장 중요한 해입니다. 사양을 작성하고, 컴파일러를 실행하고, 코드를 읽지 않고 사양을 작성한 다음 컴파일러를 실행합니다. 커뮤니티에 떠오른 슬로건은 "**code is cheap**"(코드는 싸다)이다. 즉, 어쨌든 AI는 초당 10,000줄을 더 생성할 수 있는데 왜 신경써야 하느냐는 뜻이다.
Matt Pocock은 이에 대해 공개적으로 반대하는 몇 안되는 사람 중 한 명입니다. 그는 AI 코딩이 매우 강력하다는 사실을 부정하지 않지만 실제로 자신의 수업 "실제 엔지니어를 위한 클로드 코드"에서 사양 대 코드를 테스트했으며 결론은 매우 가슴 뭉클했습니다. **실행할 때마다 코드가 점점 더 나빠집니다**. 이것이 바로 Pragmatic Programmer에서 언급된 "소프트웨어 엔트로피", 즉 소프트웨어 엔트로피 증가입니다.
그래서 그는 두 가지 일을 했습니다:
1. 이 관찰 내용을 18분짜리 강연으로 정리하세요. *소프트웨어 기본은 그 어느 때보다 중요합니다*.
2. 해당 해독제를 GitHub 저장소에 패키징합니다. [`mattpocock/skills`](https://github.com/mattpocock/skills) - "Skills for Real Engineers. Straight from my .claude 디렉터리."
창고는 2026년 2월 3일에 출시되었으며 4개월 만에 **61.1,000개의 별과 5.3,000개의 포크**에 도달했습니다. 같은 기간 동안 가장 빠르게 성장하는 AI 프로그래밍 웨어하우스 중 하나였습니다.
***
## 맷 포콕은 누구인가?
TypeScript를 작성해 본 적이 있다면 아마도 이 유형을 접했을 것입니다. 그는 최근 몇 년간 중국어와 영어계에서 가장 많은 TypeScript 교육을 펼친 사람 중 한 명입니다.
* 영어계에서 매우 인기 있는 유료 강좌 시리즈인 **TotalTypeScript.com**의 창립자
* **aihero.dev** 뉴스레터 구독자 60,000명 이상, 주제가 TS에서 AI 코딩으로 변경됨
* Twitter [`@mattpocockuk`](https://twitter.com/mattpocockuk) 및 YouTube `@mattpocockuk`에 짧은 동영상 튜토리얼이 많이 있습니다.
* OpenAI/Anthropic 사람이 아니며 순수 독립 개발자 + 교육자 배경입니다.
그의 성격은 매우 명확합니다. **선임 엔지니어의 관점에서 본 AI 코딩**. 우리는 “AGI가 온다”라고 외치지도 않고, “프로그래머들이 일자리를 잃을 것이다”라고 외치지도 않습니다. 그가 외친 것은 "구세대 소프트웨어 엔지니어의 비법은 여전히 매우 유용합니다. LLM이 실행할 수 있는 형식으로 변환하기만 하면 됩니다."였습니다.
***
## 핵심 주장: 코드는 저렴하지 않습니다
전체 연설에는 단 하나의 주장만 있으며 각 기술은 그에 대한 각주입니다.
> 코드 베이스 구조가 나쁘면 AI는 나쁜 코드 베이스에만 나쁜 코드를 작성합니다. 따라서 **좋은 코드 기반이 그 어느 때보다 중요하며 기본적인 소프트웨어 기술이 그 어느 때보다 중요합니다**.
Matt는 인간과 AI의 역할을 매우 간단하게 설명하기 위해 군사적 비유를 사용했습니다.
전략적 계층은 무엇을 하는가? 디자인 개념, 통합 언어, 모듈 경계 - 이 세 가지는 "코드 작성"이 아닌 "문제 정의"이며 LLM이 가장 잘 수행할 수 없는 작업입니다.
***
## 다섯 가지 실패 패턴 → 다섯 가지 고서 → 다섯 가지 기술
Matt는 연설에서 전체 방법론을 매핑 테이블로 압축했습니다. 실패 모드에 직면할 때마다 그는 20년 전에 해결된 고전 이론을 다시 지적한 다음 Markdown 형식의 Skill 파일을 제공합니다.
| # | AI 프로그래밍 실패 모드 | 고전 이론 및 출처 | 해당 스킬 |
| - | ------------------------- | ------------------------------------------------ | --------------------------------------------------------------------------------------------------- |
| 1 | AI는 당신이 원하는 것을 하지 않습니다 | *디자인의 디자인*(브룩스)—— 디자인 컨셉, 디자인 트리 | [`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me) |
| 2 | AI가 여러 가지 장황한 용어로 말을 건다 | *도메인 기반 디자인*(Evans) - 유비쿼터스 언어 | [`/grill-with-docs`](/ko/docs/notes/matt-pocock-skills/grill-with-docs) |
| 3 | AI는 제대로 작동하지만 실행할 수 없습니다 | *실용주의 프로그래머*(Hunt & Thomas)—— "피드백 속도가 속도 제한입니다" | [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd) |
| 4 | AI가 잘못된 코드 기반을 돌아다닌다 | *소프트웨어 디자인 철학*(Ousterhout) - 심층 모듈, 삭제 테스트 | [`/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture) |
| 5 | 당신의 두뇌는 AI 출력을 따라갈 수 없습니다 | 켄트 벡 —— 매일 디자인에 투자 | "인터페이스 디자인, 구현 위임" |
항목 5 저장소에는 별도의 기술이 없으며(예전에는 `design-an-interface`이 있었지만 더 이상 사용되지 않음) 그 정신은 [`/to-prd`](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues) 및 [`/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)에 흡수되었습니다. 둘 다 코드를 작성하기 전에 모듈 인터페이스에 대해 생각해야 합니다.
***
## 매일 실제로 사용하는 기술 5가지
연설은 철학적 뼈대였습니다. Matt는 나중에 aihero.dev에 "내가 매일 사용하는 5가지 상담원 기술"이라는 기사를 게시하여 이 뼈대를 일일 작업 흐름으로 전환했습니다. 이 5개는 앞으로 이 시리즈에서 하나씩 해체될 개체입니다.
```
/grill-me ← 先和 AI 谈清楚要做什么
↓
/to-prd ← 把对话凝固成 PRD
↓
/to-issues ← 把 PRD 切成可独立领取的 vertical slice
↓
/tdd ← 每个 slice 用红绿重构跑通
↓
/improve-codebase-architecture ← 周期性检查,把 shallow 模块改成 deep
```
이 5가지 기술이 함께 결합되어 Matt의 완전한 연구 개발 프로세스를 구성합니다. 각 단계에 해당하는 실패 모드는 이전 섹션의 표에 나와 있습니다.
각 품목의 자세한 분해(이 시리즈의 후속 페이지):
* [Grill Me: AI가 당신의 필요에 대해 고문하게 하세요](/ko/docs/notes/matt-pocock-skills/grill-me)
* [Grill With Docs: 프로젝트 언어 및 ADR 유지](/ko/docs/notes/matt-pocock-skills/grill-with-docs)
* [to-PRD + to-이슈: 대화에서 실행 티켓으로](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [TDD: AI가 적록 리팩토링을 통해 작은 단계를 밟도록 강제](/ko/docs/notes/matt-pocock-skills/tdd)
* [코드베이스 아키텍처 개선: 얕은 모듈을 깊은 모듈로 재구성](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)
***
## 설치 방법
웨어하우스 README에서는 설치를 위한 한 줄 명령을 제공합니다.
```bash
npx skills@latest add mattpocock/skills
```
이 명령은 다음을 수행합니다.
1. 어떤 스킬을 설치하고 싶은지 확인해보세요
2. 설치할 에이전트를 선택할 수 있습니다. (Claude Code, Codex, Cursor 등 모두 지원)
3. 해당 SKILL.md 파일을 `.claude/skills/`(또는 에이전트에 해당하는 디렉터리)에 넣습니다.
\*\*`/setup-matt-pocock-skills`\*\*을 동시에 확인하는 것이 좋습니다. 이는 세 가지 질문을 묻는 일회성 구성 기술입니다.
* **이슈 트래커**는 어떤 용도로 사용되나요? (GitHub / GitLab / 로컬 마크다운 / 기타)
* **Triage label**에는 어떤 단어가 사용되나요? (필요한 분류 또는 기타)
* **도메인 문서**를 어디에 둘 것인가? (CONTEXT.md/ADR 경로)
`/setup-matt-pocock-skills`을 한 번 실행하면 프로젝트 루트 디렉터리의 `AGENTS.md` 또는 `CLAUDE.md`에 기록됩니다. 그 후에는 모든 엔지니어링 기술(to-prd, to-issues, triage, tdd 등)이 자동으로 이 구성을 읽습니다. 이 단계는 생략되며 이후의 모든 스킬은 동일한 질문을 계속해서 묻습니다.
`/grill-me`(가장 가볍고 순수한 생산성 클래스)을 시도하고 싶다면 문제 추적기에 의존하지 않으므로 설정을 건너뛸 수 있습니다.
***
## 본 스킬셋과 BMAD / Spec-Kit / GSD의 차이점
[BMAD](/ko/docs/notes/speckit/concept), Spec-Kit, GSD와 같은 사양 기반 프레임워크를 이미 사용하고 있다면 "왜 Matt의 세트가 여전히 필요한가요?"라고 물을 수 있습니다.
Matt는 README에 다음과 같이 직접적으로 썼습니다.
**핵심 차이점**:
* BMAD/Spec-Kit/GSD는 사양에서 코드까지 완전한 파이프라인을 지정하는 **프레임워크**입니다. 그 과정을 따라야 합니다.
* Matt 이 세트는 **구성요소**입니다. 각 스킬에는 몇 줄에서 수십 줄에 이르는 마크다운 파일이 있습니다. 언제든지 분해하고 수정할 수 있습니다.
예: `grill-me`의 실제 전체 텍스트는 이 정도로 짧습니다.
```markdown
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
전체 스킬은 7라인입니다. 하지만 Claude가 결정을 내리기 전에 20, 50, 심지어 100가지 질문을 하게 만드는 것은 바로 이 7줄입니다. **아주 적은 텍스트를 사용하여 큰 행동 변화를 활용**한다는 디자인 철학이 이 기술 세트가 인기를 얻는 근본적인 이유입니다.
***
## 이 시리즈를 읽는 방법
이전에 Matt의 세트를 접해본 적이 없다면 Meta.json의 순서대로 읽는 것이 좋습니다.
1. **개요**(현재 보고 있는 기사) - 전체 그림 보기
2. **Grill Me** - 개별적으로 설치하여 먼저 사용해 보세요. 임계값이 가장 낮습니다.
3. **Grill With Docs** - grill-me의 고급 버전, CONTEXT.md 도입 시작
4. **to-PRD + to-Issues** - 대화를 실행 가능한 티켓으로 전환
5. **TDD** - Matt 자신이 "에이전트 출력 품질을 향상시키기 위해 사용해 본 방법 중 가장 안정적인 방법"이라고 말했습니다.
6. **코드베이스 아키텍처 개선** - AI를 장기간 사용할 수 있도록 정기적인 유지 관리
이미 Claude Code를 사용하여 실제 프로젝트를 작성하고 있다면 가장 직관적인 기사인 [`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me) + [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)로 바로 이동하세요.
가르치거나 글을 쓰고 있다면 [`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me)을 읽는 것만으로도 충분합니다. 이는 코드에만 국한되지 않는 일반적인 "설계 대화" 도구입니다.
***
## 내 사용 제안
이 기술 세트를 직접 설치한 후 가장 큰 물리적 변화는 다음과 같습니다.
**첫 번째**: 서두르지 말고 코드 작성을 시작하세요. 과거에는 AI가 "로그인 추가"를 수신하면 500줄을 레이아웃하기 시작했습니다. 이제 `/grill-me`이(가) 먼저 "기기를 기억하시겠습니까?"라는 20가지 질문을 할 것입니다. "세션이 만료되는 데 얼마나 걸리나요?" "계정 잠금에 몇 번이나 실패하셨나요?" 30분 후에 쓰도록 하세요. 당신이 절약하게 될 것은 후속 2시간의 재작업입니다.
**두 번째**: CLAUDE.md가 더 이상 부풀어오르지 않습니다. 예전에는 CLAUDE.md에 "코드를 작성하기 전에 요구사항을 이해해주세요", "지나치게 추상화하지 마세요" 등 금지사항이 많이 적혀 있었는데, 클로드는 그래도 이를 지켰습니다. Matt의 세트로 전환한 후 CLAUDE.md에는 도메인 지식(설계 시스템, 구성요소 사양, 배포)만 넣고 일반적인 방법론은 스킬에 넘겨줍니다. 양측의 책임은 분명합니다.
**셋째**: 심층적인 모듈 사고는 기술 자체보다 더 가치가 있습니다. `/improve-codebase-architecture`을 설치하지 않더라도 SKILL.md에서 "**삭제 테스트**"를 읽는 것만으로도(이 모듈을 삭제한 후 복잡성이 사라지면 통과되었음을 의미함) 이미 PR 검토 중에 다시 살펴보게 됩니다.
**비용 참고**:
* 스킬 5개를 설치하면 AI가 추가 질문을 하게 됩니다. "한 문장에 500줄을 생성하는 것"에 익숙한 사람들은 불편하다고 느낄 것입니다.
* `/tdd`이 엄격하게 구현된 후에는 먼저 테스트를 작성하기 위해 간단한 스크립트도 필요합니다. 이는 탐색 코드에 친숙하지 않습니다. "이번에는 TDD 건너뛰기"라고 말할 수 있습니다.
* `/grill-with-docs`이(가) CONTEXT.md를 수정하는 데 앞장설 것입니다. 처음 실행하기 전에 시험 실행하는 것이 가장 좋습니다.
***
## 참조 리소스
**연설에 인용된 5권의 책**(나오는 순서대로):
* *소프트웨어 디자인 철학* — John Ousterhout(복잡성 정의, 심층 모듈)
* *The Pragmatic Programmer* — David Thomas & Andrew Hunt(software entropy、outrunning headlights)
* *The Design of Design* — Frederick P. Brooks(design concept、design tree)
* *Domain-Driven Design* — Eric Evans(ubiquitous language)
* *Test-Driven Development* — Kent Beck(invest in design every day)
각각 20세 이상입니다. Matt는 자신의 연설에서 다음과 같은 문장을 여러 번 반복했습니다. "**Amazon에서 구매하세요.**" - 이 문장 자체가 이 연설의 부활절 달걀입니다.
# 기타 Skills:압축 커뮤니케이션, 인수인계, 교육, Skill 작성 및 안전 가드레일
## 개별적으로 다루지 않는 이유
Matt의 README는 skills를 세 가지 범주로 나눕니다:
* Engineering: 매일 실제 코드를 작성하는 데 사용
* Productivity: 일반적인 워크플로우 도구
* Misc: 그가 비축해 둔 작은 도구들
앞선 몇 개의 글에서는 메인 라인에 해당하는 engineering skill을 다루었습니다. 이 글에서는 나머지 productivity와 misc를 함께 다룰 것입니다. 왜냐하면 이들은 대부분 완전한 개발 프로세스가 아니라, **특정 시나리오에서 유용한 작은 스위치**이기 때문입니다.
만약 5개만 설치한다면, 여전히 다음을 우선적으로 설치하는 것을 추천합니다:
* [`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me)
* [`/grill-with-docs`](/ko/docs/notes/matt-pocock-skills/grill-with-docs)
* [`/to-prd` + `/to-issues`](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)
* [`/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)
하지만 메인 라인을 이미 실행하고 있다면, 다음 도구들이 일상 경험을 더욱 원활하게 만들어 줄 것입니다.
## Productivity Skills
### caveman: 커뮤니케이션 극단적 압축
`/caveman`은 "적은 토큰 모드"입니다. 에이전트에게 인사, 채움 단어, 과도한 설명, 모호한 완충을 제거하고 기술 정보만 남기도록 요구합니다.
다음과 같은 경우에 적합합니다:
* 빈번한 반복 작업 중이며 긴 응답을 읽고 싶지 않을 때
* 디버깅 시 사실, 원인, 다음 단계만 필요할 때
* 긴 컨텍스트가 거의 찼고 출력을 압축해야 할 때
* AI가 미사여구를 줄이도록 강제하고 싶을 때
이는 AI를 무례하게 만드는 것이 아니라, 더 짧은 문법으로 완전한 기술적 정확성을 유지하도록 하는 것입니다. 이 모드는 종료를 말할 때까지 계속 유효합니다.
이를 장기적인 기본값보다는 임시 설정으로 간주할 것입니다. 고위험 작업, 보안 경고, 복잡한 다단계 지시의 경우 너무 짧으면 오해하기 쉽습니다.
### handoff: 현재 대화를 다음 에이전트에게 인계
`/handoff`의 목표는 현재 대화를 인계 문서로 압축하여 현재 작업 공간을 오염시키지 않고 시스템 임시 디렉토리에 저장하는 것입니다.
다음 내용을 포함합니다:
* 현재 목표
* 결정된 사항
* 주요 경로 및 파일
* 남은 작업
* 다음에 호출할 skill 추천
* 민감 정보 비식별화
긴 작업 중단, 컨텍스트 부족, 또는 다른 에이전트에게 작업을 계속 맡기고 싶을 때 특히 유용합니다.
핵심은 PRD, 이슈, ADR, 커밋, diff에 이미 존재하는 내용을 복사하지 않고 경로 또는 URL만 참조하는 것입니다. 인계 문서의 가치는 프로젝트 문서를 다시 만드는 것이 아니라, **대화에 흩어진 상태를 보완**하는 것입니다.
### teach: 현재 디렉토리를 학습 작업 공간으로 만들기
`/teach`는 이 그룹에서 가장 무거운 도구입니다. 현재 디렉토리를 장기 학습 작업 공간으로 취급하며 다음을 유지합니다:
* `MISSION.md`: 이 주제를 배우는 이유
* `RESOURCES.md`: 고품질 리소스 목록
* `learning-records/*.md`: 학습 기록 (ADR과 유사)
* `lessons/*.html`: 각 세션별 인터랙티브 강의
* `reference/*.html`: 빠른 참조 자료
* `NOTES.md`: 교육 선호도 및 작업 노트
이것의 장점은 학습을 일회성 질문이 아닌 장기적인 시스템으로 본다는 것입니다. 특히 강조하는 점은:
* 미션 우선: 무엇을 배우는지보다 왜 배우는지가 더 중요합니다.
* 검색 연습: 회상 연습을 통해 장기 기억을 구축합니다.
* 간격/교차 학습: 단기적인 유창함에 속지 마십시오.
* 높은 신뢰도의 자료: 먼저 자료를 찾고, 모델의 기억에 의존하여 말하지 마십시오.
단순히 "X에 대해 설명해 줘"라고 묻는 경우에는 필요하지 않지만, 몇 주 동안 특정 주제를 배우고 싶다면 매우 적합합니다.
### write-a-skill: 새 skill 작성을 위한 스캐폴딩
`/write-a-skill`은 Matt가 skill 구조 자체를 추상화한 것입니다.
이것은 skill이 최소한 다음을 포함해야 한다고 요구합니다:
```text
skill-name/
├── SKILL.md
├── REFERENCE.md
├── EXAMPLES.md
└── scripts/
```
물론 뒤의 세 가지는 필수가 아니며, 내용이 너무 길거나, 예시가 가치 있거나, 작업이 스크립트화될 수 있을 때만 추가합니다.
가장 중요한 판단은 `description`이 에이전트가 skill을 로드할지 결정할 때 유일하게 먼저 보는 정보라는 것입니다. 따라서 description은 "문서 처리를 돕는다"와 같은 공허한 말로 작성해서는 안 되며, 다음을 명확히 설명해야 합니다:
* 어떤 기능을 제공하는가
* 언제 트리거되는가
* 트리거 단어 또는 컨텍스트는 무엇인가
이는 제가 skill을 작성하는 경험과 일치합니다. 많은 skill이 실패하는 이유는 본문이 잘못 작성되었기 때문이 아니라, description이 너무 광범위하게 작성되어 에이전트가 언제 로드해야 할지 알지 못하기 때문입니다.
## Misc Skills
### git-guardrails-claude-code: 위험한 git 명령 차단
이 skill은 Claude Code에 `PreToolUse` 훅을 추가하여 Bash 실행 전에 위험한 git 명령을 차단합니다.
기본적으로 다음을 차단합니다:
* `git push`
* `git reset --hard`
* `git clean -f` / `git clean -fd`
* `git branch -D`
* `git checkout .` / `git restore .`
그 가치는 매우 직접적입니다: 에이전트가 귀하의 승인 없이 푸시하거나, 하드 리셋하거나, 추적되지 않은 파일을 삭제하는 것을 방지합니다.
AI가 실제 저장소에서 작업하도록 자주 하는 경우, 이 skill은 설치할 가치가 있습니다. AI를 신뢰하지 않는 것이 아니라, 고위험 작업은 도구 계층에서 차단하고 프롬프트에만 의존하지 않는 것입니다.
### setup-pre-commit: 프로젝트에 커밋 전 검사 추가
`/setup-pre-commit`은 다음을 설정합니다:
* Husky pre-commit hook
* lint-staged + Prettier
* typecheck
* test
먼저 패키지 관리자를 감지한 다음, 프로젝트에 이미 있는 스크립트에 따라 pre-commit에서 실행해야 할 것을 결정합니다. `typecheck` 또는 `test`가 없을 경우 강제로 생성하지 않고 생략하고 알려줍니다.
이 skill의 가치는 설정 자체에 있는 것이 아니라, Matt의 품질 관점에 있습니다: **AI가 코드가 문제없다고 말하게 하는 것에만 의존하지 말고, 결정적인 검사를 통과하게 하십시오.**
### migrate-to-shoehorn: 테스트에서 `as` 줄이기
이것은 매우 Total TypeScript 스타일의 작은 도구입니다. 테스트에서 TypeScript `as` 타입 단언을 `@total-typescript/shoehorn`으로 마이그레이션합니다.
일반적인 대체:
| 이전 문법 | 새 문법 | 시나리오 |
| --------------------------- | ------------------ | -------------------------------- |
| `obj as Request` | `fromPartial(obj)` | 테스트에서 큰 객체의 몇 가지 필드만 관심 있을 때 |
| `obj as unknown as Request` | `fromAny(obj)` | 의도적으로 잘못된 타입을 전달하여 오류 경로를 테스트할 때 |
| 전체 객체 가짜 데이터 | `fromExact(obj)` | 완전한 모양을 강제해야 할 때 |
이것은 프로덕션 코드가 아닌 테스트 코드에만 사용하도록 명확히 되어 있습니다.
이 skill은 매우 제한적이지만, Matt의 엔지니어링 취향과 잘 맞습니다: 타입 시스템을 위해 테스트에서 20개의 무의미한 필드를 만들지 않고, 맨 `as`를 사용하여 타입 안전성을 완전히 비활성화하지도 않습니다.
### scaffold-exercises: 강의 저장소에 연습 디렉토리 생성
이 skill은 분명히 Matt가 강의를 만드는 작업 흐름에서 나온 것입니다. 규격에 따라 다음을 생성합니다:
```text
exercises/
└── 05-memory-skill-building/
└── 05.02-short-term-memory/
├── explainer/
├── problem/
└── solution/
```
각 하위 디렉토리는 최소한 비어 있지 않은 `readme.md`를 포함하고, 필요한 경우 `main.ts`를 포함하며, `pnpm ai-hero-cli internal lint`를 통과해야 합니다.
대부분의 엔지니어링 프로젝트에는 유용하지 않지만, 강의, 부트캠프, 연습 저장소에는 매우 실용적입니다. 더 중요한 것은, 좋은 skill의 특징을 보여준다는 것입니다: **반복적이고 기계적이며 세부 사항을 놓치기 쉬운 형식 작업을 에이전트에게 맡기는 것**입니다.
## 현재 메인 라인에 포함하지 않는 이유
상위 저장소에는 `deprecated/`, `in-progress/`, `personal/` 디렉토리가 더 있습니다.
이들을 공식 사용 가이드에 포함하지 않는 것을 권장합니다:
| 디렉토리 | 메인 라인에 포함하지 않는 이유 |
| -------------- | ------------------------------------------------- |
| `deprecated/` | 폐기되었으므로 독자가 오래된 프로세스를 계속 사용하는 것을 오도하기 쉽습니다. |
| `in-progress/` | 아직 실험 중이며, 동작과 명칭이 변경될 수 있습니다. |
| `personal/` | Matt 자신의 개인 작업 공간에 더 가깝고, 일반 독자에게 적합하지 않을 수 있습니다. |
나중에 작성해야 한다면, 안정적인 추천 사항에 혼합하는 대신 "Matt Pocock skills 저장소 고고학"과 같은 별도의 글을 작성할 수 있습니다.
## 이 작은 도구들의 공통점
이 skill들은 매우 산발적으로 보이지만, 그 뒤에는 동일한 원칙이 있습니다:
> 에이전트가 쉽게 벗어날 수 있는 것을 작고 명확한 작업 패턴으로 만듭니다.
* `caveman`은 커뮤니케이션의 벗어남을 방지합니다.
* `handoff`는 컨텍스트 손실을 방지합니다.
* `teach`는 학습이 일회성 질문으로 변하는 것을 방지합니다.
* `write-a-skill`은 skill 구조가 즉흥적으로 작성되는 것을 방지합니다.
* `git-guardrails`는 위험한 명령이 자율에 의존하는 것을 방지합니다.
* `setup-pre-commit`은 품질 검사가 AI의 자체 보고에 의존하는 것을 방지합니다.
* `migrate-to-shoehorn`은 테스트 타입 단언이 통제 불능이 되는 것을 방지합니다.
* `scaffold-exercises`는 강의 구조가 수동으로 누락되는 것을 방지합니다.
이것이 Matt의 이 repo에서 가장 배울 만한 점입니다: skill은 거대할 필요가 없습니다. 자주 발생하는 작은 편차가 20줄의 명령으로 안정적으로 수정될 수 있다면, skill로 작성할 가치가 있습니다.
## 참고 자료
# 프로토타입: 폐기 가능한 코드로 디자인 질문에 답하기
## 프로토타입은 "일단 대충 작성하는 것"이 아닙니다.
`/prototype`의 첫 번째 정의는 중요합니다.
> 프로토타입은 질문에 답하기 위해 작성된 폐기 가능한 코드입니다.
이 문장은 프로토타입을 "게으른 구현"과 구분합니다. 프로토타입은 프로덕션 코드의 전신이 아니며, 나중에 공식 버전으로 수정될 반제품도 아닙니다. 처음부터 폐기 가능한 것으로 표시되어야 합니다.
따라서 `/prototype`의 핵심은 빠르게 작성하는 것이 아니라 먼저 명확히 하는 것입니다.
> 이 프로토타입은 어떤 질문에 답해야 하는가?
## 두 가지 분기
Matt는 프로토타입을 두 가지 유형으로 나누며, 출력은 완전히 다릅니다.
| 답해야 할 질문 | 분기 | 출력 |
| ------------------------ | --------------- | ------------------------ |
| 논리, 상태 머신, 데이터 모델이 합리적인가 | Logic prototype | 실행 가능한 터미널 미니 프로그램 |
| 이 인터페이스는 어떻게 보여야 하는가 | UI prototype | 라우트 내에서 전환 가능한 여러 UI 솔루션 |
이것은 매우 실용적입니다. 많은 팀이 "프로토타입을 만들자"고 말하지만, 인터페이스 모양을 검증하고 싶은 것인지, 상태 흐름을 검증하고 싶은 것인지 명확히 하지 않습니다. 둘 다 다른 것을 필요로 합니다.
## Logic prototype: 상태를 터미널에 펼치기
질문이 "이 상태 머신이 올바른가", "이 비즈니스 규칙이 실행될 수 있는가"라면, 프로토타입은 매우 작은 명령줄 프로그램이어야 합니다.
특징:
* 메모리 내 상태이며 실제 데이터베이스에 의존하지 않습니다.
* 하나의 명령으로 시작합니다.
* 각 작업 후 관련 전체 상태를 출력합니다.
* 종이로 추론하기 어려운 분기를 다룹니다.
* 테스트를 작성하지 않고, 예외 처리를 하지 않으며, 프레임워크로 추상화하지 않습니다.
예시: 구독 상태 흐름을 설계해야 합니다.
프로덕션 코드를 직접 수정하지 마세요. 먼저 `subscription-prototype.ts`를 작성하여 사용자가 터미널에서 선택할 수 있도록 하세요.
```text
1. start trial
2. pay
3. cancel
4. expire
5. refund
6. print state
```
각 단계를 누를 때마다 현재 권한, 시험판 할당량, 유료 상태, 다음 갱신을 출력합니다. 일부 상태 조합을 전혀 생각하지 않았다는 것을 빠르게 알게 될 것입니다.
이러한 유형의 프로토타입의 가치는 다음과 같습니다. **추상 규칙을 조작 가능한 객체로 만듭니다.**
## UI prototype: 라우트 내에 여러 급진적인 솔루션 배치
질문이 "인터페이스를 어떻게 디자인해야 하는가"라면, 프로토타입은 동일한 솔루션을 세 번 미세 조정하는 것이 아니라, 충분히 다른 여러 UI 세트를 생성해야 합니다.
`/prototype`의 UI 분기는 다음을 요구합니다.
* 하나의 라우트 내에 여러 변형을 배치합니다.
* URL 검색 매개변수 또는 하단 플로팅 전환 막대를 사용하여 전환합니다.
* 솔루션 간에 명확한 차이가 있어야 합니다.
* 프로토타입 코드는 미래의 실제 페이지에 가깝지만, 명확하게 프로토타입임을 나타내는 명명 규칙을 사용합니다.
* 너무 일찍 실제 데이터 및 지속성 연결을 하지 않습니다.
이것은 일반적인 AI UI 생성과 다릅니다. AI가 한 번에 "최적의 솔루션"을 제공하도록 하는 것이 아니라, 실제 브라우저를 사용하여 여러 방향을 비교하도록 합니다.
예를 들어, 대시보드의 빈 상태에 대해 AI에게 문구만 수정하도록 하지 마세요. 다음과 같이 만들 수 있습니다.
* A: 표 형식, 높은 밀도, 다음 작업 강조
* B: 작업 중심, 왼쪽 체크리스트 + 오른쪽 미리보기
* C: 안내식, 주요 CTA 및 이전 예시 강조
그런 다음 채팅에서 세 개의 스크린샷을 보고 상상하는 대신, 동일한 라우트에서 전환합니다.
## 모든 프로토타입은 삭제 가능해야 합니다.
`/prototype`의 일반 규칙 중에서 가장 중요한 것은 "삭제 가능성"입니다.
* 파일 이름이나 경로에 프로토타입임을 명시합니다.
* 기본적으로 프로덕션 데이터베이스에 연결하지 않습니다.
* 과도한 일반 추상화를 작성하지 않습니다.
* 과도한 오류 처리를 하지 않습니다.
* 완료 후 삭제하거나 배운 내용을 공식 코드로 흡수합니다.
프로토타입을 삭제할 수 없다면, 이미 프로덕션 코드 부채가 된 것입니다.
이것은 AI 프로그래밍에서 특히 중요합니다. AI는 프로토타입을 "작동하는 것처럼 보이게" 만드는 데 능숙하며, 그러면 인간은 삭제하기를 꺼리고 결국 아무도 건드리지 않는 임시 코드 더미가 프로젝트에 남게 됩니다.
## 프로토타입 완료 후 무엇을 남겨야 하는가
프로토타입 코드는 보존할 가치가 없지만, 답변은 보존할 가치가 있습니다.
Matt는 다음 사항을 영구적인 위치에 기록할 것을 권장합니다.
* 프로토타입이 답해야 했던 질문
* 관찰된 결론
* 선택한 방향
* 포기한 방향
* 필요한 경우 ADR, 이슈, PRD 또는 커밋 메시지로 전환
즉, `/prototype`의 산출물은 코드가 아니라 **결정**입니다.
## 사용하지 말아야 할 때
다음 시나리오에서는 `/prototype`을 사용하지 마세요.
* 요구 사항이 명확하며 구현만 필요합니다.
* 버그가 재현되었으며 [`/diagnose`](/ko/docs/notes/matt-pocock-skills/diagnose-and-triage)를 사용해야 합니다.
* 리팩토링 방향이 명확하며 [`/tdd`](/ko/docs/notes/matt-pocock-skills/tdd)로 보호하면서 구현해야 합니다.
* UI가 사소한 개선일 뿐이며 여러 솔루션을 만들 가치가 없습니다.
* 삭제하거나 흡수할 시간이 없습니다.
프로토타입의 비용은 작성하는 데 있는 것이 아니라 마무리하는 데 있습니다. 마무리가 없다면 시작하지 마세요.
## 유용한 프롬프트
이렇게 호출할 수 있습니다.
```text
/prototype
이 체크아웃 상태 머신이 합리적인지 검증하고 싶습니다. logic 분기를 사용해 주세요.
폐기 가능한 터미널 프로토타입만 만들고 실제 DB는 연결하지 마세요.
각 작업 후 전체 상태를 출력해 주세요.
```
또는:
```text
/prototype
프로젝트 상세 페이지의 3가지 정보 아키텍처를 비교하고 싶습니다. UI 분기를 사용해 주세요.
기존 라우트 시스템 내의 prototype 라우트에 배치하고 하단 전환 막대를 제공해 주세요.
프로덕션 컴포넌트는 수정하지 마세요.
```
여기서 가장 중요한 것은 "답해야 할 질문"을 명확히 하는 것입니다. 이 질문이 명확하면 프로토타입이 길을 잃기 쉽지 않습니다.
## Grill Me와의 관계
[`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me)는 질문을 통해 결정을 수렴하는 데 적합하며, `/prototype`은 체험을 통해 결정을 수렴하는 데 적합합니다.
어떤 질문은 질문만으로 해결될 수 있습니다. 예를 들어 "익명 댓글은 검토해야 하는가?"입니다. 어떤 질문은 직접 만져봐야 합니다. 예를 들어 "이 드래그 앤 드롭 정렬 상태 머신이 실제로 사용하기 어려울까?"입니다. 후자는 프로토타입을 사용해야 합니다.
따라서 저는 이를 워크플로우의 분기점으로 배치할 것입니다.
```text
아이디어가 모호함
↓
/grill-me
↓
여전히 체험 또는 검증이 필요한 경우
↓
/prototype
↓
결론을 보존하고 프로토타입 삭제
↓
/to-prd 또는 /tdd
```
## 참고 자료
다음 글: [Improve Codebase Architecture: shallow 모듈을 deep 모듈로 리팩토링](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture).
# Setup Matt Pocock Skills: 먼저 프로젝트 규칙을 명확히 쓰기
## 이 Skill은 설치 문제를 해결하지 않습니다
`/setup-matt-pocock-skills`는 "설치 후 실행하는 초기화 명령"으로 오해하기 쉽습니다. 실제로는 **프로젝트 계약 생성기**에 더 가깝습니다. 후속 스킬에게 이 저장소가 어떻게 작업을 추적하고, 이슈를 어떻게 태그하며, 도메인 언어와 아키텍처 결정을 어디서 읽을지 알려줍니다.
Matt는 README의 Quickstart에서 설치 시 `/setup-matt-pocock-skills`를 선택한 다음 에이전트에서 실행하라고 특별히 강조합니다. 이유는 간단합니다. `to-prd`, `to-issues`, `triage`, `diagnose`, `tdd`, `improve-codebase-architecture`, `zoom-out` 모두 동일한 프로젝트 컨텍스트를 필요로 합니다. 각 스킬이 임시로 물어보면 프로세스가 너무 파편화될 것입니다.
이것은 "Claude의 선호도를 설정"하는 것이 아니라 세 가지 엔지니어링 질문에 답하는 것입니다.
| 질문 | 무엇을 명확히 해야 하는가 | 누가 나중에 사용하는가 |
| ------------------ | ---------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| 이슈 트래커는 어디에 있는가 | GitHub, GitLab, 로컬 markdown 또는 다른 시스템 | `to-prd`, `to-issues`, `triage` |
| 트리아지 태그는 어떻게 매핑되는가 | `needs-triage`, `needs-info`, `ready-for-agent` 등 역할에 해당하는 실제 태그 | `triage` |
| 도메인 문서는 어디에 있는가 | 단일 `CONTEXT.md` 또는 다중 컨텍스트 `CONTEXT-MAP.md` + 분할 ADR | `grill-with-docs`, `diagnose`, `tdd`, `zoom-out`, `improve-codebase-architecture` |
## 왜 중요한가
이 스킬 세트의 핵심 아이디어는 "작고 조합 가능함"입니다. 작다는 것은 다음을 의미합니다. 전체 프로젝트 프로세스를 스스로 관리하고 싶어하지 않으므로 프로젝트 내의 실제 약속을 알아야 합니다.
예시: `/to-issues`는 이슈를 생성해야 합니다. 설정이 없으면 다음을 알아야 합니다.
* `gh issue create` 호출
* `glab issue create` 호출
* `.scratch//`에 쓰기
* 또는 Linear/Jira 복사 가능한 텍스트 생성
또는 `/triage`는 이슈를 `ready-for-agent`로 이동해야 합니다. 저장소의 실제 태그가 `ai:ready`인데 스킬이 `ready-for-agent`라는 새 태그를 생성하면 이슈 트래커가 즉시 더러워집니다.
따라서 `/setup-matt-pocock-skills`의 가치는 자동화가 아니라 **암묵적인 약속을 명시적으로 만드는 것**입니다.
## 무엇을 읽는가
이 스킬은 먼저 저장소를 탐색하며 가정하지 않습니다.
* `git remote -v` 및 `.git/config`: GitHub/GitLab 프로젝트인지 판단
* 루트 디렉토리 `AGENTS.md` / `CLAUDE.md`: 이미 `## Agent skills` 섹션이 있는지 확인
* 루트 디렉토리 `CONTEXT.md` / `CONTEXT-MAP.md`: 도메인 언어 문서 형식을 판단
* `docs/adr/` 및 `src/*/docs/adr/`: ADR이 전역인지 모듈 수준인지 판단
* `docs/agents/`: 이미 설정이 실행되었는지 확인
* `.scratch/`: 로컬 markdown 이슈 약속이 이미 있는지 판단
이는 Matt의 전체 워크플로 스타일과 일치합니다. **먼저 프로젝트의 실제 상태를 보고 규칙을 작성합니다.**
## 세 가지 결정
### 1. 이슈 트래커
이것은 후속 작업 단위가 구현되는 위치입니다.
기본값은 GitHub입니다. 이 스킬 세트가 GitHub Issues를 중심으로 처음 설계되었기 때문입니다. 그러나 GitLab과 로컬 markdown도 동등한 선택지로 고려했습니다.
| 선택 | 어떤 시나리오에 적합한가 |
| -------------- | --------------------------------------------------- |
| GitHub | 오픈 소스 프로젝트, GitHub 이슈 워크플로가 이미 존재 |
| GitLab | 회사 프로젝트가 GitLab에 있고 `glab` 사용에 익숙함 |
| Local markdown | 개인 프로젝트, 임시 탐색, 원격 이슈 트래커 없음 |
| Other | Jira, Linear, Feishu, 다차원 표 등, 실제 프로세스를 텍스트로 기록해야 함 |
핵심은 무엇을 선택하느냐가 아니라 **팀이 실제로 사용하는 것**을 선택하는 것입니다. 잘못 선택하면 후속 스킬이 잘못된 시스템에서 작업을 생성합니다.
### 2. 트리아지 태그 어휘
`/triage`는 내부적으로 5가지 상태 역할을 사용합니다.
| 역할 | 의미 |
| ----------------- | --------------------- |
| `needs-triage` | 유지보수자가 판단 대기 중 |
| `needs-info` | 보고자가 정보 보충 대기 중 |
| `ready-for-agent` | AFK 에이전트에게 전달할 만큼 명확함 |
| `ready-for-human` | 인간의 판단 또는 구현 필요 |
| `wontfix` | 처리하지 않음 |
설정은 이러한 역할에 해당하는 실제 태그 이름을 묻습니다. 프로젝트에 기존 태그가 없으면 기본 이름을 사용하면 됩니다. 이미 자체 명명 체계가 있다면 스킬이 새 것을 만들지 않도록 여기서 매핑해야 합니다.
### 3. 도메인 문서
이것은 Matt의 스킬 세트와 일반 프롬프트의 가장 큰 차이점입니다. 현재 대화만 보는 것이 아니라 프로젝트의 **도메인 언어**와 **아키텍처 결정**을 읽습니다.
가장 간단한 형태:
```text
/
├── CONTEXT.md
└── docs/
└── adr/
```
대형 모노레포는 다중 컨텍스트를 사용할 수 있습니다.
```text
/
├── CONTEXT-MAP.md
├── apps/
│ └── web/
│ ├── CONTEXT.md
│ └── docs/adr/
└── services/
└── billing/
├── CONTEXT.md
└── docs/adr/
```
설정은 지금 모든 문서를 작성하도록 강요하는 것이 아니라 후속 스킬에게 어디서 찾아야 하는지, 찾지 못했을 때 어떻게 생성해야 하는지를 알려줍니다.
## 무엇을 작성하는가
최종적으로 두 가지 유형의 출력이 있습니다.
첫 번째 유형은 `AGENTS.md` 또는 `CLAUDE.md`의 `## Agent skills` 섹션입니다.
```markdown
## Agent skills
### Issue tracker
...
### Triage labels
...
### Domain docs
...
```
두 번째 유형은 `docs/agents/` 아래의 세 가지 설명입니다.
| 파일 | 내용 |
| ------------------------------ | ----------------------------------------- |
| `docs/agents/issue-tracker.md` | 이슈 시스템, 명령, 생성/업데이트 약속 |
| `docs/agents/triage-labels.md` | 표준 역할과 실제 태그 간의 매핑 |
| `docs/agents/domain.md` | `CONTEXT.md`, `CONTEXT-MAP.md`, ADR 읽기 규칙 |
기존 `CLAUDE.md`를 우선적으로 편집하며, `CLAUDE.md`가 없으면 `AGENTS.md`를 고려합니다. 이는 **프로젝트에 서로 경쟁하는 두 개의 에이전트 규칙 진입점을 만들지 않는다**는 중요한 절제를 반영합니다.
## 사용 권장 사항
Matt의 스킬 세트를 처음 설치할 때 순서는 다음과 같습니다.
1. 설치: `npx skills@latest add mattpocock/skills`
2. `/setup-matt-pocock-skills` 선택
3. `/setup-matt-pocock-skills` 실행
4. 실제 프로젝트 상태에 따라 이슈 트래커, 태그, 도메인 문서 세 가지 질문에 답합니다.
5. 생성된 `## Agent skills` 및 `docs/agents/*.md`를 확인합니다.
6. 그런 다음 [`/grill-with-docs`](/ko/docs/notes/matt-pocock-skills/grill-with-docs), [`/to-prd`](/ko/docs/notes/matt-pocock-skills/to-prd-and-issues), [`/triage`](/ko/docs/notes/matt-pocock-skills/diagnose-and-triage)를 사용합니다.
[`/grill-me`](/ko/docs/notes/matt-pocock-skills/grill-me)만 개별적으로 경험하고 싶다면 설정을 건너뛸 수 있습니다. 하지만 엔지니어링 프로세스에 진입한다면 먼저 하는 것이 좋습니다.
## 이 Skill의 설계 영감
`/setup-matt-pocock-skills`는 매우 단순해 보이지만 에이전트 워크플로에서 가장 흔한 문제 중 하나를 해결합니다. **규칙이 사람의 머릿속에 흩어져 있다는 것**입니다.
많은 팀은 "어떤 태그를 사용하는가", "어떤 이슈를 AI에게 줄 수 있는가", "CONTEXT.md는 어디에 있는가"와 같은 정보를 구두 약속으로 간주합니다. 사람은 알지만 AI는 모릅니다. AI는 모르기 때문에 계속 묻거나, 더 나쁘게는 스스로 추측합니다.
설정의 역할은 이러한 구두 약속을 읽을 수 있는 파일로 만드는 것입니다. 후속 스킬은 더 똑똑해질 필요 없이 동일한 프로젝트 계약을 안정적으로 읽기만 하면 됩니다.
이것이 제가 이 스킬에 대해 별도의 글을 쓸 가치가 있다고 생각하는 이유입니다. 이것은 과시적인 스킬이 아니라 전체 워크플로가 장기적으로 실행될 수 있도록 하는 기반입니다.
## 참고 자료
다음 글: [Grill Me: AI가 코드를 작성하기 전에 50가지 질문으로 당신을 심문하게 하세요](/ko/docs/notes/matt-pocock-skills/grill-me).
# TDD: 적록 리팩토링을 사용하여 AI가 작은 조치를 취하도록 강제
## 실패 모드: "AI가 옳은 일을 하지만 실행할 수 없습니다."
Matt의 강연에 나오는 세 번째 실패 모드: **방향은 옳지만 작동하지 않습니다**.
가장 직접적인 해결책은 AI용 피드백 인프라를 설치하는 것입니다.
* TypeScript(정적 타이핑이 없으면 *이상해*)
* LLM이 브라우저에 액세스하여 페이지 자체를 볼 수 있도록 허용
* 자동화된 테스트
그러나 Matt는 한 가지 사실을 발견했습니다. **이 피드백을 설치해도 LLM이 제대로 작동하지 않습니다**. 한 번에 500줄을 쓰다가 "아, 그거 입력해야지"라고 생각하는 경향이 있습니다. 이것이 실용주의 프로그래머가 *헤드라이트를 앞지르기*라고 부르는 것입니다. 헤드라이트가 밝힐 수 있는 것보다 더 빠르게 운전하십시오. 그리고 벽에 부딪히는 것은 시간 문제일 뿐입니다.
> "The rate of feedback is your speed limit, which means you should be testing as you go, taking small deliberate steps. **And the AI by default is really not very good at that.**"
이 문제를 해결하려면 도구 수준에서 AI를 단계별로 강제로 중지해야 합니다. Matt의 대답은 TDD입니다. **먼저 테스트하면 체크포인트가 강제 실행될 수 있습니다**.
## 고전 이론: 켄트 벡(Kent Beck)의 적록 재구성
TDD의 표준 리듬은 Kent Beck이 2003년 저서 "Test-Driven Development: By 예제"에서 정의했습니다.
1. **빨간색**: 실패한 테스트 작성(무엇을 해야 할지 설명)
2. **녹색**: 테스트를 통과할 만큼 충분히 작은 코드를 작성합니다.
3. **REFACTOR**: 테스트 보호 중인 코드 구조 개선
각 루프는 몇 분 정도로 매우 짧습니다. 모든 단계에는 자동화된 검사(테스트 통과/실패)가 있습니다.
Matt는 이 리듬을 직접 따르지만 그의 SKILL.md는 **반패턴**에 대해 이야기하는 데 많은 시간을 소비합니다. 이것이 핵심입니다.
## 주요 안티 패턴: 빨간색과 녹색을 가로로 슬라이스
많은 사람들은 TDD가 "모든 테스트를 먼저 작성한 다음 모든 구현을 작성"하는 것을 의미한다고 생각합니다. Matt는 SKILL.md에서 이것이 잘못되었다고 직접 말합니다.
```
WRONG (horizontal slicing):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical slicing via tracer bullets):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
왜 수평이 틀렸나요? SKILL.md는 세 가지 이유를 제시했습니다.
> 1. Tests written in bulk test *imagined* behavior, not *actual* behavior
> 2. You end up testing the *shape* of things (data structures, function signatures) rather than user-facing behavior
> 3. Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
인간의 말: **모든 테스트를 한 번에 작성하는 것은 실제 코드가 아니라 머릿속에 있는 것을 테스트하는 것입니다**. impl3을 작성하다 보면 test1의 디자인이 잘못되었음을 깨닫게 됩니다. 하지만 이때 test2/test3/test4는 모두 잘못된 디자인과 결합되어 있습니다. 돌아가서 변경하세요.
올바른 접근 방식은 하나의 구현을 테스트한 후 한 쌍을 작성한 후 다음 쌍을 여는 것입니다. 각 쌍이 완료되고 이 구현에서 무언가를 배운 후에는 상상이 아닌 실제 경험을 기반으로 다음 테스트 쌍을 설계할 수 있습니다.
## 스킬 전문 구조
`engineering/tdd/SKILL.md`은 TDD 자체에 많은 뉘앙스가 있기 때문에 Matt가 작성한 가장 긴 기술 중 하나입니다. 핵심 구조는 다음과 같습니다.
### 철학
> **Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
> **Good tests** are integration-style: they exercise real code paths through public APIs. They describe *what* the system does, not *how* it does it. A good test reads like a specification.
> **Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly). The warning sign: your test breaks when you refactor, but behavior hasn't changed.
진단을 기억하세요. **내부 함수의 이름을 바꾸면 테스트가 중단됩니다. 그러면 이 테스트는 동작이 아닌 구현을 테스트하는 것이므로 나쁜 테스트입니다**.
### 워크플로(체크리스트 포함)
#### 1. Planning
코드를 작성하기 전에 사용자와 일치시키십시오.
```
[ ] Confirm with user what interface changes are needed
[ ] Confirm with user which behaviors to test (prioritize)
[ ] Identify opportunities for deep modules (small interface, deep impl)
[ ] Design interfaces for testability
[ ] List the behaviors to test (not implementation steps)
[ ] Get user approval on the plan
```
핵심 질문: "**공개 인터페이스는 어떤 모습이어야 합니까? 어떤 동작을 테스트해야 합니까?**"
> "**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case."
이것은 매우 반직관적입니다. 기본적으로 AI는 모든 극단적인 경우를 소진하기를 원하지만 Matt는 **우선순위**를 강조합니다. 모든 행동을 측정할 가치가 있는 것은 아니며 핵심 경로에 화력을 집중합니다.
#### 2. Tracer Bullet
**한가지** 사항을 확인하는 \*\*테스트를 작성하세요.
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
이것은 "추적 총알"입니다. 먼저 쏘고 시력을 확인하십시오. Matt는 이 작업이 **엔드 투 엔드**여야 한다고 강조했습니다. 스키마를 먼저 작성한 다음 API를 작성한 다음 UI를 작성하는 것이 아니라 전체 스택을 통과하는 가장 얇은 경로를 자르는 것입니다.
#### 3. Incremental Loop
이후의 각 동작에 대해 RED→GREEN을 반복합니다.
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
규칙:
* 한 번에 하나의 테스트
* 현재 테스트를 통과할 수 있을 만큼만 코드를 작성하세요.
* **향후 테스트를 예측하지 마세요**
* 테스트는 관찰 가능한 동작에 중점을 둡니다.
"예측하지 마십시오"가 특히 중요합니다. AI는 어쩔 수 없이 "이 함수는 X를 지원해야 하는데, 그런데 추가하자"라고 생각하고, 이것이 수평 슬라이싱을 시작합니다.
#### 4. Refactor
모든 테스트를 통과한 후 리팩토링 기회를 찾으십시오.
```
[ ] Extract duplication
[ ] Deepen modules (move complexity behind simple interfaces)
[ ] Apply SOLID principles where natural
[ ] Consider what new code reveals about existing code
[ ] Run tests after each refactor step
```
> **Never refactor while RED.** Get to GREEN first.
빨간색으로 리팩토링 = 테스트와 코드를 동시에 변경 = 테스트가 잘못된 것인지 코드가 잘못된 것인지 알 수 없습니다. **먼저 녹색을 그린 다음 리팩터링합니다**.
### Per-Cycle Checklist
각각의 빨간색과 녹색 주기가 끝나면 Matt는 AI에게 자체 점검을 요청합니다.
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
이 5가지 포인트는 잘못된 테스트와 과도한 구현을 식별하는 데 사용됩니다. AI 자가 점검은 가장 흔히 발생하는 실수를 방지할 수 있습니다.
## 실사용 : 발행부터 홍보까지
`/tdd`은 `/to-issues`의 Matt 워크플로의 다음 단계입니다. 수직 슬라이스 문제가 있는 경우 프로세스는 다음과 같습니다.
```
你: 实现 issue #43
↓
/tdd
↓
Claude 读 issue acceptance criteria
↓
Claude 探索代码库 → 找到 CONTEXT.md → 用项目术语
↓
Planning 阶段:
- 列出准备改的接口
- 列出准备测的行为(按优先级排序)
- 让你点头
↓
Tracer Bullet:
- RED: 写第一个测试(基于 acceptance criteria 第 1 条)
- 跑测试,确认 fail
- GREEN: 写最小实现
- 跑测试,确认 pass
↓
Incremental Loop:
- 每个 acceptance criteria 一个 RED→GREEN
↓
Refactor:
- 看 deep module 提取机会
- 每次重构后跑全套测试
↓
PR
```
빨간색과 녹색 주기마다 AI가 중지되고 "테스트 실패"/"테스트 통과, 차이점은 다음과 같습니다"라는 상태를 제공합니다. **이러한 일시 중지는 헤드라이트를 앞지르기 위한 해독제입니다** - AI는 한 번에 천 줄을 배치할 가능성이 없습니다.
## Mock 정보: Matt의 강력한 의견
SKILL.md는 모의의 위험성을 구체적으로 언급합니다. 그는 또한 별도의 `mocking.md`을 제공합니다. 핵심 아이디어:
> "Bad tests... mock internal collaborators."
모의 내부 협력자 = 테스트와 구현 간의 1:1 결합 = 리팩토링 시 테스트 팀이 무릎을 꿇습니다. Matt가 선호하는 것은 **통합 스타일 테스트**입니다. 실제 데이터베이스(인 메모리 또는 테스트 컨테이너), 실제 HTTP(MSW) 및 실제 파일 시스템(tmp dir)을 사용해 보십시오. 매우 비용이 많이 들거나 불안정한 경계(예: OpenAI API 호출)에서만 모의합니다.
이는 많은 팀의 현재 상황과 반대됩니다. 대부분의 코드 라이브러리는 단위 테스트로 가득 차 있으며 실제 코드보다 모의 코드가 더 많습니다. Matt는 연설에서 다음과 같이 판단했습니다. **좋은 코드 기반 = 테스트하기 쉬운 코드 기반**. 테스트하기 위해 여러 가지를 모의해야 한다면 코드 구조에 문제가 있다는 뜻이므로 먼저 아키텍처를 변경해야 합니다([`/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)으로 이동).
## AI 시대 TDD의 새로운 의미
23년 전 Kent Beck이 이 책을 썼을 때 TDD의 핵심 이점은 "사람들이 잘못된 코드를 작성하지 않는다"는 것이었습니다. AI 시대에 TDD는 추가적인 의미를 갖습니다.
**AI가 이해할 수 있는 유일한 "성공 기준"입니다**.
Matt는 나중에 그의 연설에서 Karpathy의 지혜로운 말을 인용했습니다.
"성공 기준"의 가장 좋은 형태는 **테스트**입니다. 이는 기계로 검증할 수 있고 바이너리이며 논쟁할 수 없습니다. AI에게 테스트 스위트 제공 + "통과"는 AI에게 요구 사항 설명 + "구현하세요"보다 10배 더 안정적입니다.
따라서 `/tdd`은 단순한 품질 보증 도구가 아니라 에이전트 루프에 대한 입력 인터페이스입니다. 각각의 빨간색과 녹색 주기는 완전한 "입력 → 작업 → 피드백"입니다. AI는 사이클에서 이 구현의 실제 상황을 학습하고 다음 사이클은 더 정확해질 것입니다.
## 설치 및 사용방법
```bash
npx skills@latest add mattpocock/skills
```
`tdd` + `setup-matt-pocock-skills`을 확인하세요.
지금 Codex를 주로 사용한다면 설치 제품을 `.agents/skills/`에 인계받고, 프로젝트 수준 워크플로, 테스트 명령, 이슈 트래커 규칙을 `AGENTS.md`에 작성하세요. Matt의 `/tdd`의 핵심은 Claude Code에 묶여 있지 않은 적록 리팩토링 주기입니다.
**통화 방법**:
* 직접: `/tdd` - 현재 대화 컨텍스트에서 무엇을 측정할지 추론하도록 합니다.
* 이슈 가져오기: `/tdd implement #43` - 이슈를 가져온 다음 엽니다.
* 버그 수정: `/tdd reproduce this bug then fix it` - 먼저 버그를 재현할 수 있는 실패한 테스트를 작성한 다음 수정합니다.
## 메모
**모든 작업에 적합하지는 않습니다**. 일회성 스크립트, 놀이터 탐색 코드, UI 미세 조정 - TDD를 사용하지 마십시오. 속도가 느려집니다. Matt 자신은 TDD가 "지속적인 가치가 있고 유지 관리가 필요한" 코드에 적합하다고 말했습니다.
**먼저 테스트 인프라를 준비하세요**. 프로젝트가 테스트 프레임워크(Vitest / Jest / Playwright 등)를 설치하지 않은 경우 먼저 설치한 다음 `/tdd`을 사용하세요. 그렇지 않으면 먼저 설치해 주지만 해당 단계에서 많은 질문이 있습니다.
**e2e 테스트를 자동으로 추가하지 마세요**. e2e는 느리고 선명하며 TDD 리듬은 분 수준입니다. `/tdd`은 기본적으로 e2e가 아닌 통합 테스트로 설정되어 있지만 "e2e가 아닌 유닛 + 통합만"이라고 명시적으로 알릴 수 있습니다.
**재구성 단계는 통제에서 벗어나기 가장 쉬운 단계입니다**. AI가 GREEN 상태를 얻은 후에는 여러 가지를 신나게 리팩토링합니다. 이를 응시하고 각 리팩토링 후에 테스트를 실행합니다. 이 부분은 AI 편차가 발생할 위험이 높은 영역입니다.
## 참조 리소스
다음 기사: [코드베이스 아키텍처 개선: 얕은 모듈을 깊은 모듈로 재구성](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture) - 정기적인 유지 관리를 통해 AI가 장기적으로 코드 베이스에서 실행될 수 있습니다.
# to-PRD + to-Issues: 그릴의 대화를 실행 가능한 수직 조각으로 압축합니다.
## 워크플로에서 이 섹션의 위치
Matt의 작업 흐름 다이어그램으로 돌아가서:
```
/grill-me 或 /grill-with-docs ← 谈清楚
↓
/to-prd ← 凝固成 PRD(你在这里)
↓
/to-issues ← 切成可领取的 vertical slice(你在这里)
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture
```
`/to-prd` 및 `/to-issues`은 이전과 다음 사이의 링크입니다: 추상적인 대화 결정**을 실행 가능한 작업 단위**로 변환합니다. Matt의 실제 사용에서는 두 단계가 연속되어 있기 때문에 두 단계를 결합했습니다.
## 실패모드 : 굽고 난 뒤 아무것도 할 수 없음
많은 사람들이 `/grill-me`을 사용한 후 막히게 됩니다. 많은 결정을 내리기 위해 긴 대화를 나누지만 **코드 작성을 시작하는 방법**은 무엇입니까?
Claude에게 직접적으로 "해 주세요"라고 말하는 것은 잘못된 것입니다. 이유는 다음과 같습니다.
1. 한 번에 완전한 기능 구현 = AI 출력 1000개 이상의 라인 = 검토 어려움, 테스트 어려움, 버그 찾기 어려움
2. AI가 메모리를 잃으면 다음 세션에는 컨텍스트가 없습니다.
3. 추적 불가 - 우리가 어디에 도달했고 얼마나 남았는지 알 수 없습니다.
올바른 접근 방식은 결정을 아티팩트(PRD)로 동결한 다음 PRD를 독립적으로 완료할 수 있을 만큼 작은 작업 패키지(문제)로 자르는 것입니다. 이는 30년 동안 소프트웨어 공학에서는 상식이었지만 AI 시대에는 새로운 의미를 갖습니다.
> AFK 에이전트(당신이 없을 때 실행되는 에이전트)가 독립적으로 픽업하여 완료할 수 있도록 잘게 자릅니다.
## /to-prd: 대화를 PRD로 압축합니다.
### 기술의 주요 제약
`/to-prd`의 SKILL.md는 시작 부분에 매우 중요한 문장을 씁니다.
> "This skill takes the current conversation context and codebase understanding and produces a PRD. **Do NOT interview the user — just synthesize what you already know.**"
더 이상 질문하지 마세요. 그릴미 스테이지는 이런데, to-prd는 **합성**만 합니다. 따라서 **컨텍스트를 지우지 말고 to-prd로 실행**하세요. 이는 이전 그릴의 모든 대화에 의존합니다.
### 스킬 처리 흐름
1. **코드 베이스 탐색**(아직 탐색하지 않은 경우) - 프로젝트의 CONTEXT.md 어휘를 사용하고 기존 ADR을 존중합니다.
2. **초안 모듈**——인터페이스를 독립적으로 테스트할 수 있도록 심층 모듈로 추출할 수 있는 기회를 사전에 찾습니다.
3. **사용자에 맞게 모듈 정렬** - "이 모듈이 맞습니까? 어떤 모듈을 테스트해야 합니까?"
4. 템플릿에 따라 **PRD를 생성**하여 이슈 트래커로 보내고 `needs-triage` 태그를 지정합니다.
### PRD 템플릿
Matt가 제공한 템플릿은 다음과 같습니다.
```markdown
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list:
1. As a , I want a , so that
2. ...
## Implementation Decisions
- The modules that will be built/modified
- The interfaces of those modules
- Technical clarifications from the developer
- Architectural decisions
- Schema changes / API contracts / Specific interactions
(NO specific file paths or code snippets — they rot fast.)
## Testing Decisions
- What makes a good test (test external behavior, not internals)
- Which modules will be tested
- Prior art (similar tests in the codebase)
## Out of Scope
What's NOT in this PRD.
## Further Notes
```
몇 가지 주요 디자인:
* **사용자 스토리가 대부분을 차지함**: 요구 사항이 길고 번호가 매겨진 목록 - 전체 기능 포인트를 철저하게 열거해야 합니다. 이로써 '분명한 줄 알았는데'라는 사각지대를 피할 수 있게 됐다.
* **구현 시 파일 경로나 코드를 작성하지 않습니다**: Matt는 "매우 빨리 구식이 될 수 있습니다"라고 직접 말했습니다. 이는 LLM 시대의 고유한 고려 사항입니다. 특정 경로는 재구성 직후 쓸모 없게 되지만 "모듈 경계" 및 "인터페이스 계약"은 수명 주기가 더 깁니다.
* **범위를 벗어나야 함**: 이 단락은 대부분의 PRD 템플릿에서 무시되지만 나중에 문제를 잘라낼 때 경계 보험이 됩니다.
## /to-issues: PRD를 수직 조각으로 자릅니다.
### 수직 슬라이스란 무엇입니까?
이는 Matt의 전체 방법론에서 가장 중요한 개념 중 하나입니다. SKILL.md는 직접적으로 다음과 같이 말했습니다.
> Each issue is a thin vertical slice cutting through ALL integration layers end-to-end, NOT a horizontal slice of one layer.
가장 명확한 예는 "주석 기능"을 만드는 것입니다.
**수평 슬라이싱(잘못된 방향)**:
* 문제 1: 데이터베이스 스키마
* 문제 2: API 엔드포인트
* 이슈 3: UI 컴포넌트
* 문제 4: 테스트
**세로 절단(Tracer Bullet 방법)**:
* 이슈 1: "방문자가 익명으로 댓글을 제출할 수 있습니다" (스키마 + API + UI + 테스트가 모두 포함되지만 범위가 작아서 익명만 가능함)
* 문제 2: "로그인한 사용자의 댓글이 해당 계정과 연결되어 있습니다."
* 이슈 3: "댓글에 답글을 달 수 있습니다"
* 문제 4: "관리자는 댓글을 삭제할 수 있습니다."
수평 슬라이싱의 문제점: 각 슬라이스를 개별적으로 테스트할 수 없습니다. Issue 1이 완성된 후에는 시연할 내용이 없었고, Issue 4가 되어서야 전체 링크를 실행할 수 있었습니다. 그때서야 스키마 설계가 잘못되었음을 발견했습니다.
완성된 각 수직 슬라이스는 **종단 간 사용 가능한 기능적 하위 집합**입니다. Pragmatic Programmer의 표현으로는 **추적 총알**(추적 총알)이라고 합니다. 먼저 한 발을 쏘아 십자선을 확인한 후 다음 발을 조정하세요.
### HITL vs AFK
`/to-issues`은 또한 각 조각에 레이블을 지정합니다.
* **HITL**(Human in the Loop) - 사람들이 의사 결정에 참여해야 합니다. 건축 결정, 디자인 검토 등
* **AFK** (Away From Keyboard)——에이전트가 독립적으로 작업을 완료할 수 있으며, 다시 돌아와서 결과를 확인할 수 있습니다.
> "가능한 경우 HITL보다 AFK를 선호합니다."
이는 Matt의 작업 흐름에서 매우 획기적인 아이디어입니다. 문제 편집을 마친 후 사용자가 부재 중일 때(예: 밤이나 주말) 실행되는 에이전트에 직접 문제를 보냅니다. 다음 날 다시 돌아오면 PR이 이미 거기 누워서 검토를 기다리고 있습니다. HITL 부분은 낮 동안 머물면서 에이전트와 협력합니다.
### 슬라이싱 확인 링크
`/to-issues`은 문제가 발생하는 즉시 문제를 생성하지 않습니다. 먼저 타일링 구성표를 번호가 매겨진 목록으로 표시합니다.
```
1. Title: 访客提交匿名评论
Type: AFK
Blocked by: None
User stories covered: #1, #2
2. Title: 评论关联到登录账户
Type: AFK
Blocked by: #1
User stories covered: #3
3. Title: 评论审核流程
Type: HITL(需要确认审核 UI 设计)
Blocked by: #1
User stories covered: #4, #5
```
그런 다음 다음과 같이 질문하십시오.
* 세분성이 정확한가? 너무 두껍다 / 너무 얇다?
* 종속성이 올바른가요?
* 어떤 것을 병합/분할해야 하나요?
* HITL/AFK 표시가 정확합니까?
실제로 문제 추적기에 보내기 전에 종속성 순서(차단기 우선)로 고개를 끄덕일 때까지 반복하여 이후 문제가 첫 번째 문제의 실제 문제 ID를 참조할 수 있도록 합니다.
### 이슈 템플릿
```markdown
## Parent
A reference to the parent issue (if any).
## What to build
A concise description. Describe end-to-end behavior, NOT layer-by-layer
implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
## Blocked by
- A reference to the blocking ticket
(or "None - can start immediately")
```
수직 슬라이스의 정신과 일치하는 "종단 간 동작 설명" 줄을 참고하세요. 승인 기준은 Claude가 `/tdd` 단계에서 하나씩 테스트로 변환하는 승인 목록입니다.
## 사용방법: 전체 프로세스 예시
블로그에 댓글 기능을 추가하고 싶다고 가정해 보겠습니다. 전체 프로세스:
```
你: 我想给博客加评论功能
↓
/grill-me → Claude 问 30 个问题(要不要登录?匿名?嵌套?审核?……)
↓
你回答完毕,达成共识
↓
/to-prd → Claude 生成结构化 PRD,提交到 GitHub Issues #42
↓
/to-issues → Claude 提议切成 4 个 vertical slice
让你确认粒度和依赖
你点头
按依赖顺序发布到 GitHub Issues #43~#46
↓
你回家睡觉
↓
夜里 AFK agent 抓 #43(无依赖),跑 /tdd 完成 → 提 PR
你早上 review、merge
↓
agent 抓 #44 / #45 ……
```
전체 과정에서 화면 앞에 앉아 모든 세부 사항을 볼 필요가 없습니다. 핵심 결정은 그릴미(grill-me) 단계에서 내려집니다.
## 설치 및 전제조건
```bash
npx skills@latest add mattpocock/skills
```
`to-prd`, `to-issues`, `setup-matt-pocock-skills`을 확인하세요.
\*\*먼저 `/setup-matt-pocock-skills`\*\*을 실행해야 합니다. 문제 추적기(GitHub / GitLab / local markdown)와 분류 레이블 어휘를 AGENTS.md/CLAUDE.md에 작성합니다. 그렇지 않으면 to-prd 및 to-issue는 문제를 보낼 위치를 알 수 없습니다.
지원되는 문제 추적기:
* **GitHub 문제**(기본값, `gh` CLI 사용)
* **GitLab 문제**(`glab` CLI 사용)
* **로컬 마크다운**(`.scratch//` 아래에 파일 생성) - 개인 프로젝트 또는 원격이 없는 프로젝트에 적합
* **기타**(Jira, Linear 등) - 산문을 사용하여 워크플로를 설명하면 설명에 따라 스킬이 호출됩니다.
## FAQ
\*\*Q: 이미 완성된 PRD가 있습니다. to-prd를 건너뛰고 to-issue로 바로 이동할 수 있나요? \*\*
답: 그렇습니다. `/to-issues`은 이슈 참조를 매개 변수("문제 #42를 수직 조각으로 나누기")로 받아들이고 이슈 콘텐츠를 가져온 다음 조각화합니다.
\*\*Q: 내 프로젝트에서는 이슈 트래커를 사용하지 않습니다. 사용할 수 있나요? \*\*
답: 그렇습니다. 설정 중에 "로컬 마크다운"을 선택하면 모든 이슈가 `.scratch//001-foo.md`과 같은 로컬 파일이 됩니다.
\*\*Q: PRD가 너무 길어서 AI 자체가 처리할 수 없는 경우 어떻게 해야 하나요? \*\*
A: 이는 슬라이스 세분화의 표시입니다. PRD는 하나의 거대한 PRD가 아닌 여러 개의 독립적인 PRD로 슬라이스되어야 합니다. 그릴 단계에서 느껴야 합니다. 50번째 이슈에 대해 이야기하고 여전히 새로운 기능을 도입하는 중이라면 먼저 멈추고 PRD를 두 개로 나누어 일괄적으로 수행하세요.
\*\*Q: AFK 에이전트는 어떻게 자동으로 문제를 포착합니까? \*\*
A: Matt의 저장소는 이 부분을 제공하지 않으며 자체 에이전트 배열과 조정되어야 합니다(예를 들어 GitHub Actions는 문제를 실행하기 위해 Claude Code를 트리거합니다). 가장 쉬운 방법은 cron을 사용하여 매시간 `is:open no:assignee label:agent-ready`을 확인하는 것입니다.
## 이 과정의 진정한 가치
`/to-prd` 및 `/to-issues`은 "자동화된 프로젝트 관리"처럼 보이지만 Matt는 더 깊은 이유 때문에 이를 자신의 작업 흐름의 중심에 두었습니다.
\*\*"**완전히 작은 것이 무엇인가**"\*\*라고 생각하게 만듭니다. 기능을 수직 조각으로 잘라야 한다면 소프트웨어 엔지니어링에서 가장 어려운 일 중 하나인 이음새를 찾는 것입니다. 이 두 가지 기술보다 생각 자체가 더 가치가 있습니다.
그리고 수직 슬라이스의 크기는 AI가 한번에 처리할 수 있는 크기입니다. **작업 크기를 AI 능력의 상한과 일치시키세요** - 이것이 LLM과의 협업의 기본 리듬입니다.
## 참조 리소스
다음 기사: [TDD: 빨간색과 녹색 재구성을 사용하여 AI가 작은 단계를 거치도록 강제](/ko/docs/notes/matt-pocock-skills/tdd)——문제가 해결된 후 AI가 실제로 작은 단계를 실현하도록 만드는 방법.
# Zoom Out: 길을 잃었을 때 AI에게 먼저 지도를 그리게 하기
## 가장 짧지만 유용함
`/zoom-out`은 Matt의 안정적인 엔지니어링 스킬 중 가장 짧을 수 있습니다. 핵심 지시는 한 문장으로 요약될 수 있습니다.
> 이 코드에 익숙하지 않습니다. 추상화 수준을 한 단계 높여 관련 모듈과 호출자 지도를 프로젝트 도메인 언어로 그려주세요.
이것은 코드를 작성하거나 리팩토링하기 위한 것이 아닙니다. 이것은 매우 흔한 상태를 처리하는 데 사용됩니다: **당신과 AI 모두 특정 파일에 집중하고 있지만, 왜 이 파일이 존재하는지 잊기 시작할 때**입니다.
## 해결하는 실패 패턴
AI 코딩은 쉽게 지역 최적화에 빠질 수 있습니다.
1. 사용자가 파일을 지정합니다.
2. AI가 이 파일을 읽습니다.
3. AI가 지역 코드에 따라 의도를 추측합니다.
4. 수정 후 상위 호출자, 도메인 규칙 또는 ADR이 이 수정 사항을 지원하지 않는다는 것을 발견합니다.
인간도 마찬가지입니다. 디버깅을 오래 하면 특정 함수에 집중하게 되고 시스템 내에서의 위치를 잊어버립니다.
`/zoom-out`은 이러한 터널 시야를 차단하는 역할을 합니다. 에이전트가 구현을 일시 중지하고 먼저 다음 질문에 답하도록 합니다.
* 이 코드는 어떤 도메인 개념에 속합니까?
* 누가 그것을 호출합니까?
* 그것은 누구를 호출합니까?
* 그것 뒤에 있는 불변량은 무엇입니까?
* `CONTEXT.md`의 용어와 어떻게 대응됩니까?
* 특정 ADR에 의해 제약을 받습니까?
## `/improve-codebase-architecture`와의 차이점
`/zoom-out`과 [`/improve-codebase-architecture`](/ko/docs/notes/matt-pocock-skills/improve-codebase-architecture)는 모두 시스템 전체를 보지만 목표는 완전히 다릅니다.
| 스킬 | 목표 | 출력 |
| -------------------------------- | ---------------- | ------------------------- |
| `/zoom-out` | 익숙하지 않은 코드 이해 지원 | 지도, 호출 관계, 도메인 설명 |
| `/improve-codebase-architecture` | 개선할 아키텍처 기회 찾기 | 후보 리팩토링, 삭제 테스트, 인터페이스 설계 |
`/zoom-out`은 "이것에 대해 설명해 주세요"와 더 유사합니다. 리팩토링 제안을 서두르거나 코드를 직접 수정해서는 안 됩니다. 그 임무는 인지 부하를 줄이는 것입니다.
## 도메인 용어를 강조하는 이유
이 스킬은 프로젝트의 도메인 용어집을 사용하도록 명시적으로 요구합니다. 그 이유는 다음과 같습니다. 파일 이름만 사용하여 설명하면 AI가 다음과 같은 내용을 쉽게 출력할 수 있습니다.
```text
OrderService는 OrderRepository를 호출하고, OrderRepository는 db 클라이언트를 호출합니다.
```
이것은 설명처럼 들리지만 실제로는 설명하지 않은 것입니다. 더 유용한 지도는 다음과 같아야 합니다.
```text
결제 흐름에서 Order Draft는 사용자가 아직 결제하기 전의 임시 주문입니다.
Order Finalization은 Draft를 변경 불가능한 Order로 변환하고 재고 예약을 트리거합니다.
`OrderService.finalize()`는 이 변환의 심(seam)이며, 주로 결제 콜백 및 관리자 재시도에서 호출됩니다.
```
두 번째 설명은 코드를 비즈니스 언어로 되돌립니다. "누가 누구를 호출하는지"만 아는 것이 아니라 "왜 존재하는지"도 알게 됩니다.
## 언제 사용하기에 적합한가
다음과 같은 시나리오에서 `/zoom-out`을 적극적으로 호출하는 것이 좋습니다.
* 익숙하지 않은 모듈을 맡기 전
* 버그를 수정하지만 관련 호출 체인을 아직 확신하지 못할 때
* AI가 생성한 코드를 검토 중이며 올바른 수준을 수정했는지 알 수 없을 때
* PRD 작성을 준비 중이며 모듈 경계를 확인하고 싶을 때
* 이미 3개의 파일을 보았지만 시스템 다이어그램을 형성하지 못했을 때
* `/improve-codebase-architecture`를 실행할 준비가 되었지만 아직 후보 영역을 확신하지 못할 때
이것은 "시작하기 전 5분"으로 특히 적합합니다. 어떤 버그는 코드가 어렵기 때문이 아니라 처음부터 잘못된 수준으로 보기 때문입니다.
## 재사용 가능한 출력 형식
원본 스킬은 매우 짧지만, 사용할 때 AI가 이 형식으로 출력하도록 하는 것이 좋습니다.
```markdown
## 시스템 내 이 코드의 위치
## 주요 도메인 용어
## 주요 모듈
| 모듈 | 책임 | 호출자 | 호출 대상 |
|---|---|---|---|
## 주요 흐름
## 알려진 제약 조건 / ADR
## 먼저 봐야 할 파일
```
이 형식은 일반적인 설명보다 더 안정적이며 후속 `/to-prd` 또는 `/diagnose`의 컨텍스트로 변환하기에 더 적합합니다.
## 계획 모드로 사용하지 마세요
`/zoom-out`의 위험은 AI가 지도를 설명한 후 "이렇게 수정할 수 있습니다"라고 제안하는 것입니다. 코드를 이해하기만 하려면 명확하게 제한해야 합니다.
```text
구현 계획은 제시하지 말고, 파일을 수정하지 말고, 구조만 설명하세요.
```
그 가치는 의사 결정과 이해를 분리하는 데 있습니다. 이해가 불분명할 때 제안하는 것은 종종 오해를 더 보기 좋게 포장하는 것일 뿐입니다.
## 제 사용 제안
`/zoom-out`은 다른 스킬과 조합하기에 좋습니다.
* `/zoom-out` → `/diagnose`: 먼저 시스템 지도를 보고 피드백 루프를 구축합니다.
* `/zoom-out` → `/grill-with-docs`: 먼저 기존 도메인 언어를 이해하고 새로운 요구 사항을 질문합니다.
* `/zoom-out` → `/to-prd`: 먼저 모듈 위치를 확인하고 PRD를 작성합니다.
* `/zoom-out` → `/improve-codebase-architecture`: 먼저 지도를 그리고 얕은/깊은 문제를 찾습니다.
이것은 완전한 프로세스가 아니라 브레이크일 뿐입니다. AI가 지역 파일에서 더 많은 것을 수정하기 시작할 때 먼저 zoom out하게 하면 종종 나중에 다시 작업하는 시간을 절약할 수 있습니다.
## 참고 자료
다음 글: [Prototype: 폐기 가능한 코드로 디자인 질문에 답하기](/ko/docs/notes/matt-pocock-skills/prototype).
# Pi Agent란 무엇인가
## 서론
기능만 본다면 Pi Agent는 쉽게 저평가될 수 있습니다. 터미널에서 실행되며, 파일을 읽고 수정하고, 명령을 실행하고, 세션을 저장하며, 모델을 전환할 수도 있습니다. 마치 또 다른 Claude Code나 Codex처럼 들립니다.
하지만 제가 생각하기에 Pi가 정말 흥미로운 점은 무엇을 더 했는지가 아니라 무엇을 덜 했는지에 있습니다. Pi는 AI 프로그래밍 도구의 가장 핵심적인 계층, 즉 모델, 컨텍스트, 도구, 세션, 확장을 유지하고 사용자 워크플로우를 미리 고정시키지 않으려 노력했습니다.
그래서 저는 Pi를 이렇게 이해하고 싶습니다.
**Pi Agent는 "더 완벽한" AI 프로그래밍 제품이 아니라, 더 얇고 투명한 코딩 에이전트 하네스입니다.**
이 판단은 기능 목록보다 더 중요합니다. 왜냐하면 Pi를 어떻게 학습해야 하는지를 결정하기 때문입니다. 명령어를 외우는 것보다 코딩 에이전트가 어떤 계층으로 구성되어 있는지 먼저 이해해야 합니다.
## Pi의 위치를 올바르게 파악하기
AI 프로그래밍 도구는 일반적으로 모델, 하네스, 엔지니어링 환경의 세 가지 계층으로 크게 나눌 수 있습니다. 하지만 이 세 계층만으로는 너무 추상적입니다. Pi에서 정말 주목할 만한 점은 중간 계층인 하네스가 어떤 모듈로 다시 분해되는지입니다.
이 다이어그램에는 세 가지 중요한 포인트가 있습니다.
첫째, Pi는 단순히 "채팅 UI"만 있는 것이 아닙니다. CLI, 대화형 TUI, print/JSON, RPC, SDK는 모두 진입점일 뿐이며, 실제로 작업을 처리하는 것은 `AgentSessionRuntime`과 `AgentSession`입니다.
둘째, Pi는 모델에 요청하기 전에 리소스 로딩을 수행합니다. `ResourceLoader`는 `AGENTS.md`, `CLAUDE.md`, 스킬, 확장, 프롬프트 템플릿 등을 정리한 다음 `SystemPrompt Builder`에 전달하여 모델이 실제로 보는 컨텍스트를 구성합니다.
셋째, 모델이 도구를 호출할 때, 모델이 파일 시스템을 직접 제어하는 것이 아닙니다. `AgentHarness`와 `AgentLoop`는 도구 검증, 도구 실행, 결과 수신, 다음 라운드 계속을 담당합니다. Extensions, Tool Registry, SessionManager는 옆에서 기능을 확장하고 상태를 저장합니다.
따라서 Pi는 중간에 서 있지만, 이 "중간"이라는 말이 공허한 이야기가 아닙니다. 구체적으로 제어하는 것은 다음과 같습니다. 어떤 컨텍스트가 모델에 들어가는지, 어떤 도구를 호출할 수 있는지, 도구 결과가 세션으로 어떻게 돌아오는지, 어떤 기능이 확장으로 보완되는지.
이것이 바로 Pi를 소개하는 많은 글에서 minimal, transparent, extensible을 강조하는 이유입니다. 이들은 사실 같은 것을 말하고 있습니다. Pi는 에이전트의 핵심 실행 계층을 작게 만들고, 사용자가 볼 수 있고 수정할 수 있도록 노력합니다.
## 한 번의 작업 흐름
Pi의 한 번의 요청은 "모델에게 한 문장을 묻고, 모델이 한 문장을 답하는" 것이 아닙니다. 더 정확히 말하면, 그것은 분기가 있는 시간 순서입니다.
여기서 가장 중요한 것은 네 번째 단계와 다섯 번째 단계 사이의 왕복입니다. 모델은 파일 시스템을 직접 건드리지 않고, `tool_call`만 제안합니다. Pi는 이 호출을 받아 도구 이름과 매개변수를 검증하고, 존재할 수 있는 확장 훅을 트리거하며, 실제 작업을 실행한 다음 `tool_result`를 컨텍스트에 다시 넣습니다. 모델은 새로운 컨텍스트에 따라 다음 단계를 판단합니다.
이것이 코딩 에이전트와 일반 챗봇의 차이입니다. 챗봇은 주로 텍스트 내에서 작업을 완료합니다. 코딩 에이전트는 엔지니어링 시스템에 진입해야 하므로 도구, 컨텍스트 및 상태를 관리하는 하네스가 반드시 필요합니다.
Pi의 기본 도구는 매우 적습니다.
| 도구 | 의미 |
| ------- | ------------- |
| `read` | 파일 읽기 |
| `edit` | 기존 파일 수정 |
| `write` | 파일 생성 또는 덮어쓰기 |
| `bash` | 셸 명령 실행 |
`grep`, `find`, `ls`와 같은 읽기 전용 도구도 활성화하거나 제한할 수 있습니다. 이 도구 세트는 절제되어 보이지만, 이미 프로그래밍 폐쇄 루프를 형성했습니다. 코드 읽기, 코드 수정, 테스트 실행, 오류에 따라 계속 수정.
이러한 설계 뒤에 있는 질문은 "Pi가 더 많은 것을 할 것인가"가 아니라 "더 많은 것이 기본적으로 핵심에 포함되어야 하는가"입니다. Pi의 답변은 명확합니다. 반드시 그렇지는 않습니다.
## 왜 많은 기능을 서둘러 내장하지 않는가
많은 AI 프로그래밍 제품은 계획 모드, 할 일, 서브 에이전트, MCP, 권한 팝업, 백그라운드 작업, 브라우저 도구 등을 제품에 내장합니다. 이렇게 하면 빠르게 시작할 수 있지만, 대가가 따릅니다. 모델이 실제로 어떤 컨텍스트를 받았는지 알기 어렵고, 제품 워크플로우를 자신의 워크플로우로 변경하기 어렵습니다.
Pi의 경로는 그 반대입니다. 핵심을 매우 작게 유지하고 워크플로우를 외부에 배치합니다.
| 변경하고 싶은 것 | Pi가 어디에 맡기는가 |
| -------------- | ------------------------- |
| 프로젝트 규칙 | `AGENTS.md` / `CLAUDE.md` |
| 전문 작업 방법 | Skills |
| 사용자 정의 도구 및 UI | Extensions |
| 공유 가능한 기능 세트 | Pi Packages |
| 모델 선택 | Provider / Model 설정 |
이것은 "기능 부족"이 아니라 제품의 절충점입니다. 핵심은 에이전트 루프만 관리하고, 구체적인 워크플로우는 사용자와 팀이 직접 조합하도록 맡깁니다.
예를 들어, Pi는 기본적으로 DeepSearch를 내장하지 않습니다. 하지만 이것이 심층 검색을 할 수 없다는 의미는 아닙니다. Pi의 사고방식에 더 부합하는 방법은 확장 기능을 작성하는 것입니다. `deep_search` 도구를 등록하고, Tavily, Exa, Brave Search 또는 회사 내부 검색을 연결한 다음, 모델이 필요할 때 이를 호출하도록 하는 것입니다.
이것은 "검색 버튼"을 제품에 하드코딩하는 것과는 다릅니다. 전자는 에이전트의 기능을 확장하는 것이고, 후자는 제품이 워크플로우를 대신 결정하는 것입니다.
## 중요: 자신의 워크플로우를 확장하는 방법
Pi를 단순히 "경험하는 것"이 아니라 실제로 사용하려면 자신의 워크플로우를 확장하는 것이 중요합니다.
여기서 혼동하기 쉬운 점은 Pi가 "플러그인"이라는 한 가지 확장 방식만 있는 것이 아니라는 것입니다. Pi는 프로젝트 규칙, 작업 방법, 실제 도구, 공유 가능한 패키지의 네 가지 계층 진입점을 제공합니다. 먼저 자신이 어떤 종류의 것을 정착시키고 싶은지 판단해야 합니다.
| 정착시키려는 것 | 무엇을 사용하는가 | 어떤 시나리오에 적합한가 |
| ------------- | ------------------------- | -------------------------------------------------------------- |
| 프로젝트 습관 및 제약 | `AGENTS.md` / `CLAUDE.md` | 에이전트에게 코드를 어떻게 수정하고, 어떤 검사를 실행하며, 어떤 디렉토리를 건드리지 말아야 하는지 알려줍니다. |
| 재사용 가능한 방법 세트 | Skill | 코드 검토, 글쓰기, 릴리스, 문서 생성, 이미지 처리와 같은 "단계와 경험" |
| 실제 기능 | Extension | 도구 등록, 도구 호출 가로채기, 슬래시 명령 추가, UI 추가, 외부 API 연결 |
| 배포 가능한 기능 세트 | Pi Package | 확장, 스킬, 프롬프트 템플릿, 테마를 자신 또는 팀이 재사용할 수 있도록 패키징 |
제 이해는 다음과 같습니다. **Skill은 작업 매뉴얼이고, Extension은 실행 가능한 플러그인이며, Package는 배포 컨테이너입니다.**
예를 들어, 현재 이 블로그 워크플로우는 다음과 같이 분해할 수 있습니다.
| 워크플로우 요구 사항 | 어디에 넣을 것인가 |
| -------------------------------------------------------------------------------------- | ----------------------- |
| "콘텐츠는 한국어로만 작성하고, 이미지는 `BlogImage`를 사용해야 하며, MDX를 수정한 후 `pnpm types:check`를 실행해야 합니다." | `AGENTS.md` |
| "개념 글을 작성할 때는 오해, 정의, 메커니즘, 예시, 경계 순으로 구성해야 합니다." | `article-writing` Skill |
| "Pi에 Tavily / Exa / Brave Search를 검색할 수 있는 `deep_search` 도구를 추가합니다." | Extension |
| "글쓰기 Skill, DeepSearch Extension, 위챗 게시 명령을 여러 프로젝트에서 사용할 수 있도록 패키징합니다." | Pi Package |
이것은 단순히 "플러그인 설치"라고 말하는 것보다 더 정확합니다. 왜냐하면 많은 워크플로우는 코드를 작성할 필요 없이 좋은 규칙이나 Skill만 있으면 되기 때문입니다. 하지만 에이전트가 외부 검색, 데이터베이스 조회, CI 호출, 위험한 명령 가로채기 등 실제로 새로운 기능을 갖기를 원한다면 Extension을 작성해야 합니다.
### Extension: 진정한 플러그인 계층
Pi의 Extension은 TypeScript 모듈입니다. 몇 가지 유형의 작업을 수행할 수 있습니다.
| 기능 | 예시 |
| --------- | ------------------------------------------------ |
| 도구 등록 | `deep_search`, `query_logs`, `open_issue` |
| 명령 등록 | `/review`, `/publish`, `/checkpoint` |
| 이벤트 가로채기 | `bash`에서 `rm -rf`, `sudo`, `.env` 파일 작성 전에 확인 요청 |
| UI 변경 | TUI에 상태, 선택 상자, 확인 상자, 작업 패널 표시 |
| 상태 저장 | 할 일, 연결 풀, 마지막 검색 결과, 작업 단계 기록 |
| 외부 시스템 연결 | CI, GitHub, 로그 시스템, 회사 내부 API |
Extension은 전역적으로 또는 프로젝트 내에 배치할 수 있습니다.
```text
~/.pi/agent/extensions/ # 전역 확장, 모든 프로젝트에서 사용 가능
.pi/extensions/ # 프로젝트 확장, 현재 프로젝트에서만 사용
```
임시 확장을 테스트하려면 다음을 사용할 수 있습니다.
```bash
pi -e ./my-extension.ts
```
자동 검색 디렉토리에 배치한 후 Pi에서 다음을 사용할 수 있습니다.
```text
/reload
```
확장, 스킬, 프롬프트 및 컨텍스트 파일을 다시 로드합니다.
이것이 제가 Pi에서 가장 가치 있다고 생각하는 부분입니다. 단순히 "모델이 코드 작성을 돕도록 하는 것"이 아니라, 모델을 위해 제어 가능한 작업 환경을 설계하는 것입니다. Extension은 모델이 어떤 기능을 호출할 수 있는지 결정하고, hooks는 어떤 동작을 가로채야 하는지 결정하며, commands는 자신의 워크플로우가 어떻게 트리거되는지 결정합니다.
### Skill: 모든 것을 플러그인으로 작성하지 마세요
어떤 기능이 주로 "어떻게 하는지"에 관한 것이고, "실제 API를 호출하거나 프로그램을 실행하는 것"이 아니라면, Skill로 작성하는 것이 더 적합합니다.
Skill의 구조는 일반적으로 다음과 같습니다.
```text
my-skill/
SKILL.md
scripts/
templates/
references/
```
Pi는 시작할 때 전체 스킬을 컨텍스트에 모두 넣지 않습니다. 먼저 스킬의 이름과 설명을 로드합니다. 작업이 일치하면 모델이 전체 `SKILL.md`를 읽도록 합니다. 이를 점진적 공개라고 합니다. 장점은 복잡한 방법론을 저장할 수 있지만, 매번 컨텍스트를 오염시킬 필요가 없다는 것입니다.
예를 들어 "좋은 글쓰기", "공식 계정에 게시", "브라우저 QA 수행" 등은 Skill에 더 가깝습니다. 이들의 가치는 주로 단계, 판단 기준 및 참조 자료에 있으며, LLM이 호출할 수 있는 도구를 등록할 필요는 없습니다.
### Package: 자신의 워크플로우를 패키징하기
안정적인 기능 세트가 있다면 Pi Package로 만드는 것을 고려할 수 있습니다.
Package에는 다음이 포함될 수 있습니다.
| 내용 | 역할 |
| ---------------- | ---------------------- |
| extensions | 실행 가능한 플러그인, 도구, 명령, 훅 |
| skills | 작업 방법 및 작업 매뉴얼 |
| prompt templates | 자주 사용하는 프롬프트 템플릿 |
| themes | TUI 테마 |
설치 방법은 다음과 같습니다.
```bash
pi install npm:@scope/my-pi-package
pi install git:github.com/user/repo@v1
pi install ./relative/path/to/package
pi list
pi remove npm:@scope/my-pi-package
pi update --extensions
```
기본 설치는 개인 설정에 기록됩니다. 팀 프로젝트에서 공유하고 싶다면 프로젝트 수준 설정을 사용하여 `.pi/settings.json`에 패키지를 기록할 수 있습니다. 이렇게 하면 다른 사람이 프로젝트에 진입하여 Pi를 시작할 때 누락된 패키지를 자동으로 보완할 수 있습니다.
하지만 여기서도 매우 신중해야 합니다. Package, Extension, Skill은 모두 에이전트의 동작에 영향을 미칠 수 있습니다. 타사 패키지는 브라우저 플러그인과 같은 낮은 권한의 장식이 아니며, 코드를 실행할 수도 있고 모델이 명령을 실행하도록 지시할 수도 있습니다. 설치 전에 소스 코드를 확인해야 합니다.
따라서 저는 Pi 확장을 이 순서대로 학습할 것입니다.
1. 먼저 `AGENTS.md`를 사용하여 프로젝트 규칙을 명확히 작성합니다.
2. 다음으로 반복되는 방법을 Skill로 만듭니다.
3. 실제 도구 기능이 필요할 때 Extension을 작성합니다.
4. 여러 프로젝트에서 재사용할 때 마지막으로 Package로 만듭니다.
이렇게 학습하는 것이 더 안정적입니다. 처음부터 플러그인을 작성하는 것이 아니라, 먼저 워크플로우를 "규칙, 방법, 도구, 배포"의 네 가지 유형으로 분해한 다음, 각 유형을 Pi의 어떤 계층에 배치할지 결정합니다.
## 소스 코드에서 무엇을 볼 수 있는가
Pi 소스 코드를 볼 때, 각 함수를 따라가는 것보다 몇몇 파일이 각각 어떤 설계 계층을 나타내는지 보는 것이 가장 도움이 되었습니다.
| 소스 코드 위치 | 설명 |
| ---------------------------------------------------- | -------------------------------------------------------- |
| `packages/agent/src/agent-loop.ts` | 핵심 루프: 사용자 메시지, 모델 응답, 도구 호출 및 도구 결과를 연결합니다. |
| `packages/agent/src/harness/agent-harness.ts` | 하네스 상태: 세션, 시스템 프롬프트, 도구, 훅 및 메시지 큐를 관리합니다. |
| `packages/coding-agent/src/core/tools/index.ts` | 내장 도구 집합: 기본 코딩 도구는 `read`, `bash`, `edit`, `write`입니다. |
| `packages/coding-agent/src/core/resource-loader.ts` | 리소스 로딩: 프로젝트 지침, 확장, 스킬, 프롬프트 템플릿 및 테마를 읽습니다. |
| `packages/coding-agent/src/core/system-prompt.ts` | 시스템 프롬프트 빌드: 도구 설명, 프로젝트 컨텍스트, 스킬 및 현재 디렉토리를 프롬프트에 넣습니다. |
| `packages/coding-agent/src/core/extensions/types.ts` | 확장 시스템: 확장이 도구, 명령, 단축키, UI 및 생명 주기 이벤트를 등록할 수 있도록 합니다. |
이러한 구성 요소들이 합쳐져 Pi의 심장을 이룹니다. 먼저 컨텍스트와 도구를 조립한 다음 요청을 모델에 전달합니다. 모델이 도구를 호출해야 하면 Pi가 도구를 실행합니다. 결과가 돌아오면 루프가 계속됩니다.
따라서 Pi의 "극도로 간결함"은 빈말이 아닙니다. 소스 코드 구조 자체도 이 아이디어를 표현하고 있습니다. 에이전트 루프, 하네스, 코딩 도구, 리소스 로딩, 확장 시스템을 분리하여 각 계층이 비교적 명확합니다.
## Pi의 경계
Pi는 자유도가 높지만, 자유도가 안전을 의미하지는 않습니다.
Pi 패키지와 확장은 코드를 실행할 수 있습니다. 스킬도 모델이 스크립트를 실행하도록 지시할 수 있습니다. `bash`는 실제 시스템에 접근할 수 있습니다. 공식 문서와 보안 분석 모두 타사 패키지, 확장, 스킬은 스스로 검토해야 한다고 경고합니다.
저는 Pi의 경계를 세 가지로 이해합니다.
1. **샌드박스가 아닙니다**: 자연스러운 격리 환경으로 생각하지 마십시오. 위험한 프로젝트는 컨테이너, 임시 디렉토리 또는 깨끗한 작업 트리에 넣는 것이 가장 좋습니다.
2. **권한을 대신 판단하지 않습니다**: Pi의 핵심 철학은 수많은 팝업으로 위험을 관리하는 것이 아니라, 사용자가 도구, 컨텍스트 및 확장을 제어하도록 하는 것입니다.
3. **엔지니어링 경계를 이해하는 사람에게 적합합니다**: 에이전트가 언제 명령을 실행하도록 할지, 언제 읽기 전용 도구만 제공할지, 언제 git 체크포인트를 먼저 만들지 알아야 합니다.
이것이 Pi와 일부 더 제품화된 에이전트의 차이점입니다. 제품화된 도구는 더 많은 보안 및 상호 작용 세부 사항을 포장합니다. Pi는 더 직접적인 제어권을 제공하며, 동시에 더 많은 책임을 사용자에게 돌려줍니다.
## Pi를 어떻게 학습해야 하는가
Pi를 학습할 때 "어떤 명령어가 있는가"부터 시작하는 것은 권장하지 않습니다. 명령어는 금방 찾아볼 수 있으며, 정말 배워야 할 것은 다음 질문들입니다.
| 질문 | 왜 중요한가 |
| --------------------------- | ---------------------------------------- |
| Pi는 컨텍스트를 어떻게 구성하는가 | 모델이 실제로 무엇을 아는지 결정합니다. |
| Pi의 도구 범위가 왜 이렇게 작은가 | 에이전트의 동작이 관찰 가능한지 결정합니다. |
| Extension은 도구를 어떻게 등록하는가 | 자신의 워크플로우를 연결할 수 있는지 결정합니다. |
| Skill과 Extension의 차이점은 무엇인가 | 언제 설명을 작성하고 언제 코드를 작성할지 결정합니다. |
| Session은 어떻게 저장되고 분기되는가 | 한 번의 엔지니어링 탐색이 복구, 검토 및 계속될 수 있는지 결정합니다. |
이미 Claude Code나 Codex를 사용해 본 적이 있다면, Pi를 "분해해서 보는" 기회로 삼을 수 있습니다. 모델이 코드를 작성하도록 하는 것은 같지만, 왜 어떤 도구는 블랙박스 제품 같고, 어떤 도구는 개조 가능한 런타임 같은가?
이 질문은 "Pi가 특정 도구를 대체할 수 있는가"보다 더 가치 있는 질문입니다.
## 마지막으로
Pi Agent의 가장 가치 있는 점은 AI 프로그래밍 도구의 중간 계층을 노출했다는 것입니다.
Pi는 코딩 에이전트의 능력이 모델뿐만 아니라 하네스의 설계에서도 비롯된다는 것을 상기시켜 줍니다. 모델은 생각하는 것을 담당하고, 하네스는 모델이 실제 엔지니어링 환경에서 행동하도록 합니다. 컨텍스트가 어떻게 들어오고, 도구가 어떻게 나가고, 결과가 어떻게 돌아오고, 확장이 어떻게 삽입되는지, 이러한 세부 사항들이 에이전트가 신뢰할 수 있고 투명하며 제어 가능한지 여부를 공동으로 결정합니다.
따라서 저는 Pi를 단순히 "Claude Code의 대체품"으로 보지 않을 것입니다. Pi는 개발자가 연구하고 개조하기에 적합한 에이전트 런타임에 가깝습니다. Pi를 직접 사용하여 코드를 작성할 수도 있고, Pi를 통해 자신의 에이전트 워크플로우를 설계하는 방법을 배울 수도 있습니다.
다음 실전 글에서는 이 아이디어를 계속 이어갈 것입니다. 일반적인 튜토리얼을 작성하는 대신, Pi Extension을 사용하여 `deep_search` 도구를 만들고 외부 검색 기능을 에이전트 루프에 연결하는 방법을 살펴보겠습니다.
## 추가 자료
# Pi Agent 실천 가이드
## 빠른 요약
개념 편에서 저는 Pi Agent를 극도로 간소화된 **Agent Harness**로 이해했습니다. 이는 모델, 터미널, 파일 시스템, 셸, 세션 및 확장 시스템을 연결하지만, 무거운 워크플로우를 미리 설정하지는 않습니다.
따라서 실천 편에서는 일반적인 "Pi가 파일을 수정하도록 하는" 사례를 다루고 싶지 않습니다. 그 사례는 기본적인 루프를 설명할 수 있지만, Pi의 확장성을 충분히 보여주지는 못합니다.
Pi에 더 적합한 실전 사례는 Pi가 기본적으로 가지고 있지 않지만 많은 사람들이 실제로 필요로 하는 기능인 **DeepSearch**를 추가하는 것입니다.
여기서 DeepSearch는 단순한 웹 검색이 아니라 연구 기반 워크플로우입니다.
| 단계 | 수행할 작업 |
| ------ | -------------------------------------- |
| 문제 분해 | 모호한 질문을 검색 가능한 몇 가지 하위 질문으로 분해 |
| 다중 검색 | 공식 문서, 코드 저장소, 블로그, 토론 포럼 또는 논문을 각각 검색 |
| 출처 필터링 | 중복 제거, 낮은 품질 결과 제외, 1차 출처 우선 유지 |
| 증거 정리 | 핵심 사실, 링크, 시간, 버전 및 불확실성 추출 |
| 종합 답변 | 결론 제시, 근거 및 제한 사항 설명 |
제 판단은 다음과 같습니다. **DeepSearch는 Pi 본체에 작성되어서는 안 되며, 프롬프트만으로 해결해서도 안 됩니다. Pi Extension으로 만드는 것이 더 적합합니다.**
이유는 간단합니다. DeepSearch는 네트워크 요청, 타사 검색 API, 출처 필터링, 결과 잘라내기, 인용 형식 및 보안 경계를 포함합니다. 이 모든 것은 코딩 에이전트의 최소 핵심이 아닌 워크플로우 기능에 속합니다.
## 설계 목표
이 사례는 완벽한 연구 시스템이 아닌, 실행 가능한 최소 버전을 구현하는 것을 목표로 합니다.
목표는 다음과 같습니다.
```text
Pi에 deep_search 도구를 추가합니다.
이 도구는 다음을 받습니다.
- query: 사용자가 연구할 질문
- depth: 검색 깊이
- maxResults: 반환할 최대 후보 자료 수
이 도구는 다음을 출력합니다.
- 구조화된 검색 결과
- 각 결과의 제목, URL, 요약, 관련성
- 모델이 사용할 증거 힌트
Pi는 이 증거를 받은 후 현재 모델이 최종 결론을 생성합니다.
```
저는 의도적으로 "검색"과 "종합"을 분리할 것입니다.
| 부분 | 누가 담당하는가 | 이유 |
| --------------- | -------------------- | ---------------------------- |
| 검색 API 호출 | DeepSearch extension | 이는 확정적인 외부 기능입니다 |
| 결과 중복 제거 및 잘라내기 | DeepSearch extension | 컨텍스트가 노이즈로 가득 차는 것을 방지합니다 |
| 어떤 증거가 중요한지 판단 | Pi 현재 모델 | 추론 및 컨텍스트 이해가 필요합니다 |
| 최종 답변 작성 | Pi 현재 모델 | 사용자 질문 및 프로젝트 컨텍스트와 결합해야 합니다 |
이렇게 하는 것이 더 안정적입니다. Extension은 자체적으로 모델을 호출할 필요가 없으며, 중첩된 에이전트가 될 필요도 없습니다. 고품질 증거만 제공하여 Pi의 원래 모델이 계속 추론하도록 합니다.
## 준비 작업
Pi extension은 전역 디렉터리 또는 프로젝트 디렉터리에 배치할 수 있습니다. 여기서는 먼저 프로젝트 디렉터리에 배치하는 것을 권장합니다.
```text
.pi/extensions/deepsearch/
package.json
index.ts
```
프로젝트 로컬 extension의 장점은 경계가 명확하다는 것입니다. 이 DeepSearch 기능은 현재 프로젝트에서만 활성화되며, 모든 Pi 세션에 영향을 미치지 않습니다.
검색 서비스는 Tavily, Exa, Brave Search, SerpAPI 또는 자체 검색 백엔드를 선택할 수 있습니다. 첫 번째 버전에서는 서비스 제공업체에 얽매이지 말고, 먼저 `searchWeb()` 함수로 추상화하십시오.
예를 들어, 환경 변수에 API 키를 저장합니다.
```bash
export TAVILY_API_KEY=tvly-...
```
타사 검색 API를 연결하고 싶지 않다면, 먼저 로컬 모의 데이터를 사용하여 extension을 실행할 수 있습니다. 도구 등록, 매개변수 전달 및 결과 형식이 안정화된 후 실제 검색 서비스를 연결하십시오.
## Step 1: Extension 디렉터리 생성
먼저 디렉터리를 생성합니다.
```bash
mkdir -p .pi/extensions/deepsearch
```
extension에 종속성이 필요한 경우 `package.json`을 배치할 수 있습니다.
```json
{
"name": "pi-deepsearch-extension",
"private": true,
"dependencies": {
"typebox": "*",
"@earendil-works/pi-ai": "*",
"@earendil-works/pi-coding-agent": "*"
},
"pi": {
"extensions": ["./index.ts"]
}
}
```
그런 다음 종속성을 설치합니다.
```bash
cd .pi/extensions/deepsearch
npm install
```
Pi의 extension은 TypeScript 모듈이므로 수동으로 컴파일할 필요가 없습니다. 이 경험은 도구 실험을 빠르게 수행하는 데 매우 적합합니다.
## Step 2: deep\_search 도구 등록
핵심 파일은 `.pi/extensions/deepsearch/index.ts`입니다.
첫 번째 버전은 다음과 같이 작성할 수 있습니다.
```typescript
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { StringEnum } from "@earendil-works/pi-ai";
import { Type } from "typebox";
type SearchResult = {
title: string;
url: string;
snippet: string;
score?: number;
};
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "deep_search",
label: "DeepSearch",
description: "Search the web for source-backed evidence about a question.",
promptSnippet: "Research a question with web search and return source-backed evidence.",
promptGuidelines: [
"Use deep_search when the user asks for current facts, external sources, comparison, investigation, or source-backed research.",
"After deep_search returns results, synthesize an answer with citations and clearly separate facts, inference, and uncertainty.",
"Do not treat deep_search results as final truth; inspect source quality and mention gaps."
],
parameters: Type.Object({
query: Type.String({
description: "The research question or search query."
}),
depth: Type.Optional(StringEnum(["quick", "normal", "deep"] as const)),
maxResults: Type.Optional(Type.Number({
minimum: 3,
maximum: 10,
default: 6
}))
}),
async execute(_toolCallId, params, signal) {
const depth = params.depth ?? "normal";
const maxResults = params.maxResults ?? 6;
const results = await searchWeb(params.query, depth, maxResults, signal);
return {
content: [
{
type: "text",
text: formatResultsForModel(params.query, results)
}
],
details: {
query: params.query,
depth,
results
}
};
}
});
}
async function searchWeb(
query: string,
depth: "quick" | "normal" | "deep",
maxResults: number,
signal: AbortSignal
): Promise {
const apiKey = process.env.TAVILY_API_KEY;
if (!apiKey) {
throw new Error("Missing TAVILY_API_KEY. Set it before starting pi.");
}
const response = await fetch("https://api.tavily.com/search", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
api_key: apiKey,
query,
search_depth: depth === "quick" ? "basic" : "advanced",
max_results: maxResults,
include_answer: false,
include_raw_content: depth === "deep"
}),
signal
});
if (!response.ok) {
throw new Error(`Search failed: ${response.status} ${response.statusText}`);
}
const data = await response.json() as {
results?: Array<{
title?: string;
url?: string;
content?: string;
score?: number;
}>;
};
return dedupeByUrl((data.results ?? []).map((item) => ({
title: item.title ?? "Untitled",
url: item.url ?? "",
snippet: item.content ?? "",
score: item.score
}))).filter((item) => item.url);
}
function dedupeByUrl(results: SearchResult[]): SearchResult[] {
const seen = new Set();
const deduped: SearchResult[] = [];
for (const result of results) {
const key = normalizeUrl(result.url);
if (seen.has(key)) continue;
seen.add(key);
deduped.push(result);
}
return deduped;
}
function normalizeUrl(url: string): string {
try {
const parsed = new URL(url);
parsed.hash = "";
parsed.searchParams.delete("utm_source");
parsed.searchParams.delete("utm_medium");
parsed.searchParams.delete("utm_campaign");
return parsed.toString();
} catch {
return url;
}
}
function formatResultsForModel(query: string, results: SearchResult[]): string {
if (results.length === 0) {
return `DeepSearch found no results for: ${query}`;
}
const lines = results.map((result, index) => {
return [
`## Source ${index + 1}`,
`Title: ${result.title}`,
`URL: ${result.url}`,
result.score === undefined ? undefined : `Score: ${result.score}`,
`Snippet: ${result.snippet}`
].filter(Boolean).join("\n");
});
return [
`DeepSearch query: ${query}`,
"",
"Use these sources as evidence. Cite URLs when making factual claims.",
"Separate confirmed facts from inference and uncertainty.",
"",
...lines
].join("\n\n");
}
```
이 코드는 가장 중요한 작업만 수행합니다.
| 코드 위치 | 역할 |
| ------------------------- | --------------------------------- |
| `pi.registerTool()` | `deep_search`를 모델 호출에 노출 |
| `parameters` | 모델에 도구가 어떤 매개변수를 필요로 하는지 알림 |
| `promptGuidelines` | 모델에 언제 사용하고, 사용 후 어떻게 처리해야 하는지 알림 |
| `searchWeb()` | 실제 검색 서비스 호출 |
| `dedupeByUrl()` | 중복 URL 제거 |
| `formatResultsForModel()` | 검색 결과를 모델이 쉽게 참조할 수 있는 증거 블록으로 정리 |
첫 번째 버전에서는 너무 복잡하게 만들지 마십시오. DeepSearch의 진정한 어려움은 검색 요청을 작성하는 것이 아니라, 출처 품질, 컨텍스트 길이, 인용 형식 및 불확실성을 제어하는 것입니다.
## Step 3: /deepsearch 명령 추가
도구는 모델이 호출하도록 되어 있지만, 사용자에게도 직접적인 진입점이 필요합니다.
사용자 입력을 더 명확한 연구 작업으로 다시 작성하는 명령을 추가로 등록할 수 있습니다.
```typescript
export default function (pi: ExtensionAPI) {
pi.registerCommand("deepsearch", {
description: "Run a source-backed DeepSearch task",
handler: async (args, ctx) => {
const query = String(args ?? "").trim();
if (!query) {
ctx.ui.notify("Usage: /deepsearch ", "warning");
return;
}
pi.sendUserMessage(
[
"아래 질문에 대해 DeepSearch를 수행해 주세요.",
"",
`질문: ${query}`,
"",
"요구 사항:",
"1. 먼저 deep_search를 호출해야 하는지 판단합니다.",
"2. 질문이 복잡하면 2-4개의 하위 질문으로 분해하여 각각 검색합니다.",
"3. 최종 답변에는 반드시 출처 링크가 포함되어야 합니다.",
"4. 사실, 추론 및 아직 불확실한 부분을 구분합니다.",
"5. 검색 결과를 그대로 나열하지 말고, 종합적인 판단을 제시합니다."
].join("\n"),
{ deliverAs: "followUp" }
);
}
});
pi.registerTool({
// deep_search tool definition...
});
}
```
이제 사용자는 직접 다음과 같이 입력할 수 있습니다.
```text
/deepsearch Pi Coding Agent의 extension 메커니즘은 어떤 기능에 적합한가요?
```
`/deepsearch`는 직접 검색하지 않고, Pi에 더 완전한 작업 설명을 보냅니다. 모델은 설명에 따라 `deep_search`를 호출하고, 결과를 기반으로 종합적인 작업을 완료합니다.
저는 이러한 설계를 더 선호합니다. 에이전트의 판단 공간을 유지하기 때문입니다. 검색 도구는 증거의 진입점일 뿐, 최종 답변 생성기가 아닙니다.
## Step 4: 시작 및 검증
프로젝트 로컬 extension이 준비되면 프로젝트 루트 디렉터리에서 Pi를 직접 시작할 수 있습니다.
```bash
TAVILY_API_KEY=tvly-... pi
```
임시로 테스트하는 경우, extension을 명시적으로 지정할 수도 있습니다.
```bash
TAVILY_API_KEY=tvly-... pi -e ./.pi/extensions/deepsearch/index.ts
```
Pi에 들어간 후, 외부 사실이 필요한 질문을 먼저 합니다.
```text
/deepsearch Pi Coding Agent 최신 버전의 extension 시스템은 어떤 기능을 지원하나요?
```
허용 가능한 출력은 단순히 몇 가지 검색 결과가 아니라 다음을 포함해야 합니다.
| 체크포인트 | 적합한 성능 |
| -------- | ---------------------------------- |
| 도구 호출 여부 | `deep_search`가 호출된 것을 볼 수 있음 |
| 출처 명확성 | 각 핵심 사실 뒤에 URL이 있음 |
| 중복 제거 여부 | 동일한 페이지를 중복 인용하지 않음 |
| 판단 여부 | 자료만 나열하지 않고, 적용 가능한 시나리오를 요약할 수 있음 |
| 불확실성 여부 | 버전 변경, 타사 API, 커뮤니티 확장에 대한 경계를 유지 |
결과가 단순히 "검색 결과 목록"이라면 `promptGuidelines`가 충분히 강력하지 않다는 의미입니다. 가이드라인을 더 명확하게 변경할 수 있습니다.
```typescript
promptGuidelines: [
"Use deep_search to gather evidence, not to produce the final answer.",
"After deep_search, write a concise research brief with citations.",
"Prefer official documentation, source code, release notes, and primary sources.",
"Mention when sources disagree or when the evidence is incomplete."
]
```
## Step 5: DeepSearch를 연구 도구처럼 만들기
첫 번째 버전을 실행한 후, 세 가지 유형의 기능을 계속 추가할 수 있습니다.
### 하위 문제 분해
DeepSearch가 가장 실패하기 쉬운 부분은 큰 문제를 검색 API에 직접 던지는 것입니다.
예를 들어:
```text
Pi Agent가 Claude Code를 대체할 수 있나요?
```
이것은 좋은 검색 쿼리가 아닙니다. 적어도 다음으로 분해할 수 있습니다.
| 하위 문제 | 역할 |
| -------------------------- | -------- |
| Pi Agent의 핵심 설계는 무엇인가 | 포지셔닝 찾기 |
| Pi Agent는 어떤 도구와 확장을 지원하는가 | 기능 경계 찾기 |
| Claude Code의 기본 기능은 무엇인가 | 비교 대상 찾기 |
| 둘의 권한, 보안, 확장성 차이는 무엇인가 | 판단 형성 |
첫 번째 버전에서는 모델이 스스로 분해하도록 할 수 있습니다. 두 번째 버전에서는 `/deepsearch` 명령이 모델에게 먼저 하위 질문을 나열한 다음 `deep_search`를 하나씩 호출하도록 강제할 수 있습니다.
### 출처 품질 계층화
DeepSearch의 출력은 검색 API의 점수 순서로만 정렬되어서는 안 됩니다. 실제 기술 문서를 작성할 때는 다음을 우선적으로 고려합니다.
| 우선순위 | 출처 |
| ---- | -------------------------- |
| P0 | 공식 문서, 소스 코드, 릴리스 노트 |
| P1 | 저자 블로그, 유지 관리자 설명, 이슈 / PR |
| P2 | 고품질 튜토리얼, 기술 분석 |
| P3 | 커뮤니티 토론, Reddit, X, 포럼 |
Extension은 `formatResultsForModel()`에서 먼저 출처 유형을 표시할 수 있습니다.
```typescript
function classifySource(url: string): "official" | "source" | "community" | "other" {
const host = new URL(url).hostname;
if (host === "pi.dev") return "official";
if (host === "github.com") return "source";
if (host.includes("reddit.com")) return "community";
return "other";
}
```
이렇게 하면 모델이 종합할 때 커뮤니티 소문과 공식 문서를 동일한 증거 등급으로 취급하지 않습니다.
### 컨텍스트 잘라내기
검색 결과는 컨텍스트를 쉽게 오염시킬 수 있습니다. DeepSearch 도구 출력은 간결하고 정교해야 합니다.
제 제안은 다음과 같습니다.
| 내용 | 도구 출력에 포함 여부 |
| -------------- | -------------------------- |
| 제목 | 포함 |
| URL | 포함 |
| 200-500자 요약 | 포함 |
| 페이지 전체 텍스트 | 기본적으로 포함하지 않음 |
| 원본 HTML | 포함하지 않음 |
| 검색 API 원본 JSON | `details`에 포함, 본문에 포함하지 않음 |
전체 텍스트 읽기가 정말 필요한 경우, 두 번째 도구를 만들 수 있습니다.
```text
fetch_source(url)
```
이렇게 하면 DeepSearch는 첫 번째 단계에서 후보 출처를 찾고, 두 번째 단계에서는 가장 중요한 2-3개 페이지만 가져옵니다. 처음부터 수십 개의 웹 페이지 전체 텍스트를 모델에 모두 넣지 마십시오.
## 자주 묻는 질문
### 왜 bash로 검색 스크립트를 직접 실행하지 않나요?
가능하지만 extension만큼 안정적이지 않습니다.
bash를 사용하는 문제는 모델이 매번 명령, 매개변수, 출력 형식 및 오류 처리를 다시 결정해야 한다는 것입니다. Extension은 이러한 세부 사항을 고정하여 모델이 `deep_search`만 호출하면 되도록 합니다.
### 왜 요약도 extension에 작성하지 않나요?
첫 번째 버전에서는 권장하지 않습니다.
extension이 자체적으로 모델을 호출하여 요약하면 중첩된 모델 호출, 비용 계산, 컨텍스트 드리프트 및 인용 책임 문제가 발생합니다. 더 간단한 방법은 extension이 증거만 반환하고, Pi 현재 세션의 모델이 종합을 담당하는 것입니다.
### 이 DeepSearch는 MCP에 해당하나요?
아닙니다. 이는 Pi extension이 등록한 로컬 도구입니다.
이미 성숙한 MCP 검색 서버가 있다면 Pi의 MCP 관련 패키지 또는 extension을 통해 연결할 수도 있습니다. 하지만 이 사례는 Pi 자체의 확장 메커니즘을 이해하기 위해 직접 extension을 작성하는 것을 선택했습니다.
### 보안상 주의할 점은 무엇인가요?
적어도 네 가지 사항에 주의해야 합니다.
| 위험 | 조치 |
| ----------------- | --------------------------------------- |
| API 키 유출 | 환경 변수에서만 읽고, 저장소에 작성하지 않음 |
| 신뢰할 수 없는 웹 페이지 내용 | 웹 페이지 내용을 시스템 명령으로 간주하지 않고, 검증할 증거로만 간주 |
| 검색 결과 오염 | 공식 및 소스 코드를 우선하고, 커뮤니티 결과의 가중치를 낮춤 |
| 컨텍스트 폭발 | 결과 수 및 요약 길이 제한 |
DeepSearch는 "검색 강화"처럼 보이지만, 본질적으로 외부 웹 페이지를 에이전트 컨텍스트로 가져오는 것입니다. 외부 내용이 컨텍스트로 들어오면 프롬프트 인젝션을 실제 위험으로 간주해야 합니다.
## 요약
저는 Pi Agent의 첫 번째 실전 사례를 **DeepSearch Extension**으로 정할 것입니다. 이는 Pi의 세 가지 핵심 특징을 동시에 보여줄 수 있기 때문입니다.
* Pi의 핵심은 기본적으로 작으며, 모든 워크플로우를 내장하지 않습니다.
* 실제로 유용한 기능은 extension을 통해 추가할 수 있습니다.
* Extension은 단순히 명령을 추가하는 것이 아니라, 모델이 외부 세계로 진입하는 경계를 정의하는 것입니다.
이 사례가 성공적으로 실행되면 Pi는 단순히 로컬 코드 편집 에이전트가 아니라, 제어 가능한 연구 진입점을 갖게 됩니다. 외부 자료가 필요한 문제에 직면했을 때, 먼저 검색하고, 필터링하고, 출처를 기반으로 답변할 수 있습니다.
이는 모델이 기억에 의존하여 답변하는 것보다 더 신뢰할 수 있으며, 매번 검색 명령을 수동으로 작성하는 것보다 더 재사용 가능합니다.
## 참고 문서
# Ralph Wiggum 심층 분석
## 서론
퇴근 전에 AI에게 작업을 할당하고, 다음 날 아침에 사용 가능한 코드를 받는다 — 이 꿈은 복잡한 Agent 클러스터와 정교한 오케스트레이션 시스템이 필요할 것 같습니다. 하지만 2025년 가장 화제가 된 AI 프로그래밍 기술의 핵심은 바로 이 한 줄입니다:
```bash
while :; do cat PROMPT.md | claude ; done
```
무한 루프로, 반복적으로 작업을 Claude에게 전달합니다. 이것이 바로 **Ralph Wiggum**입니다. 민망할 정도로 단순하지만, 실제로 누군가가 이것을 사용하여 $297로 원래 견적이 $50,000이었던 프로젝트를 완성했습니다.
왜 이렇게 단순한 방법이 오히려 효과적일까요? Anthropic이 공식 플러그인을 출시한 후, 발명자 Geoffrey Huntley가 "This isn't it"이라고 말한 것은 또 무슨 일일까요?
## Ralph란 무엇인가
이름은 《심슨 가족》의 캐릭터에서 유래했습니다. Ralph Wiggum은 경찰서장의 아들로, 작품에서 가장 "순수한" 인물입니다 — 자신이 뭘 하고 있는지 잘 모르지만, 절대 멈추지 않습니다. 그의 상징적인 대사 "I'm helping!"은 의외로 이 기술의 정수를 드러냅니다: **순진하고 끈질긴 지속**(Naive and relentless persistence).
여기서 중요한 구분이 있습니다: **Ralph는 방법론이지, 도구가 아닙니다**. "애자일 개발"이 방법론이지 특정 소프트웨어가 아닌 것처럼, Ralph는 하나의 작업 방식을 설명합니다. 구현에 따라 효과가 크게 달라질 수 있으며, 이 문제는 뒤에서 자세히 다루겠습니다.
## Ralph가 필요한 이유: Context Rot 문제
Ralph가 왜 효과적인지 이해하려면, 먼저 이것이 해결하는 문제를 이해해야 합니다.
### AI는 어떻게 "멍청해지는가"
Claude로 복잡한 작업을 처리할 때, 이런 경험이 있으실 것입니다: 처음에는 대화가 원활하고, Claude가 정확하게 이해하며 실행도 잘합니다. 하지만 대화가 길어질수록 "둔해지기" 시작합니다 — 중요한 정보를 잊고, 같은 실수를 반복하며, 코드 품질이 떨어지고, 심지어 이해할 수 없는 "환각"을 만들어내기 시작합니다.
이것은 AI가 똑똑하지 못해서가 아닙니다. 문제는 **컨텍스트 윈도우가 오염되었기 때문**입니다.
이런 상황을 상상해 보십시오: Claude에게 기능 하나를 작성하라고 했는데, 첫 번째 시도가 실패했습니다. "이거 수정해"라고 했더니 다시 시도했지만 또 실패했습니다. 이렇게 열 번을 반복하면, Claude의 컨텍스트에는 아홉 번의 실패한 코드, 아홉 세트의 오류 메시지, 더 이상 관련 없는 대량의 논의가 쌓여 있습니다. 이런 잡다한 정보 속에서 핵심을 찾는 것은 점점 어려워집니다.
### Dumb Zone
Geoffrey Huntley와 커뮤니티 개발자들은 하나의 현상을 발견하고, 이를 "Dumb Zone"이라 명명했습니다:
| 컨텍스트 크기 | 성능 |
| ----------------- | ------------------ |
| 0 - 50k tokens | 최고 성능 |
| 50k - 100k tokens | 양호, 약간의 저하 |
| 100k+ tokens | 눈에 띄는 저하, 지시 무시 시작 |
| 150k+ tokens | 심각한 저하 |
정확한 임계점은 없지만, 경험 법칙으로: **컨텍스트가 절반 정도 차면 경계해야 합니다**. 200k tokens의 Claude의 경우, 100k를 넘으면 "멍청해진" AI와 대화하고 있을 수 있습니다.
### 누적된 컨텍스트는 부채입니다
여기에 직관에 반하는 통찰이 있습니다: 누적된 컨텍스트는 자산이 아니라 부채입니다.
우리는 기억력이 좋을수록, 보존하는 정보가 많을수록 좋다고 생각하는 데 익숙합니다. 하지만 대규모 언어 모델의 세계에서 이 직관은 틀렸습니다. 대화가 길어질수록 컨텍스트에 쌓이는 "부정적 정보"가 많아집니다: 실패한 코드, 더 이상 관련 없는 논의, 수정된 잘못된 이해. 이것들은 공간을 차지할 뿐만 아니라 AI의 "주의력"을 분산시킵니다.
## Ralph의 작동 원리
Context Rot을 이해했다면, Ralph의 해결책은 명확합니다: **누적 컨텍스트가 문제라면, 누적하지 않으면 됩니다**.
Ralph는 세 가지 기둥 위에 구축되어 있습니다:
### 1. 새 Session
매 루프 반복 시, **완전히 새로운 Claude 인스턴스**를 시작하여 완전히 깨끗한 컨텍스트 윈도우를 얻습니다. "대화 기록 지우기"가 아닙니다 — 그렇게 하면 누적된 상태가 여전히 남아있을 수 있습니다. 현재 프로세스를 완전히 종료하고 새로운 것을 시작하는 것입니다.
이는 매 반복이 시작될 때 Claude가 최적의 상태에 있다는 것을 의미합니다. 이전의 오류가 방해하지 않고, 오래된 논의가 주의를 분산시키지 않습니다.
**이것이 루프가 반드시 Claude Code 외부에서 실행되어야 하는 이유입니다** — bash 루프가 Claude 프로세스의 생명주기를 제어할 수 있어야 합니다.
### 2. 파일을 진실의 원천으로
매번 새로운 컨텍스트라면, AI는 이전에 무엇을 했는지 어떻게 알 수 있을까요? 답은: 대화 기록이 아닌 파일 시스템을 통해서입니다.
핵심 파일:
* **PRD/spec 파일** — 목표, 기능 목록, 성공 기준 정의
* **IMPLEMENTATION\_PLAN.md** — 작업 분해 및 진행 상황
* **progress.txt** — 자유 형식 로그, 매 반복 종료 시 학습한 내용 추가
* **Git 기록** — 코드 변경의 증거
매 반복이 시작될 때, Claude는 이 파일들을 읽어 목표와 진행 상황을 파악합니다. Claude가 보는 것은 체계적으로 정리된 상태 스냅샷이지, 혼란스러운 대화 기록이 아닙니다.
### 3. 피드백 루프
깨끗한 컨텍스트와 영속화된 상태만으로는 충분하지 않습니다. AI가 문제가 있는 코드를 작성하고 커밋하면 오류가 누적됩니다.
피드백 루프는 자동화된 품질 게이트입니다:
* **TypeScript 타입 검사** — 타입 정확성에 대한 즉각적 피드백
* **단위 테스트** — 기능이 예상대로 작동하는지 검증
* **CI/CD** — 코드 빌드 및 통합 가능 여부 확인
테스트가 실패하면 코드가 커밋되지 않고, Claude는 실패 정보를 확인합니다. 다음 반복의 새로운 Claude 인스턴스가 문제를 수정하려고 시도합니다.
> 완전한 품질 보증 체계를 구축하는 방법에 대해서는, [저의 Claude Code 품질 검사 프로세스](/ko/blog/claude-code-quality-control)에서 5단계 방어선의 실전 경험을 공유했습니다: Hooks 자동화, 테스트 전략, AI Review, Pre-commit, GitHub 통합.
## Human on the Loop
Geoffrey Huntley가 반복적으로 강조하는 개념적 구분이 있습니다:
| Human **in** the Loop | Human **on** the Loop |
| --------------------- | ------------------------- |
| 보모식 동행 | 감독식 관리 |
| AI가 매 단계마다 확인을 기다림 | 목표와 경계를 설정하면 AI가 자율적으로 실행 |
| 당신이 워크플로우의 병목 | 가끔 진행 상황을 점검 |
실제 사용에는 두 가지 모드가 있습니다:
* **AFK 모드**: 퇴근 전 시작하고, 집에 가서 잠자고, 아침에 결과 확인
* **Human-in-the-loop 모드**: 매 반복 후 일시 정지하여 확인, 복잡하거나 불확실한 작업에 적합
## 어떤 작업이 Ralph에 적합한가
Ralph는 만능이 아닙니다. 핵심 장점이 "성공할 때까지 반복"이므로, 특정 유형의 작업에 적합합니다.
### 적합한 작업
| 시나리오 | 이유 |
| ---------------- | ------------------------------------ |
| 명확한 성공 기준이 있는 작업 | 완료 여부를 자동으로 검증 가능 (테스트 통과, 타입 검사 통과) |
| 반복적 개선이 필요한 작업 | Ralph의 핵심 장점이 바로 끊임없는 시도 |
| 그린필드 프로젝트 | 기존 코드 파손 우려 없음 |
| 자동 테스트가 있는 프로젝트 | 테스트가 백프레셔 메커니즘으로 작용하여 품질 보장 |
### 적합하지 않은 작업
| 시나리오 | 이유 |
| ------------------ | -------------------------- |
| 인간의 판단이 필요한 디자인 결정 | "보기 좋은지"를 자동으로 검증할 수 없음 |
| 일회성 작업 | 반복이 필요 없는 작업에 Ralph 사용은 낭비 |
| 프로덕션 환경 디버깅 | 리스크가 너무 높아 무인 작업에 부적합 |
| 성공 기준이 불명확한 작업 | 언제 멈춰야 할지 판단 불가 |
### 세 가지 사용 모드
**완전 구현 모드**
이것은 Ralph의 가장 일반적인 사용법입니다: 완전한 기능이나 프로젝트를 처음부터 구축합니다. spec 파일과 구현 계획을 준비하고, Ralph가 모든 작업을 자동으로 실행하도록 합니다.
전형적인 시나리오:
* 새로운 REST API 구축
* CLI 도구 개발
* 새로운 기능 모듈 구현
실제 사례: 한 개발자가 이 모드로 $50,000 가치의 외주 프로젝트를 완성했으며, 총 API 비용은 $297에 불과했습니다. 전체 과정에는 MVP 개발, 테스트 작성, 코드 리뷰가 포함되었으며, 전 과정이 자동화되었습니다. 또 다른 사례로는 오래된 코드베이스를 React v16에서 v19로 업그레이드한 것으로, Ralph가 14시간 동안 실행되었고, 인간의 개입이 전혀 필요하지 않았습니다.
**탐색 모드**
모든 작업이 코드 산출물을 필요로 하는 것은 아닙니다. 때로는 이해가 필요합니다 — 새로 맡은 코드베이스의 이해, 복잡한 시스템의 아키텍처 이해, 특정 모듈의 작동 원리 이해.
전형적인 시나리오:
* 낯선 프로젝트를 인수받아 빠르게 전체적인 인식을 구축해야 할 때
* 기존 코드베이스에 대한 문서 생성
* 시스템 아키텍처 분석, 잠재적 문제 발견
이 모드에서는 프롬프트가 "X 기능을 구현하라"가 아니라 "이 코드베이스를 읽고 아키텍처 문서를 생성하라" 또는 "모든 API 엔드포인트를 찾아 그 역할을 설명하라"입니다. Claude가 매 반복마다 깊이 탐색하며, 점진적으로 더 완전한 이해를 구축합니다.
**무차별 테스트 모드**
어떤 버그는 증상도 알고, 기대하는 올바른 동작도 알지만, 근본 원인을 찾지 못합니다. 이때 Ralph에게 "무차별 대입"을 시킬 수 있습니다.
전형적인 시나리오:
* 간헐적으로 발생하는 버그, 재현이 어려움
* 특정 테스트가 가끔 실패하는데 원인 불명
* 성능 문제, 병목이 어디인지 불확실
목표를 설정합니다: "이 버그를 수정하고, 이 테스트가 안정적으로 통과하게 하라". Ralph가 다양한 수정 방안을 계속 시도하여, 효과적인 것을 찾을 때까지 반복합니다. 이 방법은 "어떻게 고쳐야 할지는 모르겠지만, 고쳐졌는지는 알 수 있다"는 유형의 문제에 특히 적합합니다.
## 구현 방식의 선택
Ralph의 방법론을 이해했다면, 다음으로 실질적인 질문에 직면합니다: 이 루프를 어떻게 구현할 것인가?
커뮤니티에서 엔지니어링 수준이 다른 두 가지 구현이 발전했습니다:
**미니멀 노선** — [snarktank/ralph](/ko/docs/notes/ralph-wiggum/snarktank): 수백 줄의 bash 스크립트, 매번 새로운 세션, 루프 자체에 집중합니다. 경량이고 시작하기 쉬우며, 빠르게 시작하기에 적합합니다.
**엔지니어링 노선** — [frankbria/ralph-claude-code](/ko/docs/notes/ralph-wiggum/frankbria): 완전한 도구 체인(모니터링 대시보드, 서킷 브레이커, 속도 제한, 세션 만료 관리). 기본적으로 `--continue`를 통해 세션을 재사용하며, `--no-continue`로 새 세션 모드로 전환할 수도 있습니다.
| 차원 | 미니멀 (snarktank) | 엔지니어링 (frankbria) |
| ------- | --------------- | ---------------------- |
| 세션 모드 | 매번 새로 생성 | 기본 재사용, 새로 생성으로 전환 가능 |
| 모니터링 | 수동 확인 | 내장 tmux 대시보드 |
| 안전 메커니즘 | max\_iterations | 서킷 브레이커 + 속도 제한 + 타임아웃 |
| 설치 복잡도 | Skill 복사 | install.sh + 마법사 |
두 구현 모두 장단점이 있으며, 선택은 엔지니어링 도구에 대한 여러분의 요구 사항에 따라 달라집니다. 자세한 사용법과 비교 분석은 각각의 실전 문서를 참조하십시오.
## 마치며
Ralph는 우리에게 중요한 교훈을 줍니다: 때로는 가장 단순한 방법이 가장 효과적입니다. 모든 사람이 더 복잡한 아키텍처를 추구할 때, 하나의 bash 루프가 게임의 규칙을 바꿨습니다.
물론 Ralph는 퍼즐의 일부에 불과합니다. 좋은 프롬프트, 적합한 프로젝트, 올바른 피드백 메커니즘이 있어야 위력을 발휘할 수 있습니다. 원리를 이해한 후, 여러분의 필요에 따라 적합한 구현 방식을 선택할 수 있습니다:
* 장기 AFK, 대량 반복이 필요하신가요? → [snarktank/ralph 실전 가이드](/ko/docs/notes/ralph-wiggum/snarktank)
* 엔지니어링 모니터링과 안전 메커니즘이 필요하신가요? → [frankbria/ralph-claude-code 실전 가이드](/ko/docs/notes/ralph-wiggum/frankbria)
***
**관련 읽을거리**:
* [snarktank/ralph 실전 가이드](/ko/docs/notes/ralph-wiggum/snarktank) — 미니멀 외부 루프, 설치부터 실전까지의 완전한 운영 매뉴얼
* [frankbria/ralph-claude-code 실전 가이드](/ko/docs/notes/ralph-wiggum/frankbria) — 엔지니어링 구현: 모니터링, 서킷 브레이커와 안전 메커니즘
* [Claude Subagent 완전 가이드](/ko/docs/notes/claude-subagent) — 컨텍스트를 깨끗하게 유지하는 또 다른 방법
* [Claude Skills란 무엇인가](/ko/docs/notes/claude-skills/concept) — Claude의 재사용 가능한 작업 매뉴얼 탐색
* [GSD 심층 분석](/ko/docs/notes/gsd/concept) — Ralph를 기반으로 구축된 완전한 컨텍스트 엔지니어링 시스템
* [Claude 시스템 아키텍처 완전 분석](/ko/docs/notes/claude-architecture) — Hooks, Subagent 등 구성 요소의 전체 아키텍처 이해
# frankbria/ralph-claude-code 실전 가이드
## 서론
[이전 글](/ko/docs/notes/ralph-wiggum/concept)에서는 Ralph의 방법론을 소개했고, [snarktank/ralph](/ko/docs/notes/ralph-wiggum/snarktank)에서는 극도로 간결한 외부 루프 구현을 살펴보았습니다. 이제 또 다른 방향을 살펴보겠습니다: [frankbria/ralph-claude-code](https://github.com/frankbria/ralph-claude-code).
snarktank/ralph의 철학이 "최소한의 코드로 최대한의 일을 하는 것"이라면, frankbria의 철학은 "**모든 것을 엔지니어링하는 것**"입니다 — 대화형 설정 마법사, 실시간 모니터링 대시보드, Circuit Breaker, 속도 제한, 세션 만료 관리. 간결함을 추구하지 않고 **제어 가능성**을 추구합니다.
두 구현 사이에 우열은 없으며, 서로 다른 사용 시나리오에 적합합니다. 이 글에서는 frankbria의 전체 도구 체인을 안내합니다.
## 설치 및 설정
### 전역 설치
```bash
# 리포지토리 클론
git clone https://github.com/frankbria/ralph-claude-code.git
cd ralph-claude-code
# 전역 설치
./install.sh
```
설치가 완료되면 다음과 같은 전역 명령어를 사용할 수 있습니다:
| 명령어 | 설명 |
| --------------- | --------------------- |
| `ralph` | Ralph 루프 시작 |
| `ralph-enable` | 기존 프로젝트에서 Ralph 활성화 |
| `ralph-setup` | 새 프로젝트를 생성하고 Ralph 설정 |
| `ralph-import` | 기존 PRD/요구사항 문서 가져오기 |
| `ralph-monitor` | 실시간 모니터링 대시보드 시작 |
### 프로젝트 초기화
기존 프로젝트에는 대화형 마법사를 사용합니다:
```bash
cd your-project
ralph-enable
```
마법사가 프로젝트 유형(Node.js, Python, Go 등)과 프레임워크(Next.js, FastAPI 등)를 자동으로 감지한 후 해당하는 설정 파일을 생성합니다.
완전히 새로운 프로젝트의 경우:
```bash
ralph-setup my-new-project
```
이 명령어는 프로젝트 디렉토리를 생성하고, Git을 초기화하며, `.ralph/` 설정 디렉토리를 생성합니다.
### 기존 요구사항 가져오기
이미 PRD 문서나 요구사항 명세서가 있는 경우:
```bash
ralph-import path/to/your-prd.md
```
Ralph이 문서를 파싱하여 작업 목록을 추출하고 구조화된 `fix_plan.md`를 생성합니다.
## .ralph/ 디렉토리 구조
frankbria의 메모리와 설정은 `.ralph/` 디렉토리에 집중되어 있습니다:
```
.ralph/
├── PROMPT.md # 프로젝트 목표와 컨텍스트
├── fix_plan.md # 작업 목록 (prd.json과 유사한 역할)
├── AGENT.md # 빌드/테스트 명령어 (자동 관리)
├── specs/ # 상세 요구사항 문서
│ ├── feature-a.md
│ └── feature-b.md
└── sessions/ # 세션 지속성 데이터
├── current.json
└── history/
```
**snarktank/ralph과의 비교**:
| frankbria | snarktank | 역할 |
| ------------- | ----------------------------------------- | --------------- |
| `PROMPT.md` | `prd.json`의 `projectName` + `description` | 프로젝트 목표 정의 |
| `fix_plan.md` | `prd.json`의 `userStories` | 작업 목록과 진행 상황 |
| `AGENT.md` | `CLAUDE.md` / `AGENTS.md` | 빌드 명령어와 프로젝트 규칙 |
| `specs/` | `prd.json`의 `notes` 필드 | 상세 요구사항 |
| `sessions/` | 없음 (매번 새 프로세스) | 세션 상태 추적 |
`AGENT.md`는 **자동으로 관리**됩니다 — Ralph이 실행 과정에서 발견한 프로젝트 규칙에 따라 이 파일을 자동으로 업데이트하며, snarktank/ralph의 `progress.txt`와 유사하지만 더 구조화되어 있습니다.
## 핵심 명령어
### 기본 실행
```bash
# Ralph 루프 시작
ralph
# 실시간 모니터링 포함
ralph --monitor
# tmux에서 시작 (장시간 실행 시 권장)
ralph --live
```
### 모니터링 대시보드
```bash
# 독립적으로 모니터링 시작
ralph-monitor
```
`ralph-monitor`는 tmux 대시보드를 열어 다음을 실시간으로 표시합니다:
* 현재 실행 중인 작업
* 완료/미완료 작업 수
* API 호출 횟수 및 비용 추정
* Circuit Breaker 상태
* 최근 오류 로그
### 주요 파라미터
| 파라미터 | 설명 | 기본값 |
| ----------------- | -------------- | ----- |
| `--resume` | 마지막 중단 지점부터 계속 | - |
| `--calls ` | 최대 API 호출 횟수 | 100 |
| `--timeout ` | 타임아웃 시간(분) | 300 |
| `--monitor` | 실시간 모니터링 활성화 | false |
| `--live` | tmux에서 실행 | false |
```bash
# API 호출 50회 제한, 2시간 타임아웃
ralph --calls 50 --timeout 120
# 마지막 중단 지점부터 계속
ralph --resume
```
## 안전 메커니즘
frankbria의 가장 큰 차별화 특성은 다계층 안전 메커니즘입니다.
### Circuit Breaker
Circuit Breaker는 "진행 없음"을 감지하면 자동으로 루프를 중단하여 무의미한 API 소모를 방지합니다:
**연속 진행 없음 감지**: 연속 N번의 반복에서 새로운 작업이 완료되지 않으면 Circuit Breaker가 작동합니다.
**동일 오류 감지**: 동일한 오류 메시지가 연속으로 발생하면 AI가 무한 루프에 빠진 것이므로 Circuit Breaker가 작동합니다.
### 속도 제한
기본적으로 100 calls/hour로 제한하여 예상치 못한 API 요금 폭증을 방지합니다. 파라미터를 통해 조정할 수 있습니다:
```bash
ralph --calls 200 # 200 calls로 상향
```
### 5시간 API 한도 3단계 감지
Anthropic API에는 5시간 슬라이딩 윈도우 사용 한도가 있습니다. frankbria에는 3단계 감지가 내장되어 있습니다:
1. **사전 감지**: 각 API 호출 전에 잔여 한도를 추정합니다
2. **응답 감지**: API 응답의 rate limit headers를 파싱합니다
3. **후퇴 전략**: 한도에 가까워지면 자동으로 호출 빈도를 낮춥니다
### 세션 만료 관리
기본 세션 유효 기간은 24시간입니다. 초과 시 세션 데이터를 자동으로 정리하여 만료된 컨텍스트가 후속 실행에 영향을 미치는 것을 방지합니다.
## 지능형 종료 감지
frankbria는 단순히 모든 작업이 완료된 후 종료하지 않습니다. **이중 조건 종료 게이트**를 사용합니다:
```
종료 조건 = completion_indicators >= 2 AND EXIT_SIGNAL: true
```
**completion\_indicators**는 AI 출력에서 감지된 완료 신호 수이며, 다음을 포함합니다:
* "모든 작업이 완료되었습니다"
* "더 이상 할 일이 없습니다"
* 테스트 전부 통과
* fix\_plan.md의 모든 항목이 done으로 표시됨
**EXIT\_SIGNAL**은 AI가 출력에서 명시적으로 선언한 종료 의도입니다.
왜 두 가지 조건이 필요할까요? **조기 종료를 방지**하기 위해서입니다. 단일 신호는 오판일 수 있습니다 — 예를 들어 AI가 "작업 완료"라고 말했지만 실제로는 현재 story만 완료한 경우입니다. 이중 조건을 통해 여러 독립적인 신호가 모두 완료를 확인한 경우에만 실제로 종료합니다.
## snarktank/ralph과의 비교
| 차원 | snarktank/ralph | frankbria/ralph-claude-code |
| ------------ | -------------------- | ------------------------------------- |
| **구현 방식** | 외부 bash 루프 (매번 새 세션) | 외부 bash 루프 (`--continue`로 세션 재사용) |
| **세션 모드** | 매번 새로 생성 | 기본 재사용 (`--no-continue`로 새로 생성 전환 가능) |
| **컨텍스트** | 매번 새로 생성 | `--continue`를 통해 반복 간 축적 |
| **설치** | Skill 복사 | install.sh + 대화형 마법사 |
| **작업 형식** | prd.json | PROMPT.md + fix\_plan.md |
| **모니터링** | 수동 `cat`/`jq` | 내장 tmux 대시보드 |
| **안전 메커니즘** | max\_iterations | Circuit Breaker + 속도 제한 + 타임아웃 |
| **작업 소스** | PRD only | beads / GitHub Issues / PRD |
| **적합한 시나리오** | 장기 AFK, 대량 반복 | 단기\~중기 반복, 모니터링 필요 |
### 핵심 차이: 엔지니어링 수준
둘 다 외부 bash 루프가 새로운 Claude 프로세스를 시작합니다. 핵심 차이는 세션 관리 방식에 있지 않으며(frankbria는 `--no-continue`로 새 세션 모드로 전환 가능), **엔지니어링 수준**에 있습니다:
* **snarktank**: 극도로 간결한 스크립트, 수백 줄의 bash, 루프 자체에 집중
* **frankbria**: 완전한 엔지니어링 도구 체인 — 모니터링 대시보드, Circuit Breaker, 속도 제한, 세션 만료 관리
frankbria는 기본적으로 `--continue` 세션 재사용을 활성화하며, 짧은 작업에 적합합니다. 긴 작업의 경우 `--no-continue`로 전환하여 새 세션 모드를 사용하면 snarktank와 동일한 Context Rot 방지 효과를 얻을 수 있으며, 동시에 frankbria의 엔지니어링 장점도 유지할 수 있습니다.
### 세션 재사용 비활성화 방법
frankbria는 `--continue`를 비활성화하는 세 가지 방법을 제공합니다:
```bash
# 방법 1: 명령줄 파라미터
ralph --no-continue
# 방법 2: 환경 변수
export CLAUDE_USE_CONTINUE=false
# 방법 3: .ralphrc 설정
SESSION_CONTINUITY=false
```
비활성화하면 frankbria의 동작은 snarktank와 동일해지지만(매번 새 세션), 모든 엔지니어링 도구(모니터링, Circuit Breaker, 속도 제한 등)는 유지됩니다.
## Context Rot의 현실적 트레이드오프
세션 재사용 방식의 선택은 본질적으로 Context Rot와 시작 오버헤드 사이의 트레이드오프입니다:
**짧은 작업 (\< 50k tokens)**: 세션 재사용이 더 유리합니다. 컨텍스트가 아직 열화되지 않았고, 처음 몇 번의 반복에서 쌓인 기억을 이후에 활용할 수 있습니다. 매번 새 세션을 생성하는 시작 오버헤드가 오히려 낭비입니다.
**긴 작업 (100k+ tokens)**: 새 세션이 더 안정적입니다. 100k tokens를 초과하면 Context Rot가 눈에 띄게 악화되어, 축적된 컨텍스트가 자산에서 부채로 변합니다. 새 세션은 시작 오버헤드가 있지만, 매번 최적의 상태를 유지합니다.
**실전 권장 사항**:
| 시나리오 | 권장 | 이유 |
| ------------------ | ---------------------------------------- | ------------------------- |
| 5개 미만의 작은 작업 | frankbria (기본 모드) | 빠른 시작, 컨텍스트 재사용 가능 |
| 10개 이상의 작업, AFK 필요 | snarktank 또는 frankbria + `--no-continue` | Context Rot 방지, 더 안정적 |
| 실시간 모니터링 필요 | frankbria | 내장 대시보드 |
| 작업량 불확실 | frankbria + `--no-continue` | 엔지니어링 도구 + Context Rot 방지 |
frankbria 사용자는 작업 규모에 따라 유연하게 선택할 수 있습니다: 짧은 작업에는 기본 `--continue` 모드를, 긴 작업에는 `--no-continue` 모드로 전환합니다. snarktank와 비교하여, frankbria의 장점은 어떤 모드에서든 완전한 엔지니어링 도구 체인을 유지한다는 것입니다.
## 요약
frankbria/ralph-claude-code는 Ralph 방법론의 엔지니어링된 구현 방향을 대표합니다. snarktank의 간결함을 일부 희생하는 대신, 더 완벽한 모니터링, 안전 및 설정 기능을 제공합니다.
어떤 구현을 선택할지는 구체적인 요구사항에 따라 달라집니다 — "더 올바른" 답은 없고, "더 적합한" 선택만 있을 뿐입니다.
***
**추가 읽을거리**:
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) — 핵심 원리와 방법론
* [snarktank/ralph 실전 가이드](/ko/docs/notes/ralph-wiggum/snarktank) — 극도로 간결한 외부 루프 구현
* [GSD 심층 분석](/ko/docs/notes/gsd/concept) — Ralph을 기반으로 구축한 완전한 컨텍스트 엔지니어링 시스템
* [Claude 시스템 아키텍처 전체 분석](/ko/docs/notes/claude-architecture) — Hooks, Subagent 등 구성 요소의 전체 아키텍처 이해
# Ralph의 실용 가이드
## 소개
[이전 글](/ko/docs/notes/ralph-wiggum/concept)에서 우리는 Ralph의 핵심 원칙인 무한 루프 + 매번 새로운 컨텍스트 + 파일을 유일한 정보 소스로 이해했습니다. 이 세 가지 기둥은 간단해 보이지만 개념을 이해하는 것부터 실제로 실행하는 것까지 세부적으로 해결해야 할 부분이 많습니다.
이 기사에서는 시작해 보겠습니다. [snarktank/ralph](https://github.com/snarktank/ralph)를 사용하여 설치부터 실행까지 전체 과정을 완료하는 방법을 배우게 됩니다. snarktank/ralph는 커뮤니티에서 가장 성숙한 Ralph 구현 중 하나입니다(별 10,000개 이상). Claude Code와 Amp라는 두 가지 도구와 PRD 생성, JSON 변환 및 자동화된 실행을 위한 완전한 도구 체인을 지원합니다.
## 전제조건
시작하기 전에 환경이 다음 요구 사항을 충족하는지 확인하세요.
| 종속성 | 설명 |
| --------------- | ------------------------------------------------------------- |
| **AI 프로그래밍 도구** | 클로드 코드(`npm install -g @anthropic-ai/claude-code`) 또는 Amp CLI |
| **jq** | JSON 처리 도구(macOS: `brew install jq`) |
| **힘내** | 프로젝트는 Git 저장소여야 합니다 |
```bash
# 检查依赖
claude --version # Claude Code CLI
jq --version # JSON 处理
git --version # Git
```
## 설치 및 구성
snarktank/ralph는 사용 시나리오에 따라 선택할 수 있는 다양한 설치 방법을 제공합니다.
### 방법 1: Claude Code에서 직접 설치(권장)
가장 쉬운 방법은 Claude Code 대화에 GitHub 링크를 붙여넣고 Claude가 자동으로 설치를 완료하도록 하는 것입니다.
```
Install this skill for me: https://github.com/snarktank/ralph
```
Claude Code는 자동으로 저장소를 복제하고 기술 파일을 올바른 위치에 복사합니다. `/prd` 및 `/ralph` 명령은 설치 후에 사용할 수 있습니다.
### 방법 2: Claude Code 마켓 설치
market 명령을 통해 설치:
```bash
# 添加并安装插件
/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace
```
설치 후 `/prd`(PRD 생성) 및 `/ralph`(JSON으로 변환) 두 가지 기술을 사용할 수 있습니다.
### 방법 3: 수동 스킬 설치(Claude Code/Amp)
기술 파일을 해당 도구의 전역 구성 디렉터리에 수동으로 복사합니다.
```bash
# 先克隆仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# Claude Code 用户
cp -r /tmp/ralph/skills/prd ~/.claude/skills/
cp -r /tmp/ralph/skills/ralph ~/.claude/skills/
# Amp 用户
cp -r /tmp/ralph/skills/prd ~/.config/amp/skills/
cp -r /tmp/ralph/skills/ralph ~/.config/amp/skills/
```
`/prd` 및 `/ralph` 명령은 설치 후에 사용할 수 있습니다.
### 방법 4: 프로젝트 수준 설치
Ralph 스크립트를 프로젝트에 직접 복사합니다. 팀 공유 또는 사용자 정의 스크립트가 필요한 시나리오에 이상적입니다.
```bash
# 克隆 Ralph 仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# 复制核心文件到项目
mkdir -p scripts/ralph
cp /tmp/ralph/ralph.sh scripts/ralph/
cp /tmp/ralph/CLAUDE.md scripts/ralph/ # Claude Code 用户
# 或
cp /tmp/ralph/prompt.md scripts/ralph/ # Amp 用户
# 赋予执行权限
chmod +x scripts/ralph/ralph.sh
```
설치가 완료되면 프로젝트 구조는 다음과 같습니다.
```
your-project/
├── scripts/ralph/
│ ├── ralph.sh # 核心循环脚本
│ └── CLAUDE.md # Claude Code 的 Prompt 模板
├── tasks/ # PRD 文件目录(执行时自动创建)
│ └── prd.json # 你的任务定义
└── ...
```
> **제안**: 첫 번째 방법이 가장 쉽습니다. GitHub 링크를 Claude Code에 전달하면 됩니다. 설치 프로세스를 수동으로 제어하려면 방법 2(market 명령) 또는 방법 3(수동 복사)을 선택하세요. 스크립트를 팀과 공유하거나 사용자 정의해야 하는 경우 방법 4를 선택하십시오.
***
## 핵심 파일 구조
Ralph의 메모리는 전적으로 파일 시스템에 의존합니다. Ralph를 잘 사용하기 위해서는 각 파일의 역할을 이해하는 것이 필수입니다.
### ralph.sh - 루프 엔진
이것이 Ralph의 핵심입니다. 새로운 AI 인스턴스를 지속적으로 생성하는 bash 스크립트입니다.
```bash
# 基本用法
./scripts/ralph/ralph.sh [max_iterations] # 默认:Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # 使用 Claude Code
```
각 반복에서 ralph.sh는 다음 단계를 수행합니다.
1. 기능 분기 생성(prd.json의 `branchName`에서)
2. 우선순위가 가장 높은 미완성 스토리(`passes: false`)를 선택하세요.
3. 이 스토리를 구현하기 위해 **새** AI 인스턴스를 생성합니다.
4. 품질 검사 실행(유형 검사, 테스트)
5. 통과 확인 → git commit; 검사 실패 → 다음 반복으로 남겨두기
6. prd.json을 업데이트하고 스토리를 `passes: true`로 표시합니다.
7. 배운 내용을 Progress.txt에 추가하세요.
8. 모든 스토리가 완료되거나 최대 반복 횟수에 도달할 때까지 반복합니다.
기본 반복 제한은 10입니다. 프로젝트 복잡성에 따라 조정합니다.
```bash
# 简单项目
./scripts/ralph/ralph.sh --tool claude 10
# 复杂项目
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json - 작업 정의
이것이 Ralph의 "두뇌"입니다. 모든 작업이 정의되는 곳입니다. 이는 플랫 JSON 파일입니다.
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists with valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve MDX component props format"
}
]
}
```
**필드 설명**:
| 필드 | 설명 |
| -------------------- | ----------------------------- |
| `projectName` | 로그 및 브랜치 이름 지정에 사용되는 프로젝트 이름 |
| `branchName` | Git 분기 이름 - Ralph가 자동으로 생성합니다 |
| `id` | 스토리 고유 식별자, 권장되는 `US-001` 형식 |
| `title` | 짧은 제목 |
| `description` | 자세한 설명 - 구체적일수록 좋습니다 |
| `acceptanceCriteria` | 승인 기준 목록 - **가장 중요한 필드입니다** |
| `priority` | 우선 순위 번호 - 숫자가 작을수록 먼저 실행됩니다 |
| `passes` | 완료 여부 - Ralph가 자동으로 업데이트됩니다 |
| `dependsOn` | 종속 스토리 ID 목록 |
| `notes` | 추가 팁 및 상황 |
### Progress.txt - 경험 기록
이것이 랄프의 '장기 기억'이다. 각 반복 후에 AI는 추가 학습 경험을 추가합니다.
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
다음 반복을 위한 새로운 Claude 인스턴스는 이 파일을 읽고 즉시 이전 경험을 모두 얻습니다. 이것이 Ralph의 실행이 점점 더 좋아지는 이유입니다. **지식은 반복 간에 축적되지만 컨텍스트는 깨끗하게 유지됩니다**.
### AGENTS.md - 지속적인 지식 기반
Ralph는 Progress.txt 외에도 프로젝트의 `AGENTS.md`(또는 `CLAUDE.md`) 파일도 업데이트합니다. Claude Code와 Amp는 시작할 때 자동으로 이 파일을 읽습니다.
Progress.txt와 달리 AGENTS.md는 안정적인 프로젝트 간 지식을 기록합니다.
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## PRD 작성
PRD(제품 요구 사항 문서)의 품질이 Ralph의 실행 결과를 직접적으로 결정합니다. 글을 잘 썼어요. 순조롭게 항해를 하게 됐어요, 랄프. 잘못 작성하면 Ralph는 같은 이야기에서 반복적으로 실패할 것입니다.
### 스킬을 사용하여 PRD 생성
snarktank/ralph 스킬이 설치되어 있으면 대화식으로 PRD를 생성할 수 있습니다.
```bash
# 在 Claude Code 或 Amp 中
/prd I want to add i18n support to the blog, translating all Chinese content to English
```
AI는 몇 가지 명확한 질문(관련된 문서, 기술 스택 제한, 품질 표준 등)을 묻고 구조화된 PRD 문서를 생성합니다.
생성 후 `/ralph` 명령을 사용하여 PRD를 `prd.json` 형식으로 변환합니다.
```bash
/ralph # 转换 PRD 为 prd.json
```
### 수동으로 PRD 작성
prd.json을 직접 작성할 수도 있습니다. 다음은 주요 설계 원칙입니다.
**원칙 1: 스토리 세분성은 보통입니다**
각 스토리는 한 번의 반복으로 완료될 수 있을 만큼 작아야 하며, 독립적으로 가치를 전달할 수 있을 만큼 커야 합니다.
```json
// ❌ 太大:一次迭代完不成
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, forgot password, OAuth, permission management..."
}
// ❌ 太小:没有独立价值
{
"id": "US-001",
"title": "Create email field on User table",
"description": "Add email field to User model"
}
// ✅ 刚好:一次迭代能完成,有独立价值
{
"id": "US-001",
"title": "Implement email/password login",
"description": "Create login API and login page with email/password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form submits successfully",
"All tests pass"
]
}
```
**경험 법칙**: 스토리에는 1~~3개의 파일 수정이 포함되며 3~~5개의 승인 기준이 있습니다.
**원칙 2: 승인 기준은 자동으로 검증 가능해야 합니다**
Ralph는 스토리가 완전한지 확인해야 하므로 승인 기준을 객관적으로 평가할 수 있어야 합니다.
```json
// ❌ 模糊的标准
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ 可验证的标准
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**원칙 3: dependencyOn을 사용하여 실행 순서 제어**
일부 이야기에는 종속성이 있습니다. `dependsOn` 필드는 Ralph가 올바른 순서로 실행되도록 보장합니다.
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**원칙 4: 메모에 맥락을 제공하세요**
메모 필드는 AI에게 추가 힌트를 제공합니다. AI가 모를 수도 있는 귀하가 알고 있는 정보를 여기에 적어 보십시오.
```json
{
"notes": "Project uses fumadocs framework, i18n files follow .en.mdx suffix naming. Reference content/docs/notes/speckit/concept.en.mdx for translation style."
}
```
***
## Ralph 루프 실행
PRD가 준비되면 루프를 실행할 차례입니다.
### 실행 시작
```bash
# 使用 Claude Code,默认 10 次迭代
./scripts/ralph/ralph.sh --tool claude
# 指定迭代次数
./scripts/ralph/ralph.sh --tool claude 30
# 使用 Amp(默认)
./scripts/ralph/ralph.sh 20
```
### 실행 과정
시작하면 다음과 유사한 출력이 표시됩니다.
```
=== Ralph Loop - Iteration 1 ===
Branch: ralph/i18n-translation
Selected story: US-001 - Translate homepage metadata
Spawning fresh Claude instance...
[Claude Code executing...]
Quality check: pnpm types:check ... PASSED
Committing: feat: [US-001] - Translate homepage metadata
Updating prd.json: US-001 passes: true
Appending to progress.txt
=== Ralph Loop - Iteration 2 ===
Selected story: US-002 - Translate blog post hello-world
Spawning fresh Claude instance...
```
각 반복은 Claude의 완전히 새로운 인스턴스입니다. prd.json을 읽어 무엇을 해야 할지 알고, Progress.txt를 읽어 이전에 배운 내용을 알고 있습니다.
\###완료 신호
모든 스토리가 `passes: true`로 표시되면 Ralph는 완료 신호를 출력하고 종료합니다.
```
All stories completed!
COMPLETE
```
### 모니터링 및 디버깅
Ralph가 실행되는 동안 다음 명령을 사용하여 진행 상황을 볼 수 있습니다.
```bash
# 查看每个 story 的完成状态
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 查看经验日志
cat progress.txt
# 查看最近的 git 提交
git log --oneline -10
# 实时跟踪 Ralph 输出
tail -f progress.txt
```
### 자동 보관
다른 `branchName`을 사용하여 새 기능을 시작하면 Ralph는 마지막 실행의 파일을 `archive/YYYY-MM-DD-feature-name/` 디렉터리에 자동으로 보관하여 작업 디렉터리를 깨끗하게 유지합니다.
***
## 피드백 루프 및 품질 게이트 제어
Ralph의 "자가 수정" 능력은 전적으로 피드백 루프의 품질에 달려 있습니다. 피드백 루프가 없으면 Ralph는 맹목적으로 반복되는 스크립트일 뿐입니다. 즉, 코드가 올바른지 확인할 수 없는 상태에서 계속해서 코드를 대량 생성합니다.
### 품질 검사 구성
CLAUDE.md(또는 Prompt.md)에서 QA 명령을 정의합니다.
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 품질 접근 제어 수준
| 계층 구조 | 도구 | 포착된 문제 |
| -------- | --------------- | ---------------------- |
| 즉각적인 피드백 | TypeScript 컴파일러 | 유형 오류, 구문 오류 |
| 기능 검증 | 단위 테스트 | 논리 오류, 극단적인 경우 |
| 통합 검증 | 빌드 명령 | 종속성 문제, 구성 오류 |
| 런타임 검증 | 개발자 브라우저 기술 | UI 렌더링 문제(프런트 엔드 프로젝트) |
> 프런트 엔드 스토리의 경우 Ralph는 "개발자 브라우저 기술을 사용하여 브라우저에서 확인"이라는 허용 기준을 추가할 것을 권장합니다. AI가 실제로 브라우저를 열어 페이지가 올바르게 렌더링되는지 확인하도록 합니다.
### 품질검사에 실패한 경우
스토리가 반복적으로 QA에 실패하는 경우 Ralph는 동일한 스토리를 무한정 재시도하지 않습니다. 반복 상한에 도달한 후 중지되고 현재 상태를 유지합니다. 다음을 수행할 수 있습니다.
1. 진행이 중단된 이유를 이해하려면 Progress.txt를 확인하세요.
2. 문제를 수동으로 수정한 후 다시 실행하세요.
3. 스토리 세분화 조정(너무 클 수도 있음)
4. 메모 필드에 더 많은 컨텍스트를 추가하세요.
***
## 프롬프트 사용자 정의
Ralph의 프롬프트 템플릿(CLAUDE.md 또는 Prompt.md)은 AI 동작을 제어하는 기본 수단입니다. 설치 후에는 프로젝트에 따라 사용자 정의해야 합니다.
### 주요 맞춤 아이템
**1. 프로젝트별 품질 명령**
```markdown
## Project-Specific Commands
- Typecheck: `pnpm types:check` (not `tsc` or `pnpm typecheck`)
- Test: `pnpm vitest run`
- Build: `pnpm build`
- Lint: `pnpm lint`
```
**2. 코드 스타일 제약**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**3. 알려진 함정**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**4. 막혔을 때 처리**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 실제 사례: Ralph로 블로그 번역하기
Ralph가 실제로 어떻게 작동하는지 보여주기 위해 다음은 실제 예입니다. Ralph 스타일 자율 에이전트를 사용하여 전체 블로그를 중국어에서 영어로 번역합니다.
### 프로젝트 설정
이 프로젝트에는 22개 이상의 콘텐츠 파일(블로그 게시물, 문서, 탐색 메타데이터)을 중국어에서 영어로 번역해야 합니다. 이 프로젝트는 fumadocs의 Next.js 블로그를 기반으로 하며 i18n을 지원합니다. 작업은 `prd.json` 파일에 정의되어 있으며 각각 명확한 승인 기준이 있는 16개의 사용자 스토리를 포함합니다.
```
scripts/ralph/
├── prd.json # 16 个 user story,带验收标准
└── progress.txt # 经验日志,每个 story 完成后更新
```
모든 사용자 스토리는 일관된 패턴을 따릅니다.
* **명시적 결과물**: "콘텐츠 만들기/blog/xxx.en.mdx"
* **검증 가능한 기준**: "Typecheck 통과", "내부 링크는 /en/ 접두사를 사용합니다."
* **기술적 제약**: "코드 블록을 번역되지 않은 상태로 유지", "QuoteCard에서 defaultLang='en' 설정"
### 실행 모드
에이전트는 Ralph 방법론의 핵심 원칙을 따릅니다.
1. **문서는 진실의 원천입니다**: `prd.json`는 어떤 이야기가 전달되었는지 추적합니다(`passes: true/false`). `progress.txt` 반복 간 경험을 얻습니다. "Typecheck 명령은 `pnpm typecheck`이 아니라 `pnpm types:check`입니다."
2. **자동화된 품질 게이트**: 각 번역 후 `pnpm types:check`이 실행되어 MDX 파일이 올바르게 컴파일되었는지 확인합니다. 유형 검사에 실패하면 제출하기 전에 수정하세요.
3. **점진적 발전**: 각 스토리는 설명 제출 정보(`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`)와 함께 독립적으로 제출되며 필요할 때 쉽게 롤백할 수 있습니다.
4. **병렬 실행**: 긴 기사의 경우 여러 하위 에이전트가 동시에 번역됩니다. 예를 들어 US-010(claude-skills 개념 + 실습), US-011(speckit 개념 + 실습) 및 US-012(claude-architecture + claude-subagent)가 병렬로 실행됩니다.
### 주요 교훈
| 체험 | 세부정보 |
| ------------------------- | -------------------------------------------------------------------------------- |
| **축적된 지식이 중요합니다** | 초기 스토리에서 발견된 패턴(QuoteCard `defaultLang`, 링크 접두어 규칙)을 통해 후속 스토리를 더 빠르게 완료할 수 있습니다 |
| **피드백 루프로서의 Typecheck** | 문제가 쌓이기 전에 누락된 가져오기 또는 잘못된 MDX를 찾아보세요 |
| **병렬화 및 확장 가능** | 6개의 번역 에이전트가 동시에 실행 중이며 완료 시간은 에이전트 1개 |
| **PRD 세분화가 중요** | 스토리당 범위 1-2 파일 - 안정적으로 완료할 수 있을 만큼 작고 의미가 있을 만큼 큽니다 |
| **진행 로그는 반복되는 실수를 방지합니다** | Progress.txt의 "코드베이스 패턴" 부분은 동일한 문제가 재발견되는 것을 방지하기 위한 지식 기반이 됩니다 |
### 결과
16개의 사용자 스토리가 모두 단일 세션에서 완료되었습니다. 8개의 Meta.en.json 탐색 파일이 생성되었고, 3개의 블로그 게시물이 번역되었으며, 12개의 문서 페이지가 번역되었으며, 전체 사이트 구축이 확인되었습니다. 승인 기준이 명확하고 피드백 루프(유형 확인)가 문제를 즉시 포착하므로 각 번역은 일관된 품질을 유지합니다.
이 프로젝트는 잘 정의된 작업 + 명확한 성공 기준 + 자동화된 검증 + 파일 시스템을 통한 증분 전달 등 Ralph의 **완전한 구현 모델**을 보여줍니다.
***
## 커뮤니티 구현 및 대안
snarktank/ralph가 유일한 옵션은 아닙니다. 이러한 각 구현에는 필요에 따라 고유한 장점과 단점이 있습니다.
| 자원 | 링크 | 지침 |
| ---------- | ----------------------------------------------------------------------------------- | -------------------------------------- |
| 스나크탱크/랄프 | [스나크탱크/랄프](https://github.com/snarktank/ralph) | 이 기사에서 사용된 가장 완벽한 기능은 |
| 랄프 오케스트라 | [mikeyobrien/ralph-orchestrator](https://github.com/mikeyobrien/ralph-orchestrator) | 더 많은 사용자 정의 옵션을 갖춘 Mickey O'Brien이 개발함 |
| 랄프 루프 에이전트 | [vercel-labs/ralph-loop-agent](https://github.com/vercel-labs/ralph-loop-agent) | AI SDK를 기반으로 한 Vercel의 구현 |
| 랄피 | [마이클시멜레스/랄피](https://github.com/michaelshimeles/ralphy) | Michael Shimeles의 경량 구현 |
### 대안: GSD
GSD는 엄밀히 말하면 Ralph의 "커뮤니티 구현"이 아니라 **대안**입니다. Ralph의 핵심 원칙(컨텍스트 관리, 원자성 작업)을 적용하지만 토론 → 계획 → 실행 → 확인이라는 보다 완전한 워크플로를 제공합니다.
| 자원 | 링크 | 지침 |
| ---------- | ------------------------------------------------------------ | -------------------------- |
| GSD(작업 완료) | [반짝이카우보이/젠장](https://github.com/glittercowboy/get-shit-done) | 아이디어부터 PRD, 실행까지 완벽한 프레임워크 |
Ralph가 너무 "원시적"이고 더 많은 프로세스 지원이 필요하다고 생각되면 GSD가 더 적합할 수 있습니다. 자세한 내용은 [GSD 심층 분석](/ko/docs/notes/gsd/concept)을 참조하세요.
***
## 추천 리소스
**공식 출처**:
| 자원 | 링크 | 지침 |
| --------------------- | --------------------------------------------------------------------------- | ---------- |
| Geoffrey Huntley의 블로그 | [ghuntley.com/ralph](https://ghuntley.com/ralph/) | 발명가의 원본 기사 |
| 랄프 위검 방법 | [ghuntley/ralph-wiggum 방법](https://github.com/ghuntley/how-to-ralph-wiggum) | 공식 사용자 가이드 |
**동영상 튜토리얼**:
| 자원 | 링크 | 지침 |
| ------------------- | ----------------------------------------------------------------------------- | ----------------------------------- |
| Ralph Wiggum 심층 토론 | [클로드 코드의 구현이 아닌 이유](https://www.youtube.com/watch?v=O2bBWDoxO4s) | Geoffrey Huntley가 공식 구현의 문제점을 설명합니다 |
| Ralph의 올바른 사용 | [Ralph Wiggum 루프를 잘못 사용하고 있습니다.](https://www.youtube.com/watch?v=I7azCAgoUHc) | Roman(Mentat) 사용법 시연 |
| Ralph에 대해 이야기해야 합니다 | [랄프에 대해 이야기 좀 해야겠어요](https://www.youtube.com/watch?v=Yr9O6KFwbW4) | 논란에 대한 테오의 분석 |
***
## 모범 사례 및 FAQ
### 비용 관리
Ralph의 자동 실행은 API 수수료가 지속적으로 발생함을 의미합니다. 여러 가지 통제 조치:
* **항상 `max_iterations`** 설정: 가장 기본적인 안전망입니다.
* **스토리 세분화를 합리적으로 유지**: 스토리가 너무 크면 여러 번의 반복이 필요합니다. 너무 상세한 이야기는 시작 오버헤드를 증가시킵니다.
* **소규모 테스트 우선**: 새 프로젝트를 먼저 3\~5회 반복 실행한 다음 프롬프트와 품질 액세스 제어가 제대로 작동하는지 확인한 후 확장합니다.
### 일반적인 함정
**함정 1: 스토리가 너무 방대함**
증상: 스토리가 반복적으로 실패하고 반복 횟수가 빠르게 소진됩니다.
해결 방법: 2\~3개의 작은 스토리로 나눕니다. "완전한 인증 시스템 구축"은 "로그인 API 구현" + "로그인 페이지 생성" + "JWT 미들웨어 추가"로 나누어집니다.
**트랩 2: 피드백 루프 없음**
증상: Ralph는 스토리가 완료되었다고 주장하지만 실제 코드에 문제가 있습니다.
해결책: 승인 기준에 실행 가능한 확인 명령을 추가하십시오. "코드가 작성되었습니다"는 승인 기준이 아닙니다. "pnpm 테스트를 모두 통과했습니다"입니다.
**트랩 3: Progress.txt가 사용되지 않습니다**
증상: 동일한 오류가 다른 반복에서 반복적으로 나타납니다.
해결 방법: 프롬프트 템플릿에서 "progress.txt를 읽고 그 안에 있는 규칙을 따르십시오"라고 명시적으로 지시하는지 확인하세요. AI가 자동으로 경험치를 추가하지 않는 경우 프롬프트에 "각 스토리가 완료된 후 Progress.txt에 경험치를 추가하세요"를 추가하세요.
**트랩 4: 잘못된 종속성 순서**
증상: 스토리가 아직 존재하지 않는 코드에 의존하여 구현이 실패합니다.
해결 방법: `dependsOn` 필드를 올바르게 설정하십시오. 인프라 이야기가 먼저 나오도록 하세요.
### FAQ
\*\*Q: Ralph와 공식 플러그인의 차이점은 무엇인가요? \*\*
핵심 차이점: snarktank/ralph는 반복할 때마다 새로운 프로세스를 생성하는 반면(완전히 새로운 컨텍스트) 공식 플러그인은 동일한 세션 내에서 반복됩니다(컨텍스트는 계속 축적됩니다). 자세한 내용은 [이전 기사 분석](/ko/docs/notes/ralph-wiggum/concept#the-problem-with-the-official-plugin)을 참조하세요.
\*\*Q: 실행 중에 prd.json을 수동으로 수정할 수 있나요? \*\*
그렇습니다. Ralph는 각 반복이 시작될 때 prd.json을 다시 읽습니다. 반복 간에 스토리 설명을 수정하거나, 새 스토리를 추가하거나, 수동으로 스토리를 `passes: true`(건너뛰기)로 표시할 수 있습니다.
\*\*Q: Ralph가 반복적으로 실패하는 스토리에 막히면 어떻게 해야 하나요? \*\*
1. Progress.txt를 확인하여 실패 이유를 파악하세요.
2. 노트에 더 많은 맥락을 추가하세요
3. 스토리 분할(너무 클 수도 있음)
4. 차단 문제를 수동으로 수정한 후 다시 실행하세요.
\*\*Q: Ralph가 실행되는 동안 다른 작업을 할 수 있나요? \*\*
그렇습니다. Ralph는 "Human on the Loop"로 설계되었으므로 쳐다볼 필요가 없습니다. AFK 모드에서는 퇴근 전 시작해 다음날 아침 결과를 확인해보세요. Ralph가 작업 중인 파일을 수정하지 마세요.
\*\*Q: 비용을 관리하는 방법은 무엇입니까? \*\*
세 가지 방법: 합리적인 `max_iterations`을 설정하고, 스토리 세분성을 적절하게 유지하고(낭비되는 반복을 줄이기 위해), 먼저 소규모 시험 실행을 수행하여 프로세스가 올바른지 확인합니다. 일반적으로 스토리가 10~~20개 있는 프로젝트의 경우 API 수수료는 $50~~100입니다.
***
## 요약
Ralph의 작업 흐름은 다섯 단계로 정리할 수 있습니다.
```
安装 → 编写 PRD → 配置质量门禁 → 运行循环 → 检查结果
```
핵심 개념은 동일하게 유지됩니다. **문서를 유일한 진실 소스로 삼고, 각 반복을 처음부터 시작하고, 품질 게이트를 제어하도록 하세요**.
이제 프로젝트로 돌아가서 prd.json을 준비하고 `./scripts/ralph/ralph.sh --tool claude`을 실행한 후 커피 한 잔을 만드세요.
### 추가 자료
* [Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept) - Ralph의 핵심 원칙을 다시 살펴봅니다.
* [GSD 심층 분석](/ko/docs/notes/gsd/concept) - Ralph를 기반으로 구축된 완벽한 컨텍스트 엔지니어링 시스템
* [클로드 스킬이란](/ko/docs/notes/claude-skills/concept) ——랄프의 PRD 스킬은 클로드 스킬입니다
* [Speckit 실용 가이드](/ko/docs/notes/speckit/practice) - 또 다른 구조화된 AI 프로그래밍 워크플로우
# snarktank/ralph 실전 가이드
## 서론
[이전 글](/ko/docs/notes/ralph-wiggum/concept)에서 Ralph의 핵심 원리인 무한 루프 + 매번 새로운 컨텍스트 + 파일을 진실의 원천으로 삼는 것을 살펴보았습니다. 세 가지 기둥은 간단해 보이지만, 이해에서 실제 구동까지는 적지 않은 세부 사항이 있습니다.
이번 글에서는 직접 실습해 보겠습니다. [snarktank/ralph](https://github.com/snarktank/ralph)는 Ralph 방법론의 **외부 루프 구현**으로, 매번 반복할 때마다 완전히 새로운 Claude 프로세스를 시작하여 Context Rot 문제를 근본적으로 해결합니다. 현재 커뮤니티에서 가장 완성도 높은 Ralph 구현 중 하나(10k+ stars)이며, Claude Code와 Amp 양 플랫폼을 지원하고 PRD 생성, JSON 변환, 자동 실행의 전체 도구 체인을 제공합니다.
> 또 다른 구현 방식은 [frankbria/ralph-claude-code](/ko/docs/notes/ralph-wiggum/frankbria)로, 완전한 엔지니어링 도구 체인(모니터링 대시보드, 서킷 브레이커, 속도 제한)을 제공하며 제어 가능성과 안전 메커니즘에 중점을 두고 있습니다. 둘의 비교는 해당 글을 참고하시기 바랍니다.
## 사전 요구 사항
시작하기 전에 다음 환경 조건이 충족되었는지 확인하십시오:
| 의존성 | 설명 |
| --------------- | ------------------------------------------------------------------- |
| **AI 프로그래밍 도구** | Claude Code (`npm install -g @anthropic-ai/claude-code`) 또는 Amp CLI |
| **jq** | JSON 처리 도구 (macOS: `brew install jq`) |
| **Git** | 프로젝트가 Git 저장소여야 합니다 |
```bash
# 의존성 확인
claude --version # Claude Code CLI
jq --version # JSON 처리
git --version # Git
```
## 설치 및 구성
가장 간단한 방법은 Claude Code 대화에서 GitHub 링크를 직접 붙여넣는 것입니다:
```
帮我安装这个 skill:https://github.com/snarktank/ralph
```
Claude Code가 자동으로 저장소를 클론하고 skill 파일을 올바른 위치에 복사합니다. 설치가 완료되면 `/prd`와 `/ralph` 명령을 사용할 수 있습니다.
> snarktank/ralph는 Marketplace 설치, 수동 skill 파일 복사, 프로젝트 수준 설치 등 다양한 방법도 지원합니다. 자세한 내용은 [GitHub 저장소 설명](https://github.com/snarktank/ralph)을 참고하십시오.
***
## 핵심 파일 구조
Ralph의 기억은 완전히 파일 시스템에 의존합니다. 각 파일의 역할을 이해하는 것이 Ralph를 잘 활용하기 위한 전제 조건입니다.
### ralph.sh — 루프 엔진
이것이 Ralph의 핵심입니다. 새로운 AI 인스턴스를 반복적으로 시작하는 bash 스크립트입니다.
```bash
# 기본 사용법
./scripts/ralph/ralph.sh [max_iterations] # 기본적으로 Amp 사용
./scripts/ralph/ralph.sh --tool claude [iterations] # Claude Code 사용
```
매번 반복할 때마다 ralph.sh는 다음과 같은 작업을 수행합니다:
1. 기능 브랜치 생성 (prd.json의 `branchName` 기반)
2. 가장 높은 우선순위의 미완료 story 선택 (`passes: false`)
3. 해당 story를 구현하기 위해 **완전히 새로운** AI 인스턴스 시작
4. 품질 검사 실행 (타입 체크, 테스트)
5. 검사 통과 → git commit; 실패 → 다음 반복으로 넘김
6. prd.json 업데이트, story를 `passes: true`로 표시
7. progress.txt에 이번에 배운 경험 추가
8. 모든 story가 완료되거나 반복 상한에 도달할 때까지 반복
기본 반복 상한은 10회입니다. 프로젝트 복잡도에 따라 조정하십시오:
```bash
# 간단한 프로젝트
./scripts/ralph/ralph.sh --tool claude 10
# 복잡한 프로젝트
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — 작업 정의
이것은 Ralph의 "두뇌"입니다. 모든 작업이 여기에 정의됩니다. 형식은 플랫한 JSON 파일입니다:
```json
{
"projectName": "블로그 i18n 번역",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "홈페이지 메타데이터 번역",
"description": "content/docs/meta.en.json을 생성하고 모든 네비게이션 항목의 영어 번역을 포함",
"acceptanceCriteria": [
"meta.en.json 파일이 존재하고 JSON 형식이 올바름",
"모든 네비게이션 제목이 영어로 번역됨",
"pnpm types:check 통과"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "기존 meta.json의 구조를 참고"
},
{
"id": "US-002",
"title": "블로그 게시글 hello-world 번역",
"description": "content/blog/hello-world.en.mdx를 생성하고 중국어에서 영어로 번역",
"acceptanceCriteria": [
"hello-world.en.mdx 파일이 존재",
"모든 QuoteCard 컴포넌트에 defaultLang='en' 설정",
"내부 링크에 /en/ 접두사 사용",
"코드 블록은 번역하지 않음",
"pnpm types:check 통과"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "MDX 컴포넌트의 props 형식을 유지할 것"
}
]
}
```
**필드 설명**:
| 필드 | 설명 |
| -------------------- | ---------------------------- |
| `projectName` | 프로젝트 이름, 로그 및 브랜치 명명에 사용 |
| `branchName` | Git 브랜치 이름, Ralph가 자동 생성 |
| `id` | Story 고유 식별자, `US-001` 형식 권장 |
| `title` | 간결한 제목 |
| `description` | 상세 설명, 구체적일수록 좋음 |
| `acceptanceCriteria` | 인수 기준 목록 — **가장 중요한 필드** |
| `priority` | 우선순위 숫자, 작을수록 먼저 실행 |
| `passes` | 완료 여부, Ralph가 자동 업데이트 |
| `dependsOn` | 의존하는 story ID 목록 |
| `notes` | 추가 메모 및 힌트 |
### progress.txt — 경험 로그
이것은 Ralph의 "장기 기억"입니다. 매번 반복이 끝난 후 AI가 이번에 배운 내용을 여기에 추가합니다:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
다음 반복의 새로운 Claude 인스턴스가 이 파일을 읽어 이전의 모든 경험을 즉시 획득합니다. 이것이 Ralph가 반복할수록 점점 원활하게 실행되는 이유입니다 — **지식은 반복 간에 축적되지만 컨텍스트는 깨끗하게 유지됩니다**.
### AGENTS.md — 영구 지식 베이스
progress.txt 외에도 Ralph는 프로젝트의 `AGENTS.md` 파일(또는 `CLAUDE.md`)도 업데이트합니다. Claude Code와 Amp 모두 시작 시 이 파일들을 자동으로 읽습니다.
progress.txt와 달리, AGENTS.md에는 **안정적이고 프로젝트 간에 범용적인 지식**이 기록됩니다:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## PRD 작성
PRD(Product Requirements Document)의 품질이 Ralph의 실행 효과를 직접적으로 결정합니다. 잘 작성하면 Ralph가 순조롭게 진행되고, 잘못 작성하면 같은 story에서 반복적으로 실패하게 됩니다.
### Skill을 사용한 PRD 생성
snarktank/ralph의 skill을 설치했다면 대화형 방식으로 PRD를 생성할 수 있습니다:
```bash
# Claude Code 또는 Amp에서
/prd 我想为博客系统添加 i18n 支持,需要将所有中文内容翻译成英文
```
AI가 일련의 명확화 질문(관련 파일, 기술 스택 제약 조건, 품질 기준 등)을 한 다음 구조화된 PRD 문서를 생성합니다.
생성 후 `/ralph` 명령으로 PRD를 `prd.json` 형식으로 변환합니다:
```bash
/ralph # PRD를 prd.json으로 변환
```
### 수동 PRD 작성
prd.json을 직접 작성할 수도 있습니다. 다음은 핵심 설계 원칙입니다.
**원칙 1: Story 단위를 적절하게**
각 story는 한 번의 반복으로 완료할 수 있을 만큼 작아야 하고, 독립적인 납품 가치가 있을 만큼 커야 합니다.
```json
// ❌ 너무 큼: 한 번의 반복으로 완료 불가
{
"id": "US-001",
"title": "완전한 사용자 인증 시스템 구축",
"description": "회원가입, 로그인, 비밀번호 찾기, OAuth, 권한 관리 구현..."
}
// ❌ 너무 작음: 독립적인 가치 없음
{
"id": "US-001",
"title": "User 테이블의 email 필드 생성",
"description": "User 모델에 email 필드 추가"
}
// ✅ 적절함: 한 번에 완료 가능, 독립적 가치 있음
{
"id": "US-001",
"title": "이메일 비밀번호 로그인 구현",
"description": "로그인 API와 로그인 페이지 생성, 이메일 비밀번호 인증 지원",
"acceptanceCriteria": [
"POST /api/auth/login이 email + password를 수신",
"JWT token 반환",
"로그인 페이지 폼 제출 가능",
"모든 테스트 통과"
]
}
```
**경험 법칙**: 하나의 story는 1~~3개의 파일 수정을 포함하고, 3~~5개의 인수 기준을 갖습니다.
**원칙 2: 인수 기준은 자동 검증 가능해야 함**
Ralph는 story의 완료 여부를 판단해야 하므로 인수 기준은 객관적으로 판정할 수 있어야 합니다:
```json
// ❌ 모호한 기준
"acceptanceCriteria": [
"코드 품질이 좋음",
"성능이 좋음",
"사용자 경험이 원활함"
]
// ✅ 검증 가능한 기준
"acceptanceCriteria": [
"pnpm types:check 통과",
"pnpm test 통과",
"API 응답 시간 < 200ms",
"파일 src/auth/login.ts가 존재하고 loginHandler 함수를 export함"
]
```
**원칙 3: dependsOn으로 순서 제어**
story 간에 의존 관계가 있는 경우 `dependsOn` 필드로 Ralph가 올바른 순서로 실행하도록 보장합니다:
```json
{
"userStories": [
{
"id": "US-001",
"title": "데이터베이스 스키마 생성",
"dependsOn": []
},
{
"id": "US-002",
"title": "사용자 등록 API 구현",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "로그인 페이지 구현",
"dependsOn": ["US-002"]
}
]
}
```
**원칙 4: notes에 컨텍스트 제공**
notes 필드는 AI에게 제공하는 추가 힌트입니다. 여러분은 알고 있지만 AI가 모를 수 있는 정보를 여기에 작성하십시오:
```json
{
"notes": "프로젝트는 fumadocs 프레임워크를 사용합니다. i18n 파일 명명 규칙은 .en.mdx 접미사입니다. content/docs/notes/speckit/concept.en.mdx의 번역 스타일을 참고하십시오."
}
```
***
## Ralph Loop 실행
PRD가 준비되면 루프를 시작합니다.
### 실행 시작
```bash
# Claude Code 사용, 기본 10회 반복
./scripts/ralph/ralph.sh --tool claude
# 반복 횟수 지정
./scripts/ralph/ralph.sh --tool claude 30
# Amp 사용 (기본값)
./scripts/ralph/ralph.sh 20
```
### 실행 과정
시작하면 다음과 같은 출력을 볼 수 있습니다:
```
Starting Ralph - Tool: claude - Max iterations: 35
===============================================================
Ralph Iteration 1 of 35 (claude)
===============================================================
## US-001 Complete
**Summary of what was done:**
1. Created meta.en.json with all navigation items translated
2. Ran pnpm types:check — PASSED
3. Committed: feat: [US-001] - Translate homepage metadata
There are still **15 user stories with `passes: false`** remaining.
The next story is **US-002: 翻译博客文章 hello-world**.
Iteration 1 complete. Continuing...
===============================================================
Ralph Iteration 2 of 35 (claude)
===============================================================
```
매번 반복은 완전히 새로운 Claude 인스턴스입니다. prd.json을 읽어 현재 무엇을 해야 하는지 파악하고, progress.txt를 읽어 이전에 무엇을 배웠는지 파악합니다.
### 완료 신호
모든 story가 `passes: true`로 표시되면 Ralph는 완료 신호를 출력하고 종료합니다:
```
All stories completed!
COMPLETE
```
### 모니터링 및 디버깅
Ralph 실행 중에 다음 명령으로 진행 상황을 확인할 수 있습니다:
```bash
# 각 story의 완료 상태 확인 (아이콘 포함, 더 직관적)
cat tasks/prd.json | python3 -c "
import json,sys
for s in json.load(sys.stdin)['userStories']:
print(f'{\"✅\" if s[\"passes\"] else \"⬜\"} {s[\"id\"]}: {s[\"title\"]}')"
# 또는 jq로 확인
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 경험 로그 확인
cat progress.txt
# 최근 git 커밋 확인
git log --oneline -10
# Ralph 출력 실시간 확인
tail -f progress.txt
# 완료 후 메인 브랜치 대비 전체 변경 사항 확인
git diff main...ralph/your-branch-name --stat
```
### 중단 및 재개
Ralph 실행 시간이 길어질 수 있지만, 중간에 중단하는 것은 완전히 안전합니다:
* **중단**: `Ctrl+C`로 직접 중단할 수 있습니다. 완료된 story(`passes: true`)는 손실되지 않으며, 이미 commit되고 prd.json에 기록되어 있습니다
* **재개**: 동일한 명령을 다시 실행하면 Ralph가 첫 번째 `passes: false` story부터 자동으로 계속합니다
```bash
# 중단 후 재개, 동일한 명령을 다시 실행하기만 하면 됨
./scripts/ralph/ralph.sh --tool claude 35
```
특정 story가 반복적으로 실패하여 차단되는 경우, 수동으로 건너뛸 수 있습니다. `prd.json`을 편집하여 해당 story의 `passes` 필드를 `true`로 변경한 다음 다시 실행하면 Ralph가 이를 건너뛰고 후속 story를 계속 처리합니다.
### 자동 아카이브
새로운 `branchName`으로 다른 기능을 시작하면, Ralph가 이전 실행의 파일을 `archive/YYYY-MM-DD-feature-name/` 디렉터리에 자동으로 아카이브하여 작업 디렉터리를 깔끔하게 유지합니다.
***
## 피드백 루프와 품질 게이트
Ralph의 "자기 교정" 능력은 전적으로 피드백 루프의 품질에 달려 있습니다. 피드백 루프가 없는 Ralph는 맹목적으로 루프하는 스크립트에 불과합니다 — 계속 코드를 생성하지만 코드가 올바른지 판단할 수 없습니다.
### 품질 검사 구성
CLAUDE.md(또는 prompt.md)에서 품질 검사 명령을 정의합니다:
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### 품질 게이트의 계층
| 계층 | 도구 | 포착하는 문제 |
| ------ | ------------------- | ---------------------- |
| 즉각 피드백 | TypeScript compiler | 타입 오류, 구문 오류 |
| 기능 검증 | 단위 테스트 | 로직 오류, 경계 조건 |
| 통합 검증 | Build 명령 | 의존성 문제, 구성 오류 |
| 런타임 검증 | dev-browser skill | UI 렌더링 문제 (프론트엔드 프로젝트) |
> 프론트엔드 story의 경우, Ralph는 인수 기준에 "Verify in browser using dev-browser skill"을 추가할 것을 권장합니다 — AI가 실제로 브라우저를 열어 페이지 렌더링이 올바른지 확인하도록 합니다.
### 품질 검사 실패 시
특정 story의 품질 검사가 반복적으로 실패해도 Ralph는 같은 story를 무한히 재시도하지 않습니다. 반복 상한에 도달하면 멈추고 현재 상태를 남깁니다. 이 경우 다음과 같이 대처할 수 있습니다:
1. progress.txt를 확인하여 AI가 어디에서 막혔는지 확인
2. 수동으로 문제를 수정한 후 다시 실행
3. story의 단위 조정 (너무 클 수 있음)
4. notes에 더 많은 컨텍스트 보충
### Prompt 커스터마이징
Ralph의 prompt 템플릿(CLAUDE.md 또는 prompt.md)은 AI 동작을 제어하는 주요 수단입니다. 설치 후 자신의 프로젝트에 맞게 커스터마이징해야 합니다. 주요 커스터마이징 방향:
**코드 스타일 제약**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**흔한 함정**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**막혔을 때의 처리 방법**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## 실전 사례: Ralph로 블로그 i18n 번역 완료하기
Ralph가 실제 프로젝트에서 어떻게 작동하는지 보여드리기 위해, 여기서 실제 사례를 공유합니다: Ralph 스타일의 자율 agent를 사용하여 전체 블로그를 중국어에서 영어로 번역한 사례입니다.
### 프로젝트 설정
프로젝트에서는 22개 이상의 콘텐츠 파일(블로그 게시글, 문서, 네비게이션 메타데이터)을 중국어에서 영어로 번역해야 했으며, 목표는 fumadocs 기반의 Next.js 블로그 i18n 지원이었습니다. 작업은 `prd.json` 파일에 정의되어 있으며, 16개의 user story가 포함되고 각각 명확한 인수 기준이 있었습니다:
```
scripts/ralph/
├── prd.json # 16개의 user story, 인수 기준 포함
└── progress.txt # 경험 로그, 각 story 완료 후 업데이트
```
각 user story는 일관된 패턴을 따릅니다:
* **명확한 산출물**: "Create content/blog/xxx.en.mdx"
* **검증 가능한 기준**: "Typecheck passes", "Internal links use /en/ prefix"
* **기술적 제약**: "Keep code blocks untranslated", "Set defaultLang='en' on QuoteCard"
### 실행 패턴
Agent는 Ralph 방법론의 핵심 원칙을 따랐습니다:
1. **파일이 진실의 원천**: `prd.json`이 각 story의 상태(`passes: true/false`)를 추적합니다. `progress.txt`가 반복 간에 경험을 축적합니다 — 예를 들어 "Typecheck 명령은 `pnpm types:check`이며 `pnpm typecheck`가 아닙니다"
2. **자동화된 품질 게이트**: 매번 번역 완료 후 `pnpm types:check`를 실행하여 MDX 파일이 올바르게 컴파일되는지 검증합니다. typecheck가 실패하면 먼저 문제를 수정한 후 커밋합니다.
3. **점진적 전진**: 각 story를 독립적으로 커밋하고 설명적인 커밋 메시지를 사용합니다(`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`). 필요 시 롤백이 용이합니다.
4. **병렬 실행**: 긴 글의 경우 여러 subagent가 동시에 번역합니다 — 예를 들어 US-010(claude-skills concept + practice), US-011(speckit concept + practice), US-012(claude-architecture + claude-subagent)가 동시에 병렬 실행됩니다.
### 핵심 경험
| 경험 | 상세 내용 |
| ------------------------ | ------------------------------------------------------------------------------ |
| **지식 축적이 중요함** | 초기 story에서 발견한 패턴(QuoteCard의 `defaultLang`, 링크 접두사 규칙)이 후속 story를 더 빠르게 완료하게 함 |
| **피드백 루프로서의 Typecheck** | 문제가 쌓이기 전에 누락된 import나 잘못된 형식의 MDX를 포착 |
| **병렬화는 확장 가능** | 6개의 번역 agent가 동시에 실행되어도 완료 시간은 1개와 거의 동일 |
| **PRD 단위가 매우 중요** | 각 story를 1\~2개 파일로 제한 — 확실히 완료할 수 있을 만큼 작고, 의미 있을 만큼 큼 |
| **진행 로그가 같은 실수의 반복을 방지** | `progress.txt`의 "Codebase Patterns" 섹션이 지식 베이스가 되어 같은 함정에 다시 빠지는 것을 방지 |
### 성과
16개의 user story가 단일 세션에서 모두 완료되었습니다: 8개의 meta.en.json 네비게이션 파일 생성, 3개의 블로그 게시글 번역, 12개의 문서 페이지 번역, 전체 사이트 빌드 검증 통과. 인수 기준이 명확하고 피드백 루프(typecheck)가 문제를 즉시 발견할 수 있었기 때문에 각 번역은 일관된 품질을 유지했습니다.
이 프로젝트는 Ralph의 \*\*전체 구현 모드(Full Implementation Mode)\*\*를 보여줍니다 — 명확한 작업 정의, 명확한 성공 기준, 자동화된 검증, 그리고 파일 시스템을 통한 점진적 납품입니다.
***
## 모범 사례 및 자주 묻는 질문
### 비용 관리
Ralph의 자동 실행은 API 비용이 지속적으로 발생한다는 것을 의미합니다. 몇 가지 제어 수단이 있습니다:
* **항상 `max_iterations`를 설정하십시오**: 이것이 가장 기본적인 안전망입니다
* **Story 단위를 합리적으로 유지**: 너무 큰 story는 여러 번의 반복을 소비하고, 너무 잘게 나눈 story는 시작 오버헤드를 증가시킵니다
* **먼저 소규모로 테스트**: 새 프로젝트는 먼저 3\~5회 반복으로 시험 실행하여 prompt와 품질 게이트가 정상 작동하는지 확인한 후 본격적으로 실행하십시오
### 흔한 함정
**함정 1: Story가 너무 큼**
증상: 하나의 story가 반복적으로 실패하고 반복 횟수가 빠르게 소진됩니다.
해결: 2\~3개의 더 작은 story로 분할합니다. "완전한 인증 시스템 구축"을 "로그인 API 구현" + "로그인 페이지 생성" + "JWT 미들웨어 추가"로 분할합니다.
**함정 2: 피드백 루프 없음**
증상: Ralph가 story 완료를 선언하지만 실제 코드에 문제가 있습니다.
해결: 인수 기준에 실행 가능한 검사 명령을 포함합니다. "코드가 작성됨"은 인수 기준이 아니며, "pnpm test가 모두 통과"가 인수 기준입니다.
**함정 3: progress.txt가 활용되지 않음**
증상: 같은 오류가 다른 반복에서 반복적으로 발생합니다.
해결: prompt 템플릿에 "progress.txt를 읽고 그 안의 경험을 따르라"는 명확한 지시가 있는지 확인합니다. AI가 자동으로 학습 내용을 추가하지 않으면 prompt에 "After each story, append learnings to progress.txt"를 추가합니다.
**함정 4: 의존 관계 순서 오류**
증상: 특정 story가 의존하는 코드가 아직 존재하지 않아 구현이 실패합니다.
해결: `dependsOn` 필드를 올바르게 설정하여 인프라 story가 먼저 오도록 합니다.
### 자주 묻는 질문
**Q: 실행 중에 prd.json을 수동으로 수정하여 개입할 수 있습니까?**
가능합니다. Ralph는 매 반복 시작 시 prd.json을 다시 읽습니다. 반복 사이에 story 설명을 수정하거나, 새로운 story를 추가하거나, 특정 story를 수동으로 `passes: true`로 표시(건너뛰기)할 수 있습니다.
**Q: Ralph가 하나의 story에서 반복적으로 실패하면 어떻게 해야 합니까?**
1. progress.txt에서 실패 원인을 확인합니다
2. notes에 추가 컨텍스트를 보충합니다
3. story를 분할합니다 (단위가 너무 클 수 있음)
4. 차단 문제를 수동으로 수정한 후 다시 실행합니다
**Q: Ralph 실행 중에 다른 일을 해도 괜찮습니까?**
괜찮습니다. Ralph는 "Human on the Loop"으로 설계되었습니다 — 계속 지켜볼 필요가 없습니다. AFK 모드에서 퇴근 전에 시작하고 다음 날 결과를 확인하면 됩니다. 실행 중에는 Ralph가 작업 중인 파일을 수정하지 않도록 주의하십시오.
**Q: 비용은 어떻게 관리합니까?**
세 가지 방법이 있습니다: 합리적인 `max_iterations` 설정, 적절한 story 단위 유지(불필요한 반복 감소), 그리고 먼저 소규모로 시험 실행하여 프로세스가 올바른지 확인하는 것입니다. 일반적으로 10~~20개의 story가 있는 프로젝트는 $50~~100의 API 비용 범위에 해당합니다.
***
## 요약
Ralph의 사용 흐름은 다섯 단계로 요약할 수 있습니다:
```
설치 → PRD 작성 → 품질 게이트 구성 → 루프 실행 → 성과 확인
```
핵심 사고방식은 항상 변하지 않습니다: **파일을 진실의 원천으로 삼고, 매번 반복을 완전히 새로운 시작으로 만들고, 품질 게이트가 대신 검증하도록 합니다**.
이제 여러분의 프로젝트로 돌아가 prd.json을 준비하고, `./scripts/ralph/ralph.sh --tool claude`를 실행한 다음, 커피 한 잔 마시러 가십시오.
### 추가 읽을거리
* 《[Ralph Wiggum 심층 분석](/ko/docs/notes/ralph-wiggum/concept)》 — Ralph의 핵심 원리 복습
* 《[frankbria/ralph-claude-code 실전 가이드](/ko/docs/notes/ralph-wiggum/frankbria)》 — 엔지니어링화된 Ralph 구현: 모니터링, 서킷 브레이커 및 안전 메커니즘
* 《[GSD 심층 분석](/ko/docs/notes/gsd/concept)》 — Ralph 위에 구축된 완전한 컨텍스트 엔지니어링 시스템
* 《[Claude Skills란 무엇인가](/ko/docs/notes/claude-skills/concept)》 — Ralph의 PRD skill은 하나의 Claude Skill입니다
* 《[Speckit 실전 가이드](/ko/docs/notes/speckit/practice)》 — 또 다른 구조화된 AI 프로그래밍 워크플로
# 개념 소개
## 서론
2025년 10월, GitHub는 Spec Kit이라는 툴킷을 오픈소스로 공개하며 「규격 주도 개발」(Spec-Driven Development)이라는 개념을 AI 프로그래밍의 영역으로 공식적으로 끌어왔습니다. 복고적으로 보일 수 있는 이 이념 -- 규격을 먼저 작성하고 코드를 나중에 작성하는 것 -- 이 AI 프로그래밍 도구를 다루는 새로운 패러다임이 되고 있습니다.
Claude Code, Cursor 또는 GitHub Copilot과 같은 AI 프로그래밍 어시스턴트를 자주 사용하신다면, 분명 이런 어려움을 겪어보셨을 것입니다: 「사용자 로그인 기능을 추가해줘」라고 말하면, AI가 의욕적으로 많은 코드를 작성하지만, 자세히 살펴보면 -- 익숙하지 않은 프레임워크를 사용하고, 보안 전략이 기대와 다르고, UI 스타일도 맞지 않습니다……그리고 수정을 반복하다가 결국 지치게 됩니다.
문제가 어디에 있을까요? AI가 충분히 똑똑하지 않은 것이 아니라, 여러분이 제공한 정보가 충분하지 않은 것입니다. 「사용자 로그인 기능 추가」라는 말은 명확해 보이지만, 실제로는 수백 가지의 명시되지 않은 결정을 숨기고 있습니다: 어떤 인증 방식을 사용할 것인가? 비밀번호 요구사항은 무엇인가? 로그인 실패는 어떻게 처리할 것인가? 로그인 상태를 기억해야 하는가? 제3자 로그인을 지원할 것인가?……AI는 추측할 수밖에 없고, 추측은 곧 편차를 의미합니다.
규격 주도 개발은 바로 이 문제를 해결하기 위해 탄생했습니다.
## Vibe Coding: 속도의 대가
2025년 초, 전 Tesla AI 디렉터 Andrej Karpathy가 「Vibe Coding」이라는 용어를 만들어, 「AI의 제안을 심층 검토 없이 수용하는」 개발 방식을 설명했습니다. 이 용어는 빠르게 유행하며 Collins 사전의 2025년 올해의 단어로 선정되기까지 했습니다.
Vibe Coding의 유혹은 분명합니다: 아이디어를 설명하면 AI가 코드를 생성하고, 실행되는 것처럼 보이면 됩니다. 빠른 프로토타입, 해커톤, 일회성 스크립트에는 이 방식이 확실히 효율적입니다. 하지만 프로덕션 시스템에 사용되면 문제가 발생합니다.
이와 유사한 사례는 업계에서 흔히 볼 수 있습니다: AI가 생성한 데이터베이스 쿼리가 소규모 테스트에서는 잘 작동했지만 실제 트래픽에서는 시스템이 느려졌고, 조합된 인증 모듈이 QA를 통과했지만 2주 후 비활성화된 계정이 여전히 관리 도구에 접근할 수 있다는 것이 발견되었습니다. Final Round AI의 2025년 조사에 따르면, **18명의 CTO 중 16명이 AI 생성 코드로 인한 프로덕션 장애를 경험했습니다**.
이것은 Vibe Coding이 전혀 가치가 없다는 의미가 아닙니다. 핵심은 **경계를 식별하는 것**입니다:
| 시나리오 | Vibe Coding | 규격 주도 개발 |
| -------- | ----------- | -------- |
| 프로토타입/데모 | ✓ 적합 | 과도함 |
| 일회성 스크립트 | ✓ 적합 | 과도함 |
| 프로덕션 기능 | 위험 높음 | ✓ 권장 |
| 보안 관련 | 위험 | ✓ 필수 |
| 팀 협업 | 유지보수 어려움 | ✓ 권장 |
규격 주도 개발은 AI의 효율성을 유지하면서 Vibe Coding의 함정을 피하기 위한 것입니다.
## 규격 주도 개발이란 무엇인가
규격 주도 개발의 핵심 이념은 한 문장으로 요약할 수 있습니다: **「무엇을 할 것인가」를 먼저 정의하고, 「어떻게 할 것인가」를 나중에 고려합니다**.
이것은 소프트웨어 엔지니어링에서 익숙한 이야기처럼 들리지만, AI 프로그래밍 시대에는 새로운 의미를 갖게 되었습니다. 전통적인 요구사항 문서는 사람을 위해 작성되었기 때문에 장황하고 모호하며 전문 용어로 가득 차 있는 경우가 많습니다. 규격 주도 개발에서의 「규격」은 AI를 위해 작성됩니다 -- 간결하고, 구조화되어 있으며, 실행 가능합니다.
집을 짓는다고 상상해 보십시오. 전통적인 AI 프로그래밍 방식은 시공팀에게 「편안한 3베드룸 집을 지어주세요」라고 말하고 알아서 하도록 맡기는 것과 같습니다. 결과가 괜찮을 수도 있지만, 여러분이 상상했던 것과 크게 다를 가능성이 더 높습니다. 규격 주도 개발은 먼저 건축 설계도를 그리는 것입니다: 몇 층인지, 각 층의 면적은 얼마인지, 창문 방향, 자재 규격……시공팀이 설계도대로 시공하면 결과는 자연히 기대에 부합합니다.
AI 프로그래밍에서 이 설계도가 바로 **규격 문서**(Specification)입니다. 어떤 프로그래밍 언어를 사용하는지, 어떤 프레임워크를 사용하는지는 관심사가 아니며, 기능이 어떤 효과를 달성해야 하는지, 사용자가 어떤 작업을 완료해야 하는지, 성공의 기준이 무엇인지만 다룹니다.
전통적인 개발 프로세스와 비교했을 때, 규격 주도 개발에는 근본적인 차이가 있습니다:
| 전통적 AI 프로그래밍 | 규격 주도 개발 |
| ---------------------- | ------------------------- |
| 요구사항 직접 설명 → AI가 코드 생성 | 요구사항 → 규격 → 계획 → 작업 → 코드 |
| AI가 많은 세부사항을 추측해야 함 | 모든 단계가 명확하며, AI는 실행만 하면 됨 |
| 재작업 빈번, 커뮤니케이션 비용 높음 | 초기 투자, 이후 순조로움 |
| 간단한 작업에 적합 | 복잡한 기능에 적합 |
이러한 「점진적 구체화」 과정이 바로 규격 주도 개발의 핵심입니다. 한 번에 완성하는 것이 아니라, 여러 단계를 거쳐 요구사항을 점진적으로 명확히 하며, 각 단계에서 검토하고 조정할 수 있습니다.
## Speckit 워크플로우 개요
GitHub의 Spec Kit과 Claude Code의 speckit 명령은 모두 유사한 워크플로우를 따르며, 대략 6단계로 나눌 수 있습니다:
```
Constitution → Specify → Clarify → Plan → Tasks → Implement
↓ ↓ ↓ ↓ ↓ ↓
프로젝트 헌법 기능 규격 모호점 명확화 기술 계획 작업 분해 실행 구현
```
**1. Constitution(프로젝트 헌법)**
프로젝트 헌법은 프로젝트 전체의 기본 원칙과 제약 조건을 정의합니다. 예를 들어 「테스트 우선」「단순함 우선」「API 우선」 등이 있습니다. 이러한 원칙은 이후 모든 단계에 적용되어, AI가 생성하는 방안이 여러분의 기술적 선호도에 부합하도록 보장합니다.
**2. Specify(기능 규격)**
이것은 핵심적인 첫 번째 단계입니다. 자연어로 원하는 기능을 설명하면, AI가 이를 구조화된 규격 문서로 정리해 줍니다. 여기에는 다음이 포함됩니다:
* 사용자 스토리: 누가 무엇을 왜 하는가
* 기능 요구사항: 시스템이 반드시 갖추어야 할 능력
* 성공 기준: 기능의 달성 여부를 어떻게 판단하는가
중요한 것은, 규격 문서는 **「무엇을 할 것인가」에만 집중하며 「어떻게 할 것인가」는 다루지 않습니다** -- 구체적인 기술 스택을 언급하지 않고, 코드 구조를 작성하지 않습니다.
**3. Clarify(모호점 명확화)**
AI가 규격의 모호한 부분을 검토하고, 최대 5개의 핵심 질문을 제시합니다. 이러한 질문은 보통 기능 범위, 사용자 유형, 보안 요구사항 등에 관한 것입니다. 질의응답을 통해 규격이 더욱 명확해집니다.
**4. Plan(기술 계획)**
명확한 규격이 완성된 후에야 기술 방안을 고려하기 시작합니다. 이 단계에서는 다음이 산출됩니다:
* 기술 선정(언어, 프레임워크, 데이터베이스)
* 데이터 모델 설계
* API 계약 정의
* 연구 보고서(기술 결정 해결)
**5. Tasks(작업 분해)**
기술 계획을 실행 가능한 작업 목록으로 분해합니다. 각 작업에는 명확한 ID, 설명, 파일 경로가 있어 AI에게 직접 맡겨 실행할 수 있습니다. 작업은 사용자 스토리별로 그룹화되며, 병렬 개발을 지원합니다.
**6. Implement(실행 구현)**
작업 목록에 따라 하나씩 실행합니다. 각 작업이 완료되면 완료로 표시하여 추적 가능성을 보장합니다.
이 6단계의 산출물은 명확한 체인을 형성합니다:
| 단계 | 산출물 | 역할 |
| ------------ | -------------------- | ------------ |
| Constitution | constitution.md | 프로젝트 원칙 정의 |
| Specify | spec.md | 기능 요구사항 기술 |
| Clarify | 업데이트된 spec.md | 모호점 제거 |
| Plan | plan.md, research.md | 기술 방안 설계 |
| Tasks | tasks.md | 실행 가능한 작업 목록 |
| Implement | 실제 코드 | 최종 산출물 |
## 이 방식이 효과적인 이유
규격 주도 개발이 「AI에게 직접 코드를 작성하게 하는 것」이 실패하는 곳에서 성공할 수 있는 이유는, AI 프로그래밍의 핵심 모순인 **정보 비대칭**을 해결하기 때문입니다.
「사진 공유 기능 추가」라고 말할 때, 여러분의 머릿속에는 완전한 그림이 있을 수 있지만, AI는 이 몇 글자만 볼 수 있습니다. AI는 추측해야 합니다: 어디에 공유하는가? 누가 볼 수 있는가? 압축이 필요한가? 워터마크가 있는가? 일괄 처리를 지원하는가?……모든 추측이 틀릴 수 있습니다.
규격 주도 개발은 「먼저 생각을 정리하도록 강제하는 것」을 통해 이 문제를 해결합니다. 사용자 스토리, 기능 요구사항, 성공 기준을 작성하도록 요구받으면, 「당연하다고 생각했던」 세부사항들이 수면 위로 떠오릅니다. 이 과정 자체가 가치가 있습니다 -- AI를 사용하지 않더라도 요구사항을 명확히 작성하면 커뮤니케이션 비용을 줄일 수 있습니다.
또한 점진적 구체화는 오류를 더 일찍 노출시킵니다. Specify 단계에서 요구사항 편차를 발견하면 수정 비용이 거의 제로입니다. 코드가 완성된 후에야 발견하면 처음부터 다시 해야 할 수도 있습니다.
물론 규격 주도 개발이 만능은 아닙니다. 명확한 적용 시나리오가 있습니다:
**적합한 경우**:
* 복잡한 기능 개발(여러 모듈, 다양한 인터랙션 포함)
* 팀 협업 프로젝트(규격 문서가 커뮤니케이션 매개체 역할)
* 높은 품질이 요구되는 시나리오(추적 가능성, 검증 가능성 필요)
**적합하지 않은 경우**:
* 간단한 버그 수정이나 작은 변경
* 탐색적 프로그래밍(무엇을 할지 아직 모를 때)
* 시간이 극도로 촉박한 경우(규격을 작성할 시간이 없을 때)
핵심은 작업의 복잡도를 식별하는 것입니다. 한 시간 안에 완료할 수 있는 일에 한 시간을 들여 규격을 작성할 필요는 없습니다. 하지만 일주일짜리 기능 개발에 두 시간을 들여 규격을 작성하는 것은 충분히 가치가 있습니다.
## 그러나 규격은 만능이 아닙니다
흔한 오해를 바로잡을 필요가 있습니다: **규격 주도 개발은 추측을 줄이지만, 검토의 필요성을 없애지는 않습니다**.
완전한 규격이 있더라도, AI는 여전히 다음과 같은 실수를 할 수 있습니다:
* 경계 케이스 누락(규격이 다루지 않은 극단적 시나리오)
* 성능 요구사항을 충족하지 못하는 코드 생성
* 잠재적 보안 취약점 도입
* 일관성 없는 스타일의 구현 산출
이것은 건축 시공과 같습니다: 상세한 설계도가 있더라도 준공 검사는 여전히 필요합니다. 시공팀이 도면대로 집을 지었다고 해서 바로 입주하지는 않을 것입니다 -- 전기 회로가 안전한지, 배관이 원활한지, 문과 창문이 견고한지 확인할 것입니다.
규격 주도 개발의 가치는 **오류를 더 쉽게 발견할 수 있게 하는 것**이지, 오류 자체를 없애는 것이 아닙니다.
## 요약
규격 주도 개발의 핵심은 간단한 이치입니다: **작업이 복잡할수록, 실행 전에 먼저 생각을 정리하는 것이 더 중요합니다**. AI 프로그래밍 도구는 이 이치의 중요성을 증폭시킵니다 -- AI는 여러분의 지시를 충실히 실행하지만, 여러분의 의도를 진정으로 이해할 수는 없기 때문입니다.
세 가지 핵심 키워드를 기억하십시오:
| 키워드 | 의미 |
| ----------------- | -------------------------------------------- |
| **규격 먼저, 코드 나중에** | 「무엇을 할 것인가」를 먼저 정의하고, 「어떻게 할 것인가」를 나중에 고려합니다 |
| **점진적 구체화** | 모호함에서 명확함으로, 매 단계마다 검토와 조정이 가능합니다 |
| **추측 감소** | 명확한 규격 = AI의 추측 여지 감소 |
이념을 이해하셨다면, 다음 글 《[Speckit 실천 가이드](/ko/docs/notes/speckit/practice)》에서 직접 실습해 보실 수 있습니다: speckit 명령을 사용하여 하나의 기능에 대한 규격 주도 개발 프로세스를 완성하는 방법을 안내합니다.
《[Claude Skills](/ko/docs/notes/claude-skills/concept)》와 함께 사용하면, 규격의 실행을 더욱 자동화하고 표준화할 수 있습니다.
# 실전 가이드
## 서론
[이전 글](/ko/docs/notes/speckit/concept)에서 스펙 주도 개발의 이념을 알아보았습니다 — 먼저 '무엇을 할 것인지'를 정의하고, 그 다음에 '어떻게 할 것인지'를 고려하는 방식입니다. 얼핏 불필요해 보이는 이 프로세스는 실제로 AI 프로그래밍에서 재작업과 커뮤니케이션 비용을 크게 줄여줍니다.
이번 글에서는 직접 실습해 보겠습니다. speckit 명령어 시리즈를 사용하여 요구사항 기술에서 코드 구현까지의 전체 프로세스를 완성하는 방법을 배우게 됩니다.
## 설치 및 구성
Speckit 명령어는 GitHub 공식 [Spec Kit](https://github.com/github/spec-kit) 프로젝트에서 제공됩니다. 사용 시나리오에 따라 몇 가지 다른 통합 방식이 있습니다.
### 새 프로젝트 초기화
새 프로젝트의 경우, 공식 specify-cli 도구를 사용하여 초기화하는 것을 권장합니다:
```bash
# uv로 specify-cli 설치
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git
# 새 프로젝트 초기화, Claude를 AI 어시스턴트로 지정
specify init my-project --ai claude
```
이 명령은 `.specify/` 구성 디렉토리와 관련 템플릿 파일을 포함한 프로젝트 디렉토리 구조를 자동으로 생성합니다.
### 기존 프로젝트 통합
Speckit 명령어를 사용하려면 구성 파일이 필요합니다. 기존 프로젝트에 speckit을 통합할 때는 specify-cli를 사용하는 것을 권장합니다:
```bash
cd your-existing-project
specify init . --ai claude # .은 현재 디렉토리를 의미합니다
```
이 명령은 프로젝트에 다음을 생성합니다:
```
your-project/
├── .specify/
│ ├── templates/ # 스펙, 계획 등의 템플릿
│ ├── scripts/ # 보조 스크립트
│ └── memory/ # constitution.md
├── .claude/
│ └── commands/ # Claude Code 명령어 구성
│ ├── speckit.specify.md
│ ├── speckit.plan.md
│ └── ...
└── specs/ # 기능 스펙 저장 디렉토리
```
초기화 시 기존 파일을 덮어쓰지 않습니다. 완료 후 Claude Code에서 `/speckit.*` 시리즈 명령어를 사용할 수 있습니다.
> **주의**: speckit 명령어는 Claude Code에 내장된 것이 아니므로, 반드시 위의 초기화 단계를 먼저 완료해야 합니다. 초기화 없이 `/speckit.specify`를 직접 실행하면 명령어가 존재하지 않는다는 메시지가 표시됩니다.
***
## 명령어 상세 설명
Speckit은 스펙 주도 개발의 각 단계를 지원하는 일련의 명령어를 제공합니다. 각 명령어는 명확한 입력과 출력을 가지며, 추적 가능한 체인을 형성합니다.
### /speckit.specify — 기능 스펙 생성
이것은 전체 프로세스의 시작점입니다. 자연어로 원하는 기능을 설명하면, AI가 구조화된 스펙 문서로 정리해 줍니다.
**기능**: 자연어 설명에서 기능 스펙 문서 생성
**입력**: 기능 설명 (자연어)
**출력**:
* `specs/[번호]-[기능명]/spec.md` — 기능 스펙 문서
* 새 git 브랜치 (예: `001-user-auth`)
**사용 예시**:
```
/speckit.specify 이메일 비밀번호 로그인을 지원하는 사용자 로그인 기능을 추가하고 싶습니다. 로그인 상태 유지 옵션이 필요합니다
```
실행 후 AI는:
1. 간단한 기능 이름을 생성합니다 (예: `user-auth`)
2. 새 기능 브랜치를 생성합니다
3. 사용자 스토리, 기능 요구사항, 성공 기준이 포함된 스펙 문서를 생성합니다
4. 불명확한 부분에 `[NEEDS CLARIFICATION]`을 표시합니다
**스펙 문서의 핵심 구조**:
```markdown
# Feature Specification: 사용자 로그인 기능
## User Scenarios & Testing
### User Story 1 - 사용자 로그인 (Priority: P1)
사용자가 이메일과 비밀번호로 시스템에 로그인합니다...
**Acceptance Scenarios**:
1. Given 사용자가 올바른 이메일과 비밀번호를 입력, When 로그인 클릭, Then 시스템에 성공적으로 진입
## Requirements
### Functional Requirements
- FR-001: 시스템은 이메일 비밀번호 로그인을 지원해야 합니다
- FR-002: 시스템은 "로그인 상태 유지" 옵션을 제공해야 합니다
## Success Criteria
- SC-001: 사용자가 30초 이내에 로그인 프로세스를 완료할 수 있어야 합니다
```
스펙 문서는 **어떤 기술적 세부사항도 포함하지 않습니다** — 어떤 프레임워크를 사용할지, 데이터베이스 구조를 어떻게 할지, API를 어떻게 정의할지 언급하지 않습니다. 이런 것들은 이후 단계에서 다룹니다.
***
### /speckit.clarify — 모호한 부분 명확화
스펙 문서를 작성한 후에도 모호한 부분이 있을 수 있습니다. 이 명령어는 스펙을 검토하고, 핵심 질문을 통해 명확화를 도와줍니다.
**기능**: 스펙의 모호한 부분을 식별하고, 질의응답을 통해 스펙을 보완
**입력**: 기존 spec.md 문서
**출력**: 업데이트된 spec.md (명확화 기록 포함)
**사용 예시**:
```
/speckit.clarify
```
실행 후 AI는:
1. 스펙 문서의 모호한 부분을 스캔합니다
2. 우선순위별로 정렬합니다 (범위 > 보안 > 사용자 경험 > 기술 세부사항)
3. 하나씩 질문하며, 한 번에 하나의 질문만 합니다
4. 답변에 따라 스펙 문서를 업데이트합니다
**질의응답 예시**:
```markdown
## Question 1: 로그인 실패 처리
**Context**: 스펙에서 사용자 로그인을 언급했지만, 로그인 실패 시 처리 방식이 명시되어 있지 않습니다.
**Recommended:** Option B - 연속 5회 실패 후 계정 잠금은 보안 모범 사례입니다
| Option | Description |
|--------|-------------|
| A | 오류 메시지만 표시하고 제한 없음 |
| B | 연속 5회 실패 후 15분간 계정 잠금 |
| C | 캡차를 사용하여 무차별 대입 공격 방지 |
옵션 문자로 답변하거나(예: "B"), "yes"로 권장 사항을 수락하거나, 직접 답변을 제공할 수 있습니다.
```
명확화할 때마다 스펙 문서가 자동으로 업데이트되고, 명확화 기록이 추가됩니다:
```markdown
## Clarifications
### Session 2025-12-20
- Q: 로그인 실패 시 어떻게 처리하나요? → A: 연속 5회 실패 후 15분간 계정 잠금
```
***
### /speckit.plan — 기술 계획 생성
스펙이 명확해지면 기술 설계 단계로 진입합니다. 이 단계에서는 기술 계획과 조사 보고서가 산출됩니다.
**기능**: 스펙을 기반으로 기술 구현 계획 생성
**입력**: spec.md 문서
**출력**:
* `plan.md` — 기술 계획 (아키텍처, 데이터 모델, API 설계)
* `research.md` — 조사 보고서 (기술 선택 결정)
* `data-model.md` — 데이터 모델 (해당되는 경우)
* `contracts/` — API 계약 (해당되는 경우)
**사용 예시**:
```
/speckit.plan 저는 Next.js + Prisma + PostgreSQL을 사용합니다
```
명령어 뒤에 기술 스택 선호도를 추가할 수 있습니다. 실행 후 AI는:
1. 스펙의 기능 요구사항을 분석합니다
2. 관련 기술의 모범 사례를 조사합니다
3. 데이터 모델과 API 구조를 설계합니다
4. 완전한 기술 계획을 생성합니다
**기술 계획의 핵심 내용**:
```markdown
# Implementation Plan: 사용자 로그인 기능
## Technical Context
**Language/Version**: TypeScript 5.x
**Primary Dependencies**: Next.js 15, Prisma, PostgreSQL
**Authentication**: NextAuth.js with credentials provider
## Project Structure
src/
├── app/
│ └── (auth)/
│ ├── login/
│ └── api/auth/
├── lib/
│ └── auth/
└── prisma/
└── schema.prisma
## Data Model
- User: id, email, passwordHash, createdAt, updatedAt
- Session: id, userId, expiresAt
```
***
### /speckit.tasks — 작업 분해
기술 계획이 완성되면, 이를 실행 가능한 작업 목록으로 분해합니다.
**기능**: 기술 계획을 실행 가능한 작업 목록으로 분할
**입력**: plan.md 문서
**출력**: `tasks.md` — 의존성 순서로 정렬된 작업 목록
**사용 예시**:
```
/speckit.tasks
```
실행 후 AI는:
1. plan.md에서 기술 방안을 추출합니다
2. spec.md에서 사용자 스토리 우선순위를 추출합니다
3. 사용자 스토리별로 그룹화하여 작업을 생성합니다
4. 병렬 실행 가능한 작업에 `[P]`를 표시합니다
5. 각 작업에 구체적인 파일 경로를 지정합니다
**작업 목록 형식**:
```markdown
## Phase 1: Setup
- [ ] T001 프로젝트 구조 생성
- [ ] T002 [P] Prisma schema 구성
- [ ] T003 [P] NextAuth 구성
## Phase 2: User Story 1 - 사용자 로그인 (P1)
- [ ] T004 [US1] User 모델 생성 in prisma/schema.prisma
- [ ] T005 [US1] 로그인 API 구현 in src/app/api/auth/[...nextauth]/route.ts
- [ ] T006 [US1] 로그인 페이지 생성 in src/app/(auth)/login/page.tsx
```
각 작업에는 다음이 포함됩니다:
* **작업 ID** (T001, T002...) — 추적용
* **\[P] 표시** — 다른 \[P] 작업과 병렬 실행 가능함을 표시
* **\[US] 태그** — 어떤 사용자 스토리에 속하는지 표시
* **파일 경로** — 어떤 파일을 조작할지 명확히 지정
***
### /speckit.implement — 구현 실행
모든 준비가 완료되면 작업 목록 실행을 시작합니다.
**기능**: 작업 목록에 따라 순차적으로 구현 실행
**입력**: tasks.md 문서
**출력**: 실제 코드
**사용 예시**:
```
/speckit.implement
```
실행 전에 AI는 체크리스트(있는 경우)를 확인합니다. 실행 시:
1. 단계 순서대로 작업을 실행합니다
2. 각 작업 완료 시 `[X]`로 표시합니다
3. 작업 의존성 관계를 준수합니다
4. 병렬 작업은 동시에 진행할 수 있습니다
**실행 과정 예시**:
```
Phase 1: Setup
✓ T001 프로젝트 구조 생성
✓ T002 Prisma schema 구성
✓ T003 NextAuth 구성
Phase 2: User Story 1
✓ T004 User 모델 생성
T005 실행 중...
```
### 구현 후 검토
`/speckit.implement` 완료 후, **코드를 바로 병합하지 마십시오**. AI가 생성한 코드는 사람이 검토해야 합니다:
**반드시 수행해야 하는 검증 단계**:
1. **테스트 스위트 실행**
```bash
npm test # 또는 사용 중인 테스트 명령어
```
AI가 기존 기능을 손상시키지 않았는지 확인합니다.
2. **코드 리뷰 포인트**
* 스펙의 의도에 부합하는지 (spec.md 대조)
* 프로젝트의 코드 스타일을 따르는지
* 잠재적 보안 문제가 있는지
3. **경계 테스트**
AI가 놓칠 수 있는 경계 케이스를 수동으로 테스트합니다:
* 빈 값 처리
* 극단적 입력
* 동시성 시나리오
* 오류 경로
4. **성능 점검**
데이터베이스 작업이나 API 호출이 관련된 경우, N+1 쿼리 등의 성능 문제가 없는지 확인합니다.
> **팁**: 스펙을 아무리 상세하게 작성하더라도, AI는 구현 세부사항에서 편차가 발생할 수 있습니다. 검토는 스펙 주도 개발을 불신하는 것이 아니라, 엔지니어링 규율의 일부입니다.
***
### /speckit.analyze — 일관성 분석
이것은 선택적인 품질 검사 단계로, 스펙, 계획, 작업 간의 일관성을 검증하는 데 사용됩니다.
**기능**: 문서 간 일관성 및 품질 분석
**입력**: spec.md, plan.md, tasks.md
**출력**: 분석 보고서 (어떤 파일도 수정하지 않음)
**사용 예시**:
```
/speckit.analyze
```
실행 후 다음을 검사합니다:
* 모든 요구사항에 대응하는 작업이 있는지
* 작업이 모든 사용자 스토리를 커버하는지
* 용어가 일관적인지
* 누락이나 중복이 있는지
***
### 기타 명령어 (선택 사항)
위의 핵심 명령어 외에도, speckit은 몇 가지 보조 명령어를 제공합니다. 이 명령어들은 주요 프로세스에 포함되지 않지만, 특정 시나리오에서 유용합니다.
**`/speckit.constitution`** — 프로젝트 헌법 생성
프로젝트의 개발 원칙과 규범을 정의하는 데 사용됩니다. 팀 프로젝트에 적합하며, 모든 구성원이 통일된 개발 표준을 따르도록 보장합니다.
* **입력**: 대화형 질의응답 또는 원칙 직접 제공
* **출력**: `.specify/constitution.md` 프로젝트 헌법 파일
* **시나리오**: 새 팀 프로젝트 초기화, 코드 스타일 및 아키텍처 결정 통일
**`/speckit.checklist`** — 품질 검사 체크리스트 생성
기능 스펙을 기반으로 맞춤형 품질 검사 체크리스트를 생성하여, 구현 전 품질 기준을 확인하는 데 사용됩니다.
* **입력**: spec.md 문서
* **출력**: `checklists/` 디렉토리에 검사 체크리스트
* **시나리오**: 중요 기능 출시 전 품질 점검, 코드 리뷰 참고
**`/speckit.taskstoissues`** — 작업을 GitHub Issues로 변환
tasks.md의 작업을 자동으로 GitHub Issues로 변환하여, 팀 협업과 작업 배정을 편리하게 합니다.
* **입력**: tasks.md 문서
* **출력**: GitHub Issues (gh CLI를 통해 생성)
* **시나리오**: 팀 협업 개발, Sprint 계획, 작업 추적
***
## 도구 생태계
이 글에서 소개한 speckit 명령어는 [GitHub Spec Kit](https://github.com/github/spec-kit) 프로젝트에서 제공됩니다. 이 외에도 2025년에 여러 주요 AI 프로그래밍 도구들이 유사한 스펙 주도 워크플로를 지원하기 시작했습니다:
| 도구 | 특징 | 적용 시나리오 |
| --------------------------------------------------------- | ------------------------------------------------------------- | -------------------- |
| **[GitHub Spec Kit](https://github.com/github/spec-kit)** | 이 글에서 사용한 도구, MIT 오픈소스, Claude Code / Copilot / Gemini CLI 지원 | 명령줄 선호자, 크로스 도구 협업 |
| **[AWS Kiro](https://kiro.dev/)** | VS Code 포크, 시각화 워크플로, EARS 표기법 | GUI 선호자, AWS 생태계 사용자 |
| **[JetBrains Junie](https://blog.jetbrains.com/junie/)** | IntelliJ 생태계 통합, Think More 추론 모드 | JetBrains IDE 사용자 |
| **Cursor Plan Mode** | 내장 계획 단계, 실행 계획 자동 생성 | 이미 Cursor를 사용 중인 개발자 |
**선택 방법**:
* Claude Code, GitHub Copilot 또는 Gemini CLI를 사용한다면 GitHub Spec Kit을 권장합니다
* 그래픽 인터페이스와 시각화 워크플로를 선호한다면 AWS Kiro를 시도해 볼 수 있습니다
* JetBrains 사용자라면 Junie가 IDE와의 통합이 더 자연스럽습니다
* 이미 Cursor를 사용 중이라면 Plan Mode가 유사한 계획 기능을 제공합니다
핵심 이념은 동일합니다 — 도구는 단지 수단일 뿐이며, 중요한 것은 **스펙을 먼저, 코드는 나중에**라는 사고방식입니다.
***
## 전체 사례 시연
실제 사례를 통해 speckit의 전체 프로세스를 진행해 보겠습니다.
**시나리오**: 블로그 시스템에 글 북마크 기능 추가
### Step 1: 스펙 생성
```
/speckit.specify 블로그에 글 북마크 기능을 추가하고 싶습니다. 사용자가 좋아하는 글을 북마크하고 개인 센터에서 북마크 목록을 볼 수 있어야 합니다
```
AI가 스펙 문서 `specs/002-article-bookmark/spec.md`를 생성합니다:
```markdown
# Feature Specification: 글 북마크 기능
## User Scenarios & Testing
### User Story 1 - 글 북마크 (Priority: P1)
사용자가 글을 읽을 때 북마크 버튼을 클릭하여 글을 북마크 목록에 추가합니다.
**Acceptance Scenarios**:
1. Given 사용자가 로그인한 상태, When 북마크 버튼 클릭, Then 글이 북마크 목록에 추가됨
2. Given 사용자가 로그인하지 않은 상태, When 북마크 버튼 클릭, Then 로그인 안내
### User Story 2 - 북마크 목록 조회 (Priority: P2)
사용자가 개인 센터에서 자신이 북마크한 모든 글을 조회합니다.
## Requirements
- FR-001: 사용자는 글을 북마크/북마크 해제할 수 있어야 합니다
- FR-002: 북마크 버튼은 현재 북마크 상태를 표시해야 합니다
- FR-003: 개인 센터에 북마크 목록이 표시되어야 합니다
## Success Criteria
- SC-001: 북마크 작업이 500ms 이내에 완료되어야 합니다
- SC-002: 북마크 목록은 페이지네이션을 지원하며 페이지당 10개의 글을 표시합니다
```
### Step 2: 요구사항 명확화
```
/speckit.clarify
```
AI가 질문합니다: 「북마크 수에 상한이 있나요?」
답변: 「최대 100개까지 북마크할 수 있습니다」
스펙이 업데이트되어 다음이 추가됩니다:
* FR-004: 각 사용자는 최대 100개의 글을 북마크할 수 있습니다
* 상한에 도달하면 안내 메시지를 표시합니다
### Step 3: 계획 생성
```
/speckit.plan Next.js + Prisma 사용
```
다음을 포함한 기술 계획이 생성됩니다:
* Bookmark 모델 (userId, articleId, createdAt)
* API 라우트 설계 (POST/DELETE /api/bookmarks)
* 컴포넌트 설계 (BookmarkButton, BookmarkList)
### Step 4: 작업 분해
```
/speckit.tasks
```
작업 목록이 생성됩니다:
```markdown
## Phase 1: Setup
- [ ] T001 Prisma schema에 Bookmark 모델 추가
## Phase 2: US1 - 글 북마크
- [ ] T002 [US1] 북마크 API 생성 in src/app/api/bookmarks/route.ts
- [ ] T003 [US1] BookmarkButton 컴포넌트 생성 in src/components/BookmarkButton.tsx
- [ ] T004 [US1] 글 페이지에 통합
## Phase 3: US2 - 북마크 목록
- [ ] T005 [US2] 북마크 목록 페이지 생성 in src/app/profile/bookmarks/page.tsx
- [ ] T006 [US2] 페이지네이션 로직 구현
```
### Step 5: 구현 실행
```
/speckit.implement
```
작업 순서에 따라 실행하며, 각 작업 완료 시 `[X]`로 표시합니다.
***
## 모범 사례 및 주의사항
### speckit을 사용해야 할 때
**적합한 시나리오**:
* 새 기능 개발 (3개 이상의 파일 관련)
* 요구사항이 완전히 명확하지 않을 때 (clarify로 명확화)
* 다인 협업 프로젝트 (스펙이 합의 기준 역할)
* 중요한 기능 (추적 가능성이 필요한 경우)
**적합하지 않은 시나리오**:
* 간단한 버그 수정
* 한 줄 코드 변경
* 긴급 핫픽스
* 순수 탐색적 실험
### 흔한 함정
speckit을 사용하는 과정에서 주의해야 할 몇 가지 흔한 함정이 있습니다:
**함정 1: 스펙이 너무 모호함**
증상: AI가 생성한 코드가 기대와 크게 다르며, 대량의 재작업이 필요합니다.
```markdown
# ❌ 모호한 스펙
사용자가 글을 검색할 수 있음
# ✓ 명확한 스펙
- FR-001: 사용자가 제목 키워드로 글을 검색할 수 있음
- FR-002: 검색 결과는 관련도순으로 정렬, 페이지당 10개 표시
- FR-003: 검색어가 결과에서 하이라이트 표시됨
- FR-004: 검색어가 비어있으면 인기 글을 표시
```
해결 방법: `/speckit.clarify`를 실행하거나, 기능 요구사항과 성공 기준을 수동으로 보충합니다.
**함정 2: 스펙이 너무 상세함**
증상: AI가 제한되어 능력을 발휘하지 못하고, 생성된 코드가 지나치게 경직되거나, 일부 지시를 무시합니다.
```markdown
# ❌ 과도하게 상세 (구현 세부사항 지정)
lodash의 debounce 함수를 사용하고, 300ms 지연,
useCallback으로 감싸고, 의존성은 [searchTerm]...
# ✓ 적절한 수준의 상세도 (무엇을 할지만 기술)
검색 입력에 디바운스를 적용하여 빈번한 요청을 방지
```
해결 방법: 스펙은 '무엇을 할지' 수준을 유지하고, '어떻게 할지'는 Plan 단계에 맡깁니다.
**함정 3: Plan 단계 건너뛰기**
증상: Tasks가 너무 대략적이거나 너무 세분화되어, 구현 시 빈번한 재작업이 발생하고, 작업 간 의존성 관계가 혼란스럽습니다.
해결 방법: 복잡한 기능은 반드시 Plan 단계를 완료합니다. Plan은 기술 방안을 산출할 뿐만 아니라 잠재적 아키텍처 문제를 식별하는 데도 도움이 됩니다.
**함정 4: 검토 없이 바로 병합**
증상: 출시 후 경계 케이스 미처리, 보안 취약점, 성능 문제가 발견됩니다.
해결 방법: 위의 '구현 후 검토' 섹션을 참고하여, 병합 전에 항상 테스트와 코드 리뷰를 수행합니다.
### 자주 묻는 질문
**Q: 모든 기능에 대해 전체 프로세스를 거쳐야 하나요?**
아닙니다. 간단한 변경은 바로 코딩할 수 있으며, 복잡한 기능은 최소한 specify + plan을 완료하는 것을 권장합니다.
**Q: 스펙을 매우 상세하게 작성했는데도 AI가 기대에 맞지 않는 코드를 생성했습니다.**
스펙이 정말 '상세'한지 확인하십시오. 많은 경우 명확하게 설명했다고 생각하지만 실제로는 모호한 부분이 남아있습니다. `/speckit.clarify`를 실행하여 누락된 부분이 없는지 확인해 보십시오.
**Q: 일부 단계를 건너뛸 수 있나요?**
가능합니다. 최소 프로세스는 specify → tasks → implement입니다. 그러나 clarify와 plan을 건너뛰면 후반 재작업 위험이 증가할 수 있습니다.
**Q: 이미 생성된 스펙을 수정하려면 어떻게 하나요?**
spec.md 파일을 직접 편집하면 됩니다. 수정 후에는 일관성을 유지하기 위해 plan과 tasks를 다시 생성하는 것을 권장합니다.
**Q: AI가 생성한 코드가 완전히 잘못되었을 때 어떻게 디버깅하나요?**
몇 가지 단계로 원인을 파악합니다:
1. **스펙 확인**: 스펙이 정말 명확한가요? `/speckit.clarify`를 실행하여 누락이 없는지 확인합니다
2. **계획 확인**: plan.md의 기술 방안이 합리적인가요? 합리적이지 않다면 직접 편집한 후 tasks를 다시 생성합니다
3. **범위 축소**: AI에게 하나의 작업만 실행하게 하고, 출력이 기대에 부합하는지 관찰합니다
4. **제약 조건 추가**: constitution.md에 더 명확한 기술 선호도를 추가합니다
**Q: Plan과 Tasks가 일치하지 않으면 어떻게 하나요?**
`/speckit.analyze`를 실행하면 불일치를 감지할 수 있습니다. 일반적인 원인:
* Plan을 업데이트한 후 Tasks를 다시 생성하는 것을 잊은 경우
* Tasks를 수동으로 편집했지만 Plan을 업데이트하지 않은 경우
* 스펙 변경 후 일부 문서만 업데이트한 경우
해결 방법: spec.md를 기준으로 삼아 plan.md와 tasks.md를 순차적으로 다시 생성합니다.
**Q: 기능 간 의존성을 어떻게 처리하나요?**
기능 B가 기능 A에 의존하는 경우, 두 가지 방법이 있습니다:
1. **스펙 통합**: A와 B를 하나의 spec.md에 작성하여 AI가 통합적으로 계획하도록 합니다
2. **단계별 개발**: A의 전체 프로세스를 먼저 완료한 후 B의 specify를 시작합니다
의존성이 있는 여러 기능을 동시에 개발하는 것은 권장하지 않습니다. 통합 문제를 야기하기 쉽습니다.
***
## 요약
Speckit의 핵심 가치는 프로세스를 추가하는 것이 아니라, **암묵적 지식을 명시적으로 만드는 것**입니다. 사용자 스토리, 기능 요구사항, 성공 기준을 작성하도록 요구받으면, '말하지 않아도 당연한' 것이라고 생각했던 세부사항들이 수면 위로 떠오릅니다.
이 프로세스를 기억하십시오:
```
Specify → Clarify → Plan → Tasks → Implement
요구사항 명확화 설계 분해 실행
```
각 단계는 다음 단계의 모호성을 줄여줍니다. 최종적으로 AI가 받는 것은 명확한 작업 목록이지, 모호한 의도 설명이 아닙니다.
이제 여러분의 프로젝트로 돌아가 `/speckit.specify`로 첫 번째 스펙 주도 개발 프로세스를 시작해 보십시오.
### 추가 읽을거리
* 《[스펙 주도 개발이란 무엇인가](/ko/docs/notes/speckit/concept)》— 핵심 이념 복습
* 《[GSD 심층 분석](/ko/docs/notes/gsd/concept)》— 스펙 주도 사고방식을 채택한 또 다른 컨텍스트 엔지니어링 시스템
* 《[Claude Skills란 무엇인가](/ko/docs/notes/claude-skills/concept)》— Speckit 자체가 하나의 Claude Skill입니다
# AI 시대의 TDD: 모델이 먼저 빨간불에 닿게 하세요
## 먼저 결론부터 얘기해보자
AI가 코드를 작성한 후 TDD는 구식이 아니지만 위치가 변경되었습니다.
과거에는 TDD에 대해 이야기할 때 프로그래머의 자기 훈련에 대해 자주 이야기했습니다. 먼저 테스트를 작성한 다음 구현을 작성하고 작은 단계로 리팩토링하는 것입니다. AI 프로그래밍의 경우 브레이크 시스템과 비슷합니다. 모델이 가장 잘하는 것이 가장 위험한 것이기도 하기 때문입니다. 모델은 완전해 보이는 큰 코드 조각을 빠르게 작성할 수 있습니다.
기능 구현을 요청하면 다음을 제공할 수 있습니다.
* 구현 파일
* 일련의 테스트
* 설명
* "완료"라는 단어
문제는 "완전해 보이는 것"이 엔지니어링 측면에서는 수행되지 않는다는 것입니다. 공학적 의미에서 완료하려면 최소한 다음과 같이 답하십시오.
> 이 동작은 명시적으로 실패한 테스트로 정의됩니까?
> 구현으로 인해 이 실패가 녹색으로 바뀌었나요?
> 녹색으로 변한 후 동작 변경 없이 코드를 정리했나요?
AI 시대에 TDD가 다시 거론되는 이유다.
프로세스를 발전된 것처럼 보이게 만드는 것이 아니라 '모델에 대한 신뢰'를 '피드백에 대한 신뢰'로 바꾸는 것입니다.
## 1. 일반적인 오해: TDD는 "먼저 테스트 작성"이 아닙니다.
많은 사람들이 TDD를 의식으로 이해하기 때문에 싫어합니다.
```text
先写测试。
再写代码。
最后跑一下。
```
이것은 확실히 지루하고 쉽게 형식주의로 바뀔 수 있습니다.
정말 유용한 TDD는 "테스트 파일이 일찍 나타나는 것"이 아니라 "실패가 충분히 일찍 나타나는 것"입니다.
### 핵심은 테스트용이 아닌 빨간색
TDD의 첫 번째 단계는 TEST가 아닌 RED라고 합니다.
빨간색은 시스템이 확실히 실패하도록 먼저 테스트를 작성한다는 의미입니다. 이것이 실패하려면 세 가지가 충족되어야 합니다.
1. 실패합니다.
2. 대상 동작이 존재하지 않기 때문에 실패합니다.
3. 예상대로 실패합니다.
빨간색이 먼저 보이지 않으면 뒤따르는 녹색은 의미가 없습니다.
예를 들어 `slugify("Hello World") -> "hello-world"`을 구현하려고 합니다. 귀중한 RED는 "내가 테스트 파일을 작성했습니다"가 아니라 다음과 같습니다.
```text
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
Failure: NameError: name 'slugify' is not defined
Reason: 目标函数还不存在,符合预期
```
테스트가 사양이 되는 시기입니다. 다음 단계를 달성하려면 이 동작을 true로 만들기만 하면 됩니다.
### 먼저 친환경을 선택한 다음 테스트를 구성하세요. 일반적으로 스토리를 구성합니다.
AI가 다른 방향으로 가는 것은 쉽습니다. 먼저 구현을 작성하고 나중에 테스트를 추가합니다.
이것은 원활한 경험입니다. 코드가 실행되고 테스트가 완료되는 것을 보면 마음속에 "거의"라는 느낌이 들 것입니다. 그러나 여기에는 치명적인 문제가 있습니다. 테스트는 현재 구현을 소급하여 구현하는 것일 뿐일 가능성이 높습니다.
"요구 사항이 무엇이어야 하는지"를 묻는 것이 아니라 "쉽게 통과할 수 있도록 현재 코드를 작성하는 방법"을 묻는 것입니다.
이것이 바로 AI 글쓰기 테스트에서 종종 다음과 같은 냄새가 나는 이유입니다.
* 주장이 현재 구현에 너무 구체적입니다.
* 모의가 너무 많아 실제 경계가 측정되지 않습니다.
* 행복한 경로만 테스트
* 기존 코드를 통과시키려면 어설션을 매우 넓게 만드십시오.
* 어떤 테스트도 이전 코드가 원래 틀렸다는 것을 증명할 수 없습니다.
TDD는 그 반대를 요구합니다. 먼저 요구 사항이 실패하도록 하고 코드가 요구 사항을 따라잡도록 합니다.
## 2. AI 시대에 TDD가 더 필요한 이유
AI 프로그래밍의 핵심 모순은 '코드가 느리게 작성된다'가 아니라 '피드백이 늦게 온다'는 것이다.
TDD가 없으면 일반적으로 다음과 같이 작업합니다.
```text
描述需求 -> AI 写一堆代码 -> 人肉看 diff -> 跑一下 -> 发现问题 -> 回头修
```
질문은 끝까지 쌓이게 됩니다. 그것이 틀렸다는 것을 알게 될 때쯤에는 세 가지 범주가 함께 섞여 있을 수 있습니다.
* 요구사항에 대한 오해
* 구현 경로가 잘못되었습니다.
* 리팩토링으로 인해 기존 동작이 중단됨
TDD의 역할은 이 긴 사슬을 줄이는 것입니다.
### 모델에 결정 가능한 목표를 제공합니다.
"우아하게 쓰기"는 목표가 아닙니다.
"사용자의 로그인 상태가 만료된 후 자동으로 로그인 페이지로 돌아가는 것"은 충분히 구체적이지 않습니다.
더 나은 목표는 다음과 같습니다.
```text
当 access token 过期时:
1. 请求返回 401。
2. 客户端清理本地 session。
3. 用户被重定向到 /login。
4. 原始目标地址被保存在 redirect 参数里。
```
한 단계 더 나아가 그 중 하나를 실패한 테스트로 바꾸십시오.
```text
given expired session
when user opens /settings
then app redirects to /login?redirect=/settings
```
이때 AI는 더 이상 "로그인 상태 만료를 처리하는 방법"을 추측하지 않고 명확한 동작을 완료합니다.
### 대규모 작업을 작은 닫힌 루프로 나눕니다.
AI가 통제력을 잃기 가장 쉬운 곳은 모든 것을 한 번에 수행하는 것입니다.
로그인, 권한, 새로 고침 토큰, 오류 프롬프트 및 경로 점프를 한 번에 구현하면 결국 큰 차이가 발생합니다. 작동할 수도 있지만 검토 비용이 높습니다. 비즈니스, 상태, 라우팅, 경계, 테스트 및 리팩토링을 동시에 판단해야 합니다.
TDD의 리듬은 다음과 같습니다.
```text
一个行为 -> 一个失败测试 -> 最小实现 -> 变绿 -> 再下一个行为
```
한 번에 소량씩만 진행하세요. 이해할 수 있을 만큼 작고, AI가 이야기를 만들어내기 어려울 만큼 작으며, 실패했을 때 빠르게 찾아낼 수 있을 만큼 작습니다.
### 모델이 "부드럽게 플레이"하도록 제한합니다.
AI의 일반적인 문제는 지나친 열성입니다.
경계 버그 수정을 요청하면 도우미가 추출됩니다. 테스트를 추가하도록 요청하면 구현이 변경됩니다. 리팩토링을 요청하면 동작이 변경됩니다.
TDD는 단계를 사용하여 이러한 작업을 구분합니다.
| 스테이지 | 해야 할 일 | 하지 말아야 할 것 |
| ---- | ---------- | -------------- |
| 레드 | 실패한 테스트 작성 | 프로덕션 구현 작성 |
| 그린 | 최소 구현 작성 | 녹색이 되도록 테스트 수정 |
| 리팩터 | 구조 정리 | 새로운 행동 도입 |
이 표는 "조심하세요"보다 더 유용합니다. 이는 모델이 어느 단계에 있는지 알려주고 인간이 경계 위반을 더 쉽게 감지할 수 있도록 해줍니다.
## 3. 레드와 그린의 재구성: 세 개의 문이 아닌 세 개의 슬로건
“Red, Green, Refactor”는 쉽게 슬로건이 될 수 있습니다. 실제 사용에서는 세 개의 문처럼 보여야 합니다. 문을 통과할 때마다 증거를 남겨야 합니다.
### 첫 번째 문: RED, 요구 사항이 충족되지 않았음을 증명
RED 단계에서 가장 중요한 질문은 다음과 같습니다.
> 이 테스트가 실패하면 아직 목표 행동이 부족하다는 것을 증명하는 것인가요?
나쁜 빨간색:
```py
assert True
```
그다지 좋은 RED도 아닙니다.
```py
assert "hello" in format_title("Hello World")
```
너무 넓습니다. 많은 잘못된 구현도 통과됩니다.
더 나은 빨간색:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
이 테스트는 작지만 명확합니다. 입력, 출력 및 동작을 지정합니다.
### 두 번째 문: 녹색, 현재 테스트만 통과하도록 놔두세요.
GREEN 단계는 최종 아키텍처를 작성하는 단계가 아닙니다.
임무는 단 하나입니다. 최소한의 코드로 현재 실패한 테스트를 통과하는 것입니다.
이 진술은 직관에 반하는 것처럼 들립니다. 많은 사람들은 "최소 구현"이 너무 추악한 것은 아닌지 걱정합니다. 네, 때로는 추악할 수도 있습니다. 그러나 그 가치는 설계 압력을 유지하는 데 있습니다.
첫 번째 테스트가 다음과 같은 경우:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
허용되는 GREEN은 다음과 같습니다.
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
중국어, 악센트, 연속 구두점, 이모티콘, SEO 특수 사례를 바로 지원할 필요는 없습니다. 이는 이후 테스트를 통해 구동되어야 합니다.
### 세 번째 문: REFACTOR, 구조만 변경하고 동작은 변경하지 않습니다.
REFACTOR 단계는 AI가 가장 혼동하기 쉬운 단계입니다.
"코드 정리"는 "그런데 코드를 향상시키다"로 해석됩니다. 이것은 작동하지 않습니다. 리팩토링의 정의는 매우 좁습니다. 외부 동작은 변경되지 않지만 내부 구조는 더 좋아집니다.
좋은 리팩토링은 다음과 같습니다:
* 변수 이름을 좀 더 정확한 이름으로 변경하세요.
* 반복되는 표현 추출
* 너무 깊은 조건 분기 제거
* 모듈 책임을 더 명확하게 하기 위해 기능 위치 이동
잘못된 리팩토링은 다음과 같습니다.
* 이제 새로운 입력이 지원됩니다.
* 오류 메시지가 쉽게 변경되었습니다.
* 종속성을 쉽게 변경했습니다.
* 테스트 어설션을 편리하게 변경했습니다.
판단 기준은 간단합니다.
> 이 커밋이 방금 `refactor:`이라고 불렸다면 테스트 전후에 녹색이 동일해야 하며 사용자 동작도 동일해야 합니다.
## 4. 좋은 테스트의 맛
TDD는 더 많은 테스트가 더 좋아지는 것에 관한 것이 아닙니다. AI는 가치가 거의 없는 일련의 테스트를 생성하는 데에도 매우 능숙합니다.
더 중요한 것은 맛을 테스트하는 것입니다.
### 좋은 테스트는 스펙과 같습니다
좋은 테스트는 비즈니스 사양처럼 읽어야 합니다.
```text
当用户没有权限时,保存按钮不可点击。
当标题为空时,表单显示错误信息。
当重复提交同一个请求时,只创建一条记录。
```
내부적으로 수행되는 작업이 아니라 외부 작업과 관련이 있습니다.
잘못된 테스트는 구현 노트처럼 보입니다.
```text
应该调用 validateInput 三次。
应该读取 state.user.flags。
应该触发 handleClick 内部函数。
```
구현 세부 사항이 묶여지면 리팩토링이 어려워집니다. 방금 내부 구조를 변경했지만 테스트가 대규모로 실패했습니다. 이러한 테스트는 코드를 보호하기보다는 코드를 동결시킵니다.
### 좋은 테스트에는 경계가 있습니다
테스트는 단 하나의 질문에만 답하도록 설계되는 것이 가장 좋습니다.
테스트가 다음을 주장하는 경우:
* 올바른 형식
* 권한이 올바른지
* 네트워크 요청이 올바른지
* 토스트 카피가 정확합니다
* 데이터베이스 상태가 올바른지
실패하면 문제가 무엇인지 알기가 어렵습니다.
AI는 특히 한 번에 많은 것을 증명하려고 하기 때문에 이러한 "대규모, 포괄적" 테스트를 작성하는 경향이 있습니다. TDD는 그 반대입니다. 즉, 작업, 실패, 구현입니다.
### 좋은 테스트는 구현의 부정 행위를 어렵게 만듭니다.
테스트가 너무 구체적인 입력만 다루는 경우 AI는 일치하는 가짜 구현을 작성할 수 있습니다.
예를 들면:
```py
def slugify(text: str) -> str:
return "hello-world"
```
첫 번째 테스트에서는 통과하지만 두 번째 테스트에서는 실제 논리가 강제로 실행됩니다.
```py
def test_slugify_handles_another_title():
assert slugify("Test Driven Development") == "test-driven-development"
```
따라서 TDD는 항상 하나의 테스트만 작성하는 것이 아니라 각 라운드마다 하나의 행동 압력만 추가합니다. 압력이 점차 증가하고 디자인이 점차 커집니다.
## 5. AI는 어떻게 TDD를 우회할 것인가?
AI는 당연히 테스트를 존중하지 않기 때문에 이 부분을 분명히 해야 합니다.
최적화 목표는 간단합니다. 방금 언급한 작업을 완료하는 것입니다. "테스트 통과"라고 말하면 인간이 원하지 않는 일부 작업을 수행할 수 있습니다.
### 첫 번째 유형: 테스트를 녹색으로 변경
가장 일반적인 것:
* `assert slugify("Hello World") == "hello-world"`을 현재 출력으로 변경
* 실패한 어설션 제거
* 테스트에 `skip` 추가
* 엄격한 주장을 느슨한 주장으로 변경
이것은 TDD가 아닙니다. 이것은 빨간불을 끄는 것입니다.
### 두 번째 유형: 과적합 구현 작성
예를 들어 테스트에는 입력이 하나만 있습니다.
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
모델은 다음과 같이 작성될 수 있습니다.
```py
def slugify(text: str) -> str:
if text == "Hello World":
return "hello-world"
return text
```
지금은 꾸짖을 필요가 없습니다. 구현이 계속해서 하드 코딩될 수 없도록 다음 동작을 계속 추가해야 합니다.
### 세 번째 방법: 모의를 사용하여 실제 경계를 덮습니다.
AI는 모의를 좋아합니다. Mock은 테스트 작성을 더 쉽게 만들고 많은 실제 문제를 사라지게 만듭니다.
조롱할 수 없는 것은 아니지만 이렇게 질문해야 합니다.
> 지금 제가 조롱하고 있는 것은 느린 의존성인가요, 아니면 정말 검증하고 싶은 경계인가요?
결제 콜백 구문 분석을 검증하고 싶지만 구문 분석 레이어를 모의하는 경우 테스트는 의미가 없습니다.
## 6. TDD를 사용하지 말아야 할 경우
TDD에는 가치가 있지만 모든 것이 가치 있는 것은 아닙니다.
### 부적합한 장면
* 순수한 시각적 미세 조정
* 일회성 스크립트
* 기술 탐색 데모
* 요구사항 자체가 명확하게 고려되지 않은 프로토타입
* 테스트 프레임워크가 아직 창고를 설정하지 않았습니다.
이러한 시나리오에서는 탐색 속도를 먼저 추구하고 프로세스에 방해를 받지 마십시오.
### 적합한 장면
* 버그 수정
* 권한, 청구, 상태 머신
* 데이터 변환 및 경계 처리
* 오랫동안 유지될 핵심 모듈
* AI가 반복적으로 수정하는 코드 경로
판단 기준은 "이 기능이 훌륭한지 아닌지"가 아니라,
> 틀리면 비용이 뻔한가요?
비용은 명백하므로 먼저 테스트를 작성해 볼 가치가 있습니다.
## 7. 실행 가능한 멘탈 메소드
AI에게 단 한 문장만 준다면 다음과 같이 말하지 않을 것입니다.
```text
请高质量实现这个功能。
```
나는 말할 것입니다 :
```text
先写一个失败测试,运行它,确认失败原因符合预期。不要写实现,直到我说 go。
```
이 문장은 모델에게 "잘 행동하라"고 요구하는 것이 아니라, 확인할 수 있는 프로세스에 들어가라고 요구하기 때문에 품질이 더 높습니다.
좀 더 완벽함:
```text
每轮只处理一个行为。
RED:写一个失败测试并运行。
GREEN:写最小实现,不改测试。
REFACTOR:只在绿色状态下整理结构。
每轮报告测试文件、命令、失败原因、通过结果。
```
이것이 AI시대 TDD의 핵심이다.
테스트나 프로세스에 대해 미신을 믿는 것이 아니라 모든 단계에 대한 증거를 확보하는 것에 대한 것입니다.
## 마무리
AI 프로그래밍에 가장 필요한 것은 더 많은 코드가 아니라 더 짧은 피드백입니다.
TDD의 가치는 다음과 같습니다. "나는 그것이 옳다고 생각했습니다"를 "여기에 실패가 있었고 그 다음에는 녹색으로 변했습니다."로 바뀌었습니다. 변화는 작지만 충분히 현실적입니다.
한 문장만 기억한다면 다음을 기억하세요.
> AI가 코드를 직접 전달하도록 하지 마세요. 먼저 빨간색 빛을 전달한 다음 빨간색 빛을 녹색으로 바꾸도록 합니다.
다음 글 [실용 가이드](/ko/docs/notes/tdd-with-ai/practice)에서는 이 리듬을 직접 복사할 수 있는 워크플로로 바꿀 것입니다.
## 추천 리소스
# AI 시대의 TDD: Codex 실용 매뉴얼
## 먼저 지도를 주세요
[개념](/ko/docs/notes/tdd-with-ai/concept) 내용: AI 프로그래밍에서 TDD의 가치는 "먼저 테스트를 작성하는" 의식이 아니라 먼저 검증 가능한 빨간색 표시등을 만든 다음 구현을 통해 이를 녹색으로 바꾸는 것입니다.
이 문서에서는 Codex를 구현하는 방법에 대해 설명합니다.
“어떤 서류를 일치시켜야 합니까?”라고 즉시 묻지 마십시오. 더 나은 질문은 다음과 같습니다.
> 매번 동일한 TDD 작업 흐름에서 Codex가 작동하도록 하려면 어떻게 해야 합니까?
이 워크플로는 네 가지 수준으로 나눌 수 있습니다.
| 계층 구조 | 뭐하는거야 | 어디 | 언제가 적절한가 |
| ----- | ---------- | ----------------------------------- | ---------------------------------- |
| L1 | 프로젝트 분야 작성 | `AGENTS.md` | 모든 프로젝트에는 |
| L2 | 응고과정 | `.agents/skills/tdd-codex/SKILL.md` | 요구 사항을 충족하기 위해 TDD를 반복적으로 사용 |
| L3 | 격리 단계 | `.codex/agents/*.toml` | 복잡한 업무, 테스트와 구현 사이의 상호 오염 우려 |
| L4 | 자동 알림 | `.codex/hooks.json` | AI가 비밀리에 테스트를 수정하는 것을 두려워하는 중요한 창고 |
사용 가능한 가장 작은 버전은 L1 + L2입니다.
완전한 방어선은 L1 + L2 + L3 + L4입니다.
## 1. 먼저 '완성'이 무엇인지 정의해보세요
완성 표준이 없다면 Codex는 "작성된 코드"를 "완료"로 쉽게 처리할 수 있습니다.
TDD 시나리오에서는 완료 기준이 더 구체적이어야 합니다.
### 매 라운드마다 증거가 전달되어야 합니다.
Codex가 매 라운드마다 다음 6가지 항목을 보고하도록 합니다.
```text
Behavior: 这一轮实现哪个行为
Test: 测试文件和测试名
Command: 跑了什么命令
RED: 失败原因是否符合预期
GREEN: 通过结果是什么
REFACTOR: 是否重构,为什么
```
이는 "완료"보다 훨씬 더 유용합니다.
먼저 구현을 작성한 다음 합리적으로 보이는 테스트를 추가하는 대신 모델이 실제로 빨간색과 녹색 주기를 거쳤음을 알려줍니다.
### 한 라운드에서는 하나의 행동만 처리됩니다.
이것은 매우 중요합니다.
Codex가 전체 테스트 매트릭스를 한 번에 생성하도록 하지 마십시오. 그것은 "수평 배치 테스트"가 될 것입니다.
```text
RED: test1, test2, test3, test4, test5
GREEN: 一次写一个大实现
```
원하는 것은 세로로 자르는 것입니다.
```text
RED test1 -> GREEN impl1 -> REFACTOR
RED test2 -> GREEN impl2 -> REFACTOR
RED test3 -> GREEN impl3 -> REFACTOR
```
구현의 첫 번째 라운드에서는 문제에 대한 이해가 바뀔 것입니다. 모든 테스트를 한 번에 작성하지 마세요.
## 2. L1: AGENTS.md에 규율을 작성합니다.
`AGENTS.md`은 Codex가 프로젝트에 들어갈 때 읽는 설명 파일입니다.
OpenAI 공식 문서에 따르면 Codex는 먼저 전역 설명을 읽은 다음 프로젝트 루트 디렉터리에서 현재 디렉터리까지 읽습니다. 각 레이어는 `AGENTS.override.md`을 먼저 읽고, 그렇지 않으면 `AGENTS.md`을 읽습니다. 현재 디렉터리에 더 가까운 설명은 나중에 나타나므로 우선 순위가 더 높습니다. 기본 병합 제한은 `32 KiB`이므로 여기에 긴 튜토리얼을 작성할 수 없습니다.
### 프로젝트 교통 규칙과 같아야 합니다.
`AGENTS.md`은 TDD가 무엇인지 Codex에 가르칠 책임이 없습니다. 이 프로젝트에서 허용되지 않는 동작은 무엇인지 명확하게 작성하는 것만 담당합니다.
이 단락을 직접 넣을 수 있습니다.
```markdown
# TDD Rules
- For new behavior and bug fixes, use red/green TDD.
- RED: write exactly one failing behavior test first.
- Run the smallest relevant test command and confirm the failure is expected.
- Do not edit production implementation during RED.
- GREEN: write the minimum production code required to pass the current failing test.
- Never modify, delete, skip, or weaken tests to make implementation pass.
- REFACTOR only after tests are green.
- Keep structural changes and behavior changes separate.
- Report Behavior, Test, Command, RED, GREEN, and REFACTOR for each cycle.
```
추가 프로젝트 명령:
```markdown
# Verification
- Use `pytest` or the smallest relevant pytest command for Python behavior tests.
- Use `npm run types:check` only when this blog site's MDX or TypeScript changes.
- Use the smallest targeted command during RED/GREEN loops.
- If a command is slow, explain what targeted command was used first and what full command remains.
```
### 백과사전으로 써서는 안 된다
잘못된 `AGENTS.md`은(는) 다음으로 채워집니다.
* TDD의 역사
* 모든 테스트 철학
* 다양한 프레임워크 튜토리얼
* 복잡한 프롬프트 템플릿
* 다양한 언어로 완전한 사양
이러한 것들은 정말로 중요한 규칙을 희석시킵니다.
내 제안은 다음과 같습니다. `AGENTS.md` 상주 규율만 적용하세요. 긴 프로세스에 기술을 사용하십시오.
## 3. L2: 프로세스를 코덱스 스킬로 만들기
`AGENTS.md`은 "기본 규율"을 해결하고 기술은 "전체 프로세스"를 해결합니다.
Codex에 "do it by TDD"라고 자주 말할 때, 이 문장을 프로젝트 수준의 기술로 업그레이드해야 합니다.
### 디렉토리 구조
여기에 넣으세요:
```text
.agents/
skills/
tdd-codex/
SKILL.md
```
Codex는 현재 디렉터리부터 `.agents/skills`까지 스캔합니다. 웨어하우스 루트 디렉터리의 기술은 팀에서 사용하는 워크플로에 적합합니다.
### 최소 사용 가능 SKILL.md
```markdown
---
name: tdd-codex
description: Implementing or fixing maintainable code with Codex using strict red-green-refactor TDD. Use for new behavior, bug reproduction, behavior tests, or safe AI coding.
---
# TDD Codex Workflow
Use one behavior slice per cycle.
## Phase 0: Scope
Identify one observable behavior.
Name the public API, user flow, or integration boundary under test.
Do not edit production code.
## Phase 1: RED
Write exactly one failing behavior test.
Prefer public behavior over implementation details.
Run the smallest relevant test command.
Confirm the failure is expected.
Stop and report:
- Behavior
- Test file
- Command
- Failure reason
## Phase 2: GREEN
Write the minimum production code to pass the current failing test.
Never modify, delete, skip, or weaken tests to pass.
Do not add speculative features or abstractions.
Run the same test command.
Report the passing result.
## Phase 3: REFACTOR
Only refactor after tests are green.
If the code is already simple, skip.
If refactoring, make one structural change at a time.
Run tests after each refactor.
Do not change behavior.
## Cycle Report
Return:
- Behavior:
- Test:
- Command:
- RED:
- GREEN:
- REFACTOR:
- Next slice:
```
### 호출 방법
이제부터 다음과 같이 말할 수 있습니다.
```text
用 tdd-codex skill 做这个需求。
每轮只处理一个行为。
先 RED,确认失败后停下来,不要直接写实现。
```
또는 더 짧게:
```text
用 tdd-codex。先写红灯,等我说 go。
```
요점은 프롬프트가 얼마나 아름다운가가 아니라 매번 Codex를 동일한 트랙으로 다시 가져오는 것입니다.
## 4. L3: 하위 에이전트를 사용하여 빨간색과 녹색 재구성을 분리합니다.
모든 작업에 하위 에이전트가 필요한 것은 아닙니다.
그러나 작업이 복잡하고, 테스트가 구현에 의해 쉽게 오염되고, 리팩토링이 통제를 벗어나기 쉬운 경우에는 RED, GREEN, REFACTOR를 여러 에이전트로 분할하는 것이 더 안정적입니다.
### 언제 해체할 가치가 있나요?
분해에 적합:
* 권한, 청구, 상태 머신
-다중 모듈 기능
* 버그가 매우 숨겨져 있으므로 먼저 재발 테스트를 작성해야 합니다.
* 모델은 항상 친환경을 얻기 위한 테스트로 변경됩니다.
* 테스트 품질 검토만 담당하는 사람을 원합니다.
분해에 적합하지 않음:
* 가젯 기능
* 카피라이팅 변경
* 순수한 시각적 미세 조정
* 일회성 스크립트
에이전트를 해체하는 데 드는 비용은 현실입니다. 격리의 이점이 통신 비용보다 클 경우에만 사용하십시오.
### RED agent
```toml
# .codex/agents/tdd-test-writer.toml
name = "tdd_test_writer"
description = "RED phase agent. Writes one failing behavior test and stops before implementation."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the RED phase.
Write exactly one behavior-focused test for the requested slice.
Prefer public APIs and user-visible behavior over implementation details.
Run the smallest relevant test command.
Confirm the test fails for the expected reason.
Do not edit production implementation.
Do not add multiple tests at once.
Return Behavior, Test, Command, and RED failure reason.
"""
```
### GREEN agent
```toml
# .codex/agents/tdd-implementer.toml
name = "tdd_implementer"
description = "GREEN phase agent. Implements the minimum production code to pass the current failing test."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the GREEN phase.
Read the failing test and relevant production code.
Write the minimum implementation required to pass the current test.
Never modify, delete, skip, or weaken tests to make them pass.
Do not add speculative features, helpers, configuration, or abstractions.
Run the relevant tests and return the command plus passing output.
"""
```
### REFACTOR agent
```toml
# .codex/agents/tdd-refactorer.toml
name = "tdd_refactorer"
description = "REFACTOR phase agent. Improves structure only after tests are green."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the REFACTOR phase.
Start by running the relevant tests to confirm the code is green.
Look for duplication, unclear names, needless branching, or misplaced responsibility.
Skip refactoring when the code is already simple.
If you refactor, make one structural change at a time.
Run tests after each refactor.
Never change behavior in this phase.
"""
```
### 메인 세션 명령 방법
```text
按三阶段 TDD 做这个 slice:
1. tdd_test_writer 只写一个失败测试,并确认 RED。
2. 等我确认后,tdd_implementer 写最小实现,并确认 GREEN。
3. tdd_refactorer 判断是否需要结构重构。
不要批量铺测试。
不要在 GREEN 阶段修改测试。
```
여기서 중요한 점은 컨텍스트를 분리하는 것입니다. 테스트를 작성하는 사람들은 구현 세부 사항에 영향을 받지 않도록 노력해야 합니다. 구현을 작성하는 사람들은 수동으로 테스트할 수 없습니다. 리팩토링하는 사람들은 새로운 행동을 도입할 수 없습니다.
## 5. L4: Hooks를 사용하여 테스트 비교에 집중
규칙에만 의존하는 경우에도 모델이 여전히 범위를 벗어날 수 있습니다.
가장 일반적인 국경 간 시도는 테스트가 빨간색이고 모델이 테스트를 녹색으로 바꾸기 위해 변경하는 것입니다.
Hooks의 가치는 "완전히 안전"하다는 것이 아니라 이 작업을 즉시 노출시키는 것입니다.
### Codex 후크 활성화
먼저 구성에서 기능 플래그를 엽니다.
```toml
# ~/.codex/config.toml 或 /.codex/config.toml
[features]
codex_hooks = true
```
Codex는 활성 구성 레이어 옆에 있는 후크를 찾습니다. 일반적인 위치:
* `~/.codex/hooks.json`
* `~/.codex/config.toml`
* `/.codex/hooks.json`
* `/.codex/config.toml`
프로젝트 레벨에서는 웨어하우스를 따라갈 수 있으므로 `/.codex/hooks.json`을 먼저 사용하는 것이 좋습니다.
### hooks.json
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/watch-test-edits.sh\"",
"timeout": 10,
"statusMessage": "Checking test file edits"
}
]
},
{
"matcher": "Bash|apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/run-fast-check.sh\"",
"timeout": 120,
"statusMessage": "Running fast checks"
}
]
}
]
}
}
```
### 테스트 파일이 수정되었는지 확인
```bash
# .codex/hooks/watch-test-edits.sh
#!/usr/bin/env bash
set -euo pipefail
changed_tests="$(
git diff --name-only |
grep -E '(^|/)(__tests__|tests?)/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.go$|test_.*\.py$' || true
)"
if [ -n "$changed_tests" ]; then
cat < "hello-world"
- trim leading/trailing spaces
- collapse repeated spaces into one hyphen
- remove punctuation
- normalize "Café" -> "cafe"
- empty input returns empty string
```
### 1단계: 첫 번째 빨간불
```text
读取 SPEC.md。
只实现第一条行为:"Hello World" -> "hello-world"。
先 RED:只写一个失败测试,运行它,确认失败。
不要写生产实现。
```
이상적인 출력:
```text
Behavior: basic title becomes lowercase hyphenated slug
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
RED: failed because slugify is not defined
```
그래야만 계속할 수 있습니다.
### 2단계: 최소한의 녹색
```text
go
```
Codex는 최소한의 구현을 작성합니다.
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
그런 다음 다음을 보고하십시오.
```text
GREEN: pytest tests/test_slugify.py -q passed
REFACTOR: skipped, implementation is still simple
Next slice: trim leading/trailing spaces
```
### 3단계: 두 번째 빨간불
```text
继续下一条:去掉首尾空格。
先 RED,只写一个测试。
```
테스트:
```py
def test_slugify_trims_spaces():
assert slugify(" Hello World ") == "hello-world"
```
현재 구현이 `-hello-world-`을 출력하면 빨간색 표시등이 true입니다.
그런 다음 녹색:
```py
def slugify(text: str) -> str:
return text.strip().lower().replace(" ", "-")
```
### 4단계: 추상화를 서두르지 마세요.
이 시점에서 많은 AI는 `normalizeInput`, `removePunctuation`, `toAscii`를 그리고 싶어할 것입니다.
아직 서두르지 마세요.
TDD의 디자인은 상상이 아니라 테스트 압력으로 밀어내야 한다. 유니코드, 구두점, 빈 문자열을 추가하고 구조적 압박이 실제로 발생할 때까지 기다린 다음 리팩토링하세요.
## 8. 일일 사용량 간편 확인
### 새로운 기능
```text
用 TDD 实现这个需求。
每轮只处理一个行为。
先写一个失败测试并运行确认 RED。
不要写生产实现,直到我说 go。
```
### 버그 수정
```text
先写一个能复现这个 bug 的失败测试。
确认它因为这个 bug 失败后,再写最小修复。
不要改测试来适配当前实现。
```
### 복잡한 기능
```text
先不要写代码。
请给出 TDD 分解计划:
- 外圈集成测试是什么
- 内圈每个行为 slice 是什么
- 每轮用什么命令验证
- 哪些地方不能 mock
等我确认后再开始 RED。
```
### Review
```text
Review 这次改动,重点看:
- 是否先有失败测试
- 测试是否测行为而不是实现
- 是否存在为了通过而弱化测试
- 结构改动和行为改动是否混在一起
- 是否缺少外圈集成测试
```
## 9. 최종 체크리스트
Codex에게 TDD를 해달라고 요청할 때마다 이 표를 이용해 마지막에 확인해보세요.
| 질문 | 자격기준 |
| ------------------ | --------------------------- |
| 정말 인기가 많은가 먼저 | 실패한 명령과 실패 이유가 있습니까? |
| 빨간색이 맞나요 | 실패 이유는 목표 행동의 부족에 해당 |
| 라운드당 하나의 행동만 수행 | 일괄 테스트 없음 |
| GREEN 시험이 변경되었나요? | 녹색으로 변경된 테스트가 없습니다 |
| 테스트는 동작을 테스트하는가 | 내부 구현 세부 사항에 의존하지 않습니다 |
| 혼합된 동작이 리팩토링되었습니까? | 구조적 변화와 행동 변화를 분리 |
| 완전한 검증이 있습니까 | 대상 테스트 및 필요한 전체 검사가 실행되었습니다 |
이 테이블이 통과할 수 없다면 서두르지 말고 병합하세요.
## 추천 리소스
# Cuando la IA permite "aprobar sin estudiar", ¿qué le queda a la universidad?
Anthropic realizó recientemente una entrevista en la que reunió a cuatro estudiantes de Princeton, Berkeley, la London School of Economics y la Universidad Estatal de Arizona para hablar sobre la situación real de la IA en el campus. La conversación duró casi 40 minutos, sin discursos promocionales: solo confusión, ansiedad y reflexión genuinas.
Esta entrevista revela un problema más profundo: **la IA no solo está cambiando la forma de aprender, sino que está desarmando la lógica fundamental de todo el sistema educativo**.
***
## Una realidad imposible de ignorar
Al inicio de la entrevista, el moderador hizo una pregunta directa: ¿cuál es el ambiente actual respecto a la IA en el campus?
La respuesta: **el 90% de los estudiantes usa IA**. No de forma ocasional, sino como parte de su flujo de trabajo diario: resumir apuntes de clase, responder cuestionarios, obtener retroalimentación sobre tareas, analizar casos de negocio, hacer investigación de mercado, completar análisis financieros. Algunos estudiantes incluso la usan para resolver exámenes, con una razón muy práctica: cuando se es estudiante de posgrado y se tienen varios trabajos simultáneamente, no siempre hay tiempo disponible.
Pero lo más interesante es que, aunque casi todos la usan, nadie sabe cuáles son las reglas. Algunos cursos prohíben explícitamente la IA, otros la fomentan activamente, y la mayoría se encuentra en una zona ambigua. Los estudiantes no saben dónde están los límites; los profesores tampoco saben cómo gestionarlo. Este estado se describe como una "zona gris": quieren usarla, pero temen violar las normas; si no la usan, sienten que se quedan atrás.
Lo más peligroso de esta zona gris no es que los estudiantes puedan infringir las reglas, sino que **impide que ocurran discusiones verdaderamente valiosas**. Los estudiantes no pueden compartir abiertamente las mejores prácticas de uso de IA, los profesores no pueden orientar a los estudiantes sobre cómo usar la herramienta de manera responsable, y toda la comunidad académica queda atrapada en un estado incómodo de prohibición superficial y uso privado.
Y cuando las reglas no pueden aplicarse de manera efectiva, se convierten en un filtro que separa a "quienes saben disimular" de "quienes no saben disimular". Una prohibición formal pero con uso generalizado en privado **no impedirá que los estudiantes usen IA; solo impedirá que discutan abiertamente cómo usarla mejor**.
***
## La IA es un espejo
En la entrevista surgió una observación muy aguda: la inteligencia artificial, y en particular la forma en que los estudiantes la usan, revela mucho sobre sus motivaciones.
La idea detrás de esta afirmación es que **la IA se ha convertido en un espejo que refleja el verdadero propósito de ir a la universidad**.
La entrevista clasifica los objetivos universitarios en tres categorías: primero, profundizar en el conocimiento especializado y dominar la comprensión profunda de un campo; segundo, prepararse para la carrera profesional, encontrar un buen empleo y construir una red de contactos; tercero, ampliar las relaciones sociales, disfrutar de la vida social y experimentar la cultura universitaria. Cada estudiante asigna un peso diferente a estos tres objetivos, y la forma en que usa la IA expone esos pesos con precisión.
Si a usted solo le importa "aprobar el examen" y "obtener el título", enviará directamente el resultado de la IA como tarea. Esto no es un juicio moral, sino una realidad: cuando la tecnología permite alcanzar un objetivo con el mínimo esfuerzo, ¿por qué tomar el camino largo? Si su objetivo siempre fue obtener un título para conseguir empleo, usar IA para completar tareas es una decisión perfectamente racional.
Pero si usted realmente quiere aprender y comprender un campo en profundidad, usará la IA como un interlocutor. Le hará preguntas, le pedirá que explique conceptos y luego reformulará las ideas con sus propias palabras. Le pedirá que escriba una primera versión del código y luego usted mismo lo refactorizará y optimizará. Se asegurará de que en cada etapa comprende verdaderamente lo que está sucediendo.
Esta polarización no solo existe entre distintos estudiantes, sino también entre distintas disciplinas. Los estudiantes de humanidades tienden a prescindir de la IA, porque su aprendizaje requiere lectura atenta: leer cuidadosamente el texto original, saborear los matices del lenguaje, comprender la intención del autor. La IA interfiere con este proceso, porque ofrece resúmenes y paráfrasis, no la experiencia directa del texto original. Los estudiantes de ingeniería y negocios, en cambio, usan la IA extensamente, porque reduce las barreras técnicas y permite que personas sin formación en informática puedan escribir código, crear sitios web y analizar datos.
**Esta polarización es, en esencia, una diferencia en la comprensión de "qué vale la pena aprender"**. Para el estudiante de humanidades, la experiencia de leer a Shakespeare en el original ya es aprendizaje en sí misma; para el estudiante de ingeniería, lo importante es si puede resolver el problema, no si escribe personalmente cada línea de código. La IA hace que esta diferencia sea aún más evidente.
***
## ¿Herramienta o muleta? Un criterio simple para distinguirlas
En la entrevista surgió una pregunta clave: ¿cómo distinguir si la IA es una herramienta o una muleta?
Los estudiantes dieron una respuesta sorprendentemente unánime: **si puede explicarlo o no**.
Si usted no puede explicar lo que creó, si no puede describir qué papel desempeñó la IA en el proceso, entonces es una muleta. Si puede explicarlo como si estuviera hablando con un estudiante de primaria, si puede ofrecer explicaciones tanto de nivel básico como avanzado, entonces es una herramienta.
Este criterio parece simple, pero toca la esencia del aprendizaje. La lógica central del método Feynman es: si no puede explicar un concepto con palabras sencillas, entonces aún no lo ha comprendido realmente. El aprendizaje en la era de la IA sigue la misma lógica: si no puede explicar lo que la IA hizo por usted, entonces solo está "externalizando el pensamiento", no "potenciándolo".
En la entrevista se mencionó un ejemplo interesante: un estudiante desarrolló una herramienta que permite cargar las diapositivas de una clase, y la IA genera anotaciones similares a las de un profesor junto a cada diapositiva. Este estudiante dijo: "Funciona bien porque yo ya le he indicado qué quiero saber: las definiciones de ciertos elementos en las diapositivas. Las diapositivas a veces son abstractas, carecen de contexto, y es necesario agregar explicaciones al margen."
La clave está en "yo ya le he indicado qué quiero saber". Este estudiante conoce sus propias lagunas de conocimiento, sabe qué tipo de ayuda necesita y luego guía activamente a la IA para que brinde esa ayuda. Esto es una herramienta. Si en cambio simplemente le entrega las diapositivas a la IA y le pide que "resuma esta clase", y luego memoriza ese resumen, eso es una muleta.
**La diferencia radica en la proactividad y la comprensión**. Quien usa una herramienta sabe lo que está haciendo y controla todo el proceso. Quien usa una muleta cede el control a la tecnología y se convierte en un receptor pasivo.
***
## El rezago de las instituciones no es lentitud de reacción, sino una incapacidad estructural para responder
En la entrevista se mencionaron algunos esfuerzos institucionales. La London School of Economics tiene un curso obligatorio que comienza a orientar a los estudiantes sobre cómo usar Claude: conversar con la herramienta, asignarle diferentes roles y luego entregar los registros de conversación para que se evalúe cómo interactuaron con la IA. El centro de gestión profesional de la Universidad Estatal de Arizona creó una biblioteca de prompts con plantillas para diferentes escenarios. Todos estos son buenos intentos, y la idea central es: no prohibir la IA, sino enseñar a los estudiantes a usarla de manera responsable.
Pero estos son la minoría. La mayoría de las universidades aún debate si "deberían permitir que los estudiantes usen IA". Algunos profesores dicen que se puede usar, pero hay que indicar cómo se utilizó en la tarea; algunos cursos la prohíben directamente; otros simplemente no la mencionan, asumiendo que los estudiantes no la usarán. No hay un marco integrado, no hay estándares unificados; todo el sistema está en un estado de caos.
El problema más profundo es: **esta cuestión, por su naturaleza, no puede resolverse con normas y reglamentos**.
La lógica de supervisión de la educación tradicional es: la institución establece reglas, los estudiantes las cumplen y quienes las violan son sancionados. Esta lógica funciona bajo la premisa de que "el comportamiento infractor puede ser detectado". Pero la IA rompe esa premisa.
Se puede prohibir que los estudiantes usen IA al entregar una tarea, pero no se puede monitorear si la usaron durante el proceso de reflexión. Se pueden usar herramientas de detección de IA, pero la precisión de estas herramientas está lejos de alcanzar un nivel que justifique sanciones: la tasa de falsos positivos es demasiado alta, y los estudiantes aprenden rápidamente a evadir la detección. Más importante aún, **fundamentalmente no se puede distinguir entre "una tarea de alta calidad realizada con ayuda de IA" y "una tarea de alta calidad realizada de forma independiente"**, porque un buen uso de la IA debería ser precisamente imperceptible.
En la entrevista se dijo algo muy directo: fundamentalmente, ningún reglamento va a cambiar la forma en que los estudiantes usan la IA. La responsabilidad está en manos de los estudiantes. Esto no es evadir responsabilidades, sino reconocer la realidad.
Cuando la tecnología hace posible "aprobar sin estudiar", las instituciones educativas no enfrentan un problema de gestión, sino un problema existencial: **si los exámenes no pueden demostrar el aprendizaje, ¿cuál es la razón de ser de la universidad?**
***
## La lógica fundamental del sistema educativo se ha roto
Esta pregunta toca la contradicción fundamental del sistema educativo.
La educación tradicional se construyó sobre varias premisas centrales: primero, el conocimiento es escaso y se necesitan instituciones especializadas (universidades) y profesionales (profesores) para transmitirlo; segundo, los resultados del aprendizaje pueden medirse mediante exámenes; tercero, un título demuestra que usted domina el conocimiento de un campo y, por lo tanto, está calificado para trabajar en él.
La IA está rompiendo estas premisas una por una.
**El conocimiento ya no es escaso**. En YouTube hay cursos gratuitos de Stanford, Claude puede responder sus preguntas en cualquier momento, en GitHub hay innumerables proyectos de código abierto para estudiar. No se necesita ir a la universidad para acceder al conocimiento, y ni siquiera es necesario pagar una suscripción para obtener tutoría básica con IA.
**Los exámenes no pueden medir el aprendizaje**. Cuando la IA puede completar la mayoría de las preguntas de un examen, este pasa de ser "una herramienta para medir el nivel de comprensión" a "una herramienta para medir si se sabe usar IA". Esto no significa que los exámenes sean completamente inútiles, sino que ya no pueden distinguir con precisión entre "quien realmente aprendió" y "quien sabe aprovechar las herramientas".
**El valor del título está disminuyendo**. Cuando los empleadores se dan cuenta de que un título no garantiza que el candidato realmente domina los conocimientos relevantes, le dan más importancia a las pruebas de capacidad real: portafolios, experiencia en proyectos, desempeño en prácticas profesionales. El título pasa de ser una "prueba de competencia" a un simple "requisito mínimo".
**Cuando estas premisas se quiebran, la propuesta de valor del sistema educativo necesita ser redefinida**.
La entrevista ofrece una respuesta: el valor de la universidad pasa de "**transmitir conocimiento**" a "**proporcionar un entorno**". Un entorno donde se puede cometer errores, explorar y confrontar ideas con otras personas. Donde se puede pasar un fin de semana con los compañeros de residencia haciendo una "lista de deseos antes de graduarse", probar ideas "probablemente tontas" en un hackathon, debatir con los profesores, discutir con los compañeros y aprender del fracaso sin arriesgar la carrera profesional.
Los proyectos estudiantiles mencionados en la entrevista ilustran bien este punto: "recordatorio automático de inscripción de materias", "buscador de aulas vacías", "ranking de deseos antes de la graduación". Ninguno de estos proyectos es técnicamente complejo, y muchos de sus creadores ni siquiera tienen formación en informática. Pero nacen de emociones humanas reales: el miedo a perderse algo, la búsqueda de comodidad, el aprecio por la vida universitaria.
**Cuando las barreras técnicas se reducen, lo importante ya no es "si sabe programar", sino "qué problema quiere resolver"**. La universidad ofrece un espacio para explorar libremente estas preguntas, un entorno donde se pueden convertir ideas en realidad y aprender del fracaso.
La IA puede completar sus tareas, pero no puede vivir por usted ese período en el que "se puede errar y se puede explorar".
***
## La paradoja del mercado laboral
La segunda mitad de la entrevista abordó el empleo, revelando otra paradoja.
Los estudiantes usan IA para redactar currículos; las empresas usan IA para filtrarlos. Todo el ciclo de contratación se convierte en hablarle a una pantalla: primero se usa IA para escribir la carta de presentación, luego se responden preguntas frente a una cámara, y finalmente se recibe una carta de rechazo generada por IA. Desde el envío del currículum hasta la recepción del rechazo, pueden pasar solo 15 minutos. La eficiencia es alta, pero la humanidad es escasa.
Esto ha creado un mercado laboral de "IA contra IA". Los estudiantes entrenan a la IA para escribir "buenos" currículos; las empresas entrenan a la IA para filtrar "buenos" candidatos. El papel de los seres humanos reales en este proceso es cada vez menor. Hablarle a una pantalla no genera química, no permite mostrar esas cualidades difíciles de cuantificar pero muy importantes: el sentido del humor, la capacidad de adaptación, las sutilezas del trabajo en equipo.
Pero la otra cara de la paradoja es que el dominio de la IA se ha convertido en una nueva ventaja competitiva. Las cuatro grandes firmas de consultoría solían contratar MBA generalistas; ahora buscan específicamente MBA con habilidades en IA. Si usted sabe cómo aplicar la IA en diferentes industrias, será su candidato preferido.
**La paradoja es esta: la IA hace que el mercado laboral sea más frío, y al mismo tiempo hace que el mercado laboral valore más las habilidades en IA. No se puede escapar; solo se puede aprender a usarla de manera efectiva**.
Esto nos lleva de vuelta a la pregunta central: ¿qué significa "usar de manera efectiva"? No es saber usar ChatGPT para escribir correos electrónicos, sino ser capaz de identificar qué problemas son aptos para ser resueltos con IA, poder diseñar prompts que guíen a la IA hacia los resultados que usted necesita, y poder evaluar la calidad del resultado de la IA y realizar las correcciones necesarias.
Esta capacidad no se cultiva prohibiendo la IA, sino a través de mucha práctica y ensayo y error. Por eso las universidades que adoptan activamente la IA, que crean Claude Builder Clubs y organizan hackathons, están ofreciendo a sus estudiantes una educación más valiosa: permiten que los estudiantes aprendan a colaborar con la IA en un entorno relativamente seguro.
***
## La transferencia de la responsabilidad
La observación más importante de la entrevista quizás sea esta: cuando la tecnología permite "aprobar sin estudiar", el sentido del aprendizaje en sí mismo se convierte en una pregunta que cada persona necesita responder por sí misma.
Se trata de una transferencia de responsabilidad. **De la institución al estudiante, de las reglas a la autodisciplina, de la motivación extrínseca a la motivación intrínseca**.
La motivación en la educación tradicional es extrínseca: se necesita aprobar los exámenes para obtener el título, y se necesita el título para conseguir empleo. Este sistema de incentivos externos impulsa a los estudiantes a estudiar. Pero cuando la IA permite aprobar sin estudiar, este sistema de incentivos deja de funcionar.
Lo único que queda es la motivación intrínseca: ¿realmente quiere aprender? ¿De verdad le interesa este campo? ¿Realmente quiere comprender en profundidad, o solo quiere obtener un título?
En la entrevista hay un detalle muy revelador. Un estudiante mencionó que el posgrado tiene su lado negativo: se tienen varios trabajos simultáneamente, no hay tiempo, así que a veces se usa la IA para completar rápidamente un examen. Pero luego añadió: el posgrado debería ser el período en el que se desarrolla el pensamiento crítico, en el que se muestra una faceta más decidida. **Era consciente de la contradicción, pero eligió la eficiencia**.
Esto no es un juicio moral. Bajo la presión de la realidad, la eficiencia suele ser más importante que los ideales. Pero esta elección revela un hecho: **cuando la presión externa (completar el examen) y la motivación interna (aprender en profundidad) entran en conflicto, muchas personas eligen la primera**.
La IA agudiza este conflicto, porque hace que "cumplir con los exámenes" sea extremadamente fácil. En la era anterior a la IA, incluso si solo se quería cumplir con el examen, era necesario aprender algo para aprobarlo. La IA elimina ese paso intermedio: se puede aprobar sin aprender absolutamente nada.
**Esto obliga a cada persona a enfrentar esa pregunta directamente: ¿para qué va usted a la universidad?**
Si la respuesta es "para obtener un título y conseguir empleo", entonces usar IA para completar las tareas es perfectamente razonable. Si la respuesta es "realmente quiero aprender este campo", entonces necesita resistir activamente la tentación de los atajos que ofrece la IA.
La universidad no puede tomar esa decisión por usted. Las reglas no pueden obligarlo a desarrollar motivación intrínseca. **Es su propia responsabilidad**.
***
## La tecnología no esperará a que usted esté listo
Al final de la entrevista, hay una actitud que se mantiene constante: "Ya encontraremos la manera de resolverlo."
¿Las reglas de la universidad no se actualizan a tiempo? Se empieza a usar la herramienta primero y luego se le dice a la institución qué funciona. ¿La IA podría usarse para hacer trampa? Poco a poco se aprende a usarla de manera responsable. ¿El mercado laboral cambió? Se adapta uno a las nuevas reglas del juego.
Esto no es optimismo ciego, sino realismo. La tecnología ya está aquí y no va a esperar a que usted esté listo para empezar a cambiar el mundo. Se puede elegir resistir o adaptarse, pero no se puede elegir detener el tiempo.
**La relación de esta generación de estudiantes con la IA no es de miedo ni de adopción ciega, sino de exploración en medio del caos y aprendizaje a través del ensayo y error**.
Los proyectos que hacen en el Claude Builder Club no son técnicamente complejos, pero resuelven problemas reales. Las ideas que prueban en los hackathons pueden ser tontas, pero al menos lo están intentando. La confusión que sienten en clase —las reglas no son claras— al menos los lleva a reflexionar.
Esta actitud de "**aprender haciendo**" probablemente sea más útil que cualquier reglamento para ayudarlos a adaptarse a la era de la IA.
Kevin Kelly propuso en *What Technology Wants* el concepto de "technium": la tecnología, como un todo, parece tener voluntad propia y querer volverse cada vez más poderosa. En la entrevista hay una observación que resuena con esta idea: en los últimos dos años, todo lo que la IA ha necesitado para seguir avanzando, lo ha conseguido. El cambio de actitud hacia la energía nuclear, la discusión sobre centros de datos espaciales: cada vez que un cuello de botella podría frenar la IA, el obstáculo es eliminado.
Pero quizás la formulación más precisa sea: **los estudiantes crearán lo que necesiten**.
La IA es solo una herramienta. Lo que determinará el futuro es cómo esta generación de estudiantes elija usarla: si para evadir el pensamiento o para potenciarlo; si para cumplir con los exámenes o para explorar el mundo; si como muleta o como herramienta.
Esa decisión no la puede controlar la universidad; solo ellos mismos pueden tomarla.
Y a juzgar por esta entrevista, al menos una parte de los estudiantes está reflexionando seriamente sobre esta cuestión. Eso probablemente sea suficiente.
# Pensar como un Agent: la filosofia de diseno de herramientas del equipo de Claude Code
Thariq es ingeniero en Anthropic y uno de los constructores principales de Claude Code. A finales de febrero publico un extenso hilo en X donde compartio cinco casos reales sobre el diseno de herramientas para agents durante la construccion de Claude Code — no se trata de un marco teorico, sino de experiencia adquirida en primera linea. El articulo obtuvo 3.51 millones de visualizaciones, 209 respuestas y 9,691 likes, generando una gran cantidad de discusiones de alta calidad.
Para presentar la pregunta central de todo el articulo, Thariq utilizo una excelente analogia: imagine que se enfrenta a un problema matematico dificil, que herramientas necesita.
**Papel y lapiz** es la configuracion minima, pero estara limitado al calculo manual. **Una calculadora** es mejor, pero necesita saber como usar las funciones avanzadas. **Una computadora** es lo mas poderoso, pero necesita saber programar.
La eleccion de herramientas depende de las capacidades del usuario. Darle una computadora a alguien que no sabe programar es peor que darle una calculadora. Darle una calculadora a un programador limita su potencial.
Con los agents sucede lo mismo. La pregunta no es "cual herramienta es la mas poderosa", sino "cual herramienta se ajusta mejor a las capacidades actuales del modelo". Lo que comparte este articulo es precisamente la experiencia del equipo de Claude Code tropezando mientras buscaba ese punto de equilibrio.
A continuacion presento cuatro temas progresivos que he destilado del texto original y de las discusiones de la comunidad.
***
## Menos es mas — la contraintuicion sobre la cantidad de herramientas
La intuicion nos dice que cuantas mas herramientas tenga un agent, mas capaz sera. Pero la experiencia del equipo de Claude Code es exactamente la opuesta.
Claude Code actualmente tiene solo alrededor de 20 herramientas, y el equipo revisa constantemente si realmente necesita todas. El umbral para agregar una nueva herramienta es alto, porque significa darle al modelo una opcion mas que considerar — cada herramienta adicional agrega una "carga cognitiva" mas al proceso de toma de decisiones del modelo.
Este hallazgo genero una fuerte resonancia en la comunidad. leon.M compartio su experiencia practica:
La investigacion CodeAct de Apple proporciono respaldo cuantitativo a esta opinion: **una unica primitiva de ejecucion de codigo (code execution primitive) supera a conjuntos complejos de herramientas especializadas en hasta un 20% en tareas complejas**. Menos, efectivamente, puede ser mas.
Las tres iteraciones de la herramienta AskUserQuestion son el mejor ejemplo de este principio. El equipo de Claude Code queria mejorar la capacidad de Claude para hacer preguntas al usuario — aunque Claude podia preguntar directamente en texto plano, responder esas preguntas se sentia demasiado laborioso. Como reducir la friccion.
**Primer intento**: agregarle un parametro a ExitPlanTool para que al generar un plan incluyera un conjunto de preguntas. Resultado — Claude se confundio. Pedirle que generara un plan y al mismo tiempo preguntas sobre el plan creaba un conflicto: que pasa si la respuesta del usuario contradice el plan.
**Segundo intento**: modificar las instrucciones de salida para que Claude usara un formato markdown especifico para preguntar, y luego el frontend parseara y formateara. Resultado — inestable. Claude agregaba oraciones adicionales, omitia opciones o simplemente usaba un formato completamente diferente.
**Tercer intento**: crear una herramienta independiente llamada AskUserQuestion. Claude puede invocarla en cualquier momento; al hacerlo, aparece una ventana emergente con la pregunta que bloquea el ciclo del agent hasta que el usuario responda. **Funciono.**
Thariq escribio una frase muy interesante en el texto original:
> Claude parece disfrutar invocando esta herramienta, y descubrimos que la calidad de su salida es muy buena. Incluso una herramienta perfectamente disenada no funcionara si Claude no entiende como invocarla.
phuong hizo una pregunta perspicaz: "'Claude parece disfrutar invocando esta herramienta' es la frase mas interesante y menos explicada aqui. Como detectan la 'afinidad' del modelo hacia una herramienta — leyendo transcripts o mediante metricas internas de frecuencia de invocacion? Si esta heuristica pudiera ser mas especifica, cambiaria la forma en que todos disenan herramientas para agents."
La pregunta no obtuvo respuesta, pero apunta a una intuicion importante: **el criterio de exito del diseno de herramientas no es "que un humano lo considere razonable", sino "que el modelo entienda como usarla y quiera usarla"**.
Emeka complemento con una leccion directa desde la perspectiva del desarrollo de herramientas empresariales: "Una vez intente controlar cada posible entrada y salida al construir herramientas para un agent, y el resultado fue que el modelo simplemente... las ignoro. Ahorre esfuerzo de ingenieria y confie en que el modelo puede manejar la ambiguedad."
***
## Las herramientas se vuelven obsoletas — el efecto cadena tras la mejora de capacidades
Si "menos es mas" es una leccion sobre la dimension espacial, "las herramientas se vuelven obsoletas" es una leccion sobre la dimension temporal.
**Las herramientas que una vez ayudaron al modelo pueden convertirse en limitaciones a medida que el modelo progresa.**
Cuando Claude Code se lanzo inicialmente, el equipo se dio cuenta de que el modelo necesitaba una lista de tareas pendientes para mantener el rumbo — podia escribir elementos al inicio y luego ir marcandolos como completados. Para esto le proporcionaron a Claude la herramienta TodoWrite. Pero aun asi, Claude frecuentemente olvidaba lo que debia hacer.
La solucion del equipo fue insertar un recordatorio del sistema cada 5 turnos de conversacion, indicandole a Claude cuales eran sus objetivos.
Pero con las mejoras del modelo, el problema se invirtio: el modelo ya no necesitaba que le recordaran la lista de tareas pendientes; por el contrario, sentia esos recordatorios como una restriccion. **Ser recordado repetidamente de la lista de tareas hacia que Claude sintiera que debia seguirla estrictamente, en lugar de ajustar de manera flexible segun las necesidades.** Al mismo tiempo, Opus 4.5 mejoro significativamente en el uso de sub-agents, pero como deberian coordinarse los sub-agents para compartir una lista de tareas pendientes.
Asi que el equipo reemplazo TodoWrite con Task Tool. La diferencia entre ambos es fundamental: los Todos servian para mantener al modelo en el camino correcto — como un jefe supervisando la lista de tareas de un empleado; los Tasks, en cambio, se enfocan mas en facilitar la comunicacion entre agents — como un tablero de colaboracion de equipo. Los Tasks soportan dependencias, comparten actualizaciones entre sub-agents, y el modelo puede modificarlos y eliminarlos.
**De TodoWrite a recordatorios cada 5 turnos a Task Tool, fueron tres redisenos.** No porque los disenos anteriores fueran "incorrectos", sino porque el modelo habia crecido.
El comentario de modi resume este fenomeno con precision:
Esto tambien da lugar a una controversia interesante. Danny Cosson considero que AskUserQuestion "en realidad es un mal diseno" — fuerza el modelo elegante de entrada de texto/salida de texto de Claude Code dentro de un patron de interaccion especifico, practicamente sin beneficios.
Pero la refutacion de @Toong fue muy convincente: el valor de AskUserQuestion radica en dos puntos — **iniciativa** (indica explicitamente al modelo que tiene derecho a preguntar) y **gestion de estado** (distingue claramente la salida de texto estandar del estado de bloqueo de "esperando intervencion del usuario").
La misma herramienta, dos evaluaciones diametralmente opuestas. Esto demuestra precisamente lo que Thariq dijo al final: **el diseno de herramientas es un arte, no una ciencia**. Depende del modelo que se utilice, del objetivo del agent y del entorno en el que opera.
***
## Dejar que el Agent encuentre sus propias respuestas — de "alimentar" a "busqueda autonoma"
Esta es la parte del articulo que considero de mayor valor practico.
Claude Code inicialmente utilizaba una base de datos vectorial RAG para buscar contexto para Claude. RAG era potente y rapido, pero tenia dos problemas: primero, requeria indexacion y configuracion, lo cual podia ser fragil en diferentes entornos; segundo, y mas fundamental — **este enfoque proporcionaba contexto a Claude, en lugar de dejarlo buscarlo por si mismo**.
El equipo realizo un cambio clave: si Claude puede buscar en la web, por que no puede buscar en su base de codigo. Al proporcionarle a Claude la herramienta Grep, le permitieron buscar archivos y construir contexto por su cuenta.
El comentario de Beacon fue directo al grano:
**En el transcurso de un ano, Claude paso de ser casi incapaz de construir contexto de forma autonoma a poder realizar busquedas anidadas a traves de multiples capas de archivos, encontrando con precision el contexto necesario.** La clave de esta evolucion no fue darle mas informacion a Claude, sino mejores capacidades de busqueda.
Cuando Claude Code introdujo Agent Skills, el equipo propuso formalmente el concepto de **divulgacion progresiva (Progressive Disclosure)**: permitir que el agent descubra gradualmente el contexto relevante a traves de la exploracion.
La implementacion concreta es elegante: Claude puede leer archivos de skill, y estos archivos a su vez pueden referenciar otros archivos que el modelo puede leer recursivamente. Un uso comun de los skills es agregar mas capacidades de busqueda a Claude — por ejemplo, proporcionarle instrucciones sobre como usar una API o consultar una base de datos.
Brian Wagner compartio su practica de tres capas, completamente alineada con el enfoque de Claude Code:
> SKILL.md se mantiene conciso (unas 100 lineas), el contexto pesado se coloca en archivos que Claude descubre solo cuando lo necesita. Yo lo llamo la tercera capa. Ustedes lo llaman divulgacion progresiva. La esencia es la misma.
PrimeLine incluso construyo un sistema cuantificado de tres capas: configuracion JSON (aprox. 500 tokens) > resumen de esquema (aprox. 300 tokens) > markdown completo (aprox. 3K tokens). Un enrutador de contexto decide que capa cargar segun las palabras clave de la tarea.
La logica central de esta estrategia por capas es: **el contexto es un recurso limitado con rendimientos marginales decrecientes**. Entregar toda la informacion al agent de una sola vez no solo desperdicia tokens, sino que diluye la informacion verdaderamente importante. Proporcionarla bajo demanda, dejando que el agent decida cuando necesita detalles mas profundos — esa es la solucion escalable.
El sub-agent Claude Code Guide es otra aplicacion ingeniosa de la divulgacion progresiva. El equipo noto que Claude no sabia lo suficiente sobre como usar el propio Claude Code — si usted le preguntaba como agregar un MCP o que hacia cierto comando slash, no podia responder.
Podrian haber metido toda la informacion en el prompt del sistema, pero los usuarios rara vez hacen este tipo de preguntas, y hacerlo habria aumentado la erosion del contexto, interfiriendo con el trabajo principal de Claude Code: escribir codigo.
Primero intentaron darle a Claude un enlace a la documentacion para que buscara por si mismo — funciono, pero Claude cargaba una gran cantidad de resultados en el contexto para encontrar la respuesta correcta. La solucion final fue construir un sub-agent dedicado: Claude Code Guide. Este sub-agent tiene instrucciones detalladas de busqueda y sabe como buscar eficientemente en la documentacion y que contenido devolver. **Sin agregar ninguna herramienta nueva, se expandio el espacio de accion de Claude.**
Lance Martin, en su articulo sobre patrones de diseno de agents, propuso una perspectiva complementaria: **en lugar de definir docenas de herramientas para un agent, es mejor darle una computadora y dejar que use codigo para orquestar herramientas.** La abstraccion central de Claude Code es el CLI — el agent vive en su computadora y completa tareas complejas a traves de primitivas basicas como bash y el sistema de archivos. Unas pocas herramientas atomicas (como la herramienta bash) son mas flexibles y consumen menos tokens que un conjunto extenso de herramientas.
***
## Empatia con el modelo — pensar como un Agent
Los tres temas anteriores — menos es mas, las herramientas se vuelven obsoletas, dejar que el agent encuentre sus propias respuestas — comparten una meta-metodologia comun. Thariq lo senalo al inicio del articulo: **ver el mundo como un agent.**
No es un conjunto de reglas, sino una forma de pensar. David Zhang le dio un nombre:
**Empatia con el modelo** — no es disenar herramientas "razonables" desde la perspectiva humana, sino pensar desde la perspectiva del modelo: que es lo que realmente ve, como lo entendera, como lo usara.
Terminally Drifting destilo en tres pasos la misma leccion que todo equipo de agents termina aprendiendo:
> 1. Disenar herramientas para humanos
> 2. El modelo las usa como un mapache con permisos de administrador
> 3. Redisenar para la economia de tokens y efectos secundarios predecibles
"Pensar como un agent" es ese avance clave.
Vish compartio su momento de revelacion: "Estabamos construyendo interfaces de herramientas que tenian sentido para los humanos, y luego no entendiamos por que el agent siempre tomaba decisiones extranas. **Una vez que se invierte el modelo mental y se piensa en lo que el modelo realmente ve en la definicion de la herramienta, todo cambia.**"
Emeka dijo lo mismo desde la perspectiva del desarrollo de herramientas empresariales: "Todo esta configurado por defecto con el modelo mental humano — schemas, nombres de campos, flujos, todo optimizado para la lectura humana. Pero para un agent, ese es el marco equivocado."
Esta "inversion del modelo mental" suena simple, pero requiere practica constante. Es necesario leer cuidadosamente la salida del agent — no para ver que hizo bien, sino para entender por que tomo ciertas decisiones, donde dudo, donde tomo un camino equivocado. Estos "comportamientos anomalos" a menudo no son bugs del modelo, sino bugs del diseno de herramientas.
En las discusiones de la comunidad tambien surgieron varias perspectivas complementarias dignas de reflexion.
OAIR senalo un punto ciego: **todas estas iteraciones de herramientas asumen que el agent no tiene estado — cada sesion comienza desde cero. Y si la "herramienta" mas importante no esta en el espacio de accion, sino en la memoria persistente sobre como funciona esta base de codigo?** Claude Code respondio parcialmente a esta cuestion mas adelante con los archivos CLAUDE.md y el sistema de memoria, pero la gestion de estado persistente sigue siendo un desafio abierto en el diseno de agents.
Clinker propuso otro principio de diseno: **optimizar la recuperabilidad (recoverability) en lugar de la capacidad bruta — precondiciones explicitas en las herramientas, estado observable y reintentos de bajo costo suelen ser mas efectivos que simplemente agregar mas herramientas rapido.** Esto esta en linea con la filosofia de ingenieria de software de "hacer que el sistema sea tolerante a fallos en lugar de infalible".
El resumen de Paradigm Collapse fue el mas incisivo:
> El diseno del Action Space es fundamentalmente diseno de poder — los permisos que usted le otorga a la IA determinan el rol que asume. Es exactamente igual que gestionar un equipo: usted cree que el cuello de botella es la capacidad de las personas, pero en realidad el cuello de botella son los limites de permisos que usted ha trazado.
***
## Reflexiones finales
Volviendo a las palabras finales de Thariq:
> Experimente mas, lea sus salidas, pruebe nuevos enfoques. Observe como un agent.
Como usuario intensivo de Claude Code, la mayor impresion que me dejo este articulo es: **esas funcionalidades que parecen "naturales" tienen detras innumerables iteraciones de "esto no funciona, probemos otra cosa".** AskUserQuestion se intento tres veces, TodoWrite se rediseno tres veces, RAG fue reemplazado por Grep. Cada mejora no surgio porque se les ocurrio una solucion mas inteligente, sino porque observaron cuidadosamente el comportamiento real del modelo.
El subtitulo de este articulo es "Seeing like an Agent" — ver el mundo como un agent. Pero viendolo desde otro angulo, esta es en realidad la esencia de toda buena practica de ingenieria: **no disenar sistemas desde su propia perspectiva, sino desde la perspectiva del usuario.** Solo que esta vez, el usuario es un modelo de IA.
En el futuro, cada desarrollador que construya agents probablemente necesitara dominar lo que David Zhang llama "empatia con el modelo". No es ninguna capacidad misteriosa; su nucleo se reduce a tres cosas: **observar el comportamiento real del modelo, leer sus salidas y luego ajustar el diseno en funcion de lo que se observa.**
Observe como un agent.
***
**Lecturas complementarias:**
# De la guerra de chips a los centros de datos espaciales: la proxima decada de la industria de IA
> Invertir es la busqueda de la verdad. Si usted encuentra la verdad primero y su juicio es correcto, esa es la forma en que genera alpha. Y debe ser una verdad que los demas aun no han visto.
Estas palabras provienen de la entrevista de Gavin Baker, fundador de Atreides Management, en el podcast "Invest Like the Best" de Patrick O'Shaughnessy. Gavin es considerado uno de los inversores con mayor pasion y perspicacia en el sector tecnologico. Esta conversacion de casi dos horas cubrio GPU, TPU, la economia de la IA, centros de datos espaciales, el futuro del SaaS e incluso su transicion personal de instructor de esqui a inversor.
Esta entrevista tiene una densidad de informacion muy alta, con demasiados momentos que hacen que uno "se detenga a pensar". A continuacion, algunos de los puntos de vista mas dignos de reflexion.
***
## Como seguir el desarrollo de la IA? Comience gastando 200 dolares
Al inicio de la entrevista, Patrick hizo una pregunta muy practica: cuando se lanza un nuevo modelo como Gemini 3, como procesa usted esa informacion?
La respuesta de Gavin fue directa: **usted tiene que usarlo personalmente**.
Pero lo clave no es "usarlo", sino que version usar. Le sorprendio que algunos inversores usaran la version gratuita de la IA y concluyeran que "la IA no es gran cosa":
> La version gratuita es como si estuviera interactuando con un nino de 10 anos, y luego basandose en el desempeno de ese nino de 10 anos, intenta predecir como sera a los 35. Puede pagar, de hecho, debe pagar para obtener la membresia de nivel mas alto, 200 dolares al mes. Esos son los verdaderos adultos de 30 a 35 anos.
Esta analogia es muy precisa. La diferencia entre los modelos de IA nacionales e internacionales funciona de manera similar: muchas personas prueban un modelo gratuito local y concluyen que "la IA no es gran cosa", pero si ha usado modelos de primer nivel como Claude 4.5 Opus, Gemini 3 Pro o GPT-5.2 Reasoning, la experiencia es completamente diferente.
Hablando de pagar, personalmente gasto alrededor de 2000 yuanes al mes suscribiendome a diversos productos de IA, siendo la mayor parte el paquete Claude Code Max de 250 dolares. Si usted es desarrollador y tiene necesidades de programacion significativas, le recomiendo encarecidamente suscribirse directamente al Claude Code oficial (desde 125 dolares), en lugar de usar sitios espejo. Por un lado, no sabe si esos sitios espejo realmente estan usando los modelos autenticos; por otro lado, el paquete Max tiene una relacion costo-beneficio muy buena, mucho mas economico que pagar por uso.
En cuanto a los canales para obtener informacion, la respuesta de Gavin podria sorprender a muchos: **X (Twitter)**.
Dijo que el desarrollo de la IA en gran medida "esta ocurriendo en tiempo real en la plataforma X". Las personas en el planeta que realmente entienden la frontera de la IA son aproximadamente entre 500 y 1000, una parte considerable esta en China, y es necesario seguirlas de cerca. Menciono especialmente a Andrej Karpathy:
> Cada articulo que escribe Andrej Karpathy, usted tiene que leerlo tres veces. Como minimo.
Como alguien que tambien sigue las novedades de la IA, me identifico profundamente con esto. Las discusiones sobre IA en Twitter son definitivamente mas inmediatas y profundas que cualquier medio de comunicacion. Los investigadores de los laboratorios publican directamente sobre los ultimos avances, e incluso "discuten acaloradamente" entre si. Gavin menciono que el equipo de PyTorch de Meta y el equipo de Jax de Google tuvieron una disputa publica en X, y finalmente los directores de ambos laboratorios tuvieron que intervenir diciendo: "Nuestra gente no puede hablar mal del laboratorio del otro."
***
## Scaling Laws: nuestro "momento del antiguo Egipto"
Tras el lanzamiento de Gemini 3, muchas personas se preguntaron que revelaba sobre las scaling laws (leyes de escala). Gavin ofrecio una perspectiva que nunca habia escuchado:
> Nuestra comprension de las scaling laws de preentrenamiento probablemente sea como la comprension que tenian los antiguos egipcios del sol. Podian medirlo con gran precision, tan precisamente que el eje este-oeste de la Gran Piramide se alinea perfectamente con los equinoccios de primavera y otono, al igual que Stonehenge. Mediciones perfectas. Pero no entendian la mecanica orbital. No sabian por que el sol sale por el este y se pone por el oeste.
Esta analogia me hizo detenerme a pensar por un buen rato. Efectivamente, podemos predecir con gran precision: si aumentamos la capacidad de computo del modelo 10 veces, cuanto mejorara el rendimiento. Pero no sabemos por que. No es una "ley", sino una "observacion empirica", una observacion empirica que medimos con extrema precision pero cuyo principio subyacente no comprendemos.
Entonces, por que es importante Gemini 3? Porque demostro que esta "observacion empirica" sigue siendo valida. En un momento en que los chips Blackwell estaban retrasados y todos se preocupaban por si "las scaling laws ya habian dejado de funcionar", Gemini 3 dio una respuesta clara: **no han dejado de funcionar**.
Pero lo mas interesante es lo que Gavin dijo a continuacion: si no hubieran aparecido los "modelos de razonamiento" (reasoning models), el desarrollo de la IA entre 2024 y 2025 deberia haberse estancado.
Por que? Porque despues de que XAI logro hacer funcionar 200,000 GPU Hopper en sincronizacion, el siguiente paso requeria esperar los chips Blackwell. No se puede mantener "coherentes" mas de 200,000 Hopper, es decir, que funcionen como un todo unificado. Y Blackwell se retraso.
> Si no hubieran existido los modelos de razonamiento, desde mediados de 2024 hasta ahora, la IA no habria tenido ningun progreso. Todo se habria estancado. Puede imaginarse lo que eso significaria para el mercado? Estariamos viviendo en un entorno completamente diferente. Los modelos de razonamiento, en cierto sentido, salvaron a la IA, porque le permitieron seguir avanzando sin Blackwell.
Esta es una perspectiva que no habia considerado antes: los modelos de razonamiento (como o1) no son solo una nueva capacidad, sino que en realidad "salvaron" el ritmo de desarrollo de toda la industria de la IA.
***
## La guerra de chips: Google esta "absorbiendo el oxigeno"
Al hablar sobre la competencia entre GPU y TPU, Gavin dijo algo que me impresiono mucho:
> Google es actualmente el productor de tokens de menor costo. Lo que ha estado haciendo, yo lo describiria como "absorber el oxigeno economico del ecosistema de IA", lo cual es una estrategia extremadamente racional para ellos.
Como productor de bajo costo, Google ha estado ofreciendo servicios de IA a precios bajos (incluso con perdidas), haciendo la vida muy dificil a los competidores. Esta es una tactica clasica en la industria tecnologica, pero Gavin senalo un cambio interesante:
> La IA es la primera vez en mi carrera que veo que ser el "productor de menor costo" realmente importa en tecnologia. Apple no vale billones por producir telefonos a bajo costo. Microsoft no vale billones por producir software a bajo costo. NVIDIA no vale billones por producir aceleradores de IA a bajo costo. Eso nunca habia importado.
Pero en la era de la IA, cuando la electricidad se convierte en el factor limitante, **cuantos tokens se pueden producir por vatio** se vuelve crucial. Si usted puede producir de 3 a 5 veces mas tokens por vatio, eso equivale a 3-5 veces mas ingresos. El precio del computo se vuelve irrelevante, porque el cuello de botella es la electricidad.
Este panorama esta a punto de cambiar. Los chips Blackwell finalmente estan comenzando a desplegarse, y Gavin predice que el primer modelo Blackwell vendra de XAI:
> Segun Jensen, nadie construye centros de datos mas rapido que Elon. Jensen lo ha dicho publicamente.
Cuando Blackwell y los posteriores chips Ruben se desplieguen a gran escala, la ventaja de Google como productor de bajo costo desaparecera. En ese momento, estaran dispuestos a seguir operando su negocio de IA con un margen bruto de -30%? Ese calculo cambiara completamente.
***
## Centros de datos espaciales: una locura, pero correcto desde primeros principios
Cuando Patrick pregunto sobre "alguna idea loca que no se discute mucho", Gavin comenzo a hablar sobre los centros de datos espaciales. Al principio pense que estaba bromeando, pero despues de escuchar su analisis, me di cuenta de que esta podria ser la parte mas visionaria de toda la entrevista.
> Desde la perspectiva de primeros principios, los centros de datos espaciales son superiores a los centros de datos terrestres en todas las dimensiones.
Su argumento es el siguiente:
**1. Energia**: En el espacio, los satelites pueden estar expuestos a la luz solar las 24 horas, y la intensidad de la radiacion solar es 6 veces mayor que en la superficie terrestre. Ademas, como siempre hay luz solar, no se necesitan baterias, que representan una parte significativa del costo. Asi que la energia de menor costo en el sistema solar es la "energia solar espacial".
**2. Refrigeracion**: En los centros de datos terrestres, la mayor parte del costo y el peso se destina a la refrigeracion. Pero en el espacio? La refrigeracion es gratuita. Se colocan los radiadores en el lado sombreado del satelite, donde la temperatura se acerca al cero absoluto.
**3. Red**: En un centro de datos, los racks se conectan con fibra optica, que esencialmente es laser viajando a traves de cables. Lo unico mas rapido que eso es laser viajando a traves del vacio. Asi que si se conectan los satelites en el espacio con laser, la red es en realidad mas rapida que en los centros de datos terrestres.
**4. Experiencia del usuario**: Actualmente, cuando le hace una pregunta a la IA, la senal viaja del telefono a la antena, a la fibra optica, a algun centro de datos, se procesa y regresa por el mismo camino. Pero si un satelite puede comunicarse directamente con el telefono (Starlink ya ha demostrado esta capacidad de conexion directa), toda la cadena seria mucho mas corta.
Por supuesto, esto requiere lanzamientos masivos de Starship para hacerse realidad, lo cual podria tomar 5-6 anos mas. Pero Gavin senalo una convergencia interesante: Tesla, SpaceX y XAI estan confluyendo. XAI sera el "modulo inteligente" de los robots Optimus, SpaceX construira centros de datos en el espacio para proporcionar potencia de computo a la IA, y estas tres empresas estan formando un ciclo de ventajas competitivas que se refuerzan mutuamente.
***
## La "plataforma en llamas" del SaaS
Si el contenido anterior lo entusiasmo sobre el futuro de la IA, esta seccion podria preocuparle sobre el futuro de muchas empresas existentes.
Gavin lo dijo directamente: **las empresas de SaaS de aplicaciones estan cometiendo exactamente los mismos errores que los minoristas fisicos cometieron frente al comercio electronico**.
Los minoristas fisicos miraban a Amazon y pensaban "el comercio electronico es un negocio de bajo margen, como podria ser mas eficiente que nosotros? Los clientes pagan para venir a nuestra tienda y se llevan los productos a casa ellos mismos." Veian claramente las necesidades del cliente, pero se negaban a invertir porque no les gustaba la estructura de margenes del comercio electronico. El resultado? El margen de beneficio del negocio minorista de Amazon en Norteamerica es ahora mayor que el de muchos minoristas tradicionales.
Las empresas de SaaS enfrentan la misma situacion ahora. El software tradicional se escribe una vez y se puede replicar y distribuir infinitamente, con margenes brutos de 80-90%. Pero la IA es diferente: cada uso requiere nuevo computo, y una buena empresa de IA podria tener margenes brutos de solo el 40%.
> Si quiere crear un agente de IA y no esta dispuesto a operar con un margen bruto inferior al 35%, nunca tendra exito. Porque las empresas nativas de IA operan con ese margen de beneficio. Si quiere proteger un margen bruto del 80%, esta garantizando que no tendra exito en IA. Absolutamente garantizado.
Gavin dice que esta es una "decision de vida o muerte", y **excepto Microsoft, casi todas las empresas estan fracasando**.
Cito el famoso memorando de la "plataforma en llamas" de Nokia: su plataforma esta en llamas. Pero al lado hay una nueva plataforma muy buena, puede saltar a ella y luego regresar a apagar el fuego en la plataforma original. Ahora tiene dos plataformas.
Salesforce, ServiceNow, HubSpot, GitLab, Atlassian: el cree que todas estas empresas pueden y deben ejecutar esta estrategia: hacer publicos sus ingresos de IA, hacer publico su margen bruto de IA (un margen bruto bajo precisamente demuestra que es "IA real"), y luego senalar a los competidores respaldados por capital de riesgo que aun estan perdiendo dinero, diciendo "yo tengo algo que ellos no tienen: un negocio que genera flujo de caja."
***
## La historia de crecimiento de un inversor
Al final de la entrevista, Patrick hizo una pregunta mas personal: como le presentaria a un joven lo que usted hace?
La respuesta de Gavin comenzo con "invertir es la busqueda de la verdad", pero lo verdaderamente interesante es su historia de vida.
Su plan original era: ser instructor de esqui en invierno, guia de rafting en verano, escalar durante la temporada baja, y de paso intentar escribir novelas y hacer fotografia de vida silvestre. Ese era su "plan de vida" en la universidad, y sus padres lo apoyaban completamente.
Pero sus padres le pusieron una pequena condicion: podria hacer una pasantia profesional, solo una, en lo que fuera?
La unica pasantia que pudo encontrar fue en el departamento de gestion de patrimonio privado de una casa de bolsa. El trabajo era simple: cada vez que la empresa publicaba un informe de investigacion, tenia que verificar cuales clientes poseian esas acciones y enviarles el informe.
Entonces comenzo a leer esos informes.
> Pense: "Dios mio, esto es lo mas interesante que puedo imaginar."
Entendio la inversion como un "juego donde coexisten la habilidad y la suerte", algo parecido al poker. Se puede perder por mala suerte, por ejemplo si un meteorito golpea la sede de la empresa en la que invirtio, pero la mayor parte del tiempo, la habilidad es lo que importa. Y la forma de obtener ventaja es tener el conocimiento historico mas profundo, combinado con la comprension mas precisa del mundo actual, para formar un juicio diferenciado sobre "que va a pasar a continuacion".
Eso fue en su tercer dia de pasantia. Fue a una libreria y compro los libros de Peter Lynch, los leyo en dos dias. Luego leyo a Buffett, leyo "Market Wizards", leyo las cartas de Buffett a los accionistas, dos veces. Luego aprendio contabilidad por su cuenta. Al volver a la universidad, cambio su carrera de ingles e historia a historia y economia.
Tambien menciono una experiencia como personal de limpieza. Mientras trabajaba en la estacion de esqui de Alta, hacia limpieza de habitaciones. Una vez, mientras limpiaba una habitacion, vio que el huesped estaba leyendo el mismo libro que el, y le dijo "es un gran libro, estoy mas o menos en la misma parte que usted." La persona lo miro como si fuera un extraterrestre, y luego, aun mas sorprendida, le pregunto: "Usted lee libros?"
> Esa experiencia cambio permanentemente la forma en que trato a las demas personas.
***
## Reflexion final: la IA obtiene lo que necesita
Hacia el final de la entrevista, Gavin dijo algo que me parecio lo mas fascinante:
> En los ultimos dos anos, todo lo que la IA ha necesitado para seguir desarrollandose, lo ha obtenido. Ha visto alguna vez que la opinion publica estadounidense cambie tan rapido sobre algun tema como lo hizo con la energia nuclear? Simplemente sucedio. Y sucedio justo cuando la IA lo necesitaba. Ahora que enfrentamos limitaciones de energia electrica en la Tierra, de repente aparece la discusion sobre centros de datos espaciales. Cada vez que algun cuello de botella podria desacelerar el desarrollo de la IA, todo se acelera en su lugar.
Esto recuerda el concepto de "technium" (el tecnium) que Kevin Kelly propuso en "What Technology Wants": la tecnologia como un todo parece tener una especie de voluntad propia, queriendo volverse cada vez mas poderosa.
Quiza sea solo coincidencia. Quiza sea simplemente mucha gente inteligente resolviendo problemas. Pero el patron que Gavin observa, que cada obstaculo que enfrenta la IA es de alguna manera eliminado, ciertamente merece ser reflexionado.
# Selecciones Curadas
# Curations - Selecciones Curadas
Aquí se recopilan mis reflexiones y resúmenes tras ver videos de expertos en tecnología y leer blogs de calidad.
No son simples notas de extracto, sino que incorporan mi propia comprensión y experiencia práctica.
## Fuentes de contenido
* Análisis de videos técnicos
* Selección de artículos de blogs
* Resúmenes de podcasts/entrevistas
## Contenido más reciente
### Pensar como un Agent: La filosofía de diseño de herramientas del equipo de Claude Code
El ingeniero de Anthropic, Thariq, compartió su experiencia en el diseño de herramientas para agentes durante la construcción de Claude Code — desde las tres iteraciones de AskUserQuestion hasta las tres refactorizaciones de TodoWrite, desde RAG hasta la divulgación progresiva. Cada caso apunta a la misma metodología central: ver el mundo como un Agent.
[Leer artículo completo →](./claude-code-seeing-like-an-agent)
### El 90% de los universitarios usa IA, pero nadie sabe cuáles son las reglas
Cuatro estudiantes de Princeton, Berkeley y LSE hablaron sobre el verdadero estado de la IA en el campus — trampas, confusión, polarización, y esa pregunta que nadie se atreve a hacer: ¿para qué sirve todavía la universidad? No es un video promocional, sino confusión y reflexión genuinas.
[Leer artículo completo →](./ai-on-campus-student-perspectives)
### De la guerra de chips a los centros de datos espaciales: La próxima década de la industria de IA
De la guerra de chips a los centros de datos espaciales, de la encrucijada del SaaS a la esencia de la inversión. Gavin Baker compartió en el podcast "Invest Like the Best" sus perspectivas más profundas sobre la industria de IA: por qué debes usar la versión de pago de la IA, por qué las Scaling Laws son como "los antiguos egipcios entendiendo el sol", y cómo los modelos de razonamiento "salvaron" el ritmo de desarrollo de toda la industria de IA.
[Leer artículo completo →](./gavin-baker-ai-economics)
# El tweet de Karpathy explotó a 62 mil estrellas: ¿Qué hizo exactamente andrej-karpathy-skills?
El 27 de enero de 2026, Andrej Karpathy publicó un tweet muy largo en X, un ensayo de programación de 11 secciones y aproximadamente 1400 palabras, que registra los obstáculos que encontró durante la transición de "80% escritura a mano + 20% agente" en noviembre a "80% agente + 20% pulido" en diciembre. El tweet finalmente alcanzó **7,69 millones de visitas, 39.000 me gusta y 36.000 marcadores**.
Tres meses después, se lanzó un repositorio de GitHub llamado `forrestchang/andrej-karpathy-skills`, que empaquetaba el tweet de Karpathy en cuatro reglas instalables. En dos semanas, alcanzó **62,7 mil estrellas y 5,5 mil bifurcaciones**, convirtiéndose en el número 1 en la lista semanal de GitHub en abril de 2026.
Ontología de almacén: **un archivo Markdown**.
***
## 1. ¿De qué se queja Karpathy?
El extenso artículo de Karpathy enumera esencialmente cuatro "condiciones crónicas" codificadas por LLM.
**La primera enfermedad: hacer suposiciones en secreto**
> "El tipo de error más común es que los modelos hacen suposiciones incorrectas para usted y luego las aplican sin validarlas. No manejan su propia confusión, no buscan aclaraciones, no muestran inconsistencias, no presentan compensaciones, no refutan cuando llega el momento de refutar y son un poco demasiado halagadores".
Se trata de un "error de colusión". Usted dice "Agregar un inicio de sesión para mí", no pregunta qué autenticación usar, si recordar el dispositivo o cómo administrar la sesión; simplemente presenta una solución que considera razonable. Para cuando termines de revisar y descubras que es diferente de lo que quieres, ya tendrá 500 líneas escritas.
**La segunda enfermedad: el exceso de ingeniería**
> "En particular, les gusta complicar demasiado su código y sus API, inflar las capas de abstracción y no limpiar el código muerto. Usarán 1000 líneas de código para implementar una estructura ineficiente, inflada y frágil, y tienes que persuadir como un niño y decir: 'Bueno, ¿por qué no haces esto?', antes de que digan: '¡Por supuesto!' y luego redúzcalo inmediatamente a 100 líneas."
Este es el efecto de observador más típico de la codificación LLM: si bien se le da un contexto "generosamente", recompensará "generosamente" la complejidad\*\*. Modo de estrategia, modo de fábrica, inyección de dependencia: todo se le proporciona.
**La tercera enfermedad: cambiar algo que no pediste cambiar**
> "A veces modifican o eliminan algunos comentarios y códigos porque no les gusta o no los entienden completamente, incluso si estos cambios no tienen nada que ver con la tarea actual".
Le pide que corrija un error y elimina convenientemente el comentario TODO inacabado que aparece junto a él con el argumento de que "parece que ya no es necesario".
**Enfermedad 4: Incluso si escribes las reglas en CLAUDE.md, aún así se romperá**
> "El problema anterior todavía existe incluso si hice algunos intentos simples de reparación en CLAUDE.md."
Esta es la frase más desgarradora de todo el tweet. Karpathy, ex miembro del equipo fundador de OpenAI y director de IA de Tesla, no pudo escribir CLAUDE.md que mantendría a Claude completamente a raya.
***
## 2. Solución a `andrej-karpathy-skills`
`forrestchang` Sistematizar las soluciones a estas cuatro enfermedades en cuatro principios y empaquetarlos en un archivo CLAUDE.md.
| Principios | Enfermedades correspondientes | Acciones principales |
| --------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Piense antes de codificar** | Asumiendo en secreto | Exprese sus suposiciones con claridad, enumere múltiples interpretaciones, deténgase y pregunte si está confundido y refute cuando sea necesario |
| **La simplicidad es lo primero** | Sobreingeniería | Escriba sólo el código mínimo requerido; no escriba flexibilidad especulativa, manejo de errores, abstracción |
| **Cambios quirúrgicos** | Cambios no autorizados | Toca sólo lo necesario; no refactorice ni cambie el estilo; solo informa otros códigos muertos pero no los eliminas |
| **Ejecución basada en objetivos** | Desalineación del método | Proporcione criterios de éxito verificables + pruebas y deje que el modelo pase por sí solo |
El cuarto principio es una referencia directa a otra frase famosa del tweet de Karpathy: "Apalancamiento":
Este es el punto de apoyo de toda la metodología del proyecto: \*\*Los primeros tres principios evitan que el LLM pierda el tiempo; el cuarto principio le dice cómo aprovechar verdaderamente sus fortalezas. \*\*
***
## 3. Cómo instalar y usar
**Método A: como complemento de Claude Code** (se recomienda utilizarlo primero)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
Después de la instalación, el diálogo del Código Claude de todos los proyectos cumplirá automáticamente con estos cuatro principios. Tiene efecto global y puede desactivarse en cualquier momento mediante `/plugin`.
**Método B: copiar manualmente CLAUDE.md**
Ingrese al repositorio → abra `CLAUDE.md` → copie → pegue en `CLAUDE.md` en el directorio raíz de su proyecto. Sólo efectivo para este proyecto.
El repositorio también proporciona adaptaciones `CURSOR.md` y `.cursor/rules/` adicionales: un conjunto de contenido que cubre los principales IDE de IA.
***
## 4. ¿Por qué puede alcanzar las 62k estrellas?
Este es un fenómeno que vale la pena analizar. 62,7 mil estrellas es un número exagerado para un "repositorio de un solo archivo"; en comparación, el descuento de Microsoft (9 mil) y las habilidades del agente de Addy Osmani (4,6 mil) en el mismo período combinados no son tantas.
Desglosado por peso de impacto:
**1. Respaldo de IP de Karpathy**: el mismo contenido no puede exceder los 10 000 si se llama `forrestchang-skills`. Karpathy viene con el capital cultural del "ex equipo fundador de OpenAI + director de IA de Tesla + instructor de CS231n", y sus tweets vienen con una etiqueta de "lectura obligada".
**2. Momento perfecto**: Opus 4.7 se lanzó el 16 de abril y las quejas por exceso de ingeniería alcanzaron su punto máximo. El repositorio apareció justo cuando todos buscaban un antídoto para "evitar que Claude se volviera tan loco".
**3. Los puntos débiles son universales**: todos los usuarios de Claude Code/Cursor han superado estos cuatro obstáculos y la tasa de empatía se acerca al 100%.
**4. El umbral es extremadamente bajo**: 1 archivo o 2 líneas de comandos. El costo de una estrella es tan bajo que puede ignorarse. "Si no lo instalas, perderás".
**5. Fuerte sentido de verificabilidad**: los 4 principios son claros y fáciles de recordar, y fáciles de capturar y reenviar. A diferencia de la guía de proyectos de 1000 líneas, que es prohibitiva.
**6. README bilingüe** ——`README.zh.md` devora directamente el tráfico del círculo de IA chino, V2EX/instantáneamente/detona simultáneamente en Weibo.
**7. Promoción cruzada por parte del autor** - La frase en la columna superior *"Mira mi nuevo proyecto Multica"* dirige el tráfico a la plataforma del agente comercial del autor `multica-ai/multica`. \*\*Este repositorio es esencialmente la parte superior del embudo de adquisición de clientes de Multica. \*\*
**8. Meta fit**: los "errores de codificación LLM" que analiza son exactamente lo que todos los lectores experimentan cuando codifican con LLM. La lectura y el uso están integrados y la tasa de conversión es extremadamente alta.
En una palabra: lo que vende no es código ni herramientas, sino **empaquetar las emociones de Karpathy en reglas instalables**; este es el caso más típico de "el contenido es producto" en el círculo de programación de IA en 2026.
***
## 5. Mis sugerencias de uso
**Primero use el método A para instalar globalmente**. Vea si mejora su experiencia al escribir herramientas y scripts, especialmente cuando deja que Claude cambie el código de otras personas, si reduce el problema de realizar cambios aleatorios.
**Después de una semana o dos decidiremos si lo fusionaremos con el proyecto CLAUDE.md**. CLAUDE.md de cada proyecto ya está lleno de conocimiento del dominio (sistema de diseño, especificación de componentes, proceso de implementación), mientras que el conjunto de Karpathy es una metodología general. Los dos no entran en conflicto y pueden superponerse, pero el momento debe esperar hasta que esté realmente seguro de que es útil.
**Preste atención al costo**: Hará que Claude haga más preguntas, lo que resultará molesto para las personas que están acostumbradas a "generar en una oración"; puede que no haga la ligera limpieza que se debería hacer (demasiado estricta); se verá limitado por tareas de exploración muy vagas.
**Valor más profundo**: Le obliga a establecer claramente sus requisitos, lo que resulta ser el requisito previo para toda ingeniería de software de alta calidad.
**Seguimiento digno de mención**: el propio forrestchang también está promocionando Multica, una "plataforma de agentes administrados de código abierto" para producir el mecanismo de habilidades. Si este conjunto de cuatro principios finalmente se convierte en el estándar de facto, Multica será su vehículo comercial. Presta atención a esta línea.
***
## Recursos de referencia
# Indie Dev
Registro del proceso completo de desarrollo independiente — cada paso desde la preparación hasta la publicación.
# Bark
# Bark
Una herramienta que permite enviar notificaciones push personalizadas al iPhone mediante simples solicitudes HTTP. Gratuita, de código abierto y autohospedable.
## Por qué la recomiendo
* **API minimalista** - Un solo comando curl para enviar una notificación push, sin configuración compleja
* **Código abierto y gratuito** - Licencia MIT, completamente de código abierto, sin ningún costo
* **Control de privacidad** - Soporte para autoalojar el servidor, los datos de las notificaciones están completamente bajo tu control
## Casos de uso
* **Notificaciones de scripts** - Recibir alertas cuando se completan tareas de larga duración, como la finalización de copias de seguridad de datos
* **Monitoreo de servicios** - Recibir alertas inmediatas cuando el servidor presenta anomalías
* **Integración de automatización** - Resultados de compilación CI/CD, notificaciones de finalización de tareas programadas
## Inicio rápido
1. Descargar Bark desde el [App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865)
2. Abrir la aplicación y copiar tu dirección de notificación push
3. Enviar tu primera notificación push:
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
Notificación push con título:
```bash
curl https://api.day.app/YOUR_KEY/Título/Contenido
```
## Información del proyecto
* GitHub: [Finb/Bark](https://github.com/Finb/Bark)
* Stars: 7.2k+
* Licencia: MIT
# Toolkit
# Toolkit
Aquí se recopilan proyectos destacados de GitHub y software útil que he descubierto.
# Guía completa de Claude Agent Teams
## Introducción
Si ha utilizado los Subagent de Claude Code, probablemente considere que el desarrollo en paralelo ya es bastante poderoso. Sin embargo, los Subagent tienen una limitación: solo pueden reportar resultados al Agent principal, pero no pueden comunicarse entre sí.
Agent Teams cambia esto por completo. Imagine lo siguiente: un Agent se encarga de la auditoría de seguridad, otro de la optimización del rendimiento y un tercero de la cobertura de pruebas, y no solo pueden trabajar en paralelo, sino que también pueden dialogar directamente, desafiarse mutuamente y llegar a consensos. Este es el valor central de Agent Teams.
## Comprender Agent Teams
La arquitectura de Agent Teams se asemeja mucho a un equipo de desarrollo real:
```
┌─────────────────────────────────────────────────────────┐
│ Usted (usuario) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (Instancia principal de Claude, coordina) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│ Aud. segur. │◄─►│ Opt. rendim. │◄─►│ Cob. pruebas │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└──────────────┼──────────────┘
▼
┌─────────────────┐
│ Lista de tareas │
│ compartida │
└─────────────────┘
```
### Diferencias con Subagent
| Característica | Subagent | Agent Teams |
| -------------------------- | --------------------------------------------------------------- | --------------------------------------------------------- |
| **Contexto** | Contexto independiente, resultados devueltos al Agent principal | Contexto independiente, ejecución completamente autónoma |
| **Comunicación** | Solo puede reportar al Agent principal | Los Teammates pueden comunicarse directamente entre sí |
| **Coordinación de tareas** | El Agent principal gestiona todo el trabajo | Lista de tareas compartida, coordinación autónoma |
| **Casos de uso** | Tareas enfocadas donde solo se necesitan resultados | Trabajos complejos que requieren discusión y colaboración |
| **Costo en tokens** | Menor: resumen de resultados devuelto al contexto principal | Mayor: cada Teammate es una instancia independiente |
En pocas palabras: **Subagent es un contratista al que envía a ejecutar tareas; Agent Teams es un equipo de proyecto que colabora en la misma sala**.
### Por qué Agent Teams es efectivo
La idea clave: **la especialización genera enfoque**.
Cuando un único Agent maneja tareas complejas de múltiples pasos, el contexto crece constantemente y con frecuencia necesita reiniciarse con `/clear`. Agent Teams permite que cada Teammate mantenga un área de enfoque reducida, con un contexto limpio y un rendimiento más estable.
## Habilitar Agent Teams
Agent Teams es actualmente una funcionalidad experimental, deshabilitada por defecto. Es necesario habilitarla manualmente:
**Opción 1: Variable de entorno**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**Opción 2: settings.json (recomendado, efecto permanente)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## Uso principal
### Crear su primer Agent Team
Una vez habilitado, simplemente indique a Claude en lenguaje natural que cree un equipo:
```
Crea un agent team para revisar el PR #142.
Genera tres revisores:
- Uno enfocado en problemas de seguridad
- Uno que verifique el impacto en el rendimiento
- Uno que valide la cobertura de pruebas
Que cada uno revise y reporte sus hallazgos.
```
**Consejo sobre palabras clave**: Use "create an agent team" o "spawn an agent team". Si solo dice "spawn agents", podría confundir Subagent con Agent Teams.
### Modos de visualización
Agent Teams soporta dos modos de visualización:
| Modo | Descripción | Requisitos |
| --------------- | -------------------------------------------------------- | ------------------------- |
| **In-process** | Todos los Teammates se ejecutan en la terminal principal | Sin requisitos especiales |
| **Split panes** | Cada Teammate en un panel independiente | Requiere tmux o iTerm2 |
El valor predeterminado es `auto`: si se ejecuta dentro de tmux se usan split panes, de lo contrario se usa in-process.
**Configurar el modo de visualización**:
```json
{
"teammateMode": "in-process"
}
```
**Especificar para una sesión individual**:
```bash
claude --teammate-mode in-process
```
### Atajos de teclado
| Acción | Atajo |
| ----------------------------------------- | ------------ |
| Cambiar entre Teammates | `Shift+Down` |
| Volver al Teammate anterior | `Shift+Up` |
| Alternar visualización de lista de tareas | `Ctrl+T` |
| Interrumpir el Teammate actual | `Escape` |
| **Habilitar Delegate Mode** | `Shift+Tab` |
| Entrar a la sesión de un Teammate | `Enter` |
### Delegate Mode (importante)
Delegate Mode es una de las funcionalidades más importantes de Agent Teams:
| Modo | Comportamiento del Lead |
| ----------------- | -------------------------------------------------------------------------- |
| **Modo normal** | El Lead puede implementar tareas y escribir código por sí mismo |
| **Delegate Mode** | El Lead solo puede coordinar, no puede escribir código ni ejecutar pruebas |
**Por qué se necesita Delegate Mode**:
Sin esta restricción, el Lead tiende a "acaparar el trabajo": aunque hay tres Teammates esperando para trabajar, el Lead comienza a escribir código por su cuenta. Al activar Delegate Mode, el Lead se ve forzado a ser un gerente de proyecto puro, pudiendo únicamente gestionar tareas, comunicarse con los Teammates y revisar los resultados.
```
# Presione Shift+Tab inmediatamente después de iniciar el equipo
```
## Casos prácticos
### Caso 1: Revisión de código en paralelo
Un único revisor tiende a profundizar demasiado en cierto tipo de problemas. Al dividir las dimensiones de revisión en áreas independientes, la seguridad, el rendimiento y la cobertura de pruebas reciben la misma atención:
```
Crea un agent team para revisar este PR. Genera tres revisores:
- Revisor de seguridad: verifica autenticación, autorización, vulnerabilidades de inyección
- Revisor de rendimiento: analiza complejidad algorítmica, consultas a base de datos, estrategias de caché
- Revisor de pruebas: valida cobertura de pruebas, condiciones límite, manejo de errores
Que revisen individualmente y luego discutan entre sí los problemas encontrados.
```
### Caso 2: Depuración con hipótesis competitivas
Cuando la causa raíz no está clara, un único Agent tiende a encontrar una explicación aparentemente razonable y detenerse. Hacer que los Teammates se desafíen mutuamente evita este problema:
```
Los usuarios reportan que la aplicación se cierra después de enviar un mensaje,
en lugar de mantener la conexión.
Genera 5 agent teammates para investigar diferentes hipótesis. Que discutan entre sí,
intenten refutar las teorías de los demás, como en un debate científico.
Actualicen los hallazgos consensuados en el informe de investigación.
```
**Mecanismo clave**: estructura de debate. Múltiples investigadores independientes intentan activamente refutar las teorías de los demás; las hipótesis que sobreviven tienen más probabilidades de ser la verdadera causa raíz.
### Caso 3: Producción de contenido en masa
Esta es una aplicación típica para tareas no técnicas: transformar una entrada en múltiples salidas:
```
Crea un agent team para convertir este guion de video en contenido para cuatro plataformas:
- Autor de artículo para LinkedIn
- Autor de hilo para Twitter
- Autor de newsletter
- Autor de artículo de blog
Ubicación del guion: /content/scripts/video-20.md
```
Cada Teammate crea de forma independiente, pero manteniendo la coherencia del contenido.
### Caso 4: Clúster de verificación de calidad QA
Una verificación de calidad para un sitio web de blog, desplegando 5 Agents para probar diferentes aspectos en paralelo:
```
Crea un agent team para una verificación de calidad integral del blog:
- Agent 1: pruebas de páginas principales (inicio, acerca de, contacto)
- Agent 2: pruebas de páginas de artículos (renderizado, navegación, metadatos SEO)
- Agent 3: verificación de enlaces (internos, externos, enlaces rotos)
- Agent 4: validación SEO (títulos, descripciones, datos estructurados)
- Agent 5: pruebas de accesibilidad (etiquetas ARIA, contraste, navegación por teclado)
Genera un informe de problemas ordenado por prioridad.
```
**Resultado**: en cuestión de minutos se completa una verificación integral que normalmente requeriría ejecución secuencial manual. Cada Agent se enfoca en su propia área y al final se consolida en una lista de problemas ordenada por prioridad.
### Caso 5: Modo de discusión en múltiples rondas
Un patrón de prompt útil: hacer que los Teammates discutan como en una reunión:
```
Usa Agent Teams para crear 4 teammates que discutan [decisión técnica],
realizando 3 rondas de discusión. Que los teammates intercambien opiniones en cada ronda.
Uno de los teammates debe encargarse específicamente de la perspectiva Red Team, planteando críticas.
```
Este patrón es especialmente adecuado para escenarios que requieren evaluar múltiples perspectivas, como decisiones de arquitectura y selección de tecnología.
### Caso 6: Proyecto del compilador de C
Anthropic utilizó 16 Agents para construir desde cero un compilador de C capaz de compilar el kernel de Linux:
| Métrica | Datos |
| ------------------ | -------------------------------------------------- |
| Cantidad de Agents | 16 instancias en paralelo |
| Sesiones | \~2,000 sesiones de Claude Code |
| Costo | \~$20,000 |
| Líneas de código | 100,000 líneas |
| Uso de tokens | 2 mil millones de entrada + 140 millones de salida |
Resultado final: un compilador en Rust capaz de construir Linux 6.9 arrancable en x86, ARM y RISC-V.
## Gestión del equipo
### Especificar Teammates y modelos
Claude decide automáticamente cuántos Teammates generar según la tarea, pero usted también puede especificarlo explícitamente:
```
Crea 4 teammates para refactorizar estos módulos en paralelo.
Cada teammate debe usar el modelo Sonnet.
```
### Solicitar aprobación del plan
Para tareas complejas o de alto riesgo, puede requerir que los Teammates elaboren un plan antes de ejecutar:
```
Genera un teammate arquitecto para refactorizar el módulo de autenticación.
Antes de que hagan cualquier cambio, requiere aprobación del plan.
```
El Teammate envía una solicitud de aprobación al Lead después de completar el plan. El Lead puede aprobarlo o devolver observaciones para modificación.
### Comunicarse directamente con los Teammates
Cada Teammate es una sesión completa de Claude Code. Usted puede enviar mensajes directamente a cualquier Teammate:
* **Modo in-process**: cambie con `Shift+Down` y luego escriba el mensaje
* **Modo split-pane**: haga clic directamente en el panel correspondiente
### Cerrar Teammates
```
Por favor cierra el teammate de auditoría de seguridad
```
El Lead envía una solicitud de cierre, y el Teammate puede aprobarla o rechazarla (explicando el motivo).
### Limpiar el equipo
Al finalizar, permita que el Lead limpie los recursos:
```
Limpia el equipo
```
**Importante**: siempre limpie a través del Lead. No permita que los Teammates ejecuten la limpieza, ya que podría causar inconsistencia en el estado de los recursos.
## Mejores prácticas
### Control del tamaño del equipo
Recomendaciones de tamaño:
| Tamaño del equipo | Caso de uso |
| ----------------- | ------------------------------------------------------------ |
| 3 | Revisiones multiperspectiva simples |
| 4-5 | Desarrollo de funcionalidades estándar o refactorización |
| 6+ | Migraciones a gran escala o tareas de arquitectura complejas |
**Regla general**: asignar 5-6 tareas a cada Teammate es adecuado. Si tiene 15 tareas independientes, 3 Teammates es un buen punto de partida.
### Granularidad de las tareas
* **Demasiado pequeñas**: la sobrecarga de coordinación supera los beneficios
* **Demasiado grandes**: los Teammates trabajan demasiado tiempo sin puntos de control, aumentando el riesgo de desperdicio
* **Adecuadas**: unidades de trabajo independientes y completas con resultados claros (una función, un archivo de pruebas, un informe de revisión)
### Evitar conflictos de archivos
Dos Teammates editando el mismo archivo provocarán sobrescrituras. Al dividir el trabajo, asegúrese de que cada Teammate sea responsable de un conjunto diferente de archivos:
```
Teammate 1: responsable del directorio src/auth/
Teammate 2: responsable del directorio src/api/
Teammate 3: responsable del directorio src/utils/
```
### Monitoreo y orientación
Verifique periódicamente el progreso de los Teammates y corrija las direcciones inadecuadas a tiempo. Dejar que el equipo se ejecute sin supervisión durante demasiado tiempo aumenta el riesgo de desperdicio.
Si el Lead comienza a implementar tareas por sí mismo en lugar de esperar a los Teammates:
```
Espera a que tus teammates completen sus tareas antes de continuar
```
### Proporcionar suficiente contexto
Los Teammates cargan automáticamente el contexto del proyecto (CLAUDE.md, MCP servers, Skills), pero no heredan el historial de conversación del Lead. Proporcione suficientes detalles de la tarea al generarlos:
```
Genera un teammate de auditoría de seguridad con el siguiente prompt:
"Revisa las vulnerabilidades de seguridad del módulo de autenticación en src/auth/.
Enfócate en el manejo de tokens, gestión de sesiones y validación de entradas.
La aplicación usa JWT tokens almacenados en httpOnly cookies.
Al reportar hallazgos, incluye una calificación de severidad."
```
### Patrón de verificación con auto-reporte
Incluya criterios de verificación claros en la descripción de la tarea para asegurar que los Teammates se auto-verifiquen al completarla:
```
Al completar la tarea, reporta al Lead:
1. Qué archivos revisaste
2. Qué problemas encontraste
3. Qué modificaciones realizaste
4. Si se cumplieron los criterios de verificación (pruebas pasaron, lint sin advertencias, etc.)
```
Este patrón reduce la carga de verificación del Lead y al mismo tiempo asegura que la tarea esté realmente completada y no solo "aparentemente completada".
## Técnicas avanzadas
### Usar Hooks para imponer puertas de calidad
Mediante Hooks, puede imponer reglas cuando los Teammates completen su trabajo:
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle`: se ejecuta cuando un Teammate está a punto de quedar inactivo. Devolver exit code 2 envía retroalimentación y hace que el Teammate continúe trabajando
* `TaskCompleted`: se ejecuta cuando una tarea se marca como completada. Devolver exit code 2 puede bloquear la finalización y enviar retroalimentación
### Pre-aprobación de permisos
Las solicitudes de permisos de los Teammates se propagan al Lead, lo que puede causar interrupciones frecuentes. Pre-apruebe las operaciones comunes antes de generarlos:
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### Combinación con Worktree
Agent Teams puede usarse junto con Worktree, donde cada Teammate trabaja en su propio worktree:
```
Crea un agent team donde cada teammate trabaje en un worktree independiente,
para evitar conflictos de archivos.
```
### Herramientas de orquestación de terceros
Además del Agent Teams nativo, la comunidad también ha desarrollado algunas herramientas de orquestación:
| Herramienta | Descripción |
| --------------- | ------------------------------------------------------------------- |
| **Gas Town** | Herramienta para gestionar múltiples sesiones de Claude en paralelo |
| **Multiclaude** | Ejecuta instancias de Claude en múltiples ventanas de terminal |
Estas herramientas ofrecen alternativas fuera de la funcionalidad experimental de Agent Teams, pero requieren más configuración manual. Si el Agent Teams nativo satisface sus necesidades, se recomienda priorizar la funcionalidad oficial.
## Limitaciones actuales
Agent Teams sigue siendo una funcionalidad experimental; es importante conocer sus limitaciones:
| Limitación | Descripción |
| ------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| No se pueden recuperar teammates in-process | `/resume` y `/rewind` no recuperan teammates in-process |
| El estado de las tareas puede retrasarse | Los Teammates a veces olvidan marcar las tareas como completadas |
| El cierre puede ser lento | Los Teammates finalizan la solicitud actual antes de cerrarse |
| Un equipo por sesión | El Lead solo puede gestionar un equipo a la vez |
| Sin equipos anidados | Los Teammates no pueden generar sus propios equipos |
| Lead fijo | La sesión que crea el equipo es el Lead, no se puede transferir |
| Split panes requiere tmux/iTerm2 | No compatible con la terminal de VS Code, Windows Terminal ni Ghostty |
| Plan mode a nivel de sesión | El estado de Plan mode del Teammate se fija al generarse y no puede cambiarse durante la sesión |
## Consideraciones de costo
El consumo de tokens de Agent Teams es significativamente mayor que el de una sesión individual:
| Escenario | Consumo de tokens | Multiplicador de costo |
| --------------------------------------- | ------------------------ | ---------------------- |
| Sesión de un solo Agent | \~200k tokens | 1x |
| 3 Teammates | \~800k tokens | \~4x |
| 5 Teammates | \~1.2M tokens | \~6x |
| 16 Teammates (caso del compilador de C) | 2 mil millones de tokens | $20,000 / 2 semanas |
**Análisis de costos**:
* Cada Teammate es una instancia completamente independiente de Claude con su propio contexto
* La comunicación entre Teammates también consume tokens
* El Lead necesita coordinar a todos los Teammates, lo que genera costos adicionales
**Cuándo vale la pena**:
* Vale la pena para tareas de investigación que requieren exploración en paralelo
* Vale la pena para revisiones multiperspectiva (seguridad, rendimiento, pruebas)
* Vale la pena para decisiones que requieren discusión y consenso
* No vale la pena para tareas rutinarias que pueden completarse secuencialmente
* No vale la pena para tareas paralelas que no necesitan comunicación entre sí (usar Subagent es más económico)
## Mi experiencia de uso
### Cuándo usar Agent Teams
Mis criterios de decisión:
1. **La tarea requiere múltiples perspectivas**: conocimiento especializado de diferentes áreas (seguridad + rendimiento + pruebas)
2. **Se necesita discusión y consenso**: hipótesis competitivas, decisiones de arquitectura
3. **La exploración en paralelo tiene valor**: comparación de múltiples enfoques de implementación
Si solo necesita ejecución en paralelo sin comunicación mutua, Subagent o Worktree son más adecuados.
### Comience con investigación y revisión
Si es nuevo en Agent Teams, comience con tareas que no requieran escribir código: revisar PRs, investigar soluciones técnicas, investigar bugs. Estas tareas tienen límites claros, demuestran el valor de la exploración en paralelo y al mismo tiempo evitan los desafíos de coordinación de la implementación en paralelo.
### Combinación con otras funcionalidades
| Combinación | Efecto |
| ---------------------- | -------------------------------------------------------- |
| Agent Teams + Worktree | Cada Teammate trabaja en un entorno aislado |
| Agent Teams + Hooks | Verificación de calidad automatizada y retroalimentación |
| Agent Teams + Skills | Cada Teammate posee capacidades especializadas |
## Conclusión
Agent Teams representa un nuevo paradigma en el desarrollo asistido por IA: de "un asistente de IA" a "un equipo de IA".
Anthropic utilizó 16 Agents durante 2 semanas, con un costo de $20,000, para escribir un compilador de C de 100 mil líneas de código. La lección clave de este proyecto es: **la calidad de las pruebas es lo más importante de todo**. Los Agents resolverán autónomamente el problema que usted les asigne, por lo que el validador de tareas debe ser casi perfecto; de lo contrario, los Agents resolverán el problema equivocado.
Recuerde tres puntos clave:
| Punto clave | Descripción |
| ---------------- | -------------------------------------------------------------------------- |
| **Colaboración** | Los Teammates pueden comunicarse directamente, no solo reportar resultados |
| **Compartir** | Coordinan el trabajo a través de una lista de tareas compartida |
| **Supervisión** | Verifique periódicamente el progreso y corrija la dirección a tiempo |
Comenzar es sencillo:
```
Crea un agent team para [su tarea]
```
***
**Lecturas relacionadas**:
* [Guía completa de Claude Worktree](/es/docs/notes/claude-worktree) -- Comprender la combinación de Worktree con Agent Teams
* [Guía completa de Claude Subagent](/es/docs/notes/claude-subagent) -- Comparar los casos de uso de Subagent y Agent Teams
* [Guía rápida de inicio de Tmux](/es/docs/notes/tmux-tutorial) -- Usar Tmux para gestionar múltiples sesiones de Agent
**Materiales de referencia**:
* [Claude Code Official Documentation - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**Tutoriales en video**:
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) -- 7 demostraciones de casos de uso no técnicos
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) -- Explicación detallada de flujos de trabajo multiagente
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) -- Análisis profundo de orquestación de equipos
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) -- Tutorial completo de configuración
# Análisis completo de la arquitectura del sistema Claude
## Introducción
En septiembre de 2025, Anthropic cerró una ronda de financiamiento de $13B con una **valoración de $183B**, convirtiéndose en la cuarta empresa privada más grande del mundo. Su producto estrella, Claude Code, ha atraído a **115,000** desarrolladores activos desde su lanzamiento en febrero, procesando **195 millones de líneas** de código por semana, con un crecimiento de usuarios del **300%**.
Aún más interesante, el CEO de Anthropic, Dario Amodei, reveló: **el 90% del código de Claude Code fue escrito por él mismo**.
**¿Cómo es posible?**
¿Cómo puede un asistente de programación con IA "escribirse a sí mismo"? ¿Qué tiene de único su diseño arquitectónico que le permite asistir — e incluso reemplazar — a los desarrolladores humanos de manera tan eficiente?
La respuesta se encuentra en la **arquitectura modular** de Claude: MCP proporciona herramientas, Skills enseña cómo usarlas, Subagents ejecuta tareas en paralelo y Hooks garantiza el control determinístico — estos componentes trabajan en conjunto para dotar a Claude de la capacidad de "trabajar como un programador".
Este documento le ofrecerá una **vista panorámica de toda la arquitectura** — la función de cada componente, cómo colaboran entre sí y ejemplos de configuración para comenzar rápidamente. Los artículos posteriores profundizarán en los detalles de cada componente.
## Panorama general de la arquitectura
El sistema Claude utiliza un diseño de **arquitectura modular**, donde los componentes se clasifican por función y trabajan como **colaboradores complementarios** en lugar de dependencias jerárquicas:
**Concepto clave**: Estos componentes son extensiones **complementarias al mismo nivel**, no dependencias jerárquicas — puede combinarlos libremente según sus necesidades:
| Usted desea... | Utilice... | Descripción en una línea |
| ------------------------------------------------------ | ------------- | ------------------------------------------------------------------------------------------- |
| Conectarse a fuentes de datos y servicios externos | **MCP** | Dotar a Claude de "manos y pies" para acceder a bases de datos, APIs y sistemas de archivos |
| Enseñar a Claude flujos de trabajo específicos | **Skills** | Hacer que Claude "sepa" cómo operar en un dominio particular |
| Procesar tareas complejas en paralelo | **Subagents** | Dividir tareas grandes en subtareas, con múltiples Agents trabajando simultáneamente |
| Activar rápidamente operaciones repetitivas | **Commands** | Lanzar flujos de trabajo comunes con un solo clic, eliminando instrucciones repetitivas |
| Garantizar que ciertas operaciones siempre se ejecuten | **Hooks** | Sin importar las decisiones de Claude, este paso debe ejecutarse |
***
## Núcleo de ejecución
### Agent SDK — Motor de ejecución
Agent SDK es el **motor de ejecución central** de todo el sistema Claude Agent, proporcionando:
* **Bucle principal (Main Loop)**: El ciclo de trabajo central del Agent
* **Gestión de contexto**: Presupuesto de tokens, compresión automática (se activa al 92% de uso)
* **Despacho de herramientas**: Decide qué herramienta usar y cómo ejecutarla
* **Sistema de permisos**: Controla los permisos de acceso a las herramientas
El modelo operativo central del Agent es un simple **ciclo de retroalimentación**:
```
Recopilar contexto → Ejecutar acciones → Verificar trabajo → Repetir
```
***
### Built-in Tools — Herramientas integradas
Claude Agent incluye más de 20 herramientas integradas, organizadas en tres categorías:
| Categoría | Herramientas | Descripción |
| --------------- | ------------------- | -------------------------------------------------------------------- |
| **Lectura** | Read, Glob, Grep | Lectura de archivos, coincidencia de patrones, búsqueda de contenido |
| **Operaciones** | Write, Edit, Bash | Escritura de archivos, edición, ejecución de comandos |
| **Red** | WebSearch, WebFetch | Búsqueda web, extracción de páginas web |
Estas herramientas están **disponibles por defecto**, sin necesidad de configuración adicional. Claude interactúa con la computadora a través de estas herramientas, de la misma manera que un programador utiliza un IDE.
***
## Configuración y contexto
### CLAUDE.md — Contexto persistente
Cada vez que se inicia una nueva conversación, es necesario repetir el contexto del proyecto, las convenciones de código, las normas arquitectónicas... CLAUDE.md permite **configurar una vez y cargar automáticamente**.
CLAUDE.md funciona como un **README para IA** — le indica a Claude los conocimientos previos, la forma de trabajo y las convenciones del proyecto.
#### Herencia jerárquica
Claude carga los archivos CLAUDE.md en el siguiente orden, donde **los más específicos tienen mayor prioridad**:
```
Enterprise (más baja)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md) (más alta)
```
#### Contenido recomendado
CLAUDE.md debe incluir la siguiente información esencial:
| Categoría | Ejemplo de contenido |
| --------------------------- | -------------------------------------------------------- |
| **Stack tecnológico** | Next.js 14 + TypeScript, Tailwind CSS |
| **Comandos de compilación** | `npm run dev`, `npm run build`, `npm run test` |
| **Estándares de código** | Convenciones de nomenclatura, configuración de Lint |
| **Estructura del proyecto** | Descripción del propósito de los directorios principales |
**Principio clave**: Mantenga la brevedad. CLAUDE.md se **carga en cada conversación** — un archivo demasiado extenso desperdicia tokens valiosos.
***
## Empaquetado y distribución
### Plugins — Unidades instalables
Las configuraciones de equipo están dispersas, son difíciles de compartir y de estandarizar. Cada persona tiene su propio conjunto de Skills, Commands, Hooks... ¿Cómo se unifica la gestión?
Plugins empaqueta **Skills + Commands + Subagents + Hooks + MCP** en **unidades instalables**, permitiendo la distribución con un solo clic y la estandarización del equipo.
#### Estructura de directorios
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # Manifiesto del plugin (requerido)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Configuración de Hooks
├── .mcp.json # Configuración de MCP Server
└── README.md # Documentación
```
#### Ejemplo de configuración
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "Kit de herramientas de desarrollo frontend",
"author": "Your Team"
}
```
```bash
# Métodos de instalación
claude plugin install github:your-org/your-plugin # Desde GitHub
claude plugin install /path/to/plugin # Desde local
```
**Recursos relacionados**
| Recurso | Descripción |
| -------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Repositorio oficial de plugins de Anthropic |
| [wshobson/agents](https://github.com/wshobson/agents) | 24.3k estrellas, colección de plantillas de Agent de alta calidad |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | Mercado comunitario de plugins para descubrir todo tipo de plugins |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | 19.3k estrellas, lista curada de recursos de Claude Code |
***
## Capacidades de extensión (módulos complementarios)
Las capacidades de extensión del sistema Claude están compuestas por múltiples **módulos complementarios**, cada uno con su función específica, trabajando en sinergia:
| Módulo | Función | Método de activación |
| ------------- | ------------------------------------------------------------------------ | -------------------------------- |
| **MCP** | Conectar datos y servicios externos (WHAT) | Disponible tras la configuración |
| **Skills** | Conocimiento procedimental — enseñar a Claude cómo hacer las cosas (HOW) | Coincidencia automática |
| **Subagents** | Contexto independiente, delegación paralela de tareas | Invocación explícita |
| **Commands** | Flujos de trabajo repetitivos | Manual `/cmd` |
| **Hooks** | Control determinístico, basado en eventos | Activación automática |
***
### MCP — Conectividad externa
#### Filosofía de diseño
En el enfoque tradicional, cada fuente de datos externa requiere una **integración personalizada**, lo que genera una pesadilla de integración N x M. MCP proporciona un protocolo estandarizado para lograr **integrar una vez, usar en todas partes**.
MCP (Model Context Protocol) fue diseñado como el **puerto USB-C para aplicaciones de IA**:
| Característica | Descripción |
| ---------------------------- | --------------------------------------------------------------------------------- |
| **Estándar abierto** | Publicado en noviembre de 2024, donado a la Linux Foundation en diciembre de 2025 |
| **Adopción de la industria** | Adoptado por OpenAI, Microsoft, Google, AWS y otros |
| **Escala del ecosistema** | Más de 97M de descargas mensuales del SDK, miles de servidores comunitarios |
#### Patrón arquitectónico
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, herramientas de IA)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (Clientes de protocolo)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (Exponen capacidades específicas)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**Casos de uso**: Conectar bases de datos, integrar servicios de terceros (GitHub, Slack, Notion), acceder a APIs privadas, procesamiento de flujos de datos en tiempo real.
#### Ejemplo de configuración
Cree `.mcp.json` en el directorio raíz del proyecto:
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — Flujos de trabajo manuales
Slash Commands proporcionan flujos de trabajo repetitivos de **activación manual**.
| Característica | Descripción |
| ------------------------ | --------------------------------------------------------- |
| **Método de activación** | Escribir manualmente `/command-name` |
| **Ubicación** | `.claude/commands/` |
| **Propósito** | Flujos de trabajo repetitivos, operaciones estandarizadas |
**Ejemplo**: Cree `.claude/commands/review.md`
```markdown
Por favor, realice una revisión de código sobre los cambios actuales, enfocándose en:
1. Estilo y consistencia del código
2. Problemas de rendimiento potenciales
3. Vulnerabilidades de seguridad
4. Cobertura de pruebas
```
Luego escriba `/review` para activarlo.
***
### Hooks — Control determinístico
Hooks es el núcleo del **control determinístico** — ciertas operaciones deben ejecutarse y no pueden depender del juicio del LLM.
| Categoría | Evento | Momento de activación |
| ---------------- | -------------------- | ----------------------------------------------------- |
| **Herramientas** | `PreToolUse` | Antes de la ejecución de la herramienta |
| | `PostToolUse` | Después de la ejecución exitosa de la herramienta |
| | `PostToolUseFailure` | Después de un fallo en la ejecución de la herramienta |
| | `PermissionRequest` | Al solicitar permisos |
| **Sesión** | `SessionStart` | Al iniciar la sesión |
| | `SessionEnd` | Al finalizar la sesión |
| | `Stop` | Cuando Claude completa una respuesta |
| **Subagentes** | `SubagentStart` | Al iniciar un subagente |
| | `SubagentStop` | Al detener un subagente |
| **Otros** | `UserPromptSubmit` | Después de que el usuario envía un prompt |
| | `Notification` | Eventos de notificación |
| | `PreCompact` | Antes de la compresión del contexto |
#### Ejemplo de configuración
Formateo automático de archivos TypeScript:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### En la práctica: Bucle autónomo
**Ralph Wiggum** es un plugin oficial de Anthropic que utiliza el Stop hook para implementar un ciclo de iteración autónomo:
```bash
/ralph-loop "Implementar una API de TODO con CRUD y pruebas" --max-iterations 20
```
**Cómo funciona**: El Stop hook intercepta la salida de Claude, reinyecta el prompt original y continúa iterando hasta que la tarea se complete o se alcance el número máximo de iteraciones.
**Escenarios adecuados**: Tareas que requieren múltiples iteraciones (pasar pruebas, refactorización de código), tareas con medios de verificación automatizada.
***
### Subagents — Delegación de tareas y ejecución paralela
#### Filosofía de diseño
Un Agent único enfrenta desafíos: ventana de contexto limitada, imposibilidad de paralelizar y responsabilidades poco claras. Subagents adopta el patrón arquitectónico **Orchestrator-Worker** para resolver estos problemas:
```
Main Agent (Orchestrator)
├── Analizar la solicitud del usuario
├── Formular el plan
├── Descomponer las tareas
└── Generar subagentes especializados
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - ejecución paralela)
│ Código │ Pruebas│ Docs │
└────────┴────────┴────────┘
↓
Consolidar resultados → El agente principal sintetiza la salida
```
#### Características principales
| Característica | Descripción |
| --------------------------------------- | ------------------------------------------------------------------------------- |
| **Aislamiento de contexto** | Cada Subagent tiene su propio contexto independiente, evitando la contaminación |
| **Especialización de tareas** | Prompts de sistema personalizados definen roles dedicados |
| **Control de permisos de herramientas** | Se puede restringir a los Subagents a usar solo herramientas específicas |
| **Ejecución paralela** | Múltiples Subagents trabajan simultáneamente |
**Datos de rendimiento**: Los sistemas multiagente superan a los de agente único en un 90.2%, y la paralelización puede reducir el tiempo de investigación en un 90% (el consumo de tokens es aproximadamente 15x, pero vale la pena para tareas complejas).
#### Ejemplo de configuración
Cree un archivo Markdown en `.claude/agents/`:
```markdown
---
name: Code Reviewer
description: Un subagente especializado en revisión de código
tools:
- Read
- Grep
- Glob
---
Usted es un experto senior en revisión de código. Enfóquese en:
1. Calidad del código y mantenibilidad
2. Bugs potenciales y casos límite
3. Oportunidades de optimización de rendimiento
4. Vulnerabilidades de seguridad
```
***
### Skills — Conocimiento procedimental
#### Filosofía de diseño
Skills son **manuales de trabajo reutilizables para IA** — paquetes modulares de conocimiento que Claude puede cargar dinámicamente según la necesidad. El principio de diseño central es la **Divulgación progresiva (Progressive Disclosure)**:
```
📚 Manual de Skills
│
├─ 📋 Índice ─────────────── [Capa de metadatos] Precargado al inicio (~30-50 tokens)
│ name: "weekly-report"
│ description: "Generar informes semanales estandarizados"
│
├─ 📖 Contenido principal ── [Capa de documento central] Se carga cuando es relevante (~cientos a miles de tokens)
│ # Weekly Report Generator
│ ## Instructions
│ Generar informes semanales siguiendo esta estructura...
│
└─ 📎 Apéndice ──────────── [Capa de recursos de referencia] Se carga cuando es necesario
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills**: MCP otorga a Claude la capacidad de acceder a herramientas (WHAT), mientras que Skills le enseña a Claude cómo usar esas herramientas de manera efectiva (HOW).
#### Ventajas principales
| Ventaja | Descripción |
| ------------------------- | --------------------------------------------------------------------------------------------- |
| **Eficiente en tokens** | Los metadatos solo ocupan 30-50 tokens; se pueden activar decenas de Skills simultáneamente |
| **Activación automática** | Coincidencia automática basada en el contexto de la tarea, sin necesidad de activación manual |
| **Componible** | Múltiples Skills trabajan juntos automáticamente |
| **Portable** | Experiencia consistente en Claude.ai, Claude Code y la API |
#### Ejemplo de configuración
Cree un directorio en `.claude/skills/`:
```
my-skill/
├── SKILL.md # Instrucciones centrales (requerido)
├── scripts/ # Scripts ejecutables (opcional)
└── references/ # Materiales de referencia (opcional)
```
Estructura central de SKILL.md:
```yaml
---
name: code-review # Nombre del Skill
description: Revisión de código para calidad y seguridad # Descripción breve (usada para coincidencia automática)
---
# Code Review Skill
## Instructions
[Instrucciones detalladas paso a paso...]
## Output Format
[Requisitos del formato de salida...]
```
**Punto clave**: La `description` en el frontmatter se utiliza para la coincidencia automática — manténgala concisa y precisa.
***
## Referencias oficiales
**Filosofía de diseño**
| Recurso | Descripción |
| --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Artículo fundamental sobre arquitectura de Agents |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Prácticas de ingeniería del Agent SDK |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Filosofía de diseño de Skills |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | Anuncio de lanzamiento de MCP |
**Documentación oficial**
| Recurso | Descripción |
| ----------------------------------------------------------------------------- | -------------------------------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Skills comparado con otros componentes |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | Guía de uso de CLAUDE.md |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Documentación oficial de Subagents |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Documentación oficial de Hooks |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | Documentación oficial de MCP |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | Especificación del protocolo MCP |
**Análisis en profundidad**
| Recurso | Descripción |
| --------------------------------------------------------------------------------------------------------- | --------------------------------------------------- |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | Incluye diagramas de arquitectura |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Análisis profundo de los principios de Skills |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Detalles internos de la construcción de Claude Code |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Comparación técnica entre Skills y MCP |
***
## Lecturas adicionales
Si desea profundizar en los conceptos y la práctica de Skills, puede consultar:
* [¿Qué son los Claude Skills?](/es/docs/notes/claude-skills/concept) — Explicación detallada de los principios fundamentales de Skills
* [Guía práctica de Claude Skills](/es/docs/notes/claude-skills/practice) — Cree su primer Skill paso a paso
* [Guía completa de Claude Subagent](/es/docs/notes/claude-subagent) — Uso y personalización de subagentes
* [Análisis profundo de GSD](/es/docs/notes/gsd/concept) — Un sistema de programación con IA basado en ingeniería de contexto
* [Mis mejores prácticas de Claude Code](/es/blog/claude-code-best-practices) — Consejos de uso diario de Claude Code
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# Guia completa de Claude Worktree
## Introduccion
Cuando utiliza Claude Code para tareas complejas, es posible que haya enfrentado esta situacion: tiene tres tareas independientes que manejar, pero ejecutar multiples instancias de Claude en el mismo directorio provoca conflictos de codigo — un Agent modifica un archivo mientras otro tambien lo esta editando, y al momento de fusionar todo se convierte en un desastre.
En febrero de 2025, Anthropic lanzo el comando `--worktree`, transformando completamente esta situacion. Ahora puede ejecutar `claude -w feature-1`, `claude -w feature-2` y `claude -w bugfix-1` en tres terminales separadas, con cada Agent trabajando en un entorno aislado sin interferir con los demas.
## Comprender Worktree
Imagine que es un arquitecto disenando tres habitaciones diferentes al mismo tiempo. El enfoque tradicional seria dibujar en el mismo plano, lo cual facilmente genera confusion al hacer cambios. El enfoque de Worktree le proporciona tres planos independientes, cada uno dedicado al diseno de una habitacion, para luego fusionarlos en el plano principal.
Desde el punto de vista tecnico, Worktree es una funcionalidad nativa de Git. El comando `--worktree` de Claude Code encapsula esta funcionalidad para hacerla mas simple — un solo comando crea el entorno aislado, inicia la instancia de Claude y realiza la limpieza automatica al finalizar.
### Por que no simplemente hacer multiples Clone
Quizas se pregunte: ¿por que no simplemente clonar el codigo multiples veces?
| Solucion | Uso de disco | Dificultad de sincronizacion | Complejidad de limpieza |
| ------------------- | --------------------------------------------- | ---------------------------------- | ----------------------------------------- |
| Multiples Clone | Cada copia es un repositorio completo | Requiere pull/push manual | Requiere eliminar directorios manualmente |
| Git Worktree | Solo copia archivos de trabajo, comparte .git | Comparte historial automaticamente | `git worktree remove` |
| Claude `--worktree` | Solo copia archivos de trabajo, comparte .git | Comparte historial automaticamente | Limpieza automatica al salir |
Worktree comparte la misma base de datos `.git`, todo el historial de commits e informacion de ramas es compartido. Esto significa que un commit creado en un worktree es inmediatamente visible desde los demas worktrees.
### Cuando utilizar Worktree
Antes de comenzar, evalue si su tarea es adecuada para usar worktree.
Regla general: **si la tarea requiere mas de 30 minutos, considere usar worktree**. Para tareas cortas, usar worktree puede ser contraproducente — crear el entorno, instalar dependencias y fusionar al final puede tomar mas tiempo que la tarea en si. Sin embargo, para tareas que requieren trabajo profundo, el aislamiento de worktree es muy valioso.
| Adecuado para Worktree | No tan adecuado |
| ----------------------------------------------------- | ----------------------------------------------- |
| Desarrollo de funcionalidades independientes | Cambios pequenos que se completan en 10 minutos |
| Refactorizacion paralela de diferentes modulos | Tareas que requieren interaccion frecuente |
| Tareas de larga duracion | Fuerte dependencia de otros cambios en progreso |
| Cambios experimentales que necesitan pruebas aisladas | Correcciones de bugs simples |
### Requisitos previos
Antes de usar worktree, asegurese de cumplir con las siguientes condiciones:
| Condicion | Descripcion |
| ---------------------- | ---------------------------------------------------------------------- |
| Git inicializado | Debe estar en un directorio de repositorio Git (con directorio `.git`) |
| Al menos un commit | No se puede crear worktree en un repositorio vacio |
| Rama remota disponible | Por defecto se hace checkout desde la rama remota (como `origin/main`) |
## Flujo de trabajo completo: de la creacion a la limpieza
A continuacion, recorreremos el flujo completo desde la creacion del Worktree hasta la limpieza final, siguiendo el orden real de desarrollo.
### Paso 1: Crear el Worktree
#### Crear desde la rama remota predeterminada
Utilice el parametro `-w` o `--worktree` para iniciar Claude:
```bash
# Crear un worktree llamado "feature-auth" e iniciar Claude
claude -w feature-auth
# Generar un nombre aleatorio automaticamente (como "bright-running-fox")
claude -w
```
Este comando realiza cuatro acciones:
1. Crea un nuevo directorio de trabajo en `/.claude/worktrees/feature-auth/`
2. Crea una nueva rama llamada `worktree-feature-auth`
3. Hace checkout del codigo desde la rama remota predeterminada (como `origin/main` u `origin/master`) — **tenga en cuenta que no es la rama en la que se encuentra actualmente**
4. Inicia Claude Code en el nuevo directorio
Todos los worktrees se ubican en el directorio `.claude/worktrees/`:
```
your-project/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← primer worktree
│ ├── bugfix-123/ ← segundo worktree
│ └── refactor-api/ ← tercer worktree
├── src/
└── package.json
```
Se recomienda agregar esta ruta a `.gitignore`:
```bash
# .gitignore
.claude/worktrees/
```
#### Crear desde la rama actual o una rama especifica
`-w` siempre hace checkout desde la rama remota predeterminada y actualmente no soporta especificar una rama base. Si desea crear un worktree basado en la rama actual (o una rama especifica), hay tres formas:
**Forma 1: Creacion manual con Git**
```bash
# Crear worktree basado en el HEAD actual
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# O basado en una rama especifica
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# Luego iniciar Claude en ese directorio
cd .claude/worktrees/my-feature && claude
```
Esta forma le da control total — puede crear worktrees basados en cualquier rama o cualquier commit, y el directorio de trabajo comienza en la rama correcta desde el inicio. La documentacion oficial tambien recomienda: "Si necesita mas control sobre ramas y ubicaciones, cree el worktree directamente con Git y luego ejecute Claude en ese directorio."
**Forma 2: Crear dentro de la conversacion (recomendado)**
En una sesion existente de Claude, simplemente pida a Claude que cree el worktree:
```
> 从当前分支开启一个worktree
> start a worktree
```
A diferencia del comando `-w`, un worktree creado dentro de la conversacion **se basara automaticamente en la rama actual**, no en la rama remota predeterminada. Claude completara automaticamente la creacion del worktree y cambiara a el, sin necesidad de ejecutar manualmente ningun comando Git. Si ya esta trabajando en una rama feature, esta es la forma mas conveniente — con una sola instruccion obtiene un entorno aislado basado en su rama actual.
**Forma 3: Crear primero con `-w`, luego cambiar de rama en la sesion**
Primero cree el worktree con `claude -w`, y una vez dentro de la sesion pida a Claude que cambie a la rama objetivo. La desventaja de esta forma es que primero descarga la rama remota predeterminada y luego cambia — un paso adicional, menos limpio que las dos formas anteriores. Ademas, si la rama objetivo ya esta en uso por otro worktree, se encontrara con un conflicto de ramas:
Como se muestra en la captura, Claude detecta el conflicto de ramas y ofrece dos opciones: volver al directorio principal para operar, o crear una nueva rama de trabajo basada en la rama objetivo. Aunque eventualmente funciona, el proceso no es tan directo como las formas 1 y 2.
**Avanzado: encapsular en un comando unico con Makefile**
Si frecuentemente necesita crear worktrees desde la rama actual, puede agregar un comando rapido al `Makefile` en la raiz del proyecto, encadenando la creacion + apertura del editor + inicio de Claude en un pipeline determinista:
```makefile
# Crear worktree desde la rama actual e iniciar el entorno de desarrollo
# Uso: make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "用法: make worktree name="; \
echo "示例: make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ 从 $$(git branch --show-current) 创建 worktree: $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ 初始化环境"
cd .claude/worktrees/$(name) && npm install
@echo "→ 在 Zed 中打开"
zed .claude/worktrees/$(name)
@echo "→ 启动 Claude"
cd .claude/worktrees/$(name) && claude
```
Su uso es muy sencillo:
```bash
# Crear worktree basado en la rama actual, abrir en Zed e iniciar Claude
make worktree name=fix-login-bug
# Abrir varias tareas en paralelo
make worktree name=feature-search
make worktree name=refactor-api
```
En comparacion con escribir manualmente multiples comandos Git/cd/claude, `make worktree name=xxx` requiere solo una linea, y el flujo ejecutado es exactamente el mismo cada vez — no olvidara ningun paso ni escribira mal una ruta. Es importante tener en cuenta que, como el Makefile utiliza `git worktree add` nativo, el Hook `WorktreeCreate` de Claude Code no se activara (dicho Hook solo se activa al usar `claude -w` o al crear worktree dentro de la conversacion). Por lo tanto, los pasos de inicializacion del entorno (instalacion de dependencias, copia de `.env`, etc.) deben escribirse directamente en el Makefile, como se muestra en el ejemplo anterior con `npm install`.
### Paso 2: Inicializar el entorno
Una vez creado el Worktree, lo primero es inicializar el entorno de desarrollo. Cada nuevo worktree es un directorio independiente; `node_modules`, entornos virtuales, archivos `.env`, etc., no se transfieren automaticamente.
Claude Code proporciona el Hook `WorktreeCreate` para automatizar la configuracion del entorno:
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
De esta forma, cada vez que se crea un worktree, las dependencias se instalan automaticamente y el archivo de variables de entorno se copia automaticamente. Pasos de inicializacion comunes:
| Tipo de proyecto | Comando de inicializacion |
| ---------------- | ----------------------------------------------------------- |
| Node.js | `npm install` o `yarn` |
| Python | `pip install -r requirements.txt` o activar entorno virtual |
| Go | `go mod download` |
| General | Copiar archivos `.env`, configurar variables de entorno |
Si no ha configurado el Hook, tambien puede ejecutar `/init` al inicio de cada sesion de worktree para asegurar que Claude comprenda correctamente el contexto del directorio de trabajo actual, releyendo la estructura del proyecto y la configuracion de CLAUDE.md.
### Paso 3: Commit y fusion
Una vez que el entorno esta listo y el desarrollo completado, el siguiente paso es fusionar los cambios de vuelta a la rama objetivo.
**Fusionar de vuelta a la rama main**
El caso mas comun — el worktree se creo desde `origin/main` y los cambios deben fusionarse de vuelta a `main`. En la sesion de Claude del worktree, simplemente diga:
```
> 提交所有改动,推送到远端,然后创建一个 PR 到 main
```
Claude completara automaticamente todo el flujo de commit, push y `gh pr create`.
**Fusionar de vuelta a una rama feature**
Si esta desarrollando en la rama `feature-x` y los cambios del worktree deben fusionarse de vuelta a `feature-x` en lugar de `main`:
```
> 提交改动并推送,然后创建一个 PR 合并到 feature-x 分支
```
Claude ejecutara `gh pr create --base feature-x`, creando un PR directamente hacia la rama feature.
Tambien puede salir de la sesion del worktree (eligiendo conservar el worktree) y volver al directorio principal para iniciar Claude:
```
> 把 worktree-my-task 分支的改动合并到当前分支
```
Si hay algunos commits en el worktree que no desea, puede hacer cherry-pick selectivamente:
```
> 帮我查看 worktree-my-task 分支的 commit 历史,然后把其中关于认证模块的那几个 commit cherry-pick 到当前分支
```
> **Nota**: Todos los worktrees comparten la misma base de datos `.git`. Los commits creados en un worktree son inmediatamente visibles en el directorio principal, sin necesidad de operaciones adicionales de push/pull.
### Paso 4: Salir y limpiar
Una vez completada la fusion del codigo, puede salir de la sesion del worktree.
Al salir de la sesion del worktree, Claude procesara automaticamente segun la situacion:
| Estado | Accion |
| ------------------------- | --------------------------------------------- |
| **Sin cambios** | Elimina automaticamente el worktree y la rama |
| **Con cambios o commits** | Le pregunta si desea conservar o eliminar |
Los worktrees conservados permanecen disponibles para que pueda continuar trabajando mas adelante.
Tambien puede configurar el Hook `WorktreeRemove` para automatizar la limpieza:
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**Comandos de gestion manual**
Si necesita gestionar worktrees manualmente, puede usar los comandos estandar de Git:
```bash
# Listar todos los worktrees
git worktree list
# Eliminar manualmente un worktree
git worktree remove .claude/worktrees/feature-auth
# Limpiar referencias de worktrees obsoletos
git worktree prune
```
> **Nota**: No elimine directamente el directorio del worktree con `rm -rf`. La forma correcta es usar `git worktree remove`, o si ya lo elimino por error, ejecute `git worktree prune` para limpiar las referencias residuales.
## Patrones de desarrollo paralelo
Una vez dominado el flujo de trabajo basico, veamos como aprovechar worktree para el desarrollo paralelo.
### Multiples terminales en paralelo
El uso mas comun es ejecutar simultaneamente en multiples pestanas de terminal:
```bash
# Terminal 1: Manejar la funcionalidad de autenticacion de usuarios
claude -w feature-auth
# Terminal 2: Corregir bug de pagos
claude -w bugfix-payment
# Terminal 3: Refactorizar modulo de API
claude -w refactor-api
```
Cada instancia de Claude trabaja en su propio worktree, y las modificaciones no se afectan entre si. Puede:
* En una terminal, dejar que Claude desarrolle una nueva funcionalidad
* En otra terminal, dejar que Claude corrija un bug
* En una tercera terminal, continuar con su propia revision de codigo
### Implementacion competitiva
Un uso muy eficiente es que multiples Agents implementen independientemente la misma funcionalidad:
```bash
# Tres terminales ejecutandose por separado
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
Proporcioneles los mismos requisitos y deje que cada uno implemente su version. Al final, compare las tres soluciones y fusione la mejor. Esto aprovecha la naturaleza no determinista de los LLM — la misma entrada puede producir diferentes salidas, y a veces la segunda version resulta ser mejor.
La exploracion de diseno de UI tambien es muy adecuada para este patron. Suponga que desea redisenar la interfaz de su aplicacion pero no esta seguro de que estilo es mejor:
```bash
# Dejar que tres Agents implementen diferentes estilos
claude -w ui-minimal # estilo minimalista
claude -w ui-colorful # colores vibrantes
claude -w ui-glassmorphism # estilo glassmorphism
```
Una vez completados, ejecute los servidores de desarrollo de las tres versiones simultaneamente (en diferentes puertos), compare los resultados lado a lado y fusione la solucion preferida a la rama principal — mucho mas eficiente que el flujo tradicional de "hacer una version, ver el resultado, no queda bien, volver a cambiar".
### Aislamiento de Subagent
Worktree no solo es util para la instancia principal de Claude, sino tambien para los Subagent. En el frontmatter de un Subagent personalizado, agregue `isolation: worktree`:
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
Tambien puede indicarlo directamente a Claude en la conversacion:
```
> 使用 worktree 来隔离你的 agents
> use worktrees for your agents
```
Cuando un Subagent esta configurado con aislamiento worktree:
```
Agent principal (directorio principal)
│
├── Inicia Migration Agent 1 ──→ worktree-migration-1/
│ └── Procesa directorio src/auth/
│
├── Inicia Migration Agent 2 ──→ worktree-migration-2/
│ └── Procesa directorio src/api/
│
└── Inicia Migration Agent 3 ──→ worktree-migration-3/
└── Procesa directorio src/utils/
```
Cada Subagent trabaja de forma independiente en su propio worktree, sin interferir entre si. Al finalizar, el worktree se limpia automaticamente (si no hay cambios sin confirmar).
### Combinacion con Tmux e IDE
Con el parametro `--tmux`, puede iniciar automaticamente en una nueva sesion de Tmux; incluso si cierra la terminal, Claude seguira ejecutandose en segundo plano:
```bash
claude -w feature-auth --tmux
```
Si utiliza VS Code o Cursor, el panel de control de codigo fuente reconocera automaticamente todos los worktrees — el repositorio principal se muestra como un repo, y cada worktree aparece como un repo independiente, permitiendo cambiar, confirmar y enviar directamente desde el IDE. Los Worktrees tambien se pueden combinar con el [ciclo Ralph](/es/docs/notes/ralph-wiggum/concept); cada ciclo Ralph se ejecuta en su propio worktree, de modo que incluso si el ciclo falla, no afectara la rama principal.
## Consideraciones y mejores practicas
### Errores comunes
1. **Confundir el origen de la rama**: Los worktrees creados con `-w` hacen checkout desde la **rama remota predeterminada**, no desde la rama en la que se encuentra actualmente. Si ejecuta `claude -w my-task` estando en la rama `feature-x`, el codigo del nuevo worktree proviene de `origin/main` y no incluira los cambios de `feature-x`. Para trabajar desde la rama actual, consulte [Crear desde la rama actual o una rama especifica](#crear-desde-la-rama-actual-o-una-rama-especifica).
2. **Los cambios no confirmados no se transfieren**: Al crear un worktree, las modificaciones no preparadas o no confirmadas del directorio principal no apareceran en el nuevo worktree. Los Worktrees se crean unicamente basandose en el historial de commits, asi que asegurese de que los cambios importantes esten confirmados.
3. **La misma rama no puede ser utilizada por multiples Worktrees**: Git no permite que dos worktrees tengan la misma rama en checkout simultaneamente. Si ya esta en la rama `feature-x` en el directorio principal, intentar hacer checkout de `feature-x` en un worktree generara un error. Cada worktree debe estar en una rama diferente.
4. **El entorno necesita reinicializarse**: Cada nuevo worktree no incluye dependencias de ejecucion como `node_modules`; se recomienda configurar el Hook `WorktreeCreate` para automatizar esto (consulte [Paso 2: Inicializar el entorno](#paso-2-inicializar-el-entorno)).
### Recomendaciones de uso
No se exceda. Aunque tecnicamente puede abrir muchos worktrees, cada instancia de Claude consume cuota de API, demasiadas tareas en paralelo son dificiles de rastrear, y los conflictos al momento de fusionar se vuelven mas complejos.
**Convencion de nombres**: Adopte buenos habitos de nomenclatura para facilitar la gestion posterior:
```bash
# Buenos nombres
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# Malos nombres
claude -w test
claude -w temp
claude -w 1
```
## Control de versiones no-Git
Si utiliza SVN, Perforce o Mercurial, puede lograr un efecto de aislamiento similar configurando los Hooks `WorktreeCreate` y `WorktreeRemove`. Al configurar estos Hooks, usar `--worktree` invocara sus comandos personalizados en lugar del comportamiento predeterminado de Git.
## Reflexiones finales
Worktree es una funcionalidad que el equipo de Claude Code utiliza todos los dias; Boris Cherny la denomina "el consejo de productividad numero uno". El valor principal es simple: **permitir que multiples Agents trabajen en paralelo sin interferir entre si**.
Comenzar es sencillo, solo necesita:
```bash
claude -w your-task-name
```
***
**Lecturas relacionadas**:
* [Guia completa de Claude Subagent](/es/docs/notes/claude-subagent) — Comprender la sinergia entre Subagent y Worktree
* [Analisis profundo de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) — Otro metodo para mejorar la eficiencia de la programacion con IA
* [Analisis completo de la arquitectura de Claude](/es/docs/notes/claude-architecture) — Comprender la posicion de Worktree en la arquitectura general
**Referencias**:
* [Documentacion oficial de Claude Code - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Anuncio de Worktree de Boris Cherny](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Documentacion oficial de Git Worktree](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**Tutoriales en video**:
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — Demostracion practica del flujo de trabajo completo con worktree
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — Metodo de creacion manual de worktree
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — Explicacion detallada de la funcionalidad nativa de worktree
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — Tutorial rapido que incluye el uso con Subagent
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — Ventajas del desarrollo paralelo con multiples worktrees
# Documentación
Bienvenido al centro de documentación. Aquí encontrará la documentación técnica y los tutoriales que he recopilado sobre Claude Code.
# Guía rápida de inicio de Tmux
## Introducción
Si ha utilizado Agent Teams de Claude Code o desea ejecutar múltiples instancias de Claude simultáneamente, Tmux es prácticamente una herramienta indispensable. Le permite ejecutar múltiples sesiones en una sola ventana de terminal, las sesiones continúan ejecutándose en segundo plano incluso si cierra la terminal, y Claude puede generar y gestionar automáticamente múltiples Agents dentro de Tmux.
Este tutorial está diseñado específicamente para usuarios de Claude Code, cubriendo tanto los fundamentos de Tmux como los consejos de integración con Claude Code.
## Comprender Tmux
Los tres conceptos fundamentales de Tmux:
```
┌─────────────────────────────────────────────────────────┐
│ Session (Sesión) │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window (Ventana) ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| Concepto | Analogía | Descripción |
| ----------- | --------------------- | ---------------------------------------------------------------------------- |
| **Session** | Espacio de trabajo | Contenedor de nivel superior, continúa ejecutándose incluso si se desconecta |
| **Window** | Pestaña del navegador | Una sesión puede contener múltiples ventanas |
| **Pane** | Pantalla dividida | Una ventana puede dividirse en múltiples paneles |
## Instalación y fundamentos
### Instalar Tmux
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
Verificar la instalación:
```bash
tmux -V
# Salida similar a: tmux 3.6a
```
### Tecla de prefijo
Todos los comandos de Tmux comienzan con una **tecla de prefijo**, que por defecto es `Ctrl+B`.
Método para ingresar comandos:
1. Presione `Ctrl+B` (sin soltar)
2. Después de soltar, presione la tecla del comando
Por ejemplo, dividir ventana: `Ctrl+B` y luego presione `%`
## Referencia rápida de comandos
### Gestión de sesiones
| Comando | Descripción |
| --------------------------- | ---------------------------------------------------------- |
| `tmux` | Crear nueva sesión |
| `tmux new -s name` | Crear sesión con nombre |
| `tmux ls` | Listar todas las sesiones |
| `tmux attach -t name` | Conectarse a una sesión |
| `tmux kill-session -t name` | Cerrar una sesión |
| `Ctrl+B d` | Desconectar la sesión actual (se ejecuta en segundo plano) |
### Gestión de ventanas
| Atajo | Descripción |
| ------------ | --------------------------------- |
| `Ctrl+B c` | Crear nueva ventana |
| `Ctrl+B n` | Siguiente ventana |
| `Ctrl+B p` | Ventana anterior |
| `Ctrl+B 0-9` | Cambiar a la ventana especificada |
| `Ctrl+B ,` | Renombrar la ventana actual |
| `Ctrl+B &` | Cerrar la ventana actual |
### Gestión de paneles
| Atajo | Descripción |
| ---------------------------- | ------------------------------------- |
| `Ctrl+B %` | División vertical (izquierda-derecha) |
| `Ctrl+B "` | División horizontal (arriba-abajo) |
| `Ctrl+B teclas de dirección` | Moverse entre paneles |
| `Ctrl+B x` | Cerrar el panel actual |
| `Ctrl+B z` | Maximizar/restaurar panel |
| `Ctrl+B {` | Mover panel hacia la izquierda |
| `Ctrl+B }` | Mover panel hacia la derecha |
### Otros comandos frecuentes
| Atajo | Descripción |
| ---------- | --------------------------------------------- |
| `Ctrl+B [` | Entrar en modo copia (permite desplazamiento) |
| `q` | Salir del modo copia |
| `Ctrl+B ?` | Mostrar todos los atajos |
## Integración con Claude Code
### Por qué Claude Code necesita Tmux
1. **Modo split-pane de Agent Teams**: cada Teammate se muestra en un panel independiente
2. **Ejecución en segundo plano**: las tareas continúan ejecutándose incluso si cierra la terminal
3. **Persistencia de sesiones**: restaura el contexto completo al reconectarse después de una desconexión
4. **Gestión de múltiples instancias**: ejecute múltiples sesiones de Claude simultáneamente
### Uso básico: ejecutar Claude en segundo plano
```bash
# Iniciar Claude en tmux
tmux new -s claude-work
claude
# Desconectar la sesión (Claude continúa ejecutándose)
# Ctrl+B d
# Reconectarse más tarde
tmux attach -t claude-work
```
### Usar el parámetro --tmux
Claude Code tiene soporte nativo para la integración con Tmux:
```bash
# Iniciar Claude en una nueva sesión de tmux
claude --tmux
# Usar junto con worktree
claude -w feature-auth --tmux
```
Esto automáticamente:
1. Crea una nueva sesión de tmux
2. Inicia Claude Code en ella
3. Nombra la sesión como `claude-{ID aleatorio}`
### Modo Tmux de Agent Teams
Agent Teams puede usar el modo de visualización split-pane, donde cada Teammate se ejecuta en un panel independiente:
```json
// settings.json
{
"teammateMode": "tmux"
}
```
O mediante la línea de comandos:
```bash
claude --teammate-mode tmux
```
Resultado:
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## Configuración práctica
### \~/.tmux.conf recomendado
Cree o edite `~/.tmux.conf`:
```bash
# Usar Ctrl+A como tecla de prefijo (más fácil de presionar)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# Habilitar soporte de ratón
set -g mouse on
# Aumentar el búfer de historial (Claude genera mucha salida)
set -g history-limit 50000
# Navegación de paneles estilo vim
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# Atajos de división más intuitivos
bind | split-window -h
bind - split-window -v
# Recarga rápida de configuración
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# Soporte de 256 colores
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# Numeración de ventanas desde 1 (el 0 está muy lejos)
set -g base-index 1
setw -g pane-base-index 1
# Optimización de la barra de estado
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
Recargar configuración:
```bash
tmux source-file ~/.tmux.conf
```
### Configuración dedicada para Claude Code
Configuración optimizada para Claude Code:
```bash
# Atajo de ventana emergente para sesión de Claude
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
El efecto de esta configuración:
1. Presione `Ctrl+A y` para abrir la ventana emergente de Claude
2. Cada directorio tiene una sesión independiente de Claude
3. La sesión continúa ejecutándose después de cerrar la ventana emergente
4. Al reabrirla, se restaura la conversación anterior
## Flujos de trabajo comunes
### Flujo de trabajo 1: Múltiples proyectos en paralelo
```bash
# Crear una sesión independiente para cada proyecto
tmux new -s project-a
# Iniciar Claude dentro
claude -w feature-x
# Desconectar y crear otra sesión
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# Cambiar entre sesiones
tmux switch -t project-a
tmux switch -t project-b
# O listar todas las sesiones para seleccionar
# Ctrl+B s
```
### Flujo de trabajo 2: Panel de desarrollo
Crear un entorno de desarrollo con múltiples paneles:
```bash
# Crear sesión
tmux new -s dev
# Dividir en tres paneles
# Ctrl+B % (división vertical)
# Ctrl+B " (división horizontal del lado derecho)
# Disposición de paneles:
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# Ejecutar Claude en el primer panel
claude
# Cambiar al segundo panel (Ctrl+B flecha derecha)
tail -f logs/app.log
# Cambiar al tercer panel
npm test -- --watch
```
### Flujo de trabajo 3: Desarrollo remoto
La característica más poderosa de Tmux es la persistencia de sesiones, especialmente adecuada para desarrollo remoto por SSH:
```bash
# Conectarse al servidor remoto
ssh user@server
# Crear sesión de tmux
tmux new -s remote-claude
# Iniciar Claude
claude
# Desconectar SSH (Claude continúa ejecutándose)
# Ctrl+B d
exit
# Reconectarse más tarde
ssh user@server
tmux attach -t remote-claude
# La sesión de Claude se restaura completamente
```
### Flujo de trabajo 4: Monitoreo de Agent Teams
Usar tmux para monitorear todos los Teammates de Agent Teams:
```bash
# Iniciar Claude en modo tmux
claude --teammate-mode tmux
# Crear Agent Team
# "Crea un agent team para revisar el código..."
# En este momento la pantalla se divide automáticamente, un panel por Teammate
# Puede hacer clic en diferentes paneles para comunicarse directamente con el Teammate correspondiente
```
## Solución de problemas
### Problemas comunes
| Problema | Solución |
| --------------------------------------- | ------------------------------------------------------------- |
| Colores que se muestran incorrectamente | Asegúrese de que `TERM=xterm-256color` |
| El ratón no funciona | Agregue `set -g mouse on` a la configuración |
| Problemas de copiar y pegar | Use `Enter` para copiar en modo copia |
| La sesión desapareció | Verifique `tmux ls`, posiblemente fue un reinicio del sistema |
### Limpiar sesiones huérfanas
Claude Code a veces deja sesiones de tmux sin limpiar:
```bash
# Listar todas las sesiones
tmux ls
# Terminar una sesión específica
tmux kill-session -t session-name
# Terminar todas las sesiones (¡use con precaución!)
tmux kill-server
```
### Usuarios de iTerm2
Si utiliza iTerm2 en macOS, puede usar su integración nativa:
```bash
# Usar el modo de integración tmux de iTerm2
tmux -CC
# O en Claude Code
claude --teammate-mode tmux
```
iTerm2 convierte automáticamente los paneles de tmux en pestañas y pantallas divididas nativas.
## Mi experiencia de uso
### Cuándo usar Tmux
| Escenario | ¿Se necesita Tmux? |
| -------------------------------------- | --------------------- |
| Conversación simple y única con Claude | No es necesario |
| Tareas de larga duración | Necesario |
| Agent Teams | Altamente recomendado |
| Desarrollo remoto | Imprescindible |
| Múltiples proyectos en paralelo | Recomendado |
### Configuración mínima
Si no desea complicarse con la configuración, solo necesita recordar estos comandos:
```bash
# Crear sesión
tmux new -s work
# Desconectar (ejecutar en segundo plano)
Ctrl+B d
# Reconectarse
tmux attach -t work
# Dividir paneles
Ctrl+B % # División izquierda-derecha
Ctrl+B " # División arriba-abajo
# Cambiar entre paneles
Ctrl+B teclas de dirección
```
### Mejores combinaciones con Claude Code
1. **Worktree + Tmux**: cada worktree en una sesión de tmux independiente
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux**: gestión visual de todos los Teammates
```bash
claude --teammate-mode tmux
```
3. **Tareas largas + desconexión**: inicie y desconéctese, vuelva más tarde a verificar
```bash
# Iniciar
tmux new -s migration
claude
# "Ejecutar migración de base de datos..."
# Ctrl+B d
# Varias horas después
tmux attach -t migration
```
## Conclusión
Tmux es una herramienta clave para el uso eficiente de Claude Code, especialmente en los siguientes escenarios:
| Punto clave | Descripción |
| ----------------- | ------------------------------------------------------- |
| **Persistencia** | Las sesiones no se pierden al desconectarse |
| **Paralelismo** | Gestione múltiples instancias de Claude simultáneamente |
| **Visualización** | Visualización split-pane de Agent Teams |
Solo tres comandos básicos para comenzar:
* `tmux new -s name` crear sesión
* `Ctrl+B d` desconectar sesión
* `tmux attach -t name` reconectarse
***
**Lecturas relacionadas**:
* [Guía completa de Claude Agent Teams](/es/docs/notes/claude-agent-teams) -- Agent Teams necesita Tmux para el modo split-pane
* [Guía completa de Claude Worktree](/es/docs/notes/claude-worktree) -- Worktree puede combinarse con Tmux para ejecución en segundo plano
**Materiales de referencia**:
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**Tutoriales en video**:
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) -- Fundamentos de Tmux
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) -- Flujo de trabajo de integración con Claude Code
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) -- Técnicas de configuración avanzada
# Sprint MVP: Completar las funciones principales en dos semanas
Este es un artículo de prueba.
# Registrarse en Apple Developer Program
Para publicar tu app en la App Store, el primer paso es registrarte en el Apple Developer Program (Programa de Desarrolladores de Apple). Cuesta ¥688 yuanes (US$99) al año, una inversión inevitable para todo desarrollador iOS independiente.
En este artículo te explicaremos las diferencias entre los tipos de cuenta, qué necesitas preparar antes del registro y el proceso completo paso a paso.
## Comparación de tipos de cuenta
El Apple Developer Program ofrece tres tipos de cuenta, diseñados para diferentes escenarios de desarrollo:
| Característica | Cuenta Individual | Cuenta de Organización | Cuenta Enterprise |
| ------------------------ | ------------------------------------------ | --------------------------------------------------- | ----------------------------------------- |
| Cuota anual | ¥688($99) | ¥688($99) | ¥1,988($299) |
| Publicar en App Store | ✅ | ✅ | ❌(solo distribución interna) |
| Nombre del desarrollador | Nombre personal | Nombre de la organización/empresa | Nombre de la organización |
| Gestión de equipo | ❌ | ✅ | ✅ |
| Número D-U-N-S | No requerido | Requerido | Requerido |
| Tiempo de revisión | Rápido(generalmente menos de 48 horas) | Más lento(requiere verificación de la organización) | Más lento |
| Ideal para | Desarrolladores independientes, individuos | Empresas, estudios | Aplicaciones internas de grandes empresas |
**La elección del desarrollador independiente**: Si eres desarrollador individual, elige directamente la **cuenta Individual**. Es el proceso más simple, la revisión más rápida y las funciones son más que suficientes. El nombre del desarrollador que aparece en la App Store será tu nombre real.
## Preparación previa al registro
### Requisitos indispensables
Antes de comenzar el registro, asegúrate de tener preparado lo siguiente:
* **Apple ID**: Si aún no tienes uno, créalo en [appleid.apple.com](https://appleid.apple.com). Se recomienda usar un correo electrónico de uso frecuente, ya que todas las notificaciones relacionadas con el desarrollo se enviarán a ese correo.
* **Autenticación de dos factores**: Tu Apple ID debe tener activada la autenticación de dos factores (Two-Factor Authentication). En iPhone, ve a "Ajustes → Apple ID → Inicio de sesión y seguridad → Autenticación de dos factores" para activarla.
* **Dispositivo Apple**: El proceso de registro requiere completar la verificación de identidad en un iPhone o iPad, donde necesitarás descargar la app Apple Developer.
### Requisitos adicionales para cuenta de Organización
Si vas a registrar una cuenta de Organización, también necesitarás:
* **Número D-U-N-S**: Solicítalo con anticipación en el sitio web de Dun & Bradstreet; la revisión toma de 5 a 14 días hábiles.
* **Identidad del representante legal**: La persona que registra debe ser el representante legal de la organización o una persona autorizada.
* **Información de la organización**: Incluyendo dirección registrada, nombre del representante legal, datos de contacto, etc.
## Preparación del equipo de desarrollo
Registrar la cuenta de desarrollador es solo el primer paso; el desarrollo para iOS también requiere algunas herramientas de hardware y software.
### Dispositivos indispensables
* **Mac** — Xcode solo funciona en macOS, este es un requisito obligatorio. Se recomienda un Mac con Apple Silicon (chips serie M), que ofrece compilación rápida y puede ejecutar el simulador de iOS directamente. Un MacBook Air serie M es suficiente para el desarrollo independiente; si el presupuesto es limitado, considera un Mac mini.
* **iPhone / iPad (recomendado pero no obligatorio)** — El simulador cubre la mayoría de los escenarios de depuración, pero las pruebas en dispositivo real son insustituibles para rendimiento, sensores (cámara/GPS/NFC), notificaciones push, etc. Incluso sin una cuenta de desarrollador de pago, puedes depurar en un dispositivo real con un Apple ID gratuito (aunque con limitaciones como la re-firma cada 7 días; consulta las preguntas frecuentes al final).
### Herramientas de desarrollo
* **Xcode** — El IDE oficial de Apple, se descarga gratis desde Mac App Store. Es bastante grande (aproximadamente 12GB+), así que ten paciencia durante la primera instalación.
* **Apple Developer App** — Para registrar tu cuenta, ver videos de la WWDC y consultar documentación.
* **TestFlight** — Herramienta de distribución para pruebas beta, el canal oficial para invitar usuarios a probar tu app.
### Consideraciones importantes
* Las versiones de macOS y Xcode deben mantenerse actualizadas. Apple lanza una nueva versión de Xcode después de cada WWDC, que generalmente requiere una de las 1-2 versiones más recientes de macOS.
* Xcode se actualiza con frecuencia y ocupa mucho espacio; se recomienda reservar suficiente espacio en disco (al menos 50GB o más).
* Si tu app utiliza funciones de hardware (cámara, Bluetooth, NFC, etc.), las pruebas en dispositivo real son indispensables.
* Si no tienes una Mac, los servicios de Mac en la nube (como MacStadium, AWS EC2 Mac) son una alternativa, aunque la experiencia no es tan buena como en un dispositivo nativo.
## Proceso de registro
### Paso 1: Descargar Apple Developer App
En tu iPhone o iPad, abre la App Store, busca "Apple Developer" y descárgala.
### Paso 2: Iniciar sesión y comenzar el registro
Abre la app Apple Developer e inicia sesión con tu Apple ID. Toca la pestaña "Cuenta" y luego selecciona "Inscribirse en el Apple Developer Program".
### Paso 3: Completar información y verificación de identidad
Sigue las instrucciones para completar tu información personal:
1. **Confirmar datos de identidad**: Nombre, dirección y otra información básica.
2. **Verificación de identidad**: Según tu región, la app puede solicitarte que fotografíes un documento de identidad oficial (pasaporte, licencia de conducir, etc.) o que te tomes una selfie para la verificación.
3. **Aceptar el acuerdo**: Lee y acepta el Acuerdo de Licencia del Apple Developer Program.
> La verificación de identidad debe realizarse en un lugar con buena iluminación para asegurar fotos nítidas. Todo el proceso de registro debe completarse en el mismo dispositivo.
### Paso 4: Pagar la cuota anual
Una vez confirmada la información, paga la cuota anual de ¥688 ($99). Se acepta el método de pago vinculado a tu Apple ID. Recibirás un correo de confirmación después del pago.
### Paso 5: Esperar la revisión
* **Cuenta Individual**: Generalmente se aprueba en menos de 48 horas. En mi caso, pagué el 14 de marzo y a la mañana siguiente del 15 de marzo ya había recibido el correo de bienvenida, en menos de 24 horas.
* **Cuenta de Organización**: Apple verificará la información de la organización y el número D-U-N-S, lo que puede tomar más tiempo.
Una vez aprobada la revisión, podrás iniciar sesión en el panel de desarrollador en [developer.apple.com](https://developer.apple.com) y acceder a todos los recursos de desarrollo.
## Gestión de suscripción y renovación
El Apple Developer Program funciona con suscripción anual, con renovación automática de ¥688 al año.
### Renovación automática
La renovación automática está activada por defecto; el cobro se realizará desde el método de pago vinculado a tu Apple ID antes del vencimiento. Se recomienda mantener la renovación automática activa para evitar que la cuenta expire y afecte las apps publicadas (consulta las preguntas frecuentes a continuación).
### Cancelar o gestionar la suscripción
Si necesitas modificar la configuración de renovación, abre "Ajustes → Apple ID → Suscripciones" en tu iPhone y busca Apple Developer Program para gestionarlo.
## Preguntas frecuentes
**P: ¿Qué pasa si olvido renovar?**
Cuando la cuenta expira, tus apps serán retiradas de la App Store, pero no se eliminarán. Al volver a pagar, las apps se restaurarán. Sin embargo, durante ese período las descargas y actualizaciones de los usuarios se verán afectadas, por lo que se recomienda activar la renovación automática.
**P: ¿Cómo solicito un número D-U-N-S?**
Visita el siguiente enlace, completa la información de tu empresa y envía la solicitud. La revisión generalmente toma de 5 a 14 días hábiles. Las cuentas individuales no necesitan este número.
**P: ¿Cómo contacto a Apple si tengo problemas?**
Visita el [Soporte de Apple Developer](https://developer.apple.com/contact/), donde puedes comunicarte por chat en línea o por teléfono. El soporte para la región de China ofrece servicio en chino con buenos tiempos de respuesta.
**P: ¿Puedo empezar a desarrollar antes de registrarme?**
Puedes empezar a experimentar, pero debes conocer las limitaciones de la cuenta gratuita. Con un Apple ID gratuito ya puedes escribir código en Xcode, depurar con el simulador y también instalar apps en tu propio dispositivo. Para aprender Swift y validar ideas básicas de UI, es más que suficiente.
Sin embargo, la cuenta gratuita tiene varias limitaciones: las apps instaladas en dispositivos reales necesitan recompilarse e instalarse cada 7 días, hay un máximo de 3 dispositivos por plataforma, y funciones como notificaciones push, iCloud, TestFlight y compras dentro de la app no están disponibles. Si tu app necesita estas funciones, necesitarás la membresía de pago desde la etapa de desarrollo, no solo para publicar en la App Store.
Recomendación: si solo estás aprendiendo Swift y haciendo demos, puedes usar la cuenta gratuita; una vez que comiences un proyecto formal, regístrate cuanto antes en la membresía de pago para evitar que las limitaciones retrasen tu progreso.
# Validación de ideas: De la inspiración difusa a una dirección ejecutable
Este es un artículo de prueba.
# Selección de tecnología: Por qué elegir Next.js + Supabase
Este es un artículo de prueba.
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# Introducción al Concepto
## Introducción
Cuando has configurado un flujo de trabajo perfecto en Claude Code —comandos personalizados, hooks de revisión de código, Skills dedicados— podrías preguntarte: ¿puedo empaquetar todo esto y compartirlo con mi equipo o la comunidad?
Ese es exactamente el problema que Plugin resuelve.
Si los Skills son "manuales de instrucciones" para la IA, entonces un Plugin es una "caja de herramientas": agrupa Skills, Commands, Hooks, servidores MCP y todas las demás configuraciones para que puedas instalar y distribuir todo con un solo comando.
## Entendiendo Plugin
Imagina que eres un artesano experimentado que ha acumulado un conjunto confiable de herramientas a lo largo de los años: martillos, sierras, reglas, diversos destornilladores. Cada vez que cambias de banco de trabajo, tienes que llevar cada herramienta una por una y reorganizar todo. Un Plugin es como una caja de herramientas bien diseñada: no solo contiene todas tus herramientas, sino que las mantiene ordenadas por categoría, listas para usar donde sea que las lleves.
Desde una perspectiva técnica, Plugin es el mecanismo de empaquetado de extensiones de Claude Code. Un Plugin puede contener:
| Componente | Función | Ubicación del archivo |
| ------------------ | --------------------------------------- | --------------------- |
| **Comandos slash** | Puntos de entrada para acciones rápidas | `commands/` |
| **Subagents** | Sub-agentes especializados | `agents/` |
| **Skills** | Paquetes de conocimiento para IA | `skills/` |
| **Hooks** | Scripts de automatización por eventos | `hooks/` |
| **Servidores MCP** | Conexiones a sistemas externos | `.mcp.json` |
| **Servidores LSP** | Configuración de servidores de lenguaje | `.lsp.json` |
Estos componentes trabajan juntos para formar una solución completa de flujo de trabajo.
## Plugin vs Configuración independiente
En Claude Code, puedes colocar configuraciones en el directorio `.claude/` del proyecto o empaquetarlas como Plugin. La diferencia fundamental radica en el **método de distribución** y el **espacio de nombres**:
| Aspecto | Configuración independiente (`.claude/`) | Plugin |
| ----------------------- | -------------------------------------------------------- | --------------------------------------------- |
| Nombre del comando | `/hello` | `/plugin-name:hello` |
| Caso de uso | Flujos personales, configuración específica del proyecto | Compartir en equipo, distribución comunitaria |
| Gestión de versiones | Gestionado con el código del proyecto | Soporta Semantic Versioning |
| Método de actualización | Sincronización manual | Soporta actualizaciones automáticas |
| Manejo de conflictos | Puede entrar en conflicto con otras configuraciones | Aislamiento por espacio de nombres |
**Cuándo elegir Plugin**:
* Necesitas compartir configuraciones de flujo de trabajo con miembros del equipo
* Quieres reutilizar el mismo conjunto de herramientas en múltiples proyectos
* Planeas distribuir configuraciones a la comunidad
* Necesitas control de versiones y actualizaciones automáticas
**Cuándo usar configuración independiente**:
* Experimentos rápidos de uso personal
* Configuraciones específicas del proyecto que no necesitan reutilización
* Comandos simples de uso único
## Estructura de directorios del Plugin
La estructura estándar de un Plugin es la siguiente:
```
my-plugin/
├── .claude-plugin/ # 元数据目录
│ └── plugin.json # 必需:插件清单
├── commands/ # 斜杠命令
│ ├── review.md
│ └── deploy.md
├── agents/ # 子代理
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # 事件钩子
│ └── hooks.json
├── scripts/ # 辅助脚本
│ └── format-code.sh
├── .mcp.json # MCP 服务器配置
└── .lsp.json # LSP 服务器配置
```
**Notas importantes**:
* `plugin.json` debe colocarse dentro del directorio `.claude-plugin/`
* Los demás directorios (commands, agents, skills, etc.) van en la raíz del plugin
* No coloques directorios funcionales dentro de `.claude-plugin/`
## Archivo de configuración principal
El núcleo de un Plugin es `.claude-plugin/plugin.json`, que define los metadatos y las rutas de los componentes del plugin:
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "一个示例插件",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| Campo | Requerido | Descripción |
| ------------- | --------- | --------------------------------------------------------------- |
| `name` | Sí | Identificador único del plugin, usa letras minúsculas y guiones |
| `version` | No | Número de versión semántica |
| `description` | No | Descripción breve del plugin |
| `author` | No | Información del autor |
| `keywords` | No | Etiquetas para descubrimiento |
| `commands` | No | Ruta del archivo o directorio de comandos |
| `agents` | No | Ruta del archivo o directorio de agentes |
| `skills` | No | Ruta del directorio de Skills |
| `hooks` | No | Ruta de configuración de hooks |
| `mcpServers` | No | Ruta de configuración MCP |
## Alcances de instalación
Plugin soporta cuatro alcances de instalación para adaptarse a diferentes casos de uso:
| Alcance | Archivo de configuración | Propósito |
| --------- | ----------------------------- | ------------------------------------------------------- |
| `user` | `~/.claude/settings.json` | Plugins personales, disponibles en todos los proyectos |
| `project` | `.claude/settings.json` | Plugins de equipo, compartidos vía control de versiones |
| `local` | `.claude/settings.local.json` | Específico del proyecto, en gitignore |
| `managed` | `managed-settings.json` | Gestión empresarial (solo lectura) |
El alcance de instalación predeterminado es `user`. Si quieres agregar la configuración del plugin a Git para uso del equipo, elige el alcance `project`.
## Ventajas principales
### Aislamiento de espacios de nombres
Los comandos del Plugin llevan un prefijo de espacio de nombres (por ejemplo, `/my-plugin:review`), evitando conflictos de nombres con otros plugins o configuraciones del proyecto. Esto es especialmente importante en la colaboración en equipo: plugins desarrollados por diferentes equipos pueden coexistir sin problemas.
### Gestión de versiones
Plugin soporta Semantic Versioning, lo que te permite:
* Rastrear el historial de cambios del plugin
* Volver a versiones anteriores cuando sea necesario
* Recibir automáticamente actualizaciones compatibles
### Distribución sencilla
A través del Plugin Marketplace, puedes:
* Alojar plugins en GitHub
* Permitir que los usuarios instalen con un comando simple
* Manejar automáticamente dependencias y actualizaciones
### Colaboración en equipo
Plugin es particularmente adecuado para escenarios de equipo:
* Unificar la cadena de herramientas de desarrollo del equipo
* Los nuevos miembros obtienen todas las herramientas con un solo comando
* La gestión centralizada de configuración reduce el trabajo duplicado
## Ecosistema de Plugin
El ecosistema de Plugin de Claude Code está creciendo rápidamente. A principios de 2025, el ecosistema ha alcanzado una escala considerable:
* **229+ plugins** activos en el ecosistema
* **239 Agent Skills** distribuidos en el marketplace
* **200+ servidores MCP** preconstruidos en el toolkit de Docker
**Recursos oficiales**:
| Recurso | Enlace | Descripción |
| ---------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Repositorio oficial de Skills |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Directorio oficial de plugins |
| Docker MCP Toolkit | [Sitio web](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ servidores MCP preconstruidos |
**Selecciones de la comunidad**:
| Recurso | Enlace | Descripción |
| ---------------------- | ----------------------------------------------------------------- | --------------------------------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 plugins recopilados automáticamente |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Compilación de mejores prácticas |
| claude-plugins.dev | [Sitio web](https://claude-plugins.dev/) | Registro comunitario y CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agentes + 15 orquestadores |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 agentes especializados |
## Relación con otras funcionalidades
Plugin es un concepto de "contenedor" que puede incluir otras funcionalidades del ecosistema de Claude Code:
```
Plugin (Contenedor)
├── Skills (Paquetes de conocimiento)
├── Commands (Comandos rápidos)
├── Agents (Sub-agentes)
├── Hooks (Hooks de eventos)
└── MCP/LSP (Conexiones externas)
```
Entender esta jerarquía es importante:
* **Skills** enseñan a Claude cómo hacer algo
* **Commands** proporcionan puntos de entrada rápidos
* **Agents** manejan tareas independientes y especializadas
* **Hooks** habilitan la automatización basada en eventos
* **Plugin** empaqueta todo esto junto para facilitar la distribución y gestión
### Skills vs Plugins
La diferencia entre Skills y Plugins puede resultar confusa al principio. Según el análisis de [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins):
| Característica | Skills | Plugins |
| ---------------- | ------------------------------------------- | ----------------------------------------- |
| **Alcance** | Todos los productos Claude (Web, API, Code) | Solo Claude Code |
| **Contenido** | Guías Markdown + scripts opcionales | Commands, Agents, Hooks, MCP, Skills |
| **Activación** | Automática (el modelo decide cuándo usar) | Variable (depende del tipo de componente) |
| **Mejor para** | Enseñar conocimiento de dominio a Claude | Extender el entorno de Claude Code |
| **Distribución** | Repositorios GitHub, sistema de archivos | Marketplace descentralizado |
**Perspectiva clave**: Los Skills son activados automáticamente por el modelo sin invocación manual; los Plugins son un mecanismo de empaquetado que resuelve el desafío de la distribución compartida. Ambos pueden usarse juntos: un Plugin puede contener Skills.
## Resumen
Claude Code Plugin es esencialmente un **mecanismo de empaquetado y distribución de flujos de trabajo**. Resuelve los problemas de reutilización de configuraciones y colaboración en equipo, permitiéndote compartir tu cadena de herramientas cuidadosamente elaborada con más personas.
Recuerda tres palabras clave:
| Palabra clave | Significado |
| ---------------- | ----------------------------------------------------------------- |
| **Empaquetado** | Integra múltiples componentes de configuración en una sola unidad |
| **Aislamiento** | Los espacios de nombres evitan conflictos |
| **Distribución** | Compartir fácilmente a través del Marketplace |
Ahora que comprendes los conceptos, el siguiente artículo, [Guía práctica de Claude Code Plugin](/es/docs/notes/claude-plugin/practice), te guiará en la práctica: crear un Plugin desde cero, publicar en el Marketplace y las mejores prácticas para la colaboración en equipo.
Si aún no estás familiarizado con los componentes que un Plugin puede contener, te recomendamos leer [¿Qué son los Claude Skills?](/es/docs/notes/claude-skills/concept) para entender los conceptos fundamentales de Skills.
# Guía Práctica
## Repaso rápido
En el artículo anterior, exploramos los conceptos centrales de los Plugins: son el mecanismo de Claude Code para empaquetar y distribuir flujos de trabajo, integrando Commands, Skills, Agents, Hooks y otros componentes en una sola unidad para compartir con el equipo y distribuir en la comunidad. Este artículo adopta un enfoque práctico y te guiará a través del proceso completo, desde la creación hasta la publicación.
## Crea tu primer Plugin
### Paso 1: Crear la estructura de directorios
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### Paso 2: Crear el manifiesto del plugin
Define los metadatos de tu plugin en `.claude-plugin/plugin.json`:
```json
{
"name": "my-first-plugin",
"description": "我的第一个 Claude Code 插件",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### Paso 3: Agregar comandos slash
Crea archivos Markdown en el directorio `commands/`. Cada archivo corresponde a un comando:
`commands/hello.md`:
```markdown
---
description: 向用户发送友好的问候
---
# Hello 命令
请热情地问候用户,并询问今天可以帮助他们做什么。
```
### Paso 4: Probar el plugin
Usa la bandera `--plugin-dir` para cargar tu plugin local y probarlo:
```bash
claude --plugin-dir ./my-first-plugin
```
Ejecuta el comando en Claude Code:
```
/my-first-plugin:hello
```
### Paso 5: Agregar argumentos al comando
Los comandos admiten recibir argumentos proporcionados por el usuario. Actualiza `hello.md`:
```markdown
---
description: 向指定用户发送个性化问候
---
# Hello 命令
请热情地问候名为 "$ARGUMENTS" 的用户,并询问今天可以帮助他们做什么。
如果用户没有提供名字,就使用"朋友"作为称呼。
```
Prueba el comando con argumentos:
```
/my-first-plugin:hello 小明
```
**Marcadores de posición admitidos para argumentos**:
* `$ARGUMENTS` - Toda la entrada del usuario
* `$1`, `$2`, `$3` - Argumentos individuales
## Agregar más componentes
### Agregar Skills
Crea un directorio `skills/`. Cada Skill es una carpeta que contiene un archivo `SKILL.md`:
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md`:
```yaml
---
name: code-review
description: 审查代码质量、安全性和可维护性
---
当审查代码时,请检查以下方面:
1. **代码组织**:结构是否清晰
2. **错误处理**:异常是否被妥善处理
3. **安全隐患**:是否存在安全漏洞
4. **测试覆盖**:关键逻辑是否有测试
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Agregar Subagents
Crea un directorio `agents/`:
`agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家。
当被调用时:
1. 运行 git diff 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(如注入、敏感信息泄露)
- 性能优化机会
```
### Agregar Hooks
Los Hooks te permiten ejecutar scripts automáticamente cuando ocurren eventos específicos. Crea `hooks/hooks.json`:
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**Importante**: Usa la variable de entorno `${CLAUDE_PLUGIN_ROOT}` para referenciar archivos dentro del directorio del plugin, asegurando que las rutas se resuelvan correctamente sin importar dónde esté instalado el plugin.
Crea el script correspondiente `scripts/format-code.sh`:
```bash
#!/bin/bash
# 格式化代码
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
Recuerda asignar permisos de ejecución:
```bash
chmod +x scripts/format-code.sh
```
### Agregar servidores MCP
Si tu plugin necesita conectarse a sistemas externos, crea `.mcp.json`:
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## Estructura completa del Plugin
Un Plugin con funcionalidad completa podría verse así:
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单
├── commands/
│ ├── review.md # 代码审查命令
│ ├── deploy.md # 部署命令
│ └── test.md # 测试命令
├── agents/
│ ├── code-reviewer.md # 代码审查代理
│ └── debugger.md # 调试代理
├── skills/
│ └── code-standards/
│ └── SKILL.md # 代码规范知识
├── hooks/
│ └── hooks.json # 事件钩子配置
├── scripts/
│ ├── format-code.sh # 格式化脚本
│ └── run-tests.sh # 测试脚本
├── .mcp.json # MCP 配置
├── LICENSE
├── README.md
└── CHANGELOG.md
```
El `plugin.json` correspondiente:
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "开发者工具箱:代码审查、测试、部署一站式解决方案",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## Publicar en el Marketplace
### Qué es el Marketplace
El Marketplace es el centro de distribución de Plugins. Puedes entenderlo como una "tienda de plugins": los usuarios pueden instalar tus plugins publicados con un simple comando.
### Crear la configuración del Marketplace
Crea `.claude-plugin/marketplace.json` en tu repositorio de GitHub:
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "开发者工具箱",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "文档生成工具"
}
]
}
```
### Tipos de fuente de plugins
El Marketplace admite múltiples tipos de fuentes:
**Ruta relativa** (plugins dentro del mismo repositorio):
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**Repositorio de GitHub**:
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**Cualquier repositorio Git**:
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### Flujo de publicación
1. **Crear un repositorio en GitHub**
2. **Subir el código**:
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **Los usuarios agregan tu Marketplace**:
```bash
/plugin marketplace add your-username/your-repo
```
4. **Los usuarios instalan plugins**:
```bash
/plugin install dev-toolkit@your-marketplace
```
## Instalar y gestionar Plugins
### A través del menú interactivo
```bash
/plugin
```
Esto abre una interfaz interactiva donde puedes explorar, instalar, habilitar y deshabilitar plugins.
### A través de la línea de comandos
**Agregar un Marketplace**:
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # 本地
```
**Instalar plugins**:
```bash
# 安装到用户范围(默认)
/plugin install formatter@my-marketplace
# 安装到项目范围(团队共享)
/plugin install formatter@my-marketplace --scope project
# 安装到本地范围(gitignored)
/plugin install formatter@my-marketplace --scope local
```
**Otros comandos de gestión**:
```bash
/plugin enable # 启用插件
/plugin disable # 禁用插件
/plugin uninstall # 卸载插件
/plugin update # 更新插件
```
### Validar plugins
Verifica que la configuración de tu plugin sea correcta antes de publicar:
```bash
claude plugin validate .
```
O dentro de Claude Code:
```
/plugin validate .
```
## Configuración para colaboración en equipo
### Compartir la configuración de plugins en un proyecto
Confirma la configuración de plugins en el control de versiones para que los miembros del equipo la obtengan automáticamente:
`.claude/settings.json`:
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
Después de que los miembros del equipo clonen el proyecto, estos plugins estarán disponibles automáticamente.
### Restricciones empresariales del Marketplace
Para entornos empresariales que requieren un control estricto, puedes restringir los Marketplaces permitidos en la configuración administrada:
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
Establecerlo como un arreglo vacío `[]` desactiva completamente los plugins externos.
## Referencia de comandos CLI
| Comando | Descripción |
| -------------------------------------- | ----------------------------------------------------------- |
| `/plugin` | Abrir la interfaz de gestión interactiva |
| `/plugin install @` | Instalar un plugin |
| `/plugin uninstall ` | Desinstalar un plugin |
| `/plugin enable ` | Habilitar un plugin |
| `/plugin disable ` | Deshabilitar un plugin |
| `/plugin update ` | Actualizar un plugin |
| `/plugin validate .` | Validar la configuración del plugin en el directorio actual |
| `/plugin marketplace add ` | Agregar un Marketplace |
| `/plugin marketplace list` | Listar los Marketplaces agregados |
| `/plugin marketplace update` | Actualizar la caché del Marketplace |
| `/plugin marketplace remove ` | Eliminar un Marketplace |
## Mejores prácticas
### Mejores prácticas de desarrollo
1. **Mantén los Skills enfocados**: Cada Skill debe hacer bien una sola cosa; evita diseños que intenten abarcar todo
2. **Escribe descripciones claras**: Ayuda a Claude a entender cuándo usar tus componentes
3. **Prueba primero con tu equipo**: Valida internamente antes de distribuir a la comunidad
4. **Documenta los cambios de versión**: Registra los cambios de cada versión en CHANGELOG.md
### Mejores prácticas de estructura de directorios
* Coloca `commands/`, `agents/`, `skills/` en el directorio raíz del plugin
* Solo coloca `plugin.json` dentro del directorio `.claude-plugin/`
* Usa `${CLAUDE_PLUGIN_ROOT}` para referenciar archivos dentro del plugin
* Nunca uses `../` para acceder a archivos fuera del plugin
### Mejores prácticas de Hooks
1. Los scripts deben ser ejecutables: `chmod +x script.sh`
2. Usa un shebang para declarar el intérprete: `#!/bin/bash`
3. Usa la variable `${CLAUDE_PLUGIN_ROOT}` para asegurar que las rutas sean correctas
4. Prueba los scripts de forma independiente antes de integrarlos en Hooks
### Mejores prácticas de gestión de versiones
Sigue el Semantic Versioning:
* **MAJOR** (1.0.0 → 2.0.0): Cambios que rompen la compatibilidad
* **MINOR** (1.0.0 → 1.1.0): Nuevas funcionalidades (compatible con versiones anteriores)
* **PATCH** (1.0.0 → 1.0.1): Corrección de errores (compatible con versiones anteriores)
## Solución de problemas comunes
| Problema | Causa posible | Solución |
| ----------------------- | ------------------------------------ | ---------------------------------------------------------------------------- |
| El plugin no se carga | plugin.json con formato incorrecto | Valida con `claude plugin validate` |
| El comando no aparece | Estructura de directorios incorrecta | Asegúrate de que `commands/` esté en la raíz, no dentro de `.claude-plugin/` |
| Los Hooks no se activan | El script no es ejecutable | Ejecuta `chmod +x script.sh` |
| Ruta no encontrada | Se usaron rutas relativas | Cambia a `${CLAUDE_PLUGIN_ROOT}` |
| El servidor MCP falla | Variables de entorno no configuradas | Verifica la configuración de rutas en `.mcp.json` |
## Migrar desde una configuración existente
Si ya tienes configuraciones en el directorio `.claude/`, sigue estos pasos para migrarlas a un Plugin:
1. **Crear la estructura del Plugin**:
```bash
mkdir my-plugin/.claude-plugin
```
2. **Crear plugin.json**:
```json
{
"name": "my-plugin",
"description": "从现有配置迁移的插件",
"version": "1.0.0"
}
```
3. **Copiar los archivos existentes**:
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **Migrar Hooks**:
Copia la configuración de `hooks` desde `.claude/settings.json` a `hooks/hooks.json`
5. **Probar**:
```bash
claude --plugin-dir ./my-plugin
```
## Recursos de aprendizaje
### Documentación oficial
| Recurso | Enlace | Descripción |
| --------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------- |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | Documentación de referencia de plugins |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | Guía de creación de plugins |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | Mejores prácticas oficiales |
| Agent Skills Standard | [agentskills.io](https://agentskills.io) | Especificación del estándar abierto |
### Repositorios oficiales
| Recurso | Enlace | Descripción |
| ---------------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------ |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Repositorio oficial de Skills |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Catálogo oficial de plugins |
| Docker MCP Toolkit | [Sitio web](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | Más de 200 MCPs preconstruidos |
### Recursos de la comunidad
| Recurso | Enlace | Descripción |
| --------------------------- | ---------------------------------------------------------------------------- | ------------------------------------------ |
| claude-plugins.dev | [Sitio web](https://claude-plugins.dev/) | Registro comunitario y CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | Colección de 243 plugins |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Compilación de mejores prácticas |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agentes + 15 orquestadores |
| Tutorial de jeremylongshore | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | Cientos de plugins + tutoriales de Jupyter |
### Lectura recomendada
| Artículo | Fuente |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## Perspectivas a futuro
El sistema de Plugins representa un salto cualitativo en la capacidad de extensión de Claude Code. A medida que la comunidad crece, podemos esperar:
* **Un ecosistema de plugins más rico**: Cubriendo una amplia variedad de escenarios de desarrollo y flujos de trabajo
* **Funcionalidades de nivel empresarial**: Gestión de permisos y capacidades de auditoría más completas
* **Compatibilidad multiplataforma**: El estándar abierto de Skills ya ha sido adoptado por múltiples proveedores
Ahora es el momento perfecto para participar. Puedes comenzar con comandos simples, agregar gradualmente Skills y Hooks, y eventualmente construir una solución completa de flujo de trabajo.
Si deseas conocer más sobre el componente Subagent que los Plugins pueden incluir, lee [Qué son los Claude Code Subagents](/es/docs/notes/claude-subagent/concept).
# Introducción al concepto de Claude Code Subagent
## Introducción
Al utilizar Claude Code para manejar tareas complejas, es posible que se haya encontrado con un dilema de este tipo: el contexto de la conversación principal se hace cada vez más largo, la IA comienza a "olvidar" información importante anterior y la calidad de la respuesta disminuye gradualmente.
Subagent nació para solucionar este problema.
Si Skills es el "manual de trabajo" para Claude, entonces el Subagente es el "empleado de tiempo completo" que usted contrata: tiene su propia estación de trabajo independiente (contexto), se concentra en un tipo específico de trabajo y le informa los resultados una vez finalizado.
## Comprensión del subagente
Imagina que eres el director ejecutivo de una empresa. Cuando la empresa es pequeña, te encargas de todo tú mismo. Pero a medida que su negocio se expande, comienza a contratar empleados de tiempo completo: contadores para finanzas, recursos humanos para contratación e ingenieros para desarrollo. Cada empleado trabaja en su propia estación y le informa a usted después de completar la tarea.
El subagente desempeña exactamente este papel en Claude Code.
Desde una perspectiva técnica, los Subagentes son asistentes de IA especializados que cuentan con las siguientes características:
| Características | Descripción |
| ------------------------------ | -------------------------------------------------------------------- |
| **Contexto independiente** | Cada Subagente se ejecuta en su propia ventana contextual |
| **Capacidades especializadas** | Optimizado para tipos de tareas específicas |
| **Herramientas configurables** | Solo puede acceder a un conjunto específico de herramientas |
| **Mensajes personalizados** | Hay indicaciones especiales del sistema para guiar el comportamiento |
## ¿Por qué necesitamos un contexto independiente?
Este es el concepto de diseño central de Subagent y merece una comprensión profunda.
En una conversación normal, toda la información se acumula en el mismo contexto. Cuando Claude busca en la base del código, analiza los archivos y luego realiza modificaciones, todo este procesamiento intermedio ocupa espacio de contexto. A medida que avanza la conversación y el contexto se vuelve más complejo, Claude puede comenzar a "olvidar" información importante anterior.
El subagente cambia que:
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
El proceso de análisis del subagente no contamina la conversación principal. El diálogo principal obtuvo sólo resultados refinados, manteniendo la claridad y el enfoque.
## Tipo de subagente incorporado
Claude Code proporciona tres potentes subagentes integrados que cubren los escenarios de uso más comunes:
### Explorar subagente
**Orientación**: exploración rápida y de solo lectura del código base.
**Características**:
* Utilice el modelo Haiku (rápido, baja latencia)
* Estrictamente de solo lectura: los archivos no se pueden crear, modificar ni eliminar
* Herramientas disponibles: Glob, Grep, Read, Bash (operaciones de solo lectura)
**Cuándo utilizar**:
Cuando hace preguntas exploratorias como "¿Dónde se implementa esta función?" y "¿Cómo se manejan los errores?" Claude llamará automáticamente al subagente Explorar.
**Nivel de detalle**:
| Nivel | Descripción | Escenarios aplicables |
| ------------- | -------------------------------------- | ----------------------------------------------------------- |
| Rápido | Búsqueda rápida con exploración mínima | Consultas sencillas y específicas |
| Medio | Exploración moderada | Equilibrando velocidad e integridad |
| Muy minucioso | Análisis completo | Cuestiones complejas que requieren una comprensión profunda |
### Subagente del plan
**Posicionamiento**: Estudiar el código base y preparar un plan de implementación.
**Características**:
* Utilice el modelo Sonnet (mayores capacidades de inferencia)
* Sólo herramientas de exploración: Read, Glob, Grep, Bash
* Llamado automáticamente en modo de planificación.
**Cuándo utilizar**:
Cuando ingresa al modo de planificación y necesita que Claude realice una investigación antes de proponer un plan, Plan Subagent recopilará información automáticamente y luego dará sugerencias de planes basadas en los resultados de la investigación.
### Subagente de uso general
**Posicionamiento**: maneje tareas complejas de varios pasos.
**Características**:
* Utilice el modelo Soneto
* Acceso a todas las herramientas (incluidas lectura y escritura)
* Adecuado para tareas complejas que requieren exploración y modificación.
**Cuándo utilizar**:
Cuando la tarea implica varios pasos, requiere una búsqueda antes de modificarla, o la búsqueda inicial puede fallar y es necesario probar múltiples estrategias.
## Mi comprensión y práctica.
Si observa detenidamente los tres subagentes oficiales integrados, encontrará una cosa en común: **Todos son tareas de investigación y planificación**. Explore es responsable de explorar la base del código, Plan es responsable de hacer planes e incluso el propósito general se utiliza principalmente para investigación y análisis. Ninguno de ellos está diseñado específicamente para escribir código.
Esto confirma mi comprensión de Subagent: **El valor central de Subagent no es el "contexto limpio", sino permitir que el agente principal se concentre en hacer las cosas**.
### Modo de división del trabajo
Mi método de uso es muy simple: el subagente es responsable del trabajo de "recopilación de información", como investigación, planificación y revisión, y el agente principal es responsable de la ejecución real.
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### ¿Por qué no dejar que Subagent escriba código?
A algunas personas les gusta que el agente principal programe varios subagentes para escribir código. Creo que esto no es confiable. La razón es simple: **Falta mucho contexto**.
El contexto del subagente es independiente. No sabe qué se discutió, qué decisiones se tomaron y qué limitaciones se impusieron en la conversación principal. Pedirle que escriba código es como pedirle a un nuevo empleado que complete una tarea sin ninguna información previa: es probable que el código producido no coincida con sus expectativas.
Por el contrario, es mucho más razonable posicionar a Subagent como un "investigador":
* La tarea de investigación en sí no requiere mucho contexto.
* Se devuelve información en lugar de código, que puede ser utilizado por el agente principal según el contexto completo.
* Incluso si los resultados de la encuesta están sesgados, el agente principal puede corregirlos.
### Mi uso diario
1. **Antes de comenzar una nueva tarea**: permita que el agente de Explore comprenda rápidamente la estructura del código relevante
2. **Planificación de tareas complejas**: permita que el agente del plan analice los requisitos y formule los pasos de implementación.
3. **Revisión de código**: permita que el agente de revisión verifique la calidad del código y los problemas de seguridad.
4. **Codificación real**: el agente principal escribe código según el contexto recopilado.
Aquí está la lista de agentes que uso actualmente:
La ventaja de esto es que la ventana de contexto del agente principal permanece limpia, con sólo "la información que necesito saber" en lugar de "un montón de resultados intermedios generados durante el proceso de búsqueda del subagente".
## Comparación con otras funciones
### Subagent vs Skills
Ésta es la confusión más común. Diferencia fundamental: **Las habilidades inyectan conocimiento en Claude; Subagente crea trabajadores independientes**.
| Dimensiones | Habilidades | Subagente |
| ------------------------------- | ----------------------------------------------------- | -------------------------------------------------- |
| **Características principales** | Proporcionar experiencia e instrucciones | Agentes que realizan tareas de forma independiente |
| **Contexto** | Compartir contexto de conversación principal | Tener contexto independiente |
| **Método de activación** | Coincidencia automática basada en descripción | Delegación automática o llamada manual |
| **Escenarios aplicables** | Hacer que Claude sea mejor en ciertos tipos de tareas | Tareas independientes complejas de varios pasos |
Para decirlo en sentido figurado: las habilidades son como materiales de formación que le permiten a Claude aprender a hacer algo; El subagente es como un empleado de tiempo completo, que completa la tarea de forma independiente en su puesto de trabajo e informa los resultados.
Los dos se pueden combinar: un subagente de revisión de código puede cargar la habilidad de especificación del código para lograr el efecto combinado de "conocimiento experto + profesional".
### Subagente vs comando de barra diagonal
| Dimensiones | Subagente | Comando de barra diagonal |
| ------------------------ | ----------------------------------------- | ----------------------------------- |
| **Método de activación** | Delegación automática o llamada explícita | Entrada del manual de usuario |
| **Contexto** | Contexto independiente | Conversación principal compartida |
| **Complejidad** | Adecuado para tareas complejas | Adecuado para operaciones sencillas |
El comando de barra diagonal es una tecla de acceso directo y usted ingresa `/review` para activar una operación predefinida; El subagente es un trabajador independiente que puede completar tareas complejas de varios pasos de forma autónoma.
### Subagent vs Plugin
El complemento es un concepto de "contenedor", que puede contener subagente:
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
Puede definir un subagente en el directorio `agents/` del complemento y distribuirlo con el complemento.
## Patrón de diseño agente
Anthropic resume seis patrones de diseño Agentic principales en su documentación oficial. Comprender estos patrones puede ayudar a diseñar mejor los sistemas de subagente:
| Patrón | Idea central | Solicitud de subagente |
| ---------------------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------- |
| **Encadenamiento rápido** | Descomponer tareas complejas en múltiples pasos secuenciales | Llamada en cadena a varios subagentes |
| **Enrutamiento** | Distribuido a procesadores especializados según el tipo de entrada | Se delegan diferentes tipos de tareas a Subagentes especializados |
| **Paralelización** | Ejecutar múltiples subtareas independientes al mismo tiempo | Iniciar múltiples subagentes en paralelo |
| **Orquestador-Trabajadores** | Coordinador central asigna tareas a los trabajadores | Claude como coordinador, Subagente como trabajador |
| **Evaluador-Optimizador** | Salida del generador, optimización del evaluador | Generar Subagente + Revisar Subagente |
| **Agentes** | Agentes independientes que toman decisiones autónomas | Cada Subagente se ejecuta de forma independiente |
Estos modos se pueden utilizar en combinación. Por ejemplo, un sistema de calidad de código podría utilizar ambos:
* **Paralelización**: ejecute análisis de seguridad y análisis de rendimiento simultáneamente
* **Orquestador-Trabajadores**: el Maestro Claude coordina múltiples Subagentes especializados
* **Evaluador-Optimizador**: Revisar el código inmediatamente después de la generación
## Ventajas principales
### Protección de contexto
El mayor valor de Subagent radica en proteger el contexto de la conversación principal. Los procesos intermedios, como la búsqueda de código y el análisis de archivos, no se acumularán en el diálogo principal, lo que permitirá que el diálogo principal se centre siempre en objetivos de alto nivel.
### Capacidades de especialización
Puede crear subagentes especializados para dominios específicos, configurados con instrucciones detalladas y herramientas adecuadas. Un subagente especializado se desempeña mejor en una tarea específica que un Claude de propósito general.
### Control de permisos flexible
Cada Subagente puede tener diferentes derechos de acceso a las herramientas. Por ejemplo, la clase de exploración Subagent solo otorga permisos de solo lectura y la clase de modificación Subagent solo otorga permisos de escritura. Este control detallado mejora la seguridad.
### Reutilizabilidad
Una vez creado, el Subagent se puede reutilizar en todos los proyectos o compartir con equipos a través de complementos.
## Cuándo utilizar el subagente
**Escenarios adecuados para usar Subagent**:
* Requiere contexto independiente para ejecutar tareas.
* Las tareas son flujos de trabajo complejos de varios pasos.
* Requiere un conjunto de herramientas diferente al de la conversación principal.
* Las tareas pueden tardar mucho en ejecutarse
**Escenarios no adecuados para utilizar Subagent**:
* Consulta simple y única
* Requiere una estrecha interacción con el diálogo principal.
* Las misiones se pueden completar rápidamente
## Escenarios de aplicación típicos
### Revisión de código
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
Cuando completa un fragmento de código, puede hacer que el subagente de revisión de código lo revise en un contexto independiente sin interferir con su trabajo de desarrollo principal.
### Análisis de depuración
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
Cuando se encuentra un error, el subagente de depuración puede analizar en profundidad la causa del error, probar varias hipótesis y, finalmente, proporcionar sugerencias de reparación.
### Exploración de la base de código
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
Cuando eres nuevo en un proyecto nuevo, Explore Subagent puede mapear rápidamente tu base de código sin saturar tu conversación principal con toneladas de resultados de búsqueda.
## Recursos de aprendizaje
### Recursos oficiales
| Recursos | Enlaces | Instrucciones |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------- |
| Documentación del Código Claude | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Entrada de Documentación Oficial |
| Guía de subagente | [Documentos del Código Claude](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Documentación Oficial Subagente |
| Investigación de sistemas multiagente | [Ingeniería Antrópica](https://www.anthropic.com/engineering/multi-agent-research-system) | Detalles de la investigación sobre una mejora del rendimiento del 90,2% |
| Patrón de diseño agente | [Documentos antrópicos](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Explicación detallada de seis patrones de diseño principales |
### Recursos comunitarios
| Recursos | Enlaces | Instrucciones |
| --------------------------- | ----------------------------------------------------------------- | ------------------------------------------- |
| wshobson/agentes | [GitHub](https://github.com/wshobson/agents) | 99 Agentes + 15 Orquestadores |
| Ingeniería compuesta | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Complementos para 17 agentes especializados |
| código-claude-impresionante | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Resumen de mejores prácticas |
## Resumen
Claude Code Subagent es esencialmente un asistente de IA especializado e independiente del contexto. Resuelve el problema de la sobrecarga de información en tareas complejas mediante el aislamiento del contexto, manteniendo la conversación principal clara y enfocada en todo momento.
Recuerde tres palabras clave:
| Palabras clave | Significado |
| ----------------- | --------------------------------------------------------------------- |
| **Independiente** | Cada Subagente tiene su propia ventana contextual |
| **Especializado** | Optimizado para tipos de tareas específicas |
| **Delegación** | Claude puede delegar tareas al Subagente de forma automática o manual |
Después de comprender el concepto, el siguiente artículo "[Guía práctica del subagente de Claude Code](/es/docs/notes/claude-subagent/practice)" lo llevará a practicar: crear un subagente personalizado, configurar permisos de herramientas y mejores prácticas en proyectos reales.
Si desea conocer las Skills que puede cargar el Subagent, lea "[Qué son las Skills de Claude](/es/docs/notes/claude-skills/concept)". Si desea empaquetar Subagent para su distribución, lea "[Qué es el complemento Claude Code](/es/docs/notes/claude-plugin/concept)".
# Guía práctica del subagente de Claude Code
## Revisión rápida
En el artículo anterior, aprendimos sobre el concepto central de Subagent: es un asistente de IA especializado independiente del contexto que resuelve el problema de la sobrecarga de información en tareas complejas mediante el aislamiento del contexto. Claude Code tiene tres subagentes integrados: Explorar, Planificar y Propósito general. Este artículo lo llevará desde una perspectiva práctica a crear un subagente personalizado y dominar el uso avanzado.
## Administrar subagente
### A través del comando /agentes
La forma más sencilla es utilizar la interfaz interactiva:
```bash
/agents
```
Esto abrirá un menú donde podrá:
* Ver todos los subagentes (integrados + personalizados)
-Crear nuevo Subagente
* Editar configuración y permisos de herramientas para el subagente existente
* Eliminar subagente innecesario
* Ver qué subagente está activo cuando hay un conflicto de nombre
### A través de la gestión de archivos
El subagente se almacena como un archivo Markdown. También puede crear y editar archivos directamente.
**Ubicación de almacenamiento**:
| ubicación | camino | alcance |
| ----------------- | ------------------------------------ | --------------------------------------------------- |
| Nivel de proyecto | `.claude/agents/` | Dedicado al proyecto actual y se puede enviar a Git |
| Nivel de usuario | `~/.claude/agents/` | Disponible en todos los proyectos |
| Complemento | Directorio `agents/` del complemento | Instalado con complemento |
**Prioridad**: Nivel de proyecto > Nivel de usuario > Nivel de complemento
Cuando existe un subagente con el mismo nombre en varias ubicaciones, el que tenga mayor prioridad sobrescribirá al que tenga menor prioridad.
## Crea tu primer Subagente
### Paso 1: crear un directorio
```bash
mkdir -p .claude/agents
```
### Paso 2: crear un archivo Markdown
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Paso 3: Subagente de prueba
En Código Claude:
```
> 用 code-reviewer 代理审查我最近的修改
```
O dejar que Claude elija automáticamente:
```
> 帮我审查一下代码质量
```
Si `description` está escrito con suficiente claridad, Claude reconocerá y llamará automáticamente a su subagente.
## Explicación detallada de los campos de configuración
El archivo de configuración del subagente consta de dos partes: el frontmatter de YAML y el cuerpo de Markdown.
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| Campo | Requerido | Descripción |
| ---------------- | --------- | ---------------------------------------------------------------------------------------------------- |
| `name` | es | un identificador único, utilice letras minúsculas y guiones |
| `description` | Sí | Descripción en lenguaje natural (Claude usa esto para determinar cuándo llamar) |
| `tools` | No | Lista de herramientas separadas por comas. Si se omite, se heredan todas las herramientas |
| `model` | No | Selección de modelo: `sonnet`, `opus`, `haiku` o `inherit` |
| `permissionMode` | No | Modo de permiso (ver más abajo) |
| `skills` | No | Habilidades cargadas automáticamente (el subagente no hereda las habilidades de la sesión principal) |
### Modo de permiso
| Modo | Descripción |
| ------------------- | ---------------------------------------------- |
| `default` | Comprobación de permisos normales |
| `acceptEdits` | Aceptar automáticamente operaciones de edición |
| `bypassPermissions` | Saltar todas las comprobaciones de permisos |
| `plan` | Sólo proponer un plan, no ejecutarlo |
| `ignore` | Ignorar este subagente |
### Texto de rebajas
El texto es el mensaje del sistema del Subagente. Cuanto más detallado escriba, mejor se desempeñará el Subagent.
Un buen aviso del sistema debe incluir:
* Definición clara de roles
* Pasos de trabajo específicos
* Lista de verificación clave
* Requisitos de formato de salida
## Mecanismo de disparo
### Delegación automática
Claude decidirá automáticamente si delegar según el contenido de la tarea y el `description` del subagente.
**Consejo para fomentar el uso automático**: Utilice palabras desencadenantes en `description`:
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
o:
```yaml
description: MUST BE USED when encountering errors or test failures
```
### Llamada explícita
Dígale directamente a Claude qué subagente usar:
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
## Configuración de herramientas
### Lista de herramientas de uso común
| Herramientas | Instrucciones |
| ------------ | ------------------------------------ |
| `Read` | Leer el contenido del archivo |
| `Write` | Escribir al archivo |
| `Edit` | Editar archivo |
| `Glob` | Coincidencia de patrones de archivos |
| `Grep` | Búsqueda de expresiones regulares |
| `Bash` | Ejecutar comando de shell |
| `WebFetch` | Obtener contenido web |
| `WebSearch` | Buscar en la web |
### Estrategia de configuración de herramientas
**Subagente de solo lectura** (exploración, análisis):
```yaml
tools: Read, Grep, Glob, Bash
```
NOTA: Incluso si se incluye Bash, Subagent solo debe usarse para comandos de solo lectura (ls, git status, git log, etc.).
**Lectura y escritura de subagente** (reparación, refactorización):
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**Principio de privilegio mínimo**: Otorgue solo las herramientas necesarias para evitar operaciones accidentales.
## Plantilla práctica de subagente
### Revisor de código
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### Experto en depuración
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### Corredor de prueba
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### Generador de documentos
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### Escáner de seguridad
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## Uso avanzado
### Patrones de diseño a nivel de producción
En entornos de producción, existen varios patrones probados para la colaboración entre múltiples agentes:
#### Modo 3 Amigos
Un modelo de colaboración compuesto por tres roles: producto, arquitectura e implementación:
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| Funciones | Responsabilidades | Configuración de herramientas |
| ----------------- | -------------------------------------------------- | ----------------------------- |
| Agente PM | Definición de función, clasificación de requisitos | Leer, Búsqueda web |
| Agente Arquitecto | Diseño de soluciones técnicas | Leer, Glob, Grep |
| Código Claude | Implementación de código | Todas las herramientas |
#### Tubería de tres etapas
Divida las tareas complejas en tres etapas claras:
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
Cada etapa es responsable de un subagente dedicado y la salida sirve como entrada para la siguiente etapa.
#### Estrategia de orquestación del modelo
Se utilizan diferentes modelos en diferentes etapas para optimizar costos y efectos:
| Etapa | Modelo recomendado | Razón |
| --------------------- | ------------------ | ------------------------------------ |
| Fase de planificación | Soneto | Se requiere un razonamiento profundo |
| Fase de ejecución | haikus | Rápido, bajo costo |
| Etapa de revisión | Soneto | Se requiere juicio integral |
Ejemplo de configuración:
```yaml
---
name: quick-executor
model: haiku
---
```
### Enlace de subagente
Para flujos de trabajo complejos, se pueden vincular varios subagentes:
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### Ejecución reanudable
La ejecución del subagente se puede pausar y reanudar, manteniendo el contexto anterior completo:
**Llamada inicial**:
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**Agente de recuperación**:
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**Escenario de uso**:
* Estudios de larga duración, completados en múltiples sesiones.
* Mejoras iterativas, manteniendo el contexto.
* Flujo de trabajo de varios pasos para procesar tareas relacionadas en secuencia
### Configurar habilidades para el subagente
El subagente no hereda automáticamente las habilidades de la sesión principal. Si es necesario, declararlo explícitamente:
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### Definición dinámica CLI
No es necesario guardar el archivo, defina el subagente temporal directamente en la línea de comando:
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
Adecuado para pruebas rápidas o uso único.
## Mejores prácticas
### 1. Manténgase concentrado
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
Un Subagente que hace bien una cosa es mejor que un Subagente que hace muchas cosas.
### 2. Escribe una descripción clara
Claude usa `description` para decidir cuándo usar Subagent. Una buena descripción debería responder:
1. \*\*¿Qué hace este Subagente? \*\* Listar habilidades específicas
2. \*\*¿Cuándo se debe utilizar? \*\* Contiene palabras desencadenantes
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. Restringir el acceso a la herramienta
Otorga solo las herramientas que necesitas:
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
Esto evita que Subagent modifique archivos accidentalmente y le permite concentrarse más en su trabajo de revisión.
### 4. Escriba indicaciones detalladas del sistema
Cuanto más detalladas sean las indicaciones del sistema, mejor se desempeñará el Subagent:
* Definición clara de roles
* Pasos de trabajo específicos
* Lista de verificación clave
* Requisitos de formato de salida
### 5. Control de versiones
Confirme el subagente a nivel de proyecto en Git:
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
Los miembros del equipo obtienen automáticamente el mismo subagente después de clonar el proyecto.
## Solución de problemas comunes
| Problema | Posible causa | Solución |
| ------------------------------------- | ------------------------------------------------------ | ------------------------------------------------------------------------------------------------------- |
| El subagente no se llama | la descripción no es lo suficientemente clara | Agregue palabras desencadenantes para hacerlo más específico |
| No se llama al subagente | Ubicación incorrecta del archivo | Asegúrese de que el archivo esté en `.claude/agents/` o `~/.claude/agents/` |
| Las herramientas no están disponibles | Error de configuración del campo de herramientas | Compruebe la ortografía de los nombres de las herramientas y asegúrese de que estén separados por comas |
| La salida es inestable | El mensaje del sistema es demasiado vago | Agregue pasos específicos y requisitos de formato de salida |
| Contexto perdido | Sesión finalizada | Usando ejecución reanudable |
| Conflicto de nombres | Subagente con el mismo nombre en múltiples ubicaciones | Utilice `/agents` para ver cuál está activo |
## Compartir con el equipo
### Método 1: a través de Git
Coloque Subagent en el directorio `.claude/agents/` y envíelo al repositorio del proyecto. Los miembros del equipo se obtienen automáticamente después de la clonación.
### Método 2: a través del complemento
Coloque el Subagent en el directorio `agents/` del complemento y distribúyalo a través del mecanismo del complemento.
### Método 3: compartir a nivel de usuario
Coloque los subagentes de uso común en `~/.claude/agents/` para que estén disponibles en todos los proyectos. La sincronización entre varias máquinas se puede gestionar mediante archivos de puntos.
## Recursos de aprendizaje
### Documentación oficial
| Recursos | Enlaces | Instrucciones |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------- |
| Documentación del Código Claude | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Entrada de Documentación Oficial |
| Guía de subagente | [Documentos del Código Claude](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Detalles de configuración del subagente |
| Patrones de diseño agentes | [Documentos antrópicos](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Seis patrones de diseño básicos |
| Estudio multiagente | [Ingeniería Antrópica](https://www.anthropic.com/engineering/multi-agent-research-system) | Detalles del estudio sobre la mejora del rendimiento del 90,2% |
### Recursos comunitarios
| Recursos | Enlaces | Instrucciones |
| --------------------------- | ----------------------------------------------------------------- | ----------------------------------------------- |
| wshobson/agentes | [GitHub](https://github.com/wshobson/agents) | 99 agentes + 15 plantillas de orquestador |
| Ingeniería compuesta | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Complementos para 17 agentes especializados |
| código-claude-impresionante | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Resumen de las mejores prácticas de Claude Code |
### Lectura recomendada
| Artículo | Fuente | Tema |
| ------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------ |
| Construyendo agentes eficaces | [Antrópico](https://www.anthropic.com/research/building-effective-agents) | Principios de diseño de agentes |
| Cómo construimos nuestro sistema de investigación multiagente | [Antrópico](https://www.anthropic.com/engineering/multi-agent-research-system) | Práctica de arquitectura multiagente |
| Habilidades, comandos, subagentes y complementos de Claude | [Tecnología de jóvenes líderes](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Análisis de comparación de funciones |
## Resumen
Claude Code Subagent es una poderosa herramienta para mejorar la eficiencia de la programación de IA. Hace que las tareas complejas sean manejables a través de contextos independientes y configuraciones especializadas.
Inicio rápido:
1. Ejecute `/agents` para abrir la interfaz de administración.
2. Cree un subagente simple (como un revisor de código)
3. Pruebe la delegación automática y las llamadas explícitas.
4. Ajuste la configuración según sea necesario
A medida que profundices en su uso, podrás gradualmente:
* Crea un subagente exclusivo para tu equipo.
* Configurar enlaces de subagente para manejar flujos de trabajo complejos
* Manejar tareas a largo plazo con ejecución reanudable.
Si desea empaquetar y distribuir Subagent con otras configuraciones, lea la "[Guía práctica del complemento Claude Code](/es/docs/notes/claude-plugin/practice)".
# Consejos avanzados
## Notificaciones de terminal: alertas al completar tareas
¿Quieres recibir una notificación cuando Claude termine una tarea?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
Combínalo con las notificaciones de iTerm2, o usa `terminal-notifier` para notificaciones personalizadas (consulta la configuración de Hooks en [Mejores prácticas](/es/blog/claude-code-best-practices)).
## Uso avanzado de Hooks
Hooks no solo ejecuta comandos shell. En realidad hay cuatro tipos:
1. **command**:Shell 命令(最常见)
2. **http**:POST JSON 到 URL(支持自定义 headers 和环境变量展开)
3. **prompt**:发给 Claude 评估(比如「所有任务都完成了吗?」)
4. **agent**:启动一个有工具访问权限的子代理来验证
Algunos eventos avanzados de Hook que vale la pena conocer:
* `PostCompact`:压缩完成后触发,适合注入提醒让 Claude 重新读取关键文件
* `SessionStart`:写入 `$CLAUDE_ENV_FILE` 可以给整个会话持久化环境变量
* `PreToolUse`:可以修改工具输入(`updatedInput`),甚至自动批准或拒绝操作
## Ecosistema de plugins
Usa `/plugin` para explorar e instalar plugins de la comunidad. Algunos destacados:
* **dx**(by ykdojo):提供 `/handoff`(自动写交接文档)、`/clone`(克隆对话)、`/half-clone`(只克隆最近的对话减少上下文)
* **mine**(by anipotts):把所有 Claude Code 会话数据导入 SQLite,支持成本追踪、缓存分析、错误记忆等查询
## Agent Teams: colaboración multi-agente
Configura una variable de entorno para habilitar la función experimental de Agent Teams:
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
Una vez habilitada, una sesión puede actuar como Team Lead, coordinando múltiples agentes que trabajan simultáneamente a través de git worktree. Cada agente se ejecuta de forma independiente en su propia ventana de contexto, ideal para el desarrollo en paralelo de proyectos grandes.
Ten en cuenta que el consumo de tokens aumenta entre 4 y 15 veces, así que úsalo con moderación.
## Estrategias de prompts
Los siguientes consejos provienen del hilo de Twitter de Boris Cherny sobre prácticas de equipo, básicamente las mejores prácticas de "ingeniería de prompts" aplicadas a Claude Code.
### Usa Claude como tu revisor de código
No solo le pidas a Claude que escriba código, también haz que revise el tuyo:
```
Grill me on these changes and don't make a PR until I pass your test.
```
O pídele que demuestre que el código funciona:
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### No reformules cuando no estés satisfecho
El tip #6 de Boris: si Claude da una respuesta mediocre, no la reformules y vuelvas a preguntar. Mejor di "Esta solución no es suficientemente buena, dime específicamente qué se puede mejorar". Iterar sobre la respuesta existente funciona mejor que empezar de cero.
### Deja que Claude actualice su propio CLAUDE.md
Después de corregir un error, agrega:
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris dice que Claude es sorprendentemente bueno escribiendo reglas para sí mismo. Con el tiempo, CLAUDE.md se vuelve cada vez más preciso y la calidad de las conversaciones mejora continuamente.
### Solo di "fix"
Con Slack MCP habilitado, pega un reporte de bug de Slack y di una sola palabra: **fix**. Cero cambio de contexto.
O cuando CI falla, simplemente di:
```
Go fix the failing CI tests.
```
No necesitas analizar logs manualmente ni explicar cuál es el problema — deja que Claude revise los logs, diagnostique el problema y lo solucione.
## Para finalizar
Claude Code evoluciona muy rápido, y estos consejos se están refinando constantemente. Te recomendamos seguir el Changelog oficial para mantenerte al día.
Si aún no has leído mis artículos anteriores, te sugiero comenzar con los flujos de trabajo básicos:
### Lectura adicional
* [Mis mejores prácticas con Claude Code](/es/blog/claude-code-best-practices) — Consejos esenciales de flujo de trabajo y guía de comandos slash
* [Control de calidad en programación con IA: 5 líneas de defensa](/es/blog/claude-code-quality-control) — Sistema de aseguramiento de calidad para programación con Claude Code
* [Arquitectura del sistema Claude explicada](/es/docs/notes/claude-architecture) — Comprendiendo MCP, Skills, Subagents, Hooks y más
# Comandos Prácticos y Automatización
## `/diff`: Visor Interactivo de Diff
Escribe `/diff` para abrir una vista interactiva de diff:
* **Flechas izquierda/derecha**: Alterna entre el git diff completo (todos los cambios) y los cambios por turno de Claude
* **Flechas arriba/abajo**: Navega entre diferentes archivos
Mucho mejor que ejecutar `git diff` en la terminal, especialmente cuando los cambios abarcan múltiples archivos.
## `/simplify`: Revisión de Código Multi-Agente
Ejecutar `/simplify` lanza 3 agentes de revisión en paralelo:
* Agente de **reutilización de código**: Busca patrones duplicados
* Agente de **calidad de código**: Verifica legibilidad y estructura
* Agente de **eficiencia**: Analiza sobrecarga de rendimiento innecesaria
Los tres agentes trabajan de forma independiente y luego consolidan los resultados, corrigiendo automáticamente los problemas válidos y omitiendo los falsos positivos.
## `/security-review`: Escaneo de Seguridad
Realiza una auditoría de seguridad sobre los cambios de la rama actual, verificando inyección SQL, XSS, fallas de autenticación, problemas de manejo de datos y vulnerabilidades en dependencias. Cada hallazgo pasa por una validación adversarial para reducir falsos positivos.
## Funciones Ocultas de `/copy`
`/copy` no solo copia la última respuesta. Cuando la respuesta contiene bloques de código, muestra un selector interactivo que te permite elegir un bloque de código específico en lugar de copiar toda la respuesta. También puedes pasar un número para copiar respuestas anteriores: `/copy 2` copia la penúltima, `/copy 3` copia la antepenúltima, sin necesidad de desplazarte para seleccionar manualmente.
## `/batch`: Refactorización Paralela a Gran Escala
```
/batch 把 src/ 下所有组件从 Class 组件迁移到函数组件
```
Esta es una funcionalidad de gran alcance. `/batch` analiza el código base, divide la tarea en 5–30 unidades independientes, lanza un agente independiente para cada unidad que trabaja en un git worktree aislado, y finalmente cada agente hace commit y abre un PR.
Ideal para migraciones a gran escala, anotaciones de tipos en lote, renombramientos globales y escenarios similares.
## `/loop`: Tareas Programadas
```
/loop 5m 检查部署是否完成
/loop 1h /review-pr 1234
```
Crea una tarea programada dentro de la sesión que se repite en el intervalo especificado. Útil para consultar el estado del despliegue, revisar PRs periódicamente, etc. Es a nivel de sesión (desaparece al salir), con un máximo de 50 tareas y expiración automática a los 3 días.
## Entrada por Pipe: Alimenta a Claude con Cualquier Cosa
```bash
# 让 Claude 分析错误日志
cat error.log | claude -p "分析这个错误日志,找出根本原因"
# 让 Claude 总结最近的改动
git diff HEAD~3 | claude -p "总结这三次提交的改动"
# 让 Claude 解读命令输出
kubectl get pods | claude -p "哪些 pod 状态异常?"
```
`-p` es el modo headless (no interactivo), ideal para usar en scripts y pipelines de CI/CD.
## Parámetros Ocultos del Modo Headless
El modo `-p` tiene algunos parámetros extremadamente poderosos pero poco conocidos:
```bash
# 设置花费上限(超过就停)
claude -p --max-budget-usd 5.00 "重构认证模块"
# 限制对话轮数
claude -p --max-turns 3 "修复这个测试"
# 输出 JSON 格式(方便程序解析)
claude -p --output-format json "分析这个项目"
# 要求输出符合特定 JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "总结项目"
# 多轮 headless 对话(用 session-id 保持上下文)
claude -p --session-id my-task "第一步:分析代码"
claude -p --session-id my-task "第二步:生成测试"
# 指定备用模型(主模型过载时自动切换)
claude -p --fallback-model sonnet "复杂分析"
# 限制可用工具
claude -p --tools "Read,Grep,Glob" "只读分析,不要改代码"
# 完全替换系统提示词
claude -p --system-prompt "你是一个 Python 专家" "优化这段代码"
```
# Configuración y diagnóstico
## `/statusline`: Barra de estado personalizada
Usa `/statusline` para personalizar la información que se muestra en la barra de estado inferior mediante descripciones en lenguaje natural. También puedes crear manualmente un script `~/.claude/statusline.sh`.
La información que puedes mostrar incluye: modelo actual, rama de git, número de archivos sin confirmar, progreso del uso de contexto, costo de la sesión, entre otros. Cuando tienes varias ventanas de Claude abiertas para diferentes tareas, la barra de estado te ayuda a identificar rápidamente qué hace cada ventana.
## Autocompletado de settings.json
Agrega `$schema` al inicio de tu settings.json y VS Code / Cursor proporcionará autocompletado y validación de las opciones de configuración:
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## Algunas configuraciones ocultas útiles
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration`: Muestra la duración de cada turno de conversación
* `DISABLE_AUTOUPDATER`: Desactiva la verificación automática de actualizaciones, reduciendo la sobrecarga de contexto
## `/stats` y `/insights`: Análisis de uso
* `/stats`: Visualiza el uso diario, historial de sesiones, rachas de uso y preferencias de modelo, con filtrado por rango de fechas
* `/insights`: Analiza todo tu historial de Claude Code, te dice qué flujos de trabajo son efectivos, dónde están los cuellos de botella y genera sugerencias de optimización
## history.jsonl: Historial de prompts
Claude guarda cada prompt que envías en `~/.claude/history.jsonl`. Puedes pedirle a Claude que analice este archivo para encontrar patrones en los prompts y oportunidades de optimización.
## `/doctor`: Chequeo de salud
Cuando encuentres problemas extraños, ejecuta `/doctor` (o `claude doctor` en la terminal). Verifica el estado de instalación, versión, estado de autenticación y dependencias del sistema para ayudarte a identificar problemas rápidamente.
## Herramientas de la comunidad
La herramienta comunitaria `ccusage` puede rastrear el uso de tokens:
```bash
npx ccusage daily
```
Si has usado `--dangerously-skip-permissions` o has aprobado muchos comandos, puedes usar `cc-safe` para escanear posibles riesgos:
```bash
npx cc-safe .
```
Verifica `.claude/settings.json` en busca de comandos de alto riesgo como `sudo`, `rm -rf`, `chmod 777`, `git reset --hard`, entre otros.
## Más comandos slash que vale la pena conocer
| Comando | Función |
| --------------------- | -------------------------------------------------------------------------------------------------- |
| `/export [filename]` | Exportar conversación como texto plano |
| `/pr-comments [PR]` | Obtener comentarios de PR (detecta automáticamente la rama actual) |
| `/release-notes` | Ver el registro de cambios de la versión actual |
| `/plugin` | Explorar e instalar plugins de la comunidad |
| `/fast` | Alternar modo rápido |
| `claude --debug` | Activar logs de depuración al iniciar (soporta filtrado por categoría, ej., `--debug "api,hooks"`) |
| `/install-github-app` | Instalar GitHub App para revisiones automáticas de PR |
# Gestión del contexto
## `/compact` acepta argumentos
Muchos saben que `/compact` puede comprimir el contexto, pero pocos saben que acepta argumentos para especificar qué conservar:
```
/compact 保留所有关于数据库 schema 的讨论,以及当前的重构方案
```
De esta forma, la compresión priorizará el contenido que especificaste, evitando la pérdida de contexto crítico.
## Escribe instrucciones de supervivencia de compactación en CLAUDE.md
Añade una sección `## Compact Instructions` en tu CLAUDE.md para indicarle a Claude qué debe preservar durante la compactación:
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
Así, incluso la compactación automática no perderá información crítica.
## Evita que Claude abandone prematuramente por presupuesto de tokens
Añade esto en tu CLAUDE.md:
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
A veces Claude se detiene proactivamente cuando el contexto está casi lleno, diciendo "el contexto está casi lleno". Añadir esto evita que abandone prematuramente.
## Protocolo Handoff: traspaso de sesión
Cuando el contexto está casi lleno pero la tarea no ha terminado, pídele a Claude que escriba un documento de traspaso:
```
把剩余的计划写到 HANDOFF.md 里,说明你尝试了什么、什么有效、什么没效。
```
Luego abre una nueva sesión y simplemente usa `@HANDOFF.md` para restaurar el contexto completo. Esto comprime más de 10K tokens de contexto a menos de 2K, mucho más preciso que `/compact`.
## Compacta proactivamente al 70-80%
Un punto fácil de pasar por alto: cuando el contexto se acerca a su límite, Claude activa automáticamente la compactación. Pero cuando la compactación automática ocurre a mitad de una tarea, puede perder información crítica y degradar la calidad de las respuestas posteriores.
Un mejor enfoque es la **gestión proactiva**: ejecuta `/compact` manualmente cuando el contexto alcance el 70-80%; funciona mucho mejor que esperar la compactación automática. Ejecuta `/clear` inmediatamente después de completar una tarea; no dejes que el contexto crezca indefinidamente.
También puedes activar la compactación automática antes mediante una variable de entorno:
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context`: diagnóstico del contexto
¿No estás seguro de cuánto espacio queda en la ventana de contexto? `/context` te lo dirá:
* Qué herramientas o servicios MCP consumen más contexto
* Porcentaje de uso de la capacidad actual
* Sugerencias de optimización específicas
He descubierto que a veces solo tener ciertos servicios MCP registrados (sin siquiera usarlos) puede consumir más del 30% de la ventana de contexto. Usa `/context` para verificar; limpiar los MCP que no uses puede liberar bastante espacio.
## Carga diferida automática de herramientas MCP
Cuando las definiciones de herramientas MCP superan el 10% del contexto, Claude Code activa automáticamente Tool Search, cargando un índice de búsqueda ligero en lugar de las definiciones completas de herramientas. Esto reduce el consumo de contexto MCP en más del 85% (por ejemplo, de 77K tokens a 8.7K). Esta función está **activada por defecto** y no requiere configuración manual.
Ten en cuenta: Tool Search solo es compatible con los modelos Sonnet 4+ y Opus 4+, no con Haiku. Si tu `ANTHROPIC_BASE_URL` apunta a un proxy no oficial, Tool Search se desactivará automáticamente (porque la mayoría de los proxies no reenvían los bloques `tool_reference`).
Si deseas personalizar el comportamiento, configúralo en settings.json:
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
Valores de configuración soportados:
* **Sin configurar**: Activado por defecto
* **`true`**: Forzar activación (incluyendo escenarios de proxy no oficial)
* **`auto`**: Se activa cuando el contexto supera el 10% (equivalente al comportamiento predeterminado)
* **`auto:`**: Umbral personalizado, por ejemplo, `auto:5` significa activar al superar el 5%
* **`false`**: Desactivado, todas las herramientas MCP se precargan
# Trucos ocultos de Claude Code
Tips prácticos de Claude Code recopilados de los tweets del fundador Boris Cherny, la comunidad y el changelog. Muchos son verdaderas joyas escondidas -- atajos, funciones ocultas, trucos de la línea de comandos y más -- que una vez que empiezas a usarlos, no hay vuelta atrás.
## Contenido
* [Atajos de teclado](./shortcuts) -- Cambio de modo con Shift+Tab, historial con Esc+Esc, Ctrl+S para guardar temporalmente, y más
* [Entrada e interacción](./input-interaction) -- Comandos de terminal con `!`, inyección de archivos con `@`, pegado de URLs, /btw para interrupciones, modo Vim
* [Pensamiento y control del modelo](./thinking-model) -- Palabras clave think/ultrathink, /effort, subagents, opusplan
* [Gestión de sesiones](./session-management) -- /rename, /branch, /color, control remoto
* [Gestión de contexto](./context-management) -- Parámetros de /compact, directivas de compresión, protocolo Handoff, carga diferida de MCP
* [Comandos y automatización](./commands-automation) -- /diff, /simplify, /batch, /loop, modo Headless
* [Configuración y diagnósticos](./config-diagnostics) -- statusline, settings.json, /stats, /doctor
* [Avanzado](./advanced) -- Hooks a fondo, ecosistema de plugins, Agent Teams, filosofía de prompts
### Lecturas adicionales
* [Arquitectura del sistema Claude al detalle](/es/docs/notes/claude-architecture) -- Entendiendo MCP, Skills, Subagents, Hooks y otros componentes
* [Guía completa de Claude Subagents](/es/docs/notes/claude-subagent) -- Conceptos y práctica de subagentes
# Entrada e interacción
## `!`:直接运行终端命令
在输入框以 `!` 开头,可以直接在 Claude Code 内执行终端命令,不需要切换到另一个终端窗口:
```
! git status
! npm run build
! docker ps
```
输入 `!` 加命令前缀后按 Tab 还能自动补全历史命令。
## `@` + 文件路径:注入文件上下文
在输入时用 `@` 加文件路径,可以把文件内容直接注入到上下文中:
```
帮我看看 @src/auth/login.ts 和 @src/auth/middleware.ts 之间的逻辑有没有问题
```
支持 Tab 键自动补全路径,不需要手动输入完整路径。比让 Claude 自己去读文件更快,因为省去了工具调用的开销。
## 直接粘贴 URL
直接把 URL 粘贴到输入中,Claude 会自动抓取网页内容作为上下文:
```
参考这个 API 文档 https://docs.example.com/api/v2 来写客户端代码
```
## 喂 `/llms-full.txt` 让 Claude 自己查文档
很多开源项目的文档站点会提供 `/llms-full.txt` 文件(LLM 友好的完整文档)。遇到某个库的问题时,把这个文件的 URL 粘贴给 Claude,它能自己查文档解决绝大部分问题:
```
参考 https://docs.astro.build/llms-full.txt 帮我解决这个路由问题
```
## `/btw`:在 Claude 工作时插嘴
这是 2026 年 3 月刚加的新功能。当 Claude 正在执行任务时,你可以用 `/btw` 发起一个旁路对话——问问它在想什么、给它补充信息,而不需要打断当前任务。
正如 Anthropic 工程师 @trq212 在推特上说的:「没人会用 Ctrl+C 打断同事,你只需要说一句 'btw',他们就会抬头看你。」
## Vim 模式
输入 `/vim` 开启 Vim 模式,支持:
* 模式切换(Normal/Insert)
* 导航(h/j/k/l, w/b/e, 0/$)
* 编辑操作(d, c, y, p)
* 文本对象(iw, aw, i", a())
如果你是 Vim 用户,这比默认的输入体验好太多。用 `/config` 可以设置为永久开启。
## 语音模式
输入 `/voice` 激活语音模式,长按空格键说话,松开发送。适合不想打字但又需要给 Claude 交代任务的时候。按键可以在 `keybindings.json` 中自定义。
# Gestión de sesiones
## `/rename`:Nombrar la sesión
```
/rename my-auth-refactor
```
Dale un nombre a tu sesión actual. La ventaja es que, en el selector interactivo de sesiones (`claude --resume`), las sesiones con nombre pueden seleccionarse y restaurarse directamente sin necesidad de pulsar Enter para confirmar; también puedes iniciar directamente desde el terminal con `claude --resume my-auth-refactor`.
En el selector, basta con escribir para buscar y filtrar. También admite estos atajos de teclado: `Ctrl+V` para previsualizar la sesión, `Ctrl+R` para renombrar, `Ctrl+A` para alternar la visualización de todos los proyectos y `Ctrl+B` para filtrar por rama.
## `/branch`:Bifurcar la conversación
Al igual que las ramas de git, `/branch` crea una bifurcación en el punto actual de la conversación. Puedes probar diferentes enfoques en la bifurcación sin afectar la conversación original. Si no te convence, simplemente vuelve a la rama original y continúa.
## `/color`:Colorear la ventana
Establece un color para la barra de prompt de la sesión actual. Admite red, blue, green, yellow, purple, orange, pink y cyan.
## Kit completo de gestión de sesiones por línea de comandos
```bash
# 恢复当前目录最近的会话
claude --continue
# 打开会话选择器,或按名称恢复
claude --resume
claude --resume my-auth-refactor
# 启动时直接命名会话
claude -n "auth-refactor"
# Fork 上一次会话(保留上下文,创建新分支)
claude -c --fork-session
# 恢复与特定 PR 关联的会话
claude --from-pr 123
# 在隔离的 git worktree 中启动
claude -w
```
Las sesiones de Claude Code se guardan automáticamente (no necesitas Ctrl+S), así que puedes reanudar tu último trabajo con `--continue` cada vez que abras una terminal.
## `claude --remote`:Continuar desde otro dispositivo
```bash
claude --remote "your task description"
```
Inicia una sesión web que puedes continuar en claude.ai o en la aplicación móvil.
## `/remote-control`:Controlar Claude local desde el móvil
Escribe `/remote-control` en Claude Code en tu ordenador y se generará un código de conexión. Luego, introduce ese código en la aplicación Claude de tu móvil para controlar remotamente la sesión local de Claude Code: da instrucciones a Claude en tu ordenador desde tu teléfono.
# Atajos de Teclado
El sistema de atajos de Claude Code es mucho mas completo de lo que la mayoria imagina -- presiona `?` para ver todos los atajos disponibles en tu contexto actual.
## Shift+Tab: Cambio ciclico de modos
Este es probablemente el atajo mas importante. Presionar `Shift+Tab` cicla entre tres modos:
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
No necesitas escribir `/plan` ni `/auto-accept` manualmente -- una sola tecla lo resuelve todo. Mi flujo de trabajo: cuando recibo una tarea nueva, presiono dos veces para saltar a Plan Mode, confirmo el enfoque, y luego presiono una vez mas para cambiar a Auto-Accept y dejar que Claude ejecute por su cuenta.
## Esc + Esc: La maquina del tiempo
Presiona `Esc` dos veces seguidas y aparece el menu de retroceso (Rewind):
* **Restaurar codigo y conversacion**: Vuelves a un punto de control anterior, tanto los archivos como el historial de chat se revierten
* **Solo conversacion**: Reviertes los mensajes pero conservas los cambios de codigo actuales
* **Solo codigo**: Deshaces las modificaciones de archivos pero conservas el historial de conversacion
Claude rastrea automaticamente cada edicion de archivo como un punto de control. Esto es mucho mas granular que `git checkout .` porque puedes regresar a cualquier edicion individual, no solo al ultimo commit.
Un detalle importante: solo se rastrean los archivos que Claude edito directamente mediante herramientas. Los archivos que modificaste manualmente, `git push` u otras operaciones externas no se pueden revertir.
## Ctrl+S: Guardado temporal de prompts (Prompt Stash)
¿Estas a mitad de escribir un prompt y necesitas atender otra cosa primero? Presiona `Ctrl+S` y tu entrada actual se guarda temporalmente:
Luego puedes escribir otro comando o pregunta. Cuando envies ese mensaje, el contenido guardado se **restaura automaticamente** en el campo de entrada para que continues donde lo dejaste.
Piensa en esto como `git stash` pero para prompts. Ejemplo: estas escribiendo una descripcion larga de refactorizacion y te das cuenta de que quieres que Claude revise un archivo primero -- presiona `Ctrl+S` para guardar la descripcion, haz tu pregunta sobre el archivo, y cuando te responda, tu descripcion vuelve automaticamente.
## Ctrl+B: Enviar tareas al segundo plano
¿Claude esta procesando una tarea que toma mucho tiempo (como una refactorizacion grande) y quieres trabajar en otra cosa? Presiona `Ctrl+B` para enviar la tarea actual al segundo plano -- tu terminal queda libre de inmediato para nuevas instrucciones.
Usa `Ctrl+T` para ver la lista de tareas en segundo plano, y presiona `Ctrl+F` dos veces para terminar todos los agentes en segundo plano.
> Usuarios de tmux: la tecla de prefijo por defecto de tmux tambien es `Ctrl+B`, asi que necesitas presionarla dos veces para activar la funcion de segundo plano de Claude.
## Ctrl+G: Escribe prompts largos en tu editor
A veces necesitas darle a Claude un conjunto extenso de instrucciones y escribir en la terminal es incomodo. Presiona `Ctrl+G` para abrir tu `$EDITOR` predeterminado (VS Code, Vim, etc.), escribe tu prompt ahi, y se envia a Claude automaticamente cuando guardas y cierras.
Para cambiar el editor predeterminado, configuralo en tu archivo de shell (`~/.zshrc` o `~/.bashrc`):
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
El parametro `--wait` es importante -- le indica al editor que espere hasta que cierres el archivo; de lo contrario, Claude recibe contenido vacio de inmediato. Los editores de terminal como Vim bloquean naturalmente, asi que no lo necesitan.
Muy util para descripciones de requerimientos de varios parrafos o para pegar material de referencia extenso. En Plan Mode, incluso puedes usar `Ctrl+G` para editar directamente en tu editor el plan generado por Claude.
## Cmd+T: Alternar pensamiento extendido
El atajo predeterminado es `Cmd+T` (o `Meta+T` en Windows/Linux) y activa o desactiva el modo de pensamiento extendido (Extended Thinking). Cuando esta activado, Claude razona con mayor profundidad antes de responder -- ideal para decisiones complejas de arquitectura o investigacion de bugs dificiles.
Ojo: la mayoria de las terminales (iTerm2, Terminal.app, Warp, etc.) interceptan `Cmd+T` como "nueva pestana", asi que este atajo a menudo no funciona en la practica. Dos alternativas: usa `/keybindings` para reasignarlo a una tecla que no genere conflicto, o simplemente usa el comando `/effort` para cambiar la profundidad de razonamiento (mismo efecto, y ademas permite control preciso del nivel).
## Atajos de Readline
El campo de entrada de Claude Code soporta los atajos estandar de Readline -- los veteranos de la terminal se sentiran como en casa:
| Atajo | Funcion |
| ----------------- | ---------------------------------------- |
| Ctrl+A | Ir al inicio de la linea |
| Ctrl+E | Ir al final de la linea |
| Ctrl+W | Eliminar la palabra anterior |
| Ctrl+U | Eliminar hasta el inicio de la linea |
| Ctrl+K | Eliminar hasta el final de la linea |
| Ctrl+Y | Pegar el ultimo texto eliminado |
| Alt+Y | Ciclar por el historial de eliminaciones |
| Option+Left/Right | Saltar por palabras (Mac) |
## Atajos de aprobacion: `y/n/d/e`
Cuando Claude propone una modificacion de archivo y espera tu confirmacion, cuatro atajos de una sola tecla controlan el flujo:
* `y`: Aceptar
* `n`: Rechazar
* `d`: Ver el diff completo
* **`e`: Editar antes de aceptar**
`e` es el mas ignorado pero el mas util -- te permite hacer ajustes finos sobre los cambios de Claude antes de aplicarlos. ¿No te convencen algunas lineas? No necesitas rechazar y empezar de nuevo, solo presiona `e` y corrigelo.
## Referencia rapida
| Atajo | Funcion |
| ----------- | ------------------------------------------------------------------------------------------------------------- |
| Shift+Tab | Ciclar modos: Normal → Auto-Accept → Plan |
| Esc+Esc | Abrir menu de retroceso |
| Ctrl+S | Guardar entrada actual, se restaura automaticamente tras el siguiente envio |
| Ctrl+B | Enviar tarea actual al segundo plano |
| Ctrl+T | Ver lista de tareas en segundo plano |
| Ctrl+F (x2) | Terminar todos los agentes en segundo plano |
| Ctrl+G | Escribir prompt en editor externo |
| Ctrl+O | Alternar vista detallada de herramientas |
| Cmd+T | Alternar pensamiento extendido (puede ser interceptado por la terminal; considera reasignar o usar `/effort`) |
| `\` + Enter | Entrada multilinea (sin configuracion adicional) |
| Shift+Enter | Entrada multilinea (requiere ejecutar `/terminal-setup` primero) |
| Up / Down | Navegar historial de entradas |
| Ctrl+R | Buscar en historial de comandos |
| Ctrl+L | Limpiar pantalla (historial conservado) |
| Ctrl+C | Cancelar generacion actual |
| Ctrl+D | Salir de Claude Code |
| `?` | Mostrar todos los atajos disponibles |
## Atajos personalizados
Si los atajos por defecto no se ajustan a tus preferencias, usa `/keybindings` para abrir `~/.claude/keybindings.json` y personalizarlos. Los cambios toman efecto de inmediato -- no necesitas reiniciar.
Se soporta la sintaxis de teclas combinadas (por ejemplo, `ctrl+shift+c`) y el modo Chord (por ejemplo, `ctrl+k ctrl+s` -- presiona Ctrl+K, suelta, y luego presiona Ctrl+S). Hay 16 contextos de enlace diferentes (Chat, Autocomplete, Confirmation, DiffDialog, etc.), cada uno con su propio conjunto de acciones asignables.
# Pensamiento y control de modelos
## Controlar la profundidad de pensamiento con palabras clave
Agregar palabras clave específicas en tus prompts activa diferentes niveles de presupuesto de pensamiento. Esta es una función exclusiva de Claude Code (no está disponible en claude.ai web):
| Palabra clave | Presupuesto de pensamiento | Caso de uso |
| ----------------------------- | -------------------------- | -------------------------------------------- |
| `think` | \~4,000 tokens | Problemas de programación cotidianos |
| `think hard` / `megathink` | \~10,000 tokens | Lógica compleja, dependencias entre archivos |
| `think harder` / `ultrathink` | \~31,999 tokens | Diseño de arquitectura, bugs difíciles |
En la práctica, generalmente agrego `think hard` cuando Claude da una respuesta superficial y vuelvo a preguntar. Para problemas particularmente complejos (como depurar errores a través de múltiples servicios), voy directo a `ultrathink`.
## `/effort`: Controlar la profundidad de pensamiento
Además de las palabras clave (think / ultrathink), puedes usar `/effort` para configurar directamente la profundidad de pensamiento:
```
/effort low # 简单任务,跳过深度思考,更快更省
/effort high # 复杂任务,深度推理
/effort max # 最大思考预算(仅 Opus)
/effort auto # 让 Claude 自己判断
```
La configuración persiste durante toda la sesión. Usa `low` para ediciones simples de archivos y `max` para diseño de arquitectura compleja — así ahorras dinero sin sacrificar calidad.
## La palabra clave `use subagents`
Agrega `use subagents` al final de cualquier solicitud, y Claude dividirá la tarea en múltiples sub-agentes que se ejecutan en paralelo. Esto no solo es más rápido, sino que también mantiene limpia la ventana de contexto del agente principal.
Boris mencionó esto específicamente en Twitter: delegar tareas individuales a sub-agentes para mantener el contexto del agente principal enfocado.
## `opusplan`: La mejor estrategia de modelos en relación costo-beneficio
Resumen en una línea: **Opus piensa, Sonnet ejecuta**.
## `/model`: Cambiar de modelo
Usa `/model` para cambiar de modelo en cualquier momento durante una sesión. Por ejemplo, usa Sonnet en el día a día, cambia temporalmente a Opus para problemas complejos y vuelve cuando termines.
## Control de estilo de salida
En `/config`, selecciona "Output style" — hay dos modos poco comunes pero muy útiles:
* **Modo Explanatory**: Claude inserta "puntos de conocimiento" entre las tareas, explicando frameworks y patrones de código relevantes — ideal para aprender un proyecto nuevo
* **Modo Learning**: Un modo de aprendizaje colaborativo donde Claude agrega marcadores `TODO(human)` en el código para que los implementes tú mismo, en lugar de darte la respuesta directamente
También puedes crear archivos personalizados de estilo de salida (en formato Markdown) en `~/.claude/output-styles/` para modificar directamente el prompt del sistema. Nota: los estilos de salida personalizados **reemplazarán completamente** el prompt de sistema de programación predeterminado, a menos que configures `keep-coding-instructions: true`.
# Introducción conceptual
## Introducción
En octubre de 2025, Anthropic lanzó discretamente una nueva funcionalidad llamada Claude Skills. Esta actualización, aparentemente modesta, fue calificada por el reconocido bloguero tecnológico Simon Willison como "posiblemente más importante que MCP", prediciendo que desencadenaría una "explosión cámbrica" en el ámbito de las herramientas de IA.
Esta valoración no carece de fundamento. Si utiliza frecuentemente asistentes de IA, seguramente ha enfrentado esta frustración: cada vez que inicia una nueva conversación, debe repetir las mismas instrucciones de flujo de trabajo; cuando finalmente logra que la IA funcione a su satisfacción, al cambiar de ventana de conversación debe empezar desde cero. Los Skills fueron creados precisamente para resolver este punto de dolor.
## Comprender los Claude Skills
Imagine que usted es el director de una empresa. Cuando un nuevo empleado se incorpora, le entrega un manual de trabajo que detalla los procesos laborales de la empresa, las normas de marca y cómo manejar problemas comunes. Claude Skills es ese "manual de trabajo" para el asistente de IA — permite que complete tareas específicas de manera repetible y estandarizada.
Desde una perspectiva técnica, los Skills son carpetas que contienen instrucciones, scripts y recursos, que Claude puede cargar dinámicamente cuando los necesita. Cada Skill le enseña a Claude cómo completar cierto tipo de tarea de manera consistente, y este conocimiento se preserva de forma persistente entre conversaciones. Esto significa que solo necesita "capacitarlo" una vez, y a partir de entonces, sin importar cuándo lo use, Claude recordará cómo hacerlo.
### Tres componentes principales
Un Skill completo se compone de las siguientes tres partes:
| Componente | Función | ¿Requerido? |
| ---------------------------- | ------------------------------------------------------------------------------------- | ----------- |
| **SKILL.md** | Documento de instrucciones principal, contiene metadatos e instrucciones detalladas | Requerido |
| **Materiales de referencia** | Guías de marca, documentos de políticas, plantillas y otra información complementaria | Opcional |
| **Scripts** | Código Python/JavaScript para cálculos complejos u operaciones con archivos | Opcional |
SKILL.md es el "alma" de todo el Skill. Su estructura básica es la siguiente:
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
El frontmatter YAML al inicio del archivo contiene dos campos clave: `name` es el nombre identificador del skill, con un máximo de 64 caracteres; `description` le indica a Claude para qué sirve este skill y cuándo debe usarlo, con un máximo de 200 caracteres. Claude se basa precisamente en esta descripción para determinar cuándo debe invocar un Skill, por lo que cuanto más clara y precisa sea, mayor será la probabilidad de que el Skill se active correctamente.
### Escenarios de aplicación
Los escenarios de aplicación de los Skills son muy amplios y cubren diversas tareas repetitivas del trabajo diario:
**Procesamiento de documentos**: Creación masiva de hojas de cálculo Excel, presentaciones PowerPoint, documentos Word e informes PDF. Anthropic ofrece oficialmente un conjunto de skills de documentos, listos para usar.
**Cumplimiento de marca**: Empaquete los colores de marca de su empresa, reglas de uso de logotipos, especificaciones de espaciado y estilo de tono en un Skill, asegurando que todo el contenido generado por la IA cumpla con los estándares de marca.
**Actas de reuniones**: Resuma automáticamente los registros de reuniones, extraiga elementos de acción, asigne responsables y genere correos de seguimiento.
**Análisis de datos**: Ejecute flujos de análisis estandarizados, como escaneo de inteligencia competitiva (extracción estructurada de actualizaciones de productos, cambios de precios, comentarios de analistas), análisis financiero (análisis de informes financieros, construcción de modelos financieros).
**Gestión de proyectos**: Construya planes de proyecto a partir de objetivos, sugiera hitos y genere informes semanales o resúmenes para inversionistas.
## Arquitectura de divulgación progresiva
El diseño más ingenioso de los Skills radica en su forma de cargar la información. Las descripciones tradicionales de herramientas MCP pueden consumir miles o incluso decenas de miles de tokens, mientras que los metadatos de los Skills solo ocupan unas pocas decenas de tokens. Esto significa que puede habilitar una gran cantidad de Skills simultáneamente sin preocuparse de que las descripciones de herramientas llenen la ventana de contexto.
Esta eficiencia proviene de un diseño arquitectónico llamado **divulgación progresiva** (Progressive Disclosure). Los Skills emplean una estructura de información de tres capas, cargando bajo demanda capa por capa — como un manual con índice:
```
📚 Manual de trabajo de Skills
│
├─ 📋 Índice ─────────────────────────── 【Capa de metadatos】Precargada al inicio
│ │
│ │ name: "weekly-report"
│ │ description: "Generar informes semanales estandarizados basados en el contenido de trabajo"
│ │
│ │ ✓ Solo ocupa 30-50 tokens
│ │ ✓ Los índices de todos los Skills son visibles simultáneamente
│ │
│
├─ 📖 Capítulos principales ───────────── 【Capa de documento central】Se carga cuando es relevante
│ │
│ │ # Weekly Report Generator
│ │
│ │ ## Instructions
│ │ Generar informe semanal con la siguiente estructura...
│ │
│ │ ## Examples
│ │ Entrada: Esta semana completé la función de inicio de sesión...
│ │ Salida: ### Completado esta semana ...
│ │
│ │ ⚡ Se expande solo cuando Claude lo considera necesario
│ │ 📊 Consume cientos a miles de tokens
│ │
│
└─ 📎 Anexos ─────────────────────────── 【Capa de recursos de referencia】Se carga cuando se necesita
│
│ references/
│ ├── brand-guide.md Guía de marca
│ ├── template.xlsx Plantilla de informe
│ └── examples/ Informes semanales históricos
│
│ 🔍 Se carga solo cuando es explícitamente necesario
│ 📦 Puede contener gran cantidad de material de referencia
```
Primero mira el índice para saber qué capítulos hay (capa de metadatos), encuentra el capítulo que necesita y lo abre para leerlo (capa de documento central), y finalmente, si necesita más detalles, consulta los anexos (capa de recursos de referencia).
| Capa | Contenido | Momento de carga | Consumo de tokens |
| ---------------------------------- | ---------------------------------------- | ---------------------------- | ----------------- |
| **Capa de metadatos** | name + description | Precargada al inicio | 30-50 |
| **Capa de documento central** | Texto completo de SKILL.md | Se carga cuando es relevante | Cientos a miles |
| **Capa de recursos de referencia** | Archivos de referencia, plantillas, etc. | Se carga cuando se necesita | Bajo demanda |
Esto se alinea perfectamente con la esencia de los modelos de lenguaje — "ingresar texto para que el modelo comprenda". Los Skills no introducen protocolos complejos ni llamadas a API, sino que, mediante una estructura de texto cuidadosamente organizada, permiten que la IA obtenga y aplique conocimiento de manera eficiente. Simon Willison calificó este diseño como "increíblemente simple", precisamente porque resuelve un problema complejo de la forma más sencilla.
## Ventajas principales
### Eficiencia de tokens
La arquitectura de divulgación progresiva de los Skills proporciona una eficiencia de tokens extremadamente alta. Podemos entenderlo con una simple comparación:
| Enfoque | Consumo de tokens al inicio | Consumo total con 100 skills |
| ------------------------------------ | --------------------------- | ------------------------------------ |
| Enfoque tradicional (carga completa) | Miles a decenas de miles | Puede exceder la ventana de contexto |
| Skills (divulgación progresiva) | 30-50 | 3,000-5,000 |
Dado que los metadatos de cada skill solo ocupan unas pocas decenas de tokens, puede habilitar decenas o incluso cientos de Skills simultáneamente; el contenido completo se carga bajo demanda, sin desperdiciar valioso espacio de contexto.
### Composabilidad
Múltiples Skills pueden trabajar de forma colaborativa automáticamente. Cuando plantea una tarea compleja, Claude identifica de manera inteligente qué Skills necesita invocar y los coordina para completar la tarea en conjunto.
Por ejemplo, si dice "Genera un informe trimestral basado en estos datos de ventas", Claude podría:
1. Invocar el Skill de análisis de datos para procesar los datos crudos
2. Invocar el Skill de generación de gráficos para crear visualizaciones
3. Invocar el Skill de documentos para generar el informe final
Durante todo el proceso, no necesita especificar manualmente qué Skill usar; Claude selecciona y combina automáticamente según las necesidades de la tarea.
### Portabilidad
Un mismo Skill puede utilizarse en todas las plataformas del ecosistema de Anthropic:
| Plataforma | Descripción |
| ----------- | --------------------------------------------------------------- |
| Claude.ai | Versión web, adecuada para usuarios generales |
| Claude Code | Herramienta de línea de comandos, adecuada para desarrolladores |
| API | Integración programática, adecuada para desarrollo de sistemas |
El Skill de redacción de marca que crea para su equipo puede mantener un comportamiento consistente en todas estas plataformas, logrando verdaderamente **construir una vez, usar en cualquier lugar**.
> **Estrategias de otras plataformas de IA**: Actualmente, los Skills son una funcionalidad exclusiva de Anthropic. OpenAI emplea una estrategia dual de Custom GPTs + Assistants API (dos sistemas que no están unificados); Microsoft Copilot y Google Gemini se enfocan en la integración profunda con sus respectivos ecosistemas, en lugar de módulos de habilidades reutilizables. Claude Skills se considera una característica diferenciadora significativa.
### Datos de eficiencia
Según pruebas de referencia internas de Anthropic, los equipos que utilizan Skills **redujeron en un 73% el tiempo dedicado a ingeniería de prompts repetitivos**. Esto no solo significa un aumento en la eficiencia, sino más importante aún, la estandarización y reutilización de los flujos de trabajo — los miembros del equipo ya no necesitan mantener individualmente un conjunto de prompts, sino que comparten el mismo conjunto de Skills verificados.
## Resumen
Claude Skills son esencialmente **manuales de trabajo reutilizables** para asistentes de IA. Mediante la arquitectura de divulgación progresiva, logran una eficiencia de tokens extremadamente alta, permitiendo que la IA domine una gran cantidad de conocimiento especializado sin ocupar valioso espacio de contexto.
Recuerde tres palabras clave y habrá comprendido la esencia de los Skills:
| Palabra clave | Significado |
| -------------- | -------------------------------------------------------------------------- |
| **Eficiente** | Los metadatos ocupan solo unas pocas decenas de tokens, carga bajo demanda |
| **Composable** | Múltiples Skills trabajan colaborativamente de forma automática |
| **Portátil** | Experiencia consistente entre plataformas |
Ahora que comprende los conceptos, el siguiente artículo, la «[Guía práctica de Claude Skills](/es/docs/notes/claude-skills/practice)», le guiará en la práctica: cómo habilitar e instalar Skills, crear su primer Skill personalizado y evitar los errores comunes.
Si desea formalizar aún más sus flujos de trabajo, puede consultar «[Qué es el desarrollo dirigido por especificaciones](/es/docs/notes/speckit/concept)» para aprender cómo llevar la programación con IA de la "intuición" a la "ingeniería".
# Guía práctica
## Repaso rápido
En el [artículo anterior](/es/docs/notes/claude-skills/concept), conocimos los conceptos fundamentales de los Skills: son manuales de trabajo reutilizables para asistentes de IA que, mediante una arquitectura de divulgación progresiva, logran una eficiencia de tokens extremadamente alta, con tres características principales: eficientes, composables y portátiles. Este artículo parte desde una perspectiva práctica para ayudarle a comprender las diferencias entre Skills y otras funcionalidades, aprender a habilitar, instalar y crear Skills, y dominar las mejores prácticas mientras evita los errores comunes.
## Comparación de funcionalidades
El ecosistema de Claude cuenta con múltiples funcionalidades que pueden generar confusión al primer contacto. La siguiente tabla le ayudará a distinguirlas rápidamente:
| Funcionalidad | Qué es | Más adecuada para | Persistencia |
| ------------- | -------------------------------------- | --------------------------------------------------- | ------------------------------------------ |
| **Skills** | Paquetes de conocimiento especializado | Tareas repetitivas, flujos estandarizados | Persistente entre conversaciones |
| **Prompts** | Instrucciones instantáneas | Solicitudes puntuales | Solo la conversación actual |
| **Projects** | Base de conocimientos | Información de contexto, documentación del proyecto | Dentro del espacio de trabajo del proyecto |
| **MCP** | Conector | Datos externos, llamadas a API | Conexión continua |
| **Subagents** | Subagentes | Delegación de tareas, procesamiento paralelo | Entre sesiones |
### Skills vs MCP
Esta es la confusión más común. La diferencia fundamental: **MCP conecta a Claude con los datos, Skills le enseña a Claude cómo procesar los datos**. Ambos son complementarios, no sustitutos.
| Dimensión | Skills | MCP |
| ----------------------- | ------------------------------------------------------------- | ---------------------------------------------------------------- |
| **Función principal** | Enseña a Claude cómo ejecutar tareas | Conecta a Claude con sistemas externos |
| **Consumo de tokens** | Muy bajo (decenas de tokens) | Alto (miles a decenas de miles de tokens) |
| **Complejidad técnica** | Simple (Markdown + YAML) | Compleja (especificación completa de protocolo) |
| **Escenarios típicos** | Redacción de marca, generación de informes, flujos de trabajo | Consultas a bases de datos, llamadas a API, servicios en la nube |
| **Portabilidad** | Entre Claude.ai/Code/API | Adoptado por múltiples compañías de modelos |
Comprendiendo esta diferencia, sabrá cuándo usar cada uno. Cuando necesite consultar bases de datos, llamar a API o acceder a servicios en la nube, use MCP; cuando necesite seguir un estilo de escritura específico, ejecutar flujos estandarizados o reutilizar conocimiento especializado, use Skills.
La mejor práctica es combinar ambos: use MCP para conectar su sistema CRM y obtener datos de clientes, y Skills para definir cómo analizar esos datos y generar informes.
### Skills vs Subagents
La diferencia fundamental: **Skills hacen que Claude sea más competente en cierto tipo de tareas; Subagents permiten que Claude delegue tareas a "empleados expertos" independientes**.
| Dimensión | Skills | Subagents |
| ---------------------- | ------------------------------------------------------------ | ----------------------------------------------------- |
| **Función principal** | Proporcionar conocimiento especializado e instrucciones | Subagentes que ejecutan tareas de forma independiente |
| **Contexto** | Se inyecta en el contexto de la conversación principal | Posee su propia ventana de contexto independiente |
| **Escenarios de uso** | Hacer que Claude sea más competente en cierto tipo de tareas | Tareas complejas de múltiples pasos e independientes |
| **Modo de activación** | Coincidencia automática según la descripción | Invocación manual o delegación automática de Claude |
| **Portabilidad** | Entre Claude.ai/Code/API | Solo Claude Code y Agent SDK |
De forma ilustrativa, los Skills son como material de capacitación — enseñan a Claude cómo hacer algo; los [Subagents](/es/docs/notes/claude-subagent) son como empleados especializados — tienen su propio escritorio (contexto) y permisos (herramientas), completan tareas de forma independiente y luego reportan los resultados.
Ambos pueden usarse en combinación: por ejemplo, un subagente de revisión de código puede cargar un Skill de mejores prácticas específico para un lenguaje, logrando el efecto combinado de "experto + conocimiento especializado". Según [investigaciones de Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system), los sistemas multiagente (Claude Opus 4 como agente principal + Claude Sonnet 4 como subagentes) superaron en un 90.2% al agente único en evaluaciones internas.
### Skills vs comandos de barra
Si ha usado Claude Code, seguramente conoce [comandos de barra](/es/blog/claude-code-best-practices) como `/commit` y `/review`. La diferencia fundamental: **Skills se activan automáticamente según el contexto; los comandos de barra requieren entrada manual para activarse**.
| Dimensión | Skills | Comandos de barra (Slash Commands) |
| --------------------------- | ------------------------------------------------------ | -------------------------------------------- |
| **Modo de activación** | Activación automática (según coincidencia de contexto) | Entrada manual (como `/commit`) |
| **Condición de activación** | Claude determina la relevancia según la description | El usuario ingresa explícitamente el comando |
| **Escenarios de uso** | Mejora de capacidades "siempre activa" | Operaciones claras y repetibles |
| **Percepción del usuario** | Sin percepción, se activa automáticamente | Necesita recordar el nombre del comando |
Ejemplo: cuando ingresa `/commit`, Claude ejecuta un flujo de commit predefinido — esto es un comando de barra; cuando dice "ayúdame a escribir un informe semanal", Claude identifica automáticamente y carga el Skill de generación de informes semanales, sin necesidad de que ingrese ningún comando — esto son Skills.
Para recordarlo fácilmente: los comandos de barra son atajos de teclado que requieren activación manual; los Skills son conocimiento de fondo que Claude usa automáticamente cuando lo considera necesario.
### Skills vs Plugins
Los Plugins son el mecanismo de paquetes de extensión de Claude Code. La diferencia fundamental: **Skills son extensiones de capacidades que se activan automáticamente; Plugins son configuraciones de flujo de trabajo completas empaquetadas para distribución**.
| Dimensión | Skills | Plugins |
| ----------------------------- | --------------------------------------- | ----------------------------------------------- |
| **Función principal** | Extensión de capacidades especializadas | Empaquetado y distribución de flujos de trabajo |
| **Modo de activación** | Activación automática según contexto | Los componentes se fusionan tras la instalación |
| **Alcance de uso** | Multiplataforma (Claude.ai/Code/API) | Solo Claude Code |
| **Contenido incluido** | Instrucciones + scripts + recursos | Comandos de barra + hooks + skills |
| **Mecanismo de distribución** | Carpeta individual | Instalación mediante marketplace |
Punto clave: los Plugins pueden contener Skills (en el directorio `skills/`), siendo una unidad de empaquetado más grande. Al instalar un Plugin, los Skills incluidos se activan automáticamente, los comandos de barra aparecen en el autocompletado y los hooks se fusionan con la configuración existente.
En resumen: use Skills para expandir las capacidades de Claude, y Plugins para distribuir configuraciones de flujo de trabajo estandarizadas entre equipos.
## Tutorial práctico
### Método 1: Habilitar Skills integrados
Esta es la forma más sencilla de comenzar. Anthropic ofrece oficialmente un conjunto de skills de documentos prácticos:
| Skill | Funcionalidad |
| --------------------- | ------------------------------------------------------------------------------- |
| **Excel (xlsx)** | Crear hojas de cálculo, analizar datos, generar informes con gráficos |
| **PowerPoint (pptx)** | Crear presentaciones, editar diapositivas, analizar contenido de presentaciones |
| **Word (docx)** | Crear documentos, editar contenido, formatear texto |
| **PDF (pdf)** | Generar documentos PDF formateados e informes |
**Pasos para habilitar**:
1. Inicie sesión en [Claude.ai](https://claude.ai)
2. Haga clic en el avatar en la esquina superior derecha y acceda a **Settings**
3. Encuentre la opción **Capabilities**
4. Habilite los skills que necesite
Una vez habilitados, puede probar directamente: "Ayúdame a crear una hoja de cálculo Excel con el presupuesto de ventas del Q3, incluyendo desglose mensual y totales".
> **Nota**: Requiere un plan Pro, Max, Team o Enterprise, y la función de ejecución de código debe estar habilitada.
### Método 2: Instalar Skills de la comunidad
Si utiliza Claude Code, puede instalar Skills contribuidos por la comunidad mediante comandos.
**Instalación desde el marketplace de plugins**:
```bash
# Agregar el repositorio oficial de Skills
/plugin marketplace add anthropics/skills
# Instalar el paquete de skills de documentos
/plugin install document-skills@anthropic-agent-skills
# Instalar el paquete de skills de ejemplo
/plugin install example-skills@anthropic-agent-skills
```
**Ubicación de almacenamiento de Skills**:
| Ubicación | Ruta | Descripción |
| ------------------ | ------------------- | ---------------------------------------------- |
| Skills personales | `~/.claude/skills/` | Disponibles solo para usted |
| Skills de proyecto | `.claude/skills/` | Versionados con git, compartidos con el equipo |
### Método 3: Crear un Skill personalizado
Aquí es donde reside el verdadero poder de los Skills — crear flujos de trabajo exclusivos para usted.
**Paso 1: Crear la estructura de carpetas**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
Una carpeta de Skill completa podría verse así:
```
weekly-report/
├── SKILL.md # Instrucciones principales (requerido)
├── template.md # Plantilla de informe semanal (opcional)
└── examples/ # Informes semanales de ejemplo (opcional)
├── good-example.md
└── bad-example.md
```
**Paso 2: Escribir el SKILL.md**
SKILL.md es el núcleo de todo el Skill. Se compone de dos partes: el frontmatter YAML (metadatos) y el cuerpo en Markdown (instrucciones detalladas).
**Metadatos requeridos**:
| Campo | Requisito | Descripción |
| ------------- | --------------------- | --------------------------------------------------------- |
| `name` | Máximo 64 caracteres | Nombre identificador único del skill |
| `description` | Máximo 200 caracteres | Indica a Claude cuándo usar este skill (¡muy importante!) |
**Metadatos opcionales**:
| Campo | Descripción |
| --------------- | ------------------------------------------------------------------ |
| `dependencies` | Paquetes de software requeridos, como `python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | Lista de herramientas permitidas |
| `model` | Anulación opcional del modelo |
Ejemplo completo de un Skill de generación de informes semanales:
```yaml
---
name: weekly-report
description: 根据本周工作内容生成标准化的周报,包含进展、问题和下周计划
---
# 周报生成助手
## 使用场景
当用户需要生成周报、工作总结或进度汇报时,使用此技能。
## 输出格式
请按以下结构生成周报:
### 本周完成
- 列出已完成的主要工作项
- 每项包含简短说明和成果
### 进行中
- 列出正在进行的工作
- 标注当前进度和预期完成时间
### 遇到的问题
- 列出阻碍进展的问题
- 如果有,说明需要的支持
### 下周计划
- 列出下周的主要任务
- 按优先级排序
## 风格要求
- 使用简洁的表达
- 避免过于技术化的术语
- 突出成果和影响
## 示例
**输入**:这周完成了用户登录功能,修复了 3 个 bug,参加了产品评审。
**输出**:
### 本周完成
- 用户登录功能开发:完成前后端联调,支持邮箱和手机号登录
- Bug 修复:解决了 3 个高优先级问题,提升系统稳定性
### 进行中
- (无)
### 遇到的问题
- (无)
### 下周计划
- 开始用户注册功能开发
- 编写单元测试用例
```
**Paso 3: Probar**
Pruebe en Claude: "Ayúdame a generar el informe semanal. Esta semana completé el desarrollo de la función de inicio de sesión de usuarios, corregí 3 bugs y participé en dos reuniones de revisión de producto."
### Uso de Skill Creator
Si no desea escribir el SKILL.md desde cero, Claude tiene integrado un skill-creator que puede guiarlo de forma interactiva:
```
Help me create a skill for [your workflow]
```
Claude le hará una serie de preguntas para ayudarle a organizar sus necesidades y luego generará un borrador inicial de SKILL.md.
## Principios técnicos
### Skills como sistema de meta-herramientas
Los Skills son esencialmente un **sistema de meta-herramientas** — no ejecutan código directamente, sino que inyectan instrucciones especializadas en el contexto de la conversación, modificando la forma en que Claude razona.
Cuando activa un Skill, ocurren dos cosas:
1. **Mensaje de metadatos**: Un indicador de estado visible que muestra qué Skill se está cargando
2. **Prompt del skill**: Las instrucciones completas del SKILL.md se envían a Claude, pero permanecen ocultas para el usuario
### Mecanismo de descubrimiento y selección
¿Cómo sabe Claude qué Skill invocar? La respuesta es: **depende completamente de la comprensión del lenguaje**.
Los campos name y description de todos los Skills habilitados se formatean como una lista dinámica e incluyen en el prompt del sistema. Cuando usted envía un mensaje, Claude utiliza su capacidad nativa de comprensión del lenguaje para asociar su intención y decidir si debe invocar algún Skill.
Por eso el campo `description` es tan importante — es la única base de juicio de Claude. No hay algoritmos de enrutamiento complejos; la decisión se completa enteramente dentro del proceso de razonamiento de Claude.
## Mejores prácticas
Tras una amplia experimentación, la comunidad ha resumido cuatro reglas de oro para la creación de Skills:
**1. Mantener el enfoque**
Un Skill debe hacer solo una cosa. Múltiples Skills enfocados son mucho más útiles que un solo Skill que intenta abarcarlo todo; esto no solo facilita el mantenimiento sino también la composición.
**2. Descripción clara**
El campo description determina cuándo Claude invocará su Skill; es esencial describir claramente los escenarios de uso. "Generar un informe de análisis trimestral basado en datos de ventas" es una buena descripción; "procesar datos" es demasiado genérico.
**3. Proporcionar ejemplos**
Incluir ejemplos de entrada y salida en SKILL.md mejora significativamente la estabilidad de los resultados, especialmente para tareas con requisitos de formato específicos.
**4. Comenzar con lo simple**
Empiece con instrucciones básicas en Markdown puro, valide los resultados y luego considere agregar scripts, aumentando la complejidad gradualmente.
### Solución de problemas comunes
| Problema | Causa posible | Solución |
| --------------------- | -------------------------------------------- | ----------------------------------------------------------------- |
| El Skill no se activa | La description no es suficientemente precisa | Reescribir con una descripción de escenario de uso más específica |
| El Skill no se activa | El Skill no está correctamente instalado | Verificar la ruta de archivos y nombres |
| Resultados inestables | Faltan ejemplos | Agregar más ejemplos de entrada/salida |
| Resultados inestables | Instrucciones demasiado vagas | Agregar restricciones y requisitos de formato |
| Carga lenta | Archivo demasiado grande | Mover archivos grandes al subdirectorio references |
### Consideraciones de seguridad
Los Skills pueden ejecutar código, por lo que la seguridad es muy importante:
* **Fuentes confiables**: Use solo Skills de canales confiables
* **Revisar scripts**: Examine el código de los scripts incluidos en los Skills antes de instalarlos
* **Proteger información sensible**: No codifique claves de API o contraseñas directamente en los Skills
* **Gestión de permisos**: Al usar en equipo, preste atención al alcance de compartición de los Skills
## Limitaciones actuales
Como funcionalidad emergente, los Skills actualmente tienen algunas limitaciones:
| Limitación | Descripción |
| ----------------------------------- | ------------------------------------------------------------------------------------------------- |
| ~~**Solo ecosistema Anthropic**~~ | ✅ **Resuelto** - ver explicación abajo |
| **Falta de mecanismo de auditoría** | Aún no hay flujo de revisión o auditoría integrado |
| **Curva de aprendizaje** | Los equipos necesitan ajustar sus flujos de trabajo y establecer procesos de gestión de versiones |
| **Etapa emergente** | El ecosistema aún está en desarrollo |
> **Actualización importante (18 de diciembre de 2025)**: Anthropic lanzó oficialmente los Agent Skills como [estándar abierto](https://agentskills.io). La especificación y el SDK de referencia están disponibles públicamente en [agentskills.io](https://agentskills.io).
>
> **Empresas/productos que lo han adoptado**:
> 
>
> * **Microsoft**: VS Code y GitHub ya lo integran
> * **OpenAI**: ChatGPT y Codex CLI adoptan la misma arquitectura
> * **Herramientas de programación**: Cursor, Goose, Amp, OpenCode
> * **Skills de socios**: Atlassian, Figma, Canva, Stripe, Notion, Zapier
>
> Simultáneamente, Anthropic, OpenAI y Block cofundaron la [Agentic AI Foundation](https://www.linuxfoundation.org/) (alojada por la Linux Foundation), a la que también se han unido Google, Microsoft y AWS. Esto significa que los Skills están evolucionando de una funcionalidad de un solo proveedor a un estándar de la industria; los Skills escritos para Claude Code pueden interoperar con OpenAI Codex CLI.
>
> Fuentes de referencia:
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## Recursos de aprendizaje
### Recursos oficiales
| Recurso | Enlace | Descripción |
| -------------------------------- | -------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
| Repositorio de Skills en GitHub | [anthropics/skills](https://github.com/anthropics/skills) | Ejemplos oficiales, 22k+ Stars |
| Documentación de Claude Code | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Guía de uso de Skills |
| Centro de ayuda | [support.claude.com](https://support.claude.com) | Preguntas frecuentes |
| Blog técnico | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Análisis técnico en profundidad |
| Inicio rápido de API | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | Guía de integración para desarrolladores |
| Estándar abierto de Agent Skills | [agentskills.io](https://agentskills.io) | Especificación oficial y SDK |
### Selección de la comunidad
| Recurso | Enlace | Descripción |
| --------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------ |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Colección curada de Skills |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148+ comandos de barra, 54 agentes de IA |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | Skills para crear y editar documentos de oficina |
### Lecturas recomendadas
| Artículo | Autor |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic (oficial) |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## Perspectivas
La aparición de los Skills representa una dirección importante en el desarrollo de herramientas de IA — permitir que la IA no solo ejecute tareas, sino que también aprenda y memorice formas de trabajo específicas. Simon Willison predice que los Skills desencadenarán una "explosión cámbrica" en el ámbito de las herramientas de IA, y esta predicción no es una exageración.
A medida que más desarrolladores y equipos comiencen a construir y compartir Skills, es posible que veamos:
* **Mercado de Skills especializados**: Expertos de diversas industrias empaquetarán su conocimiento en Skills reutilizables
* **Integración profunda entre Skills y MCP**: Formando flujos de trabajo completos de extremo a extremo
* **Plataformas empresariales de Skills**: Colaboración en equipo, gestión de versiones, control de permisos
Ahora es un buen momento para comenzar. De inmediato, puede iniciar sesión en Claude.ai y habilitar los skills de documentos; esta semana puede intentar instalar un Skill de la comunidad y crear su primer Skill simple; a largo plazo, identificar las tareas repetitivas en su equipo y construir gradualmente una biblioteca de skills propia será una vía efectiva para mejorar la eficiencia.
### Lecturas adicionales
* «[Análisis completo de la arquitectura de Claude](/es/docs/notes/claude-architecture)» — La posición de los Skills dentro del sistema completo de Claude
* «[Guía completa de Claude Subagent](/es/docs/notes/claude-subagent)» — Comprender en profundidad el mecanismo de Subagent
* «[Mis mejores prácticas con Claude Code](/es/blog/claude-code-best-practices)» — Consejos de uso diario de Claude Code
# Análisis en profundidad de Skill-Creator: utilice datos para impulsar el desarrollo de sus habilidades
## Introducción
Este artículo está basado en información de marzo de 2026 y corresponde a Claude Code v2.1+.
Si ha leído [Concepto](/es/docs/notes/claude-skills/concept) y [Práctica](/es/docs/notes/claude-skills/practice), ya debería saber cómo escribir manualmente un archivo SKILL.md: defina frontmatter, escriba un comando, guárdelo en el directorio `.claude/skills/` y listo.
Pero aquí hay una pregunta fundamental: \*\*¿Cómo sabes que tus habilidades son realmente útiles? \*\*
Es posible que haya cambiado la redacción de un párrafo y sienta que funciona mejor, pero ese es sólo su sentimiento subjetivo. Quizás con una palabra clave diferente, la nueva versión sería peor. Tal vez tus habilidades no mejoren en absoluto en comparación con no tener habilidades: Claude puede hacerlo igual de bien por sí solo.
En los capítulos conceptual y práctico, el proceso de desarrollo de habilidades es el siguiente: **Escrito → Probado → Sentirse bien → En línea**. Todo el proceso se basa en la intuición, no hay cuantificación y no hay forma de responder "¿Cuánto mejor es esta habilidad que ninguna habilidad en absoluto?" Y Skill-Creator convirtió esto en ingeniería: **Escrito → Pruebas paralelas con/sin habilidades → Comparación de prueba ciega A/B → Puntuación cuantitativa → Iteración de retroalimentación → Verificación de datos**.
Por eso existe Skill-Creator. No solo le ayuda a "generar un SKILL.md", sino que también proporciona un conjunto completo de bucles crear → probar → evaluar → optimizar, lo que le permite dejar que sus datos hablen por sí mismos.
## ¿Qué es Skill-Creator?
Skill-Creator en sí también es una habilidad: un archivo SKILL.md de 33 KB que además admite archivos de guía de subagente, secuencias de comandos Python y visores HTML. Su estructura de directorios se ve así:
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
La instalación también es muy sencilla:
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## Siga esto nuevamente: Evaluar y optimizar una habilidad existente
Repasemos el proceso completo de Skill-Creator usando las habilidades que realmente uso. Mantengo un mercado de complementos de Claude Code [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub), en el que la habilidad `yux-video-summary` se usa para convertir subtítulos de video en resúmenes estructurados, lo que admite detección de idiomas chino e inglés, dos modos de salida DUAL\_FILE/SINGLE\_FILE, limpieza de palabras de relleno, etc. El SKILL.md de la habilidad tiene este aspecto:
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
La habilidad ha sido escrita, pero ¿cómo sabes que es realmente útil? \*\* Aquí es donde Skill-Creator entra en escena.
> Hay un principio de escritura importante en el código fuente de Skill-Creator: *"Intenta explicar el **por qué** detrás de todo. Si te encuentras escribiendo SIEMPRE o NUNCA en mayúsculas, eso es una señal de alerta: reformula y explica el razonamiento".* Significado: una buena habilidad debe **explicar el por qué**, en lugar de acumular reglas rígidas.
### Paso 1: crear casos de prueba y ejecutar la evaluación
Pregunta central: \*\*¿Es esta habilidad realmente mejor que ninguna habilidad? \*\*
Abra Claude Code e ingrese directamente:
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator primero leerá definiciones y esquemas de habilidades y luego generará automáticamente casos de prueba y afirmaciones cuantitativas. Mi ejecución generó 3 casos de prueba y 39 afirmaciones:
Tenga en cuenta que no compila casos de prueba de manera casual: comprende los dos modos de salida de DUAL\_FILE y SINGLE\_FILE definidos en la habilidad y diseña específicamente escenarios que cubren diferentes tipos de videos (tutoriales, entrevistas de podcasts, intercambio de tecnología) y combinaciones de idiomas. El diseño de Assertions también es muy particular, desde la detección de idioma, la selección del modo de salida hasta la calidad del contenido y la limpieza de palabras de relleno en chino e inglés, es mucho más completo de lo que quiero probar las dimensiones yo mismo.
Luego, el sistema inicia dos subagentes independientes simultáneamente para cada caso de prueba: with\_skill (cargando habilidades) y \*\* without\_skill \*\* (línea base, no se cargan habilidades). Se iniciaron **6 agentes paralelos** (3 casos de prueba × 2 versiones) al mismo tiempo, cada uno ejecutándose en un **árbol de trabajo independiente** sin interferir entre sí.
> La habilidad PDF de Anthropic anteriormente tenía problemas para manejar formularios que no se pueden completar: Claude necesitaba colocar texto en coordenadas precisas sin definir campos. El punto de falla se aisló mediante Eval y posteriormente el equipo arregló la lógica de posicionamiento. Ese es el valor de Eval: convertir "algo que no se siente bien" en "qué es exactamente lo que está mal aquí".
### Paso 2: tres subagentes retransmiten la puntuación
Una vez completadas todas las operaciones, los tres subagentes profesionales **automáticamente** aparecen en secuencia:
**Calificador** Verifica las afirmaciones una por una. Verificará si el resumen de la versión with\_skill contiene la tabla de descripción general, si el modo DUAL\_FILE está seleccionado correctamente, si la palabra de relleno se ha limpiado y luego registrará el aprobado/reprobado y la evidencia de cada elemento, generando `grading.json`:
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**Comparador** hace una comparación ciega A/B: recibe dos resúmenes, pero **no sabe cuál es la versión de habilidad y cuál es la versión básica**. Solo ve la "Salida A" y la "Salida B" y las juzga de forma independiente según sus propios estándares de calidad para determinar el ganador.
**Analizador** combina los resultados anteriores para hacer un diagnóstico: qué afirmaciones pasaron independientemente de las habilidades o no (lo que indica que esta afirmación no tiene diferenciación y debe ser reemplazada por una mejor afirmación), qué resultados tienen una alta variación (la prueba es inestable) y cuál es la compensación entre tiempo y token. Finalmente, se dan sugerencias de mejora.
### Paso 3: Revisar los resultados en Eval Viewer
Una vez que se completa la puntuación, Skill-Creator abrirá automáticamente un visor HTML en su navegador.
**Pestaña Resultados** Puede ver el resultado de cada caso de prueba uno por uno. Hay un cuadro de texto de comentarios en la parte inferior: escriba lo que crea que no es lo suficientemente bueno, como "el resumen carece de una línea de tiempo" y "la palabra de relleno no está limpia". Después de leer todos los casos de uso, haga clic en **Enviar todas las reseñas** y los comentarios se guardarán en `feedback.json`.
**Pestaña Resultados de referencia** Puede ver la comparación cuantitativa: la tasa de aprobación, el consumo de tiempo, el consumo de tokens de with\_skill y without\_skill, así como la comparación elemento por elemento de cada afirmación.
### Paso 4: Iterar y mejorar hasta estar satisfecho
Vuelva a Claude Code y dígale que ha terminado de enviar comentarios. Skill-Creator leerá `feedback.json` y brindará análisis y sugerencias de mejora basadas en los datos de referencia:
Mi habilidad funcionó bien con una tasa de aprobación del 97%. Skill-Creator identificó con precisión un pequeño problema: el video de la entrevista carecía de párrafos de Citas notables e hizo sugerencias para solucionarlo.
La clave es que no parchea casos de prueba individuales: generaliza sus comentarios, comprende los requisitos detrás de ellos y ajusta la estructura general de la habilidad, luego reescribe SKILL.md, vuelve a ejecutar todas las pruebas en el directorio `iteration-2/` y abre un nuevo Visor de evaluación para que pueda comparar el resultado de las dos rondas. Este ciclo continúa hasta que esté satisfecho.
> Una filosofía de mejora notable en el código fuente de Skill-Creator: *"Estamos tratando de crear habilidades que se puedan usar un millón de veces en muchos mensajes diferentes. En lugar de realizar cambios complicados y excesivos o DEBE opresivamente restrictivos, si hay algún problema persistente, intente diversificarse y usar diferentes metáforas".* Idea central: **Evitar el sobreajuste** para probar casos y buscar capacidades de generalización.
### Paso 5 (opcional): Optimice la descripción para que la habilidad se active en el momento adecuado
Se verifica la calidad de la habilidad, pero hay otra cuestión que fácilmente se pasa por alto: el campo `description` de la habilidad determina cuándo Claude la llamará.
Entrada:
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator genera automáticamente alrededor de 20 consultas de evaluación (la mitad debe activarse, la otra mitad no debe activarse) y la interfaz de revisión se abre en el navegador:
Tenga en cuenta que estas consultas están disponibles tanto en chino como en inglés y cubren una variedad de expresiones reales. Las consultas "No deberían activar" no deberían ser demasiado escandalosas; un buen contraejemplo es "Ayúdame a resumir las actas de esta reunión", que comparte la palabra clave "resumen" con el resumen en video, pero en realidad requiere habilidades de procesamiento de documentos en lugar de resumen en video.
Puede editar el texto de la consulta directamente en la página, hacer clic en **+ Agregar consulta** para agregar nuevas consultas, usar el botón Eliminar para eliminar las inapropiadas y también puede alternar el interruptor Debería activar para cada consulta. Después de confirmar que es correcto, haga clic en **Exportar conjunto de evaluación** para exportar el archivo JSON. Vuelve a Claude Code y dile que lo has exportado. El sistema ejecutará automáticamente el ciclo de optimización en segundo plano:
Todo el proceso está completamente automatizado: divida la consulta en un conjunto de entrenamiento y un conjunto de prueba 60/40, optimice iterativamente la descripción en el conjunto de entrenamiento (hasta 5 rondas) y use los resultados del conjunto de prueba para seleccionar la mejor versión para evitar el sobreajuste. Después de ejecutar, se generará la comparación de descripciones antes y después de la optimización:
La descripción optimizada se vuelve más específica: aclara los tipos de archivos admitidos (.vtt/.srt), enfatiza las características de la canalización (limpieza de relleno, lógica DUAL/SINGLE\_FILE) y utiliza MUST USE para excluir escenarios que no deberían activarse. Anthropic utilizó internamente este conjunto de optimizadores para ejecutar sus propias habilidades de creación de documentos. Como resultado, se mejoró la precisión de activación de 5 de 6 habilidades públicas.
### Uso avanzado: inyección de contexto dinámico
Si desea que la habilidad inyecte contexto automáticamente al cargar, puede incrustar un comando de shell en SKILL.md usando la sintaxis `!` de Skills 2.0:
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
Estos comandos se ejecutan antes de que Claude vea la habilidad y los datos se incrustan directamente en el mensaje. En comparación con dejar que Claude explore los archivos uno por uno, ahorra mucho tiempo y tokens.
## Dos tipos de habilidades: ¿Cuál deberías crear?
Antes de utilizar Skill-Creator, es necesario comprender los dos tipos de habilidades definidos por Anthropic:
**Tipo de mejora de habilidad**: deja que el modelo haga cosas que antes no podía o no podía hacer bien. Por ejemplo:
* Habilidades de generación de imágenes: Claude no puede generar imágenes de forma nativa, pero puede lograrlo llamando a herramientas como nanobanner a través de habilidades.
* Habilidades de diseño front-end: los diseños de IA predeterminados suelen tener mucho "sabor a IA" y unas buenas habilidades de diseño pueden mejorar considerablemente la calidad.
**Preferencia de codificación**: solidifique su flujo de trabajo específico. El modelo ya tiene capacidades individuales, pero necesita un orden preciso de ejecución. Por ejemplo:
* Habilidades de revisión de relaciones públicas: verificar la seguridad del código de acuerdo con procedimientos fijos y generar informes de nivel de riesgo.
* Habilidades de resumen de vídeo: salida según una estructura de plantilla específica, detección automática de idioma y limpieza de palabras de relleno
Las razones por las que es necesario probar estos dos tipos de habilidades son diferentes: **Tipo de mejora de capacidad** puede volverse innecesario a medida que el modelo evoluciona: si la línea base (sin\_skill) también puede pasar todas las afirmaciones, significa que el modelo es lo suficientemente nativo y esta habilidad se puede retirar; **El tipo de codificación** es más duradero, pero debes verificar si es realmente fiel a tu flujo de trabajo.
Las capacidades de evaluación de Skill-Creator le permiten verificar continuamente si una habilidad sigue siendo valiosa, en lugar de utilizar ciegamente una habilidad que puede estar desactualizada.
## Lo que dice la comunidad
La actualización de Skill-Creator generó mucha discusión, desde X/Twitter hasta Reddit y blogs independientes, y los comentarios reales son más valiosos que la documentación oficial.
### ¿Es realmente útil? Los datos hablan
La pregunta más directa: ¿Es realmente mejor agregar habilidades que no agregar habilidades? \*\* Varias mediciones reales dan una respuesta clara.
Reddit u/hashpanak realizó una evaluación de la habilidad de generación de títulos y obtuvo una tasa de aprobación del 100 % con\_skill y solo del 60 % sin\_skill. Cuando se le preguntó si el costo del token valía la pena, respondió: "Por supuesto. Después de la optimización, las tareas repetidas se pueden convertir en scripts, lo que ahorrará tokens". u/spences10 es aún más extremo: ejecutó 250 evaluaciones de sandbox y aumentó la tasa de activación de habilidades del 84 % al 100 %. La sección de comentarios u/Manfluencer10kultra dijo: "**Esto debería convertirse en una práctica estándar.**"
El blogger [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/) comparó las habilidades de seguridad de WordPress: se aprobaron las 21 afirmaciones (la línea de base fue solo el 90,5%) y la velocidad fue un 9,9% más rápida. Su resumen: **"Las habilidades solían ser arte, ahora son ingeniería".**
@0zhuxiaofeng dio cifras más específicas desde la perspectiva del flujo de trabajo real: "Después de usarlo durante un mes, el mayor cambio es que run\_eval permite que las habilidades se califiquen por sí mismas. El agente que ejecuto las operaciones de contenido ahora evalúa automáticamente el efecto después de cada lanzamiento, y las habilidades deficientes se eliminan y reescriben directamente. **La intervención manual se ha reducido de 3 horas por día a media hora**".
### Punto ciego pasado por alto: Activador ≠ Calidad
El blogger [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/) señaló un punto ciego que nadie mencionó: **Las habilidades pueden pasar la evaluación de calidad pero fallan en la evaluación desencadenante**: la calidad del resultado es muy buena, pero nunca será calificada. Después de tres rondas de optimización `run_loop.py`, activó la evaluación al 13/13. Información básica: "La descripción de una habilidad no son metadatos, sino un parámetro que se puede aprender; es necesario optimizar el comportamiento de enrutamiento real".
Esto coincide con la sugerencia de @DrWang5257: "No reescriba todo de una vez. Primero divídalo en tres secciones: condiciones de activación, plantillas de entrada y respaldo de fallas, e itere paso a paso. De esta manera, la velocidad de actualización es rápida y la tasa de renovación es baja".
### Puntos débiles reales
Aunque el efecto es bueno, también existen muchos inconvenientes:
* **el consumo de tokens es enorme**. @konghao10 dijo sin rodeos que "el consumo de tokens es enorme": ejecutar 6 agentes paralelos al mismo tiempo no es realmente barato. Reddit u/munkymead también dijo que "es caro hacerse una prueba seria".
* **Si tienes demasiadas habilidades, pelearás**. [El blogger de RoboRhythms, Noah Albert](https://www.roborhythms.com/best-claude-code-skills-2026/) descubrió que **empezar a tener problemas cuando las habilidades alcanzan 8-10**: Claude se autocuestionará el resultado, generará prefacios más detallados y ocasionalmente tendrá conflictos de comando entre habilidades. Sin embargo, Reddit u/Specialist\_Solid523 respondió: "Las habilidades mal escritas solo consumen contexto. **Las habilidades bien escritas casi siempre hacen que el uso de tokens sea más eficiente.**"
* **SKILL.md se alarga con más iteraciones**. Reddit u/IulianHI señaló una contradicción: con mejoras iterativas, los archivos de habilidades continúan expandiéndose, \*\* pero desplazan la ventana de contexto para hacer las cosas realmente \*\*. Los casos de prueba que solo cubren el camino feliz no alcanzan el 5% crítico.
* **Falta la gestión de versiones**. @fengqve se queja "¿Por qué Skill **no tiene el concepto de versión**? Se ha actualizado tantas veces que es difícil describir qué actualización es". Esto es especialmente doloroso después de múltiples rondas de iteraciones.
* **El modo sin cabeza tiene un error**. Hay un problema clave en GitHub: la habilidad nunca se activa en el modo `claude -p`, lo que hace que la recuperación que describe el ciclo de optimización sea siempre del 0% ([#36570](https://github.com/anthropics/claude-code/issues/36570)).
### Pensando más allá: superación personal recursiva
@vista8 compartió un artículo relacionado \[Memento-Skills: Let Agents Design Agents] ([https://github.com/Memento-Teams/Memento-Skills](https://github.com/Memento-Teams/Memento-Skills)), y alguien en el área de comentarios lo resumió con precisión: "El principal cuello de botella de Skill es la iteración: es fácil escribir la primera versión, pero es difícil mejorarla y usarla mejor en escenarios reales. Si puedes automatizar este ciclo de 'usar → evaluar → mejorar', es equivalente a instalar un motor de autoevolución para Agent".
Un hilo similar al 104 en Reddit r/ClaudeAI también analiza esta dirección. Pero el comentario principal le echó agua fría: u/Tatrions dijo: "El bucle recursivo funciona, pero la parte difícil es saber cuándo confiar en las mejoras. Descubrimos que tenemos que buscar evidencia: no realizar cambios a menos que ocurra una falla al menos dos veces. De lo contrario, cada ciclo está 'arreglando' algo que no está roto en primer lugar, y termina siendo peor".
## Instalación y Ecología
Skill-Creator, como una de las habilidades mantenidas oficialmente por Anthropic, está incluido en el almacén [anthropics/skills](https://github.com/anthropics/skills), que contiene más de 17 habilidades de nivel de producción.
El ecosistema de habilidades más amplio también está creciendo rápidamente: [skills.sh](https://skills.sh) El mercado ofrece una experiencia conveniente de descubrimiento e instalación, y la comunidad ha mantenido más de 1234 habilidades de agentes.
## Escribe al final
El problema central que resuelve Skill-Creator es: \*\*¿Cómo sabes que tus habilidades son realmente efectivas? \*\*
En ausencia de ello, el desarrollo de habilidades se basa en "escribir → intentar → sentirse bien". Con Skill-Creator puedes:
* Prueba efectos tanto para expertos como para no expertos con **Parallel Agent**
* Elimine el sesgo de evaluación con **Comparación ciega A/B**
* Visualice resultados y deje comentarios con **Eval Viewer**
* Utilice **Optimizador de descripciones** para controlar con precisión el tiempo de activación de las habilidades.
* Utilice **bucles iterativos** para mejorar continuamente hasta que esté satisfecho
Esto está en línea con el concepto de desarrollo basado en pruebas en ingeniería de software: no "simplemente escribir el código y pensar que se puede ejecutar", sino "utilizar pruebas para demostrar que realmente funciona como se espera".
Anthropic presentó una perspectiva interesante en el blog oficial: a medida que las capacidades del modelo mejoran, SKILL.md puede evolucionar de un "plan de implementación" (decirle a Claude **cómo**) a una "descripción de la especificación" (decirle a Claude **qué** y dejar que el modelo lo resuelva por sí solo). El marco Eval es el primer paso en esta dirección: Eval describe "qué hacer". Si algún día esta descripción es suficiente para convertirse en una habilidad, entonces el sistema de prueba establecido por Skill-Creator será aún más importante.
Si ya utiliza Habilidades, intente usar `/skill-creator` para realizar una evaluación de sus habilidades más utilizadas. Es posible que se sorprenda al descubrir que algunas habilidades en realidad no son mejores que ninguna habilidad, y ahí es donde comienza la optimización.
Lectura relacionada:
* [Qué son las Habilidades de Claude](/es/docs/notes/claude-skills/concept) — Comprender los principios básicos de las Habilidades
* [Guía Práctica](/es/docs/notes/claude-skills/practice) — Crea tu primera Skill
# Introducción al concepto
## Introducción
En el [Análisis en Profundidad de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept), exploramos un problema fundamental: **Context Rot** — a medida que las conversaciones se alargan, la ventana de contexto de Claude se llena de código fallido, discusiones obsoletas e información irrelevante, provocando que la calidad de las respuestas se degrade progresivamente.
La solución de Ralph fue "reiniciar todo": usar un bucle infinito de bash para iniciar una instancia nueva de Claude cada vez, pasando el estado a través del sistema de archivos. Simple, efectivo, pero con limitaciones evidentes — es solo una metodología sin comprensión del proyecto, sin planificación por fases ni verificación de calidad. Usted necesita escribir las especificaciones, orquestar las tareas y juzgar "¿está terminado?" por su cuenta.
Como Chase AI resumió con precisión en su video: **El Ralph Loop es un arma increíblemente poderosa, pero la mayoría de las personas no necesitan un arma — necesitan un arsenal completo.** El bucle de Ralph depende enteramente de la preparación previa: ¿Su PRD es suficientemente bueno? ¿Las definiciones de funcionalidades son lo bastante precisas? ¿Sabe cómo se ve "terminado"? Si las respuestas a estas preguntas no son precisas, sin importar cuántas veces se ejecute el bucle, será simplemente basura entra, basura sale.
¿Y si usted quiere un sistema que **no solo ejecute Claude en un bucle, sino que realmente comprenda su proyecto y entregue código de manera confiable**?
Eso es exactamente lo que **GSD (Get Shit Done)** está diseñado para hacer.
## ¿Qué es GSD?
GSD fue creado por **TÂCHES** (GitHub: glittercowboy), un desarrollador independiente. Su motivación fue directa:
> "No soy una empresa de software de 50 personas. No quiero jugar al teatro empresarial. Soy simplemente una persona creativa que quiere hacer cosas geniales."
En su transmisión en vivo, TÂCHES demostró un hecho impactante: **nunca escribe código a mano**. Usando GSD, construyó una aplicación nativa de macOS para generación de música con IA (Sample Digger) desde cero en 4 horas — cero código escrito a mano. Se posiciona no como programador, sino como "gerente de proyecto de alto nivel" — describiendo la visión, tomando decisiones clave y validando resultados. GSD hace posible esta forma de trabajar.
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> —— TÂCHES
Otras herramientas de desarrollo dirigido por especificaciones — BMAD, SpecKit — tienen su propio valor, pero tienden a introducir flujos de trabajo empresariales complejos: ceremonias de sprint, story points, sincronizaciones con stakeholders. Para desarrolladores independientes o equipos pequeños, estos procesos son una carga en sí mismos. Como dijo Chase AI: "It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way."
La filosofía de diseño de GSD es **ocultar la complejidad dentro del sistema**. Los usuarios solo necesitan unos pocos comandos simples mientras el sistema se encarga de toda la gestión de contexto, la orquestación de tareas y la verificación de calidad tras bambalinas. En el primer mes desde su lanzamiento, el proyecto obtuvo cerca de 3,000 estrellas en GitHub y 14,000 instalaciones de npm, con TÂCHES publicando actualizaciones 15-20 veces casi todos los días.
### Posición de GSD en el ecosistema de herramientas
| Dimensión | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| ------------------------------ | --------------------------------- | ------------------------------------- | ------------------------------ | --------------------------------------------- |
| Posicionamiento central | Técnica de ejecución (bucle bash) | Kit de generación de especificaciones | Framework empresarial | **Ingeniería de contexto + especificaciones** |
| Capacidad de planificación | Ninguna (traiga su propia spec) | Fuerte (spec→plan→tareas) | Fuerte (proceso ágil completo) | **Fuerte (investigación→discusión→plan)** |
| Autonomía de ejecución | Máxima (modo AFK) | Activación manual por paso | Activación manual por paso | **Activación manual por paso** |
| Modelo de participación humana | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Manejo de Context Rot | Reinicio de nueva sesión | Sin solución integrada | Sin solución integrada | **Contexto fresco de subagentes** |
| Verificación de calidad | Depende de pruebas externas | Verificaciones de compilación | Proceso QA integrado | **Verificación automática + UAT** |
| Complejidad para el usuario | Mínima | Media | Alta | **Baja** |
| Complejidad del sistema | Mínima | Media | Alta | **Alta** |
Esta tabla revela una compensación clave: **Ralph intercambia complejidad mínima del sistema por máxima autonomía de ejecución** — inícielo y váyase a dormir; mientras que **GSD intercambia alta complejidad del sistema por calidad de planificación y supervisión humana** — usted tiene la oportunidad de intervenir en cada etapa. SpecKit y BMAD se ubican en un punto intermedio, ofreciendo capacidades de planificación pero sin la ingeniería de contexto de GSD ni la ejecución autónoma de Ralph.
GSD y Ralph no son contradictorios. GSD hereda los principios fundamentales de Ralph — contexto fresco, archivos como fuente de verdad — pero construye un marco completo de comprensión y ejecución de proyectos sobre esa base. Si Ralph es "darle una tarea a la IA y dejar que siga intentando", GSD es "entender lo que usted quiere, investigar cómo hacerlo, planificar los pasos, ejecutar y verificar".
El resumen de Chase AI lo captura perfectamente: **El bucle de Ralph asume que usted llega con un plano completo — GSD le ayuda a construir ese plano.** GSD toma su idea a medio formar, hace preguntas profundas, investiga en su nombre, genera un PRD completo, lo descompone en tareas atómicas y entrega el proyecto de principio a fin. Y al ejecutar código, utiliza exactamente los mismos principios fundamentales que hacen poderoso al bucle de Ralph: contexto fresco para los subagentes y tareas lo más pequeñas y precisas posible.
## Flujo de trabajo principal
El flujo de trabajo de GSD es un ciclo de **discusión → planificación → ejecución → verificación**, donde cada etapa tiene entradas y salidas claramente definidas.
### 1. Inicializar el proyecto
```text
/gsd:new-project
```
Un solo comando inicia todo el proceso. El sistema hará lo siguiente:
1. **Hacer preguntas** — Seguirá indagando hasta comprender completamente su idea (objetivos, restricciones, preferencias tecnológicas, casos límite)
2. **Investigar** — Enviar agentes en paralelo para investigar los dominios relevantes (opcional pero recomendado)
3. **Extraer requisitos** — Distinguir entre elementos de v1, v2 y fuera de alcance
4. **Hoja de ruta** — Crear un plan por fases alineado con los requisitos
Usted aprueba la hoja de ruta y luego comienza a construir. La experiencia de TÂCHES indica: cuanto más detallada sea la descripción inicial que proporcione, menos preguntas de seguimiento hará el sistema; cuanto más vaga sea, más preguntará. Recomienda preparar un documento de visión aproximado antes de comenzar — no necesita conocer el stack tecnológico ni los detalles de implementación, solo describir lo que quiere.
**Archivos de salida**: `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md`
> ¿Ya tiene una base de código? Ejecute `/gsd:map-codebase` primero — el sistema enviará agentes en paralelo para analizar su stack tecnológico, arquitectura, convenciones y problemas potenciales. Luego `/gsd:new-project` podrá planificar basándose en su código existente.
### 2. Fase de discusión
```text
/gsd:discuss-phase 1
```
Cada fase en la hoja de ruta tiene solo una o dos oraciones de descripción — eso no es suficiente para construir lo que usted quiere. La fase de discusión existe para **capturar sus preferencias de implementación** antes de la investigación y planificación.
El sistema analiza la fase actual e identifica "zonas grises" — puntos de decisión donde existen múltiples enfoques razonables de implementación:
* Funcionalidades visuales → diseño, interacciones, manejo de estados vacíos
* API/CLI → formato de respuesta, manejo de errores, nivel de detalle
* Sistemas de contenido → estructura, tono, profundidad, flujo
Cada decisión que tome aquí afecta directamente la calidad de la investigación y planificación posteriores. Omitir este paso es válido (el sistema usará valores predeterminados razonables), pero una discusión más profunda permite al sistema construir algo más cercano a sus expectativas.
**Archivos de salida**: `{phase}-CONTEXT.md`
### 3. Fase de planificación
```text
/gsd:plan-phase 1
```
El sistema hará lo siguiente:
1. **Investigar** — Examinar cómo implementar la fase actual, guiado por las decisiones de la fase de discusión
2. **Planificar** — Crear 2-3 planes de tareas atómicas usando formato estructurado XML
3. **Verificar** — Comprobar si el plan satisface los requisitos, iterando hasta que pase
Un principio de diseño clave es el **Goal-Backward Planning** (planificación retroactiva desde el objetivo). En lugar de partir de "¿qué deberíamos construir?", pregunta "¿qué condiciones deben cumplirse para lograr el objetivo?" — y luego trabaja hacia atrás para derivar el plan y las tareas. TÂCHES afirma que este enfoque "mejoró masivamente la calidad de los resultados" porque cada tarea entiende su relación con las demás, en lugar de ser simplemente un elemento en una lista de pendientes.
Cada plan es lo suficientemente pequeño para ejecutarse dentro de una única ventana de contexto nueva. Esto es crucial — **no habrá degradación de calidad**.
**Archivos de salida**: `{phase}-RESEARCH.md`, `{phase}-{N}-PLAN.md`
### 4. Fase de ejecución
```text
/gsd:execute-phase 1
```
El sistema hará lo siguiente:
1. **Ejecución por oleadas** — Las tareas independientes se ejecutan en paralelo; las dependientes se ejecutan secuencialmente
2. **Contexto fresco** — Cada plan se ejecuta en un contexto completamente nuevo de 200k tokens, sin basura acumulada
3. **Commits atómicos** — Cada tarea obtiene un commit de git independiente
4. **Verificación de objetivos** — Comprobar si la base de código entrega la funcionalidad prometida por la fase
En la demostración en vivo de TÂCHES, completó 3 fases completas de desarrollo con su **ventana de contexto principal manteniéndose en apenas el 24%**. El subagente GSD Executor necesita cargar menos de 1,000 líneas de contexto para completar una fase entera — puede ejecutar 10 planes seguidos y el contexto aún se mantiene por debajo del 50%. Esta es una experiencia completamente diferente a trabajar directamente en Claude Code: ya no es "jugar a la ruleta rusa, apostando a cuándo chocarás contra el muro de la ventana de contexto".
**Archivos de salida**: `{phase}-{N}-SUMMARY.md`, `{phase}-VERIFICATION.md`
### 5. Fase de verificación
```text
/gsd:verify-work 1
```
La verificación automatizada puede comprobar si el código existe y si las pruebas pasan. Pero, ¿la funcionalidad **trabaja como usted espera**? Eso requiere su confirmación.
El sistema hará lo siguiente:
1. **Extraer entregables verificables** — Listar las cosas que ahora debería poder hacer
2. **Guiar la verificación una por una** — "¿Puede iniciar sesión con correo electrónico?" Sí/no, o describa el problema
3. **Diagnosticar fallos automáticamente** — Enviar un agente de depuración para encontrar la causa raíz
4. **Crear un plan de corrección** — Una corrección directamente ejecutable
Si todo pasa, continúe a la siguiente fase. Si hay problemas, ejecute `/gsd:execute-phase` nuevamente para ejecutar el plan de corrección.
Esta es la mayor diferencia filosófica entre GSD y el bucle de Ralph: **Ralph es de manos libres — inícielo y déjelo correr; GSD tiene un paso de verificación humana después de cada fase.** Chase AI señala que el bucle de Ralph es estilo "ir a conquistar" — funciona por su cuenta sin mirar atrás; GSD asegura que usted pueda corregir el rumbo en cada punto de control crítico, evitando que los errores se acumulen sin supervisión.
Además, GSD proporciona un flujo de trabajo dedicado para depuración. Cuando la verificación encuentra problemas, `/gsd:debug` lanza un **subagente de depuración aislado** con su propio flujo de trabajo de hipótesis-evidencia-resolución, creando documentación de depuración independiente para rastrear todo el proceso de investigación sin contaminar el contexto principal.
**Archivos de salida**: `{phase}-UAT.md`
### Repetir el ciclo
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
Cada fase pasa por el ciclo completo de **discusión → planificación → ejecución → verificación**. El contexto se mantiene fresco, la calidad se mantiene consistente.
Cuando todas las fases se completan, `/gsd:complete-milestone` archiva el hito y etiqueta la versión. Luego `/gsd:new-milestone` inicia la construcción de la siguiente versión.
## Por qué funciona: principios técnicos
La confiabilidad de GSD no es accidental — cuatro pilares técnicos clave la sustentan.
### Context Engineering
Claude Code es extremadamente poderoso cuando recibe el contexto adecuado. La mayoría de las personas no saben cómo darle el contexto adecuado. GSD se encarga de esto por usted.
| Archivo | Propósito |
| ----------------- | --------------------------------------------------------------------------------------- |
| `PROJECT.md` | Visión del proyecto, siempre cargado |
| `research/` | Conocimiento del ecosistema (stack tecnológico, funcionalidades, arquitectura, trampas) |
| `REQUIREMENTS.md` | Requisitos versionados con trazabilidad por fases |
| `ROADMAP.md` | Dirección y progreso |
| `STATE.md` | Decisiones, bloqueadores, posición — memoria entre sesiones |
| `PLAN.md` | Tareas atómicas + estructura XML + pasos de verificación |
| `SUMMARY.md` | Registros de ejecución, comprometidos al historial |
Cada archivo tiene **límites de tamaño** basados en los umbrales de degradación de calidad de Claude. Mantenerse dentro de los límites garantiza resultados consistentemente de alta calidad. La ventana de contexto principal se mantiene en 30-40%, mientras que el trabajo real ocurre en los contextos frescos de 200k de los subagentes.
Chase AI tiene una explicación intuitiva del context rot: **Sin importar cuán grande sea la ventana de contexto — Sonnet, Opus, incluso ventanas de un millón de tokens — los tokens en la primera mitad son más efectivos que los de la segunda mitad.** Esto no es un error; es una propiedad inherente de los LLM. El autocompact integrado de Claude Code solo puede mitigar esto parcialmente. El enfoque de GSD es más exhaustivo: cada tarea atómica se ejecuta en un subagente nuevo, asegurando que cada tarea obtenga el mejor rendimiento de Claude.
Los propios datos de TÂCHES lo confirman: con el plan Max de $200/mes, consume aproximadamente $30,000 en tokens de Opus al mes. Eso suena como mucho, pero dado que cada tarea se ejecuta en contexto fresco, el retrabajo es mínimo — la eficiencia real es muy superior a reparar cosas repetidamente en un contexto degradado.
### XML Prompt Formatting
Cada plan es XML estructurado optimizado para Claude:
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
Instrucciones precisas, sin adivinanzas, verificación integrada en cada tarea.
### Orquestación multiagente
Cada etapa utiliza el mismo patrón: un orquestador ligero envía agentes especializados, recopila resultados y encamina al siguiente paso.
| Etapa | Lo que hace el orquestador | Lo que hacen los agentes |
| ------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------ |
| Investigación | Coordina, presenta hallazgos | 4 investigadores en paralelo analizan stack tecnológico, funcionalidades, arquitectura y trampas |
| Planificación | Valida, gestiona iteraciones | El planificador crea planes, el verificador valida, el ciclo continúa hasta que pase |
| Ejecución | Agrupa en oleadas, rastrea progreso | Los ejecutores implementan en paralelo, cada uno con un contexto fresco de 200k |
| Verificación | Presenta resultados, encamina siguiente paso | El verificador revisa la base de código, el depurador diagnostica fallos |
El orquestador nunca hace el trabajo pesado. Envía agentes, espera e integra resultados. El resultado: puede ejecutar una fase completa — investigación profunda, creación y validación de múltiples planes, miles de líneas de código escritas en paralelo, verificación automatizada — **mientras su ventana de contexto principal se mantiene en 30-40%**.
### Commits atómicos de Git
Cada tarea se registra de forma independiente inmediatamente después de completarse:
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
Beneficios: `git bisect` puede localizar la tarea exacta que falló, cada tarea se puede revertir de forma independiente, y un historial limpio ayuda a Claude a entender la evolución del código en sesiones futuras.
## Limitaciones de GSD
GSD es poderoso, pero entender lo que **no puede hacer** es igualmente importante.
### GSD es un flujo de trabajo guiado por humanos, no un agente autónomo
GSD no puede ejecutarse de manera persistente. Cada frontera entre etapas — de `discuss` a `plan` a `execute` a `verify` — requiere que usted ingrese un comando manualmente. No puede decir "constrúyame una aplicación" e irse a dormir.
Esto contrasta marcadamente con el modo AFK de Ralph. Ralph está diseñado para "iniciarlo e irse a dormir" — el bucle infinito de bash sigue ejecutándose hasta que la tarea se complete o falle. GSD requiere que usted esté presente en cada punto de control crítico: aprobando la hoja de ruta, respondiendo preguntas de discusión, activando la planificación, lanzando la ejecución, confirmando los resultados de verificación.
Durante su transmisión en vivo de 4 horas, TÂCHES estuvo continuamente escribiendo comandos: `new-project`, `discuss-phase 1`, `plan-phase 1`, `execute-phase 1`, `verify-work 1`, `discuss-phase 2`... Cada transición requería que presionara Enter. Esto no es accidental — es una decisión de diseño deliberada.
### Una compensación de diseño deliberada
Ralph sacrificó capacidad de planificación por autonomía de ejecución; GSD sacrificó autonomía de ejecución por calidad de planificación y supervisión humana. **Esta es una compensación de diseño, no una deficiencia.**
* **Ventaja de Ralph**: Puede dejarlo ejecutar una funcionalidad completa mientras duerme. Pero si la especificación no es suficientemente buena, avanzará a toda velocidad en la dirección equivocada.
* **Ventaja de GSD**: Puede corregir el rumbo después de cada fase. Pero debe estar presente en todo momento — no puede alejarse.
¿Cómo sería lo ideal? Si las fases de discusión, planificación, ejecución y verificación de GSD pudieran encadenarse en un ciclo automatizado — como el bucle bash de Ralph pero con la planificación estructurada y la verificación de calidad de GSD — eso sería lo mejor de ambos mundos. Pero aún no existe tal herramienta. Quizás esa sea la próxima dirección que vale la pena explorar.
## Recursos en video
Los siguientes videos pueden ayudarle a comprender de manera más intuitiva cómo se usa GSD y qué puede lograr.
## Reflexiones finales
GSD representa una dirección en la evolución de las herramientas de programación con IA: de "dejar que la IA escriba código" a "dejar que la IA entregue proyectos de manera confiable".
Ralph Wiggum demostró una perspectiva clave — el contexto fresco es más valioso que el contexto acumulado. GSD construye sobre esta base agregando comprensión del proyecto (new-project), captura de decisiones (discuss), planificación estructurada (plan), ejecución en paralelo (execute) y verificación de calidad (verify), formando un ciclo cerrado completo.
Para desarrolladores independientes y equipos pequeños, el valor de GSD radica en empaquetar prácticas complejas de ingeniería en unos pocos comandos simples. No necesita entender la orquestación de subagentes ni la ingeniería de prompts XML — solo necesita describir lo que quiere y dejar que el sistema se encargue.
Chase AI lo expresó bien: GSD es para personas que "no vienen de un trasfondo técnico pero aún quieren construir proyectos de principio a fin en Claude Code de manera sostenible y repetible". Y la transmisión en vivo de TÂCHES lo demostró — alguien que se describe a sí mismo como "probablemente solo capaz de escribir una página HTML de Hello World por mi cuenta" usó GSD para construir una aplicación de escritorio nativa completa.
Esto no es magia. Es **poner la complejidad correcta en el lugar correcto** — el sistema maneja la complejidad de la orquestación mientras los humanos se concentran en la creatividad y las decisiones. Y sus limitaciones también merecen respeto: la decisión de GSD de mantener a los humanos presentes en todo momento es tanto su restricción como la fuente de su confiabilidad.
¿Listo para ponerse manos a la obra? Continúe con la [Guía Práctica de GSD](/es/docs/notes/gsd/practice) — que cubre la referencia completa de comandos, detalles de configuración, demostraciones de flujos de trabajo y preguntas frecuentes.
***
**Lecturas relacionadas**:
* [Análisis en Profundidad de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) — Análisis completo del problema de Context Rot y la metodología Ralph
* [¿Qué es el Desarrollo Dirigido por Especificaciones?](/es/docs/notes/speckit/concept) — El cambio de paradigma de Vibe Coding al desarrollo dirigido por especificaciones
* [Guía Completa de Claude Subagent](/es/docs/notes/claude-subagent) — Otro enfoque para mantener limpio el contexto
* [Arquitectura del Sistema Claude](/es/docs/notes/claude-architecture) — La arquitectura general de Hooks, Subagents y otros componentes
* [Mis Mejores Prácticas con Claude Code](/es/blog/claude-code-best-practices) — Consejos prácticos para el uso diario de Claude Code
# Guía práctica
## Introducción
En el [artículo anterior](/es/docs/notes/gsd/concept), exploramos los principios fundamentales de GSD: ingeniería de contexto, orquestación de subagentes, planificación por objetivos inversos y commits atómicos. Estos conceptos suenan elegantes, pero entre "comprender la teoría" y "entregar un proyecto real" hay muchos detalles operativos que resolver.
En este artículo, pasamos a la práctica. Aprenderá el sistema completo de comandos de GSD, las opciones de configuración, la estructura de archivos generados y cómo utilizarlo para entregar una funcionalidad completa desde cero.
## Instalación y configuración
### Instalación
```bash
npx get-shit-done-cc@latest
```
El instalador le pedirá que seleccione:
1. **Entorno de ejecución** — Claude Code, OpenCode, Gemini CLI o todos
2. **Alcance** — Global (todos los proyectos) o local (proyecto actual)
Después de la instalación, escriba `/gsd:help` en su entorno de ejecución para verificar que la instalación fue exitosa.
### Recomendado: modo sin permisos
GSD está diseñado para automatización sin fricciones. La forma recomendada de ejecutar Claude Code:
```bash
claude --dangerously-skip-permissions
```
Si prefiere no usar esta opción, puede configurar permisos detallados en `.claude/settings.json`.
### Actualización
```text
/gsd:update
```
Las actualizaciones de GSD son muy frecuentes (TÂCHES publica entre 15 y 20 actualizaciones casi todos los días). Se recomienda ejecutar este comando regularmente para mantenerse en la última versión.
## Referencia completa de comandos
Todas las interacciones con GSD se realizan mediante comandos de barra (slash commands) con el prefijo `/gsd:`. A continuación se presenta la referencia completa organizada por función.
### Comandos del flujo de trabajo principal
Estos cinco comandos forman el ciclo principal de GSD y se usan en secuencia.
| Comando | Descripción |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:new-project` | Inicializa un proyecto. El sistema hace preguntas hasta comprender su idea, luego investiga, extrae requisitos y crea una hoja de ruta |
| `/gsd:discuss-phase [N]` | Discute las zonas grises de la fase N. Captura sus preferencias de implementación para orientar la planificación |
| `/gsd:plan-phase [N]` | Crea planes de tareas atómicas para la fase N. Incluye subpasos de investigación, planificación y verificación |
| `/gsd:execute-phase ` | Ejecuta la fase N. Los subagentes implementan tareas en paralelo, cada una con un commit independiente |
| `/gsd:verify-work [N]` | Verifica los entregables de la fase N. Le guía en la confirmación uno por uno y diagnostica problemas automáticamente |
> `[N]` indica un parámetro opcional — el sistema detecta automáticamente la fase actual si se omite. `` indica un parámetro obligatorio.
### Gestión de hitos
| Comando | Descripción |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `/gsd:audit-milestone` | Audita el progreso del hito actual — verifica el estado de todas las fases e identifica elementos incompletos |
| `/gsd:complete-milestone` | Archiva el hito actual, etiqueta la versión y prepara para el siguiente ciclo |
| `/gsd:new-milestone [name]` | Crea un nuevo hito. Opcionalmente puede proporcionar un nombre; el sistema planifica basándose en el trabajo completado |
### Gestión de fases
| Comando | Descripción |
| --------------------------------- | --------------------------------------------------------------------------------------------------- |
| `/gsd:add-phase` | Agrega una nueva fase al final de la hoja de ruta |
| `/gsd:insert-phase [N]` | Inserta una fase urgente en la posición N; las fases posteriores se renumeran automáticamente |
| `/gsd:remove-phase [N]` | Elimina una fase y borra en cascada todos los archivos de salida relacionados |
| `/gsd:list-phase-assumptions [N]` | Lista todas las suposiciones y dependencias de una fase, ayudando a identificar riesgos potenciales |
### Quick Mode y herramientas
| Comando | Descripción |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:quick [--full]` | Modo rápido — omite investigación, verificación de plan y validación. Ideal para tareas pequeñas. `--full` activa todas las protecciones |
| `/gsd:debug [desc]` | Lanza un subagente de depuración aislado. Opcionalmente describe el problema; el sistema formula hipótesis, recopila evidencia y resuelve |
| `/gsd:add-todo [desc]` | Registra una idea en la lista de pendientes sin modificar la hoja de ruta |
| `/gsd:check-todos` | Muestra la lista de pendientes actual |
| `/gsd:map-codebase` | Analiza una base de código existente — stack tecnológico, arquitectura, convenciones y problemas potenciales |
### Gestión de sesión y configuración
| Comando | Descripción |
| ------------------ | --------------------------------------------------------------------------------------------------- |
| `/gsd:pause-work` | Pausa el trabajo. Guarda el estado actual en STATE.md para facilitar la reanudación |
| `/gsd:resume-work` | Reanuda el trabajo. Lee el último estado desde STATE.md y continúa donde lo dejó |
| `/gsd:progress` | Muestra el progreso general del proyecto — fases completadas, posición actual, elementos pendientes |
| `/gsd:help` | Muestra todos los comandos disponibles con descripciones breves |
| `/gsd:settings` | Consulta y modifica la configuración de GSD |
| `/gsd:set-profile` | Cambia el perfil del modelo (quality / balanced / budget) |
| `/gsd:update` | Actualiza GSD a la última versión |
## Configuración detallada
### Perfiles de modelo
GSD soporta tres perfiles de modelo, que se cambian con `/gsd:set-profile`:
| Perfil | Planificación | Ejecución | Verificación | Caso de uso |
| ------------------------- | ------------- | --------- | ------------ | ----------------------------------------------------------------- |
| quality | Opus | Opus | Sonnet | Proyectos complejos, funcionalidades críticas, primer uso |
| balanced (predeterminado) | Opus | Sonnet | Sonnet | Desarrollo diario, mejor equilibrio para la mayoría de escenarios |
| budget | Sonnet | Sonnet | Haiku | Funcionalidades simples, presupuesto limitado, iteración rápida |
### Configuración principal
Consulte y modifique mediante `/gsd:settings`:
| Configuración | Valor predeterminado | Descripción |
| ------------------------ | -------------------- | ------------------------------------------------------------------------------------------ |
| `mode` | `balanced` | Selección de perfil de modelo |
| `depth` | `standard` | Profundidad de investigación: `quick` (rápida) / `standard` (estándar) / `deep` (profunda) |
| `git.branching_strategy` | `feature` | Estrategia de ramas Git: `feature` (por funcionalidad) / `phase` (por fase) / `none` |
### Interruptores de flujo de trabajo
Los siguientes agentes pueden activarse o desactivarse individualmente para equilibrar velocidad y calidad:
| Interruptor | Predeterminado | Descripción |
| -------------- | -------------- | ---------------------------------------------------------------- |
| `research` | Activado | Si se realiza investigación automática antes de la planificación |
| `plan_check` | Activado | Si se verifica automáticamente después de crear el plan |
| `verifier` | Activado | Si se verifica automáticamente después de la ejecución |
| `auto_advance` | Desactivado | Si se avanza automáticamente a la siguiente fase al completar |
> Desactivar `research` y `plan_check` puede acelerar significativamente el proceso, pero podría reducir la calidad de la planificación. Se recomienda considerar desactivarlos solo después de familiarizarse con el proyecto.
## Estructura de archivos generados
Todo el estado y las salidas de GSD se almacenan en el directorio `.planning/`. Comprender esta estructura ayuda con la depuración y la intervención manual.
### Archivos a nivel de proyecto
| Archivo | Propósito | Momento de creación |
| ----------------- | ------------------------------------------------------------------ | ------------------------------------- |
| `PROJECT.md` | Visión y alcance del proyecto | `new-project` |
| `REQUIREMENTS.md` | Documentación de requisitos versionada, con trazabilidad por fases | `new-project` |
| `ROADMAP.md` | Planificación de fases y progreso | `new-project` |
| `STATE.md` | Estado actual — decisiones, bloqueos, posición | `new-project`, actualización continua |
### Archivos a nivel de fase
Cada fase produce los siguientes archivos (usando la Fase 1 como ejemplo):
| Archivo | Propósito | Momento de creación |
| -------------------- | -------------------------------------------------- | ------------------- |
| `01-CONTEXT.md` | Registro de decisiones de la fase de discusión | `discuss-phase 1` |
| `01-RESEARCH.md` | Hallazgos de investigación e investigación técnica | `plan-phase 1` |
| `01-01-PLAN.md` | Plan de la primera tarea atómica | `plan-phase 1` |
| `01-02-PLAN.md` | Plan de la segunda tarea atómica | `plan-phase 1` |
| `01-01-SUMMARY.md` | Registro de ejecución del primer plan | `execute-phase 1` |
| `01-02-SUMMARY.md` | Registro de ejecución del segundo plan | `execute-phase 1` |
| `01-VERIFICATION.md` | Resultados de verificación automática | `execute-phase 1` |
| `01-UAT.md` | Registro de pruebas de aceptación del usuario | `verify-work 1` |
### Ejemplo de estructura de directorios
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## Demostración del flujo de trabajo
A continuación, se demuestra el flujo completo desde la inicialización hasta la entrega, usando como ejemplo "agregar una funcionalidad de comentarios a un sistema de blog".
### Paso 1: Inicializar el proyecto
```text
/gsd:new-project
```
El sistema comenzará a hacer preguntas:
```
> ¿Qué desea construir?
"Quiero agregar un sistema de comentarios a mi blog en Next.js. Que soporte
comentarios anónimos y con inicio de sesión, renderizado de Markdown y un
panel de administración. Stack tecnológico: Prisma + PostgreSQL."
```
Cuanto más detallada sea su descripción, menos preguntas de seguimiento hará el sistema. TÂCHES recomienda preparar un documento de visión aproximado describiendo lo que desea, sin necesidad de conocer los detalles técnicos.
Una vez completado, el sistema genera cuatro archivos y le solicita aprobar la hoja de ruta. Una vez aprobada, comienza la fase de construcción.
> **¿Ya tiene una base de código?** Ejecute primero `/gsd:map-codebase`. El sistema analizará su arquitectura y convenciones existentes, y luego `new-project` podrá planificar basándose en su código actual.
### Paso 2: Fase de discusión
```text
/gsd:discuss-phase 1
```
El sistema identifica las zonas grises y formula preguntas una por una:
```
> Anidación de comentarios: ¿soportar múltiples niveles o solo un nivel de respuesta?
> Comentarios anónimos: ¿requiere CAPTCHA o envío directo?
> Panel de administración: ¿necesita operaciones masivas o moderación individual?
```
Cada decisión que tome aquí afecta directamente la calidad de la planificación. Si no está seguro, puede dejar que el sistema use valores predeterminados, pero una discusión más profunda reduce significativamente el retrabajo durante la ejecución.
### Paso 3: Fase de planificación
```text
/gsd:plan-phase 1
```
El sistema:
1. Investiga cómo implementar un sistema de comentarios con Prisma + PostgreSQL
2. Crea 2-3 planes de tareas atómicas (por ejemplo: modelo de datos, rutas API, componentes del frontend)
3. Verifica automáticamente que los planes cubran todos los requisitos
Cada plan es lo suficientemente pequeño para completarse en una sola ventana de contexto nueva.
### Paso 4: Fase de ejecución
```text
/gsd:execute-phase 1
```
El sistema comienza la ejecución por oleadas:
* **Oleada 1** (sin dependencias): Esquema de base de datos, modelos Prisma — ejecución en paralelo
* **Oleada 2** (depende de la Oleada 1): Rutas API, CRUD de comentarios — ejecución en paralelo
* **Oleada 3** (depende de la Oleada 2): Componente de comentarios del frontend — ejecución independiente
Cada tarea se ejecuta en un contexto nuevo de 200k tokens y recibe un commit de git independiente.
### Paso 5: Fase de verificación
```text
/gsd:verify-work 1
```
El sistema le guía en la confirmación:
```
> ✅ Tablas de base de datos creadas
> ✅ Las rutas API devuelven códigos de estado correctos
> ❓ ¿Puede ver el campo de entrada de comentarios debajo de las publicaciones? [sí/no/describir problema]
> ❓ ¿La página se actualiza en tiempo real después de enviar un comentario? [sí/no/describir problema]
```
Si alguna verificación falla, el sistema diagnostica automáticamente y crea un plan de corrección. Ejecute `/gsd:execute-phase 1` nuevamente para aplicar la corrección.
### Escenarios comunes
**Insertar una fase urgente**: Los requisitos cambiaron y necesita insertar nuevo trabajo antes de la fase actual.
```text
/gsd:insert-phase 2
```
Las fases posteriores se renumeran automáticamente (la Fase 2 original se convierte en Fase 3, y así sucesivamente).
**Pausar y reanudar**: Necesita interrumpir el trabajo para atender otro asunto.
```text
/gsd:pause-work # Guardar estado actual
# ... atender otro asunto ...
/gsd:resume-work # Reanudar desde donde lo dejó
```
**Revertir resultados insatisfactorios**:
```bash
git reset --hard HEAD~3 # Volver al estado previo a la ejecución
```
```text
/gsd:remove-phase 2 # Eliminar en cascada todos los archivos de salida de esta fase
```
TÂCHES demostró esta operación múltiples veces en sus transmisiones en vivo: si no le gusta, revierta. Limpio y decisivo.
## Flujo de depuración
Cuando la verificación encuentra problemas, o cuando se encuentra con errores durante el desarrollo, GSD proporciona un flujo de depuración dedicado.
```text
/gsd:debug La página no se actualiza en tiempo real después de enviar un comentario
```
El sistema lanza un **subagente de depuración aislado** con el siguiente flujo de trabajo:
1. **Hipótesis** — Genera múltiples hipótesis de causas raíz basadas en la descripción del problema
2. **Recopilación de evidencia** — Verifica las hipótesis una por una, revisando código, registros y solicitudes de red
3. **Resolución** — Después de identificar la causa raíz, crea un plan de corrección
Características clave:
* **Aislamiento de contexto**: El agente de depuración tiene su propia ventana de contexto y no contamina su contexto principal de desarrollo
* **Documentación**: Crea documentos de depuración independientes que registran todo el proceso de investigación
* **Plan de corrección**: Produce un plan de corrección directamente ejecutable después del diagnóstico
Esto es mucho más eficiente que depurar en el contexto principal: la información de depuración no se acumula en su ventana principal.
## Experiencia práctica
Basándonos en las transmisiones en vivo de TÂCHES y la experiencia práctica de Chase AI, a continuación se presentan algunas recomendaciones prácticas.
### Ir más lento para ir más rápido
TÂCHES reconoce que cuando comenzó a usar GSD, su mentalidad era "rápido, rápido, rápido", pero luego descubrió que **dedicar más tiempo a las fases de investigación y discusión en realidad reducía el retrabajo durante la ejecución**. Las versiones más recientes de GSD agregaron los pasos `research-project` y `define-requirements` precisamente para asegurar la dirección correcta antes de escribir código.
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### Limpiar el contexto entre fases
El hábito de TÂCHES es **ejecutar `clear` entre cada fase** para mantener el contexto principal ligero. Utiliza el terminal Warp, con cada ventana a pantalla completa (Command+Shift+Enter), ejecutando la fase actual en una ventana mientras investiga la siguiente fase en otra.
### La compensación del costo de tokens
El enfoque de subagentes de GSD efectivamente consume más tokens que usar Claude Code directamente. Pero Chase AI presenta un argumento convincente: **"plan twice, prompt once" (planificar dos veces, ejecutar una vez) es más económico a largo plazo que "ejecutar una vez y luego parchear y parchear".** Hacer las cosas bien en un contexto nuevo es mucho más eficiente que reparar repetidamente en un contexto degradado.
### Manejo de resultados insatisfactorios
Si no está satisfecho con los resultados de una fase, puede usar `git reset --hard` y luego `/gsd:remove-phase` para eliminar en cascada todos los archivos de salida de esa fase. TÂCHES demostró esto en vivo: no le gustó un efecto visual particular, así que revirtió al último estado satisfactorio, limpio y decisivo.
### El sistema de pendientes
`/gsd:add-todo` le permite registrar ideas en una lista de pendientes en cualquier momento sin modificar la hoja de ruta. Estas ideas pueden recuperarse durante `/gsd:discuss-milestone` como entrada para el siguiente hito. La estrategia de TÂCHES es "primero construir las funcionalidades, pulir la interfaz en el hito 2".
## Preguntas frecuentes y mejores prácticas
### Mejores prácticas
**Proporcione descripciones iniciales detalladas.** La calidad de `/gsd:new-project` depende de la calidad de su entrada. Prepare un documento de visión aproximado: describa objetivos, usuarios, funcionalidades principales y restricciones conocidas. Cuanto más precisa sea su descripción, menos preguntas de seguimiento y mejor la planificación.
**Limpie el contexto entre fases.** Después de completar cada fase, ejecute `clear` o `/compact` para mantener la ventana de contexto principal ligera. El hábito de TÂCHES es mantener el contexto principal entre el 30-40%.
**Pruebe primero con Quick Mode.** Para funcionalidades pequeñas de las que no esté seguro, use `/gsd:quick` para probar. Si funciona bien, incorpórelo en la hoja de ruta formal.
**Mapee primero las bases de código existentes.** Antes de usar GSD en una base de código existente, ejecute `/gsd:map-codebase`. El sistema analizará el stack tecnológico, la arquitectura y las convenciones, lo que hará que la planificación posterior se alinee mejor con el código existente.
### Preguntas frecuentes
**P: ¿Qué entornos de ejecución soporta GSD?**
R: Claude Code, OpenCode y Gemini CLI. Puede elegir uno o todos durante la instalación.
**P: ¿Cuál es la diferencia entre Quick Mode y el modo completo?**
R: Quick Mode proporciona las protecciones básicas de GSD (commits atómicos, seguimiento de estado), pero omite los pasos de investigación, verificación del plan y validación. Es ideal para correcciones de errores, funcionalidades pequeñas y cambios de configuración que no necesitan una planificación completa.
**P: ¿Se puede pausar durante la ejecución?**
R: Sí. `/gsd:pause-work` guarda el estado actual en STATE.md. La próxima vez que ejecute `/gsd:resume-work`, el sistema continuará desde donde lo dejó.
**P: ¿Cómo se controla el costo de tokens?**
R: Tres enfoques: (1) Cambiar al perfil `budget`: `/gsd:set-profile budget`; (2) Desactivar los agentes `research` o `plan_check`; (3) Usar `/gsd:quick` para tareas simples.
**P: ¿Se puede usar GSD junto con Ralph?**
R: Sí. GSD y Ralph resuelven problemas diferentes: GSD se encarga de la planificación y la ejecución estructurada, Ralph se encarga de la ejecución en ciclo autónomo. Puede usar `new-project` y `plan-phase` de GSD para generar un plan completo, y luego usar los ciclos de Ralph para ejecutar las fases que no requieren intervención humana.
**P: ¿Qué pasa con la colaboración entre múltiples personas?**
R: El directorio `.planning/` puede incluirse en Git. Varias personas pueden ejecutar diferentes fases y fusionar los resultados a través de Git. Sin embargo, se recomienda evitar ejecutar la misma fase simultáneamente.
## Resumen
El valor fundamental de GSD radica en **ocultar la complejidad dentro del sistema mientras mantiene la simplicidad para el usuario**. Solo necesita unos pocos comandos — `new-project`, `discuss-phase`, `plan-phase`, `execute-phase`, `verify-work` — mientras el sistema se encarga en segundo plano de toda la gestión de contexto, la orquestación de subagentes y la verificación de calidad.
Desde la instalación hasta la entrega, GSD proporciona un camino claro: describa lo que desea, discuta los detalles de implementación, genere planes atómicos, ejecute en paralelo y verifique los entregables. Cada paso le da la oportunidad de intervenir, y cada paso queda documentado.
Esto no es magia de "presione un botón y listo". Es un sistema que requiere su participación pero asume la mayor parte de la carga cognitiva. Como dice TÂCHES: usted es el gerente de proyecto de alto nivel, GSD es su equipo de ejecución.
***
**Lectura adicional**:
* [Análisis profundo de GSD](/es/docs/notes/gsd/concept) — Principios fundamentales, flujos de trabajo y arquitectura técnica
* [Análisis profundo de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) — Context Rot y la metodología Ralph
* [Guía práctica de snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank) — Instalación de Ralph, redacción de PRD y guía práctica
* [¿Qué es el desarrollo dirigido por especificaciones?](/es/docs/notes/speckit/concept) — De Vibe Coding al desarrollo dirigido por especificaciones
* [Guía práctica de Speckit](/es/docs/notes/speckit/practice) — Referencia de comandos de Speckit y ejemplos completos
# gstack: Cuando el CEO de YC pone su experiencia empresarial en Claude Code
## Introducción
En notas anteriores, exploramos varias "soluciones de mejora" en el ecosistema de Claude Code, desde el bucle infinito de [Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) hasta el desarrollo basado en especificaciones de [GSD](/es/docs/notes/gsd/concept). Todos intentan responder la misma pregunta: \*\* ¿Cómo cambiar la programación de IA de "adaptación" a "entrega confiable"? \*\*
La respuesta de Ralph es "reiniciar todo": utilice un nuevo proceso cada vez para evitar el deterioro del contexto. La respuesta de GSD es "basada en las especificaciones": garantizar la calidad a través de ciclos estructurados de planificación y validación de fases. Pero, ¿qué sucede si desea no sólo un sistema de ejecución, sino un equipo de ingeniería virtual completo? El director ejecutivo toma decisiones sobre el producto, el gerente de ingeniería revisa la arquitectura, el diseñador controla la experiencia, el control de calidad ejecuta pruebas reales del navegador y el ingeniero de lanzamiento gestiona el lanzamiento... todo lo juega la IA y usted lo controla.
Esta es la idea central de gstack.
## ¿Qué es gstack?
**Garry Tan**, el creador de gstack, tiene una rica experiencia técnica y empresarial: comenzó a escribir código a la edad de 14 años, se graduó en Ingeniería Informática de Stanford, es el décimo empleado de Palantir, cofundó Posterous (luego adquirido por Twitter) y se ha desempeñado como presidente y director ejecutivo de Y Combinator desde 2023.
Usó gstack para publicar más de 600 000 líneas de código de producción (35 % de prueba) en 60 días, con un promedio de más de 10 000 líneas por día, mientras seguía ejecutando YC a tiempo completo. Uno de los proyectos, garylist.org, se lanzó en 21 días, con 150.000 líneas de código y una cobertura de prueba del 35%. Según sus propias palabras, la calidad del código supera el proyecto empresarial anterior en el que gastó 5 millones de dólares, dos años y 10 ingenieros.
Desde que el proyecto fue de código abierto el 11 de marzo de 2026, pasó de v0 a v0.15.1.0 en 3 semanas y GitHub recibió más de 60 500 estrellas. Licencia MIT, completamente de código abierto.
## La posición de gstack en el ecosistema de herramientas
| Dimensiones | Código Claude nativo | Ralph Wiggum | GSD | Kit de especificaciones | Superpoderes | **gpila** |
| ------------------------ | ----------------------------------------- | --------------------------- | -------------------------------------------------- | -------------------------------------- | -------------------------------------------- | -------------------------------------------------------- |
| Posicionamiento central | Asistente de codificación universal de IA | Iteración de bucle infinito | Ingeniería contextual + basada en especificaciones | Requisitos → Especificaciones → Tareas | Disciplina de Procesos + TDD | **Equipo virtual basado en roles** |
| Patrón central | Programación conversacional | Bash Loop + Nuevo proceso | Hoja de ruta basada en fases | Especificaciones → Plan → Tareas | Estricto plan de desarrollo | **Proceso de siete pasos de Sprint** |
| Participación humana | Conversaciones en vivo | Sin intervención (AFK) | Verificación por etapa | Aprobación de especificaciones | Validación por paso | **Revisión de roles por etapa** |
| Capacidades únicas | Codificación básica | Iteración ilimitada | Gestión de la descomposición del contexto | Seguimiento de requisitos | TDD forzado | **Automatización del navegador + revisión multifunción** |
| Adecuado para escenarios | Tareas sencillas | Iteración continua | Gestión de proyectos a gran escala | Proyectos con requisitos rigurosos | Aseguramiento de la calidad de la ingeniería | **Desarrollo de productos de proceso completo** |
Se puede ver un patrón clave en la tabla: \*\*Estas herramientas no compiten entre sí, sino que resuelven problemas de programación de IA en diferentes dimensiones. \*\*
Superpowers utiliza **disciplina de proceso** para garantizar la calidad del código (TDD obligatorio, diálogo estructurado, plan de implementación); GSD utiliza **ingeniería de contexto** para gestionar proyectos complejos (planificación de fases, contexto nuevo del subagente, estado del sistema de archivos); gstack utiliza **descomposición de roles** para mejorar la calidad de la toma de decisiones (la perspectiva del CEO revisa los productos, los gerentes de ingeniería revisan la arquitectura, el control de calidad ejecuta navegadores reales).
En pocas palabras, Superpowers se basa en barreras de protección de procesos y gstack se basa en el diseño de roles: el primero es adecuado para la implementación de proyectos de 1 a N y el segundo es adecuado para la construcción de productos de 0 a 1. \*\*Los dos son productos complementarios en lugar de competidores. \*\*
## Flujo de trabajo principal: los siete pasos del Sprint
gstack organiza todo el proceso de desarrollo en un ciclo de **Pensar → Planificar → Construir → Revisar → Probar → Enviar → Reflexionar**, llamado "El Sprint"; no es un Sprint ágil, sino un ritmo de desarrollo de "roles que aparecen en secuencia".
### 1. Piense: Clínica de productos
```text
/office-hours
```
Esta es la habilidad más distintiva de gstack. La inspiración proviene directamente del horario de oficina de YC: los empresarios van a conocer a los socios de YC y se someten a un examen de conciencia. La IA te hará **6 preguntas forzadas**:
1. ¿Quién necesita esto específicamente?
2. ¿Qué pasa si no lo tienen hoy?
3. ¿Por qué es urgente este asunto ahora?
4. ¿Cómo sabes que funciona?
5. ¿Qué pasa si no haces nada?
6. ¿Cuál es la versión más pequeña que puedes lanzar?
El propósito no es ayudarlo a escribir código, sino **reexaminar el problema en sí** antes de escribir código.
### 2. Plan: revisión de funciones múltiples
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
CEO Review es esencialmente un "Modo Fundador": en lugar de ejecutar los requisitos literalmente, das un paso atrás y preguntas "¿Cuál es el verdadero propósito de este producto?" Admite cuatro modos: Ampliar alcance, Ampliar selectivamente, Mantener alcance y Reducir alcance.
### 3. Compilación: implementación de codificación
Comience a codificar según el plan aprobado. Este paso utiliza capacidades estándar de Claude Code.
### 4. Revisión: revisión paralela de expertos
```text
/review
```
Esta habilidad envía **7 subagentes paralelos** a la vez para revisar el código desde 7 perspectivas: pruebas, mantenibilidad, seguridad, rendimiento, migración de datos, contrato de API y ataque del equipo rojo. Los problemas obvios se solucionarán automáticamente.
### 5. Prueba: control de calidad real del navegador
```text
/qa
```
No es un examen de práctica. La habilidad de control de calidad inicia un **navegador Chromium real sin cabeza**, abre su aplicación, hace clic en botones, completa formularios y toma capturas de pantalla, tal como lo haría un evaluador real. Corrija errores automáticamente, genere pruebas de regresión y vuelva a verificar después de que se descubran errores.
### 6. Enviar: publicación con un solo clic
```text
/ship
```
Sincronice automáticamente la rama maestra, ejecute pruebas, revise diferencias, actualice números de versión y CHANGELOG, confirme, envíe y cree relaciones públicas. Si el proyecto no tiene un marco de prueba, incluso creará uno primero.
### 7. Reflexionar: revisar y aprender
```text
/retro
```
Informe semanal estilo gerente de ingeniería: analice el historial de confirmaciones, la proporción de pruebas y las tendencias de calidad del código. Apoye el análisis de equipos de varias personas y realice un seguimiento de indicadores como "número de días de lanzamiento consecutivos".
## Por qué funciona: principios técnicos
### Explorar Daemon: Pon tus ojos en la IA
La contribución técnica más exclusiva de gstack es Browse Daemon, una instancia persistente de Chromium sin cabeza que se comunica a través de HTTP del host local. La primera llamada inicia el navegador (\~3 segundos) y cada comando posterior tarda solo entre 100 y 200 ms. Esto significa que la IA realmente puede ver su aplicación, en lugar de adivinar la estructura DOM.
También presenta el **Sistema de referencia** (referencia de elemento `@e1`, `@e2`) para ubicar elementos a través del árbol de accesibilidad sin escribir selectores CSS. Se trata de una "contribución verdaderamente técnica" que es generalmente reconocida por la comunidad (incluidos los críticos).
### Desglose de roles: no un agente, sino un equipo
Lo que hace gstack es desmontar todos los roles en archivos de aviso independientes, lo que permite a Claude Code cambiar a las perspectivas de diferentes roles en diferentes etapas para revisar el código. Se trata esencialmente de una ingeniería rápida refinada.
La idea central es: \*\*La planificación no es igual a la revisión, la revisión no es igual al lanzamiento, y el gusto del fundador y el rigor de la ingeniería son modos de pensar completamente diferentes. \*\* En lugar de que un agente general haga todo, cambie los "modos cerebrales" cuando sea necesario: pensamiento fundador, rigor de ingeniería, revisión paranoica, ejecución rápida.
### Tres filosofías principales
ETHOS.md de gstack registra tres conceptos centrales:
1. **Boil the Lake**: cuando la IA reduce el costo marginal de la integridad a cero, elija siempre una implementación completa: cobertura de prueba del 100 %, todos los casos extremos, todas las rutas de error. Los "atajos de liberación" son una idea antigua.
2. **Buscar antes de construir**: tres capas de conocimiento: patrones probados en el tiempo, soluciones nuevas y populares y primeros principios. Empiece por comprender lo que todos hacen, cuestionar sus suposiciones y descubrir por qué las soluciones habituales son incorrectas.
3. **Soberanía del usuario**: recomendación de IA, toma de decisiones humana. Incluso si dos modelos de IA llegan a un consenso, el juicio del usuario sigue teniendo prioridad, porque el usuario tiene conocimiento del dominio, perspectiva estratégica y gusto.
## Los límites y controversias de gstack
La reacción de la comunidad a gstack es probablemente la más polarizadora de cualquier herramienta de programación de IA.
**El lado positivo**: los fundadores y los constructores no técnicos generalmente están de acuerdo, especialmente las habilidades de "pensamiento de producto" como `/office-hours` y `/plan-ceo-review`, que han ayudado a muchos desarrolladores independientes a reexaminar la dirección del producto antes de comenzar a codificar. La revisión de ingeniería (`/review`) puede descubrir algunas vulnerabilidades de seguridad ocultas. Este modelo de revisión paralela de múltiples ángulos tiene valor práctico.
El **lado cuestionador** también es muy directo:
* **El indicador LOC tiene poca importancia**: 600.000 líneas de código en 60 días. El número de líneas de código nunca es un indicador de calidad. Una gran cantidad de código puede ser sólo un andamiaje y un texto repetitivo.
* **Esencialmente una plantilla de aviso**: cada habilidad es un archivo SKILL.md y el umbral técnico no es alto. El valor real no está en el archivo en sí, sino en la calidad del diseño del mensaje.
* **Limitaciones del código de autorrevisión de AI**: `/review` Dejar que AI revise el código escrito por AI equivale a corregir su propia tarea. El paralelismo de múltiples funciones puede aliviar este problema, pero sigue siendo el mismo modelo.
* **Bono de efecto de celebridad**: si el fundador no es el CEO de YC, existe una alta probabilidad de que este proyecto no reciba tanta atención.
**Mi opinión**: Dejando a un lado las controversias, las partes realmente valiosas de gstack son dos: la tecnología de automatización del navegador de Browse Daemon y el patrón de diseño de descomposición de roles. Nada de esto depende de quién sea Garry Tan. La importancia central de la roleización no está en el nivel técnico, sino en el nivel de comportamiento: le ayuda a organizar su flujo de trabajo de IA de forma más consciente, en lugar de dejarlo todo en manos de un agente general.
gstack es adecuado para bifurcar y personalizar. Puede obtener las habilidades que necesita y cambiar las indicaciones que desee, en lugar de copiarlas todas.
## Recursos de vídeo
## Escribe al final
gstack representa una dirección interesante para las herramientas de programación de IA: no hacer que la IA sea más autónoma (la ruta de Ralph), ni hacer el proceso más rígido (la ruta de los Superpoderes), sino dejar que la IA desempeñe diferentes roles para mejorar la calidad de las decisiones. Su controversia simplemente ilustra la riqueza del ecosistema de programación de IA: ninguna solución se adapta a todos.
Si está interesado en gstack, el siguiente paso es leer el [Capítulo práctico](/es/docs/notes/gstack/practice), un tutorial paso a paso desde la instalación hasta la ejecución del flujo de trabajo completo.
***
**Lectura relacionada**:
* [Introducción a los conceptos de GSD](/es/docs/notes/gsd/concept) — Otra solución estructurada de programación de IA
* [Análisis en profundidad de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) — Comprender el punto de partida de la iteración de bucle infinito
* [Concepto de Habilidades de Claude](/es/docs/notes/claude-skills/concept) — Comprender el mecanismo subyacente de las Habilidades
# Panorama de habilidades front-end de gstack: flujo de trabajo de IA desde el diseño hasta el lanzamiento
## Introducción
En las notas anteriores, hablamos de [Qué es gstack](/es/docs/notes/gstack/concept), [Cómo ejecutar el flujo de trabajo](/es/docs/notes/gstack/practice) y [Estructura de ingeniería de habilidades](/es/docs/notes/gstack/skill-architecture). Pero hay una pregunta que no se ha discutido: entre las más de 60 habilidades que se obtienen después de la instalación de gstack, ¿cuáles están relacionadas con el diseño de interfaz de usuario/UI? ¿En qué orden? \*\*
Esta nota hace dos cosas: primero, clasifica las \~27 habilidades relacionadas con el front-end por función, y luego usa un pequeño proyecto interesante (la página de aniversario de cuenta regresiva) para recorrer el flujo de trabajo completo de principio a fin, con capturas de pantalla en cada etapa, para que puedas ver el efecto real.
## Panorama del kit de habilidades de front-end
Las habilidades de front-end de gstack se pueden dividir en 6 capas funcionales, desde los cimientos hasta el techo, y cada capa resuelve problemas en diferentes etapas.
### Diseño de infraestructura
Una configuración única a nivel de proyecto para determinar el lenguaje de diseño, y todas las Habilidades posteriores se referirán a estos puntos de referencia.
| Habilidad | Qué hacer | Cuándo utilizar |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- |
| `/design-consultation` | Consulta completa del sistema de diseño, combinación de colores de salida, fuentes, espaciado, dirección de textura | Al iniciar un nuevo proyecto, o si deseas redefinir el estilo visual |
| `/teach-impeccable` | Recopile preferencias de diseño a la vez y escríbalas en el archivo de configuración de AI | Ejecútelo una vez después de instalar gstack para que la IA recuerde su estética |
| `/brand-guidelines` | Aplicar especificaciones de fuente y combinación de colores de marca existentes | Aplicar directamente cuando exista un manual de marca existente |
> Si el proyecto ya tiene `DESIGN.md`, se puede omitir este nivel.
### Exploración del diseño
Cuando no esté seguro de la dirección, compare rápidamente varias opciones.
| Habilidad | Qué hacer | Cuándo utilizar |
| ------------------ | --------------------------------------------------------------------- | ------------------------------------------------------------------- |
| `/design-shotgun` | Genere de 3 a 5 soluciones visuales, abra el panel de comparación | No estoy seguro de qué estilo quieres, quiero ver las posibilidades |
| `/frontend-design` | Genere código de interfaz front-end reconocible a nivel de producción | Trabaje directamente después de que la dirección esté clara |
| `/canvas-design` | Generar carteles, arte visual (PNG/PDF) | Requiere un diseño visual estático en lugar de componentes web |
### Implementación del diseño
Convierta el plan en código verdaderamente ejecutable y maneje la composición tipográfica, el diseño y la capacidad de respuesta.
| Habilidad | Qué hacer | Cuándo utilizar |
| ------------------------ | ----------------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| `/design-html` | Convierta el borrador de diseño confirmado en HTML/CSS de nivel de producción | Tengo maquetas que quiero implementar directamente |
| `/mobile-responsiveness` | Diseño responsivo para dispositivos móviles e interacción táctil | Adaptación móvil desde cero |
| `/adapt` | Adaptación de puntos de interrupción entre dispositivos y tamaños de pantalla | Existe una versión de escritorio y hay que adaptarla a móviles/tablets |
| `/typeset` | Selección de fuente, nivel, tamaño, grosor y optimización de legibilidad | El diseño del texto parece "casi sin sentido" |
| `/arrange` | Espaciado de diseño, ritmo visual, reparación de alineación | Espaciado inconsistente, el diseño se siente abarrotado o disperso |
### Mejoras de diseño
Sobre la base de la finalización funcional, inyecta efectos dinámicos, personalidad y detalles emocionales.
| Habilidad | Qué hacer | Cuándo utilizar |
| ------------ | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| `/animate` | Agregue microinteracciones y animaciones útiles | La funcionalidad de la página está bien, pero se siente "rígida" |
| `/delight` | Añade detalles sorpresa y toques personalizados | Quiere que los usuarios recuerden esta página |
| `/bolder` | Amplificar el impacto visual | El diseño es demasiado sencillo y demasiado seguro |
| `/colorize` | Agregue color estratégico a la monótona interfaz | La página es demasiado gris, demasiado sencilla y le falta calidez |
| `/overdrive` | Efectos técnicos a nivel de explosión: sombreador, física de resorte, animación de desplazamiento | Un área determinada quiere un efecto sorpresa |
| `/onboard` | Nuevo proceso de guía del usuario, diseño de estado vacío | Experiencia de usuario por primera vez |
Estas cuatro habilidades de mejora están en una **relación progresiva**: `animate` es el efecto dinámico básico, `delight` es emocional, `bolder` es amplificación y `overdrive` es explosión. Apílalos paso a paso según las necesidades del proyecto, no es necesario utilizarlos todos.
### Optimización del diseño
Convergencia y refinamiento: eliminar el exceso, alinear las desviaciones y pulir las asperezas.
| Habilidad | Qué hacer | Cuándo utilizar |
| ------------ | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| `/polish` | Pulido de calidad final: alineación, espaciado, consistencia | Un último pase antes del lanzamiento |
| `/quieter` | Reducir la intensidad de la estimulación visual | El diseño es demasiado sofisticado y ruidoso |
| `/distill` | Minimizar y eliminar complejidad innecesaria | Hay demasiados elementos en la página y quiero reducirlos |
| `/normalize` | Alinear los estándares del sistema de diseño (token, espaciado, color) | El estilo se desvía de las especificaciones de DESIGN.md |
| `/clarify` | Mejore la redacción publicitaria de UX, los mensajes de error y la redacción de las etiquetas | La redacción publicitaria es confusa, los mensajes de error no son amigables |
### Revisión y verificación del diseño
Inspección sistemática antes de conectarse, encontrar problemas, calificarlos y solucionarlos.
| Habilidad | Qué hacer | Cuándo utilizar |
| --------------------- | ------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| `/plan-design-review` | Revisión del plan de diseño antes de la implementación (puntaje 0-10) | Quiero que AI revise el plan desde la perspectiva de un diseñador |
| `/design-review` | Control de calidad visual después de la implementación, comparación y reparación automática de capturas de pantalla | Una vez escrito el código, verifique el grado de restauración visual |
| `/critique` | Evaluación UX: Jerarquía Visual, Carga Cognitiva, Resonancia Emocional | ¿Quiere un informe de revisión de diseño estructurado? |
| `/audit` | Revisión técnica: accesibilidad, rendimiento, temas, capacidad de respuesta | Controles sistemáticos antes de la puesta en marcha |
| `/benchmark` | Pruebas de referencia de rendimiento, comparación antes/después | Quiere cuantificar el impacto de los cambios en el rendimiento |
## Demostración práctica: utilice la página de aniversario de cuenta regresiva para recorrer todo el proceso
Solo mirar la tabla de clasificación es demasiado abstracto. Usamos un pequeño proyecto para unir las habilidades anteriores: hacer una **página única de cuenta regresiva/aniversario**: elija una fecha significativa y cree una pantalla de cuenta regresiva con animación digital y efectos de fondo.
Este proyecto es pequeño pero completo, lo suficiente para cubrir la mayoría de los 6 niveles de habilidades. El proceso completo es:
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> No es necesario ejecutar los 7 pasos cada vez. Una vez que domine, los enlaces de uso común son solo `/frontend-design → /animate → /polish → /ship` cuatro pasos. Para mostrar la habilidad completa, aquí se dan todos los pasos.
### Etapa 1: Infraestructura: determinar el lenguaje de diseño
**Skill**:`/design-consultation` + `/teach-impeccable`
Sólo es necesario hacerlo una vez al inicio del proyecto. Genera `DESIGN.md`, lo que permite a la IA recordar sus preferencias de diseño. Si el proyecto ya tiene `DESIGN.md`, omítelo directamente.
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* TODO: Captura de pantalla: fragmento DESIGN.md producido por design-consultation */}
### Etapa 2: Exploración - Comparación de múltiples opciones
**Skill**:`/design-shotgun`
Cuando no esté seguro de la dirección, deje que la IA genere de 3 a 5 soluciones visuales y abra el panel de comparación para elegir.
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* TODO: Captura de pantalla: 3 paneles de comparación de soluciones generados por design-shotgun */}
Elija una dirección entre 3 opciones. Si sabe exactamente lo que quiere, omita este paso y vaya directamente a la Etapa 3.
### Etapa 3: Construir - producir código a nivel de producción
**Skill**:`/frontend-design` + `/adapt`
enlace central. código mientras se garantiza la capacidad de respuesta desde el principio.
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* TODO: Captura de pantalla: el efecto de la página de escritorio una vez completada la construcción */}
{/* TODO: Captura de pantalla: efecto de la versión móvil (después de la adaptación de /adapt) */}
### Etapa 4: Mejorar – Inyectar movimiento y personalidad
**Habilidad**: `/animate` → `/delight` (bajo demanda `/overdrive`)
Estos tres están en una relación progresiva: `animate` es el efecto dinámico básico, `delight` es el detalle emocional y `overdrive` es el efecto explosivo. Agregue capas según sea necesario.
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
Tenga en cuenta las restricciones de animación a las que se hace referencia `DESIGN.md`: si el sistema de diseño solo permite una transición de desplazamiento de 150 ms, `/overdrive` no se aplicará. Este es un ejercicio de buen juicio.
{/* TODO: Captura de pantalla o GIF: antes y después de la mejora del movimiento */}
### Etapa 5: Ajuste - Convergencia y pulido
**Habilidad**: `/typeset` + `/polish` (bajo demanda `/distill`, `/normalize`)
Alineación de espacios, jerarquía de fuentes, ritmo visual. Si descubre que ha agregado demasiado, use `/distill` para restar.
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* TODO: Captura de pantalla: comparación detallada del pulido antes y después del pulido */}
### Etapa 6: Revisión – Verificación sistemática
**Skill**:`/design-review` + `/audit`
Control de calidad visual + revisión técnica. `/design-review` tomará capturas de pantalla automáticamente para comparar y solucionar problemas, `/audit` verificará la accesibilidad y el rendimiento.
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* TODO: Captura de pantalla: informe de puntuación elaborado por la auditoría */}
### Etapa 7: Lanzamiento
**Skill**:`/ship`
Proceso de lanzamiento estándar de gstack: pruebas, revisión de diferencias, creación de relaciones públicas.
```text
> /ship
```
***
**Resultados esperados**: una página de cuenta regresiva visualmente exquisita, que utiliza entre 8 y 10 habilidades de front-end en el proceso. Lo que es más importante es establecer la intuición de "qué habilidad utilizar y en qué etapa".
## Hoja de referencia diaria
Lo anterior es el proceso completo. Si encuentra problemas específicos en el desarrollo diario, simplemente consulte esta tabla:
| Mi pregunta actual | Qué usar |
| ------------------------------------------------------------- | ---------------------------------------------- |
| No sé qué estilo quiero | `/design-shotgun` |
| La función de la página es buena pero se siente "casi inútil" | `/polish` |
| Siento que algo anda mal pero no puedo explicarlo | `/design-review` |
| La fuente/diseño parece incómodo | `/typeset` |
| Espaciado desordenado y diseño abarrotado | `/arrange` |
| El estilo se desvía del sistema de diseño | `/normalize` |
| Quiere extraer componentes públicos | `/extract` |
| La página es demasiado compleja y quiero restar | `/distill` |
| El diseño es demasiado sencillo y demasiado seguro | `/bolder` o `/colorize` |
| El diseño es demasiado sofisticado y ruidoso | `/quieter` |
| El texto del mensaje de error no es amigable | `/clarify` |
| Hay un problema con la pantalla del teléfono móvil | `/adapt` |
| Quiere agregar efectos de animación | `/animate` (básico) o `/overdrive` (explosión) |
| Inspección sistemática antes de conectarse | `/audit` |
| debug | `/investigate` |
## Resumen
Esta nota hace dos cosas:
1. **Panorama**: las 27 habilidades de front-end de gstack se clasifican en 6 capas (infraestructura→exploración→implementación→mejora→optimización→revisión)
2. **Demostración práctica**: use una página de aniversario de cuenta regresiva para recorrer el flujo de trabajo completo, mostrando qué habilidades se utilizan en cada etapa y por qué.
Conclusión clave: el uso más poderoso de estas habilidades no es llamarlas individualmente, sino combinarlas en un proceso: explorar direcciones, crear implementaciones, mejorar el pulido y revisar lanzamientos, con una selección clara de habilidades en cada etapa.
Pero no se deje atar por el proceso: una vez que domine, `/frontend-design → /animate → /polish → /ship` cuatro pasos son suficientes la mayor parte del tiempo.
***
**Lectura relacionada**:
* [gstack Concepts](/es/docs/notes/gstack/concept) — ¿Qué es gstack y qué problemas resuelve?
* [Capítulo práctico de gstack](/es/docs/notes/gstack/practice) — Flujo de trabajo completo desde la instalación hasta la ejecución
* [Desmontaje de la arquitectura de habilidades de gstack](/es/docs/notes/gstack/skill-architecture) — ¿Qué pueden aprender los desarrolladores de habilidades?
* [Concepto de Habilidades de Claude](/es/docs/notes/claude-skills/concept) — Comprender el mecanismo subyacente de las Habilidades
# Práctica de gstack: flujo de trabajo completo desde la instalación hasta la ejecución
## Introducción
En [Concepto](/es/docs/notes/gstack/concept), aprendimos sobre el posicionamiento central de gstack, un conjunto de habilidades basado en roles que convierte a Claude Code en un equipo de ingeniería virtual, y su posicionamiento diferenciado en el ecosistema de herramientas de programación de IA en comparación con GSD, Superpowers, Ralph y otras soluciones.
Este artículo práctico se centra en **cómo utilizar**: desde la instalación y configuración hasta la ejecución del flujo de trabajo completo, ayudándole a empezar a usar gstack en 30 minutos.
## Instalación y configuración
### Condiciones previas
* **Código Claude** está instalado y disponible
* **Git** instalado
* **Bun v1.0+** instalado (gstack está basado en Bun)
* Los usuarios de Windows también necesitan Node.js
### Instalación global (recomendada, completada en 30 segundos)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
El script de instalación hace tres cosas:
1. Agregue la información de habilidades de gstack a su archivo `CLAUDE.md`
2. Coloque todos los archivos de habilidades en el directorio de habilidades.
3. Instale Playwright y el navegador Chromium correspondiente (para `/browse` y `/qa`)
### Instalación a nivel de proyecto (compartir en equipo)
Si desea que los miembros del equipo obtengan gstack automáticamente después de clonar el repositorio:
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
\###Soporte multiagente
gstack no se limita a Claude Code y actualmente admite **10 agentes de programación de IA**. `./setup` detecta automáticamente los hosts instalados de forma predeterminada:
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
La ruta de instalación de habilidades de cada host tiene la forma `~/./skills/gstack-*/` y no interfiere entre sí.
> 💡 **Opciones adicionales para usuarios de OpenClaw**: Además de llamar a través de ACP, OpenClaw también puede instalar directamente 4 habilidades de metodología nativa (`gstack-openclaw-office-hours`, `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`, `gstack-openclaw-retro`) a través de ClawHub, que se pueden usar de manera conversacional sin una sesión de Claude Code.
### Modo Equipo (Compartir equipo + Actualizaciones automáticas, recomendado)
v1.x presenta el Modo Equipo: cada desarrollador instala gstack globalmente y el almacén solo registra "usamos gstack" y las actualizaciones se realizan automáticamente:
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
Reemplazar `required` con `optional` es un "recordatorio amable" en lugar de obligatorio. Cada vez que inicie Claude Code, ejecutará automáticamente una verificación de actualización (aceleración una vez por hora, segura y silenciosa si falla la red). No hay archivos vendidos en el almacén y no hay cambios de versión.
### Actualización
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
O use `/gstack-upgrade` directamente en Claude Code.
## Referencia de comando completa
### Proceso de sprint
| Comando | Rol | Descripción |
| --------------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/office-hours` | Horario de oficina de YC | 6 preguntas forzadas para reconstruir la dirección del producto y generar documentos de diseño |
| `/plan-ceo-review` | CEO / Fundador | Buscando productos de 10 estrellas, disponibles en cuatro modelos de gama |
| `/plan-eng-review` | Gerente de Ingeniería | Arquitectura de bloqueo, flujo de datos, casos extremos, matriz de prueba |
| `/plan-design-review` | Diseñador sénior | Puntuación de la dimensión de diseño 0-10, explique cómo lograr 10 puntos |
| `/plan-devex-review` | Líder de experiencia del desarrollador | Explore retratos de desarrolladores, compare TTHW y diseñe momentos mágicos; tres modos (EXPANSIÓN DX / POLACO / TRIAJE), 20-45 preguntas forzadas |
| `/autoplan` | Revisar el proceso | Ejecute automáticamente CEO → Diseño → Ingeniería → Revisión DX en secuencia, decida automáticamente de acuerdo con los principios de toma de decisiones de codificación y solo le presente "decisiones de gusto" |
### Diseño
| Comando | Descripción |
| ---------------------- | -------------------------------------------------------------------------------------- |
| `/design-consultation` | Cree un sistema de diseño completo desde cero y genere DESIGN.md |
| `/design-shotgun` | Genere múltiples variantes de diseño de IA y compare selecciones en el navegador |
| `/design-html` | Genere HTML/CSS de nivel de producción, admita la detección de marcos React/Svelte/Vue |
### Revisión y seguridad
| Comando | Rol | Descripción |
| ---------------- | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/review` | Ingeniero de personal | Encuentre errores que puedan pasar la CI pero que explotarán en producción, solucionarán automáticamente problemas obvios y marcarán brechas de integridad |
| `/investigate` | Experto en depuración | Depuración sistemática de causas raíz. Regla de hierro: no solucione el error hasta que encuentre la causa raíz; detener después de 3 correcciones fallidas |
| `/design-review` | Diseñador que puede escribir código | Auditoría visual + reparación automática, envío atómico, capturas de pantalla comparativas antes y después |
| `/devex-review` | Probador DX | Realice la incorporación: explore documentos, ejecute el proceso de entrada, cronometre TTHW, errores de captura de pantalla, compare con la puntuación `/plan-devex-review` |
| `/cso` | Oficial de seguridad | OWASP Top 10 + modelado de amenazas STRIDE, 17 reglas de exclusión de falsos positivos, umbral de confianza 8/10, cada hallazgo va acompañado de escenarios de utilización específicos |
### Pruebas y control de calidad
| Comando | Descripción |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/qa` | Abra la prueba del navegador real y busque el error → Corrección de confirmación atómica → Generar prueba de regresión → Volver a verificar |
| `/qa-only` | Igual que el anterior pero solo informes, sin modificaciones de código |
| `/benchmark` | Prueba de rendimiento de referencia: carga de páginas, Core Web Vitals, tamaño de recursos, soporte antes y después de la comparación |
| `/browse` | Comandos de navegador de nivel \~100 ms, Chromium real, capturas de pantalla, llenado de formularios, clics en elementos |
| `/open-gstack-browser` | Inicie el navegador GStack: control de IA visible Chromium, viene con extensión de barra lateral, sigilo anti-rastreo, enrutamiento automático de modelos (operación Sonnet/análisis Opus), admite importación de cookies con un solo clic |
| `/setup-browser-cookies` | Importe cookies de navegadores reales (Chrome/Arc/Brave/Edge) a sesiones sin cabeza para probar páginas que requieren inicio de sesión |
| `/pair-agent` | Emparejamiento de navegadores de agentes entre IA: comparta el mismo navegador GStack con OpenClaw / Hermes / Codex / Cursor, etc., cada agente tiene una pestaña independiente, viene con túnel ngrok para admitir agentes remotos, token de alcance + aislamiento de pestañas + límite de velocidad + atribución de comportamiento |
### Lanzamiento y operación y mantenimiento.
| Comando | Descripción |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/ship` | Sincronizar la rama principal → Ejecutar pruebas → Auditar cobertura → Actualizar versión → Enviar push → Crear PR; Bootstrap automático cuando el proyecto no tiene un framework de prueba |
| `/land-and-deploy` | Fusionar PR → Esperar CI → Implementar → Verificar el estado del entorno de producción |
| `/canary` | Monitoreo canary posterior a la implementación: errores de consola, regresiones de rendimiento, fallas de página |
| `/setup-deploy` | `/land-and-deploy` Configuración única: plataforma de detección automática (Fly.io/Render/Vercel/Netlify/Heroku/GitHub Actions/custom) + URL de producción + comando de implementación |
| `/setup-gbrain` | Comience con la base de datos GBrain con un solo clic (en 5 minutos): PGLite local, URL existente de Supabase o cree automáticamente un nuevo proyecto de Supabase a través de la API de administración; Registro MCP + permisos de lectura-escritura/solo lectura/denegación a nivel de almacén |
### Revisa y aprende
| Comando | Descripción |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/retro` | Informe semanal de percepción del equipo: desmontaje per cápita, estadísticas de racha ganadora, tendencias de salud de las pruebas, oportunidades de crecimiento; `/retro global` en todos los proyectos + herramientas de IA (Claude Code / Codex / Gemini) |
| `/document-release` | Actualizar automáticamente la documentación del proyecto para que coincida con el código publicado (README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS); `/ship` ahora se llama automáticamente |
| `/learn` | Administre memorias de aprendizaje entre sesiones: ver, buscar, podar, exportar, acumular por proyecto |
| `/context-save` `/context-restore` | Paquete de modo de punto de control continuo: confirmación WIP automática para guardar contexto, use `/context-restore` para reconstruir la sesión después de una falla/cambio |
### Protección de seguridad
| Comando | Descripción |
| ----------------------- | ------------------------------------------------------------------------ |
| `/careful` | Advertencia de operación peligrosa: rm -rf, DROP TABLE, force-push, etc. |
| `/freeze` / `/unfreeze` | Bloquear/desbloquear el alcance de edición en un directorio específico |
| `/guard` | Combinación `/careful` + `/freeze`, modo de máxima seguridad |
| `/checkpoint` | Guardar/restaurar instantánea del estado de trabajo |
### Integración de herramientas
| Comando | Descripción |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/codex` | Integración de OpenAI Codex CLI: revisión de código independiente (puerta de aprobación/falla), modo de confrontación, modo de consulta; El análisis de superposición entre modelos se realizará después de ejecutar con `/review` |
| `/health` | Panel de calidad del código: tsc + bioma + knip + shellcheck + pruebas → puntuación general 0-10 |
| `/skillify` | Consolidar el flujo de trabajo actual en una habilidad reutilizable |
| `/scrape` | Flujo de trabajo de raspado web |
| `/landing-report` | Informe de experiencia y rendimiento de la página de destino |
| `/make-pdf` | Generar documento PDF |
| `/benchmark-models` `/model-overlays` `/plan-tune` | Comparación entre modelos, superposición de cobertura, optimización de planes |
### Standalone CLI(v0.19+)
Además del comando de barra diagonal, gstack también viene con un conjunto de CLI independientes (no se ejecutan dentro de la sesión de Claude Code):
| Comando | Descripción |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `gstack-model-benchmark` | Evaluación entre modelos: ejecute Claude / GPT (a través de Codex CLI) / Gemini al mismo tiempo, compare el retraso, el token, el costo y (opcional) el puntaje de calidad del juez LLM; proveedor no disponible se salta automáticamente |
| `gstack-taste-update` | Aprendizaje de gustos de diseño: escriba la aprobación/desaprobación de `/design-shotgun` en el archivo de gustos a nivel de proyecto, disminuya en un 5% cada semana y retroalimente a la generación de variantes posterior |
## Detalles de configuración
### CLAUDE.md Agregar contenido
Después de la instalación, gstack agregará una lista y una breve descripción de todas las habilidades disponibles a su `CLAUDE.md`. Esto le permite a Claude Code saber qué comandos están disponibles.
### Estructura del directorio de habilidades
La entrada principal es el `~/.claude/skills/gstack/SKILL.md` de nivel superior, cada subcomando existe en forma de un directorio plano y el núcleo es el archivo `SKILL.md`:
```text
~/.claude/skills/gstack/
├── SKILL.md # 主入口 skill
├── browse/ # 浏览器 daemon
├── qa/ # QA 测试
├── review/ # 代码审查
├── ship/ # 发布流程
├── plan-ceo-review/ # CEO 审查
├── office-hours/ # 产品门诊
├── pair-agent/ # 跨 Agent 浏览器配对
├── open-gstack-browser/ # GStack Browser 启动器
├── setup-gbrain/ # GBrain 数据库一键上手
├── hosts/ # 10 个 host 配置(claude/codex/cursor/...)
├── bin/ # standalone CLI(gstack-model-benchmark 等)
└── ... # 当前 v1.x 共 50 个 skill 目录
```
Eres libre de modificar cualquier `SKILL.md` para personalizar el comportamiento; esta es la ventaja de "bifurcar y personalizar".
### Browse Daemon
Browse Daemon es una instancia permanente de Chromium. Configuración clave:
* **Puerto**: 10000-60000 seleccionado aleatoriamente, admite más de 10 espacios de trabajo paralelos
* **Seguridad**: solo vincula localhost, usa autenticación de token de portador para cada sesión
* **Cookie**: use `/setup-browser-cookies` para importar desde Chrome/Arc/Brave/Edge
## Demostración práctica del flujo de trabajo
A continuación se muestra un flujo de trabajo típico de gstack. Los comandos y resultados se basan en casos reales de la documentación y los vídeos.
> 💡 **Nota**: El siguiente resultado es un ejemplo general compilado en base a investigaciones. En el futuro se agregarán capturas de pantalla de proyectos específicos según la práctica real.
### Paso 1: Clínica del producto
```text
> /office-hours
[YC Office Hours] 6 forcing questions:
1. Who specifically needs this?
2. What do they do today without it?
3. Why is this urgent right now?
4. How will you know it works?
5. What happens if you do nothing?
6. What is the smallest version you can ship?
→ Design doc generated
```
No se apresure a escribir código, primero deje que la IA atormente sus ideas desde la perspectiva del horario de oficina de YC.
### Paso 2: Plan de revisión de funciones múltiples
```text
> /autoplan
[CEO Review] Finding the 10-star product...
[Design Review] Rating dimensions 0-10...
[Eng Review] Locking architecture + edge cases...
→ Fully reviewed plan ready
```
`/autoplan` ejecuta automáticamente tres rondas de revisiones de CEO → Diseño → Ingeniería para producir un plan completo posterior a la revisión.
### Paso 3: implementación de codificación
Codifique normalmente según el plan aprobado. Puede utilizar la conversación estándar de Claude Code.
### Paso 4: revisión del código por parte de múltiples expertos
```text
> /review
Dispatching 7 specialist reviewers...
- Testing coverage ✓
- Maintainability ✓
- Security: Found 1 issue (auto-fixing)
- Performance ✓
- Data migration ✓
- API contract ✓
- Red team: No vulnerabilities found
→ Review complete, 1 auto-fix applied
```
### Paso 5: Control de calidad del navegador
```text
> /qa
Opening headless browser...
Testing user flows:
- Login flow ✓
- Dashboard load ✓
- Form submission: Bug found → fixing → re-testing ✓
- Image upload ✓
→ 4 flows tested, 1 bug fixed, regression test generated
```
### Paso 6: Publicar
```text
> /ship
Syncing with main...
Running tests: 42 passed, 0 failed
Reviewing diff: 3 files changed
Updating VERSION: 1.2.0 → 1.3.0
Creating PR: "Add screenshot feature"
→ PR #47 created, ready for merge
```
## Consejos prácticos y experiencia comunitaria
### Sugerencia de Garry Tan
ETHOS.md de gstack, tres principios básicos:
1. **Boil the Lake**: la IA hace que la integridad sea casi gratuita: siempre completa las cosas y no tomes atajos
2. **Buscar antes de construir**: buscar primero, comprender primero y luego comenzar después de la verificación de conocimientos de tres capas.
3. **Soberanía del usuario**: recomendación de IA, tú decides. Incluso si ambos modelos de IA coinciden, su criterio sigue teniendo prioridad
El archivo README de gstack comienza con una cita de Karpathy; este es también el punto de partida para que el propio Garry Tan explique por qué quiere construir gstack:
### Experiencias comunitarias positivas
* **`/office-hours` para solicitudes de YC**: varios solicitantes de S26 en Reddit r/ycombinator informaron que usar el horario de oficina de gstack para realizar pruebas de estrés en sus materiales de solicitud es muy efectivo.
* **La auditoría de seguridad encontró vulnerabilidades reales**: Hubo comentarios del CTO. `/review` descubrió una vulnerabilidad XSS que el equipo no conocía.
* **`/browse` Pruebas reales de navegador**: Reconocido por la comunidad (incluidos los críticos) como una "contribución verdaderamente técnica"
### Errores comunes
* **Solicitudes de permiso frecuentes**: algunos usuarios informaron que "las solicitudes de permiso deben aprobarse cada 30 segundos, lo que hace imposible dormir". Se recomienda configurar reglas de aprobación automática apropiadas en la configuración de Claude Code
* **Alto consumo de tokens**: los mensajes caracterizados aumentarán el consumo de contexto. Si es sensible a los costos, puede utilizar selectivamente las habilidades que más necesita
* **Agent Loop**: Hay casos en HN en los que los usuarios informaron que el agente quedó atrapado en un bucle de 70 minutos. Se recomienda establecer tiempos de espera y puntos de control razonables.
* **No para todos**: los desarrolladores experimentados pueden sentir que la mayoría de las habilidades son envoltorios innecesarios. gstack es más adecuado para **fundadores independientes y equipos pequeños** que para equipos con procesos de ingeniería maduros.
## Preguntas frecuentes y mejores prácticas
\*\*P: ¿Se pueden usar gstack y Superpowers al mismo tiempo? \*\*
Sí. Los dos se complementan entre sí: Superpowers es bueno en disciplina de procesos y garantía de TDD, y gstack es bueno en pensamiento de productos y revisiones de múltiples funciones. Muchos equipos utilizan Superpowers para la disciplina de codificación diaria y gstack para la planificación de productos y el control de calidad.
\*\*P: ¿El token es caro? \*\*
Superior al Código Claude nativo. El mensaje de rol de cada habilidad ocupa la ventana de contexto. Pero si su tiempo vale más que la tarifa simbólica, este suele ser un buen negocio.
\*\*P: ¿Para qué tipo de proyectos es adecuado? \*\*
Ideal para **desarrollo de productos de proceso completo**, desde la idea hasta el lanzamiento. Si simplemente corrige errores o crea pequeñas funciones, el código nativo de Claude es suficiente. El valor de gstack se maximiza en el "proceso completo".
\*\*P: ¿Cómo personalizar la habilidad? \*\*
Cada habilidad es un archivo `SKILL.md`. Simplemente edítelo directamente:
1. Busque el directorio de habilidades: `~/.claude/skills/gstack//`
2. Editar `SKILL.md`
3. Vuelva a ejecutar `./setup`
La comunidad recomienda bifurcar el repositorio y personalizarlo en lugar de modificar directamente la instalación global.
### Mejores prácticas
1. **Primero `/office-hours` luego código**: acostúmbrese a realizar clínicas de productos antes de escribir cualquier código.
2. **Haz un buen uso de la verificación `/browse`**: no te limites a mirar el código, deja que la IA realmente "vea" tu aplicación.
3. **Periódico `/retro`**: mantener la visibilidad de la calidad del código y el ritmo de trabajo.
4. **Adopción gradual**: No es necesario utilizar todas las habilidades a la vez. A partir de `/office-hours` + `/review` + `/ship`
5. **Personalización de bifurcación**: si encuentra un mensaje inapropiado, cámbielo directamente. Esta es la ventaja del código abierto.
## Resumen
El valor central de gstack no radica en cuán poderosa es una habilidad específica, sino en que proporciona un **modo de colaboración de IA estructurado**: a través del cambio de roles, puedes obtener diferentes tipos de asistencia de IA en diferentes etapas. Primero revise la dirección del producto desde la perspectiva del CEO, luego revise la arquitectura con el rigor de un gerente de ingeniería y finalmente verifique los resultados con el navegador real de control de calidad.
A continuación, puede intentar instalarlo usted mismo y comenzar su primer proyecto de gstack desde `/office-hours`.
***
**Lectura ampliada**:
* [gstack Concepts](/es/docs/notes/gstack/concept) — Comprender los conceptos centrales y el posicionamiento ecológico de herramientas de gstack
* [Capítulo práctico de GSD](/es/docs/notes/gsd/practice) — Otra guía práctica para soluciones estructuradas de programación de IA
* [Capítulo práctico de Habilidades de Claude](/es/docs/notes/claude-skills/skill-creator) — Comprender el mecanismo de creación de Habilidades
# Teardown gstack: qué habilidades pueden aprender los desarrolladores
## Introducción
En [Concepto](/es/docs/notes/gstack/concept) y [Práctico](/es/docs/notes/gstack/practice), aprendimos qué es gstack y cómo usarlo desde la perspectiva del usuario. Esta nota es desde una perspectiva diferente: \*\* Como desarrollador de habilidades \*\*, después de leer el archivo del almacén gstack, qué diseños de ingeniería vale la pena aprender y aprender.
gstack es más que una simple colección de 23 archivos de mensajes. Hay un sistema de ingeniería completo detrás: generación de plantillas, actualización automática, aprendizaje y memoria, guía progresiva, adaptación multiplataforma, pruebas en capas: estas son las claves para convertir un proyecto de habilidades de "utilizable" a "fácil de usar".
***
## 1. SKILL.md no está escrito a mano: sistema de generación de plantillas
El diseño más contrario a la intuición de gstack: \*\*Cada SKILL.md se genera automáticamente y no se puede editar directamente. \*\*
```text
SKILL.md.tmpl (人写) → gen-skill-docs → SKILL.md (机生)
```
La plantilla `.tmpl` escrita por humanos contiene lógica de flujo de trabajo y mejores prácticas, además de marcadores de posición `{{PLACEHOLDER}}`. El script de compilación extrae la referencia del comando, la lista de indicadores del navegador, el código de inicio del preámbulo, etc. del código fuente y los completa en los marcadores de posición para generar el SKILL.md final.
```text
{{PREAMBLE}} ← 从 resolvers/preamble.ts 生成的启动代码
{{BROWSE_SETUP}} ← 浏览器初始化指令
{{COMMAND_REFERENCE}} ← 从 commands.ts 提取的命令文档
{{SNAPSHOT_FLAGS}} ← 从源代码常量提取的快照选项
```
\*\*¿Por qué hacer esto? \*\*
* La documentación y el código nunca estarán desincronizados: la referencia del comando se genera a partir del código fuente y la documentación se actualiza automáticamente cuando cambia el código fuente.
* 23 habilidades comparten el mismo preámbulo (alrededor de 220 líneas) y todas las habilidades se actualizan simultáneamente
* CI puede `--dry-run` comprobar si el archivo generado ha caducado para evitar olvidarse de regenerar
**Conclusión**: Si mantiene varias habilidades, cualquier contenido compartido entre habilidades debe extraerse en plantillas y usarse en los pasos de compilación para generar los archivos finales. Sincronizar manualmente varias copias del mismo contenido causará problemas tarde o temprano.
***
## 2. Mecanismo de actualización: enlace completo desde la detección hasta la ejecución
El sistema de actualización de gstack está exquisitamente diseñado y dividido en tres capas:
### Primera capa: detección de versión
`bin/gstack-update-check` es un script bash independiente que hace lo siguiente:
1. Lea el archivo `VERSION` local
2. Verifique el caché `~/.gstack/last-update-check` (cachés UP\_TO\_DATE durante 60 minutos, cachés UPGRADE\_AVAILABLE durante 720 minutos)
3. Si el caché caduca, solicite HTTP `raw.githubusercontent.com/.../VERSION` de GitHub.
4. Compare el número de versión y genere `UPGRADE_AVAILABLE <旧> <新>`
### Segunda capa: integración del preámbulo
**La primera línea del código de inicio de SKILL.md de cada habilidad es la detección de versión**:
```bash
_UPD=$(~/.claude/skills/gstack/bin/gstack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
```
Esto significa que las actualizaciones se detectarán automáticamente cuando los usuarios llamen a cualquier habilidad; no es necesario ejecutar un comando de actualización específicamente, presencia cero pero cobertura del 100 %.
### El tercer nivel: recordatorio progresivo + actualización automática
Después de detectar una nueva versión, no molestará inmediatamente al usuario, sino que utilizará el mecanismo de repetición (Snooze) para un retroceso progresivo:
* Primer recordatorio: vuelva a mencionarlo después de 24 horas
* Segundo recordatorio: mencione nuevamente después de 48 horas
* Tercera vez y después: mencione nuevamente después de 7 días
* El lanzamiento de la nueva versión restablece el contador de repetición
Los usuarios pueden `gstack-config set auto_upgrade true` habilitar la actualización automática y omitir la confirmación para ejecutarla directamente.
Al realizar la actualización se distinguirán 5 tipos de instalación (git global, git local, suministrada, etc.). La instalación de git utiliza `git fetch + reset`, la instalación proporcionada primero realiza una copia de seguridad y luego reemplaza y restaura desde `.bak` en caso de falla. Después de la actualización, la copia del proyecto suministrada localmente también se sincronizará automáticamente.
**Puntos de los que vale la pena aprender**:
* El modo "detectar en cada llamada" tiene una cobertura extremadamente alta y es imperceptible para los usuarios
* El retroceso gradual evita interrupciones frecuentes
* Diferenciar los tipos de instalación e implementar diferentes estrategias de actualización en lugar de una solución única para todos
* La copia de seguridad y la restauración garantizan que una falla en la actualización no provoque que se cuelgue toda la habilidad.
***
## 3. Sistema de aprendizaje: haga que las habilidades sean más inteligentes cuanto más las use
gstack implementa un **sistema de memoria de sesiones cruzadas** liviano pero efectivo.
### Almacenamiento
Cada proyecto tiene un registro de aprendizaje independiente: `~/.gstack/projects/$SLUG/learnings.jsonl`, que se escribe adicionalmente.
```json
{
"skill": "review",
"type": "pitfall",
"key": "n-plus-one",
"insight": "这个项目的 User model 有 N+1 查询问题,findAll 要加 include",
"confidence": 8,
"source": "observed",
"files": ["src/models/user.ts"],
"ts": "2026-04-01T14:30:00Z"
}
```
### Colección automática
Antes de completar cada habilidad, hay un enlace de "mejora personal operativa", que refleja si se descubrieron fallas inesperadas, desvíos o peculiaridades del proyecto durante la ejecución, y cualquiera se registrará automáticamente en learnings.jsonl. El usuario no requiere activación manual.
### Carga automática
Cada vez que comienza una nueva sesión, el preámbulo cargará las primeras 3 entradas de aprendizaje de alta confianza para inyectar contexto, permitiendo que la nueva sesión herede el conocimiento histórico.
### Caída de la confianza
Las entradas de fuentes `observed` y `inferred` decaen 1 punto cada 30 días. No es necesario limpiar manualmente la base de conocimientos: los conocimientos obsoletos se desvanecen de forma natural y nuevas observaciones ocupan su lugar de forma natural.
### Interfaz de gestión
```text
/learn # 显示最近 20 条
/learn search # 搜索
/learn prune # 检测过期条目(引用的文件已删除)
/learn export # 导出为 markdown 可加入 CLAUDE.md
```
**Puntos de los que vale la pena aprender**:
* El diseño adicional de solo escritura es simple y confiable, y la concurrencia es segura
* La pérdida de confianza es una gestión del envejecimiento del conocimiento que requiere poco mantenimiento y es mucho más eficiente que la limpieza manual.
* Utilice la URL remota de git en lugar de la ruta para identificar el proyecto (a través de `gstack-slug`), que se puede clonar en diferentes ubicaciones y reutilizar.
* Admite consultas entre proyectos, pero está aislado de forma predeterminada.
***
## 4. Inyección de preámbulo: "capa de middleware" de la habilidad
Este es uno de los diseños arquitectónicos más inteligentes de gstack. Cada SKILL.md comparte un código de preámbulo de aproximadamente 220 líneas, que funciona como el middleware de un marco web:
```text
┌─ 更新检测 ──────────────────────────────────┐
│ 会话追踪 (sessions/$PPID) │
│ 配置读取 (proactive, skill_prefix, telemetry)│
│ 学习历史加载 (前 3 条高置信度) │
│ 上下文恢复 (最近的 checkpoint + timeline) │
│ 路由规则检测 │
│ 首次使用引导流程 │
└──────────────────────────────────────────────┘
↓
Skill 特有逻辑
```
El script bash de Preamble genera pares clave-valor (`BRANCH: main`, `PROACTIVE: true`), y luego la plantilla usa condiciones de lenguaje natural para permitir que Claude ajuste su comportamiento en consecuencia:
```text
If PROACTIVE is false, do not invoke skills automatically.
Instead suggest: "I think /skillname might help here -- want me to run it?"
```
Básicamente, se trata de tratar la salida de bash como las "variables de entorno" de Claude: utilizar bash para la detección del tiempo de ejecución y el lenguaje natural para el enrutamiento del comportamiento.
**Puntos de los que vale la pena aprender**: si tiene varias habilidades, la lógica compartida (carga de configuración, recuperación de estado, detección de versión) debe extraerse en un preámbulo unificado en lugar de escribir una copia para cada habilidad.
***
## 5. Arranque progresivo: modo de archivo Sentinel
La experiencia del usuario primerizo de gstack está diseñada con mucho cuidado. Asegúrese de que cada paso de inicio ocurra solo una vez mediante un archivo táctil (archivo centinela):
```text
~/.gstack/.completeness-intro-seen ← "Boil the Lake" 理念介绍
~/.gstack/.telemetry-prompted ← 遥测选择(community/anonymous/off)
~/.gstack/.proactive-prompted ← 主动触发开关
~/.gstack/.routing-prompted ← CLAUDE.md 路由规则写入
~/.gstack/.welcome-seen ← 安装欢迎消息
```
Compruebe si estos archivos existen cada vez que se inicia la habilidad. De lo contrario, muestre los archivos táctiles y de inicio correspondientes. Los pasos que ya han sido vistos nunca volverán a aparecer.
**Puntos de los que vale la pena aprender**: en comparación con mantener el estado de `"onboarding_step": 3` en la configuración, los archivos centinela son más simples y confiables: no se verán afectados por la corrupción del archivo de configuración y cada paso se controla de forma independiente.
***
## 6. Diseño estructural de SKILL.md: arquitectura de tres niveles
Cada SKILL.md sigue una estructura estándar de tres capas:
### Primera capa: YAML Frontmatter
```yaml
---
name: qa
preamble-tier: 3
version: 0.15.1.0
description: |
Systematically QA test a web application...
Use when asked to "qa", "test this site", "find bugs"...
benefits-from: [office-hours]
allowed-tools:
- Bash
- Read
- Write
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "bash ${CLAUDE_SKILL_DIR}/bin/check-careful.sh"
---
```
Campos clave:
* `allowed-tools`: lista blanca de permisos a nivel de herramienta, cada habilidad declara solo las herramientas que necesita
* `benefits-from`: Declarar explícitamente la habilidad previa a la dependencia
* `hooks`: gancho PreToolUse, que puede interceptar antes de que se llame a la herramienta (como la intercepción cuidadosa `rm -rf`)
* `description`: contiene todas las palabras desencadenantes en lenguaje natural
### Capa 2: Preámbulo compartido + Reglas generales
Código de inicio de preámbulo + Definición de voz + recuperación de contexto + principio de integridad + prioridad de búsqueda + protocolo de estado de finalización + reglas de actualización, etc. Todas las habilidades son idénticas y se generan a partir de plantillas.
### La tercera capa: lógica específica de habilidades
Esta es el "alma" de cada habilidad: definición del flujo de trabajo, configuración de roles, inyección de modelos cognitivos, activación de interacciones, etc.
**Puntos de los que vale la pena aprender**: la separación de tres capas permite que cada habilidad se centre únicamente en su propia lógica única, y el marco garantiza las partes compartidas para mantener la coherencia.
***
## 7. Colección rápida de consejos de ingeniería
Después de leer todo SKILL.md, estas son las técnicas de diseño rápido que vale la pena aprender:
### Reglas anti-adulación
El modo de inicio del horario de oficina prohíbe explícitamente el comportamiento común "y confuso" de la IA:
```text
Never say:
- "That's an interesting approach" → take a position instead
- "There are many ways to think about this" → pick one
- "You might want to consider..." → say "This is wrong because..."
- "That could work" → say whether it WILL work
```
### Lista de palabras prohibidas
La sección Voz tiene palabras y frases claramente prohibidas:
* Palabras prohibidas: profundizar, crucial, robusto, integral, matizado, fundamental, paisaje...
* Frases prohibidas: "aquí está el truco", "giro de la trama", "déjame desglosar esto"...
* Formato deshabilitado: em guión (reemplazar con coma/punto)
Estas son palabras comunes con "sabor a IA" en LLM, y el resultado es obviamente más natural después de desactivarlas.
### Inyección de modelo cognitivo
Cada habilidad de revisión inyecta un marco de pensamiento diferente:
* **Revisión del CEO**: 18 modelos cognitivos (la toma de decisiones de puerta unidireccional y bidireccional de Bezos, el pensamiento inverso de Munger, el enfoque y la resta de Jobs...)
* **Revisión en inglés**: 15 patrones de gestión de ingeniería ("aburrido por defecto", intuición del radio de explosión, ley de Conway...)
* **Revisión de diseño**: 12 patrones de cognición de diseño (Jerarquía como servicio, Adoración de las restricciones, Pruebas "¿Me daría cuenta?"...)
Estos modos no permiten que la IA funcione mecánicamente, sino que le proporcionan un **marco de pensamiento**, como si le dieran a un recién llegado inteligente una lista de experiencias de sus predecesores.
### Estándares de especificación
```text
Not "you should test this"
but `bun test test/billing.test.ts`
Not "this might be slow"
but "this queries N+1, ~200ms per page load with 50 items"
Not "there's an issue in the auth flow"
but "auth.ts:47, the token check returns undefined"
```
### Calibración de confianza
La habilidad de revisión requiere que cada descubrimiento vaya acompañado de una puntuación de confianza, y los hallazgos de baja confianza se degradan u ocultan automáticamente:
| Puntuación | Significado | Procesamiento |
| ---------- | ------------------------------------------ | ------------------------------ |
| 9-10 | Lea el código específico y verifique | Visualización normal |
| 7-8 | Coincidencia de patrones de alta confianza | Visualización normal |
| 5-6 | Moderada, posible falsa alarma | Display con instrucciones |
| 3-4 | Confianza baja | Ocultar de los informes |
| 1-2 | Puras conjeturas | Sólo se muestra en el nivel P0 |
### Puerta interactiva
La habilidad del barco define con precisión cuándo detenerse y esperar al usuario y cuándo continuar automáticamente:
```text
Only stop for:
- Tests failing with no obvious fix
- Merge conflicts requiring human judgment
- Unclear which changes to include
Never stop for:
- Normal git operations
- CHANGELOG/VERSION updates
- PR creation
```
**Puntos de los que vale la pena aprender**: una buena habilidad no es "la IA hace todo", sino una definición precisa de la frontera entre humanos y máquinas.
***
## 8. Gestión de estado: el sistema de archivos es una base de datos
Toda la persistencia de gstack se realiza a través del sistema de archivos, almacenado en `~/.gstack/`:
| Camino | Propósito | Formato |
| ------------------------------------- | ------------------------------ | ------------- |
| `config.yaml` | Configuración global | YAML |
| `sessions/$PPID` | sesión activa | tocar archivo |
| `projects/$SLUG/learnings.jsonl` | Registro de aprendizaje | JSONL |
| `projects/$SLUG/timeline.jsonl` | Cronología de habilidades | JSONL |
| `projects/$SLUG/checkpoints/*.md` | Punto de control | Rebaja |
| `projects/$SLUG/health-history.jsonl` | Historial de controles médicos | JSONL |
| `analytics/skill-usage.jsonl` | Usando telemetría | JSONL |
| `last-update-check` | caché de versión | texto plano |
Casi todos los datos de series temporales se escriben de forma anexa utilizando **JSONL** (un objeto JSON por fila). Esta elección es inteligente:
* Se agregó seguridad de concurrencia natural de escritura.
* No se requieren dependencias de bases de datos
* Puede utilizar `grep` / `jq` para realizar consultas directamente
* Corrupto hasta faltar la última línea.
***
## 9. Modo de integración entre habilidades
### Producto de transferencia de archivos
Transferir productos de trabajo entre habilidades a través del sistema de archivos:
```text
/office-hours → design doc → /plan-ceo-review 读取
/plan-ceo-review → ceo-plans/*.md → /autoplan 读取
/review → reviews.jsonl → /ship 读取并展示 Dashboard
/qa → qa-reports/ → /retro 读取
```
### Review Readiness Dashboard
La habilidad del barco dice `reviews.jsonl`, lo que muestra el estado de revisión de habilidades cruzadas antes de publicar:
```text
| Review | Runs | Last Run | Status | Required |
| Eng Review | 1 | 2026-03-16 | CLEAR | YES |
| CEO Review | 0 | — | — | no |
| Design Review | 0 | — | — | no |
```
### Sugerencias previas a la dependencia
Cuando plan-ceo-review detecta que no hay ningún documento de diseño, recomendará activamente ejecutar `/office-hours` primero:
```text
"No design doc found. /office-hours produces a structured problem statement...
Takes about 10 minutes."
Options: A) Run /office-hours now B) Skip
```
### Usar predicción de secuencia
Context Recovery analizará la secuencia reciente de uso de habilidades y predecirá el siguiente paso:
```text
If pattern repeats (e.g., review → ship → review),
suggest: "Based on your recent pattern, you probably want /ship."
```
***
## 10. Otros diseños destacables
### Sistema de gancho
Las tres habilidades: cuidado, congelación y guardia usan ganchos `PreToolUse`; este es el único mecanismo que puede interceptar antes de que se llame a la herramienta:
* **cuidado**: Interceptar Bash, marcar `rm -rf`, `DROP TABLE`, `git push --force`
* **congelar**: intercepta Editar/Escribir y comprueba si la ruta está dentro del rango permitido
* **guardia**: combina los dos anteriores
### Adaptación multiplataforma
El mismo conjunto de plantillas genera archivos de habilidades para diferentes plataformas a través del parámetro `--host`:
```bash
bun run gen:skill-docs --host claude # Claude Code 格式
bun run gen:skill-docs --host codex # OpenAI Codex 格式
bun run gen:skill-docs --host kiro # AWS Kiro 格式
bun run gen:skill-docs --host factory # Factory Droid 格式
```
El camino y el tema inicial se adaptan automáticamente y la lógica de la habilidad permanece sin cambios.
### Protocolo de estado de finalización
Se debe generar un estado de finalización estandarizado al final de cada habilidad:
```text
DONE — 全部完成,提供证据
DONE_WITH_CONCERNS — 完成但有顾虑
BLOCKED — 无法继续
NEEDS_CONTEXT — 需要更多信息
```
### Tres reglas de actualización fallidas
```text
If you have attempted a task 3 times without success, STOP and escalate.
```
Evite que la IA se quede atrapada en un ciclo de reintentos infinito.
### Selección de prueba basada en diferencias
Las pruebas E2E cuestan alrededor de $4 cada una (requiere iniciar el agente Claude), por lo que gstack declara los archivos fuente de los que depende cada prueba hasta `touchfiles.ts`, y solo ejecuta las pruebas afectadas de acuerdo con `git diff`:
```typescript
// test/helpers/touchfiles.ts
{
"qa-workflow": ["qa/SKILL.md.tmpl", "browse/src/server.ts"],
"ship-flow": ["ship/SKILL.md.tmpl", "scripts/resolvers/preamble.ts"]
}
```
***
## Resumen: principios de diseño que puedes aprender
De la práctica de ingeniería de gstack, extraje los siguientes principios de diseño que son más valiosos para los desarrolladores de Skills:
1. **Generación de plantillas > Sincronización manual**: el contenido compartido entre habilidades se genera automáticamente usando plantillas + pasos de compilación, no copiar y pegar
2. **Detección pasiva > Detección activa**: la detección de actualización está integrada en cada llamada de habilidad, el usuario no lo sabe pero la tasa de cobertura es del 100 %.
3. **Agregar registro> Base de datos compleja**: el sistema de archivos JSONL + puede cubrir la mayoría de las necesidades de persistencia, es simple y confiable
4. **Arranque progresivo > Una configuración**: utilice archivos centinela para controlar los pasos de arranque, cada uno de los cuales aparece solo una vez.
5. **Puerta precisa > Totalmente automática**: defina claramente los límites entre "detener y esperar al usuario" y "continuar automáticamente"
6. **Cuantificación de la confianza > Juicio difuso**: cada juicio de IA viene con una puntuación de confianza, y la confianza baja se degrada automáticamente.
7. **Decaimiento del tiempo > Limpieza manual**: la confianza en los registros de aprendizaje decae con el tiempo y el conocimiento obsoleto se desvanece naturalmente
8. **LISTA DE PALABRAS PROHIBIDAS > GUÍA DE ESTILO**: Una lista directa de palabras prohibidas es mucho más efectiva que "por favor use un tono natural".
***
**Lectura relacionada**:
* [gstack Concepts](/es/docs/notes/gstack/concept) — ¿Qué es gstack y qué problemas resuelve?
* [Capítulo práctico de gstack](/es/docs/notes/gstack/practice) — Flujo de trabajo completo desde la instalación hasta la ejecución
* [gstack Front-end Skill](/es/docs/notes/gstack/frontend-skills) — Panorama de habilidades de diseño de front-end/UI y flujo de trabajo recomendado
* [Concepto de Habilidades de Claude](/es/docs/notes/claude-skills/concept) — Comprender el mecanismo subyacente de las Habilidades
# Diagnose 与 Triage:先建立反馈回路,再决定交给谁
## 为什么把两个 Skill 放在一起讲
`/diagnose` 和 `/triage` 在 README 里是两个独立 skill,但它们解决的是同一个工程问题的两半:
* `/diagnose` 关心:**这个 bug 到底是什么、怎么复现、怎么证明修好了**
* `/triage` 关心:**这个 issue 现在该等信息、给 agent、给人,还是不做**
一个负责事实,一个负责流程。真实项目里这两个经常连在一起:先 triage 一个 bug issue,发现信息不够就 `needs-info`;信息够了就用 diagnose 建反馈回路;复现清楚后再决定是 `ready-for-agent` 还是 `ready-for-human`。
## /diagnose 的核心:反馈回路就是全部
`/diagnose` 最值得记住的一句话是:**先建立一个 agent 能运行的 pass/fail 信号**。
Matt 把诊断拆成 6 个阶段:
| 阶段 | 目标 |
| --------------------- | ------------------- |
| Build a feedback loop | 搭一个快速、确定、可反复运行的失败信号 |
| Reproduce | 让这个信号复现用户描述的同一个 bug |
| Hypothesise | 列 3-5 个可证伪假设 |
| Instrument | 用最少探针验证假设 |
| Fix + regression test | 在正确测试面写回归测试,再修 |
| Cleanup + post-mortem | 清理临时探针,记录真实根因 |
这和很多人调 bug 的顺序相反。普通调试常见流程是:看代码、猜原因、改一改、刷新页面。Matt 反过来:先把 bug 变成一个可重复机器信号,再谈假设。
## 什么算好反馈回路
`/diagnose` 给了一组优先级,从最好到最兜底:
| 回路 | 适合场景 |
| ----------------------------- | --------------------- |
| 失败测试 | 有合适测试面,能直接表达 bug |
| curl / HTTP script | API bug、服务端行为可用请求复现 |
| CLI + fixture | 命令行工具、解析器、转换器 |
| Headless browser | UI bug、控制台错误、网络行为 |
| Replay captured trace | 线上真实 payload、事件流、日志链路 |
| Throwaway harness | 只启动系统一小块,隔离复杂依赖 |
| Property / fuzz loop | 偶现错误输出,需要提高触发率 |
| Bisection / differential loop | 某版本后坏了,需要二分或对比旧版 |
| HITL script | 只能人手点时,也要让人按脚本提供稳定输出 |
这里有一个很硬的判断:**没有回路,不要进入假设阶段**。因为没有信号,所有分析都会变成「看起来像」。
## 非确定性 bug 怎么办
`/diagnose` 对偶发 bug 的态度也很实用:目标不是一开始就 100% 复现,而是先把复现率提高到可调试。
比如:
* 循环触发 100 次
* 并发触发
* 注入 sleep 拉大竞态窗口
* 固定随机种子或时间
* 缩小环境变量和外部依赖
1% 的偶发 bug 很难调;50% 的偶发 bug 就已经是可调试对象。这个思路对前端异步、消息队列、支付回调、流式输出都很有用。
## 假设必须可证伪
Matt 要求在动手验证前先列 3-5 个假设,并且每个假设都要写出预测:
```text
如果 X 是原因,那么改变 Y 后 bug 应该消失;
或者观察 Z 时应该出现某个特征。
```
这会防止 agent 被第一个看起来合理的解释锁死。更重要的是,它让你能判断某个实验到底有没有信息量。
一个坏假设:
```text
可能是缓存问题。
```
一个可证伪假设:
```text
如果是浏览器缓存导致旧脚本执行,那么禁用缓存并强刷后,console 里的旧 bundle hash 应该消失,按钮点击事件也应该恢复。
```
后者才值得验证。
## 修复阶段最容易犯的错
`/diagnose` 要求:如果有正确测试面,就先把最小复现转成失败测试,再修代码。
关键是「正确测试面」。不是随便补一个 unit test 就算回归测试。正确测试面必须覆盖真实 bug 模式:
* bug 是多个调用者组合触发的,就不能只测单个函数
* bug 是真实 payload 结构触发的,就不能只测一个手写 toy object
* bug 是浏览器事件顺序触发的,就不能只测纯函数
如果找不到正确测试面,这本身就是结论:代码结构没有给你留下可锁定 bug 的地方。修完之后应该把这个信息交给 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture)。
## /triage 的核心:issue 是状态机
`/triage` 不是让 AI 随便帮你「看一下 issue」。它把 issue 当作一个小状态机。
每个 issue 应该同时有:
* 一个 category:`bug` 或 `enhancement`
* 一个 state:`needs-triage`、`needs-info`、`ready-for-agent`、`ready-for-human`、`wontfix`
这套状态的价值在于让维护者可以快速回答:
* 哪些还没人看?
* 哪些等报告者补信息?
* 哪些已经清楚到可以交给 AFK agent?
* 哪些必须人类自己做?
* 哪些应该关掉,并把原因沉淀下来?
## ready-for-agent 的标准
`ready-for-agent` 是这套流程里最关键的状态。它不是「这个任务可以让 AI 试试」,而是:
> 任务已经清楚到一个不在场的 agent 可以独立领取、实现、验证。
这通常意味着 issue 里至少有:
* 背景和问题陈述
* 相关代码路径或模块
* 明确的验收标准
* 已知约束
* 如果是 bug,最好有复现方式
* 不需要额外产品/设计判断
如果缺这些,应该是 `needs-info` 或 `ready-for-human`,而不是勉强丢给 agent。
## needs-info 要问具体问题
`/triage` 给 `needs-info` 的模板很朴素,但重点是问题必须具体:
```markdown
## Triage Notes
**What we've established so far:**
- ...
**What we still need from you (@reporter):**
- ...
```
坏问题:
```text
请提供更多信息。
```
好问题:
```text
请提供触发问题的浏览器版本、出错页面 URL、点击顺序,以及 Network 面板里 `/api/orders/:id` 的 response body。
```
AI 很容易写礼貌废话,这个 skill 强迫它把「我们已经知道什么」和「还缺什么」分开。
## wontfix 也要沉淀
`/triage` 对 enhancement 的 `wontfix` 有一个有趣设计:不要只是关 issue,而是把拒绝理由写进 `.out-of-scope/` 知识库,再在评论里链接。
这样下次类似需求出现时,AI 不会重新展开同一场讨论。它可以先读 `.out-of-scope/`,提醒维护者:「这个方向之前拒绝过,理由是 X。」
这和 ADR 的精神很像:不是记录所有决定,只记录未来会让人疑惑、并且会反复出现的决定。
## 两者怎么配合
一个典型 bug issue 可以这样走:
1. `/triage` 读取 issue、评论、标签和相关代码
2. 它判断这是 `bug + needs-triage`
3. 先尝试复现;如果步骤不足,转 `needs-info`
4. 信息足够后,启动 `/diagnose`
5. `/diagnose` 建立复现回路,列假设,定位根因
6. 如果修复路径清楚、测试面明确,issue 变 `ready-for-agent`
7. 如果需要产品判断、外部权限、人工验证,issue 变 `ready-for-human`
8. 修完后把根因和回归测试写回 issue 或 PR
这套流程的关键不是「AI 自动修 bug」,而是让 issue 从模糊描述变成可执行工作包。
## 我的使用建议
如果你只记一条:
> `/diagnose` 先问「我怎么证明它坏了」;`/triage` 先问「它现在该处在哪个状态」。
这两个问题能挡住大量低质量 AI 编程:
* 没复现就修
* 没验收就开工
* 没根因就重构
* 没信息就甩给 agent
Matt 这两个 skill 并不花哨,但很像真实团队里资深工程师会做的事:先把事实收束,再推进流程。
## 参考资源
下一篇:[TDD:用红绿重构强迫 AI 走小步](/docs/notes/matt-pocock-skills/tdd)。
# Grill Me: deja que la IA te haga 50 preguntas antes de escribir código
## Modo de fallo: “La IA no hizo lo que yo quería”
El primer modo de falla del que Matt habló en su discurso es: crees que los requisitos en tu mente son muy claros y dejas que la IA los escriba; ese no es el caso en absoluto.
> "I would run it, and I would try not to look at the code, but I would look at the code, and I realized I would get worse code. I did it again, I got even worse code... I did it again, kept running the compiler, and I would just end up with garbage."
Mucha gente está familiarizada con este sentimiento: si dices "Agrega un inicio de sesión para mí", la IA no te preguntará "¿Quieres recordar el dispositivo?". "¿Cuántas veces no has podido bloquear la cuenta?" "¿Cuánto tiempo tarda en expirar la sesión?" Presenta directamente un plan que considera razonable. Cuando lo revisas, se han escrito 500 líneas: dos horas de reelaboración.
## ¿Por qué sucede esto? El concepto de diseño se desvía
Matt cita el **concepto de diseño** (concepto de diseño) de Frederick P. Brooks en “El diseño del diseño”:
> Cuando varias personas colaboran para diseñar algo, algo se crea entre ustedes: flota en su mente, una "teoría sobre esto" invisible. No es un activo, no es un activo metido en un archivo de rebajas, es un consenso invisible.
La IA escribe código tan pronto como aparece, lo que significa que no comparte el mismo concepto de diseño contigo en absoluto. Lo que está mal al escribir código no es la sintaxis, sino la premisa.
Para solucionar este problema, primero debe alinear el concepto de diseño antes de comenzar. La herramienta que proporcionó Brooks se llama **árbol de diseño**: divide una decisión en varias ramas y luego divide cada rama. No puede omitir las decisiones ascendentes y tomar decisiones descendentes directamente; de lo contrario, todo tendrá que rehacerse una vez que el ascendente cambie al descendente.
## La habilidad de Matt texto completo
La implementación de Matt de esta teoría en [`mattpocock/skills`](https://github.com/mattpocock/skills) es `productivity/grill-me/SKILL.md`, y el archivo completo más el texto frontal tiene menos de 15 líneas:
```markdown
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
---
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
Interpretar frase por frase:
* **"entrevístame sin descanso"** - La palabra clave es *sin descanso* (no soltarme). De forma predeterminada, LLM tiende a hacer 1 o 2 preguntas y luego sentirse "casi" y comenzar a actuar. Esta palabra suprime por la fuerza esta tendencia.
* **"caminar por cada rama del árbol de diseño"** —— Concepto de árbol de diseño de Brooks. Obligue a Claude a tratar sus requisitos como un árbol, resolviendo primero los ascendentes y luego los descendentes. Si dice "Iniciar sesión", primero le preguntará "Método de autenticación" (raíz del árbol) y luego expandirá "Cómo administrar la sesión"/"Cómo guardar el token" (subnodo) según su respuesta.
* **"resolver dependencias entre decisiones una por una"** - Prohibir explícitamente las preguntas empaquetadas. A menudo existen dependencias entre las decisiones (si elige SSO, el flujo descendente no necesita problemas de política de contraseñas), primero asegúrese de que el flujo ascendente pueda eliminar muchos problemas descendentes.
* **"para cada pregunta, proporcione la respuesta recomendada"** - puntos de bonificación clave. La IA no sólo hace preguntas, sino que también recomienda respuestas. Simplemente asiente o no y ahorra un 80 % del tiempo de escritura.
* **"haz las preguntas una a la vez"** - Evita que la IA te dé 10 preguntas a la vez.
* **"si se puede responder una pregunta explorando el código base, explore el código base en su lugar"** - Si es un hecho que ya existe en el proyecto (como "Qué marco de prueba se utiliza en el proyecto"), deje que Claude lo vea él mismo, no le pregunte.
7 líneas, pero cada oración corresponde a un sesgo de comportamiento específico de LLM.
## Cómo instalar y usar
**Instalación**:
```bash
npx skills@latest add mattpocock/skills
```
Marque `grill-me` y `setup-matt-pocock-skills` (grill-me no depende de este último, pero otras habilidades dependen de él, por lo que se recomienda instalarlos juntos).
**Llamar**: Ingrese `/grill-me` en el cuadro de diálogo Claude Code.
**Proceso típico**:
1. Describes lo que quieres hacer, lo cual puede ser muy vago ("Quiero agregar una función de comentarios a mi blog")
2. Ingrese `/grill-me`
3. Claude comenzó a hacer preguntas una por una y recomendó respuestas para cada pregunta.
4. Respondes cada pregunta una por una (asiente/no/correcta)
5. Generalmente, se llega a un consenso después de 20 a 50 preguntas y Claude le dará un resumen.
6. El resumen puede enviarse directamente a [`/to-prd`](/es/docs/notes/matt-pocock-skills/to-prd-and-issues) para convertirse en PRD, o pasarse directamente a [`/tdd`](/es/docs/notes/matt-pocock-skills/tdd) para comenzar a escribir.
## Caso real: ¿Cuánto cuesta una función de editor de vídeo?
Matt dio algunos números específicos en ["5 habilidades de agente que uso todos los días"](https://www.aihero.dev/5-agent-skills-i-use-every-day):
* **Nueva función de edición de vídeo** - 16 preguntas para llegar a un consenso
* **Funciones complejas** - 30\~50 preguntas
* **EXTREMADAMENTE COMPLEJO** - 100 preguntas, sesión de hasta 45 minutos
Pregunta de ejemplo (restaurada del video/publicación de blog de Matt):
* "Should video clips be reorderable, or only added/removed in sequence?"
* "When a clip is deleted, do we keep its source file, or delete the file too?"
* "Does the editor need undo/redo? How many steps deep?"
* "Should we render previews in the browser, or rely on a backend service?"
Ninguno de estos problemas era técnico: todos eran decisiones de producto. Pero **cada decisión determina la forma de cientos de líneas de código**. Si omite estas preguntas y deja que la IA las escriba directamente, generará un conjunto de respuestas por sí solo y usted volverá a rechazarlas una por una después de escribirlas.
## Diferencias entre el modo Plan y el modo Plan integrado de Claude Code
Claude Code viene con `plan mode` (presione Shift+Tab para ingresar). En la superficie, se parece a grill-me: discuta primero antes de actuar. Pero Matt dijo directamente en su discurso que prefiere grill-me:
Diferencias específicas:
| Dimensiones | Modo de planificación | /asarme |
| --------------------------------------- | ------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
| Objetivo predeterminado | Producir un plan ejecutable lo antes posible | Primero llegue a un consenso, el plan es un subproducto |
| Número de preguntas | 0\~5 | 20\~100 |
| Formato de pregunta | Pregunta un párrafo a la vez | Haga una pregunta a la vez |
| Quieres dar una respuesta recomendada | No | Sí |
| Ya sea para explorar la base del código | De vez en cuando | Activamente (instrucciones explícitas) |
| Adecuado para escenarios | Ya lo pensé claramente y quiero confirmar el plan de implementación | Aún no he pensado con claridad, es necesario que me obliguen a pensar con claridad |
La mayor diferencia práctica es "**urgente o no**". El modo Plan tiene prisa por empezar, grill-me no tiene prisa: considera "pensar con claridad" como la tarea principal y no como el prólogo.
## Uso avanzado
### 1. Escenarios sin programación
`grill-me` no vincula código y también se puede utilizar para conversaciones puramente de toma de decisiones sobre productos. Matt lo usa él mismo:
* Diseño del programa del curso.
* Redacción de artículos
* Documentos de comunicación interna.
Siempre que tengas una idea vaga en mente y quieras que te obliguen a pensar en ella, puedes usarla.
### 2. Cooperar con [`/to-prd`](/es/docs/notes/matt-pocock-skills/to-prd-and-issues)
Una vez finalizada la sesión de interrogatorio, simplemente diga `/to-prd` y Claude condensará toda la conversación en un PRD estructurado (que incluye historia de usuario, división de módulos y estrategia de prueba) y lo enviará a su rastreador de problemas. **Punto clave: no borre el contexto en el medio** - to-prd se extrae directamente del contexto de la conversación y no le volverá a preguntar.
### 3. Cooperar con [`/grill-with-docs`](/es/docs/notes/matt-pocock-skills/grill-with-docs)
Si el proyecto ya tiene `CONTEXT.md` (lenguaje de dominio) y `docs/adr/` (decisiones arquitectónicas), use `/grill-with-docs` en lugar de `/grill-me`. **Actualizará CONTEXT.md** sincrónicamente mientras lo torturan: se toman decisiones mientras se actualizan los documentos y ya no existe el problema de que "los documentos estén siempre desactualizados".
### 4. Personaliza la profundidad de la pregunta
Si tiene poco tiempo, puede agregar una oración directamente después de `/grill-me`: "Limitar a 10 preguntas, centrarse solo en decisiones arquitectónicas". Convergirá según tu límite. Pero Matt no lo recomienda: cree que "hacer más preguntas" es exactamente el valor de esta habilidad, y eliminarla es casi como un modo de plan.
## Notas
**Será molesto la primera vez que lo ejecutes**. Las personas que están acostumbradas a "generar 500 líneas en una oración" sentirán que es una pérdida de tiempo que la IA les pregunte 30 veces por primera vez. El consejo de Matt es aferrarse a las primeras 5 preguntas; las primeras 5 preguntas a menudo revelan cosas en las que ni siquiera habías pensado. Una vez que superes ese umbral, serás adicto.
**No apto para tareas extremadamente pequeñas**. Cambie un error tipográfico, agregue un archivo console.log; no use grill-me. Es adecuado para "hacer algo nuevo" o "cambiar algo viejo con efectos secundarios".
**A veces la IA solicitará detalles técnicos**. Si no te importa y quieres dejar que juzgue, simplemente responde "tu llamada" y "tú decides", y lo aceptará y continuará.
## ¿Por qué es popular esta habilidad?
`/grill-me` es la habilidad capturada y reenviada con más frecuencia entre las habilidades de Matt. La razón no es complicada:
1. **Extremadamente minimalista**: 7 líneas de rebajas, solo copie y pegue
2. **Efecto inmediato**: Puedes sentir el cambio en la “densidad de problemas” de la IA durante la primera ejecución.
3. **Portátil**: No depende de Claude Code, se pueden usar Codex, Cursor y Aider
4. **Viene con comportamiento anti-LLM predeterminado**: cada palabra es una desviación anti-LLM, con una estética de alta ingeniería
Su éxito también se ha convertido en el mejor argumento para "**la habilidad no tiene por qué ser larga**".
## Recursos de referencia
Artículo siguiente: [Grill With Docs: Mantenimiento del lenguaje del proyecto y ADR](/es/docs/notes/matt-pocock-skills/grill-with-docs) - una versión avanzada de grill-me, para proyectos con complejidad de dominio.
# Grill With Docs: Equipar la IA con memoria de proyecto utilizando lenguaje de dominio y ADR
## Modo de error: "La IA es demasiado detallada"
El segundo modo de fracaso en el discurso de Matt:
> La IA expresa algo simple con mucha palabrería. Es como hablarte dos idiomas.
Esto no tiene nada que ver con la cantidad de código, es **ubicación incorrecta del vocabulario**. La IA utilizará términos generales ("elemento", "datos", "controlador") de forma predeterminada, y los términos reales en el proyecto que tiene en mente pueden ser "Curso", "Versión borrador", "Lección fantasma". La IA no sabe que estas palabras tienen significados específicos en su proyecto, por lo que creará un montón de palabras nuevas sinónimas a su alrededor. El resultado es:
* Proceso de pensamiento largo (evite las palabras exclusivas)
* La implementación no está alineada con el diseño que tienes en mente (porque no están en el mismo espacio semántico)
* No reutilizable entre sesiones (el contexto debe restablecerse para cada conversación)
## Teoría clásica: el lenguaje ubicuo de DDD
Matt citó el "Diseño basado en dominios" de Eric Evans. Este libro fue publicado en 2003 y propuso el concepto de **lenguaje ubicuo**:
> Utilice el mismo conjunto de términos para conectar expertos en el dominio, desarrolladores y códigos. Una palabra en discusiones sobre productos, comentarios de código, nombres de variables, documentación: debe significar lo mismo.
El objetivo de DDD es hacer que el código parezca el cerebro de un experto en el dominio. En la era de la IA, hay un nuevo rol: **LLM también debe estar en este idioma**. El LLM no está en la reunión de pie, no puede ver la reunión de requisitos del producto y no puede entender la jerga de su grupo; solo puede aprender de los documentos que usted le proporciona.
Matt convirtió esto en una habilidad: escanear la base del código para extraer términos, generar un archivo de rebajas `CONTEXT.md` y luego alinearlo tanto con los humanos como con la IA.
## La evolución de las habilidades: del lenguaje ubicuo a la parrilla con documentos
La primera habilidad se llamó `ubiquitous-language`; solo hacía una cosa: escanear la base del código para generar un glosario. Pero Matt descubrió más tarde que simplemente generar un documento no era suficiente:
* **La documentación estará desactualizada**: se genera hoy, el código se cambia mañana y el glosario no se ha actualizado.
* **La gente no tomará la iniciativa de mirarlo**: Estaría muerto si lo pones ahí
Lo refactorizó en `grill-with-docs`, que combinaba tres cosas:
1. **Requisitos de tortura** (todas las habilidades heredadas de grill-me)
2. **Desafía el glosario existente**: ¿La palabra que dijiste no coincide con lo que está escrito en CONTEXT.md? Señalalo inmediatamente
3. **Actualizar documentos simultáneamente al tomar decisiones**: Las nuevas conclusiones alcanzadas durante el proceso de tortura se escriben en línea en CONTEXT.md o se crea un nuevo ADR.
Se trata de un cambio de paradigma de "generar documentos estáticos" a "diálogo es mantener documentos".
## Texto completo de la habilidad
Estructura central de `engineering/grill-with-docs/SKILL.md`:
```markdown
---
name: grill-with-docs
description: Grilling session that challenges your plan against the
existing domain model, sharpens terminology, and updates
documentation (CONTEXT.md, ADRs) inline as decisions crystallise.
---
Interview me relentlessly about every aspect of this plan until we
reach a shared understanding. Walk down each branch of the design
tree, resolving dependencies between decisions one-by-one. For each
question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question
before continuing.
If a question can be answered by exploring the codebase, explore the
codebase instead.
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
If a CONTEXT-MAP.md exists at the root, the repo has multiple contexts.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in
CONTEXT.md, call it out immediately.
"Your glossary defines 'cancellation' as X, but you seem to mean Y —
which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise
canonical term.
"You're saying 'account' — do you mean the Customer or the User?
Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with
specific scenarios.
### Cross-reference with code
When the user states how something works, check whether the code
agrees. If you find a contradiction, surface it.
### Update CONTEXT.md inline
When a term is resolved, update CONTEXT.md right there. Don't batch
these up — capture them as they happen.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. Hard to reverse
2. Surprising without context
3. The result of a real trade-off
```
## ¿Cómo se ve el CONTEXT.md real?
El repositorio [`course-video-manager`](https://github.com/mattpocock/course-video-manager/blob/main/CONTEXT.md) de Matt ofrece un ejemplo completo de CONTEXT.md. Escojamos algunos términos para tener una idea:
| Terminología | Definición |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| **Course** | The primary domain entity: a structured collection of versions, sections, lessons, and videos |
| **Draft Version** | The single mutable CourseVersion that is currently being edited; always the latest by `createdAt` |
| **Published Version** | An immutable CourseVersion with a name and description, created by the Publish flow |
| **Ghost Lesson** | A lesson that exists in the database but not yet on the file system (`fsStatus = "ghost"`) |
| **Export Hash** | A SHA256 hash derived from a video's clip filenames, timestamps, clip order |
| **Unexported Video** | A video whose current Export Hash does not match any file on disk; blocks publishing |
| **Materialization Cascade** | The chain reaction when materializing a lesson inside a ghost course |
| **Clip** | A timestamped segment of source footage within a video |
| **Fractional Index** | A string-based ordering value that allows inserting items between existing items |
| **Purge** | The deliberate deletion of an Exported Video's `.mp4` file from disk |
Tenga en cuenta algunas cosas:
1. **Cada término es un gerundio o un nombre propio**, no una frase descriptiva como “estado del pedido”
2. **Cada definición hace referencia a otros términos** (Curso → Versión → Lección → Video) formando una red de ontologías
3. **El campo de código aparece directamente** (`fsStatus = "ghost"`) - Mapeo 1:1 de documentos y código
4. **Incluya descripciones de las decisiones** ("bloquea la publicación", "reacción en cadena"): no solo sustantivos, sino reglas
Al escribir código, cuando AI vea este documento, utilizará "Lección fantasma" en lugar de "lección sin archivo". El código, las conversaciones y los mensajes de confirmación están todos unificados.
## ADR: cuando se crea
Hay una restricción importante en la habilidad:
> Only offer to create an ADR when all three are true:
>
> 1. **Hard to reverse** — the cost of changing your mind later is meaningful
> 2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
> 3. **The result of a real trade-off** — there were genuine alternatives
La cultura de ADR (Architecture Decision Record) se originó en el blog de Michael Nygard en 2011, pero muchos equipos la utilizan para escribir ADR para todas las decisiones: 18 de cada 20 ADR están ejecutando cuentas. Matt Esta triangulación es una gran herramienta: **ADR solo vale la pena si se cumplen tres condiciones simultáneamente**. De lo contrario, deje que se digiera en CONTEXT.md y se digiera en el código.
## Cómo instalar y usar
**Condiciones previas**: primero ejecute `/setup-matt-pocock-skills` (le preguntará dónde colocar CONTEXT.md y dónde colocar el directorio ADR).
**Llamar**: `/grill-with-docs`
**Proceso típico**:
1. Describe lo que quieres hacer
2. `/grill-with-docs`
3. Claude **escanea** CONTEXT.md y docs/adr/ primero, cargando los términos y decisiones existentes en contexto
4. Iniciar la tortura, durante el proceso:
* La palabra que usaste entra en conflicto con CONTEXT.md → Indícalo en el acto
* Usó palabras vagas (por ejemplo, "Cuenta" podría ser Cliente o Usuario) → Le permite elegir una de las dos y colocarla en el documento.
* El comportamiento que mencionaste es inconsistente con el código existente → señala el conflicto
5. **Actualice CONTEXT.md** sincrónicamente cuando se tome la decisión (sin retrasos ni procesamiento por lotes)
6. Decisiones clave irreversibles → Preguntar si se debe generar un ADR
Si aún no hay CONTEXT.md y docs/adr/ en el proyecto, se creará de forma diferida: los archivos no se generarán hasta que sea necesario escribir el primer término y crear el primer ADR. No le daremos una plantilla en blanco desde el principio.
## Proyectos de contexto múltiple (CONTEXT-MAP.md)
Si el proyecto es demasiado grande para caber en un CONTEXT.md (por ejemplo, pedidos y facturación son dos contextos delimitados independientes), puede colocar `CONTEXT-MAP.md` en el directorio raíz como directorio general:
```
/
├── CONTEXT-MAP.md ← 总目录
├── docs/adr/ ← 系统级决策
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← 模块级决策
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
`/grill-with-docs` reconocerá automáticamente la existencia de CONTEXT-MAP.md y saltará al subdirectorio correspondiente. Esta es una implementación directa del concepto de **contexto limitado** en DDD: las "órdenes" en cada contexto pueden tener significados diferentes y se mantienen por separado para evitar la contaminación.
La diferencia entre ## y grill-me
| dimensiones | /asarme | /parrilla-con-docs |
| -------------------------------------------- | ------------------------------------ | ------------------------------------------------------------------------------------ |
| Capacidad de cuestionamiento | ✅ | ✅ (heredar todo) |
| Verificación de la terminología del proyecto | ❌ | ✅ |
| Actualizaciones en tiempo real CONTEXT.md | ❌ | ✅ |
| Sentencia de activación de ADR | ❌ | ✅ |
| Etapa aplicable | Primeras ideas, proyectos personales | Proyectos reales con complejidad de dominio |
| Costo inicial | 0 | Requiere configuración + Tener/estar dispuesto a construir CONTEXT.md en el proyecto |
Juicio simple y crudo:
* **Guiones personales, redacción de artículos, impartición de cursos** → `/grill-me`
* **Proyectos reales que requieren mantenimiento a largo plazo** → `/grill-with-docs`
## Un beneficio contrario a la intuición: dejar que la IA aprenda a "callarse"
CONTEXT.md no es sólo para IA, es para futuras sesiones de IA. Cada vez que comienza una nueva conversación, Claude puede ingresar instantáneamente al contexto del proyecto leyendo CONTEXT.md, lo que le ahorra una larga incorporación.
Lo que es aún más sutil es: Matt dijo en su discurso que después de agregar CONTEXT.md podía ver en el *rastro de pensamiento* de AI——
> "Permite que la IA piense de una manera menos detallada".
¿Por qué? Porque sin CONTEXT.md, la IA tiene que definir constantemente sus propios términos cuando piensa: "el usuario, es decir, la persona que ordenó el artículo, en lo sucesivo denominado...". Con CONTEXT.md dice directamente "Cliente", la cadena de pensamiento es mucho más corta y la respuesta es más rápida.
**La economía simbólica de LLM determina: acortar el camino del pensamiento = producción más rápida, más precisa y más barata**. CONTEXT.md es la palanca oculta para esta eficiencia.
## Notas
**CONTEXT.md cambiará significativamente durante la primera ejecución**. Si ya hay un CONTEXT.md escrito a mano en el proyecto, primero git stash o déjelo funcionar en seco antes de ejecutarlo (puede agregar una oración en el mensaje "Enumere el contenido que se cambiará primero, no escriba el archivo directamente").
**La moderación ADR es moderación real**. No te emociones y haz que cada decisión genere un ADR, tus docs/adr/ estarán llenos de basura en 6 meses. Los tres criterios de Matt deben seguirse estrictamente.
**CONTEXT.md No ponga detalles de implementación**. Hay un dicho en Skill: "No combine CONTEXT.md con los detalles de implementación. Incluya sólo términos que sean significativos para los expertos en el dominio". Escribir "PostgreSQL" en CONTEXT.md es incorrecto: a los expertos en dominios no les importa la selección de la base de datos, eso es una cuestión de ADR.
## Recursos de referencia
Artículo siguiente: [a-PRD + a-Issues: Del diálogo al ticket ejecutable](/es/docs/notes/matt-pocock-skills/to-prd-and-issues)——Después de la tortura, cómo solidificar el diálogo en una unidad de trabajo ejecutable.
# Mejorar la arquitectura de la base de código: reestructurar módulos poco profundos en módulos profundos
## Modo de falla: "IA deambulando en una base de código incorrecta"
El cuarto modo de fracaso en la charla de Matt es una metáfora pictórica:
> "Los módulos superficiales en una base de código se ven así: tienes un montón de pequeñas manchas y la IA tiene que revisar un montón de módulos y comprender todas las dependencias antes de poder corregirlas".
> "AI is really good at creating codebases like this. So you'll have a situation where AI doesn't understand what your code is doing. It will attempt to explore the code, but because it's poorly laid out, filled with shallow modules, it doesn't get to the right module in time, or doesn't understand all the dependencies."
Este es un círculo vicioso exclusivo de la programación de IA:
```
AI 写代码倾向于产生 shallow 模块(小、多、互相依赖)
↓
代码库变得 shallow
↓
下次 AI 进来探索更难,更容易写错
↓
更多 shallow 模块被加进去
↓
代码库越来越烂,AI 越来越无能
```
Para romper este ciclo, se debe realizar periódicamente una refactorización inversa manual, fusionando módulos poco profundos en módulos profundos. Eso es exactamente lo que hace `/improve-codebase-architecture`.
## Teoría clásica: los módulos profundos de Ousterhout
John Ousterhout es profesor de informática en Stanford (y autor de los artículos sobre lenguaje Tcl y Raft). Su libro de 2018 "A Philosophy of Software Design" propone una regla simple pero poderosa:
**La "profundidad" del módulo = la complejidad oculta de la interfaz**
| Tipo | Interfaz | Implementación | Imagen |
| --------------- | -------- | -------------- | ------------------------------------ |
| **Profundo** | Sencillo | Rico | Un rectángulo: estrecho y profundo |
| **Superficial** | Complejo | Sencillo | Un rectángulo: ancho y poco profundo |
El módulo ideal es profundo: los usuarios sólo necesitan ver la interfaz corta y la complejidad está oculta en su interior. Un contraejemplo extremo es el módulo superficial: la interfaz es casi tan compleja como la implementación, lo que significa que no hay encapsulación. Los usuarios también podrían mirar la implementación directamente.
Juicio de Ousterhout: \*\* Las buenas bases de código se componen de una pequeña cantidad de módulos profundos; Las bases de código incorrecto se componen de una gran cantidad de módulos poco profundos\*\*. Esto es completamente opuesto al dogma tradicional de "mantener las funciones lo más pequeñas posible, los archivos lo más cortos posible y los módulos tantos como sea posible"; él cree que ese tipo de dogma produce módulos exactamente superficiales.
## Extensión de Matt: Prueba de eliminación
Matt tradujo la teoría de Ousterhout en una prueba de ingeniería operativa, a la que llamó **prueba de eliminación**:
> **Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.**
Palabras humanas:
* **Elimínelo, la complejidad desaparece** → Este módulo es originalmente un paso (tránsito), no funciona, así que córtelo
* **Elimínalo y la complejidad se extenderá a N personas que llaman** → Originalmente te ayuda a ocultar la complejidad, es muy profunda, déjalo
Lo bueno de esta prueba es que es bidireccional: puede identificar tanto "envoltorios delgados que deben eliminarse" como "lógica común que debe extraerse". Si descubre que después de eliminar un fragmento de código, la complejidad se extenderá a 5 lugares, significa que vale la pena extraer este código en un módulo profundo.
## Términos clave (definición precisa de Matt)
Hay un glosario en `improve-codebase-architecture/SKILL.md` que requiere **el uso estricto de estas palabras**; no se desvíe hacia "componente", "servicio", "API" y "límite":
| Terminología | Definición |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| **Módulo** | Cualquier cosa con una interfaz e implementación (función/clase/paquete/sección) |
| **Interfaz** | Todo lo que la persona que llama debe saber: tipos, invariantes, modos de error, orden, configuración (no solo firmas de funciones) |
| **Implementación** | Código dentro del módulo |
| **Profundidad** | La palanca en la interfaz. Profundo = alto apalancamiento, superficial = la interfaz es casi tan compleja como la implementación |
| **Costura** (Costura) | La ubicación de una interfaz, donde el comportamiento se puede cambiar sin modificaciones in situ. **Utilice "costura", no "límite"** |
| **Adaptador** | Implementar la implementación específica de una interfaz en costura |
| **Apalancamiento** | Los beneficios que obtiene la persona que llama de "profundo" |
| **Localidad** | Los beneficios que los mantenedores obtienen de la "profundidad": cambios, errores y conocimiento se concentran en un solo lugar |
Varios principios básicos:
* **Prueba de eliminación**: ver arriba
* **La interfaz es la superficie de prueba**: las pruebas solo se pueden ejecutar a través de la interfaz; esta es la base para la capacidad de prueba profunda del módulo.
* **Un adaptador = costura hipotética. Dos adaptadores = costura real.**: Una interfaz con una sola implementación es una costura falsa. **Las uniones reales requieren al menos dos adaptadores**
El último es particularmente contrario a la intuición: muchos equipos abstraerán una interfaz de antemano "para una futura expansión", pero en realidad solo hay una implementación. Juicio de Matt: **inútil, eliminar**. Espere hasta que el segundo se haga realidad. Este tiene el mismo origen que YAGNI.
## Flujo de trabajo de habilidades
### 1. Explorar
La habilidad primero permite que AI lea `CONTEXT.md` y `docs/adr/`, y luego usa `subagent_type=Explore` para enviar un subagente a la base del código.
En lugar de una inspiración rígida, utiliza **fricción** como señal:
> * Where does understanding one concept require bouncing between many small modules?
> * Where are modules **shallow** — interface nearly as complex as the implementation?
> * Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
> * Where do tightly-coupled modules leak across their seams?
> * Which parts of the codebase are untested, or hard to test through their current interface?
Cada vez que encuentre un punto sospechoso, aplique una prueba de eliminación: ¿eliminarlo hará que la complejidad desaparezca o se extienda? La respuesta "dispersarse" es una candidata digna de profundizar.
### 2. Candidatos presentes (Candidatos de declaración)
Presentar lista numerada de candidatos:
```
1. Files: src/orders/parser.ts, src/orders/validator.ts, src/orders/normalizer.ts
Problem: 三个文件互相调用,理解 Order 入站需要在三处跳转
Solution: 合并为单一 OrderIntake 模块,对外只暴露 parse(raw) → ValidatedOrder
Benefits:
- Locality: Order 入站的所有逻辑、错误处理、bug 修复集中一处
- Leverage: 调用方从理解 3 个接口降为 1 个
- Tests: 只需测 parse() 的输入输出,不再需要 mock 内部协作
```
Requisitos:
* Utilice **vocabulario CONTEXT.md** para hablar sobre dominios ("el módulo de admisión de pedidos", no "el FooBarHandler")
* Hablar sobre arquitectura usando **vocabulario del glosario** ("costura", "profundidad", "localidad")
* **No proponga diseños de interfaz de inmediato**: permita que los usuarios elijan candidatos interesantes primero
Si un candidato entra en conflicto con un ADR existente, menciónelo únicamente si el conflicto justifica revisar el ADR y márquelo claramente:
> "contradicts ADR-0007 — but worth reopening because…"
No desenterres todas las refactorizaciones prohibidas por ADR.
### 3. Circuito para asar
Después de que el usuario seleccione un candidato, pase al modo de parrilla (heredado de [`/grill-with-docs`](/es/docs/notes/matt-pocock-skills/grill-with-docs)):
* Recorrer el árbol de diseño: limitaciones, dependencias, la forma del módulo después de profundizarlo, qué se esconde detrás de las costuras, qué pruebas pueden sobrevivir.
* **Los efectos secundarios ocurren inmediatamente**:
* Asigne al módulo de profundización un nombre que no esté en CONTEXT.md → Agréguelo a CONTEXT.md inmediatamente
* Se agudizó un término ambiguo en la tortura → Actualice CONTEXT.md inmediatamente
* El usuario rechaza al candidato por motivos importantes (crítico, algo que los futuros exploradores deben saber) → Proponer generar ADR
* Quiere explorar los diversos diseños de interfaz de los módulos de profundización → Saltar al proceso separado `INTERFACE-DESIGN.md`
El mantenimiento de documentos y la transformación de la arquitectura ocurren en la misma conversación: no hay dos rondas.
## Caso real: la práctica de Mejba Ahmed
El desarrollador externo Mejba Ahmed escribió un artículo \["Módulos profundos: la habilidad del código Claude guardando mi base de código"] ([https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules](https://www.mejba.me/blog/improve-codebase-architecture-skill-deep-modules)) para registrar en detalle su experiencia usando esta habilidad. Puntos clave:
* Originalmente tenía más de 50 archivos en un proyecto, cada archivo tenía menos de 100 líneas - **Biblioteca poco profunda típica**
* `/improve-codebase-architecture` se quedó sin 8 candidatos de profundización
* Seleccionó 3 profundizaciones (dos módulos de procesamiento de datos fusionados, un conjunto de herramientas fusionado)
* Resultado: el número de archivos se redujo de 50+ a 30+, pero **el tamaño total del código sigue siendo básicamente el mismo** - la complejidad se reduce a una pequeña cantidad de módulos profundos
* La tasa de aciertos de los cambios de código posteriores de Claude en esta biblioteca mejoró significativamente (dijo "del 60% al 90%", lo cual no se midió estrictamente, pero se sintió fuerte)
Mejba también tiene un recordatorio: **No profundices de 8 a la vez**. Elija solo uno a la vez, ejecute la prueba + confirmar + observar y luego elija el siguiente. De lo contrario, no hay forma de retroceder una vez que lo complete.
## Cómo instalar y usar
```bash
npx skills@latest add mattpocock/skills
```
Marque `improve-codebase-architecture` + `setup-matt-pocock-skills`.
**Llamar**: `/improve-codebase-architecture`
**Ritmo recomendado**:
* **Corre una vez por semana o al final de cada sprint**
* O \*\*ejecútelo una vez después de completar una ola de desarrollo intensivo (es especialmente fácil acumular módulos poco profundos después de escribir código de IA con alta frecuencia)
* **No corras cuando tengas prisa**: sugerirá grandes cambios que no tendrás tiempo de digerir cuando tengas prisa
**Proceso típico**:
1. `/improve-codebase-architecture`
2. Exploración de IA + lista de N candidatos (con argumento de prueba de eliminación)
3. Eliges el que más te sienta
4. Caída en el diseño de alineación del bucle de parrilla.
5. Refactorización de implementación de IA (se recomienda ejecutarla junto con [`/tdd`](/es/docs/notes/matt-pocock-skills/tdd) - la refactorización debe tener protección de prueba)
6. comprometerse + observar
7. Vuelve en una semana.
## ¿Por qué esta habilidad es un "circuito cerrado" del flujo de trabajo de Matt?
Volvamos al diagrama de flujo de trabajo de Matt:
```
/grill-me → /to-prd → /to-issues → /tdd → /improve-codebase-architecture → 回到 /grill-me
```
Observe que regresa al punto inicial. `/improve-codebase-architecture` no es una herramienta que se usa una sola vez, es un **mantenimiento periódico** - porque:
1. La IA continúa agregando módulos poco profundos al código base (esta es su tendencia predeterminada y se acumulará si escribe demasiado)
2. A medida que el negocio siga evolucionando, las viejas costuras quedarán obsoletas.
3. Los términos en CONTEXT.md continúan perfeccionándose y la denominación anterior no se mantendrá.
**Cada vez que ejecutas esta habilidad, la compatibilidad con la IA del código base se actualiza**. Esta es la **única** forma de mantener saludable una base de código a largo plazo con LLM: si no la actualiza, la IA morirá en su base de código después de tres meses.
## Este conjunto de pensamientos es más valioso que la habilidad misma.
Incluso si no instala `/improve-codebase-architecture` en absoluto, recuerde las siguientes tres cosas y la calidad de la revisión de relaciones públicas mejorará un nivel:
1. **prueba de eliminación**: cada vez que vea un módulo nuevo, pregúntese: "Si lo elimina, ¿la complejidad desaparecerá o se extenderá?".
2. **Uniones verdaderas al menos dos adaptadores**: interfaz de implementación única = resumen falso, eliminar
3. **La interfaz es la superficie de prueba**: no se puede medir = hay un problema con el diseño de la interfaz
Estos tres elementos no requieren inteligencia artificial ni habilidad; son la moneda fuerte de la estética de la ingeniería. Matt los integra en habilidades para la ejecución por lotes, pero la verdadera ventaja son los tres principios mismos.
## Notas
**No profundices demasiado**. El propio Ousterhout dijo que el módulo profundo es un objetivo más que un dogma: una clase Util grande que agrupa todas las funciones no es un módulo profundo, sino un módulo divino. El criterio de juicio es "interfaz simple + implementación coherente", los cuales deben cumplirse.
**la profundización debe tener protección de prueba**. Los cambios estructurales son operaciones de alto riesgo y atreverse a refactorizar sin realizar pruebas = esperar a asumir la culpa. Si no hay pruebas actualmente, vaya a [`/tdd`](/es/docs/notes/matt-pocock-skills/tdd) para agregar pruebas a la ruta crítica y luego regrese.
**Las decisiones de ADR no deben tomarse por capricho**. Cuando rechaza a un candidato durante el interrogatorio, la IA sugerirá fácilmente generar un ADR y solo lo aceptará si el motivo es realmente "las personas futuras necesitan saberlo". De lo contrario, docs/adr/ se completará con entradas de diario.
**No importa si parte del código es superficial**. Un contenedor de registrador, un archivo constante, un script único: son superficiales, no hay problema. Esta habilidad busca esos módulos superficiales que pretenden ayudarte con la abstracción pero que en realidad añaden caos.
## Recursos de referencia
***
## Conclusión de la serie
He leído los 6 artículos hasta ahora. Revise todo el flujo de trabajo:
```
/grill-me 或 /grill-with-docs ← 谈清楚要做什么
↓
/to-prd ← 凝固成 PRD
↓
/to-issues ← 切成 vertical slice
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture ← 周期性深化
↓
回到 /grill-me
```
El espíritu de este proceso se puede condensar en una frase:
> \*\*La IA es el soldado táctico en el terreno y tú eres la capa estratégica. Recupere las tres cosas de "definir el problema", "deconstruir el problema" y "probar el problema" y hágalo usted mismo, y deje "escribir código" a la IA: esta es la verdadera posición de los ingenieros en la era de la IA. \*\*
El conjunto de habilidades de Matt no es la respuesta definitiva, sino la mejor práctica en la etapa actual. Puede que haya algo mejor tres meses después, pero el aspecto espiritual no cambiará: una buena base de código siempre es más importante que una mala, y las habilidades básicas de software siempre son valiosas.
Vuelve a [Overview](/es/docs/notes/matt-pocock-skills/overview), o elige la habilidad más útil para instalar y probar.
# Los fundamentos del software son más importantes que nunca: el conjunto de habilidades de Claude Code de Matt Pocock
## Una persona que está abrumada por la ola de especificaciones a código pero aún tranquila
2026 es el año más destacado de la narrativa "de especificaciones a código" de la programación de IA: escriba la especificación, ejecute el compilador, no lea el código, luego escriba la especificación y luego ejecute el compilador. El lema que ha surgido en la comunidad es "**el código es barato**" (el código es barato), lo que significa: De todos modos, la IA puede generar otras 10,000 líneas por segundo, ¿por qué debería importarle?
Matt Pocock es una de las pocas personas que se pronuncia públicamente en contra de esto. No niega que la codificación con IA es muy poderosa, pero en realidad probó especificaciones a código en su clase "Código Claude para ingenieros reales", y la conclusión es muy desgarradora: **Cada vez que lo ejecuto, el código empeora cada vez más**. Ésta es exactamente la "entropía del software" de la que se ha hablado en Pragmatic Programmer: aumento de la entropía del software.
Entonces hizo dos cosas:
1. Incluya esta observación en una charla de 18 minutos: *Los fundamentos del software importan más que nunca*.
2. Empaquete el antídoto correspondiente en un repositorio de GitHub: [`mattpocock/skills`](https://github.com/mattpocock/skills) - "Habilidades para ingenieros reales. Directamente desde mi directorio .claude".
El almacén se lanzó el 3 de febrero de 2026 y en 4 meses alcanzó **61,1 mil estrellas y 5,3 mil bifurcaciones**. Fue uno de los almacenes de programación de IA de más rápido crecimiento durante el mismo período.
***
## ¿Quién es Matt Pocock?
Si ha escrito TypeScript, probablemente lo haya encontrado. Es uno de los educadores de TypeScript más prolíficos en los círculos chino e inglés en los últimos años:
* Fundador de **TotalTypeScript.com**, una serie de cursos pagos que son muy populares en el círculo inglés.
* **aihero.dev** Boletín informativo con más de 60 000 suscripciones, el tema cambió de TS a AI Coding
* Hay muchos videos tutoriales cortos en Twitter [`@mattpocockuk`](https://twitter.com/mattpocockuk) y YouTube `@mattpocockuk`.
* No soy una persona OpenAI/Anthropic, puramente un desarrollador independiente + experiencia en educación.
Su personalidad es muy clara: **Codificación AI desde la perspectiva de un ingeniero senior**. No gritamos “viene AGI”, ni gritamos “los programadores van a perder sus trabajos”. Lo que gritó fue: "Los trucos de la generación anterior de ingenieros de software siguen siendo muy útiles, sólo necesitan traducirse a una forma que LLM pueda ejecutar".
***
## Argumento central: el código no es barato
Solo hay un argumento en todo el discurso, y cada habilidad es una nota a pie de página:
> Si la estructura de su base de código es mala, la IA solo escribirá código incorrecto en una base de código incorrecta. Entonces **una buena base de código es más importante que nunca y las habilidades básicas de software son más importantes que nunca**.
Matt utilizó una analogía militar para explicar los roles de los humanos y la IA de manera muy sencilla:
¿Qué hace la capa estratégica? Conceptos de diseño, lenguaje unificado, límites de módulos: estas tres cosas son "definir problemas" en lugar de "escribir código", y resultan ser lo que LLM hace menos bien por usted.
***
## Cinco patrones de fracaso → Cinco libros antiguos → Cinco habilidades
Matt comprimió toda la metodología en una tabla de mapeo en su discurso. Cada vez que encuentre un modo de falla, él le indicará la teoría clásica que se resolvió hace 20 años y luego le dará un archivo de habilidad en formato Markdown:
| # | Modo de falla de programación de IA | Teoría clásica y fuente | Habilidad correspondiente |
| - | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------- |
| 1 | La IA no hace lo que quieres | *El Diseño del Diseño* (Brooks) - concepto de diseño, árbol de diseño | [`/grill-me`](/es/docs/notes/matt-pocock-skills/grill-me) |
| 2 | La IA te habla en muchos términos detallados | *Diseño basado en dominios* (Evans) - lenguaje ubicuo | [`/grill-with-docs`](/es/docs/notes/matt-pocock-skills/grill-with-docs) |
| 3 | La IA lo hace bien pero no puede funcionar | *El programador pragmático* (Hunt & Thomas)—— "la tasa de retroalimentación es su límite de velocidad" | [`/tdd`](/es/docs/notes/matt-pocock-skills/tdd) |
| 4 | La IA deambula por bases de códigos incorrectos | *Una filosofía del diseño de software* (Ousterhout): módulos profundos, prueba de eliminación | [`/improve-codebase-architecture`](/es/docs/notes/matt-pocock-skills/improve-codebase-architecture) |
| 5 | Tu cerebro no puede seguir el ritmo de la producción de IA | Kent Beck —— invierte en diseño todos los días | "diseñar la interfaz, delegar la implementación" |
Elemento 5 No hay una habilidad separada en el repositorio (solía haber `design-an-interface` pero está en desuso), su espíritu ha sido absorbido en [`/to-prd`](/es/docs/notes/matt-pocock-skills/to-prd-and-issues) y [`/improve-codebase-architecture`](/es/docs/notes/matt-pocock-skills/improve-codebase-architecture), los cuales te obligan a pensar en las interfaces de los módulos antes de escribir código.
***
## 5 Habilidades que realmente usas todos los días
El discurso fue un esqueleto filosófico. Más tarde, Matt publicó un artículo "Cinco habilidades de agente que uso todos los días" en aihero.dev para traducir el esqueleto en un flujo de trabajo diario. Estos 5 son los objetos que serán desmantelados uno a uno en el futuro de esta serie:
```
/grill-me ← 先和 AI 谈清楚要做什么
↓
/to-prd ← 把对话凝固成 PRD
↓
/to-issues ← 把 PRD 切成可独立领取的 vertical slice
↓
/tdd ← 每个 slice 用红绿重构跑通
↓
/improve-codebase-architecture ← 周期性检查,把 shallow 模块改成 deep
```
Estas cinco habilidades se unen para formar el proceso completo de investigación y desarrollo de Matt. Los modos de fallo correspondientes a cada paso se muestran en la tabla del apartado anterior.
Desmontaje detallado de cada artículo (páginas siguientes de esta serie):
* [Grill Me: Deja que la IA te torture sobre tus necesidades](/es/docs/notes/matt-pocock-skills/grill-me)
* [Grill With Docs: Mantenimiento del lenguaje del proyecto y ADR](/es/docs/notes/matt-pocock-skills/grill-with-docs)
* [a-PRD + a-Issues: de conversación a ticket ejecutable](/es/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [TDD: Usa reconstrucción rojo-verde para obligar a la IA a dar pequeños pasos](/es/docs/notes/matt-pocock-skills/tdd)
* [Mejorar la arquitectura de la base de código: reestructurar módulos superficiales en módulos profundos](/es/docs/notes/matt-pocock-skills/improve-codebase-architecture)
***
## Cómo instalar
El archivo README del almacén proporciona un comando de una línea para instalar:
```bash
npx skills@latest add mattpocock/skills
```
Este comando:
1. Le permite verificar qué habilidades desea instalar
2. Le permite seleccionar qué agentes instalar (se admiten Claude Code, Codex, Cursor, etc.)
3. Coloque el archivo SKILL.md correspondiente en `.claude/skills/` (o el directorio correspondiente al agente)
**Se recomienda encarecidamente marcar `/setup-matt-pocock-skills`** al mismo tiempo; esta es una habilidad de configuración única que le hará tres preguntas:
* ¿Para qué se utiliza el **Rastreador de problemas**? (GitHub / GitLab / rebajas locales / otros)
* ¿Qué palabra se utiliza para **etiqueta de triaje**? (triaje de necesidades u otro)
* ¿Dónde poner **Documento de dominio**? (ruta CONTEXT.md/ADR)
Ejecute `/setup-matt-pocock-skills` una vez y escribirá en `AGENTS.md` o `CLAUDE.md` en el directorio raíz de su proyecto. Después de eso, todas las habilidades de ingeniería (to-prd, to-issues, triage, tdd, etc.) leerán automáticamente esta configuración. Este paso se omite y cada habilidad posterior le hará la misma pregunta una y otra vez.
Si solo desea probar `/grill-me` (la clase de productividad pura y más liviana), puede omitir la configuración, ya que no depende del rastreador de problemas.
***
## La diferencia entre este conjunto de habilidades y BMAD/Spec-Kit/GSD
Si ya está utilizando marcos basados en especificaciones como [BMAD](/es/docs/notes/speckit/concept), Spec-Kit y GSD, puede preguntar: "¿Por qué todavía necesitas el conjunto de Matt?"
Matt escribe muy directamente en el README:
**Diferencias principales**:
* BMAD/Spec-Kit/GSD es un **marco** que especifica un proceso completo desde la especificación hasta el código. Tienes que seguir su proceso.
* Matt Este conjunto es **componente**. Cada habilidad tiene un archivo de rebajas, que va desde unas pocas líneas hasta docenas de líneas. Puedes desmontarlo y modificarlo en cualquier momento.
Ejemplo: el texto completo real de `grill-me` es solo así de breve——
```markdown
Interview me relentlessly about every aspect of this plan until we reach
a shared understanding. Walk down each branch of the design tree,
resolving dependencies between decisions one-by-one. For each question,
provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the
codebase instead.
```
La habilidad completa tiene 7 líneas. Pero son esas siete líneas las que hacen que Claude te haga 20, 50 o incluso 100 preguntas antes de tomar una decisión. Esta filosofía de diseño de **usar muy poco texto para aprovechar grandes cambios de comportamiento** es la razón fundamental de la popularidad de este conjunto de habilidades.
***
## Cómo leer esta serie
Si no se ha encontrado con el conjunto de Matt antes, se recomienda leerlo en el orden en meta.json:
1. **Descripción general** (el artículo en el que se encuentra actualmente): obtenga la imagen completa
2. **Grill Me** - Instale uno individualmente y pruébelo primero, el umbral es el más bajo
3. **Grill With Docs**: una versión avanzada de grill-me, que comienza a presentar CONTEXT.md
4. **a-PRD + a-Issues** - Convierte la conversación en un ticket ejecutable
5. **TDD** - El propio Matt dijo "el método más estable que he usado jamás para mejorar la calidad de la producción del agente".
6. **Mejorar la arquitectura de la base de código**: mantenimiento periódico para que la IA esté disponible a largo plazo
Si ya estás usando Claude Code para escribir proyectos reales, salta directamente a [`/grill-me`](/es/docs/notes/matt-pocock-skills/grill-me) + [`/tdd`](/es/docs/notes/matt-pocock-skills/tdd). Estos dos artículos son los más intuitivos.
Si está enseñando o escribiendo, simplemente leer [`/grill-me`](/es/docs/notes/matt-pocock-skills/grill-me) es suficiente; es una herramienta general de "conversación de diseño", que no se limita al código.
***
## Mis sugerencias de uso
Después de que instalé este conjunto de habilidades, los mayores cambios físicos fueron:
**Primero**: deja de apresurarte para empezar a escribir código. En el pasado, cuando AI recibía "Agrégame un inicio de sesión", comenzaba a diseñar 500 líneas. Ahora `/grill-me` te hará 20 preguntas primero: "¿Quieres recordar el dispositivo?" "¿Cuánto tiempo tarda en expirar la sesión?" "¿Cuántas veces no has podido bloquear tu cuenta?" Que se escriba después de 30 minutos. Lo que se ahorrará son las dos horas siguientes de retrabajo.
**Segundo**: CLAUDE.md ya no está hinchado. En el pasado, había muchas prohibiciones escritas en CLAUDE.md, como "Comprenda los requisitos antes de escribir código" y "No sea demasiado abstracto", pero Claude aún así lo cometió. Después de cambiar al conjunto de Matt, CLAUDE.md solo pone el conocimiento del dominio (sistema de diseño, especificaciones de componentes, implementación) y las metodologías generales se entregan a las habilidades. Las responsabilidades de ambas partes son claras.
**Tercero**: El pensamiento profundo del módulo es más valioso que la habilidad en sí. Incluso si no instala `/improve-codebase-architecture`, simplemente leer la "**prueba de eliminación**" en su SKILL.md (si la complejidad desaparece después de eliminar este módulo, significa que es de transferencia) ya le hará echar un segundo vistazo durante la revisión de relaciones públicas.
**Tenga en cuenta el costo**:
* Después de instalar 5 habilidades, la IA hará más preguntas. A las personas que están acostumbradas a "generar 500 líneas en una frase" les resultará molesto.
* Después de implementar estrictamente `/tdd`, también se requerirán scripts simples para escribir pruebas primero, lo cual no es compatible con el código exploratorio; puede decirle "omitir TDD esta vez".
* `/grill-with-docs` tomará la iniciativa de modificar su CONTEXT.md. Es mejor ejecutarlo en seco antes de utilizarlo por primera vez.
***
## Recursos de referencia
**5 libros citados en el discurso** (en orden de aparición):
* *Una filosofía del diseño de software* — John Ousterhout (definición de complejidad, módulos profundos)
* *The Pragmatic Programmer* — David Thomas & Andrew Hunt(software entropy、outrunning headlights)
* *The Design of Design* — Frederick P. Brooks(design concept、design tree)
* *Domain-Driven Design* — Eric Evans(ubiquitous language)
* *Test-Driven Development* — Kent Beck(invest in design every day)
Cada uno tiene más de 20 años. Matt repitió una frase muchas veces en su discurso: "**Vaya a Amazon, consígalo.**"; esta frase en sí misma es el huevo de Pascua de este discurso.
# 其他 Skills:压缩沟通、交接、教学、写 Skill 与安全护栏
## 为什么不逐个展开
Matt 的 README 把 skills 分成三类:
* Engineering:每天写真实代码用
* Productivity:通用工作流工具
* Misc:他自己留着备用的小工具
前面几篇已经覆盖了主线 engineering skill。这一篇把剩下的 productivity 和 misc 合在一起讲,因为它们多数不是完整研发流程,而是**在特定场景下很好用的小开关**。
如果你只装 5 个,我仍然建议优先装:
* [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me)
* [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)
* [`/to-prd` + `/to-issues`](/docs/notes/matt-pocock-skills/to-prd-and-issues)
* [`/tdd`](/docs/notes/matt-pocock-skills/tdd)
* [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture)
但如果你已经把主线跑起来,下面这些会让日常体验更顺。
## Productivity Skills
### caveman:极限压缩沟通
`/caveman` 是一个「少 token 模式」。它要求 agent 去掉寒暄、填充词、过度解释和模糊缓冲,只保留技术信息。
它适合:
* 你已经在高频迭代,不想读长回复
* debug 时只要事实、原因、下一步
* 长上下文快满了,需要压缩输出
* 你想强制 AI 少说漂亮话
它不是让 AI 变粗鲁,而是让它用更短的语法保留完整技术精度。注意它会持续生效,直到你说退出。
我会把它当作一个临时档位,而不是长期默认。对于高风险操作、安全警告、复杂多步骤指令,太短反而容易误读。
### handoff:把当前会话交给下一个 agent
`/handoff` 的目标是把当前会话压缩成一个交接文档,并保存到系统临时目录,而不是污染当前 workspace。
它会包含:
* 当前目标
* 已做决策
* 关键路径和文件
* 剩余任务
* 建议下个 agent 调用哪些 skill
* 敏感信息脱敏
它特别适合长任务中断、上下文快满、或你想换一个 agent 继续做的时候。
关键点是:不要复制已经存在于 PRD、issue、ADR、commit、diff 里的内容,只引用路径或 URL。交接文档的价值是**把散在对话里的状态补齐**,不是再造一份项目文档。
### teach:把当前目录变成学习工作区
`/teach` 是这组里最重的一个。它把当前目录当作一个长期学习 workspace,维护:
* `MISSION.md`:你为什么要学这个主题
* `RESOURCES.md`:高质量资源列表
* `learning-records/*.md`:学习记录,类似 ADR
* `lessons/*.html`:每次一节的互动课程
* `reference/*.html`:速查资料
* `NOTES.md`:教学偏好和工作笔记
它的亮点是把学习看成长期系统,而不是一次问答。尤其强调:
* mission 先行:为什么学,比学什么更重要
* retrieval practice:用回忆练习建立长期记忆
* spacing / interleaving:不要被短期流畅感骗了
* 高信任资源:先找资料,不凭模型记忆硬讲
如果你只是问「解释一下 X」,不需要它;如果你想连续几周学一个主题,它就很合适。
### write-a-skill:写新 skill 的脚手架
`/write-a-skill` 是 Matt 对 skill 结构本身的抽象。
它要求一个 skill 至少有:
```text
skill-name/
├── SKILL.md
├── REFERENCE.md
├── EXAMPLES.md
└── scripts/
```
当然后三个不是必需,只有内容太长、示例有价值、或操作可脚本化时才加。
它最重要的判断是:`description` 是 agent 决定是否加载 skill 时唯一先看到的信息。因此 description 不能写成「帮助处理文档」这种空话,必须说明:
* 它提供什么能力
* 什么时候触发
* 触发词或上下文是什么
这和我自己写 skill 的经验一致:很多 skill 失效不是因为正文写得差,而是 description 写得太泛,agent 根本不知道该加载它。
## Misc Skills
### git-guardrails-claude-code:拦危险 git 命令
这个 skill 会给 Claude Code 配一个 `PreToolUse` hook,在执行 Bash 前拦截危险 git 命令。
默认会挡:
* `git push`
* `git reset --hard`
* `git clean -f` / `git clean -fd`
* `git branch -D`
* `git checkout .` / `git restore .`
它的价值很直接:防止 agent 在你没授权时推送、硬重置、清掉未跟踪文件。
如果你经常让 AI 在真实仓库里工作,这个 skill 很值得装。它不是不信任 AI,而是把高破坏性操作放到工具层拦截,而不是靠 prompt 祈祷。
### setup-pre-commit:给项目加提交前检查
`/setup-pre-commit` 会设置:
* Husky pre-commit hook
* lint-staged + Prettier
* typecheck
* test
它会先检测包管理器,再按项目已有 script 决定 pre-commit 里该跑什么。没有 `typecheck` 或 `test` 时不会硬造,而是省略并告诉你。
这个 skill 的价值不在配置本身,而在 Matt 的质量观:**不要只让 AI 自己说代码没问题,要让它过确定性检查**。
### migrate-to-shoehorn:测试里少写 `as`
这是一个很 Total TypeScript 风格的小工具。它把测试里的 TypeScript `as` 类型断言迁移到 `@total-typescript/shoehorn`。
典型替换:
| 旧写法 | 新写法 | 场景 |
| --------------------------- | ------------------ | -------------- |
| `obj as Request` | `fromPartial(obj)` | 测试里只关心大对象的几个字段 |
| `obj as unknown as Request` | `fromAny(obj)` | 故意传错类型测错误路径 |
| 完整对象假数据 | `fromExact(obj)` | 需要强制完整形状 |
它明确只用于测试代码,不用于生产代码。
这个 skill 很窄,但很符合 Matt 的工程品味:不要为了类型系统在测试里造 20 个无意义字段,也不要用裸 `as` 把类型安全完全关掉。
### scaffold-exercises:给课程仓库生成练习目录
这个 skill 明显来自 Matt 自己做课程的工作流。它会按规范创建:
```text
exercises/
└── 05-memory-skill-building/
└── 05.02-short-term-memory/
├── explainer/
├── problem/
└── solution/
```
每个子目录至少有非空 `readme.md`,必要时有 `main.ts`,并且要通过 `pnpm ai-hero-cli internal lint`。
它对大多数工程项目没用,但对课程、训练营、练习仓库非常实用。更重要的是,它展示了一个好 skill 的特征:**把重复、机械、容易漏细节的格式工作交给 agent**。
## 不建议现在写进主线的目录
上游仓库里还有 `deprecated/`、`in-progress/`、`personal/`。
我建议暂时不要把它们写成正式使用指南:
| 目录 | 为什么不放主线 |
| -------------- | -------------------------- |
| `deprecated/` | 已废弃,容易误导读者继续采用旧流程 |
| `in-progress/` | 还在实验,行为和命名都可能变 |
| `personal/` | 更像 Matt 自己的私人工作区,不一定适合通用读者 |
如果以后要写,可以单独做一篇「Matt Pocock skills 仓库考古」,而不是混在稳定推荐里。
## 这组小工具的共同点
这些 skill 看起来很散,但背后有同一个原则:
> 把 agent 容易漂移的事情,变成小而明确的工作模式。
* `caveman` 防止沟通漂移
* `handoff` 防止上下文丢失
* `teach` 防止学习变成一次性问答
* `write-a-skill` 防止 skill 结构随手写
* `git-guardrails` 防止危险命令靠自觉
* `setup-pre-commit` 防止质量检查靠 AI 自述
* `migrate-to-shoehorn` 防止测试类型断言失控
* `scaffold-exercises` 防止课程结构手工漏项
这也是 Matt 这套 repo 最值得学习的地方:skill 不需要宏大。一个高频小偏差,如果能被 20 行指令稳定纠正,就值得写成 skill。
## 参考资源
# Prototype:用可丢弃代码回答一个设计问题
## 原型不是「先随便写一个」
`/prototype` 的第一句定义很重要:
> 原型是用来回答一个问题的可丢弃代码。
这句话把原型和「偷懒版实现」分开了。原型不是生产代码的前身,不是以后慢慢改成正式版的半成品。它从第一天开始就应该被标记为 throwaway。
所以 `/prototype` 的关键不是写得快,而是先问清楚:
> 这个原型到底要回答什么问题?
## 两条分支
Matt 把原型分成两类,输出完全不同。
| 要回答的问题 | 分支 | 输出 |
| --------------- | --------------- | ---------------- |
| 逻辑、状态机、数据模型是否合理 | Logic prototype | 一个可运行的终端小程序 |
| 这个界面应该长什么样 | UI prototype | 一个路由里多套可切换 UI 方案 |
这点很实用。很多团队说「做个 prototype」,但没说清楚是想验证交互外观,还是验证状态流转。两者需要的东西完全不同。
## Logic prototype:把状态摊在终端里
如果问题是「这个状态机对不对」「这个业务规则能不能跑通」,原型应该是一个很小的命令行程序。
它的特点:
* 内存态,不依赖真实数据库
* 一个命令启动
* 每个操作后打印完整相关状态
* 覆盖那些纸面上难以推演的分支
* 不写测试,不做异常兜底,不抽象成框架
例子:你要设计订阅状态流。
不要直接改生产代码。先写一个 `subscription-prototype.ts`,让用户可以在终端里选择:
```text
1. start trial
2. pay
3. cancel
4. expire
5. refund
6. print state
```
每按一步,打印当前 entitlement、trial quota、paid state、next renewal。你会很快发现有些状态组合根本没想清楚。
这类原型的价值是:**让抽象规则变成可操作对象**。
## UI prototype:在一个路由里放多套激进方案
如果问题是「界面应该怎么设计」,原型应该生成几套差异足够大的 UI,而不是把同一个方案微调三次。
`/prototype` 的 UI 分支要求:
* 在一个路由里放多个 variation
* 用 URL search param 或底部浮动切换条切换
* 方案之间要有明显差异
* 原型代码靠近未来真实页面,但命名要清楚表示 prototype
* 不要过早接真实数据和持久化
这和普通 AI 生成 UI 的区别在于:它不是让 AI 一次给「最佳方案」,而是让你用真实浏览器比较几种方向。
比如一个 dashboard 空状态,不要只让 AI 改文案。可以让它做:
* A:表格式、密度高,强调下一步操作
* B:任务导向,左侧 checklist + 右侧预览
* C:引导式,突出一个主 CTA 和历史示例
然后你在同一路由切换,而不是在聊天里看三张截图脑补。
## 所有原型都必须可删除
`/prototype` 的通用规则里,最重要的是「可删除」:
* 文件名或路径要表明这是 prototype
* 不要默认接生产数据库
* 不要写通用抽象
* 不要做过度错误处理
* 结束后要删除,或者把学到的结论吸收到正式代码
如果一个原型不能删,它就已经变成生产代码负债。
这点在 AI 编程里尤其重要。AI 很擅长把 prototype 写得「看起来能用」,然后人类懒得删,最后项目里多出一堆没人敢碰的临时代码。
## 原型结束后要留下什么
原型代码不值得保留,但答案值得保留。
Matt 建议把下面几件事写到一个持久位置:
* 原型要回答的问题
* 观察到的结论
* 选择了哪个方向
* 放弃了哪些方向
* 如果需要,转成 ADR、issue、PRD 或 commit message
也就是说,`/prototype` 的产物不是代码,而是**决策**。
## 什么时候不该用
不要把 `/prototype` 用在这些场景:
* 需求已经明确,只需要实现
* bug 已经有复现,应该用 [`/diagnose`](/docs/notes/matt-pocock-skills/diagnose-and-triage)
* 重构方向已经清楚,应该用 [`/tdd`](/docs/notes/matt-pocock-skills/tdd) 保护后实施
* UI 只是小 polish,不值得做多套方案
* 你没有时间删除或吸收原型
原型的成本不在写,而在收尾。没有收尾,就不要开。
## 一个好用的提示词
可以这样调用:
```text
/prototype
我想验证这个 checkout 状态机是否合理。请走 logic 分支。
只做可丢弃终端原型,不接真实 DB。
每次操作后打印完整状态。
```
或者:
```text
/prototype
我想比较项目详情页的 3 种信息架构。请走 UI 分支。
放在现有路由体系下的 prototype route,提供底部切换条。
不要动生产组件。
```
这里最重要的是明确「要回答的问题」。只要这个问题清楚,原型就不容易跑偏。
## 和 Grill Me 的关系
[`/grill-me`](/docs/notes/matt-pocock-skills/grill-me) 适合通过提问收束决策;`/prototype` 适合通过试玩收束决策。
有些问题靠问就能解决,比如「匿名评论要不要审核」。有些问题必须摸一下,比如「这个拖拽排序状态机到底会不会难用」。后者就该 prototype。
所以我会把它放在工作流的一个分叉位置:
```text
想法模糊
↓
/grill-me
↓
如果仍然需要体验或验证
↓
/prototype
↓
保留结论,删除原型
↓
/to-prd 或 /tdd
```
## 参考资源
下一篇:[Improve Codebase Architecture:把 shallow 重构成 deep modules](/docs/notes/matt-pocock-skills/improve-codebase-architecture)。
# Setup Matt Pocock Skills:先把项目规则写清楚
## 这个 Skill 解决的不是安装问题
`/setup-matt-pocock-skills` 容易被误解成「装完之后跑一下的初始化命令」。实际上它更像一个**项目契约生成器**:告诉后续 skill 这个仓库怎么追踪任务、怎么标记 issue、在哪里读取领域语言和架构决策。
Matt 在 README 的 Quickstart 里特别提醒:安装时要选中 `/setup-matt-pocock-skills`,然后在 agent 里运行它。原因很简单:`to-prd`、`to-issues`、`triage`、`diagnose`、`tdd`、`improve-codebase-architecture`、`zoom-out` 都需要同一批项目上下文。如果每个 skill 都临时问一遍,流程会变得很碎。
它做的不是「配置 Claude 偏好」,而是回答三个工程问题:
| 问题 | 它要写清楚什么 | 后续谁会用 |
| ---------------- | --------------------------------------------------------- | ----------------------------------------------------------------------------- |
| Issue tracker 在哪 | GitHub、GitLab、本地 markdown,还是其他系统 | `to-prd`、`to-issues`、`triage` |
| Triage 标签怎么映射 | `needs-triage`、`needs-info`、`ready-for-agent` 等角色对应哪些真实标签 | `triage` |
| 领域文档在哪里 | 单一 `CONTEXT.md`,还是多上下文 `CONTEXT-MAP.md` + 分区 ADR | `grill-with-docs`、`diagnose`、`tdd`、`zoom-out`、`improve-codebase-architecture` |
## 它为什么重要
这套 skills 的核心思路是「小而可组合」。小的代价是:它们不想自己接管整个项目流程,所以必须知道你项目里的真实约定。
举例:`/to-issues` 要创建 issue。没有 setup,它不知道应该:
* 调 `gh issue create`
* 调 `glab issue create`
* 写到 `.scratch//`
* 还是给你生成一段 Linear/Jira 可复制文本
再比如 `/triage` 要把 issue 移到 `ready-for-agent`。如果你的仓库里真实标签叫 `ai:ready`,而 skill 自己创建了一个新标签 `ready-for-agent`,issue tracker 会立刻变脏。
所以 `/setup-matt-pocock-skills` 的价值不是自动化,而是**把隐含约定外显化**。
## 它会读哪些东西
这个 skill 开始时会先探索仓库,而不是假设:
* `git remote -v` 和 `.git/config`:判断是不是 GitHub/GitLab 项目
* 根目录 `AGENTS.md` / `CLAUDE.md`:看是否已有 `## Agent skills` 区块
* 根目录 `CONTEXT.md` / `CONTEXT-MAP.md`:判断领域语言文档形态
* `docs/adr/` 和 `src/*/docs/adr/`:判断 ADR 是全局还是模块级
* `docs/agents/`:看是否已经跑过 setup
* `.scratch/`:判断是否已有本地 markdown issue 约定
这符合 Matt 整套工作流的风格:**先看项目真实状态,再写规则**。
## 三个决策
### 1. Issue tracker
这是后续工作单元落地的位置。
默认倾向是 GitHub,因为这套 skill 最早围绕 GitHub Issues 设计。但它已经把 GitLab 和本地 markdown 也当作一等选择:
| 选择 | 适合什么场景 |
| -------------- | ---------------------------------- |
| GitHub | 开源项目、GitHub issue workflow 已经存在 |
| GitLab | 公司项目在 GitLab,习惯用 `glab` |
| Local markdown | 个人项目、临时探索、没有远程 issue tracker |
| Other | Jira、Linear、飞书、多维表格等,需要用文字记录你的真实流程 |
重点不是选哪个,而是选**团队真的在用的那个**。写错了,后续 skill 会在错误系统里创建任务。
### 2. Triage label vocabulary
`/triage` 内部使用 5 个状态角色:
| 角色 | 含义 |
| ----------------- | ------------------- |
| `needs-triage` | 等维护者判断 |
| `needs-info` | 等报告者补信息 |
| `ready-for-agent` | 已经清楚到可以交给 AFK agent |
| `ready-for-human` | 需要人类判断或实现 |
| `wontfix` | 不处理 |
setup 会问你这些角色对应的真实标签名。如果项目里没有既有标签,用默认名就行;如果已经有自己的命名体系,应该在这里映射,而不是让 skill 新造一套。
### 3. Domain docs
这是 Matt 这套 skill 和普通 prompt 最大的区别:它不是只看当前对话,还会读项目里的**领域语言**和**架构决策**。
最简单的形态:
```text
/
├── CONTEXT.md
└── docs/
└── adr/
```
大型 monorepo 可以用多上下文:
```text
/
├── CONTEXT-MAP.md
├── apps/
│ └── web/
│ ├── CONTEXT.md
│ └── docs/adr/
└── services/
└── billing/
├── CONTEXT.md
└── docs/adr/
```
setup 不是强迫你现在写完所有文档,而是告诉后续 skill:应该去哪里找、找不到时应该如何创建。
## 它会写什么
最终会有两类输出。
第一类是 `AGENTS.md` 或 `CLAUDE.md` 里的 `## Agent skills` 区块:
```markdown
## Agent skills
### Issue tracker
...
### Triage labels
...
### Domain docs
...
```
第二类是 `docs/agents/` 下的三份说明:
| 文件 | 内容 |
| ------------------------------ | --------------------------------------- |
| `docs/agents/issue-tracker.md` | issue 系统、命令、创建/更新约定 |
| `docs/agents/triage-labels.md` | canonical role 到真实标签的映射 |
| `docs/agents/domain.md` | `CONTEXT.md`、`CONTEXT-MAP.md`、ADR 的读取规则 |
注意它会优先编辑已有的 `CLAUDE.md`;没有 `CLAUDE.md` 才考虑 `AGENTS.md`。这体现了一个很重要的克制:**不要在项目里制造两份互相竞争的 agent 规则入口**。
## 我建议怎么用
第一次装 Matt 这套 skill 时,顺序应该是:
1. 安装:`npx skills@latest add mattpocock/skills`
2. 选中 `/setup-matt-pocock-skills`
3. 运行 `/setup-matt-pocock-skills`
4. 按真实项目状态回答 issue tracker、标签、领域文档三个问题
5. 看它生成的 `## Agent skills` 和 `docs/agents/*.md`
6. 再开始用 [`/grill-with-docs`](/docs/notes/matt-pocock-skills/grill-with-docs)、[`/to-prd`](/docs/notes/matt-pocock-skills/to-prd-and-issues)、[`/triage`](/docs/notes/matt-pocock-skills/diagnose-and-triage)
如果只是想单独体验 [`/grill-me`](/docs/notes/matt-pocock-skills/grill-me),可以跳过 setup;但只要进入工程流,最好先做。
## 这个 Skill 的设计启发
`/setup-matt-pocock-skills` 看起来很朴素,但它解决了 agent workflow 里最常见的一个问题:**规则散落在人的脑子里**。
很多团队把「我们用哪个标签」「哪些 issue 可以给 AI」「CONTEXT.md 在哪」这些信息当作口头约定。人知道,AI 不知道。AI 不知道就会反复问,或者更糟糕:自己猜。
setup 的作用是把这些口头约定变成可读取文件。后续 skill 不需要更聪明,只需要稳定地读同一份项目契约。
这也是我觉得它值得单独写一篇的原因:它不是炫技的 skill,但它是整套 workflow 能长期跑起来的地基。
## 参考资源
下一篇:[Grill Me:让 AI 在你写代码前拷问你 50 个问题](/docs/notes/matt-pocock-skills/grill-me)。
# TDD: utilice la refactorización rojo-verde para obligar a la IA a dar pequeños pasos
## Modo de fallo: "La IA hace lo correcto, pero no puede ejecutarse"
El tercer modo de fracaso en la charla de Matt: **La dirección es correcta, pero no funciona**.
La solución más directa es instalar una infraestructura de retroalimentación para la IA:
* TypeScript (sin escritura estática *es una locura*)
* Permitir que LLM acceda al navegador y vea la página por sí mismo
* Pruebas automatizadas
Pero Matt observó una cosa: **Incluso con estos comentarios instalados, LLM no funciona bien**. Tiende a escribir 500 líneas a la vez y luego piensa "oh, debería escribir, verifique eso". Esto es lo que el Programador Pragmático llama *dejar atrás a tus faros*: conduce más rápido de lo que los faros pueden iluminar y es sólo cuestión de tiempo antes de que choques contra la pared.
> "The rate of feedback is your speed limit, which means you should be testing as you go, taking small deliberate steps. **And the AI by default is really not very good at that.**"
Para solucionar este problema, debe obligar a la IA a detenerse paso a paso en el nivel de la herramienta. La respuesta de Matt es TDD: **Probar primero puede forzar puntos de control**.
## Teoría clásica: la reconstrucción rojo-verde de Kent Beck
El ritmo estándar de TDD lo define Kent Beck en su libro de 2003 "Desarrollo basado en pruebas: con el ejemplo":
1. **ROJO**: Escribe una prueba fallida (describe qué hacer)
2. **VERDE**: Escriba el código lo suficientemente pequeño como para que la prueba pase
3. **REFACTOR**: Mejora la estructura del código bajo protección de prueba
Cada bucle es extremadamente corto, del orden de minutos. Hay comprobaciones automáticas (prueba aprobada/fallada) en cada paso.
Matt sigue directamente este ritmo, pero su SKILL.md pasa mucho tiempo hablando de un **antipatrón**: este es el núcleo.
## Antipatrón clave: rojo y verde cortados horizontalmente
Mucha gente piensa que TDD significa "escribir todas las pruebas primero y luego escribir todas las implementaciones". Matt dice directamente que esto está mal en SKILL.md:
```
WRONG (horizontal slicing):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical slicing via tracer bullets):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
¿Por qué está mal la horizontal? SKILL.md dio tres razones:
> 1. Tests written in bulk test *imagined* behavior, not *actual* behavior
> 2. You end up testing the *shape* of things (data structures, function signatures) rather than user-facing behavior
> 3. Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
Dicho humano: **Escribir todas las pruebas de una sola vez es probar lo que hay en tu cabeza, no el código real**. Cuando escribe impl3, se da cuenta de que el diseño de test1 es incorrecto, pero en este momento, test2/test3/test4 están todos acoplados al diseño incorrecto. Regrese y haga cambios.
El enfoque correcto es probar una implementación y luego abrir el siguiente par después de escribir un par. Después de completar cada par y haber aprendido algo de esta implementación, el siguiente par de pruebas se puede diseñar basándose en la experiencia real, no en la imaginación.
## Estructura de texto completo de habilidad
`engineering/tdd/SKILL.md` es una de las habilidades más largas que Matt ha escrito porque TDD en sí tiene muchos matices. La estructura central es la siguiente:
### Filosofía
> **Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
> **Good tests** are integration-style: they exercise real code paths through public APIs. They describe *what* the system does, not *how* it does it. A good test reads like a specification.
> **Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly). The warning sign: your test breaks when you refactor, but behavior hasn't changed.
Recuerde un diagnóstico: **Si cambia el nombre de una función interna, la prueba se arrodillará; entonces esta prueba está probando la implementación en lugar del comportamiento, lo cual es una mala prueba**.
### Flujo de trabajo (con lista de verificación)
#### 1. Planning
Alinearse con los usuarios antes de escribir código:
```
[ ] Confirm with user what interface changes are needed
[ ] Confirm with user which behaviors to test (prioritize)
[ ] Identify opportunities for deep modules (small interface, deep impl)
[ ] Design interfaces for testability
[ ] List the behaviors to test (not implementation steps)
[ ] Get user approval on the plan
```
Pregunta clave: "**¿Cómo debería verse la interfaz pública? ¿Qué comportamientos es más importante probar?**"
> "**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case."
Este es muy contrario a la intuición. De forma predeterminada, la IA querrá agotar todos los casos extremos, pero Matt enfatiza la **prioridad**: no vale la pena medir todos los comportamientos, y centra la potencia de fuego en el camino principal.
#### 2. Tracer Bullet
Escribe **una** prueba que verifique **una** cosa:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
Esta es una "bala trazadora": dispárela primero y compruebe la mira. Matt enfatizó que este trabajo debe ser **de un extremo a otro**: no escribir el esquema primero, luego escribir la API y luego escribir la interfaz de usuario, sino cortar la ruta más delgada que recorra toda la pila.
#### 3. Incremental Loop
Repita para cada comportamiento posterior ROJO→VERDE:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
Reglas:
* Una prueba a la vez
* Escribe sólo el código suficiente para pasar la prueba actual.
* **No predecir pruebas futuras**
* Las pruebas se centran en el comportamiento observable.
"No predecir" es particularmente importante. La IA no puede evitar pensar: "Esta función debe ser compatible con X de todos modos, agreguémosla por cierto", y esto inicia el corte horizontal.
#### 4. Refactor
Una vez pasadas todas las pruebas, busque oportunidades de refactorización:
```
[ ] Extract duplication
[ ] Deepen modules (move complexity behind simple interfaces)
[ ] Apply SOLID principles where natural
[ ] Consider what new code reveals about existing code
[ ] Run tests after each refactor step
```
> **Never refactor while RED.** Get to GREEN first.
Refactorizar en rojo = cambiar pruebas y código al mismo tiempo = no sabes si la prueba es incorrecta o el código es incorrecto. **Primero verde, luego refactorizar**.
### Per-Cycle Checklist
Al final de cada ciclo rojo y verde, Matt le pide a la IA que se autoverifique:
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
Estos cinco puntos se utilizan para identificar malas pruebas y una implementación excesiva. La autocomprobación de la IA puede evitar los errores más comunes.
## Uso real: del problema al PR
`/tdd` es el siguiente paso en el flujo de trabajo de Matt desde `/to-issues`. Dado un problema de corte vertical, el proceso es:
```
你: 实现 issue #43
↓
/tdd
↓
Claude 读 issue acceptance criteria
↓
Claude 探索代码库 → 找到 CONTEXT.md → 用项目术语
↓
Planning 阶段:
- 列出准备改的接口
- 列出准备测的行为(按优先级排序)
- 让你点头
↓
Tracer Bullet:
- RED: 写第一个测试(基于 acceptance criteria 第 1 条)
- 跑测试,确认 fail
- GREEN: 写最小实现
- 跑测试,确认 pass
↓
Incremental Loop:
- 每个 acceptance criteria 一个 RED→GREEN
↓
Refactor:
- 看 deep module 提取机会
- 每次重构后跑全套测试
↓
PR
```
En cada ciclo rojo y verde, la IA se detendrá y le dará un estado: "la prueba falla"/"la prueba pasa, aquí está la diferencia". **Estas pausas son el antídoto para dejar atrás los faros**: la IA no tiene ninguna posibilidad de trazar mil líneas de una sola vez.
## Acerca de Mock: la fuerte opinión de Matt
SKILL.md menciona específicamente los peligros de las burlas; también proporciona un `mocking.md` por separado. Ideas centrales:
> "Bad tests... mock internal collaborators."
Colaboradores internos simulados = acoplamiento 1:1 entre pruebas e implementación = el equipo de pruebas se arrodilla al refactorizar. La preferencia de Matt es **pruebas de estilo de integración**: intente utilizar una base de datos real (en memoria o contenedores de prueba), HTTP real (MSW) y un sistema de archivos real (tmp dir). Solo burlándose de límites realmente costosos o inestables (como llamadas a la API de OpenAI).
Esto es contrario a la situación actual de muchos equipos: la mayoría de las bibliotecas de códigos están llenas de pruebas unitarias y hay más simulaciones que código real. Matt emitió un juicio en su discurso: **Buena base de código = base de código fácil de probar**. Si tienes que simular un montón de cosas para probar, significa que hay un problema con la estructura del código y primero debes cambiar la arquitectura (ve a [`/improve-codebase-architecture`](/es/docs/notes/matt-pocock-skills/improve-codebase-architecture)).
## El nuevo significado de TDD en la era de la IA
Cuando Kent Beck escribió ese libro hace 23 años, el beneficio principal de TDD era que "la gente no escribía código incorrecto". En la era de la IA, TDD tiene un significado adicional:
**Es el único "criterio de éxito" que la IA puede entender**.
Matt citó las sabias palabras de Karpathy más adelante en su discurso:
La mejor forma de "criterios de éxito" es **prueba**: es verificable por máquina, binario y no se puede cuestionar. Darle a la IA un conjunto de pruebas + "déjelo pasar" es diez veces más confiable que darle a la IA una descripción de los requisitos + "implemente".
Por lo tanto, `/tdd` no es sólo una herramienta de control de calidad: es la interfaz de entrada al bucle del agente. Cada ciclo rojo y verde es una "entrada → acción → retroalimentación" completa. La IA aprende la verdadera situación de esta implementación en el ciclo y el siguiente ciclo será más preciso.
## Cómo instalar y usar
```bash
npx skills@latest add mattpocock/skills
```
Marque `tdd` + `setup-matt-pocock-skills`.
Si ahora utiliza principalmente Codex, tome el control del producto de instalación en `.agents/skills/` y escriba el flujo de trabajo a nivel de proyecto, los comandos de prueba y las reglas de seguimiento de emisiones en `AGENTS.md`. La esencia de `/tdd` de Matt es el ciclo de refactorización rojo-verde, no vinculado al Código Claude.
**Método de llamada**:
* Directo: `/tdd` - deja que infiera qué medir a partir del contexto de la conversación actual
* Recuperar problema: `/tdd implement #43` - buscará el problema y luego lo abrirá
* Corrección de error: `/tdd reproduce this bug then fix it` - Primero escribirá una prueba fallida que pueda reproducir el error y luego lo solucionará.
## Notas
**No apto para todas las tareas**. Scripts únicos, código de exploración del área de juegos, ajuste de la interfaz de usuario: no use TDD, ralentizará el ritmo. El propio Matt dijo que TDD es adecuado para código que "tiene un valor duradero y necesita ser mantenido".
**Primero prepare la infraestructura de prueba**. Si el proyecto no ha instalado el marco de prueba (Vitest / Jest / Playwright, etc.), instálelo primero y luego use `/tdd`; de lo contrario, lo instalará primero, pero hay muchas preguntas en ese paso.
**No permita que agregue automáticamente pruebas e2e**. e2e es lento y nítido, y el ritmo TDD es de nivel de minutos. `/tdd` tiene como valor predeterminado la prueba de integración en lugar de e2e, pero puede indicarle explícitamente "solo unidad + integración, no e2e".
**La fase de reconstrucción es la más fácil de salir de control**. Después de que la IA obtenga el estado VERDE, refactorizará un montón de cosas con entusiasmo: mírelas fijamente y ejecutará pruebas después de cada refactorización. Esta parte es un área de alto riesgo para la desviación de la IA.
## Recursos de referencia
Siguiente artículo: [Mejorar la arquitectura de la base de código: reconstruir módulos superficiales en módulos profundos](/es/docs/notes/matt-pocock-skills/improve-codebase-architecture) - El mantenimiento periódico permite que la IA se ejecute en su base de código a largo plazo.
# to-PRD + to-Issues: Condensar la conversación de la parrilla en una porción vertical ejecutable
## La posición de esta sección en el flujo de trabajo.
Volvamos al diagrama de flujo de trabajo de Matt:
```
/grill-me 或 /grill-with-docs ← 谈清楚
↓
/to-prd ← 凝固成 PRD(你在这里)
↓
/to-issues ← 切成可领取的 vertical slice(你在这里)
↓
/tdd ← 一个 slice 一个 slice 跑红绿
↓
/improve-codebase-architecture
```
`/to-prd` y `/to-issues` son un vínculo entre lo anterior y lo siguiente: traducir decisiones abstractas de diálogo\*\* en unidades de trabajo ejecutables\*\*. Los combino porque en el uso real de Matt son dos pasos consecutivos.
## Modo de falla: nada que hacer después de asar
Muchas personas se quedan estancadas después de usar `/grill-me`: tienen una larga conversación con un montón de decisiones, pero **¿cómo empezar a escribir código**?
Darle a Claude un "por favor hazlo" directamente está mal, porque:
1. Implementar funciones completas a la vez = salida de IA de más de 1000 líneas = difícil de revisar, difícil de probar, difícil de localizar errores
2. Después de que la IA pierde la memoria, la siguiente sesión no tiene contexto.
3. Sin seguimiento: no hay forma de saber dónde hemos llegado y cuánto queda
El enfoque correcto es congelar la decisión en un artefacto (PRD) y luego dividir el PRD en paquetes de trabajo (problemas) que sean lo suficientemente pequeños como para completarse de forma independiente. Esto ha sido de conocimiento común en la ingeniería de software durante 30 años, pero adquiere un nuevo significado en la era de la IA:
> Píquelo lo suficientemente fino como para que el agente AFK (el agente que se ejecuta cuando usted no está presente) pueda recogerlo y completarlo de forma independiente.
## /to-prd: comprime la conversación en PRD
### Limitaciones clave de las habilidades
SKILL.md de `/to-prd` escribe una oración muy importante al principio:
> "This skill takes the current conversation context and codebase understanding and produces a PRD. **Do NOT interview the user — just synthesize what you already know.**"
No hagas más preguntas. Esto es lo que hace la etapa grill-me, to-prd solo hace **síntesis**. Así que **no borre el contexto y ejecute-prd**: se basa en todo el diálogo de la parrilla anterior.
### Flujo de procesamiento de habilidades
1. **Explore la base del código** (si aún no la ha explorado): utilice el vocabulario CONTEXT.md del proyecto y respete las ADR existentes.
2. **Borrador del módulo**——Busque proactivamente oportunidades que puedan extraerse como módulos profundos para que la interfaz se pueda probar de forma independiente
3. **Alinear módulos con usuarios** - "¿Son correctos estos módulos? ¿Cuáles deben probarse?"
4. **Genere PRD** según la plantilla, envíelo al rastreador de problemas y etiquételo con `needs-triage`
### Plantilla PRD
La plantilla proporcionada por Matt es la siguiente:
```markdown
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list:
1. As a , I want a , so that
2. ...
## Implementation Decisions
- The modules that will be built/modified
- The interfaces of those modules
- Technical clarifications from the developer
- Architectural decisions
- Schema changes / API contracts / Specific interactions
(NO specific file paths or code snippets — they rot fast.)
## Testing Decisions
- What makes a good test (test external behavior, not internals)
- Which modules will be tested
- Prior art (similar tests in the codebase)
## Out of Scope
What's NOT in this PRD.
## Further Notes
```
Varios diseños clave:
* **Las historias de usuarios representan la mayoría**: requisitos LARGA lista numerada, lo que le obliga a enumerar exhaustivamente los puntos de función completos. Esto evita el punto ciego de "pensé que estaba claro".
* **La implementación no escribe rutas de archivo ni código**: Matt dijo directamente "pueden terminar quedando obsoletos muy rápidamente". Esta es una consideración única en la era LLM: la ruta específica quedará obsoleta inmediatamente después de la reconstrucción, pero el "límite del módulo" y el "contrato de interfaz" tienen un ciclo de vida más largo.
* **Debe tener fuera de alcance**: este párrafo se ignora en la mayoría de las plantillas del PRD, pero es el seguro de límites cuando se cortan temas más adelante.
## /to-issues: Cortar el PRD en rodajas verticales
### ¿Qué es el corte vertical?
Este es uno de los conceptos más importantes de toda la metodología de Matt. SKILL.md dijo directamente:
> Each issue is a thin vertical slice cutting through ALL integration layers end-to-end, NOT a horizontal slice of one layer.
El ejemplo más claro es crear una “función de comentario”.
**Corte horizontal (forma incorrecta)**:
* Problema 1: esquema de base de datos
* Problema 2: puntos finales API
* Problema 3: componentes de la interfaz de usuario
* Número 4: Pruebas
**Corte longitudinal (método Tracer Bullet)**:
* Problema 1: "Los visitantes pueden enviar un comentario anónimo" (se incluyen esquema + API + UI + prueba, pero el alcance es tan pequeño que solo puede ser anónimo)
* Problema 2: "Los comentarios de los usuarios que han iniciado sesión están asociados con la cuenta"
* Problema 3: "Se pueden responder los comentarios"
* Problema 4: "Los administradores pueden eliminar comentarios"
El problema del corte horizontal: cada corte no se puede probar individualmente. Después de completar el Número 1, no había nada que demostrar, y no fue hasta el Número 4 que se pudo ejecutar todo el enlace; fue sólo entonces que descubrí que el diseño del esquema era incorrecto.
Cada segmento vertical completado es un **subconjunto funcional disponible de extremo a extremo**, denominado **bala rastreadora** (bañera rastreadora) en palabras del programador pragmático. Dispare primero una ronda para comprobar la mira y luego ajuste la siguiente ronda.
### HITL vs AFK
`/to-issues` también etiqueta cada segmento:
* **HITL** (Human in the Loop): requiere que las personas participen en la toma de decisiones. Como decisiones arquitectónicas, revisiones de diseño.
* **AFK** (Ausente del teclado): el agente puede completar el trabajo de forma independiente y usted puede regresar para ver los resultados.
> "Prefiere AFK a HITL siempre que sea posible".
Esta es una idea muy radical en el flujo de trabajo de Matt: después de terminar de cortar el problema, lo envía directamente a un agente que se ejecuta cuando usted no está presente (como por la noche o los fines de semana). Cuando regresa al día siguiente, el PR ya está allí esperando ser revisado. La parte HITL permanece durante el día y trabaja con el agente.
### Enlace de confirmación de corte
`/to-issues` no generará un problema tan pronto como aparezca; primero le mostrará el esquema de ordenamiento en mosaico como una lista numerada:
```
1. Title: 访客提交匿名评论
Type: AFK
Blocked by: None
User stories covered: #1, #2
2. Title: 评论关联到登录账户
Type: AFK
Blocked by: #1
User stories covered: #3
3. Title: 评论审核流程
Type: HITL(需要确认审核 UI 设计)
Blocked by: #1
User stories covered: #4, #5
```
Entonces pregúntale:
* ¿La granularidad es correcta? ¿Demasiado grueso/demasiado fino?
* ¿Son correctas las dependencias?
* ¿Cuáles deberían fusionarse/dividirse?
* ¿Son correctas las marcas HITL/AFK?
Itere hasta que asienta antes de enviarlo al rastreador de problemas, en orden de dependencia (bloqueador primero), para que los problemas posteriores puedan hacer referencia al ID de problema real del primer problema.
### Plantilla de problema
```markdown
## Parent
A reference to the parent issue (if any).
## What to build
A concise description. Describe end-to-end behavior, NOT layer-by-layer
implementation.
## Acceptance criteria
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
## Blocked by
- A reference to the blocking ticket
(or "None - can start immediately")
```
Tenga en cuenta la línea "describir el comportamiento de un extremo a otro", que es coherente con el espíritu del corte vertical. Los criterios de aceptación son una lista de aceptación, que Claude convertirá en pruebas una por una en la etapa `/tdd`.
## Cómo utilizar: ejemplo de proceso completo
Suponga que desea agregar funcionalidad de comentarios a su blog. Proceso completo:
```
你: 我想给博客加评论功能
↓
/grill-me → Claude 问 30 个问题(要不要登录?匿名?嵌套?审核?……)
↓
你回答完毕,达成共识
↓
/to-prd → Claude 生成结构化 PRD,提交到 GitHub Issues #42
↓
/to-issues → Claude 提议切成 4 个 vertical slice
让你确认粒度和依赖
你点头
按依赖顺序发布到 GitHub Issues #43~#46
↓
你回家睡觉
↓
夜里 AFK agent 抓 #43(无依赖),跑 /tdd 完成 → 提 PR
你早上 review、merge
↓
agent 抓 #44 / #45 ……
```
Todo el proceso no requiere que te sientes frente a la pantalla y observes cada detalle. Las decisiones clave se toman en la etapa de interrogarme.
## Instalación y requisitos previos
```bash
npx skills@latest add mattpocock/skills
```
Marque `to-prd`, `to-issues`, `setup-matt-pocock-skills`.
**Primero debe ejecutar `/setup-matt-pocock-skills`**: escribirá su rastreador de problemas (GitHub/GitLab/markdown local) y el vocabulario de etiquetas de clasificación en AGENTS.md/CLAUDE.md; de lo contrario, to-prd y to-issues no sabrán dónde enviar los problemas.
Rastreadores de problemas admitidos:
* **Problemas de GitHub** (predeterminado, usa `gh` CLI)
* **Problemas de GitLab** (usando `glab` CLI)
* **Rebaja local** (crear archivos en `.scratch//`): adecuado para proyectos personales o proyectos sin control remoto
* **Otros** (Jira, Linear, etc.) - Utilice una prosa para describir el flujo de trabajo y la habilidad se llamará según su descripción.
## Preguntas frecuentes
\*\*P: Ya existe un PRD ya preparado. ¿Puedo saltar a prd e ir directamente a issues? \*\*
R: Sí. `/to-issues` acepta una referencia de problema como parámetro ("Dividir el problema n.° 42 en sectores verticales"), buscará el contenido del problema y luego lo dividirá.
\*\*P: Mi proyecto no utiliza el rastreador de problemas, ¿se puede utilizar? \*\*
R: Sí. Seleccione "Rebaja local" durante la configuración y todos los problemas se convertirán en archivos locales como `.scratch//001-foo.md`.
\*\*P: ¿Qué debo hacer si el PRD es demasiado largo y la IA por sí sola no puede manejarlo? \*\*
R: Esto es un signo de granularidad de corte: el PRD debe dividirse en múltiples PRD independientes en lugar de un PRD gigante. Deberías sentirlo durante la etapa de parrilla: si hablas del número 50 y todavía estás introduciendo nuevas funciones, detente primero, córtalo en dos PRD y hazlo en lotes.
\*\*P: ¿Cómo detecta automáticamente el agente AFK los problemas? \*\*
R: El repositorio de Matt no proporciona esta parte y debe coordinarse con su propio acuerdo de agente (por ejemplo, GitHub Actions activa Claude Code para ejecutar problemas). Lo más fácil es hacer cron para verificar `is:open no:assignee label:agent-ready` cada hora.
## El verdadero valor de este proceso
`/to-prd` y `/to-issues` parecen "gestión de proyectos automatizada", pero Matt los sitúa en el centro de su flujo de trabajo por una razón más profunda:
**Te obliga a pensar "**Qué es una cosita completa**"**. Cuando te ves obligado a cortar la funcionalidad en cortes verticales, estás haciendo una de las cosas más difíciles en ingeniería de software: encontrar uniones. El pensamiento en sí es más valioso que estas dos habilidades.
Y el tamaño del corte vertical es el tamaño que la IA puede manejar al mismo tiempo. **Deje que el tamaño de la tarea coincida con el límite superior de las capacidades de la IA**: este es el ritmo fundamental de colaboración con LLM.
## Recursos de referencia
Siguiente artículo: [TDD: Utilice la reconstrucción roja y verde para obligar a la IA a dar pequeños pasos](/es/docs/notes/matt-pocock-skills/tdd)——Una vez solucionado el problema, cómo hacer que la IA realmente dé pequeños pasos.
# Zoom Out:当你迷路时,让 AI 先画地图
## 最短,但很有用
`/zoom-out` 可能是 Matt 这套稳定 engineering skill 里最短的一个。它的核心指令可以概括成一句:
> 我不熟悉这块代码,请上升一层抽象,用项目领域语言给我画出相关模块和调用者地图。
它不是用来写代码的,也不是用来重构的。它用来处理一种很常见的状态:**你和 AI 都已经钻进某个文件,但开始忘记这个文件为什么存在**。
## 它解决的失败模式
AI 编程很容易进入局部最优:
1. 用户指了一个文件
2. AI 读这个文件
3. AI 根据局部代码猜意图
4. 改完之后才发现上游调用者、领域规则或 ADR 不支持这个改法
人类也一样。我们 debug 久了会盯住一个函数,忘掉它在系统里的位置。
`/zoom-out` 的作用是打断这种隧道视野。它让 agent 暂停实现,先回答:
* 这块代码属于哪个领域概念?
* 谁调用它?
* 它调用谁?
* 它背后有哪些不变量?
* 它和 `CONTEXT.md` 里的术语怎么对应?
* 它是否受某个 ADR 约束?
## 和 /improve-codebase-architecture 的区别
`/zoom-out` 和 [`/improve-codebase-architecture`](/docs/notes/matt-pocock-skills/improve-codebase-architecture) 都会看系统全局,但目标完全不同。
| Skill | 目标 | 输出 |
| -------------------------------- | ---------- | ----------------------- |
| `/zoom-out` | 帮你理解一块陌生代码 | 地图、调用关系、领域解释 |
| `/improve-codebase-architecture` | 找可深化的架构机会 | 候选重构、deletion test、接口设计 |
`/zoom-out` 更像「请给我讲讲这块」。它不应该急着提出重构方案,更不应该直接改代码。它的任务是降低认知负担。
## 为什么要强调领域词汇
这个 skill 明确要求使用项目的 domain glossary。原因是:如果只用文件名解释,AI 很容易输出这种东西:
```text
OrderService 调用 OrderRepository,然后 OrderRepository 调用 db client。
```
这听起来像解释了,其实没解释。更有用的地图应该长这样:
```text
Checkout flow 里,Order Draft 是用户尚未支付前的临时订单。
Order Finalization 会把 Draft 转成不可变 Order,并触发 Inventory Reservation。
`OrderService.finalize()` 是这个转换的 seam,调用者主要来自 Payment Callback 和 Admin Retry。
```
第二种解释把代码放回了业务语言里。你不只是知道「谁调用谁」,还知道「它为什么存在」。
## 适合什么时候用
我建议在这些场景下主动调用 `/zoom-out`:
* 接手一个陌生模块前
* 改一个 bug,但还不确定相关调用链
* review AI 生成代码,看不出它是否改到了正确层级
* 准备写 PRD,想确认模块边界
* 已经看了 3 个文件还没形成系统图
* 准备跑 `/improve-codebase-architecture`,但还不确定候选区域
它特别适合作为「动手前的 5 分钟」。有些 bug 不是因为代码难,而是因为一开始就看错了层级。
## 一个可复用的输出格式
虽然原始 skill 极短,但我建议在使用时让 AI 按这个格式输出:
```markdown
## 这块代码在系统里的位置
## 关键领域术语
## 主要模块
| 模块 | 责任 | 调用者 | 被调用对象 |
|---|---|---|---|
## 关键流程
## 已知约束 / ADR
## 我建议你先看的文件
```
这个格式比普通解释更稳定,也更适合转成后续 `/to-prd` 或 `/diagnose` 的上下文。
## 不要把它用成计划模式
`/zoom-out` 的危险是:AI 讲完地图后,顺手开始建议「可以这么改」。如果你只是想理解代码,应该明确限制:
```text
只解释结构,不提出实现方案,不改文件。
```
因为它的价值就在于把决策和理解分开。理解不清时提方案,往往只是把误解包装得更漂亮。
## 我的使用建议
`/zoom-out` 很适合和其他 skill 组合:
* `/zoom-out` → `/diagnose`:先看系统地图,再建反馈回路
* `/zoom-out` → `/grill-with-docs`:先理解现有领域语言,再拷问新需求
* `/zoom-out` → `/to-prd`:先确认模块位置,再写 PRD
* `/zoom-out` → `/improve-codebase-architecture`:先画地图,再找 shallow/deep 问题
它不是完整流程,只是一个刹车。AI 开始在局部文件里越改越多时,先让它 zoom out,通常能省掉后面一轮返工。
## 参考资源
下一篇:[Prototype:用可丢弃代码回答一个设计问题](/docs/notes/matt-pocock-skills/prototype)。
# Pi Agent 是什么
## 引言
如果只看功能,Pi Agent 很容易被低估:它在终端里运行,能读文件、改文件、执行命令、保存会话,也能切换模型。听起来像另一个 Claude Code 或 Codex。
但我觉得 Pi 真正有意思的地方,不是它多做了什么,而是它少做了什么。它把 AI 编程工具最核心的那层保留下来:模型、上下文、工具、会话、扩展,然后尽量不把用户的工作流提前写死。
所以我更愿意这样理解 Pi:
**Pi Agent 不是一个“更全”的 AI 编程产品,而是一个更薄、更透明的 coding agent harness。**
这个判断比功能清单更重要。因为它决定了你应该怎么学习 Pi:不是先背命令,而是先理解一个 coding agent 到底由哪几层组成。
## 先把 Pi 放对位置
一个 AI 编程工具通常可以先粗略分成三层:模型、harness、工程环境。但如果只画这三层,还是太抽象。Pi 真正值得看的,是中间这层 harness 里面又拆成了哪些模块:
这张图里有三个重点。
第一,Pi 不是只有一个“聊天 UI”。CLI、交互 TUI、print/JSON、RPC、SDK 都只是入口,真正承接任务的是 `AgentSessionRuntime` 和 `AgentSession`。
第二,Pi 在请求模型之前会先做资源加载。`ResourceLoader` 会把 `AGENTS.md`、`CLAUDE.md`、skills、extensions、prompt templates 这些东西整理出来,再交给 `SystemPrompt Builder` 组装成模型真正看到的上下文。
第三,模型调用工具时,并不是模型直接控制文件系统。`AgentHarness` 和 `AgentLoop` 负责校验工具、执行工具、接住结果、继续下一轮。Extensions、Tool Registry、SessionManager 则在旁边扩展能力和保存状态。
所以 Pi 站在中间,但这个“中间”不是一句空话。它具体控制的是:哪些上下文进入模型,哪些工具可以被调用,工具结果如何回到会话,哪些能力由扩展补进来。
这也是为什么很多介绍 Pi 的文章都会强调 minimal、transparent、extensible。它们其实在说同一件事:Pi 试图把 agent 的核心运行层做小,让用户能看见,也能改。
## 一次工作是怎么流动的
Pi 的一次请求不是“问模型一句话,模型回一句话”。更准确地说,它是一段带分支的时序:
这里最关键的是第四步到第五步之间的来回。模型不直接碰你的文件系统,它只提出 `tool_call`;Pi 接住这个调用,校验工具名和参数,触发可能存在的 extension hooks,执行真实操作,再把 `tool_result` 放回上下文。模型再根据新的上下文判断下一步。
这就是 coding agent 和普通聊天机器人的区别。聊天机器人主要在文本里完成任务;coding agent 要进入工程系统,所以它必须有 harness 来管理工具、上下文和状态。
Pi 的默认工具很少:
| 工具 | 含义 |
| ------- | ----------- |
| `read` | 读取文件 |
| `edit` | 修改已有文件 |
| `write` | 创建或覆盖文件 |
| `bash` | 执行 shell 命令 |
还有 `grep`、`find`、`ls` 这类只读工具可以被启用或限制。这个工具集看起来克制,但已经形成了编程闭环:读代码、改代码、跑测试、根据错误继续修。
这套设计背后的问题不是“Pi 会不会做更多”,而是“更多东西是否应该默认进核心”。Pi 的回答很明确:不一定。
## 为什么它不急着内置很多功能
很多 AI 编程产品会把计划模式、todo、子代理、MCP、权限弹窗、后台任务、浏览器工具都做进产品里。这样上手快,但也带来一个代价:你很难知道模型实际收到了什么上下文,也很难把产品工作流改成自己的工作流。
Pi 的路线相反。它把核心保持得很小,然后把工作流放到外面:
| 你想改变什么 | Pi 交给哪里 |
| --------- | ------------------------- |
| 项目规则 | `AGENTS.md` / `CLAUDE.md` |
| 专门任务方法 | Skills |
| 自定义工具和 UI | Extensions |
| 一组可分享能力 | Pi Packages |
| 模型选择 | Provider / Model 配置 |
这不是“功能不够”,而是一种产品取舍:核心只管 agent loop,具体工作流交给用户和团队自己组合。
举个例子,Pi 没有默认内置 DeepSearch。但这并不意味着它不能做深度搜索。更符合 Pi 思路的做法,是写一个 extension:注册一个 `deep_search` 工具,把 Tavily、Exa、Brave Search 或公司内部搜索接进去,再让模型在需要时调用它。
这和把“搜索按钮”硬编码进产品不同。前者是你在扩展 agent 的能力,后者是产品替你决定工作流。
## 重点:怎么扩展自己的工作流
如果要真正用 Pi,而不只是“体验一下”,重点一定是扩展自己的工作流。
这里容易混淆的是,Pi 不是只有“插件”这一种扩展方式。它更像给你四层入口:项目规则、任务方法、真实工具、可分享包。你要先判断自己想沉淀的到底是哪一种东西。
| 你要沉淀的东西 | 用什么 | 适合什么场景 |
| ------- | ------------------------- | ------------------------------------------------------ |
| 项目习惯和约束 | `AGENTS.md` / `CLAUDE.md` | 告诉 agent 怎么改代码、跑什么检查、哪些目录不能碰 |
| 一套可复用方法 | Skill | 代码审查、写文章、发版、生成文档、图片处理这类“步骤和经验” |
| 一个真实能力 | Extension | 注册工具、拦截工具调用、加 slash command、加 UI、接外部 API |
| 一组可分发能力 | Pi Package | 把 extensions、skills、prompt templates、themes 打包给自己或团队复用 |
我的理解是:**Skill 是工作手册,Extension 是可执行插件,Package 是分发容器。**
比如我现在这个博客工作流,可以这样拆:
| 工作流需求 | 放进哪里 |
| ----------------------------------------------------------- | ----------------------- |
| “写内容只写中文,图片必须用 `BlogImage`,改 MDX 后跑 `pnpm types:check`” | `AGENTS.md` |
| “写概念文章时按误解、定义、机制、例子、边界来组织” | `article-writing` Skill |
| “给 Pi 增加一个 `deep_search` 工具,能查 Tavily / Exa / Brave Search” | Extension |
| “把写作 Skill、DeepSearch Extension、微信发布命令打包给多个项目用” | Pi Package |
这就比单纯说“装插件”更准确。因为很多工作流不需要写代码,只需要一份好的规则或 Skill;但只要你希望 agent 真的多一个能力,比如查外部搜索、查数据库、调用 CI、拦截危险命令,就应该写 Extension。
### Extension:真正的插件层
Pi 的 Extension 是 TypeScript 模块。它可以做几类事:
| 能力 | 例子 |
| ----- | ------------------------------------------ |
| 注册工具 | `deep_search`、`query_logs`、`open_issue` |
| 注册命令 | `/review`、`/publish`、`/checkpoint` |
| 拦截事件 | 在 `bash` 执行 `rm -rf`、`sudo`、写 `.env` 前要求确认 |
| 改 UI | 在 TUI 里显示状态、选择框、确认框、任务面板 |
| 保存状态 | 记录 todo、连接池、上次搜索结果、任务阶段 |
| 接外部系统 | CI、GitHub、日志系统、公司内部 API |
Extension 可以放在全局,也可以放在项目里:
```text
~/.pi/agent/extensions/ # 全局扩展,所有项目可用
.pi/extensions/ # 项目扩展,只在当前项目里用
```
测试一个临时扩展,可以用:
```bash
pi -e ./my-extension.ts
```
放到自动发现目录后,可以在 Pi 里用:
```text
/reload
```
重新加载 extensions、skills、prompts 和 context files。
这就是我觉得 Pi 最有价值的地方:你不只是“让模型帮你写代码”,而是在给模型设计一个可控的工作环境。Extension 决定模型能调用什么能力,hooks 决定哪些行为要被拦截,commands 决定你自己的工作流如何被触发。
### Skill:不要把所有东西都写成插件
如果一个能力主要是“怎么做”,而不是“调用一个真实 API 或执行一段程序”,那它更适合写成 Skill。
Skill 的结构通常是:
```text
my-skill/
SKILL.md
scripts/
templates/
references/
```
Pi 启动时不会把完整 skill 全塞进上下文。它会先加载 skill 的名字和描述;当任务匹配时,再让模型读取完整的 `SKILL.md`。这叫渐进式披露。好处是:你可以保存复杂方法论,但不用每次都污染上下文。
比如“写一篇好文章”“发布到公众号”“做一次浏览器 QA”,这些都更像 Skill。它们的价值主要在步骤、判断标准和参考材料,不一定需要注册一个 LLM 可调用工具。
### Package:把自己的工作流打包
当你已经有一组稳定能力,就可以考虑把它做成 Pi Package。
Package 可以包含:
| 内容 | 作用 |
| ---------------- | ----------------- |
| extensions | 可执行插件、工具、命令、hooks |
| skills | 工作方法和任务手册 |
| prompt templates | 常用 prompt 模板 |
| themes | TUI 主题 |
安装方式大概是:
```bash
pi install npm:@scope/my-pi-package
pi install git:github.com/user/repo@v1
pi install ./relative/path/to/package
pi list
pi remove npm:@scope/my-pi-package
pi update --extensions
```
默认安装会写到个人设置里。如果你想让团队项目共享,可以用项目级设置,让包记录在 `.pi/settings.json`。这样别人进入项目启动 Pi 时,也能自动补齐缺失的 package。
但这里也要非常谨慎:Package、Extension、Skill 都可能影响 agent 的行为。第三方 package 不是浏览器插件那种低权限装饰,它可能运行代码,也可能指导模型执行命令。安装前应该看源码。
所以我会按这个顺序学习 Pi 的扩展:
1. 先用 `AGENTS.md` 写清项目规则。
2. 再把重复方法做成 Skill。
3. 需要真实工具能力时,再写 Extension。
4. 多项目复用时,最后再打成 Package。
这样学习比较稳。你不是一上来就写插件,而是先把工作流拆成“规则、方法、工具、分发”四类,再决定每一类放到 Pi 的哪一层。
## 从源码里能看到什么
我看 Pi 源码时,最有帮助的不是追每个函数,而是看几个文件各自代表哪层设计:
| 源码位置 | 说明 |
| ---------------------------------------------------- | ------------------------------------------------------- |
| `packages/agent/src/agent-loop.ts` | 核心循环:把用户消息、模型响应、工具调用和工具结果串起来 |
| `packages/agent/src/harness/agent-harness.ts` | Harness 状态:管理 session、system prompt、tools、hooks 和消息队列 |
| `packages/coding-agent/src/core/tools/index.ts` | 内置工具集合:默认 coding tools 是 `read`、`bash`、`edit`、`write` |
| `packages/coding-agent/src/core/resource-loader.ts` | 资源加载:读取项目指令、extensions、skills、prompt templates 和 themes |
| `packages/coding-agent/src/core/system-prompt.ts` | 系统提示词构建:把工具说明、项目上下文、skills 和当前目录放进 prompt |
| `packages/coding-agent/src/core/extensions/types.ts` | 扩展系统:允许扩展注册工具、命令、快捷键、UI 和生命周期事件 |
这几块合起来,基本就是 Pi 的心脏:它先组装上下文和工具,再把请求交给模型;模型如果要调用工具,Pi 执行工具;结果回来后,循环继续。
所以 Pi 的“极简”不是空口号。源码结构本身也在表达这个想法:把 agent loop、harness、coding tools、资源加载、扩展系统分开,每层都相对清楚。
## Pi 的边界
Pi 的自由度很高,但自由度不等于安全。
Pi packages 和 extensions 可以运行代码;skills 也可能指导模型执行脚本;`bash` 能触达你的真实系统。官方文档和安全分析都提醒过:第三方 package、extension、skill 需要自己审查。
我会把 Pi 的边界理解成三点:
1. **它不是沙箱**:不要把它当成天然隔离环境。危险项目最好放进容器、临时目录或干净 worktree。
2. **它不替你判断权限**:Pi 的核心哲学不是靠一堆弹窗管理风险,而是让你控制工具、上下文和扩展。
3. **它适合懂工程边界的人**:你需要知道什么时候让 agent 跑命令,什么时候只给 read-only 工具,什么时候先建 git checkpoint。
这也是 Pi 和一些更产品化 agent 的差异。产品化工具会替你包装更多安全和交互细节;Pi 给你更直接的控制权,同时也把更多责任还给你。
## 应该怎么学习 Pi
学习 Pi,我不建议从“有哪些命令”开始。命令很快能查到,真正值得学的是这几个问题:
| 问题 | 为什么重要 |
| ----------------------- | ------------------- |
| Pi 怎么组装上下文 | 决定模型实际知道什么 |
| Pi 的工具面为什么这么小 | 决定 agent 的行为是否可观察 |
| Extension 怎么注册工具 | 决定你能不能把自己的工作流接进去 |
| Skill 和 Extension 有什么区别 | 决定什么时候写说明,什么时候写代码 |
| Session 怎么保存和分叉 | 决定一次工程探索能不能恢复、回看和继续 |
如果你已经用过 Claude Code 或 Codex,可以把 Pi 当成一次“拆开看”的机会:同样是让模型写代码,为什么有的工具像黑箱产品,有的工具像可改造的运行时?
这个问题比“Pi 能不能替代某个工具”更值得问。
## 写在最后
Pi Agent 最有价值的地方,是它把 AI 编程工具的中间层暴露出来了。
它提醒我们:一个 coding agent 的能力,不只来自模型,也来自 harness 的设计。模型负责想,harness 负责让模型在真实工程环境中行动。上下文怎么进来,工具怎么出去,结果怎么回来,扩展怎么插入,这些细节共同决定了 agent 是否可靠、透明、可控。
所以我不会把 Pi 简单看成“Claude Code 平替”。它更像一个适合开发者研究和改造的 agent runtime。你可以直接用它写代码,也可以用它学习怎样设计自己的 agent 工作流。
下一篇实战我会按这个思路继续:不写一个普通教程,而是用 Pi Extension 做一个 `deep_search` 工具,看看怎样把外部搜索能力接进 agent loop。
## 延伸阅读
# Pi Agent 实践指南
## 快速回顾
在概念篇里,我把 Pi Agent 理解成一个极简的 **Agent Harness**:它连接模型、终端、文件系统、shell、会话和扩展系统,但不替你预设一整套厚重工作流。
所以实践篇我不想再做一个普通的"让 Pi 改文件"案例。那个案例能说明基础闭环,但不够体现 Pi 的可扩展性。
更适合 Pi 的实战案例,是给它补一个它默认没有、但很多人真实需要的能力:**DeepSearch**。
这里的 DeepSearch 不是简单的联网搜索,而是一套研究型工作流:
| 阶段 | 要做什么 |
| ---- | ----------------------- |
| 问题拆解 | 把一个模糊问题拆成几个可检索子问题 |
| 多轮检索 | 分别搜索官方文档、代码仓库、博客、讨论区或论文 |
| 来源筛选 | 去重、排除低质量结果、优先保留一手来源 |
| 证据整理 | 摘出关键事实、链接、时间、版本和不确定性 |
| 综合回答 | 给出结论,同时说明依据和限制 |
我的判断是:**DeepSearch 不应该写进 Pi 本体,也不应该只靠 prompt 硬凑。它更适合做成一个 Pi Extension。**
原因很简单:DeepSearch 涉及网络请求、第三方搜索 API、来源过滤、结果截断、引用格式和安全边界。这些都属于工作流能力,而不是 coding agent 的最小核心。
## 设计目标
这个案例要实现的不是一个完美的研究系统,而是一个可跑通的最小版本。
目标如下:
```text
给 Pi 增加一个 deep_search 工具。
它接收:
- query:用户要研究的问题
- depth:检索深度
- maxResults:最多返回多少条候选资料
它输出:
- 结构化搜索结果
- 每条结果的标题、URL、摘要、相关性
- 给模型使用的证据提示
Pi 拿到这些证据后,再由当前模型生成最终结论。
```
我会刻意把"检索"和"综合"拆开:
| 部分 | 由谁负责 | 原因 |
| --------- | -------------------- | ------------- |
| 搜索 API 调用 | DeepSearch extension | 这是确定性的外部能力 |
| 结果去重和截断 | DeepSearch extension | 避免上下文被噪声塞满 |
| 判断哪些证据重要 | Pi 当前模型 | 需要推理和上下文理解 |
| 最终答案写作 | Pi 当前模型 | 需要结合用户问题和项目语境 |
这样做更稳。Extension 不需要自己再调用一个模型,也不需要变成一个嵌套 agent。它只提供高质量证据,让 Pi 原本的模型继续推理。
## 准备工作
Pi extension 可以放在全局目录,也可以放在项目目录。这里我建议先放项目目录:
```text
.pi/extensions/deepsearch/
package.json
index.ts
```
项目本地 extension 的好处是边界清楚。这个 DeepSearch 能力只在当前项目里启用,不会影响所有 Pi 会话。
搜索服务可以选 Tavily、Exa、Brave Search、SerpAPI,甚至你自己的搜索后端。第一版不要纠结服务商,先抽象成一个 `searchWeb()` 函数。
例如用环境变量保存 API Key:
```bash
export TAVILY_API_KEY=tvly-...
```
如果你不想接第三方搜索 API,也可以先用本地 mock 数据把 extension 跑通。等工具注册、参数传递和结果格式都稳定后,再接真实搜索服务。
## Step 1: 创建 Extension 目录
先创建目录:
```bash
mkdir -p .pi/extensions/deepsearch
```
如果 extension 需要依赖,可以放一个 `package.json`:
```json
{
"name": "pi-deepsearch-extension",
"private": true,
"dependencies": {
"typebox": "*",
"@earendil-works/pi-ai": "*",
"@earendil-works/pi-coding-agent": "*"
},
"pi": {
"extensions": ["./index.ts"]
}
}
```
然后安装依赖:
```bash
cd .pi/extensions/deepsearch
npm install
```
Pi 的 extension 是 TypeScript 模块,不需要你先手动编译。这个体验很适合快速做工具实验。
## Step 2: 注册 deep\_search 工具
核心文件是 `.pi/extensions/deepsearch/index.ts`。
第一版可以这样写:
```typescript
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { StringEnum } from "@earendil-works/pi-ai";
import { Type } from "typebox";
type SearchResult = {
title: string;
url: string;
snippet: string;
score?: number;
};
export default function (pi: ExtensionAPI) {
pi.registerTool({
name: "deep_search",
label: "DeepSearch",
description: "Search the web for source-backed evidence about a question.",
promptSnippet: "Research a question with web search and return source-backed evidence.",
promptGuidelines: [
"Use deep_search when the user asks for current facts, external sources, comparison, investigation, or source-backed research.",
"After deep_search returns results, synthesize an answer with citations and clearly separate facts, inference, and uncertainty.",
"Do not treat deep_search results as final truth; inspect source quality and mention gaps."
],
parameters: Type.Object({
query: Type.String({
description: "The research question or search query."
}),
depth: Type.Optional(StringEnum(["quick", "normal", "deep"] as const)),
maxResults: Type.Optional(Type.Number({
minimum: 3,
maximum: 10,
default: 6
}))
}),
async execute(_toolCallId, params, signal) {
const depth = params.depth ?? "normal";
const maxResults = params.maxResults ?? 6;
const results = await searchWeb(params.query, depth, maxResults, signal);
return {
content: [
{
type: "text",
text: formatResultsForModel(params.query, results)
}
],
details: {
query: params.query,
depth,
results
}
};
}
});
}
async function searchWeb(
query: string,
depth: "quick" | "normal" | "deep",
maxResults: number,
signal: AbortSignal
): Promise {
const apiKey = process.env.TAVILY_API_KEY;
if (!apiKey) {
throw new Error("Missing TAVILY_API_KEY. Set it before starting pi.");
}
const response = await fetch("https://api.tavily.com/search", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
api_key: apiKey,
query,
search_depth: depth === "quick" ? "basic" : "advanced",
max_results: maxResults,
include_answer: false,
include_raw_content: depth === "deep"
}),
signal
});
if (!response.ok) {
throw new Error(`Search failed: ${response.status} ${response.statusText}`);
}
const data = await response.json() as {
results?: Array<{
title?: string;
url?: string;
content?: string;
score?: number;
}>;
};
return dedupeByUrl((data.results ?? []).map((item) => ({
title: item.title ?? "Untitled",
url: item.url ?? "",
snippet: item.content ?? "",
score: item.score
}))).filter((item) => item.url);
}
function dedupeByUrl(results: SearchResult[]): SearchResult[] {
const seen = new Set();
const deduped: SearchResult[] = [];
for (const result of results) {
const key = normalizeUrl(result.url);
if (seen.has(key)) continue;
seen.add(key);
deduped.push(result);
}
return deduped;
}
function normalizeUrl(url: string): string {
try {
const parsed = new URL(url);
parsed.hash = "";
parsed.searchParams.delete("utm_source");
parsed.searchParams.delete("utm_medium");
parsed.searchParams.delete("utm_campaign");
return parsed.toString();
} catch {
return url;
}
}
function formatResultsForModel(query: string, results: SearchResult[]): string {
if (results.length === 0) {
return `DeepSearch found no results for: ${query}`;
}
const lines = results.map((result, index) => {
return [
`## Source ${index + 1}`,
`Title: ${result.title}`,
`URL: ${result.url}`,
result.score === undefined ? undefined : `Score: ${result.score}`,
`Snippet: ${result.snippet}`
].filter(Boolean).join("\n");
});
return [
`DeepSearch query: ${query}`,
"",
"Use these sources as evidence. Cite URLs when making factual claims.",
"Separate confirmed facts from inference and uncertainty.",
"",
...lines
].join("\n\n");
}
```
这段代码只做最关键的事:
| 代码位置 | 作用 |
| ------------------------- | ----------------------- |
| `pi.registerTool()` | 把 `deep_search` 暴露给模型调用 |
| `parameters` | 告诉模型工具需要哪些参数 |
| `promptGuidelines` | 告诉模型什么时候用、用完后怎么处理 |
| `searchWeb()` | 调用真实搜索服务 |
| `dedupeByUrl()` | 去掉重复 URL |
| `formatResultsForModel()` | 把搜索结果整理成模型容易引用的证据块 |
第一版先不要做太复杂。DeepSearch 真正难的不是写一个搜索请求,而是把来源质量、上下文长度、引用格式和不确定性控制住。
## Step 3: 加一个 /deepsearch 命令
工具是给模型调用的,但用户也需要一个直接入口。
可以再注册一个命令,把用户输入改写成更明确的研究任务:
```typescript
export default function (pi: ExtensionAPI) {
pi.registerCommand("deepsearch", {
description: "Run a source-backed DeepSearch task",
handler: async (args, ctx) => {
const query = String(args ?? "").trim();
if (!query) {
ctx.ui.notify("Usage: /deepsearch ", "warning");
return;
}
pi.sendUserMessage(
[
"请对下面的问题做 DeepSearch。",
"",
`问题:${query}`,
"",
"要求:",
"1. 先判断是否需要调用 deep_search。",
"2. 如果问题较复杂,先拆成 2-4 个子问题分别检索。",
"3. 最终答案必须包含来源链接。",
"4. 区分事实、推断和仍不确定的部分。",
"5. 不要把搜索结果原样堆出来,要给出综合判断。"
].join("\n"),
{ deliverAs: "followUp" }
);
}
});
pi.registerTool({
// deep_search tool definition...
});
}
```
这样用户就可以直接输入:
```text
/deepsearch Pi Coding Agent 的 extension 机制适合做哪些能力?
```
`/deepsearch` 不直接搜索,而是给 Pi 发送一条更完整的任务说明。模型会根据说明调用 `deep_search`,再基于结果完成综合。
我更喜欢这种设计,因为它保留了 agent 的判断空间。搜索工具只是证据入口,不是最终答案生成器。
## Step 4: 启动和验证
项目本地 extension 放好以后,可以直接在项目根目录启动 Pi:
```bash
TAVILY_API_KEY=tvly-... pi
```
如果你只是临时测试,也可以显式指定 extension:
```bash
TAVILY_API_KEY=tvly-... pi -e ./.pi/extensions/deepsearch/index.ts
```
进入 Pi 后,先问一个需要外部事实的问题:
```text
/deepsearch Pi Coding Agent 最新版本的 extension 系统支持哪些能力?
```
一个可接受的输出不应该只是几条搜索结果,而应该包含:
| 检查点 | 合格表现 |
| ------- | ---------------------- |
| 是否调用工具 | 能看到 `deep_search` 被调用 |
| 来源是否清楚 | 每个关键事实后面有 URL |
| 是否去重 | 不重复引用同一个页面 |
| 是否有判断 | 不只罗列资料,还能归纳适用场景 |
| 是否有不确定性 | 对版本变化、第三方 API、社区扩展保持边界 |
如果结果只是"搜索结果列表",说明 promptGuidelines 不够强。可以把 guideline 改得更明确:
```typescript
promptGuidelines: [
"Use deep_search to gather evidence, not to produce the final answer.",
"After deep_search, write a concise research brief with citations.",
"Prefer official documentation, source code, release notes, and primary sources.",
"Mention when sources disagree or when the evidence is incomplete."
]
```
## Step 5: 让 DeepSearch 更像研究工具
跑通第一版以后,可以继续加三类能力。
### 子问题拆解
DeepSearch 最容易失败的地方,是把一个大问题直接丢给搜索 API。
比如:
```text
Pi Agent 能不能替代 Claude Code?
```
这不是一个好 search query。它至少可以拆成:
| 子问题 | 作用 |
| -------------------- | ----- |
| Pi Agent 的核心设计是什么 | 找定位 |
| Pi Agent 支持哪些工具和扩展 | 找能力边界 |
| Claude Code 的默认能力有哪些 | 找对比对象 |
| 两者在权限、安全、可扩展性上有什么区别 | 形成判断 |
第一版可以让模型自己拆;第二版可以让 `/deepsearch` 命令强制要求模型先列子问题,再逐个调用 `deep_search`。
### 来源质量分层
DeepSearch 的输出不能只按搜索 API 的分数排序。实际写技术文章时,我会优先看:
| 优先级 | 来源 |
| --- | --------------------- |
| P0 | 官方文档、源码、release note |
| P1 | 作者博客、维护者说明、issue / PR |
| P2 | 高质量教程、技术分析 |
| P3 | 社区讨论、Reddit、X、论坛 |
Extension 可以在 `formatResultsForModel()` 里先标注来源类型:
```typescript
function classifySource(url: string): "official" | "source" | "community" | "other" {
const host = new URL(url).hostname;
if (host === "pi.dev") return "official";
if (host === "github.com") return "source";
if (host.includes("reddit.com")) return "community";
return "other";
}
```
这样模型综合时就不会把社区传言和官方文档放在同一个证据等级上。
### 上下文截断
搜索结果很容易污染上下文。DeepSearch 的工具输出应该少而精。
我的建议是:
| 内容 | 是否放进工具输出 |
| -------------- | ------------------- |
| 标题 | 放 |
| URL | 放 |
| 200-500 字摘要 | 放 |
| 页面全文 | 默认不放 |
| 原始 HTML | 不放 |
| 搜索 API 原始 JSON | 放进 `details`,不要放进正文 |
如果确实需要全文阅读,可以再做第二个工具:
```text
fetch_source(url)
```
这样 DeepSearch 第一步负责找候选来源,第二步只抓最重要的 2-3 个页面。不要一上来把十几个网页全文都塞给模型。
## 常见问题
### 为什么不直接用 bash 跑搜索脚本?
可以,但不如 extension 稳。
用 bash 的问题是:模型每次都要重新决定命令、参数、输出格式和错误处理。Extension 把这些细节固定下来,模型只需要调用 `deep_search`。
### 为什么不把总结也写在 extension 里?
第一版不建议。
如果 extension 自己再调用一个模型做总结,你就会遇到嵌套模型调用、成本统计、上下文漂移和引用责任的问题。更简单的方式是:extension 只返回证据,Pi 当前会话里的模型负责综合。
### 这个 DeepSearch 算不算 MCP?
不算。它是 Pi extension 注册出来的本地工具。
如果你已经有成熟的 MCP 搜索服务器,也可以通过 Pi 的 MCP 相关 package 或 extension 接进来。但这个案例选择直接写 extension,是为了看清 Pi 本身的扩展机制。
### 安全上要注意什么?
至少注意四件事:
| 风险 | 做法 |
| ---------- | ------------------- |
| API Key 泄露 | 只从环境变量读取,不写进仓库 |
| 不可信网页内容 | 不把网页内容当系统指令,只当待核验证据 |
| 搜索结果污染 | 优先官方和源码,降低社区结果权重 |
| 上下文爆炸 | 限制结果数量和摘要长度 |
DeepSearch 看起来是"搜索增强",本质上是让外部网页进入 agent 上下文。只要外部内容进入上下文,就要把 prompt injection 当成真实风险。
## 小结
我会把 Pi Agent 的第一个实战案例定为 **DeepSearch Extension**,因为它能同时体现 Pi 的三个关键特点:
* Pi 的核心默认很小,不内置所有工作流。
* 真正有用的能力可以通过 extension 补上。
* Extension 不只是加命令,更是在定义模型进入外部世界的边界。
这个案例跑通以后,Pi 就不只是一个本地代码编辑 agent,而是有了一个可控的研究入口:遇到需要外部资料的问题,它可以先检索、再筛选、再带来源地回答。
这比让模型凭记忆回答更可靠,也比每次手写搜索命令更可复用。
## 参考文档
# Ralph Wiggum: Análisis en Profundidad
## Introducción
Asignar tareas a la IA antes de salir del trabajo y despertar a la mañana siguiente con código funcional — este sueño parece requerir clústeres complejos de agentes y sofisticados sistemas de orquestación. Sin embargo, la técnica de programación con IA más popular de 2025 se reduce a esta única línea:
```bash
while :; do cat PROMPT.md | claude ; done
```
Un bucle infinito que alimenta tareas a Claude repetidamente. Esto es **Ralph Wiggum**. Es tan simple que resulta casi vergonzoso, pero alguien realmente lo utilizó para completar un proyecto originalmente cotizado en $50,000 por apenas $297 en costos de API.
¿Por qué funciona un enfoque tan simple? Y cuando Anthropic lanzó un plugin oficial, ¿por qué el inventor Geoffrey Huntley dijo "This isn't it" (esto no es lo correcto)?
## ¿Qué es Ralph?
El nombre proviene de un personaje de Los Simpson. Ralph Wiggum es el hijo del jefe de policía — la persona más "inocente" de toda la serie. No entiende muy bien lo que está haciendo, pero nunca se detiene. Su frase icónica "I'm helping!" (¡Estoy ayudando!) captura inesperadamente la esencia de esta técnica: **Persistencia ingenua e incansable** (Naive and relentless persistence).
Hay una distinción importante aquí: **Ralph es una metodología, no una herramienta**. Así como "Agile" es una metodología y no un software específico, Ralph describe una forma de trabajar. Las diferentes implementaciones pueden variar enormemente en su efectividad — discutiremos esto en detalle más adelante.
## Por qué se necesita Ralph: El problema del Context Rot
Para entender por qué Ralph funciona, primero hay que entender el problema que resuelve.
### Cómo la IA "se vuelve más torpe"
Al usar Claude para tareas complejas, es posible que haya experimentado esto: la conversación comienza de forma fluida — Claude comprende con precisión y ejecuta bien. Pero a medida que la conversación se alarga, se vuelve "lento" — olvida información importante, repite los mismos errores, la calidad del código disminuye e incluso comienza a producir "alucinaciones" inexplicables.
Esto no se debe a que la IA no sea suficientemente inteligente. El problema es que **la ventana de contexto se ha contaminado**.
Imagine este escenario: le pide a Claude que escriba una función y falla en el primer intento. Usted dice "corrige esto", lo intenta pero falla de nuevo. Después de diez intercambios, el contexto de Claude está repleto de: nueve intentos de código fallidos, nueve conjuntos de mensajes de error y una montaña de discusión que ya no es relevante. Encontrar la información clave entre todo este ruido se vuelve cada vez más difícil.
### La Zona de Torpeza (Dumb Zone)
Geoffrey Huntley y la comunidad de desarrolladores descubrieron un fenómeno que denominaron "Dumb Zone" (Zona de Torpeza):
| Tamaño del contexto | Rendimiento |
| ------------------- | ----------------------------------------------------- |
| 0 - 50k tokens | Rendimiento óptimo |
| 50k - 100k tokens | Bueno, ligera degradación |
| 100k+ tokens | Degradación notable, comienza a ignorar instrucciones |
| 150k+ tokens | Degradación severa |
No existe un umbral preciso, pero la regla general es: **empiece a preocuparse cuando el contexto alcance aproximadamente la mitad de su capacidad**. Para la ventana de 200k tokens de Claude, más allá de 100k tokens es posible que esté trabajando con una IA que "se ha vuelto más torpe".
### El contexto acumulado es un pasivo
He aquí una perspectiva contraintuitiva: el contexto acumulado no es un activo — es un pasivo.
Estamos acostumbrados a pensar que una mejor memoria siempre es mejor y que retener más información siempre es mejor. Pero en el mundo de los modelos de lenguaje grandes, esta intuición es incorrecta. Cuanto más larga es la conversación, más "información negativa" llena el contexto: código fallido, discusiones irrelevantes, malentendidos ya corregidos. Estos elementos no solo ocupan espacio, sino que también dispersan la "atención" de la IA.
## Cómo funciona Ralph
Una vez que se comprende el Context Rot, la solución de Ralph queda clara: **si el contexto acumulado es el problema, entonces no lo acumule**.
Ralph se sustenta en tres pilares:
### 1. Sesión nueva
En cada iteración del bucle, se lanza una **instancia completamente nueva de Claude** con una ventana de contexto totalmente limpia. No se trata simplemente de "borrar el historial de conversación" — de esa forma el estado acumulado podría persistir. Se trata de terminar completamente el proceso actual e iniciar uno nuevo.
Esto significa que Claude se encuentra en su rendimiento óptimo al inicio de cada iteración. Sin errores anteriores que lo perturben, sin discusiones obsoletas que lo distraigan.
**Por eso el bucle debe ejecutarse fuera de Claude Code** — el bucle bash necesita poder controlar el ciclo de vida del proceso de Claude.
### 2. Archivos como fuente de verdad
Si cada iteración comienza con un contexto limpio, ¿cómo sabe la IA qué se hizo antes? La respuesta: a través del sistema de archivos, no del historial de conversación.
Archivos clave:
* **Archivo PRD/spec** — Define objetivos, listas de funcionalidades, criterios de éxito
* **IMPLEMENTATION\_PLAN.md** — Desglose de tareas y seguimiento del progreso
* **progress.txt** — Registro de formato libre; cada iteración agrega lo que aprendió
* **Historial de Git** — Evidencia de los cambios en el código
Al inicio de cada iteración, Claude lee estos archivos para comprender los objetivos y el progreso. Lo que ve es una instantánea del estado cuidadosamente organizada, no un historial de conversación caótico.
### 3. Bucle de retroalimentación
El contexto limpio y el estado persistente por sí solos no son suficientes. Si la IA escribe código con errores y lo confirma, los errores se acumularán.
El bucle de retroalimentación actúa como una puerta de calidad automatizada:
* **Verificación de tipos en TypeScript** — Retroalimentación inmediata sobre la corrección de tipos
* **Pruebas unitarias** — Verifican que las funcionalidades cumplan con lo esperado
* **CI/CD** — Asegura que el código se compile e integre correctamente
Si las pruebas fallan, el código no se confirma y Claude ve los mensajes de fallo. La siguiente iteración, con una instancia nueva de Claude, intentará corregir el problema.
> Para más información sobre cómo construir un sistema integral de aseguramiento de calidad, compartí mi enfoque de cinco capas de defensa en [Mi flujo de control de calidad con Claude Code](/es/blog/claude-code-quality-control): automatización con Hooks, estrategia de pruebas, AI Review, Pre-commit e integración con GitHub.
## Human on the Loop
Geoffrey Huntley enfatiza repetidamente una distinción conceptual:
| Human **in** the Loop | Human **on** the Loop |
| -------------------------------------------------- | ------------------------------------------------------------------ |
| Acompañamiento tipo niñera | Gestión supervisora |
| La IA espera su confirmación en cada paso | Usted establece objetivos y límites, la IA opera de forma autónoma |
| Usted es el cuello de botella del flujo de trabajo | Usted revisa el progreso de vez en cuando |
En la práctica, existen dos modos:
* **Modo AFK**: Inícielo antes de salir del trabajo, vaya a casa a dormir y revise los resultados por la mañana
* **Modo Human-in-the-loop**: Pausa para revisar después de cada iteración, adecuado para tareas complejas o inciertas
## ¿Qué tareas son adecuadas para Ralph?
Ralph no es una solución universal. Su fortaleza principal es "iterar hasta tener éxito", lo que lo hace apropiado para tipos específicos de tareas.
### Tareas adecuadas
| Escenario | Razón |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| Tareas con criterios de éxito claros | La finalización se puede verificar automáticamente (pruebas aprobadas, verificación de tipos aprobada) |
| Tareas que requieren mejora iterativa | La fortaleza principal de Ralph es el reintento persistente |
| Proyectos nuevos (greenfield) | Sin riesgo de dañar código existente |
| Proyectos con pruebas automatizadas | Las pruebas sirven como mecanismo de contrapresión para asegurar la calidad |
### Tareas no adecuadas
| Escenario | Razón |
| ------------------------------------------------ | ------------------------------------------------------------------ |
| Decisiones de diseño que requieren juicio humano | "¿Se ve bien?" no se puede verificar automáticamente |
| Operaciones únicas | Las tareas que no necesitan iteración son un desperdicio con Ralph |
| Depuración en producción | Demasiado riesgoso para operación no supervisada |
| Tareas con criterios de éxito poco claros | No hay forma de determinar cuándo detenerse |
### Tres patrones de uso
**Modo de implementación completa**
Este es el uso más común de Ralph: construir una funcionalidad o proyecto completo desde cero. Se prepara un archivo de especificaciones y un plan de implementación, y se deja que Ralph ejecute todas las tareas automáticamente.
Escenarios típicos:
* Construir una nueva API REST
* Desarrollar una herramienta CLI
* Implementar un nuevo módulo de funcionalidades
Ejemplos reales: Un desarrollador utilizó este patrón para completar un proyecto originalmente cotizado en $50,000, con un costo total de API de apenas $297. Todo el proceso — desarrollo del MVP, escritura de pruebas y revisión de código — fue completamente automatizado. Otro caso involucró la actualización de un código heredado de React v16 a v19; Ralph se ejecutó durante 14 horas sin necesidad de intervención humana alguna.
**Modo de exploración**
No todas las tareas requieren producir código. A veces lo que se necesita es comprensión — comprender una base de código recién heredada, la arquitectura de un sistema complejo o cómo funciona un módulo particular.
Escenarios típicos:
* Hacerse cargo de un proyecto desconocido y necesitar construir rápidamente un modelo mental
* Generar documentación para una base de código existente
* Analizar la arquitectura del sistema para identificar problemas potenciales
En este modo, su prompt no es "implementar la funcionalidad X", sino "leer esta base de código y generar documentación de arquitectura" o "encontrar todos los endpoints de la API y describir su propósito". Claude profundiza más con cada iteración, construyendo gradualmente una comprensión más completa.
**Modo de prueba por fuerza bruta**
Algunos errores — se conocen los síntomas, se conoce el comportamiento correcto esperado, pero simplemente no se puede encontrar la causa raíz. En estos casos, se puede dejar que Ralph lo resuelva "por fuerza bruta".
Escenarios típicos:
* Un error intermitente difícil de reproducir
* Una prueba que falla ocasionalmente por razones desconocidas
* Un problema de rendimiento donde el cuello de botella no está claro
Establezca el objetivo: "Corregir este error y hacer que esta prueba pase de forma consistente". Ralph seguirá probando diferentes enfoques de corrección hasta encontrar uno que funcione. Este método es especialmente adecuado para problemas del tipo "no sé cómo corregirlo, pero sé cuándo estará corregido".
## Elección de la implementación
Una vez comprendida la metodología de Ralph, surge una pregunta práctica: ¿cómo implementar este bucle?
La comunidad ha desarrollado dos implementaciones con diferentes niveles de ingeniería:
**Ruta minimalista** — [snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank): unos cientos de líneas de script bash, sesión completamente nueva cada vez, enfocado en el bucle en sí. Ligero, fácil de comenzar, ideal para arrancar rápidamente.
**Ruta de ingeniería** — [frankbria/ralph-claude-code](/es/docs/notes/ralph-wiggum/frankbria): cadena de herramientas completa (panel de monitoreo, disyuntor, limitación de tasa, gestión de expiración de sesiones). Por defecto reutiliza sesiones mediante `--continue`, pero también permite cambiar a sesiones nuevas con `--no-continue`.
| Dimensión | Minimalista (snarktank) | Ingeniería (frankbria) |
| -------------------------- | ----------------------- | ------------------------------------------------- |
| Modo de sesión | Nueva cada vez | Reutilización por defecto, configurable a nueva |
| Monitoreo | Revisión manual | Panel tmux integrado |
| Mecanismos de seguridad | max\_iterations | Disyuntor + limitación de tasa + tiempo de espera |
| Complejidad de instalación | Copia de Skill | install.sh + asistente |
Ambas implementaciones tienen sus ventajas y desventajas; la elección depende de sus necesidades de herramientas de ingeniería. Los métodos de uso detallados y el análisis comparativo se encuentran en los artículos prácticos de cada una.
## Reflexión final
Ralph nos enseña una lección importante: a veces el enfoque más simple es el más efectivo. Mientras todos perseguían arquitecturas más complejas, un bucle bash cambió las reglas del juego.
Por supuesto, Ralph es solo una pieza del rompecabezas. Necesita buenos prompts, el proyecto adecuado y mecanismos de retroalimentación correctos para desplegar todo su potencial. Ahora que comprende los principios, puede elegir la implementación más adecuada según sus necesidades:
* ¿Necesita AFK prolongado y muchas iteraciones? → [Guía práctica de snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank)
* ¿Necesita monitoreo de ingeniería y mecanismos de seguridad? → [Guía práctica de frankbria/ralph-claude-code](/es/docs/notes/ralph-wiggum/frankbria)
***
**Lecturas relacionadas**:
* [Guía práctica de snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank) — Bucle externo minimalista, manual operativo completo desde la instalación hasta el uso práctico
* [Guía práctica de frankbria/ralph-claude-code](/es/docs/notes/ralph-wiggum/frankbria) — Implementación de ingeniería: monitoreo, disyuntor y mecanismos de seguridad
* [Guía completa de Claude Subagent](/es/docs/notes/claude-subagent) — Otro enfoque para mantener el contexto limpio
* [¿Qué son los Claude Skills?](/es/docs/notes/claude-skills/concept) — Explorando los manuales reutilizables de Claude
* [GSD: Análisis en profundidad](/es/docs/notes/gsd/concept) — Un sistema completo de ingeniería de contexto construido sobre las bases de Ralph
* [Arquitectura del sistema Claude: Visión completa](/es/docs/notes/claude-architecture) — Comprender la arquitectura general de Hooks, Subagent y otros componentes
# Guía práctica de frankbria/ralph-claude-code
## Introducción
[El artículo anterior](/es/docs/notes/ralph-wiggum/concept) presentó la metodología Ralph, y [snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank) mostró una implementación minimalista del bucle externo. Ahora veamos otra ruta: [frankbria/ralph-claude-code](https://github.com/frankbria/ralph-claude-code).
Si la filosofía de snarktank/ralph es "hacer más con menos código", la filosofía de frankbria es "**ingeniería para todo**" — asistente de configuración interactivo, panel de monitoreo en tiempo real, circuit breaker, limitación de tasa, gestión de expiración de sesiones. No busca la simplicidad, sino la **controlabilidad**.
Las dos implementaciones no son mejores ni peores entre sí, simplemente se adaptan a diferentes escenarios de uso. Este artículo le guiará a través de la cadena de herramientas completa de frankbria.
## Instalación y configuración
### Instalación global
```bash
# Clonar el repositorio
git clone https://github.com/frankbria/ralph-claude-code.git
cd ralph-claude-code
# Instalación global
./install.sh
```
Una vez completada la instalación, obtendrá los siguientes comandos globales:
| Comando | Descripción |
| --------------- | --------------------------------------------- |
| `ralph` | Iniciar el ciclo Ralph |
| `ralph-enable` | Habilitar Ralph en un proyecto existente |
| `ralph-setup` | Crear un nuevo proyecto y configurar Ralph |
| `ralph-import` | Importar documentos PRD/requisitos existentes |
| `ralph-monitor` | Iniciar el panel de monitoreo en tiempo real |
### Inicialización del proyecto
Para proyectos existentes, utilice el asistente interactivo:
```bash
cd your-project
ralph-enable
```
El asistente detectará automáticamente el tipo de proyecto (Node.js, Python, Go, etc.) y el framework (Next.js, FastAPI, etc.), y luego generará los archivos de configuración correspondientes.
Para proyectos completamente nuevos:
```bash
ralph-setup my-new-project
```
Esto creará el directorio del proyecto, inicializará Git y generará el directorio de configuración `.ralph/`.
### Importar requisitos existentes
Si ya tiene documentos PRD o especificaciones de requisitos:
```bash
ralph-import path/to/your-prd.md
```
Ralph analizará el documento, extraerá la lista de tareas y generará un `fix_plan.md` estructurado.
## Estructura del directorio .ralph/
La memoria y configuración de frankbria se centralizan en el directorio `.ralph/`:
```
.ralph/
├── PROMPT.md # Objetivos y contexto del proyecto
├── fix_plan.md # Lista de tareas (similar a prd.json)
├── AGENT.md # Comandos de build/test (mantenido automáticamente)
├── specs/ # Documentos de requisitos detallados
│ ├── feature-a.md
│ └── feature-b.md
└── sessions/ # Datos de persistencia de sesiones
├── current.json
└── history/
```
**Comparación con snarktank/ralph**:
| frankbria | snarktank | Función |
| ------------- | ------------------------------------------- | --------------------------------------------- |
| `PROMPT.md` | `projectName` + `description` de `prd.json` | Definir objetivos del proyecto |
| `fix_plan.md` | `userStories` de `prd.json` | Lista de tareas y progreso |
| `AGENT.md` | `CLAUDE.md` / `AGENTS.md` | Comandos de build y convenciones del proyecto |
| `specs/` | campo `notes` de `prd.json` | Requisitos detallados |
| `sessions/` | Ninguno (nuevo proceso cada vez) | Seguimiento del estado de sesión |
Note que `AGENT.md` se **mantiene automáticamente** — Ralph lo actualiza automáticamente durante la ejecución según las convenciones del proyecto que descubre, similar al `progress.txt` de snarktank/ralph, pero más estructurado.
## Comandos principales
### Ejecución básica
```bash
# Iniciar el ciclo Ralph
ralph
# Con monitoreo en tiempo real
ralph --monitor
# Iniciar en tmux (recomendado para ejecuciones prolongadas)
ralph --live
```
### Panel de monitoreo
```bash
# Iniciar el monitoreo de forma independiente
ralph-monitor
```
`ralph-monitor` abrirá un panel tmux que muestra en tiempo real:
* La tarea que se está ejecutando actualmente
* Conteo de tareas completadas/pendientes
* Número de llamadas a la API y estimación de costos
* Estado del circuit breaker
* Registros de errores recientes
### Parámetros comunes
| Parámetro | Descripción | Valor por defecto |
| ----------------- | -------------------------------------- | ----------------- |
| `--resume` | Continuar desde la última interrupción | - |
| `--calls ` | Número máximo de llamadas a la API | 100 |
| `--timeout ` | Tiempo de espera (minutos) | 300 |
| `--monitor` | Habilitar monitoreo en tiempo real | false |
| `--live` | Ejecutar en tmux | false |
```bash
# Limitar a 50 llamadas API, 2 horas de tiempo de espera
ralph --calls 50 --timeout 120
# Continuar desde la última interrupción
ralph --resume
```
## Mecanismos de seguridad
La mayor característica diferenciadora de frankbria son sus mecanismos de seguridad multicapa.
### Circuit Breaker (Disyuntor)
El circuit breaker detiene automáticamente el ciclo cuando detecta "falta de progreso", previniendo el consumo inútil de la API:
**Detección de falta de progreso consecutiva**: Si durante N iteraciones consecutivas no se completa ninguna tarea nueva, se activa el circuit breaker.
**Detección de errores repetidos**: Si aparece consecutivamente el mismo mensaje de error, indica que la IA está atrapada en un ciclo infinito, y se activa el circuit breaker.
### Limitación de tasa
Límite por defecto de 100 calls/hour, para prevenir facturas de API inesperadamente altas. Se puede ajustar mediante parámetros:
```bash
ralph --calls 200 # Aumentar a 200 calls
```
### Detección de tres capas para el límite de 5 horas de API
La API de Anthropic tiene un límite de uso con ventana deslizante de 5 horas. frankbria incorpora detección en tres capas:
1. **Pre-detección**: Estima la cuota restante antes de cada llamada a la API
2. **Detección de respuesta**: Analiza los headers de rate limit en la respuesta de la API
3. **Estrategia de retroceso**: Reduce automáticamente la frecuencia de llamadas al acercarse al límite
### Gestión de expiración de sesiones
La sesión tiene una validez por defecto de 24 horas. Después de ese período, los datos de sesión se limpian automáticamente para evitar que el contexto obsoleto afecte ejecuciones posteriores.
## Detección inteligente de salida
frankbria no simplemente sale cuando todas las tareas están completas. Utiliza una **puerta de salida de doble condición**:
```
Condición de salida = completion_indicators >= 2 AND EXIT_SIGNAL: true
```
**completion\_indicators** es el número de señales de finalización detectadas en la salida de la IA, incluyendo:
* "Todas las tareas están completas"
* "No hay más elementos pendientes"
* Todas las pruebas pasaron
* Todos los elementos en fix\_plan.md marcados como done
**EXIT\_SIGNAL** es la declaración explícita de intención de salida de la IA en su salida.
¿Por qué se necesitan dos condiciones? Para prevenir **salidas prematuras**. Una señal única podría ser un falso positivo — por ejemplo, la IA dice "tarea completada" pero en realidad solo completó la story actual. La doble condición asegura que solo se sale realmente cuando múltiples señales independientes confirman la finalización.
## Comparación con snarktank/ralph
| Dimensión | snarktank/ralph | frankbria/ralph-claude-code |
| --------------------------- | ------------------------------------------ | --------------------------------------------------------------- |
| **Implementación** | Bucle bash externo (sesión nueva cada vez) | Bucle bash externo (`--continue` reutiliza sesión) |
| **Modo de sesión** | Siempre nueva | Reutilización por defecto (cambiar a nueva con `--no-continue`) |
| **Contexto** | Siempre nuevo | Acumulado entre iteraciones con `--continue` |
| **Instalación** | Copiar Skill | install.sh + asistente interactivo |
| **Formato de tareas** | prd.json | PROMPT.md + fix\_plan.md |
| **Monitoreo** | Manual `cat`/`jq` | Panel tmux integrado |
| **Mecanismos de seguridad** | max\_iterations | Circuit breaker + limitación de tasa + timeout |
| **Origen de tareas** | Solo PRD | beads / GitHub Issues / PRD |
| **Escenario ideal** | AFK prolongado, muchas iteraciones | Iteraciones cortas a medianas, necesidad de monitoreo |
### Diferencia principal: nivel de ingeniería
Ambos son bucles bash externos que inician nuevos procesos de Claude. La diferencia principal no está en el modo de gestión de sesiones (frankbria puede cambiar a modo de sesión nueva con `--no-continue`), sino en el **nivel de ingeniería**:
* **snarktank**: Script minimalista, unos cientos de líneas de bash, enfocado en el ciclo en sí
* **frankbria**: Cadena de herramientas completamente ingenieril — panel de monitoreo, circuit breaker, limitación de tasa, gestión de expiración de sesiones
frankbria tiene habilitado `--continue` por defecto para reutilizar sesiones, ideal para tareas cortas. Para tareas largas, se puede cambiar al modo de sesión nueva con `--no-continue`, obteniendo la misma protección contra Context Rot que snarktank, mientras se conservan las ventajas de ingeniería de frankbria.
### Cómo deshabilitar la reutilización de sesiones
frankbria ofrece tres formas de deshabilitar `--continue`:
```bash
# Forma uno: parámetro de línea de comandos
ralph --no-continue
# Forma dos: variable de entorno
export CLAUDE_USE_CONTINUE=false
# Forma tres: configuración .ralphrc
SESSION_CONTINUITY=false
```
Una vez deshabilitado, el comportamiento de frankbria es equivalente al de snarktank (sesión nueva cada vez), pero conserva todas las herramientas de ingeniería (monitoreo, circuit breaker, limitación de tasa, etc.).
## La realidad del Context Rot y sus compensaciones
La elección del modo de reutilización de sesiones es esencialmente una compensación entre Context Rot y el costo de inicio:
**Tareas cortas (\< 50k tokens)**: Reutilizar la sesión es más ventajoso. El contexto aún no ha tenido tiempo de degradarse, y la memoria de las primeras iteraciones aún puede ser aprovechada por las posteriores. El costo de inicio de crear una sesión nueva cada vez es más bien un desperdicio.
**Tareas largas (100k+ tokens)**: Las sesiones nuevas son más confiables. Después de superar los 100k tokens, el Context Rot se intensifica notablemente, y el contexto acumulado pasa de ser un activo a una carga. Las sesiones nuevas tienen costo de inicio, pero cada vez se parte del estado óptimo.
**Recomendaciones prácticas**:
| Escenario | Recomendación | Razón |
| --------------------------------- | --------------------------------------- | ---------------------------------------------------------- |
| Menos de 5 tareas pequeñas | frankbria (modo por defecto) | Inicio rápido, contexto reutilizable |
| 10+ tareas, necesita AFK | snarktank o frankbria + `--no-continue` | Evitar Context Rot, más confiable |
| Necesita monitoreo en tiempo real | frankbria | Panel integrado |
| Cantidad de tareas incierta | frankbria + `--no-continue` | Herramientas de ingeniería + protección contra Context Rot |
Los usuarios de frankbria pueden elegir flexiblemente según la escala de las tareas: para tareas cortas usar el modo `--continue` por defecto, para tareas largas cambiar al modo `--no-continue`. En comparación con snarktank, la ventaja de frankbria es que en cualquiera de los dos modos conserva la cadena de herramientas de ingeniería completa.
## Resumen
frankbria/ralph-claude-code representa la ruta de implementación ingenieril de la metodología Ralph. Sacrifica algo de la simplicidad de snarktank a cambio de capacidades de monitoreo, seguridad y configuración más completas.
La elección de qué implementación usar depende de sus necesidades específicas — no hay una respuesta "más correcta", solo una elección "más adecuada".
***
**Lectura complementaria**:
* [Análisis profundo de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) — Principios fundamentales y metodología
* [Guía práctica de snarktank/ralph](/es/docs/notes/ralph-wiggum/snarktank) — Implementación minimalista de bucle externo
* [Análisis profundo de GSD](/es/docs/notes/gsd/concept) — Sistema completo de ingeniería de contexto construido sobre Ralph
* [Arquitectura completa del sistema Claude](/es/docs/notes/claude-architecture) — Comprender la arquitectura general de componentes como Hooks, Subagent, etc.
# Guía práctica de Ralph
## Introducción
En [artículo anterior](/es/docs/notes/ralph-wiggum/concept), entendimos los principios básicos de Ralph: bucle infinito + contexto nuevo cada vez + archivo como única fuente de verdad. Estos tres pilares parecen simples, pero desde comprender el concepto hasta ejecutarlo, hay muchos detalles que deben resolverse.
En este artículo, comencemos. Aprenderá a utilizar [snarktank/ralph](https://github.com/snarktank/ralph) para completar el proceso completo desde la instalación hasta la ejecución. snarktank/ralph es una de las implementaciones de Ralph más maduras en la comunidad (más de 10.000 estrellas). Admite dos herramientas, Claude Code y Amp, y una cadena de herramientas completa para la generación de PRD, conversión JSON y ejecución automatizada.
## Condiciones previas
Antes de comenzar, asegúrese de que su entorno cumpla los siguientes requisitos:
| Dependencias | Descripción |
| -------------------------------------- | -------------------------------------------------------------------- |
| **Herramientas de programación de IA** | Código Claude (`npm install -g @anthropic-ai/claude-code`) o Amp CLI |
| **jq** | Herramienta de procesamiento JSON (macOS: `brew install jq`) |
| **Git** | El proyecto debe ser un repositorio Git |
```bash
# 检查依赖
claude --version # Claude Code CLI
jq --version # JSON 处理
git --version # Git
```
## Instalación y configuración
snarktank/ralph proporciona una variedad de métodos de instalación, que se pueden seleccionar según el escenario de uso.
### Método 1: instalar directamente en Claude Code (recomendado)
La forma más sencilla: pegue el enlace de GitHub en la conversación de Claude Code y deje que Claude complete automáticamente la instalación:
```
Install this skill for me: https://github.com/snarktank/ralph
```
Claude Code clonará automáticamente el repositorio y copiará los archivos de habilidades en la ubicación correcta. Los comandos `/prd` y `/ralph` están disponibles después de la instalación.
### Método 2: Instalación en el mercado del Código Claude
Instalar mediante comando de mercado:
```bash
# 添加并安装插件
/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace
```
Después de la instalación, puede utilizar las dos habilidades `/prd` (generar PRD) y `/ralph` (convertir a JSON).
### Método 3: Instalación manual de habilidades (Claude Code/Amp)
Copie manualmente el archivo de habilidades al directorio de configuración global de la herramienta correspondiente:
```bash
# 先克隆仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# Claude Code 用户
cp -r /tmp/ralph/skills/prd ~/.claude/skills/
cp -r /tmp/ralph/skills/ralph ~/.claude/skills/
# Amp 用户
cp -r /tmp/ralph/skills/prd ~/.config/amp/skills/
cp -r /tmp/ralph/skills/ralph ~/.config/amp/skills/
```
Los comandos `/prd` y `/ralph` están disponibles después de la instalación.
### Método 4: instalación a nivel de proyecto
Copie el script de Ralph directamente en el proyecto, ideal para escenarios en los que se requiere compartir en equipo o scripts personalizados:
```bash
# 克隆 Ralph 仓库
git clone https://github.com/snarktank/ralph.git /tmp/ralph
# 复制核心文件到项目
mkdir -p scripts/ralph
cp /tmp/ralph/ralph.sh scripts/ralph/
cp /tmp/ralph/CLAUDE.md scripts/ralph/ # Claude Code 用户
# 或
cp /tmp/ralph/prompt.md scripts/ralph/ # Amp 用户
# 赋予执行权限
chmod +x scripts/ralph/ralph.sh
```
Una vez completada la instalación, la estructura del proyecto es la siguiente:
```
your-project/
├── scripts/ralph/
│ ├── ralph.sh # 核心循环脚本
│ └── CLAUDE.md # Claude Code 的 Prompt 模板
├── tasks/ # PRD 文件目录(执行时自动创建)
│ └── prd.json # 你的任务定义
└── ...
```
> **Sugerencia**: El primer método es el más fácil: simplemente envíe el enlace de GitHub a Claude Code. Si desea controlar manualmente el proceso de instalación, elija el método 2 (comando de mercado) o el método 3 (copia manual). Elija el método 4 cuando necesite compartir el script con el equipo o personalizarlo.
***
## Estructura de archivos principal
La memoria de Ralph depende completamente del sistema de archivos. Comprender la función de cada archivo es un requisito previo para utilizar bien Ralph.
### ralph.sh - motor de bucle
Este es el núcleo de Ralph: un script bash que genera constantemente nuevas instancias de IA.
```bash
# 基本用法
./scripts/ralph/ralph.sh [max_iterations] # 默认:Amp
./scripts/ralph/ralph.sh --tool claude [iterations] # 使用 Claude Code
```
En cada iteración, ralph.sh realiza los siguientes pasos:
1. Cree una rama de funciones (desde `branchName` en prd.json)
2. Seleccione la historia inacabada con mayor prioridad (`passes: false`)
3. Genere una **nueva** instancia de IA para implementar esta historia.
4. Ejecute controles de calidad (verificaciones de tipo, pruebas)
5. Verificación aprobada → git commit; verificación fallida → déjelo para la siguiente iteración
6. Actualice prd.json y marque la historia como `passes: true`
7. Adjunte las lecciones aprendidas al archivo Progress.txt.
8. Repita hasta que se completen todas las historias o se alcance el número máximo de iteraciones.
El límite de iteración predeterminado es 10. Ajústelo según la complejidad del proyecto:
```bash
# 简单项目
./scripts/ralph/ralph.sh --tool claude 10
# 复杂项目
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json - definición de tarea
Este es el "cerebro" de Ralph, donde se definen todas las tareas. Es un archivo JSON plano:
```json
{
"projectName": "Blog i18n Translation",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Translate homepage metadata",
"description": "Create content/docs/meta.en.json with English translations for all navigation items",
"acceptanceCriteria": [
"meta.en.json file exists with valid JSON format",
"All navigation titles are translated to English",
"pnpm types:check passes"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Reference the existing meta.json structure"
},
{
"id": "US-002",
"title": "Translate blog post hello-world",
"description": "Create content/blog/hello-world.en.mdx, translated from Chinese to English",
"acceptanceCriteria": [
"hello-world.en.mdx file exists",
"All QuoteCard components have defaultLang='en'",
"Internal links use /en/ prefix",
"Code blocks remain untranslated",
"pnpm types:check passes"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Preserve MDX component props format"
}
]
}
```
**Descripción del campo**:
| Campo | Descripción |
| -------------------- | --------------------------------------------------------------------- |
| `projectName` | Nombre del proyecto, utilizado para nombrar registros y sucursales |
| `branchName` | Nombre de la rama de Git: Ralph creará automáticamente |
| `id` | Identificador único de historia, formato recomendado `US-001` |
| `title` | Título corto |
| `description` | Descripción detallada: cuanto más específica, mejor |
| `acceptanceCriteria` | Lista de criterios de aceptación - **Este es el campo más crítico** |
| `priority` | Número de prioridad: cuanto menor sea el número, primero se ejecutará |
| `passes` | Completo o no: Ralph se actualizará automáticamente |
| `dependsOn` | Lista de identificación de historias dependientes |
| `notes` | Consejos y contexto adicionales |
### progreso.txt - registro de experiencia
Esta es la "memoria a largo plazo" de Ralph. Después de cada iteración, la IA agregará experiencia aprendida adicional:
```text
=== Iteration 1 (US-001) ===
- Discovered: typecheck command is `pnpm types:check`, not `pnpm typecheck`
- Discovered: meta.en.json needs to mirror exact structure of meta.json
- Pattern: fumadocs i18n uses `.en.` suffix convention
=== Iteration 2 (US-002) ===
- Discovered: QuoteCard requires both `quote` and `quoteZh` props
- Gotcha: internal links must use /en/ prefix for English pages
- Pattern: code blocks should never be translated
```
Una nueva instancia de Claude para la próxima iteración leerá este archivo e inmediatamente obtendrá toda la experiencia anterior. Esta es la razón por la que las ejecuciones de Ralph son cada vez mejores: **el conocimiento se acumula entre iteraciones, pero el contexto permanece limpio**.
### AGENTS.md - Base de conocimientos persistente
Además de Progress.txt, Ralph también actualiza el archivo `AGENTS.md` (o `CLAUDE.md`) en el proyecto. Tanto Claude Code como Amp leerán automáticamente estos archivos cuando se inicien.
A diferencia de Progress.txt, AGENTS.md registra conocimiento estable entre proyectos:
```markdown
# AGENTS.md
## Codebase Conventions
- Use fumadocs for documentation framework
- MDX files use custom components: QuoteCard, BlogImage, GlossaryCard
- i18n files use `.en.mdx` suffix
## Gotchas
- Always run `pnpm types:check` after modifying MDX files
- QuoteCard: set `defaultLang='en'` in English translations
```
***
## Escribe PRD
La calidad del PRD (Documento de requisitos del producto) determina directamente los resultados de ejecución de Ralph. Bien escrito, todo viento en popa, Ralph. Mal escrito, Ralph fracasará repetidamente en la misma historia.
### Usa habilidad para generar PRD
Si tienes instalada la habilidad snarktank/ralph, puedes generar PRD de forma interactiva:
```bash
# 在 Claude Code 或 Amp 中
/prd I want to add i18n support to the blog, translating all Chinese content to English
```
La IA le hará algunas preguntas aclaratorias (qué documentos están involucrados, limitaciones de la pila de tecnología, estándares de calidad, etc.) y luego generará un documento PRD estructurado.
Después de la generación, use el comando `/ralph` para convertir el PRD al formato `prd.json`:
```bash
/ralph # 转换 PRD 为 prd.json
```
### Escribir PRD manualmente
También puedes escribir prd.json directamente. Los siguientes son principios clave de diseño.
**Principio 1: la granularidad de la historia es moderada**
Cada historia debe ser lo suficientemente pequeña como para completarse en una iteración y lo suficientemente grande como para ofrecer valor de forma independiente.
```json
// ❌ 太大:一次迭代完不成
{
"id": "US-001",
"title": "Build complete user authentication system",
"description": "Implement registration, login, forgot password, OAuth, permission management..."
}
// ❌ 太小:没有独立价值
{
"id": "US-001",
"title": "Create email field on User table",
"description": "Add email field to User model"
}
// ✅ 刚好:一次迭代能完成,有独立价值
{
"id": "US-001",
"title": "Implement email/password login",
"description": "Create login API and login page with email/password authentication",
"acceptanceCriteria": [
"POST /api/auth/login accepts email + password",
"Returns JWT token",
"Login page form submits successfully",
"All tests pass"
]
}
```
**Regla general**: Una historia implica de 1 a 3 modificaciones de archivos y tiene de 3 a 5 criterios de aceptación.
**Principio 2: Los criterios de aceptación deben ser verificables automáticamente**
Ralph necesita determinar si la historia está completa, por lo que los criterios de aceptación deben poder evaluarse objetivamente:
```json
// ❌ 模糊的标准
"acceptanceCriteria": [
"Code quality is good",
"Performance is decent",
"User experience is smooth"
]
// ✅ 可验证的标准
"acceptanceCriteria": [
"pnpm types:check passes",
"pnpm test passes",
"API response time < 200ms",
"File src/auth/login.ts exists and exports loginHandler function"
]
```
**Principio 3: Utilice depende de para controlar el orden de ejecución**
Algunas historias tienen dependencias. El campo `dependsOn` garantiza que Ralph se ejecute en el orden correcto:
```json
{
"userStories": [
{
"id": "US-001",
"title": "Create database schema",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implement user registration API",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implement login page",
"dependsOn": ["US-002"]
}
]
}
```
**Principio 4: Proporcionar contexto en las notas**
El campo de notas le da a la IA pistas adicionales. Escriba aquí información que sepa y que la IA tal vez no conozca:
```json
{
"notes": "Project uses fumadocs framework, i18n files follow .en.mdx suffix naming. Reference content/docs/notes/speckit/concept.en.mdx for translation style."
}
```
***
## Ejecutar bucle Ralph
Una vez que el PRD esté listo, es hora de ejecutar el ciclo.
### Iniciar ejecución
```bash
# 使用 Claude Code,默认 10 次迭代
./scripts/ralph/ralph.sh --tool claude
# 指定迭代次数
./scripts/ralph/ralph.sh --tool claude 30
# 使用 Amp(默认)
./scripts/ralph/ralph.sh 20
```
### Proceso de ejecución
Después de comenzar, verá un resultado similar a este:
```
=== Ralph Loop - Iteration 1 ===
Branch: ralph/i18n-translation
Selected story: US-001 - Translate homepage metadata
Spawning fresh Claude instance...
[Claude Code executing...]
Quality check: pnpm types:check ... PASSED
Committing: feat: [US-001] - Translate homepage metadata
Updating prd.json: US-001 passes: true
Appending to progress.txt
=== Ralph Loop - Iteration 2 ===
Selected story: US-002 - Translate blog post hello-world
Spawning fresh Claude instance...
```
Cada iteración es una instancia completamente nueva de Claude. Sabe qué hacer leyendo prd.json y qué ha aprendido previamente leyendo Progress.txt.
\###Señal completa
Cuando todas las historias están marcadas como `passes: true`, Ralph emite una señal de finalización y sale:
```
All stories completed!
COMPLETE
```
### Monitoreo y depuración
Mientras Ralph se ejecuta, puede utilizar el siguiente comando para ver el progreso:
```bash
# 查看每个 story 的完成状态
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# 查看经验日志
cat progress.txt
# 查看最近的 git 提交
git log --oneline -10
# 实时跟踪 Ralph 输出
tail -f progress.txt
```
### Archivado automático
Cuando inicia una nueva función (usando un `branchName` diferente), Ralph archivará automáticamente los archivos de la última ejecución en el directorio `archive/YYYY-MM-DD-feature-name/`, manteniendo limpio el directorio de trabajo.
***
## Bucles de retroalimentación y control de puerta de calidad
La capacidad de Ralph para "autocorregirse" depende enteramente de la calidad del circuito de retroalimentación. Sin un ciclo de retroalimentación, Ralph es solo un script que se repite a ciegas: sigue generando código sin poder decir si es correcto.
### Configurar control de calidad
Defina comandos de control de calidad en CLAUDE.md (o Prompt.md):
```markdown
## Quality Commands
After implementing each story, run these checks IN ORDER:
1. `pnpm types:check` — TypeScript type checking
2. `pnpm test` — Unit tests
3. `pnpm build` — Full build verification
If any check fails:
- DO NOT commit
- Fix the issue
- Re-run all checks
- Only commit when all checks pass
```
### Nivel de control de acceso de calidad
| Jerarquía | Herramientas | Problemas capturados |
| ----------------------------------- | ------------------------------------- | --------------------------------------------------------------------------- |
| Comentarios instantáneos | Compilador de TypeScript | Errores tipográficos, errores de sintaxis |
| Verificación funcional | Pruebas unitarias | Errores lógicos, casos extremos |
| Verificación de integración | Comandos de compilación | Problemas de dependencia, errores de configuración |
| Verificación en tiempo de ejecución | habilidad del navegador de desarrollo | Problemas de representación de la interfaz de usuario (proyectos front-end) |
> Para historias de front-end, Ralph recomienda agregar este criterio de aceptación: "Verificar en el navegador usando la habilidad del navegador de desarrollo": permita que la IA abra el navegador para confirmar que la página se muestra correctamente.
### Cuando falla el control de calidad
Si una historia falla repetidamente en el control de calidad, Ralph no volverá a intentar la misma historia indefinidamente. Se detendrá después de alcanzar el límite superior de iteración y conservará el estado actual. Puedes:
1. Vea el archivo Progress.txt para comprender el motivo del bloqueo.
2. Solucione el problema manualmente y luego ejecútelo nuevamente.
3. Ajuste la granularidad de la historia (quizás demasiado grande)
4. Agregue más contexto al campo de notas.
***
## Personalización inmediata
La plantilla de aviso de Ralph (CLAUDE.md o aviso.md) es su principal medio para controlar el comportamiento de la IA. Después de la instalación, debes personalizarlo según tu proyecto.
### Elementos clave de personalización
**1. Comandos de calidad específicos del proyecto**
```markdown
## Project-Specific Commands
- Typecheck: `pnpm types:check` (not `tsc` or `pnpm typecheck`)
- Test: `pnpm vitest run`
- Build: `pnpm build`
- Lint: `pnpm lint`
```
**2. Restricciones de estilo de código**
```markdown
## Code Conventions
- Use TypeScript strict mode
- Prefer named exports over default exports
- Use fumadocs components for MDX content
- Follow existing file naming patterns (kebab-case)
```
**3. Errores conocidos**
```markdown
## Known Gotchas
- MDX files: always import components at the top
- i18n: English files use `.en.mdx` suffix
- Links: English pages must use `/en/` prefix
- QuoteCard: set `defaultLang` to match the file language
```
**4. Manejo cuando está atascado**
```markdown
## When Stuck
If you cannot complete a story after 3 attempts within the same iteration:
1. Document what's blocking in progress.txt
2. Move to the next story if possible
3. Do NOT modify files unrelated to the current story
```
***
## Caso práctico: Traducir blogs con Ralph
Para mostrar cómo funciona Ralph en la práctica, he aquí un ejemplo de la vida real: utilizar un agente autónomo al estilo Ralph para traducir un blog completo del chino al inglés.
### Configuración del proyecto
Este proyecto requiere traducir más de 22 archivos de contenido (publicaciones de blog, documentos, metadatos de navegación) del chino al inglés. El proyecto se basa en el blog Next.js de fumadocs y es compatible con i18n. La tarea está definida en el archivo `prd.json` y contiene 16 historias de usuario, cada una con criterios de aceptación claros:
```
scripts/ralph/
├── prd.json # 16 个 user story,带验收标准
└── progress.txt # 经验日志,每个 story 完成后更新
```
Cada historia de usuario sigue un patrón consistente:
* **Entregable explícito**: "Crear contenido/blog/xxx.en.mdx"
* **Criterios verificables**: "Pases de verificación de tipo", "Los enlaces internos usan el prefijo /en/"
* **Restricciones técnicas**: "Mantener los bloques de código sin traducir", "Establecer defaultLang='en' en QuoteCard"
### Modo de ejecución
El agente sigue los principios básicos de la metodología de Ralph:
1. **Los documentos son la fuente de la verdad**: `prd.json` rastrea qué historias pasaron (`passes: true/false`). `progress.txt` Adquiera experiencia entre iteraciones, p. "El comando Typecheck es `pnpm types:check`, no `pnpm typecheck`"
2. **Puerta de calidad automatizada**: después de cada traducción, se ejecuta `pnpm types:check` para verificar que el archivo MDX se haya compilado correctamente. Si la verificación de tipo falla, corríjala antes de enviarla.
3. **Avance incremental**: cada historia se envía de forma independiente, con información descriptiva de envío (`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`) y se puede revertir fácilmente cuando sea necesario.
4. **Ejecución paralela**: para artículos más largos, se traducen varios subagentes simultáneamente; por ejemplo, US-010 (concepto de habilidades de claude + práctica), US-011 (concepto de speckit + práctica) y US-012 (arquitectura de claude + subagente de claude) se ejecutan en paralelo.
### Lecciones clave
| Experiencia | Detalles |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **El conocimiento acumulado es importante** | Los patrones descubiertos en las primeras historias (QuoteCard `defaultLang`, reglas de prefijo de enlace) permiten que las historias posteriores se completen más rápido |
| **Verificación de tipos como circuito de retroalimentación** | Detecte importaciones faltantes o MDX con formato incorrecto antes de que se acumulen los problemas |
| **Paralelizado y escalable** | 6 agentes de traducción ejecutándose simultáneamente, el tiempo de finalización es aproximadamente el mismo que el de 1 agente |
| **La granularidad del PRD es crítica** | Alcance 1 o 2 archivos por historia: lo suficientemente pequeños para completarse de manera confiable, lo suficientemente grandes para que sean significativos |
| **El registro de progreso evita errores repetidos** | La parte "Patrones de base de código" de Progress.txt se convierte en una base de conocimientos para evitar que se vuelva a descubrir el mismo problema. |
### Resultados
Las 16 historias de usuarios se completaron en una sola sesión: se crearon 8 archivos de navegación meta.en.json, se tradujeron 3 publicaciones de blog, se tradujeron 12 páginas de documentación y se verificó la construcción completa del sitio. Cada traducción mantiene una calidad constante porque los criterios de aceptación son claros y un circuito de retroalimentación (verificación de tipos) detecta los problemas de inmediato.
Este proyecto demuestra el **modelo de implementación completo** de Ralph: tareas bien definidas + criterios de éxito claros + verificación automatizada + entrega incremental a través del sistema de archivos.
***
## Implementación comunitaria y alternativas
snarktank/ralph no es la única opción. Cada una de estas implementaciones tiene sus propias fortalezas y debilidades, según sus necesidades:
| Recursos | Enlaces | Instrucciones |
| ----------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| snarktank/ralph | [tanque snark/ralph](https://github.com/snarktank/ralph) | Utilizada en este artículo, la función más completa |
| orquestador ralph | [mikeyobrien/ralph-orquestador](https://github.com/mikeyobrien/ralph-orchestrator) | Desarrollado por Mickey O'Brien, con más opciones de personalización |
| agente-ralph-loop | [vercel-labs/ralph-loop-agent](https://github.com/vercel-labs/ralph-loop-agent) | Implementación de Vercel basada en AI SDK |
| ralphy | [michaelshimeles/ralphy](https://github.com/michaelshimeles/ralphy) | Implementación ligera por Michael Shimeles |
### Alternativa: GSD
GSD no es estrictamente una "implementación comunitaria" de Ralph; es una **alternativa**. Aplica los principios básicos de Ralph (gestión del contexto, tareas atómicas) pero proporciona un flujo de trabajo más completo: discutir → planificar → ejecutar → verificar.
| Recursos | Enlaces | Instrucciones |
| --------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------- |
| GSD (Hacer las cosas) | [glittercowboy/hacer-mierda](https://github.com/glittercowboy/get-shit-done) | Marco completo desde la idea hasta el PRD y la ejecución |
Si cree que Ralph es demasiado "crudo" y necesita más apoyo en el proceso, GSD puede ser más adecuado. Consulte [Análisis en profundidad de GSD](/es/docs/notes/gsd/concept) para obtener más detalles.
***
## Recursos recomendados
**Fuente oficial**:
| Recursos | Enlaces | Instrucciones |
| ------------------------ | ----------------------------------------------------------------------------- | ------------------------------ |
| Blog de Geoffrey Huntley | [ghuntley.com/ralph](https://ghuntley.com/ralph/) | Artículo original del inventor |
| cómo-ralph-wiggum | [ghuntley/cómo-ralph-wiggum](https://github.com/ghuntley/how-to-ralph-wiggum) | Guía oficial del usuario |
**Tutorial en vídeo**:
| Recursos | Enlaces | Instrucciones |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------- |
| Discusión en profundidad de Ralph Wiggum | [Por qué la implementación de Claude Code no lo es](https://www.youtube.com/watch?v=O2bBWDoxO4s) | Geoffrey Huntley explica los problemas con la implementación oficial |
| Uso correcto de Ralph | [Estás usando los bucles de Ralph Wiggum MAL](https://www.youtube.com/watch?v=I7azCAgoUHc) | Demostración del uso del romano (Mentat) |
| Necesitamos hablar de Ralph | [Necesitamos hablar de Ralph](https://www.youtube.com/watch?v=Yr9O6KFwbW4) | El análisis de Theo de la controversia |
***
## Mejores prácticas y preguntas frecuentes
### Control de costos
La ejecución automatizada de Ralph significa que se incurre en tarifas API de forma continua. Varias medidas de control:
* **Siempre configurado `max_iterations`**: esta es la red de seguridad más básica
* **Mantenga la granularidad de la historia razonable**: una historia que es demasiado grande consumirá múltiples iteraciones; una historia demasiado detallada aumentará los gastos generales de puesta en marcha.
* **Primero pruebas a pequeña escala**: los nuevos proyectos se ejecutarán primero durante 3 a 5 iteraciones y luego se ampliarán después de confirmar que las indicaciones y el control de acceso de calidad funcionan correctamente.
### Errores comunes
**Trampa 1: La historia es demasiado grande**
Síntomas: una historia falla repetidamente y el número de iteraciones se agota rápidamente.
Solución: divídalo en 2 o 3 pisos más pequeños. "Crear un sistema de autenticación completo" se divide en "Implementar la API de inicio de sesión" + "Crear la página de inicio de sesión" + "Agregar middleware JWT".
**Trampa 2: Sin circuito de retroalimentación**
Síntoma: Ralph afirma que la historia está completa, pero hay algún problema con el código real.
Solución: agregue comandos de verificación ejecutables en los criterios de aceptación. "El código está escrito" no es un criterio de aceptación; "pnpm prueba todos los pases" sí lo es.
**Trampa 3: Progress.txt no se utiliza**
Síntomas: el mismo error aparece repetidamente en diferentes iteraciones.
Resolución: asegúrese de que su plantilla de aviso indique explícitamente "lea el archivo Progress.txt y siga las reglas que contiene". Si la IA no agrega experiencia automáticamente, agregue "Después de completar cada historia, agregue experiencia a Progress.txt" en el mensaje.
**Trampa 4: Orden de dependencia incorrecto**
Síntoma: una historia se basa en un código que aún no existe, lo que provoca que falle la implementación.
Resolución: establezca el campo `dependsOn` correctamente. Asegúrese de que la historia de la infraestructura sea lo primero.
### FAQ
\*\*P: ¿Cuál es la diferencia entre Ralph y el complemento oficial? \*\*
Diferencia principal: snarktank/ralph genera nuevos procesos en cada iteración (verdaderamente un contexto completamente nuevo), mientras que el complemento oficial se repite dentro de la misma sesión (el contexto continúa acumulándose). Ver [análisis del artículo anterior](/es/docs/notes/ralph-wiggum/concept#the-problem-with-the-official-plugin) para más detalles.
\*\*P: ¿Se puede modificar prd.json manualmente durante la ejecución? \*\*
Sí. Ralph vuelve a leer prd.json al comienzo de cada iteración. Puede modificar las descripciones de las historias entre iteraciones, agregar nuevas historias o marcar manualmente una historia como `passes: true` (omitirla).
\*\*P: ¿Qué debería hacer Ralph si se queda atrapado en una historia que falla repetidamente? \*\*
1. Consulte el archivo Progress.txt para comprender el motivo del error.
2. Añade más contexto en las notas.
3. Divida la historia (quizás demasiado grande)
4. Solucione manualmente el problema de bloqueo y luego ejecútelo nuevamente.
\*\*P: ¿Puedo hacer otras cosas mientras Ralph está corriendo? \*\*
Sí. Ralph está diseñado para ser "Human on the Loop": no es necesario que lo mires fijamente. En el modo AFK, comience antes de salir del trabajo y verifique los resultados a la mañana siguiente. Simplemente no modifique el archivo en el que está trabajando Ralph.
\*\*P: ¿Cómo controlar los costos? \*\*
Tres métodos: establecer un `max_iterations` razonable, mantener la granularidad de la historia adecuada (para reducir las iteraciones desperdiciadas) y primero realizar una prueba a pequeña escala para confirmar que el proceso es correcto. En términos generales, para proyectos con 10 a 20 pisos, la tarifa API está entre $50 y $100.
***
## Resumen
El flujo de trabajo de Ralph se puede resumir en cinco pasos:
```
安装 → 编写 PRD → 配置质量门禁 → 运行循环 → 检查结果
```
El concepto central sigue siendo el mismo: **Deje que el documento sea la única fuente de verdad, deje que cada iteración comience desde cero y deje que la puerta de calidad controle por usted**.
Ahora, regresa a tu proyecto, prepara prd.json, ejecuta `./scripts/ralph/ralph.sh --tool claude` y prepara una taza de café.
### Lectura adicional
* [Análisis en profundidad de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) - Revisando los principios fundamentales de Ralph
* [Análisis en profundidad de GSD](/es/docs/notes/gsd/concept) - Un completo sistema de ingeniería de contexto construido sobre la base de Ralph
* [¿Qué son las habilidades de Claude?](/es/docs/notes/claude-skills/concept) ——La habilidad PRD de Ralph es una habilidad de Claude
* [Guía práctica de Speckit](/es/docs/notes/speckit/practice) - Otro flujo de trabajo estructurado de programación de IA
# Guía práctica de snarktank/ralph
## Introducción
En el [artículo anterior](/es/docs/notes/ralph-wiggum/concept) conocimos el principio central de Ralph: bucle infinito + contexto completamente nuevo en cada iteración + archivos como fuente de verdad. Estos tres pilares suenan simples, pero entre comprenderlos y lograr que funcionen en la práctica hay muchos detalles.
En este artículo, pasaremos a la práctica. [snarktank/ralph](https://github.com/snarktank/ralph) es la **implementación del bucle externo** de la metodología Ralph, que inicia un proceso de Claude completamente nuevo en cada iteración, resolviendo de raíz el problema del Context Rot. Es una de las implementaciones de Ralph más completas de la comunidad (10k+ stars), compatible con las plataformas Claude Code y Amp, y ofrece una cadena de herramientas completa para generación de PRD, conversión a JSON y ejecución automática.
> Otra línea de implementación es [frankbria/ralph-claude-code](/es/docs/notes/ralph-wiggum/frankbria), que proporciona una cadena de herramientas de ingeniería completa (panel de monitoreo, circuit breaker, limitación de velocidad), enfocada en la controlabilidad y los mecanismos de seguridad. La comparación entre ambas se encuentra en ese artículo.
## Requisitos previos
Antes de comenzar, asegúrese de que su entorno cumple las siguientes condiciones:
| Dependencia | Descripción |
| -------------------------------------- | ------------------------------------------------------------------ |
| **Herramienta de programación con IA** | Claude Code (`npm install -g @anthropic-ai/claude-code`) o Amp CLI |
| **jq** | Herramienta de procesamiento JSON (macOS: `brew install jq`) |
| **Git** | El proyecto debe ser un repositorio Git |
```bash
# Verificar dependencias
claude --version # Claude Code CLI
jq --version # Procesamiento JSON
git --version # Git
```
## Instalación y configuración
La forma más sencilla: pegue el enlace de GitHub directamente en una conversación de Claude Code:
```
Ayúdame a instalar este skill: https://github.com/snarktank/ralph
```
Claude Code clonará automáticamente el repositorio y copiará los archivos del skill en la ubicación correcta. Una vez completada la instalación, puede utilizar los comandos `/prd` y `/ralph`.
> snarktank/ralph también admite instalación desde el Marketplace, copia manual de archivos del skill, instalación a nivel de proyecto, entre otros métodos. Consulte las instrucciones en el [repositorio de GitHub](https://github.com/snarktank/ralph) para más detalles.
***
## Estructura de archivos principales
La memoria de Ralph depende completamente del sistema de archivos. Comprender la función de cada archivo es un requisito previo para utilizar Ralph de manera efectiva.
### ralph.sh — Motor del bucle
Este es el núcleo de Ralph: un script bash que se encarga de iniciar repetidamente nuevas instancias de IA.
```bash
# Uso básico
./scripts/ralph/ralph.sh [max_iterations] # Usa Amp por defecto
./scripts/ralph/ralph.sh --tool claude [iterations] # Usa Claude Code
```
En cada iteración, ralph.sh realiza lo siguiente:
1. Crea una rama de funcionalidad (basada en el `branchName` de prd.json)
2. Selecciona el story de mayor prioridad que no ha sido completado (`passes: false`)
3. Inicia una instancia de IA **completamente nueva** para implementar ese story
4. Ejecuta verificaciones de calidad (verificación de tipos, pruebas)
5. Si las verificaciones pasan, hace git commit; si fallan, lo deja para la siguiente iteración
6. Actualiza prd.json y marca el story como `passes: true`
7. Agrega la experiencia aprendida en esta iteración a progress.txt
8. Repite el proceso hasta que todos los stories estén completados o se alcance el límite de iteraciones
El límite de iteraciones predeterminado es 10. Ajústelo según la complejidad del proyecto:
```bash
# Proyecto simple
./scripts/ralph/ralph.sh --tool claude 10
# Proyecto complejo
./scripts/ralph/ralph.sh --tool claude 50
```
### prd.json — Definición de tareas
Este es el "cerebro" de Ralph: todas las tareas se definen aquí. El formato es un archivo JSON plano:
```json
{
"projectName": "Blog i18n traducción",
"branchName": "ralph/i18n-translation",
"userStories": [
{
"id": "US-001",
"title": "Traducir metadatos de la página principal",
"description": "Crear content/docs/meta.en.json con la traducción al inglés de todos los elementos de navegación",
"acceptanceCriteria": [
"El archivo meta.en.json existe y tiene formato JSON válido",
"Todos los títulos de navegación han sido traducidos al inglés",
"pnpm types:check pasa correctamente"
],
"priority": 1,
"passes": false,
"dependsOn": [],
"notes": "Consultar la estructura del meta.json existente"
},
{
"id": "US-002",
"title": "Traducir el artículo del blog hello-world",
"description": "Crear content/blog/hello-world.en.mdx, traduciendo del chino al inglés",
"acceptanceCriteria": [
"El archivo hello-world.en.mdx existe",
"Todos los componentes QuoteCard tienen configurado defaultLang='en'",
"Los enlaces internos usan el prefijo /en/",
"Los bloques de código se mantienen sin traducir",
"pnpm types:check pasa correctamente"
],
"priority": 2,
"passes": false,
"dependsOn": ["US-001"],
"notes": "Prestar atención a mantener el formato de props de los componentes MDX"
}
]
}
```
**Descripción de los campos**:
| Campo | Descripción |
| -------------------- | ---------------------------------------------------------------------- |
| `projectName` | Nombre del proyecto, utilizado para logs y nomenclatura de ramas |
| `branchName` | Nombre de la rama Git; Ralph la crea automáticamente |
| `id` | Identificador único del story; se recomienda el formato `US-001` |
| `title` | Título breve |
| `description` | Descripción detallada; cuanto más específica, mejor |
| `acceptanceCriteria` | Lista de criterios de aceptación — **este es el campo más importante** |
| `priority` | Número de prioridad; cuanto menor, se ejecuta primero |
| `passes` | Indica si está completado; Ralph lo actualiza automáticamente |
| `dependsOn` | Lista de IDs de stories de los que depende |
| `notes` | Notas y sugerencias adicionales |
### progress.txt — Registro de experiencias
Esta es la "memoria a largo plazo" de Ralph. Al final de cada iteración, la IA agrega aquí lo aprendido durante esa ronda:
```text
=== Iteración 1 (US-001) ===
- Descubierto: el comando de verificación de tipos es `pnpm types:check`, no `pnpm typecheck`
- Descubierto: meta.en.json necesita reflejar exactamente la estructura de meta.json
- Patrón: fumadocs i18n usa la convención de sufijo `.en.`
=== Iteración 2 (US-002) ===
- Descubierto: QuoteCard requiere ambas propiedades `quote` y `quoteZh`
- Advertencia: los enlaces internos deben usar el prefijo /en/ para páginas en inglés
- Patrón: los bloques de código nunca deben traducirse
```
La nueva instancia de Claude en la siguiente iteración lee este archivo y obtiene inmediatamente toda la experiencia previa. Esta es la razón por la que Ralph funciona cada vez mejor con el tiempo: **el conocimiento se acumula entre iteraciones, pero el contexto se mantiene limpio**.
### AGENTS.md — Base de conocimientos persistente
Además de progress.txt, Ralph también actualiza el archivo `AGENTS.md` (o `CLAUDE.md`) del proyecto. Tanto Claude Code como Amp leen automáticamente estos archivos al iniciarse.
A diferencia de progress.txt, AGENTS.md registra **conocimiento estable y universalmente aplicable entre proyectos**:
```markdown
# AGENTS.md
## Convenciones del código base
- Usar fumadocs como framework de documentación
- Los archivos MDX usan componentes personalizados: QuoteCard, BlogImage, GlossaryCard
- Los archivos i18n usan el sufijo `.en.mdx`
## Advertencias
- Siempre ejecutar `pnpm types:check` después de modificar archivos MDX
- QuoteCard: establecer `defaultLang='en'` en las traducciones al inglés
```
***
## Redacción del PRD
La calidad del PRD (Product Requirements Document) determina directamente la efectividad de la ejecución de Ralph. Si está bien escrito, Ralph avanza sin problemas; si está mal escrito, Ralph fallará repetidamente en el mismo story.
### Uso del Skill para generar el PRD
Si ha instalado el skill de snarktank/ralph, puede generar el PRD de forma interactiva:
```bash
# En Claude Code o Amp
/prd Quiero agregar soporte i18n al sistema del blog, necesito traducir todo el contenido en chino al inglés
```
La IA le hará una serie de preguntas de clarificación (qué archivos están involucrados, restricciones del stack tecnológico, estándares de calidad, etc.) y luego generará un documento PRD estructurado.
Una vez generado, utilice el comando `/ralph` para convertir el PRD al formato `prd.json`:
```bash
/ralph # Convierte el PRD a prd.json
```
### Redacción manual del PRD
También puede escribir el prd.json directamente. A continuación se presentan los principios clave de diseño.
**Principio uno: la granularidad del Story debe ser adecuada**
Cada story debe ser lo suficientemente pequeño como para completarse en una sola iteración, pero lo suficientemente grande como para tener un valor de entrega independiente.
```json
// ❌ Demasiado grande: no se puede completar en una iteración
{
"id": "US-001",
"title": "Construir un sistema completo de autenticación de usuarios",
"description": "Implementar registro, inicio de sesión, recuperación de contraseña, OAuth, gestión de permisos..."
}
// ❌ Demasiado pequeño: no tiene valor independiente
{
"id": "US-001",
"title": "Crear el campo email de la tabla User",
"description": "Agregar el campo email al modelo User"
}
// ✅ Adecuado: se puede completar en una vez y tiene valor independiente
{
"id": "US-001",
"title": "Implementar inicio de sesión con correo electrónico y contraseña",
"description": "Crear la API de inicio de sesión y la página de login, con soporte para verificación por correo electrónico y contraseña",
"acceptanceCriteria": [
"POST /api/auth/login acepta email + password",
"Devuelve un token JWT",
"El formulario de la página de login se puede enviar",
"Todas las pruebas pasan"
]
}
```
**Regla práctica**: un story implica la modificación de 1 a 3 archivos y tiene de 3 a 5 criterios de aceptación.
**Principio dos: los criterios de aceptación deben ser verificables automáticamente**
Ralph necesita determinar si un story está completado, por lo que los criterios de aceptación deben ser objetivamente evaluables:
```json
// ❌ Criterios vagos
"acceptanceCriteria": [
"La calidad del código es buena",
"El rendimiento es aceptable",
"La experiencia de usuario es fluida"
]
// ✅ Criterios verificables
"acceptanceCriteria": [
"pnpm types:check pasa correctamente",
"pnpm test pasa correctamente",
"El tiempo de respuesta de la API es < 200ms",
"El archivo src/auth/login.ts existe y exporta la función loginHandler"
]
```
**Principio tres: utilizar dependsOn para controlar el orden**
Algunos stories tienen relaciones de dependencia entre sí. El campo `dependsOn` asegura que Ralph los ejecute en el orden correcto:
```json
{
"userStories": [
{
"id": "US-001",
"title": "Crear el schema de la base de datos",
"dependsOn": []
},
{
"id": "US-002",
"title": "Implementar la API de registro de usuarios",
"dependsOn": ["US-001"]
},
{
"id": "US-003",
"title": "Implementar la página de inicio de sesión",
"dependsOn": ["US-002"]
}
]
}
```
**Principio cuatro: proporcionar contexto en notes**
El campo notes sirve como indicaciones adicionales para la IA. Escriba aquí la información que usted conoce pero que la IA podría desconocer:
```json
{
"notes": "El proyecto utiliza el framework fumadocs. La regla de nomenclatura para archivos i18n es el sufijo .en.mdx. Consulte el estilo de traducción de content/docs/notes/speckit/concept.en.mdx."
}
```
***
## Ejecución del Ralph Loop
Con el PRD listo, es hora de ejecutar el bucle.
### Iniciar la ejecución
```bash
# Usar Claude Code, 10 iteraciones por defecto
./scripts/ralph/ralph.sh --tool claude
# Especificar número de iteraciones
./scripts/ralph/ralph.sh --tool claude 30
# Usar Amp (por defecto)
./scripts/ralph/ralph.sh 20
```
### Proceso de ejecución
Después de iniciar, verá una salida similar a esta:
```
Iniciando Ralph - Herramienta: claude - Máximo de iteraciones: 35
===============================================================
Ralph Iteración 1 de 35 (claude)
===============================================================
## US-001 Completado
**Resumen de lo realizado:**
1. Se creó meta.en.json con todos los elementos de navegación traducidos
2. Se ejecutó pnpm types:check — APROBADO
3. Commit: feat: [US-001] - Traducir metadatos de la página principal
Todavía quedan **15 user stories con `passes: false`**.
El siguiente story es **US-002: Traducir el artículo del blog hello-world**.
Iteración 1 completada. Continuando...
===============================================================
Ralph Iteración 2 de 35 (claude)
===============================================================
```
Cada iteración es una instancia de Claude completamente nueva. Sabe qué hacer leyendo prd.json y conoce las lecciones previas a través de progress.txt.
### Señal de finalización
Cuando todos los stories están marcados como `passes: true`, Ralph emite la señal de finalización y se detiene:
```
¡Todos los stories completados!
COMPLETE
```
### Monitoreo y depuración
Mientras Ralph está en ejecución, puede usar estos comandos para verificar el progreso:
```bash
# Ver el estado de finalización de cada story (con iconos, más visual)
cat tasks/prd.json | python3 -c "
import json,sys
for s in json.load(sys.stdin)['userStories']:
print(f'{\"✅\" if s[\"passes\"] else \"⬜\"} {s[\"id\"]}: {s[\"title\"]}')"
# O usar jq para verlo
cat tasks/prd.json | jq '.userStories[] | {id, title, passes}'
# Ver el registro de experiencias
cat progress.txt
# Ver los commits recientes de git
git log --oneline -10
# Ver la salida de Ralph en tiempo real
tail -f progress.txt
# Después de completar, ver todos los cambios respecto a la rama principal
git diff main...ralph/your-branch-name --stat
```
### Interrupción y reanudación
Ralph puede funcionar durante mucho tiempo, e interrumpirlo a mitad del proceso es completamente seguro:
* **Interrupción**: simplemente presione `Ctrl+C`. Los stories completados (`passes: true`) no se perderán; ya se han hecho commit y se han registrado en prd.json
* **Reanudación**: ejecute el mismo comando nuevamente y Ralph continuará automáticamente desde el primer story con `passes: false`
```bash
# Después de una interrupción, para reanudar solo ejecute el mismo comando
./scripts/ralph/ralph.sh --tool claude 35
```
Si un story falla repetidamente y causa un bloqueo, puede omitirlo manualmente: edite `prd.json`, cambie el campo `passes` de ese story a `true` y vuelva a ejecutar. Ralph lo omitirá y continuará procesando los stories siguientes.
### Archivado automático
Cuando inicia una funcionalidad diferente con un nuevo `branchName`, Ralph archiva automáticamente los archivos de la ejecución anterior en el directorio `archive/YYYY-MM-DD-nombre-funcionalidad/`, manteniendo el directorio de trabajo ordenado.
***
## Bucle de retroalimentación y controles de calidad
La capacidad de "autocorrección" de Ralph depende completamente de la calidad del bucle de retroalimentación. Sin bucle de retroalimentación, Ralph es simplemente un script que itera ciegamente: produce código continuamente, pero no puede determinar si el código es correcto.
### Configuración de verificaciones de calidad
Defina sus comandos de verificación de calidad en CLAUDE.md (o prompt.md):
```markdown
## Comandos de calidad
Después de implementar cada story, ejecute estas verificaciones EN ORDEN:
1. `pnpm types:check` — Verificación de tipos TypeScript
2. `pnpm test` — Pruebas unitarias
3. `pnpm build` — Verificación de compilación completa
Si alguna verificación falla:
- NO haga commit
- Corrija el problema
- Vuelva a ejecutar todas las verificaciones
- Solo haga commit cuando todas las verificaciones pasen
```
### Niveles de controles de calidad
| Nivel | Herramienta | Problemas detectados |
| ----------------------------------- | --------------------- | ------------------------------------------------------------ |
| Retroalimentación inmediata | Compilador TypeScript | Errores de tipo, errores de sintaxis |
| Verificación funcional | Pruebas unitarias | Errores de lógica, casos límite |
| Verificación de integración | Comando Build | Problemas de dependencias, errores de configuración |
| Verificación en tiempo de ejecución | Skill dev-browser | Problemas de renderizado de la interfaz (proyectos frontend) |
> Para stories de frontend, Ralph recomienda agregar en los criterios de aceptación: "Verify in browser using dev-browser skill", para que la IA abra realmente el navegador y confirme que la página se renderiza correctamente.
### Cuando las verificaciones de calidad fallan
Si las verificaciones de calidad de un story fallan repetidamente, Ralph no reintenta el mismo story indefinidamente. Cuando se alcanza el límite de iteraciones, se detiene y deja el estado actual. Usted puede:
1. Revisar progress.txt para ver dónde se atascó la IA
2. Corregir el problema manualmente y volver a ejecutar
3. Ajustar la granularidad del story (tal vez sea demasiado grande)
4. Agregar más contexto en las notas
### Personalización del Prompt
La plantilla de prompt de Ralph (CLAUDE.md o prompt.md) es el medio principal para controlar el comportamiento de la IA. Después de la instalación, debe personalizarla según su propio proyecto. Direcciones clave de personalización:
**Restricciones de estilo de código**
```markdown
## Convenciones de código
- Usar el modo estricto de TypeScript
- Preferir exportaciones nombradas sobre exportaciones por defecto
- Usar componentes de fumadocs para contenido MDX
- Seguir los patrones de nomenclatura de archivos existentes (kebab-case)
```
**Errores comunes**
```markdown
## Advertencias conocidas
- Archivos MDX: siempre importar componentes al inicio
- i18n: los archivos en inglés usan el sufijo `.en.mdx`
- Enlaces: las páginas en inglés deben usar el prefijo `/en/`
- QuoteCard: establecer `defaultLang` para que coincida con el idioma del archivo
```
**Manejo cuando se queda atascado**
```markdown
## Cuando se quede atascado
Si no puede completar un story después de 3 intentos dentro de la misma iteración:
1. Documente lo que está bloqueando en progress.txt
2. Pase al siguiente story si es posible
3. NO modifique archivos que no estén relacionados con el story actual
```
***
## Caso práctico: completar la traducción i18n de un blog con Ralph
Para mostrar cómo opera Ralph en un proyecto real, aquí compartimos un caso real: el uso de un agente autónomo estilo Ralph para traducir un blog completo del chino al inglés.
### Configuración del proyecto
El proyecto requería traducir más de 22 archivos de contenido (artículos de blog, documentación, metadatos de navegación) del chino al inglés, con el objetivo de lograr soporte i18n en un blog Next.js basado en fumadocs. Las tareas se definieron en un archivo `prd.json` que contenía 16 user stories, cada uno con criterios de aceptación claros:
```
scripts/ralph/
├── prd.json # 16 user stories con criterios de aceptación
└── progress.txt # Registro de experiencias, actualizado después de cada story
```
Cada user story seguía un patrón consistente:
* **Entregable claro**: "Create content/blog/xxx.en.mdx"
* **Criterios verificables**: "Typecheck passes", "Internal links use /en/ prefix"
* **Restricciones técnicas**: "Keep code blocks untranslated", "Set defaultLang='en' on QuoteCard"
### Modo de ejecución
El agente siguió los principios fundamentales de la metodología Ralph:
1. **Los archivos como fuente de verdad**: `prd.json` rastrea el estado de cada story (`passes: true/false`). `progress.txt` acumula experiencia entre iteraciones, por ejemplo: "El comando Typecheck es `pnpm types:check`, no `pnpm typecheck`"
2. **Controles de calidad automatizados**: después de cada traducción, se ejecuta `pnpm types:check` para verificar que los archivos MDX se compilen correctamente. Si el typecheck falla, se corrige el problema antes de hacer commit.
3. **Avance incremental**: cada story se hace commit de forma independiente, con mensajes de commit descriptivos (`feat: [US-003] - Translate blog/claude-code-quality-control.mdx`), facilitando la reversión cuando sea necesario.
4. **Ejecución en paralelo**: para artículos más extensos, múltiples subagentes traducen simultáneamente; por ejemplo, US-010 (claude-skills concept + practice), US-011 (speckit concept + practice) y US-012 (claude-architecture + claude-subagent) se ejecutaron en paralelo al mismo tiempo.
### Lecciones clave
| Lección | Detalles |
| --------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **La acumulación de conocimiento es importante** | Los patrones descubiertos en los stories iniciales (el `defaultLang` de QuoteCard, las reglas de prefijo de enlaces) permitieron completar los stories posteriores más rápido |
| **Typecheck como bucle de retroalimentación** | Detecta imports faltantes o MDX con formato incorrecto antes de que los problemas se acumulen |
| **La paralelización escala** | Con 6 agentes de traducción ejecutándose simultáneamente, el tiempo de finalización fue similar al de un solo agente |
| **La granularidad del PRD es crucial** | Cada story se limitó a 1-2 archivos: lo suficientemente pequeño para completarse de forma confiable, lo suficientemente grande para ser significativo |
| **El registro de progreso evita errores repetidos** | La sección "Codebase Patterns" de `progress.txt` se convirtió en una base de conocimiento que evitó volver a cometer los mismos errores |
### Resultados
Los 16 user stories se completaron en una sola sesión: se crearon 8 archivos meta.en.json de navegación, se tradujeron 3 artículos de blog, se tradujeron 12 páginas de documentación y la compilación completa del sitio pasó la verificación. Cada traducción mantuvo una calidad consistente porque los criterios de aceptación eran claros y el bucle de retroalimentación (typecheck) detectaba problemas de inmediato.
Este proyecto demostró el **Modo de Implementación Completa (Full Implementation Mode)** de Ralph: tareas claramente definidas, criterios de éxito explícitos, verificación automatizada y entrega incremental a través del sistema de archivos.
***
## Mejores prácticas y preguntas frecuentes
### Control de costos
La ejecución automatizada de Ralph significa que los costos de API se generan de forma continua. Algunas medidas de control:
* **Siempre establecer `max_iterations`**: esta es la red de seguridad más básica
* **Granularidad razonable de los stories**: stories demasiado grandes consumen múltiples iteraciones; stories demasiado fragmentados aumentan la sobrecarga de inicio
* **Probar a pequeña escala primero**: en proyectos nuevos, haga una prueba con 3 a 5 iteraciones para confirmar que el prompt y los controles de calidad funcionan correctamente antes de escalar
### Errores comunes
**Error 1: Story demasiado grande**
Síntoma: un story falla repetidamente y las iteraciones se agotan rápidamente.
Solución: divídalo en 2-3 stories más pequeños. "Construir un sistema de autenticación completo" se divide en "Implementar la API de inicio de sesión" + "Crear la página de inicio de sesión" + "Agregar middleware JWT".
**Error 2: sin bucle de retroalimentación**
Síntoma: Ralph afirma que el story está completado, pero el código en realidad tiene problemas.
Solución: agregue comandos de verificación ejecutables en los criterios de aceptación. "El código está listo" no es un criterio de aceptación; "pnpm test pasa completamente" sí lo es.
**Error 3: progress.txt no se utiliza**
Síntoma: el mismo error aparece repetidamente en diferentes iteraciones.
Solución: confirme que la plantilla del prompt tiene instrucciones claras de "leer progress.txt y seguir las lecciones allí registradas". Si la IA no agrega contenido de aprendizaje automáticamente, añada en el prompt: "After each story, append learnings to progress.txt".
**Error 4: orden de dependencias incorrecto**
Síntoma: un story depende de código que aún no existe, lo que hace que la implementación falle.
Solución: configure correctamente el campo `dependsOn` y asegúrese de que los stories de infraestructura se ejecuten primero.
### Preguntas frecuentes
**P: ¿Se puede modificar manualmente prd.json para intervenir a mitad del proceso?**
Sí. Ralph vuelve a leer prd.json al inicio de cada iteración. Puede modificar la descripción de un story entre iteraciones, agregar nuevos stories o marcar manualmente un story como `passes: true` (para omitirlo).
**P: ¿Qué hacer si Ralph se queda atascado fallando repetidamente en un story?**
1. Revise progress.txt para ver la causa del fallo
2. Agregue más contexto en las notas
3. Divida el story (puede que la granularidad sea demasiado grande)
4. Corrija manualmente el problema de bloqueo y vuelva a ejecutar
**P: ¿Puedo hacer otras cosas mientras Ralph está en ejecución?**
Sí. Ralph está diseñado como "Human on the Loop": no necesita supervisarlo constantemente. En modo AFK, inícielo antes de salir del trabajo y revise los resultados al día siguiente. Simplemente no modifique los archivos que Ralph está operando mientras se ejecuta.
**P: ¿Cómo controlar los costos?**
Tres métodos: establecer un `max_iterations` razonable, mantener una granularidad adecuada de los stories (reducir iteraciones innecesarias) y hacer pruebas a pequeña escala primero para confirmar que el flujo funciona correctamente. En general, un proyecto con 10 a 20 stories está en el rango de $50 a $100 de costo de API.
***
## Resumen
El flujo de uso de Ralph se puede resumir en cinco pasos:
```
Instalación → Redacción del PRD → Configuración de controles de calidad → Ejecución del bucle → Revisión de resultados
```
La idea central permanece invariable: **hacer que los archivos sean la fuente de verdad, que cada iteración sea un comienzo completamente nuevo y que los controles de calidad se encarguen de la supervisión por usted**.
Ahora, regrese a su proyecto, prepare el prd.json, ejecute `./scripts/ralph/ralph.sh --tool claude` y vaya a tomarse un café.
### Lecturas complementarias
* [Análisis profundo de Ralph Wiggum](/es/docs/notes/ralph-wiggum/concept) -- Repaso de los principios fundamentales de Ralph
* [Guía práctica de frankbria/ralph-claude-code](/es/docs/notes/ralph-wiggum/frankbria) -- Implementación de Ralph orientada a ingeniería: monitoreo, circuit breaker y mecanismos de seguridad
* [Análisis profundo de GSD](/es/docs/notes/gsd/concept) -- Sistema completo de ingeniería de contexto construido sobre Ralph
* [Qué son los Claude Skills](/es/docs/notes/claude-skills/concept) -- El skill PRD de Ralph es un Claude Skill
* [Guía práctica de Speckit](/es/docs/notes/speckit/practice) -- Otro flujo de trabajo estructurado para programación con IA
# Introducción conceptual
## Introducción
En octubre de 2025, GitHub publicó como código abierto un conjunto de herramientas llamado Spec Kit, introduciendo oficialmente el concepto de "Desarrollo Basado en Especificaciones" (Spec-Driven Development) en el panorama de la programación con IA. Esta idea aparentemente retro — escribir especificaciones antes del código — se está convirtiendo en el nuevo paradigma para aprovechar las herramientas de programación con IA.
Si utiliza regularmente asistentes de programación con IA como Claude Code, Cursor o GitHub Copilot, probablemente haya experimentado esta frustración: usted dice "agregue una función de inicio de sesión de usuario", la IA produce con entusiasmo un montón de código, pero al examinarlo de cerca — usa un framework que usted no conoce, la estrategia de seguridad difiere de lo esperado, el estilo de la interfaz no coincide... Entonces comienza ronda tras ronda de correcciones hasta quedar agotado.
¿Dónde está el problema? No es que la IA no sea lo suficientemente inteligente, sino que usted no ha proporcionado suficiente información. "Agregar una función de inicio de sesión de usuario" parece claro, pero en realidad oculta cientos de decisiones no declaradas: ¿Qué método de autenticación? ¿Cuáles son los requisitos de contraseña? ¿Cómo manejar los inicios de sesión fallidos? ¿Se debe recordar el estado de sesión? ¿Soportar inicio de sesión con terceros?... La IA no tiene más opción que adivinar, y adivinar significa desviación.
El desarrollo basado en especificaciones nació precisamente para resolver este problema.
## Vibe Coding: el costo de la velocidad
A principios de 2025, el ex Director de IA de Tesla, Andrej Karpathy, acuñó el término "Vibe Coding" para describir un enfoque de desarrollo que consiste en "aceptar las sugerencias de la IA sin revisión profunda". El término se volvió viral e incluso fue nombrado Palabra del Año 2025 por el Diccionario Collins.
El atractivo del Vibe Coding es evidente: usted describe una idea, la IA genera código, y si parece que funciona, es suficiente. Para prototipado rápido, hackatones y scripts desechables, este enfoque es genuinamente eficiente. Pero cuando se aplica a sistemas de producción, surgen los problemas.
Historias como esta son comunes en la industria: consultas de base de datos generadas por IA funcionan bien en pruebas a pequeña escala, pero el sistema se arrastra bajo tráfico real; un módulo de autenticación improvisado pasa el QA, pero dos semanas después descubren que las cuentas desactivadas aún pueden acceder a las herramientas de administración. Según una encuesta de 2025 de Final Round AI, **16 de 18 CTOs han experimentado desastres en producción causados por código generado por IA**.
Esto no significa que el Vibe Coding sea inútil. La clave está en **conocer sus límites**:
| Escenario | Vibe Coding | Desarrollo basado en especificaciones |
| ------------------------- | ------------------- | ------------------------------------- |
| Prototipos / demos | Adecuado | Excesivo |
| Scripts desechables | Adecuado | Excesivo |
| Funciones de producción | Alto riesgo | Recomendado |
| Relacionado con seguridad | Peligroso | Esencial |
| Colaboración en equipo | Difícil de mantener | Recomendado |
El desarrollo basado en especificaciones busca mantener la eficiencia de la IA mientras evita las trampas del Vibe Coding.
## Qué es el desarrollo basado en especificaciones
La idea central del desarrollo basado en especificaciones se puede resumir en una frase: **Definir "qué construir" antes de considerar "cómo construirlo"**.
Esto suena como un lugar común de la ingeniería de software, pero en la era de la programación con IA, adquiere un nuevo significado. Los documentos de requisitos tradicionales están escritos para humanos, a menudo extensos, vagos y llenos de jerga. La "especificación" en el desarrollo basado en especificaciones está escrita para la IA: concisa, estructurada y ejecutable.
Imagine que va a construir una casa. El enfoque tradicional de programación con IA es como decirle al equipo de construcción "constrúyanme una casa cómoda de tres habitaciones" y dejar que ellos lo resuelvan. El resultado podría estar bien, pero es más probable que sea muy diferente de lo que usted imaginó. El desarrollo basado en especificaciones significa dibujar primero el plano arquitectónico: cuántos pisos, metros cuadrados por piso, orientación de las ventanas, especificaciones de materiales... El equipo construye según el plano, y el resultado naturalmente cumple con las expectativas.
En la programación con IA, este plano es la **especificación**. No se preocupa por qué lenguaje de programación o framework usar, solo por qué debe lograr la función, qué tareas necesitan completar los usuarios y cuáles son los criterios de éxito.
Comparado con los flujos de trabajo de desarrollo tradicionales, el desarrollo basado en especificaciones tiene una diferencia fundamental:
| Programación tradicional con IA | Desarrollo basado en especificaciones |
| -------------------------------------------------------- | -------------------------------------------------------- |
| Describir requisitos directamente -> La IA genera código | Requisitos -> Especificación -> Plan -> Tareas -> Código |
| La IA debe adivinar muchos detalles | Cada paso es explícito; la IA solo necesita ejecutar |
| Retrabajo frecuente, alto costo de comunicación | Inversión inicial, ejecución fluida después |
| Adecuado para tareas simples | Adecuado para funciones complejas |
Este proceso de "refinamiento progresivo" es la esencia del desarrollo basado en especificaciones. Usted no pasa de cero al código de un salto, sino que aclara progresivamente los requisitos a través de múltiples etapas, cada una de las cuales puede ser revisada y ajustada.
## Resumen del flujo de trabajo de Speckit
El Spec Kit de GitHub y el comando speckit en Claude Code siguen un flujo de trabajo similar, que se puede dividir aproximadamente en seis etapas:
```
Constitution -> Specify -> Clarify -> Plan -> Tasks -> Implement
| | | | | |
Carta del Especificación Resolver Plan Desglose Ejecutar
Proyecto de función ambigüedad técnico de tareas implementación
```
**1. Constitution (Carta del Proyecto)**
La constitución del proyecto define los principios y restricciones fundamentales de todo el proyecto, como "pruebas primero", "simplicidad ante todo", "API primero", etc. Estos principios se aplican a todas las etapas posteriores, asegurando que las soluciones generadas por la IA se alineen con sus preferencias técnicas.
**2. Specify (Especificación de función)**
Este es el primer paso crítico. Usted describe la función deseada en lenguaje natural, y la IA le ayuda a organizarla en un documento de especificación estructurado, que incluye:
* Historias de usuario: quién necesita hacer qué, y por qué
* Requisitos funcionales: capacidades que el sistema debe tener
* Criterios de éxito: cómo determinar si la función cumple con los estándares
Es importante destacar que el documento de especificación **se enfoca únicamente en "qué construir", no en "cómo construirlo"** — sin stacks tecnológicos específicos, sin estructura de código.
**3. Clarify (Resolver ambigüedad)**
La IA revisa la especificación en busca de puntos ambiguos y formula hasta 5 preguntas clave. Estas generalmente involucran límites de funcionalidad, tipos de usuario, requisitos de seguridad, etc. A través de este intercambio de preguntas y respuestas, la especificación se vuelve mucho más clara.
**4. Plan (Plan técnico)**
Solo después de tener una especificación clara se comienza a considerar el enfoque técnico. Este paso produce:
* Selección de tecnología (lenguaje, framework, base de datos)
* Diseño del modelo de datos
* Definiciones de contratos de API
* Informes de investigación (resolviendo decisiones técnicas)
**5. Tasks (Desglose de tareas)**
El plan técnico se desglosa en una lista de tareas ejecutables. Cada tarea tiene un ID claro, descripción y rutas de archivos, lista para ser entregada a la IA para su ejecución. Las tareas se agrupan por historia de usuario y soportan desarrollo en paralelo.
**6. Implement (Ejecutar)**
Las tareas se ejecutan una por una de la lista. Cada tarea completada se marca como terminada, asegurando la trazabilidad.
Los entregables de estas seis etapas forman una cadena clara:
| Etapa | Entregable | Propósito |
| ------------ | -------------------- | -------------------------------- |
| Constitution | constitution.md | Definir principios del proyecto |
| Specify | spec.md | Describir requisitos funcionales |
| Clarify | spec.md actualizado | Eliminar ambigüedades |
| Plan | plan.md, research.md | Diseño de solución técnica |
| Tasks | tasks.md | Lista de tareas ejecutables |
| Implement | Código real | Entregable final |
## Por qué funciona este enfoque
El desarrollo basado en especificaciones tiene éxito donde "simplemente dejar que la IA escriba código" fracasa porque aborda la contradicción central de la programación con IA: la **asimetría de información**.
Cuando usted dice "agregar una función para compartir fotos", puede tener una imagen completa en su mente, pero la IA solo ve esas pocas palabras. Debe adivinar: ¿compartir dónde? ¿Quién puede verlo? ¿Necesita compresión? ¿Marcas de agua? ¿Soporte por lotes?... Cada suposición puede ser incorrecta.
El desarrollo basado en especificaciones resuelve esto "obligándole a pensar las cosas primero". Cuando se le requiere escribir historias de usuario, requisitos funcionales y criterios de éxito, los detalles que usted asumía como "obvios" salen a la superficie. Este proceso en sí mismo es valioso — incluso sin IA, escribir los requisitos claramente reduce los costos de comunicación.
Además, el refinamiento progresivo expone los errores más temprano. Descubrir una desviación en los requisitos en la etapa de Specify no cuesta prácticamente nada de corregir; descubrirlo después de que el código está escrito podría significar empezar de nuevo.
Por supuesto, el desarrollo basado en especificaciones no es una solución mágica. Tiene casos de uso claros:
**Buena opción**:
* Desarrollo de funciones complejas (que involucran múltiples módulos e interacciones)
* Proyectos de colaboración en equipo (los documentos de especificación sirven como medio de comunicación)
* Escenarios de alta calidad (que requieren trazabilidad y verificabilidad)
**No es buena opción**:
* Correcciones simples de errores o cambios menores
* Programación exploratoria (cuando aún no sabe qué construir)
* Situaciones con tiempo extremadamente limitado (sin tiempo para escribir especificaciones)
La clave está en reconocer la complejidad de la tarea. Algo que toma una hora en completarse no necesita una hora de escritura de especificaciones; una función que requiere una semana de desarrollo absolutamente justifica dos horas de escritura de especificaciones.
## Pero las especificaciones no son una solución mágica
Es necesario aclarar un malentendido común: **el desarrollo basado en especificaciones reduce las suposiciones, pero no elimina la necesidad de revisión**.
Incluso con una especificación completa, la IA aún puede:
* Pasar por alto casos extremos (escenarios que la especificación no cubrió)
* Generar código que no cumple con los requisitos de rendimiento
* Introducir vulnerabilidades de seguridad potenciales
* Producir implementaciones con estilo inconsistente
Esto es como la construcción: incluso con planos detallados, la inspección sigue siendo necesaria. Usted no se mudaría a una casa solo porque el equipo terminó de construirla según los planos — verificaría que el cableado eléctrico sea seguro, que la plomería funcione y que las puertas y ventanas estén firmes.
El valor del desarrollo basado en especificaciones radica en **hacer que los errores sean más fáciles de encontrar**, no en eliminar los errores en sí.
## Resumen
El núcleo del desarrollo basado en especificaciones es una verdad simple: **cuanto más compleja sea la tarea, más necesita pensar las cosas antes de comenzar**. Las herramientas de programación con IA amplifican la importancia de esta verdad — porque la IA ejecutará fielmente sus instrucciones, pero no puede realmente comprender su intención.
Recuerde tres puntos clave:
| Punto clave | Significado |
| ----------------------------------- | --------------------------------------------------------------------- |
| **Especificación antes del código** | Definir "qué construir" antes de considerar "cómo construirlo" |
| **Refinamiento progresivo** | De lo vago a lo claro, con revisión y ajuste en cada paso |
| **Reducir las suposiciones** | Especificaciones claras = menos espacio para la especulación de la IA |
Ahora que comprende la filosofía, el siguiente artículo, [Guía práctica de Speckit](/es/docs/notes/speckit/practice), le guiará en la práctica: cómo usar el comando speckit para completar un flujo de trabajo de desarrollo basado en especificaciones para una función.
Combinado con [Claude Skills](/es/docs/notes/claude-skills/concept), puede automatizar y estandarizar aún más la ejecución de especificaciones.
# Guía práctica
## Introducción
En el [artículo anterior](/es/docs/notes/speckit/concept), exploramos la filosofía del desarrollo dirigido por especificaciones — definir "qué construir" antes de pensar en "cómo construirlo". Aunque este paso adicional pueda parecer innecesario, reduce drásticamente el retrabajo y los costos de comunicación en la programación asistida por IA.
En este artículo, pasamos a la práctica. Aprenderá a utilizar la suite de comandos speckit para completar el flujo de trabajo completo, desde los requisitos hasta el código funcional.
## Instalación y configuración
Los comandos de Speckit provienen del proyecto oficial [Spec Kit](https://github.com/github/spec-kit) de GitHub. Dependiendo de su caso de uso, existen varias formas de integrarlos.
### Inicialización de un proyecto nuevo
Para proyectos nuevos, el enfoque recomendado es utilizar la herramienta oficial specify-cli:
```bash
# Instalar specify-cli usando uv
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git
# Inicializar un nuevo proyecto, especificando Claude como asistente de IA
specify init my-project --ai claude
```
Esto crea automáticamente la estructura de directorios del proyecto, incluyendo el directorio de configuración `.specify/` y los archivos de plantilla relacionados.
### Integración con un proyecto existente
Los comandos de Speckit requieren archivos de configuración para funcionar. Para integrar speckit en un proyecto existente, utilice specify-cli:
```bash
cd your-existing-project
specify init . --ai claude # Nota: . se refiere al directorio actual
```
Esto crea lo siguiente en su proyecto:
```
your-project/
├── .specify/
│ ├── templates/ # Plantillas de especificaciones, planes, etc.
│ ├── scripts/ # Scripts auxiliares
│ └── memory/ # constitution.md
├── .claude/
│ └── commands/ # Configuraciones de comandos de Claude Code
│ ├── speckit.specify.md
│ ├── speckit.plan.md
│ └── ...
└── specs/ # Directorio de almacenamiento de especificaciones
```
La inicialización no sobrescribirá sus archivos existentes. Una vez completada, podrá utilizar la suite de comandos `/speckit.*` en Claude Code.
> **Nota**: Los comandos de speckit no están integrados en Claude Code — primero debe completar los pasos de inicialización descritos anteriormente. Ejecutar `/speckit.specify` sin la inicialización resultará en un error de "comando no encontrado".
***
## Referencia de comandos
Speckit proporciona un conjunto de comandos que soportan cada fase del desarrollo dirigido por especificaciones. Cada comando tiene entradas y salidas bien definidas, formando una cadena trazable.
### /speckit.specify — Crear una especificación de funcionalidad
Este es el punto de partida de todo el flujo de trabajo. Usted describe la funcionalidad deseada en lenguaje natural y la IA la organiza en un documento de especificación estructurado.
**Propósito**: Crear una especificación de funcionalidad a partir de una descripción en lenguaje natural
**Entrada**: Descripción de la funcionalidad (lenguaje natural)
**Salida**:
* `specs/[número]-[nombre-funcionalidad]/spec.md` — Documento de especificación
* Una nueva rama de git (por ejemplo, `001-user-auth`)
**Ejemplo de uso**:
```
/speckit.specify Quiero agregar una funcionalidad de inicio de sesión con autenticación por correo electrónico/contraseña y una opción de "recordarme"
```
Después de la ejecución, la IA:
1. Genera un nombre corto para la funcionalidad (por ejemplo, `user-auth`)
2. Crea una nueva rama de funcionalidad
3. Produce un documento de especificación con historias de usuario, requisitos funcionales y criterios de éxito
4. Marca las áreas poco claras con `[NEEDS CLARIFICATION]`
**Estructura central de un documento de especificación**:
```markdown
# Feature Specification: Inicio de sesión de usuario
## User Scenarios & Testing
### User Story 1 - Inicio de sesión (Priority: P1)
Los usuarios inician sesión en el sistema usando correo electrónico y contraseña...
**Acceptance Scenarios**:
1. Given correo y contraseña válidos, When se hace clic en iniciar sesión, Then se ingresa exitosamente al sistema
## Requirements
### Functional Requirements
- FR-001: El sistema debe soportar inicio de sesión con correo/contraseña
- FR-002: El sistema debe proporcionar una opción de "recordarme"
## Success Criteria
- SC-001: Los usuarios pueden completar el flujo de inicio de sesión en 30 segundos
```
Note que el documento de especificación **no contiene detalles técnicos** — no menciona frameworks, esquemas de bases de datos ni definiciones de API. Eso viene en fases posteriores.
***
### /speckit.clarify — Resolver ambigüedades
Después de redactar la especificación, pueden quedar áreas ambiguas. Este comando revisa la especificación y formula preguntas clave para ayudar a aclararlas.
**Propósito**: Identificar ambigüedades en la especificación y refinarla mediante preguntas y respuestas
**Entrada**: Documento spec.md existente
**Salida**: spec.md actualizado (con registros de aclaraciones)
**Ejemplo de uso**:
```
/speckit.clarify
```
Después de la ejecución, la IA:
1. Escanea la especificación en busca de puntos ambiguos
2. Los prioriza (Alcance > Seguridad > Experiencia de usuario > Detalles técnicos)
3. Formula una pregunta a la vez
4. Actualiza la especificación según sus respuestas
**Ejemplo de preguntas y respuestas**:
```markdown
## Question 1: Manejo de fallos de inicio de sesión
**Context**: La especificación menciona el inicio de sesión pero no especifica cómo se deben manejar los fallos.
**Recommended:** Opción B - Bloquear la cuenta después de 5 intentos fallidos consecutivos es una mejor práctica de seguridad
| Option | Description |
|--------|-------------|
| A | Mostrar solo mensaje de error, sin restricciones |
| B | Bloquear la cuenta durante 15 minutos después de 5 fallos consecutivos |
| C | Usar CAPTCHA para prevenir ataques de fuerza bruta |
Puede responder con una letra de opción (por ejemplo, "B"), decir "yes" para aceptar la recomendación, o proporcionar su propia respuesta.
```
Después de cada aclaración, el documento de especificación se actualiza automáticamente con un registro de aclaración:
```markdown
## Clarifications
### Session 2025-12-20
- Q: ¿Cómo se deben manejar los fallos de inicio de sesión? → A: Bloquear la cuenta durante 15 minutos después de 5 fallos consecutivos
```
***
### /speckit.plan — Generar un plan técnico
Una vez que la especificación está clara, se pasa a la fase de diseño técnico. Este paso produce un plan técnico y un informe de investigación.
**Propósito**: Generar un plan de implementación técnica a partir de la especificación
**Entrada**: Documento spec.md
**Salida**:
* `plan.md` — Plan técnico (arquitectura, modelos de datos, diseño de API)
* `research.md` — Informe de investigación (decisiones de selección tecnológica)
* `data-model.md` — Modelo de datos (si aplica)
* `contracts/` — Contratos de API (si aplica)
**Ejemplo de uso**:
```
/speckit.plan Estoy usando Next.js + Prisma + PostgreSQL
```
Puede agregar sus preferencias de stack tecnológico después del comando. Después de la ejecución, la IA:
1. Analiza los requisitos funcionales de la especificación
2. Investiga las mejores prácticas para las tecnologías relevantes
3. Diseña modelos de datos y estructuras de API
4. Produce un plan técnico completo
**Contenido central de un plan técnico**:
```markdown
# Implementation Plan: Inicio de sesión de usuario
## Technical Context
**Language/Version**: TypeScript 5.x
**Primary Dependencies**: Next.js 15, Prisma, PostgreSQL
**Authentication**: NextAuth.js with credentials provider
## Project Structure
src/
├── app/
│ └── (auth)/
│ ├── login/
│ └── api/auth/
├── lib/
│ └── auth/
└── prisma/
└── schema.prisma
## Data Model
- User: id, email, passwordHash, createdAt, updatedAt
- Session: id, userId, expiresAt
```
***
### /speckit.tasks — Descomponer tareas
Con el plan técnico listo, el siguiente paso es descomponerlo en una lista de tareas ejecutables.
**Propósito**: Dividir el plan técnico en una lista de tareas ejecutables
**Entrada**: Documento plan.md
**Salida**: `tasks.md` — Lista de tareas ordenada por dependencias
**Ejemplo de uso**:
```
/speckit.tasks
```
Después de la ejecución, la IA:
1. Extrae el enfoque técnico de plan.md
2. Extrae las prioridades de historias de usuario de spec.md
3. Genera tareas agrupadas por historia de usuario
4. Marca las tareas paralelizables con `[P]`
5. Especifica rutas de archivo concretas para cada tarea
**Formato de la lista de tareas**:
```markdown
## Phase 1: Setup
- [ ] T001 Crear estructura del proyecto
- [ ] T002 [P] Configurar esquema de Prisma
- [ ] T003 [P] Configurar NextAuth
## Phase 2: User Story 1 - Inicio de sesión (P1)
- [ ] T004 [US1] Crear modelo User en prisma/schema.prisma
- [ ] T005 [US1] Implementar API de inicio de sesión en src/app/api/auth/[...nextauth]/route.ts
- [ ] T006 [US1] Crear página de inicio de sesión en src/app/(auth)/login/page.tsx
```
Cada tarea incluye:
* **ID de tarea** (T001, T002...) — para seguimiento
* **Marcador \[P]** — indica que puede ejecutarse en paralelo con otras tareas \[P]
* **Etiqueta \[US]** — indica a qué historia de usuario pertenece
* **Ruta de archivo** — especifica exactamente qué archivo se debe modificar
***
### /speckit.implement — Ejecutar la implementación
Todo está listo — es hora de ejecutar la lista de tareas.
**Propósito**: Ejecutar las tareas de la lista una por una
**Entrada**: Documento tasks.md
**Salida**: Código real
**Ejemplo de uso**:
```
/speckit.implement
```
Antes de la ejecución, la IA verifica la lista de verificación (si existe una). Durante la ejecución:
1. Las tareas se ejecutan en orden de fase
2. Cada tarea completada se marca como `[X]`
3. Se respetan las dependencias entre tareas
4. Las tareas paralelas pueden ejecutarse simultáneamente
**Ejemplo de ejecución**:
```
Phase 1: Setup
✓ T001 Crear estructura del proyecto
✓ T002 Configurar esquema de Prisma
✓ T003 Configurar NextAuth
Phase 2: User Story 1
✓ T004 Crear modelo User
Ejecutando T005...
```
### Revisión posterior a la implementación
Después de que `/speckit.implement` finalice, **no fusione el código directamente**. El código generado por IA requiere revisión humana:
**Pasos de verificación obligatorios**:
1. **Ejecutar la suite de pruebas**
```bash
npm test # o su comando de pruebas
```
Asegúrese de que la IA no haya roto la funcionalidad existente.
2. **Lista de verificación para revisión de código**
* ¿El código coincide con la intención de la especificación (comparar con spec.md)?
* ¿Sigue el estilo de codificación del proyecto?
* ¿Existen posibles problemas de seguridad?
3. **Pruebas de casos límite**
Pruebe manualmente los casos límite que la IA pueda haber pasado por alto:
* Manejo de valores nulos
* Entradas extremas
* Escenarios de concurrencia
* Rutas de error
4. **Verificación de rendimiento**
Si hay operaciones de base de datos o llamadas a API involucradas, verifique si existen consultas N+1 y problemas de rendimiento similares.
> **Consejo**: Incluso con una especificación exhaustiva, la IA aún puede desviarse en los detalles de implementación. La revisión no es una señal de desconfianza en el desarrollo dirigido por especificaciones — es parte de la disciplina de ingeniería.
***
### /speckit.analyze — Análisis de consistencia
Este es un paso opcional de verificación de calidad que valida la consistencia entre la especificación, el plan y las tareas.
**Propósito**: Análisis de consistencia y calidad entre documentos
**Entrada**: spec.md, plan.md, tasks.md
**Salida**: Informe de análisis (no se modifica ningún archivo)
**Ejemplo de uso**:
```
/speckit.analyze
```
Después de la ejecución, verifica:
* Si cada requisito tiene una tarea correspondiente
* Si las tareas cubren todas las historias de usuario
* Si la terminología es consistente
* Si hay omisiones o duplicaciones
***
### Otros comandos (opcionales)
Además de los comandos principales descritos anteriormente, speckit proporciona varios comandos auxiliares. Estos no forman parte del flujo de trabajo principal, pero son útiles en escenarios específicos.
**`/speckit.constitution`** — Crear una constitución del proyecto
Se utiliza para definir los principios y estándares de desarrollo del proyecto. Ideal para proyectos de equipo para garantizar que todos los miembros sigan estándares de desarrollo unificados.
* **Entrada**: Preguntas y respuestas interactivas o principios proporcionados directamente
* **Salida**: Archivo de constitución del proyecto `.specify/constitution.md`
* **Caso de uso**: Inicialización de nuevos proyectos de equipo, unificación de estilo de código y decisiones de arquitectura
**`/speckit.checklist`** — Generar una lista de verificación de calidad
Genera una lista de verificación de calidad personalizada basada en la especificación de la funcionalidad, utilizada para asegurar la calidad antes de la implementación.
* **Entrada**: Documento spec.md
* **Salida**: Listas de verificación en el directorio `checklists/`
* **Caso de uso**: Controles de calidad antes del lanzamiento de funcionalidades importantes, referencia para revisión de código
**`/speckit.taskstoissues`** — Convertir tareas en GitHub Issues
Convierte automáticamente las tareas de tasks.md en GitHub Issues para la colaboración en equipo y la asignación de tareas.
* **Entrada**: Documento tasks.md
* **Salida**: GitHub Issues (creados a través del CLI gh)
* **Caso de uso**: Colaboración en equipo, planificación de sprints, seguimiento de tareas
***
## Ecosistema de herramientas
Los comandos speckit presentados en este artículo provienen del proyecto [GitHub Spec Kit](https://github.com/github/spec-kit). Más allá de esto, en 2025, varias herramientas importantes de programación con IA comenzaron a soportar flujos de trabajo similares dirigidos por especificaciones:
| Herramienta | Características | Ideal para |
| --------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- |
| **[GitHub Spec Kit](https://github.com/github/spec-kit)** | La herramienta utilizada en este artículo, licencia MIT, soporta Claude Code / Copilot / Gemini CLI | Usuarios de línea de comandos, colaboración entre herramientas |
| **[AWS Kiro](https://kiro.dev/)** | Fork de VS Code, flujo de trabajo visual, notación EARS | Usuarios orientados a GUI, ecosistema AWS |
| **[JetBrains Junie](https://blog.jetbrains.com/junie/)** | Integración con el ecosistema IntelliJ, modo de razonamiento Think More | Usuarios de JetBrains IDE |
| **Cursor Plan Mode** | Fase de planificación integrada, planes de ejecución generados automáticamente | Desarrolladores que ya usan Cursor |
**Cómo elegir**:
* Si utiliza Claude Code, GitHub Copilot o Gemini CLI, se recomienda GitHub Spec Kit
* Si prefiere interfaces gráficas y flujos de trabajo visuales, pruebe AWS Kiro
* Si es usuario de JetBrains, Junie ofrece una integración más natural con el IDE
* Si ya está usando Cursor, su Plan Mode proporciona capacidades de planificación similares
La filosofía central es la misma — la herramienta es solo el vehículo; lo que importa es la mentalidad de **primero la especificación, después el código**.
***
## Demostración de extremo a extremo
Recorramos un ejemplo del mundo real utilizando el flujo de trabajo completo de speckit.
**Escenario**: Agregar una funcionalidad de marcadores de artículos a un sistema de blog
### Paso 1: Crear la especificación
```
/speckit.specify Quiero agregar una funcionalidad de marcadores para que los usuarios puedan guardar sus artículos favoritos y verlos en su perfil
```
La IA genera el documento de especificación `specs/002-article-bookmark/spec.md`:
```markdown
# Feature Specification: Marcadores de artículos
## User Scenarios & Testing
### User Story 1 - Marcar un artículo (Priority: P1)
Mientras lee un artículo, el usuario hace clic en el botón de marcador para agregarlo a su lista de marcadores.
**Acceptance Scenarios**:
1. Given el usuario ha iniciado sesión, When hace clic en el botón de marcador, Then el artículo se agrega a la lista de marcadores
2. Given el usuario no ha iniciado sesión, When hace clic en el botón de marcador, Then se muestra un aviso de inicio de sesión
### User Story 2 - Ver lista de marcadores (Priority: P2)
El usuario visualiza todos sus artículos marcados en su perfil.
## Requirements
- FR-001: Los usuarios deben poder marcar/desmarcar artículos
- FR-002: El botón de marcador debe reflejar el estado actual
- FR-003: El perfil debe mostrar la lista de marcadores
## Success Criteria
- SC-001: Las operaciones de marcado se completan en menos de 500ms
- SC-002: La lista de marcadores soporta paginación con 10 artículos por página
```
### Paso 2: Aclarar requisitos
```
/speckit.clarify
```
La IA pregunta: "¿Existe un límite superior en la cantidad de marcadores?"
Respuesta: "Un máximo de 100 artículos marcados por usuario"
La especificación se actualiza con:
* FR-004: Cada usuario puede marcar un máximo de 100 artículos
* Se muestra una notificación cuando se alcanza el límite
### Paso 3: Generar el plan
```
/speckit.plan Usando Next.js + Prisma
```
El plan técnico generado incluye:
* Modelo Bookmark (userId, articleId, createdAt)
* Diseño de rutas de API (POST/DELETE /api/bookmarks)
* Diseño de componentes (BookmarkButton, BookmarkList)
### Paso 4: Descomponer tareas
```
/speckit.tasks
```
La lista de tareas generada:
```markdown
## Phase 1: Setup
- [ ] T001 Agregar modelo Bookmark al esquema de Prisma
## Phase 2: US1 - Marcar artículo
- [ ] T002 [US1] Crear API de marcadores en src/app/api/bookmarks/route.ts
- [ ] T003 [US1] Crear componente BookmarkButton en src/components/BookmarkButton.tsx
- [ ] T004 [US1] Integrar en la página de artículos
## Phase 3: US2 - Lista de marcadores
- [ ] T005 [US2] Crear página de lista de marcadores en src/app/profile/bookmarks/page.tsx
- [ ] T006 [US2] Implementar lógica de paginación
```
### Paso 5: Ejecutar la implementación
```
/speckit.implement
```
Las tareas se ejecutan en orden y cada tarea completada se marca como `[X]`.
***
## Mejores prácticas y consideraciones
### Cuándo usar Speckit
**Escenarios adecuados**:
* Desarrollo de nuevas funcionalidades (que involucren 3+ archivos)
* Cuando los requisitos no están completamente claros (use clarify para resolverlos)
* Proyectos colaborativos de múltiples personas (las especificaciones sirven como entendimiento compartido)
* Funcionalidades críticas (donde se necesita trazabilidad)
**Escenarios no adecuados**:
* Correcciones simples de errores
* Cambios de una sola línea de código
* Correcciones de emergencia
* Experimentos puramente exploratorios
### Errores comunes
Hay varios errores comunes a los que debe prestar atención al usar speckit:
**Error 1: Especificaciones demasiado vagas**
Síntoma: El código generado por la IA difiere significativamente de las expectativas, requiriendo retrabajo extensivo.
```markdown
# ❌ Especificación vaga
Los usuarios pueden buscar artículos
# ✓ Especificación clara
- FR-001: Los usuarios pueden buscar artículos por palabras clave del título
- FR-002: Los resultados de búsqueda se ordenan por relevancia, mostrando 10 por página
- FR-003: Los términos de búsqueda se resaltan en los resultados
- FR-004: Un término de búsqueda vacío muestra artículos populares
```
Solución: Ejecute `/speckit.clarify`, o agregue manualmente requisitos funcionales y criterios de éxito.
**Error 2: Especificaciones demasiado detalladas**
Síntoma: La IA está demasiado restringida y no puede aprovechar sus fortalezas, produciendo código rígido — o simplemente ignora partes de las instrucciones.
```markdown
# ❌ Sobre-especificado (dictando detalles de implementación)
Usar la función debounce de lodash con un retraso de 300ms,
envuelto en useCallback con [searchTerm] como dependencia...
# ✓ Nivel de detalle apropiado (solo indicar qué, no cómo)
La entrada de búsqueda debe tener debounce para evitar solicitudes excesivas
```
Solución: Mantenga las especificaciones en el nivel de "qué" y deje el "cómo" para la fase de Plan.
**Error 3: Omitir la fase de Plan**
Síntoma: Las tareas son demasiado generales o demasiado fragmentadas, lo que lleva a retrabajo frecuente durante la implementación y dependencias enredadas entre tareas.
Solución: Siempre complete la fase de Plan para funcionalidades complejas. La planificación no solo produce un enfoque técnico, sino que también ayuda a identificar posibles problemas de arquitectura.
**Error 4: Fusionar sin revisión**
Síntoma: Se descubren casos límite, vulnerabilidades de seguridad o problemas de rendimiento después del despliegue.
Solución: Consulte la sección "Revisión posterior a la implementación" anterior — siempre ejecute pruebas y realice revisión de código antes de fusionar.
### Preguntas frecuentes
**P: ¿Necesito seguir el flujo de trabajo completo para cada funcionalidad?**
No necesariamente. Los cambios simples pueden ir directamente al código. Para funcionalidades complejas, se recomienda completar al menos specify + plan.
**P: La especificación es muy detallada, pero la IA aún generó código inesperado.**
Verifique si la especificación es verdaderamente "detallada". A menudo pensamos que hemos sido claros, pero quedan ambigüedades. Intente ejecutar `/speckit.clarify` para ver si se omitió algo.
**P: ¿Puedo omitir ciertos pasos?**
Sí. El flujo de trabajo mínimo es specify → tasks → implement. Sin embargo, omitir clarify y plan puede aumentar el riesgo de retrabajo posterior.
**P: ¿Cómo modifico una especificación ya generada?**
Simplemente edite el archivo spec.md directamente. Después de realizar cambios, se recomienda volver a ejecutar plan y tasks para mantener la consistencia.
**P: El código generado por la IA es completamente incorrecto — ¿cómo lo depuro?**
Investigue por etapas:
1. **Verificar la especificación**: ¿La especificación es realmente clara? Intente ejecutar `/speckit.clarify` para ver si se omitió algo
2. **Verificar el plan**: ¿El enfoque técnico en plan.md es razonable? Si no lo es, edítelo directamente y regenere las tareas
3. **Reducir el alcance**: Haga que la IA ejecute solo una tarea y observe si el resultado cumple con las expectativas
4. **Agregar restricciones**: Agregue preferencias técnicas más explícitas en constitution.md
**P: ¿Qué hago si el Plan y las Tareas son inconsistentes?**
Ejecute `/speckit.analyze` para detectar inconsistencias. Causas comunes:
* El Plan se actualizó pero las Tareas no se regeneraron
* Las Tareas se editaron manualmente sin actualizar el Plan
* La especificación cambió pero solo se actualizaron algunos documentos
Solución: Tome spec.md como fuente de verdad y regenere plan.md y tasks.md en secuencia.
**P: ¿Cómo manejo las dependencias entre funcionalidades?**
Si la Funcionalidad B depende de la Funcionalidad A, hay dos enfoques:
1. **Fusionar especificaciones**: Escriba A y B en el mismo spec.md para que la IA las planifique juntas
2. **Desarrollar por fases**: Complete primero el flujo completo de la Funcionalidad A, luego comience el specify de la Funcionalidad B
No se recomienda desarrollar simultáneamente múltiples funcionalidades con dependencias, ya que esto fácilmente conduce a problemas de integración.
***
## Resumen
El valor central de speckit no es agregar proceso por sí mismo — se trata de **hacer explícito el conocimiento implícito**. Cuando se le requiere escribir historias de usuario, requisitos funcionales y criterios de éxito, los detalles que usted pensaba que eran "obvios" salen a la superficie naturalmente.
Recuerde este flujo de trabajo:
```
Specify → Clarify → Plan → Tasks → Implement
Definir Refinar Diseñar Descomponer Ejecutar
```
Cada paso reduce la ambigüedad para el siguiente. Al final, la IA recibe una lista de tareas clara en lugar de una descripción vaga de intenciones.
Ahora, regrese a su proyecto e intente iniciar su primer flujo de trabajo de desarrollo dirigido por especificaciones con `/speckit.specify`.
### Lectura adicional
* [Qué es el desarrollo dirigido por especificaciones](/es/docs/notes/speckit/concept) — Revisar la filosofía central
* [Análisis profundo de GSD](/es/docs/notes/gsd/concept) — Otro sistema de ingeniería de contexto basado en principios de especificaciones
* [Qué son los Claude Skills](/es/docs/notes/claude-skills/concept) — Speckit en sí mismo es un Claude Skill
# TDD en la era de la IA: deje que el modelo llegue primero a la luz roja
## Hablemos primero de la conclusión.
Después de que AI escribe código, TDD no está desactualizado, pero ha cambiado de posición.
En el pasado, cuando hablábamos de TDD, a menudo hablábamos de la autodisciplina de los programadores: escribir pruebas primero, luego escribir la implementación y refactorizar en pequeños pasos. Cuando se trata de programación de IA, se parece más a un sistema de frenos. Porque lo que el modelo hace mejor es también lo más peligroso: puede escribir rápidamente un gran fragmento de código que parece completo.
Si le pide que implemente una función, puede darle:
* un archivo de implementación
* un conjunto de pruebas
* una explicación
* Una palabra "hecho"
La cuestión es que "parecer completo" no se hace en el sentido de ingeniería. Para completar en un sentido de ingeniería, al menos responda:
> ¿Este comportamiento está definido por una prueba fallida explícita?
> ¿Esta falla se volvió verde debido a la implementación?
> Después de ponerse verde, ¿hemos ordenado el código sin cambiar el comportamiento?
Es por eso que se vuelve a hablar de TDD en la era de la IA.
No se trata de hacer que el proceso parezca avanzado, sino de reemplazar la "confianza en el modelo" por "confianza en la retroalimentación".
## 1. Malentendido común: TDD no es "escribir pruebas primero"
Mucha gente odia el TDD porque lo entiende como un ritual:
```text
先写测试。
再写代码。
最后跑一下。
```
Esto es ciertamente aburrido y puede convertirse fácilmente en formalismo.
El TDD verdaderamente útil no es "los archivos de prueba aparecen temprano", sino "las fallas aparecen lo suficientemente temprano".
### La clave no está en probar, sino en rojo.
El primer paso de TDD se llama ROJO, no PRUEBA.
ROJO significa: escriba primero una prueba para que el sistema falle claramente. Para que esto falle deben ser ciertas tres cosas:
1. Falla.
2. Falla porque la conducta objetivo no existe.
3. Falla de la forma esperada.
Si el rojo no se ve primero, el verde que sigue no tiene sentido.
Por ejemplo, desea implementar `slugify("Hello World") -> "hello-world"`. Un RED valioso no es "Escribí un archivo de prueba", sino:
```text
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
Failure: NameError: name 'slugify' is not defined
Reason: 目标函数还不存在,符合预期
```
Aquí es cuando la prueba se convierte en especificación. Te dice: Para lograr el siguiente paso, sólo necesitas hacer realidad este comportamiento.
### Primero se pone verde y luego inventa la prueba, normalmente inventa la historia.
Es fácil para la IA ir por el otro lado: escribir la implementación primero y agregar las pruebas más tarde.
Esta es una experiencia fluida. Cuando vea que el código se ha ejecutado y las pruebas están realizadas, se sentirá "casi" en su corazón. Pero tiene un problema fatal: es probable que la prueba sea sólo una implementación retroactiva de la implementación actual.
No se trata de preguntar "cuáles deberían ser los requisitos", sino "cómo escribir el código actual para que pueda aprobarse fácilmente".
Es por eso que las pruebas de escritura de IA suelen tener estos olores:
* Las afirmaciones son demasiado específicas para la implementación actual.
* Hay demasiadas burlas y no se miden los límites reales.
* solo prueba el camino feliz
* Para hacer que el código existente pase, haga afirmaciones muy amplias.
* Ninguna prueba puede demostrar que el código antiguo era originalmente incorrecto.
TDD requiere lo contrario: primero dejar que los requisitos fallen y luego dejar que el código se ponga al día con los requisitos.
## 2. Por qué TDD es más necesario en la era de la IA
La principal contradicción de la programación de IA no es que "el código se escribe lentamente", sino que "la retroalimentación llega tarde".
Sin TDD, normalmente trabajas así:
```text
描述需求 -> AI 写一堆代码 -> 人肉看 diff -> 跑一下 -> 发现问题 -> 回头修
```
Las preguntas se acumularán hasta el final. Para cuando descubras que está mal, es posible que se hayan mezclado tres categorías de cosas:
* Malentendido de los requisitos.
* La ruta de implementación es incorrecta.
* La refactorización rompe el comportamiento antiguo.
La función de TDD es acortar esta larga cadena.
### Le da al modelo un objetivo decidible.
"Escribir con elegancia" no es el objetivo.
"Volver automáticamente a la página de inicio de sesión después de que expire el estado de inicio de sesión del usuario" no es lo suficientemente específico.
Un mejor objetivo sería:
```text
当 access token 过期时:
1. 请求返回 401。
2. 客户端清理本地 session。
3. 用户被重定向到 /login。
4. 原始目标地址被保存在 redirect 参数里。
```
Vaya un paso más allá y convierta uno de ellos en una prueba fallida:
```text
given expired session
when user opens /settings
then app redirects to /login?redirect=/settings
```
En este momento, la IA ya no está adivinando "cómo manejar la expiración del estado de inicio de sesión", sino completando un comportamiento claro.
### Divide tareas grandes en pequeños bucles cerrados
El lugar más fácil para que la IA pierda el control es hacerlo todo de una vez.
Deje que implemente el inicio de sesión, los permisos, el token de actualización, los mensajes de error y los saltos de ruta a la vez, y al final obtendrá una gran diferencia. Podría funcionar, pero el costo de la revisión es alto. Tienes que juzgar el negocio, el estado, el enrutamiento, los límites, las pruebas y la refactorización al mismo tiempo.
El ritmo de TDD se parece más a esto:
```text
一个行为 -> 一个失败测试 -> 最小实现 -> 变绿 -> 再下一个行为
```
Avance sólo una pequeña cantidad a la vez. Es lo suficientemente pequeño como para que puedas entenderlo, lo suficientemente pequeño como para que a la IA le resulte difícil inventar historias y lo suficientemente pequeño como para poder localizar rápidamente cuando falla.
### Limita el modelo para "jugar sin problemas"
Un problema común con la IA es el exceso de entusiasmo.
Le pides que solucione un error en los límites y extrae el ayudante; le pide que agregue una prueba y cambia la implementación; le pides que refactorice y cambia su comportamiento.
TDD utiliza fases para separar estas acciones:
| Etapas | Qué hacer | Qué no hacer |
| -------- | -------------------------------- | ------------------------------------------- |
| ROJO | Escribe una prueba fallida | Escribir una implementación de producción |
| VERDE | Escriba la implementación mínima | Modifique la prueba para que se ponga verde |
| REFACTOR | Limpiar estructura | Introducir nuevos comportamientos |
Esta tabla es más útil que "Tenga cuidado". Le permite al modelo saber en qué etapa se encuentra y facilita que los humanos detecten violaciones de límites.
## 3. Reconstrucción del rojo y del verde: tres puertas, no tres lemas
“Rojo, Verde, Refactor” podría ser fácilmente un eslogan. En uso real, debería verse como tres puertas. Cada vez que pases por una puerta, deberás dejar constancia.
### La primera puerta: ROJA, lo que demuestra que no se han cubierto las necesidades.
Las preguntas más importantes durante la fase RED son:
> Si esta prueba falla, ¿prueba que todavía nos falta una conducta objetivo?
UN MAL ROJO:
```py
assert True
```
Tampoco es un muy buen ROJO:
```py
assert "hello" in format_title("Hello World")
```
Es demasiado ancho. También pasan muchas implementaciones defectuosas.
Mejor ROJO:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
Esta prueba es pequeña, pero está clara. Especifica entradas, salidas y comportamiento.
### Segunda puerta: VERDE, solo deja pasar la prueba actual
La fase VERDE no se trata de escribir la arquitectura final.
Sólo tiene una misión: pasar la prueba actualmente fallida con la menor cantidad de código.
Esta afirmación suena contraintuitiva. A mucha gente le preocupa si una "implementación mínima" será demasiado fea. Sí, a veces puede ser feo. Pero su valor radica en mantener la presión del diseño.
Si la primera prueba es:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
Un VERDE aceptable podría ser simplemente:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
No es necesario que admitas chino, acentos, puntuación continua, emojis y casos especiales de SEO de inmediato. Estos deberían ser impulsados por pruebas posteriores.
### La tercera puerta: REFACTOR, solo cambia la estructura, no el comportamiento
La etapa REFACTOR es la más fácil de confundir para la IA.
Interpretará "ordenar el código" como "mejorarlo por cierto". Esto no funcionará. La definición de refactorización es muy limitada: el comportamiento externo permanece sin cambios, pero la estructura interna mejora.
Una buena refactorización se ve así:
* Cambiar el nombre de la variable por uno más preciso.
* Extraer expresiones repetidas
* Eliminar ramas condicionales que sean demasiado profundas.
* Mover las ubicaciones de las funciones para aclarar las responsabilidades del módulo
Una mala refactorización se ve así:
* Ahora se admiten nuevas entradas
* El mensaje de error se ha cambiado fácilmente.
* Se cambiaron las dependencias fácilmente
* Cambió convenientemente la afirmación de la prueba.
Los criterios de juicio son simples:
> Si esta confirmación se llamó simplemente `refactor:`, debería ser del mismo color verde antes y después de la prueba, y el comportamiento del usuario debería ser el mismo.
## 4. El sabor de las buenas pruebas
TDD no se trata de que más pruebas sean mejores. La IA también es muy buena para generar un montón de pruebas que tienen poco valor.
Lo más importante es probar el sabor.
### Las buenas pruebas son como especificaciones
Una buena prueba debería leerse como una especificación comercial:
```text
当用户没有权限时,保存按钮不可点击。
当标题为空时,表单显示错误信息。
当重复提交同一个请求时,只创建一条记录。
```
Se ocupa del comportamiento externo, no de lo que se hace internamente.
Las malas pruebas parecen notas de implementación:
```text
应该调用 validateInput 三次。
应该读取 state.user.flags。
应该触发 handleClick 内部函数。
```
Una vez que los detalles de implementación estén resueltos, la refactorización será dolorosa. Acabas de cambiar la estructura interna, pero las pruebas fallaron a gran escala. En lugar de proteger el código, estas pruebas lo congelan.
### Las buenas pruebas tienen límites
Es mejor diseñar una prueba para responder solo una pregunta.
Si una prueba también afirma:
* Formato correcto
* Los permisos son correctos.
* La solicitud de red es correcta.
* La copia del brindis es correcta.
* El estado de la base de datos es correcto.
Cuando falla, es difícil saber cuál es el problema.
La IA es especialmente propensa a escribir pruebas tan “grandes y completas” porque quiere demostrar muchas cosas a la vez. TDD es todo lo contrario: una acción, un fracaso y una implementación.
### Las buenas pruebas dificultan que las implementaciones hagan trampa
Si la prueba solo cubre una entrada que es demasiado específica, la IA puede escribir una implementación falsa que simplemente coincida.
Por ejemplo:
```py
def slugify(text: str) -> str:
return "hello-world"
```
La primera prueba le permitirá pasar, pero la segunda prueba eliminará la lógica real:
```py
def test_slugify_handles_another_title():
assert slugify("Test Driven Development") == "test-driven-development"
```
Por lo tanto, TDD no siempre escribe solo una prueba, sino que solo agrega una presión conductual en cada ronda. La presión aumenta gradualmente y el diseño crece gradualmente.
## 5. ¿Cómo evitará la IA el TDD?
Esta parte debe quedar clara porque la IA no respeta naturalmente las pruebas.
Su objetivo de optimización es simple: completar la tarea que acabas de mencionar. Si dice "dejar pasar la prueba", es posible que se realicen algunas acciones que los humanos no quieren.
### El primer tipo: cambiar la prueba para que se ponga verde
Los más típicos:
* Cambie `assert slugify("Hello World") == "hello-world"` a la salida actual
* Eliminar afirmaciones fallidas
* Añade `skip` a la prueba.
* Cambiar afirmaciones estrictas por afirmaciones flexibles
Esto no es TDD, esto es apagar la luz roja.
### El segundo tipo: escribir implementación de sobreajuste
Por ejemplo, la prueba tiene sólo una entrada:
```py
def test_slugify_lowercases_and_uses_hyphen():
assert slugify("Hello World") == "hello-world"
```
El modelo podría escribirse como:
```py
def slugify(text: str) -> str:
if text == "Hello World":
return "hello-world"
return text
```
No es necesario que lo regañes en este momento. Debe continuar agregando el siguiente comportamiento para que la implementación no pueda continuar codificada.
### El tercer método: use simulacro para cubrir el límite real
A la IA le encantan las burlas. Los simulacros facilitan la redacción de pruebas y hacen desaparecer muchos problemas reales.
No es que no puedas burlarte, pero hay que preguntar:
> ¿De lo que me estoy burlando ahora es de una dependencia lenta, o es un límite que realmente quiero verificar?
Si desea verificar el análisis de devolución de llamada de pago pero simular la capa de análisis, la prueba no tendrá sentido.
## 6. Cuándo no utilizar TDD
TDD tiene valor, pero no todo vale la pena.
### Escena inadecuada
* Puro ajuste visual
* guión único
* Demostración de exploración tecnológica.
* Prototipos cuyos requisitos no han sido claramente pensados
* El marco de prueba aún no ha creado un almacén.
En estos escenarios, busque primero la velocidad de exploración y no se deje frenar por el proceso.
### Escena adecuada
* correcciones de errores
* Permisos, facturación, máquina de estados.
* Transformación de datos y manejo de límites.
* Módulos principales que se mantendrán durante mucho tiempo.
* Rutas de código que la IA modificará repetidamente
El criterio de valoración no es "si esta función es excelente o no", sino:
> Si está mal, ¿son obvios los costos?
El costo es obvio, por lo que vale la pena escribir la prueba primero.
## 7. Un método mental ejecutable
Si tuviera que darle a AI solo una frase, no diría:
```text
请高质量实现这个功能。
```
Yo diría:
```text
先写一个失败测试,运行它,确认失败原因符合预期。不要写实现,直到我说 go。
```
Esta frase es de mayor calidad porque no le pide al modelo que "se porte bien", sino que le pide que entre en un proceso que pueda comprobarse.
Un poco más completo:
```text
每轮只处理一个行为。
RED:写一个失败测试并运行。
GREEN:写最小实现,不改测试。
REFACTOR:只在绿色状态下整理结构。
每轮报告测试文件、命令、失败原因、通过结果。
```
Este es el núcleo de TDD en la era de la IA.
No supersticioso sobre pruebas o procesos, sino sobre tener evidencia para cada paso.
## Cierre
Lo que más necesita la programación de IA no es más código, sino comentarios más breves.
Este es el valor de TDD: convierte "Pensé que debería ser correcto" en "Aquí hubo un error y luego se puso verde". El cambio es pequeño, pero bastante real.
Si sólo recuerdas una frase, recuerda esto:
> No permita que la IA entregue el código directamente. Primero déjelo emitir una luz roja y luego deje que la luz roja se vuelva verde.
La próxima [Guía Práctica](/es/docs/notes/tdd-with-ai/practice) convertirá este ritmo en un flujo de trabajo que se puede copiar directamente.
## Recursos recomendados
# TDD en la era de la IA: Manual práctico del Codex
## Dame un mapa primero
[Concepto](/es/docs/notes/tdd-with-ai/concept) De qué se trata: En la programación de IA, el valor de TDD no es el ritual de “escribir una prueba primero”, sino crear primero una luz roja verificable y luego dejar que la implementación la ponga en verde.
Este artículo habla sobre cómo implementar Codex.
No preguntes "qué documentos necesito que coincidan" de inmediato. Una mejor pregunta es:
> ¿Cómo hago para que Codex actúe siempre con el mismo flujo de trabajo TDD?
Este flujo de trabajo se puede dividir en cuatro niveles:
\| Jerarquía | ¿Qué estás haciendo? Dónde | ¿Cuándo es apropiado?
\|------|------------|----------|--------------|
\| L1 | Escribir disciplinas del proyecto | `AGENTS.md` | Todos los proyectos deberían tener |
\| L2 | Proceso de solidificación | `.agents/skills/tdd-codex/SKILL.md` | Utilice TDD repetidamente para cumplir con los requisitos |
\| L3 | Etapa de aislamiento | `.codex/agents/*.toml` | Tareas complejas, miedo a una contaminación mutua entre las pruebas y la implementación |
\| L4 | Recordatorio automático | `.codex/hooks.json` | Importante almacén, temeroso de que la IA modifique en secreto la prueba |
Las versiones más pequeñas disponibles son L1 + L2.
La línea de defensa completa es L1 + L2 + L3 + L4.
## 1. Primero defina qué es "finalización"
Si no existen estándares de finalización, el Codex puede fácilmente tratar el "código escrito" como "hecho".
En un escenario TDD, los criterios de finalización deberían ser más específicos.
### Se deben entregar pruebas en cada ronda
Haga que Codex informe estos seis elementos en cada ronda:
```text
Behavior: 这一轮实现哪个行为
Test: 测试文件和测试名
Command: 跑了什么命令
RED: 失败原因是否符合预期
GREEN: 通过结果是什么
REFACTOR: 是否重构,为什么
```
Esto es mucho más útil que un "hecho".
Le permite saber que el modelo realmente ha pasado por los ciclos rojo y verde, en lugar de escribir la implementación primero y luego agregar una prueba que parece razonable.
### Solo se procesa un comportamiento en una ronda
Esto es fundamental.
No permita que Codex genere toda la matriz de prueba a la vez. Sería una "prueba de colocación horizontal":
```text
RED: test1, test2, test3, test4, test5
GREEN: 一次写一个大实现
```
Lo que quieres es cortarlo a lo largo:
```text
RED test1 -> GREEN impl1 -> REFACTOR
RED test2 -> GREEN impl2 -> REFACTOR
RED test3 -> GREEN impl3 -> REFACTOR
```
La primera ronda de implementación cambiará su comprensión del problema. No escribas todas tus pruebas de una sola vez.
## 2. L1: escribir disciplina en AGENTS.md
`AGENTS.md` es el archivo de descripción que Codex leerá al ingresar al proyecto.
La documentación oficial de OpenAI indica que Codex primero leerá la descripción global y luego la leerá desde el directorio raíz del proyecto hasta el directorio actual. Cada capa lee primero `AGENTS.override.md`; de lo contrario, lee `AGENTS.md`. Las descripciones más cercanas al directorio actual aparecen más tarde y, por lo tanto, tienen mayor prioridad. El límite de combinación predeterminado es `32 KiB`, por lo que no se puede escribir un tutorial largo aquí.
### Debería ser como las reglas de tráfico del proyecto.
`AGENTS.md` no es responsable de enseñarle a Codex qué es TDD. Sólo es responsable de escribir claramente: qué comportamiento no está permitido en este proyecto.
Puedes poner este párrafo directamente:
```markdown
# TDD Rules
- For new behavior and bug fixes, use red/green TDD.
- RED: write exactly one failing behavior test first.
- Run the smallest relevant test command and confirm the failure is expected.
- Do not edit production implementation during RED.
- GREEN: write the minimum production code required to pass the current failing test.
- Never modify, delete, skip, or weaken tests to make implementation pass.
- REFACTOR only after tests are green.
- Keep structural changes and behavior changes separate.
- Report Behavior, Test, Command, RED, GREEN, and REFACTOR for each cycle.
```
Comandos de proyecto adicionales:
```markdown
# Verification
- Use `pytest` or the smallest relevant pytest command for Python behavior tests.
- Use `npm run types:check` only when this blog site's MDX or TypeScript changes.
- Use the smallest targeted command during RED/GREEN loops.
- If a command is slow, explain what targeted command was used first and what full command remains.
```
### No debería escribirse como una enciclopedia.
Un `AGENTS.md` incorrecto estaría lleno de:
* Historia de TDD
* Todas las filosofías de prueba.
* Un montón de tutoriales de marco.
* Plantillas de mensajes complejas
* Especificaciones completas en diferentes idiomas.
Estas cosas diluyen las reglas que realmente importan.
Mi sugerencia es: `AGENTS.md` Poner sólo disciplinas residentes. Utilice habilidades para procesos largos.
## 3. L2: Convertir el proceso en Codex Skill
`AGENTS.md` resuelve "disciplina predeterminada" y la habilidad resuelve "proceso completo".
Cuando suele decir "hágalo por TDD" al Codex, debe actualizar esta oración a una habilidad a nivel de proyecto.
### Estructura del directorio
Ponlo aquí:
```text
.agents/
skills/
tdd-codex/
SKILL.md
```
Codex escaneará desde el directorio actual hasta `.agents/skills`. Las habilidades en el directorio raíz del almacén son adecuadas para los flujos de trabajo utilizados por el equipo.
### HABILIDAD mínima disponible.md
```markdown
---
name: tdd-codex
description: Implementing or fixing maintainable code with Codex using strict red-green-refactor TDD. Use for new behavior, bug reproduction, behavior tests, or safe AI coding.
---
# TDD Codex Workflow
Use one behavior slice per cycle.
## Phase 0: Scope
Identify one observable behavior.
Name the public API, user flow, or integration boundary under test.
Do not edit production code.
## Phase 1: RED
Write exactly one failing behavior test.
Prefer public behavior over implementation details.
Run the smallest relevant test command.
Confirm the failure is expected.
Stop and report:
- Behavior
- Test file
- Command
- Failure reason
## Phase 2: GREEN
Write the minimum production code to pass the current failing test.
Never modify, delete, skip, or weaken tests to pass.
Do not add speculative features or abstractions.
Run the same test command.
Report the passing result.
## Phase 3: REFACTOR
Only refactor after tests are green.
If the code is already simple, skip.
If refactoring, make one structural change at a time.
Run tests after each refactor.
Do not change behavior.
## Cycle Report
Return:
- Behavior:
- Test:
- Command:
- RED:
- GREEN:
- REFACTOR:
- Next slice:
```
### Método de llamada
A partir de ahora puedes decir:
```text
用 tdd-codex skill 做这个需求。
每轮只处理一个行为。
先 RED,确认失败后停下来,不要直接写实现。
```
O más corto:
```text
用 tdd-codex。先写红灯,等我说 go。
```
El punto no es lo hermoso que es el mensaje, sino que hace que Codex vuelva al mismo camino cada vez.
## 4. L3: Utilice subagentes para aislar la reconstrucción roja y verde
No todas las tareas requieren subagentes.
Sin embargo, cuando las tareas son complejas, las pruebas se contaminan fácilmente con la implementación y la refactorización es fácil de salirse de control, será más estable dividir ROJO, VERDE y REFACTOR en diferentes agentes.
### ¿Cuándo vale la pena desmantelar?
Adecuado para desmontaje:
* Permisos, facturación, máquina de estados.
-Funcionalidad de múltiples módulos
* El error está muy oculto, por lo que primero debes escribir una prueba de recurrencia.
* El modelo siempre se cambia para probar y ponerse verde.
* Quiere a alguien que solo sea responsable de revisar la calidad de las pruebas.
No apto para desmontaje:
* Funciones de gadget
* Cambios de redacción
* Puro ajuste visual
* guión único
El costo de derribar a un agente es real. Úselo sólo cuando los beneficios del aislamiento superen los costos de la comunicación.
### RED agent
```toml
# .codex/agents/tdd-test-writer.toml
name = "tdd_test_writer"
description = "RED phase agent. Writes one failing behavior test and stops before implementation."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the RED phase.
Write exactly one behavior-focused test for the requested slice.
Prefer public APIs and user-visible behavior over implementation details.
Run the smallest relevant test command.
Confirm the test fails for the expected reason.
Do not edit production implementation.
Do not add multiple tests at once.
Return Behavior, Test, Command, and RED failure reason.
"""
```
### GREEN agent
```toml
# .codex/agents/tdd-implementer.toml
name = "tdd_implementer"
description = "GREEN phase agent. Implements the minimum production code to pass the current failing test."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the GREEN phase.
Read the failing test and relevant production code.
Write the minimum implementation required to pass the current test.
Never modify, delete, skip, or weaken tests to make them pass.
Do not add speculative features, helpers, configuration, or abstractions.
Run the relevant tests and return the command plus passing output.
"""
```
### REFACTOR agent
```toml
# .codex/agents/tdd-refactorer.toml
name = "tdd_refactorer"
description = "REFACTOR phase agent. Improves structure only after tests are green."
sandbox_mode = "workspace-write"
developer_instructions = """
You own only the REFACTOR phase.
Start by running the relevant tests to confirm the code is green.
Look for duplication, unclear names, needless branching, or misplaced responsibility.
Skip refactoring when the code is already simple.
If you refactor, make one structural change at a time.
Run tests after each refactor.
Never change behavior in this phase.
"""
```
### Cómo ordenar la sesión principal
```text
按三阶段 TDD 做这个 slice:
1. tdd_test_writer 只写一个失败测试,并确认 RED。
2. 等我确认后,tdd_implementer 写最小实现,并确认 GREEN。
3. tdd_refactorer 判断是否需要结构重构。
不要批量铺测试。
不要在 GREEN 阶段修改测试。
```
El punto aquí es aislar el contexto. Las personas que escriben pruebas deberían intentar no verse afectadas por los detalles de implementación; las personas que escriben implementaciones no pueden realizar pruebas manualmente; las personas que refactorizan no pueden introducir nuevos comportamientos.
## 5. L4: Utilice ganchos para centrarse en la diferencia de prueba
Si confía únicamente en las reglas, es posible que el modelo aún se salga de los límites.
El transfronterizo más común es: la prueba es roja y el modelo cambia la prueba para volverla verde.
El valor de los ganchos no es ser "absolutamente seguro", sino exponer esta acción de inmediato.
### Habilitar ganchos del Codex
Primero abra la bandera de función en la configuración:
```toml
# ~/.codex/config.toml 或 /.codex/config.toml
[features]
codex_hooks = true
```
Codex busca enlaces junto a la capa de configuración activa. Ubicaciones comunes:
* `~/.codex/hooks.json`
* `~/.codex/config.toml`
* `/.codex/hooks.json`
* `/.codex/config.toml`
A nivel de proyecto, se recomienda utilizar `/.codex/hooks.json` primero, porque puede seguir el almacén.
### hooks.json
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/watch-test-edits.sh\"",
"timeout": 10,
"statusMessage": "Checking test file edits"
}
]
},
{
"matcher": "Bash|apply_patch|Edit|Write",
"hooks": [
{
"type": "command",
"command": "bash \"$(git rev-parse --show-toplevel)/.codex/hooks/run-fast-check.sh\"",
"timeout": 120,
"statusMessage": "Running fast checks"
}
]
}
]
}
}
```
### Compruebe si el archivo de prueba ha sido modificado
```bash
# .codex/hooks/watch-test-edits.sh
#!/usr/bin/env bash
set -euo pipefail
changed_tests="$(
git diff --name-only |
grep -E '(^|/)(__tests__|tests?)/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.go$|test_.*\.py$' || true
)"
if [ -n "$changed_tests" ]; then
cat < "hello-world"
- trim leading/trailing spaces
- collapse repeated spaces into one hyphen
- remove punctuation
- normalize "Café" -> "cafe"
- empty input returns empty string
```
### Paso 1: Primera luz roja
```text
读取 SPEC.md。
只实现第一条行为:"Hello World" -> "hello-world"。
先 RED:只写一个失败测试,运行它,确认失败。
不要写生产实现。
```
Salida ideal:
```text
Behavior: basic title becomes lowercase hyphenated slug
Test: tests/test_slugify.py
Command: pytest tests/test_slugify.py -q
RED: failed because slugify is not defined
```
Sólo entonces podrás continuar.
### Paso 2: Verde mínimo
```text
go
```
Codex escribe la implementación mínima:
```py
def slugify(text: str) -> str:
return text.lower().replace(" ", "-")
```
Luego informe:
```text
GREEN: pytest tests/test_slugify.py -q passed
REFACTOR: skipped, implementation is still simple
Next slice: trim leading/trailing spaces
```
### Paso 3: Segunda luz roja
```text
继续下一条:去掉首尾空格。
先 RED,只写一个测试。
```
Prueba:
```py
def test_slugify_trims_spaces():
assert slugify(" Hello World ") == "hello-world"
```
Si la implementación actual genera `-hello-world-`, la luz roja es verdadera.
Luego VERDE:
```py
def slugify(text: str) -> str:
return text.strip().lower().replace(" ", "-")
```
### Paso 4: No te apresures a la abstracción
En este punto, muchas IA querrán dibujar un `normalizeInput`, `removePunctuation`, `toAscii`.
No te apresures todavía.
El diseño de TDD debe ser impulsado por la presión de prueba, no por la imaginación. Espere hasta que agregue Unicode, puntuación y cadenas vacías, y realmente surja presión estructural, luego refactorice.
## 8. Comprobación rápida del uso diario
### Nuevas funciones
```text
用 TDD 实现这个需求。
每轮只处理一个行为。
先写一个失败测试并运行确认 RED。
不要写生产实现,直到我说 go。
```
### Corregir error
```text
先写一个能复现这个 bug 的失败测试。
确认它因为这个 bug 失败后,再写最小修复。
不要改测试来适配当前实现。
```
### Funciones complejas
```text
先不要写代码。
请给出 TDD 分解计划:
- 外圈集成测试是什么
- 内圈每个行为 slice 是什么
- 每轮用什么命令验证
- 哪些地方不能 mock
等我确认后再开始 RED。
```
### Review
```text
Review 这次改动,重点看:
- 是否先有失败测试
- 测试是否测行为而不是实现
- 是否存在为了通过而弱化测试
- 结构改动和行为改动是否混在一起
- 是否缺少外圈集成测试
```
## 9. Lista de verificación final
Cada vez que le pida al Codex que haga TDD, utilice esta tabla para comprobarlo al final.
| Pregunta | Criterios de elegibilidad |
| --------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| ¿Es realmente popular primero? ¿Hay comandos fallidos y motivos del fracaso? | |
| ¿Es correcto el color rojo? El motivo del fracaso corresponde a la falta de comportamiento objetivo | |
| Haz solo un comportamiento por ronda | Sin pruebas por lotes |
| VERDE ¿Se ha cambiado la prueba? | No se ha modificado ninguna prueba para ser verde |
| ¿Las pruebas prueban el comportamiento? No depende de detalles de implementación interna | |
| ¿Se han refactorizado los comportamientos mixtos? Separar cambios estructurales y cambios de comportamiento | |
| ¿Existe una verificación completa? Se han realizado pruebas específicas y las inspecciones completas necesarias | |
Si esta tabla no puede pasar, no se apresure a fusionarse.
## Recursos recomendados
# Quand l'IA permet de reussir ses examens sans etudier, que reste-t-il de l'universite ?
Anthropic a recemment publie une interview reunissant quatre etudiants issus de Princeton, Berkeley, la London School of Economics et l'Arizona State University, pour discuter de la realite de l'IA sur les campus. L'entretien a dure pres de 40 minutes, sans discours marketing — rien que de veritables interrogations, inquietudes et reflexions.
Cette interview revele un enjeu bien plus profond : **l'IA ne se contente pas de transformer les methodes d'apprentissage, elle demonte la logique fondamentale de tout le systeme educatif**.
***
## Une realite incontournable
Au debut de l'interview, l'animateur pose une question directe : quelle est l'ambiance autour de l'IA sur les campus ?
La reponse : **90 % des etudiants utilisent l'IA**. Pas de maniere occasionnelle, mais comme partie integrante de leur quotidien — resumer des notes de cours, repondre a des exercices, obtenir des retours sur leurs devoirs, analyser des etudes de cas, mener des etudes de marche, realiser des recherches financieres. Certains etudiants l'utilisent meme pour passer des quiz, avec une raison tres pragmatique : quand vous etes en master et que vous cumulez plusieurs emplois, vous n'avez pas toujours le temps.
Mais le plus frappant, c'est que si presque tout le monde l'utilise, personne ne connait les regles. Certains cours interdisent formellement l'IA, d'autres l'encouragent activement, et la majorite se situe dans une zone floue. Les etudiants ne savent pas ou se trouve la limite, les professeurs ne savent pas comment gerer la situation. C'est ce qu'on appelle la « zone grise » — on veut utiliser l'IA, mais on craint d'enfreindre les regles ; on ne l'utilise pas, mais on a l'impression de prendre du retard.
Le plus dangereux dans cette zone grise n'est pas le risque de violation des regles par les etudiants, mais le fait qu'**elle empeche les discussions vraiment constructives d'avoir lieu**. Les etudiants ne peuvent pas partager ouvertement les bonnes pratiques d'utilisation de l'IA, les professeurs ne peuvent pas guider les etudiants vers une utilisation responsable des outils, et toute la communaute academique se retrouve dans un etat embarrassant d'interdiction officielle et d'utilisation clandestine.
Et quand les regles ne peuvent pas etre appliquees efficacement, elles deviennent un filtre qui distingue « ceux qui savent faire semblant » de « ceux qui ne savent pas ». Une interdiction officielle doublée d'un usage generalise en coulisses **n'empechera pas les etudiants d'utiliser l'IA, mais les empechera de discuter ouvertement de la meilleure facon de l'utiliser**.
***
## L'IA est un miroir
L'un des points de vue les plus percutants de l'interview : l'intelligence artificielle, et en particulier la facon dont les etudiants l'utilisent, est tres revelatrice de leurs motivations.
L'idee sous-jacente : **l'IA est devenue un miroir qui reflete la veritable raison pour laquelle vous etes a l'universite**.
L'interview classe les objectifs universitaires en trois categories : premierement, approfondir les connaissances dans sa specialite et maitriser la comprehension d'un domaine ; deuxiemement, se preparer a la vie professionnelle, trouver un bon emploi et construire son reseau ; troisiemement, elargir son cercle social, profiter de la vie etudiante et vivre l'experience universitaire. Chaque etudiant accorde un poids different a ces trois objectifs, et la facon dont il utilise l'IA revele precisement ces priorites.
Si vous ne vous souciez que de « reussir les examens » et « obtenir le diplome », vous soumettez directement les reponses generees par l'IA comme devoirs. Ce n'est pas un jugement moral, c'est la realite — quand la technologie vous permet d'atteindre votre objectif au moindre cout, pourquoi feriez-vous un detour ? Si votre but est simplement d'obtenir un diplome pour trouver un emploi, utiliser l'IA pour faire vos devoirs est un choix parfaitement rationnel.
Mais si vous voulez vraiment apprendre, si vous voulez comprendre un domaine en profondeur, vous utilisez l'IA comme partenaire de dialogue. Vous lui posez des questions, vous lui demandez d'expliquer des concepts, puis vous reformulez dans vos propres mots. Vous lui faites ecrire une premiere version du code, puis vous le refactorisez et l'optimisez vous-meme. Vous vous assurez qu'a chaque etape, vous comprenez vraiment ce qui se passe.
Cette polarisation n'existe pas seulement entre differents etudiants, elle existe aussi entre differentes filieres. Les etudiants en sciences humaines tendent a se detourner de l'IA, car leur apprentissage exige une lecture approfondie — lire attentivement le texte original, savourer les subtilites du langage, comprendre l'intention de l'auteur. L'IA perturbe ce processus, car elle fournit des resumes et des reformulations, et non l'experience directe du texte original. Les etudiants en ingenierie et en commerce, en revanche, utilisent massivement l'IA, car elle abaisse les barrieres techniques et permet a des personnes sans formation informatique d'ecrire du code, de creer des sites web et d'analyser des donnees.
**Cette polarisation est fondamentalement une divergence de conceptions sur ce qui merite d'etre appris**. Pour les etudiants en lettres, l'experience de lire Shakespeare dans le texte original est en soi un apprentissage ; pour les etudiants en ingenierie, ce qui compte c'est de pouvoir resoudre un probleme, pas d'ecrire chaque ligne de code soi-meme. L'IA rend cette difference encore plus flagrante.
***
## Outil ou bequille ? Un critere de jugement simple
L'interview pose une question essentielle : comment distinguer si l'IA est un outil ou une bequille ?
Les etudiants donnent une reponse etonnamment unanime : **la capacite a expliquer**.
Si vous ne pouvez pas expliquer ce que vous avez cree, si vous ne pouvez pas decrire le role que l'IA y a joue, c'est une bequille. Si vous pouvez l'expliquer comme a un eleve de CM2, si vous etes capable de fournir des explications tant basiques qu'avancees, c'est un outil.
Ce critere parait simple, mais il touche a l'essence meme de l'apprentissage. Le principe fondamental de la methode Feynman est le suivant : si vous ne pouvez pas expliquer un concept en termes simples, c'est que vous ne l'avez pas vraiment compris. L'apprentissage a l'ere de l'IA suit la meme logique — si vous ne pouvez pas expliquer ce que l'IA a fait pour vous, vous ne faites qu'« externaliser votre reflexion », et non l'« augmenter ».
L'interview mentionne un exemple interessant : un etudiant a developpe un outil dans lequel on insere les diapositives du cours, et l'IA genere des annotations de type professoral a cote de chaque diapositive. Cet etudiant explique : « Ca fonctionne bien parce que je l'ai deja oriente pour qu'il sache ce que je veux comprendre — les definitions de certains elements sur les diapositives. Les diapositives sont parfois tres abstraites, manquent de contexte, et ont besoin d'annotations complementaires. »
Le point cle est « je l'ai deja oriente pour qu'il sache ce que je veux comprendre ». Cet etudiant sait exactement ou se situent ses lacunes, quel type d'aide il a besoin, puis il guide activement l'IA pour qu'elle fournisse cette aide. C'est un outil. S'il avait simplement envoye les diapositives a l'IA en demandant « resume-moi ce cours », puis appris le resultat par coeur, ca aurait ete une bequille.
**La difference reside dans la proactivite et la comprehension**. Celui qui utilise un outil sait ce qu'il fait et controle l'ensemble du processus. Celui qui utilise une bequille a cede le controle a la technologie et est devenu un recepteur passif.
***
## Le retard des etablissements n'est pas un manque de reactivite, c'est une incapacite structurelle
L'interview mentionne quelques initiatives d'etablissements. La London School of Economics a mis en place un cours obligatoire qui enseigne aux etudiants comment utiliser Claude — dialoguer avec lui, lui attribuer differents roles, puis demander aux etudiants de soumettre les historiques de conversation pour voir comment ils interagissent avec l'IA. Le centre de gestion de carriere de l'Arizona State University a cree une bibliotheque de prompts, fournissant des modeles pour differents scenarios. Ce sont de bonnes initiatives dont le principe central est : non pas interdire l'IA, mais apprendre aux etudiants a l'utiliser de maniere responsable.
Mais ce ne sont que des cas isoles. La plupart des etablissements debattent encore de la question « faut-il autoriser les etudiants a utiliser l'IA ? ». Certains professeurs disent oui, a condition d'indiquer comment dans les devoirs ; certains cours l'interdisent directement ; d'autres n'en parlent tout simplement pas, partant du principe que les etudiants ne l'utiliseront pas. Pas de cadre integre, pas de norme unifiee, l'ensemble du systeme est dans un etat de confusion.
Le probleme plus profond est le suivant : **cette question ne peut fondamentalement pas etre resolue par des reglements**.
La logique de supervision de l'education traditionnelle repose sur : l'etablissement fixe les regles, les etudiants les respectent, les contrevenants sont sanctionnes. Cette logique ne fonctionne que si « les infractions peuvent etre detectees ». Or l'IA brise cette premisse.
Vous pouvez interdire aux etudiants d'utiliser l'IA pour rendre leurs devoirs, mais vous ne pouvez pas surveiller s'ils l'ont utilisee dans leur processus de reflexion. Vous pouvez recourir a des outils de detection de l'IA, mais leur taux de precision est loin d'etre suffisant pour servir de base a des sanctions — le taux de faux positifs est trop eleve, et les etudiants apprendront rapidement a contourner la detection. Plus fondamentalement, **il est impossible de distinguer un devoir de qualite realise avec l'aide de l'IA d'un devoir de qualite realise de maniere autonome**, car une bonne utilisation de l'IA est censee etre transparente.
L'interview contient cette declaration tres directe : fondamentalement, aucune reglementation ne changera la facon dont les etudiants utilisent l'IA. La responsabilite est entre les mains des etudiants. Ce n'est pas un desengagement, c'est la realite.
Quand la technologie rend possible le fait de « reussir ses examens sans etudier », les etablissements ne font pas face a un probleme de gestion, mais a un probleme existentiel : **si les examens ne peuvent plus prouver l'apprentissage, quel est le sens de l'existence de l'universite ?**
***
## La logique fondamentale du systeme educatif est brisee
Cette question touche a la contradiction fondamentale du systeme educatif.
L'education traditionnelle repose sur plusieurs postulats centraux : premierement, le savoir est rare et necessite des institutions specialisees (les ecoles) et des professionnels (les professeurs) pour etre transmis ; deuxiemement, les resultats de l'apprentissage peuvent etre mesures par des examens ; troisiemement, un diplome prouve que vous maitrisez les connaissances d'un domaine et que vous etes donc qualifie pour exercer dans ce secteur.
L'IA brise ces postulats un par un.
**Le savoir n'est plus rare**. YouTube propose gratuitement des cours de Stanford, Claude peut repondre a vos questions a tout moment, GitHub regorge de projets open source a etudier. Vous n'avez pas besoin d'aller a l'ecole pour acceder au savoir, vous n'avez meme pas besoin d'un abonnement payant pour beneficier d'un tutorat IA de base.
**Les examens ne mesurent plus l'apprentissage**. Quand l'IA peut repondre a la plupart des questions d'examen, l'examen passe d'un « outil mesurant le niveau de comprehension » a un « outil mesurant la capacite a utiliser l'IA ». Cela ne signifie pas que les examens sont totalement inutiles, mais qu'ils ne permettent plus de distinguer avec precision « ceux qui ont vraiment appris » de « ceux qui savent exploiter les outils ».
**La valeur du diplome diminue**. Quand les employeurs realisent qu'un diplome ne garantit pas que le candidat maitrise reellement les connaissances concernees, ils accordent plus d'importance aux preuves de competences concretes — portfolio, experience de projets, performances en stage. Le diplome passe du statut de « preuve de competence » a celui de « seuil minimal d'entree ».
**Quand ces postulats s'effondrent, la proposition de valeur du systeme educatif doit etre redefinie.**
L'interview propose une reponse : la valeur de l'universite passe de la « **transmission du savoir** » a la « **fourniture d'un environnement** ». Un environnement ou l'on peut se tromper, explorer, confronter ses idees avec celles des autres. Vous pouvez passer un week-end avec vos colocataires a realiser une « bucket list avant la remise des diplomes », tester des idees « peut-etre stupides » lors d'un hackathon, debattre avec vos professeurs, discuter avec vos camarades, apprendre de vos echecs sans risquer votre carriere.
Les projets etudiants mentionnes dans l'interview illustrent bien ce point — un « rappel automatique d'inscription aux cours », un « detecteur de salles vides », un « classement des souhaits de fin d'etudes ». Ces projets ne sont pas techniquement complexes, et beaucoup de leurs createurs n'ont meme pas de formation en informatique. Mais ils naissent d'emotions humaines authentiques : la peur de rater quelque chose, la quete de commodite, l'attachement a la vie universitaire.
**Quand les barrieres techniques s'abaissent, l'important n'est plus « savez-vous coder ? », mais « quel probleme voulez-vous resoudre ? »** L'universite offre un espace pour explorer librement ces questions, un environnement ou l'on peut transformer ses idees en realite, puis apprendre de ses echecs.
L'IA peut faire vos devoirs a votre place, mais elle ne peut pas vivre a votre place cette periode de « liberte de se tromper et d'explorer ».
***
## Le paradoxe du marche de l'emploi
La seconde moitie de l'interview aborde l'emploi et revele un autre paradoxe.
Les etudiants utilisent l'IA pour rediger leur CV, les entreprises utilisent l'IA pour trier les CV. Tout le cycle de recrutement se transforme en monologue face a un ecran — d'abord face a l'IA pour rediger la lettre de motivation, puis face a une camera pour repondre aux questions, et enfin reception d'une lettre de refus generee par l'IA. De la soumission du CV a la reception du refus, il ne faut parfois que 15 minutes. C'est tres efficace, mais tres peu humain.
Cela cree un marche de l'emploi « IA contre IA ». Les etudiants entrainent l'IA a rediger de « bons » CV, les entreprises entrainent l'IA a selectionner de « bons » candidats. Le role des etres humains dans ce processus diminue de plus en plus. Parler a un ecran ne genere aucune alchimie et ne permet pas de reveler ces qualites difficiles a quantifier mais pourtant essentielles — le sens de l'humour, la capacite d'adaptation, les subtilites du travail en equipe.
Mais l'autre face du paradoxe : la maitrise de l'IA est elle-meme devenue un nouvel avantage competitif. Les quatre grands cabinets de conseil, qui recrutaient autrefois des MBA generalistes, recherchent desormais specifiquement des MBA possedant des competences en IA. Si vous savez appliquer l'IA a differents secteurs, vous devenez leur candidat de premier choix.
**Le paradoxe est le suivant : l'IA rend le marche de l'emploi plus froid, mais le marche de l'emploi valorise davantage les competences en IA. Vous ne pouvez pas y echapper, vous ne pouvez qu'apprendre a l'utiliser efficacement**.
Cela nous ramene a la question centrale : qu'est-ce qu'une « utilisation efficace » ? Ce n'est pas savoir utiliser ChatGPT pour rediger des e-mails, c'est etre capable d'identifier quels problemes se pretent a une resolution par l'IA, de concevoir des prompts pour guider l'IA vers les resultats souhaites, et d'evaluer la qualite des resultats de l'IA pour apporter les corrections necessaires.
Cette competence ne se developpe pas en interdisant l'IA, mais par une pratique intensive et des essais-erreurs. C'est aussi pourquoi les etablissements qui embrassent activement l'IA, qui creent des Claude Builder Clubs, qui organisent des hackathons, offrent a leurs etudiants une education plus precieuse — ils permettent aux etudiants d'apprendre a collaborer avec l'IA dans un environnement relativement securise.
***
## Le transfert de responsabilite
L'insight le plus essentiel de cette interview est peut-etre celui-ci : quand la technologie permet de « reussir ses examens sans etudier », le sens meme de l'apprentissage devient une question a laquelle chacun doit repondre par lui-meme.
C'est un transfert de responsabilite. **De l'etablissement vers l'etudiant, des regles vers la conscience individuelle, de la motivation externe vers la motivation interne.**
La motivation de l'education traditionnelle est externe : il faut reussir les examens pour obtenir le diplome, il faut le diplome pour trouver un emploi. Ce systeme d'incitation externe pousse les etudiants a apprendre. Mais quand l'IA permet de reussir les examens sans apprendre, ce systeme d'incitation s'effondre.
Il ne reste plus que la motivation interne : voulez-vous vraiment apprendre ? Ce domaine vous interesse-t-il vraiment ? Voulez-vous vraiment comprendre en profondeur, ou simplement obtenir un diplome ?
L'interview contient un detail tres revelateur. Un etudiant evoque l'aspect negatif des etudes de master — on cumule plusieurs emplois, on manque de temps, alors parfois on utilise l'IA pour terminer rapidement un quiz. Mais il ajoute ensuite : les etudes de master sont censees etre la periode ou l'on developpe sa pensee critique, ou l'on montre un cote plus affirme de soi-meme. **Il est conscient de la contradiction, mais il choisit l'efficacite**.
Ce n'est pas un jugement moral. Face aux pressions du reel, l'efficacite l'emporte souvent sur l'ideal. Mais ce choix revele une verite : **quand la pression externe (terminer un quiz) entre en conflit avec la motivation interne (apprendre en profondeur), beaucoup choisissent la premiere option**.
L'IA rend ce conflit encore plus aigu, car elle rend le fait de « se contenter de passer l'examen » extremement facile. A l'epoque pre-IA, meme si vous vouliez seulement boucler un examen, vous deviez tout de meme apprendre un minimum pour le reussir. L'IA supprime cette etape intermediaire — vous pouvez reussir sans rien apprendre du tout.
**Cela force chacun a se confronter a cette question : pourquoi etes-vous reellement a l'universite ?**
Si la reponse est « pour obtenir un diplome et trouver un emploi », alors utiliser l'IA pour faire ses devoirs est parfaitement logique. Si la reponse est « je veux vraiment apprendre ce domaine », alors vous devez activement resister a la tentation des raccourcis qu'offre l'IA.
L'etablissement ne peut pas faire ce choix a votre place. Les reglements ne peuvent pas vous imposer une motivation interne. **C'est votre responsabilite**.
***
## La technologie n'attendra pas que vous soyez pret
L'interview se conclut sur une attitude qui traverse tout l'entretien : « On trouvera une solution. »
Les regles de l'ecole ne suivent pas le rythme ? On commence par utiliser l'IA, puis on dit a l'ecole ce qui fonctionne. L'IA peut servir a tricher ? On apprend progressivement a l'utiliser de maniere responsable. Le marche de l'emploi a change ? On s'adapte aux nouvelles regles du jeu.
Ce n'est pas de l'optimisme aveugle, c'est du realisme. La technologie est deja la, elle n'attendra pas que vous soyez pret pour commencer a transformer le monde. Vous pouvez choisir de resister, vous pouvez choisir de vous adapter, mais vous ne pouvez pas choisir d'arreter le temps.
**La relation de cette generation d'etudiants avec l'IA n'est ni de la peur, ni une adoption aveugle, mais une exploration tatonnante dans le chaos, un apprentissage par l'experimentation.**
Les projets qu'ils realisent au Claude Builder Club — pas techniquement complexes, mais qui resolvent de vrais problemes. Les idees qu'ils testent lors des hackathons — peut-etre absurdes, mais au moins ils essaient. Leur perplexite en cours — les regles ne sont pas claires, mais au moins ils reflechissent.
Cette posture d'**apprentissage par la pratique** est probablement plus efficace que n'importe quelle reglementation pour les aider a s'adapter a l'ere de l'IA.
Kevin Kelly, dans *Ce que veut la technologie*, propose le concept de « technium » : la technologie, prise dans son ensemble, semble avoir sa propre volonte et vouloir devenir toujours plus puissante. L'interview fait une observation qui rejoint cette idee : au cours des deux dernieres annees, quoi que l'IA ait besoin pour continuer a se developper, elle l'a obtenu. Le revirement sur l'energie nucleaire, les discussions sur les centres de donnees spatiaux — chaque fois qu'un goulot d'etranglement menace de ralentir l'IA, l'obstacle est leve.
Mais une formulation plus juste serait peut-etre : **les etudiants creent ce dont ils ont besoin**.
L'IA n'est qu'un outil. Ce qui determinera l'avenir, c'est la facon dont cette generation d'etudiants choisira d'utiliser cet outil — pour fuir la reflexion, ou pour l'augmenter. Pour boucler des examens, ou pour explorer le monde. Pour en faire une bequille, ou un outil.
Ce choix, l'etablissement ne peut pas le controler ; c'est a eux seuls d'en decider.
Et a en juger par cette interview, au moins une partie de ces etudiants reflechit serieusement a cette question. C'est peut-etre suffisant.
# Penser comme un Agent : la philosophie de conception d'outils de l'équipe Claude Code
Thariq est ingénieur chez Anthropic et l'un des principaux contributeurs de Claude Code. Fin février, il a publié un long article sur X, partageant cinq cas concrets sur la conception d'outils pour agent issus de la construction de Claude Code — non pas un cadre théorique, mais des retours d'expérience de terrain. Cet article a récolté 3,51 millions de vues, 209 réponses et 9 691 likes, suscitant de nombreuses discussions de qualité.
Pour introduire la question centrale de l'article, Thariq utilise une excellente analogie : imaginez que vous êtes face à un problème mathématique difficile, de quels outils avez-vous besoin ?
**Un stylo et du papier** représentent le minimum — mais vous êtes limité au calcul manuel. **Une calculatrice** est mieux — mais vous devez savoir utiliser les fonctions avancées. **Un ordinateur** est le plus puissant — mais il faut savoir coder.
Le choix de l'outil dépend des capacités de l'utilisateur. Donner un ordinateur à quelqu'un qui ne sait pas programmer est moins utile que de lui donner une calculatrice. Donner une calculatrice à un programmeur limite au contraire son potentiel.
C'est la même chose pour un agent. La question n'est pas « quel outil est le plus puissant », mais « quel outil correspond le mieux aux capacités actuelles du modèle ». Cet article partage les erreurs commises par l'équipe Claude Code dans sa quête de ce point d'adéquation.
Voici quatre thèmes progressifs que j'ai extraits de l'article original et des discussions de la communauté.
***
## Moins, c'est plus — le paradoxe du nombre d'outils
L'intuition nous dit que plus un agent a d'outils, plus il est performant. Mais l'expérience de l'équipe Claude Code prouve exactement le contraire.
Claude Code ne dispose actuellement que d'environ 20 outils, et l'équipe examine constamment si tous sont réellement nécessaires. La barre pour ajouter un nouvel outil est élevée, car cela représente une option supplémentaire que le modèle doit prendre en compte — chaque outil ajouté augmente la « charge cognitive » du modèle lors de ses prises de décision.
Cette découverte a fortement résonné dans la communauté. leon.M a partagé son expérience pratique :
La recherche CodeAct d'Apple apporte un soutien quantitatif à ce point de vue : **une seule primitive d'exécution de code (code execution primitive) surpasse les ensembles d'outils spécialisés et complexes de jusqu'à 20 % sur les tâches difficiles**. Moins peut effectivement signifier plus.
Les trois itérations de l'outil AskUserQuestion constituent le meilleur exemple de ce principe. L'équipe Claude Code souhaitait améliorer la capacité de Claude à poser des questions aux utilisateurs — bien que Claude puisse poser des questions en texte brut, y répondre semblait trop fastidieux. Comment réduire cette friction ?
**Première tentative** : ajouter un paramètre à ExitPlanTool pour qu'il présente un ensemble de questions en même temps que le plan. Résultat — Claude s'est embrouillé. Lui demander simultanément de produire un plan et des questions sur ce plan posait problème : que faire si la réponse de l'utilisateur entre en conflit avec le plan ?
**Deuxième tentative** : modifier les instructions de sortie pour que Claude pose des questions dans un format markdown spécifique, puis parser et formater côté frontend. Résultat — instable. Claude ajoutait des phrases supplémentaires, omettait des options, ou utilisait un format complètement différent.
**Troisième tentative** : créer un outil AskUserQuestion indépendant. Claude peut l'appeler à tout moment ; l'appel déclenche une fenêtre contextuelle affichant la question et bloque la boucle de l'agent jusqu'à ce que l'utilisateur réponde. **Cela a fonctionné.**
Thariq a écrit une phrase particulièrement intéressante dans l'article original :
> Claude semble prendre plaisir à appeler cet outil, et nous avons constaté que la qualité de ses sorties était excellente. Même un outil parfaitement conçu ne fonctionne pas si Claude ne comprend pas comment l'appeler.
phuong a posé une question perspicace : « "Claude semble prendre plaisir à appeler cet outil" est la phrase la plus fascinante et la moins expliquée ici. Comment détectez-vous l'"affinité" du modèle pour un outil — en lisant les transcripts ou via des métriques internes de fréquence d'appel ? Si cette heuristique pouvait être plus précise, elle changerait la façon dont tout le monde conçoit les outils d'agent. »
Cette question est restée sans réponse, mais elle pointe vers une intuition importante : **le critère de succès de la conception d'outils n'est pas "l'humain trouve ça logique", mais "le modèle comprend comment l'utiliser et a envie de l'utiliser"**.
Emeka a ajouté une leçon très directe du point de vue du développement d'outils en entreprise : « J'ai essayé de contrôler chaque entrée et sortie possible en construisant des outils pour un agent, et le modèle a tout simplement... contourné le système. Économisez vos efforts d'ingénierie, faites confiance au modèle pour gérer l'ambiguïté. »
***
## Les outils deviennent obsolètes — l'effet de carcan après une montée en capacité
Si « moins c'est plus » concerne la dimension spatiale, « les outils deviennent obsolètes » est une leçon dans la dimension temporelle.
**Un outil qui aidait le modèle peut, à mesure que le modèle progresse, devenir une contrainte.**
Lors du lancement initial de Claude Code, l'équipe a réalisé que le modèle avait besoin d'une liste de tâches pour maintenir le cap — il pouvait y inscrire des éléments au démarrage, puis les cocher au fur et à mesure. Pour cela, ils ont fourni à Claude l'outil TodoWrite. Mais malgré cela, Claude oubliait fréquemment ce qu'il devait faire.
L'équipe a répondu en insérant un rappel système toutes les 5 itérations de conversation pour rappeler à Claude ses objectifs.
Mais avec les progrès du modèle, le problème s'est inversé : non seulement le modèle n'avait plus besoin qu'on lui rappelle sa liste de tâches, mais il percevait ces rappels comme une contrainte. **Les rappels répétés de la liste de tâches donnaient à Claude l'impression qu'il devait suivre strictement la liste, au lieu de s'adapter de manière flexible selon les besoins.** Parallèlement, Opus 4.5 s'était considérablement amélioré dans l'utilisation de sous-agents, mais comment coordonner une liste de tâches partagée entre sous-agents ?
L'équipe a donc remplacé TodoWrite par le Task Tool. La différence entre les deux est fondamentale : les Todos servaient à maintenir le modèle sur la bonne voie — comme un patron surveillant la liste de tâches d'un employé ; les Tasks sont davantage axés sur la communication entre agents — comme un tableau de collaboration d'équipe. Les Tasks supportent les dépendances, le partage de mises à jour entre sous-agents, et le modèle peut les modifier et les supprimer.
**De TodoWrite à un rappel toutes les 5 itérations, puis au Task Tool — ce sont trois reconceptions.** Non pas parce que les conceptions précédentes étaient « mauvaises », mais parce que le modèle avait évolué.
Le commentaire de modi résume parfaitement ce phénomène :
Cela soulève également un débat intéressant. Danny Cosson estime que AskUserQuestion est « en réalité une mauvaise conception » — il force le modèle de texte en entrée / texte en sortie très élégant de Claude Code dans un mode d'interaction spécifique, avec très peu de bénéfices.
Mais la réponse de @Toong est convaincante : la valeur d'AskUserQuestion réside en deux points — **l'initiative** (indiquer explicitement au modèle qu'il a le droit de poser des questions) et **la gestion d'état** (distinguer clairement la sortie texte standard de l'état de blocage « en attente d'intervention de l'utilisateur »).
Un même outil, deux évaluations radicalement différentes. Cela illustre parfaitement ce que Thariq dit en conclusion : **la conception d'outils est un art, pas une science**. Elle dépend du modèle que vous utilisez, des objectifs de l'agent et de l'environnement dans lequel il opère.
***
## Laisser l'Agent trouver ses propres réponses — du « gavage » à la « recherche autonome »
C'est la partie de l'article que je considère comme la plus utile en pratique.
Claude Code utilisait initialement une base de données vectorielle RAG pour fournir du contexte à Claude. RAG est puissant et rapide, mais pose deux problèmes : d'une part, il nécessite une indexation et une configuration qui peuvent être fragiles selon les environnements ; d'autre part, et c'est plus fondamental — **cette approche fournit le contexte à Claude au lieu de le laisser le trouver lui-même**.
L'équipe a effectué un changement clé : puisque Claude sait chercher sur le web, pourquoi ne pourrait-il pas chercher dans votre base de code ? En fournissant à Claude l'outil Grep, ils l'ont laissé chercher les fichiers et construire son propre contexte.
Le commentaire de Beacon va droit au but :
**En un an, Claude est passé d'une quasi-incapacité à construire son propre contexte à la capacité d'effectuer des recherches imbriquées sur plusieurs couches de fichiers pour trouver précisément le contexte nécessaire.** La clé de cette évolution n'a pas été de donner plus d'informations à Claude, mais de lui donner de meilleures capacités de recherche.
Lorsque Claude Code a introduit les Agent Skills, l'équipe a formellement proposé le concept de **divulgation progressive (Progressive Disclosure)** : permettre à l'agent de découvrir progressivement le contexte pertinent par l'exploration.
L'implémentation concrète est élégante : Claude peut lire des fichiers skill, et ces fichiers peuvent référencer d'autres fichiers que le modèle peut lire récursivement. Un usage courant des skills est de donner à Claude des capacités de recherche supplémentaires — par exemple des instructions sur l'utilisation d'une API ou l'interrogation d'une base de données.
Brian Wagner a partagé sa pratique en trois couches, parfaitement alignée avec l'approche de Claude Code :
> SKILL.md reste concis (environ 100 lignes), le contexte lourd est placé dans des fichiers que Claude découvre quand il en a besoin. J'appelle ça la troisième couche. Vous l'appelez divulgation progressive. C'est la même chose.
PrimeLine a même construit un système quantifié à trois niveaux : configuration JSON (environ 500 tokens) > résumé du schéma (environ 300 tokens) > markdown complet (environ 3K tokens). Un routeur de contexte décide quelle couche charger en fonction des mots-clés de la tâche.
La logique centrale de cette stratégie en couches est la suivante : **le contexte est une ressource limitée à rendements marginaux décroissants**. Déverser toutes les informations d'un coup dans l'agent gaspille non seulement des tokens, mais dilue aussi les informations vraiment importantes. Fournir l'information à la demande et laisser l'agent décider quand il a besoin de détails plus approfondis — voilà l'approche scalable.
Le sous-agent Claude Code Guide est une autre application ingénieuse de la divulgation progressive. L'équipe a remarqué que Claude ne connaissait pas suffisamment le fonctionnement de Claude Code lui-même — si vous lui demandiez comment ajouter un MCP ou ce que fait une commande slash, il ne pouvait pas répondre.
Ils auraient pu tout intégrer dans le prompt système, mais les utilisateurs posent rarement ce type de questions, et cela aurait augmenté l'érosion du contexte, interférant avec le travail principal de Claude Code : écrire du code.
Ils ont d'abord essayé de donner à Claude un lien vers la documentation pour qu'il cherche lui-même — ça fonctionnait, mais Claude chargeait un grand volume de résultats dans le contexte pour trouver la bonne réponse. La solution finale a été de construire un sous-agent dédié : Claude Code Guide. Ce sous-agent dispose d'instructions de recherche détaillées et sait comment chercher efficacement dans la documentation et quoi renvoyer. **Sans ajouter le moindre nouvel outil, l'espace d'action de Claude a été élargi.**
Lance Martin, dans son article sur les modèles de conception d'agents, propose une perspective complémentaire : **plutôt que de définir des dizaines d'outils pour l'agent, donnez-lui un ordinateur et laissez-le orchestrer les outils par le code.** L'abstraction fondamentale de Claude Code est le CLI — l'agent vit sur votre ordinateur et accomplit des tâches complexes à travers des primitives de base comme bash et le système de fichiers. Quelques outils atomiques (comme l'outil bash) sont plus flexibles et consomment moins de tokens qu'un vaste ensemble d'outils.
***
## Empathie envers le modèle — penser comme un Agent
Les trois thèmes précédents — moins c'est plus, les outils deviennent obsolètes, laisser l'agent trouver ses propres réponses — partagent une méta-méthodologie commune. Thariq l'a clairement énoncé dès le début de l'article : **voir le monde comme un agent.**
Ce n'est pas un ensemble de règles, mais un mode de pensée. David Zhang lui a donné un nom :
**L'empathie envers le modèle** — il ne s'agit pas de concevoir des outils « logiques » du point de vue humain, mais de réfléchir du point de vue du modèle à ce qu'il voit réellement, comment il va comprendre et comment il va utiliser l'outil.
Terminally Drifting a résumé en trois étapes la leçon que chaque équipe agent finit par apprendre :
> 1. Vous concevez des outils pour des humains
> 2. Le modèle les utilise comme un raton laveur avec les droits administrateur
> 3. Vous les reconceptualisez pour l'économie de tokens et les effets secondaires prévisibles
« Penser comme un agent » est la percée décisive.
Vish a partagé son expérience de déclic : « Nous construisions sans cesse des interfaces d'outils logiques pour les humains, sans comprendre pourquoi l'agent faisait des choix étranges. **Dès que vous inversez le modèle mental et que vous réfléchissez à ce que le modèle voit réellement dans la définition de l'outil, tout change.** »
Emeka a exprimé la même idée du point de vue du développement d'outils en entreprise : « Tout est par défaut conçu selon le modèle mental humain — les schémas, les noms de champs, les processus, tout est optimisé pour la lecture humaine. Mais pour un agent, c'est le mauvais cadre de référence. »
Cette « inversion du modèle mental » est simple à énoncer, mais demande une pratique constante. Vous devez lire attentivement les sorties de l'agent — non pas pour voir ce qu'il a bien fait, mais pour comprendre pourquoi il a fait certains choix, où il a hésité, où il a fait des détours. Ces « comportements anormaux » ne sont souvent pas des bugs du modèle, mais des bugs de conception d'outils.
Plusieurs perspectives complémentaires intéressantes ont émergé des discussions de la communauté.
OAIR a soulevé un angle mort : **toutes ces itérations d'outils supposent que l'agent est sans état — chaque session repart de zéro. Et si l'« outil » le plus important ne se trouvait pas dans l'espace d'action, mais dans une mémoire persistante sur le fonctionnement de la base de code ?** Claude Code a partiellement répondu à cette question par la suite avec les fichiers CLAUDE.md et le système de mémoire, mais la gestion de l'état persistant reste un problème ouvert dans la conception d'agents.
Clinker a proposé un autre principe de conception : **optimiser la récupérabilité (recoverability) plutôt que la capacité brute — des préconditions d'outils explicites, un état observable et des tentatives peu coûteuses sont souvent plus efficaces que l'ajout rapide de nouveaux outils.** Cela rejoint le principe du génie logiciel qui consiste à « rendre le système tolérant aux pannes plutôt qu'infaillible ».
Le résumé de 范式折叠 est le plus incisif :
> La conception de l'action space est fondamentalement une conception de pouvoir — les permissions que vous donnez à l'IA déterminent son rôle. C'est exactement comme gérer une équipe : vous pensez que le goulot d'étranglement est la compétence des gens, alors qu'en réalité c'est le périmètre de permissions que vous avez tracé.
***
## En conclusion
Revenons à la conclusion de Thariq :
> Expérimentez davantage, lisez vos sorties, essayez de nouvelles approches. Observez comme un agent.
En tant qu'utilisateur assidu de Claude Code, ma principale impression après avoir lu cet article est la suivante : **ces fonctionnalités qui semblent « naturelles » sont le fruit d'innombrables itérations de type « ça ne marche pas, essayons autrement ».** AskUserQuestion a nécessité trois tentatives, TodoWrite a été reconçu trois fois, RAG a été remplacé par Grep. Chaque amélioration n'est pas venue d'une idée plus brillante, mais d'une observation attentive du comportement réel du modèle.
Le sous-titre de cet article est « Seeing like an Agent » — voir le monde comme un agent. Mais vu sous un autre angle, c'est aussi l'essence de toute bonne pratique d'ingénierie : **ne pas concevoir un système depuis sa propre perspective, mais depuis celle de l'utilisateur.** Sauf que cette fois, l'utilisateur est un modèle d'IA.
À l'avenir, chaque développeur construisant des agents devra probablement maîtriser ce que David Zhang appelle l'« empathie envers le modèle ». Ce n'est pas une capacité mystérieuse — elle repose sur trois choses : **observer le comportement réel du modèle, lire ses sorties, puis ajuster la conception en fonction de ce que vous observez.**
Observez comme un agent.
***
**Lectures complémentaires :**
# De la guerre des puces aux centres de données spatiaux : la prochaine décennie de l'industrie de l'IA
> L'investissement est une quête de vérité. Si vous trouvez la vérité en premier et que votre jugement est correct, c'est ainsi que vous créez de l'alpha. Et il faut que ce soit une vérité que les autres n'ont pas encore vue.
Ces mots viennent de Gavin Baker, fondateur d'Atreides Management, lors de son interview dans le podcast « Invest Like the Best » de Patrick O'Shaughnessy. Gavin est considéré comme l'un des investisseurs les plus passionnés et perspicaces du secteur technologique. Cet échange de près de deux heures couvre les GPU, les TPU, l'économie de l'IA, les centres de données spatiaux, l'avenir du SaaS, et même sa transition de moniteur de ski à investisseur.
Cette interview est extrêmement dense en informations, avec de nombreux moments qui vous font « marquer une pause pour réfléchir ». Voici quelques-uns des points de vue les plus stimulants.
***
## Comment suivre l'évolution de l'IA ? Commencez par dépenser 200 dollars
Au début de l'interview, Patrick pose une question très pratique : quand un nouveau modèle comme Gemini 3 est lancé, comment traitez-vous ces informations ?
La réponse de Gavin est directe : **vous devez l'utiliser vous-même**.
Mais le point essentiel n'est pas simplement de « l'utiliser », c'est quelle version vous utilisez. Il est surpris par les investisseurs qui utilisent la version gratuite de l'IA et en concluent que « l'IA n'est pas si impressionnante que ça » :
> La version gratuite, c'est comme si vous aviez affaire à un enfant de 10 ans, et que vous prédiquez ce qu'il deviendra à 35 ans en vous basant sur ses performances à 10 ans. Vous pouvez payer — en fait, vous devez payer pour obtenir l'abonnement de niveau supérieur, 200 dollars par mois. Ce sont eux les véritables adultes de 30-35 ans.
Cette analogie est très pertinente. L'écart entre les modèles d'IA suit la même logique — beaucoup de gens essaient un modèle gratuit de base et concluent que « l'IA n'est pas si extraordinaire ». Mais si vous avez utilisé Claude 4.5 Opus, Gemini 3 Pro ou GPT-5.2 Reasoning — ces modèles de pointe — l'expérience est radicalement différente.
En ce qui concerne les abonnements payants, je dépense personnellement environ 300 euros par mois pour divers produits IA, dont la majeure partie va à l'abonnement Claude Code Max à 250 dollars. Si vous êtes développeur avec des besoins importants en programmation, je vous recommande vivement de souscrire directement à Claude Code officiel (à partir de 125 dollars), plutôt que d'utiliser des sites miroirs. D'une part, vous ne savez pas si le site miroir utilise réellement le vrai modèle ; d'autre part, l'abonnement Max offre un excellent rapport qualité-prix, bien plus avantageux que la facturation à l'usage.
Quant aux canaux d'information, la réponse de Gavin pourrait en surprendre plus d'un : **X (Twitter)**.
Il explique que le développement de l'IA se déroule en grande partie « en temps réel sur X ». Il y a probablement entre 500 et 1 000 personnes sur Terre qui comprennent véritablement la frontière de l'IA, dont une proportion significative se trouve en Chine, et vous devez suivre ces personnes de près. Il mentionne particulièrement Andrej Karpathy :
> Chaque article qu'écrit Andrej Karpathy, vous devez le lire trois fois. Au minimum.
En tant que personne qui suit également l'actualité de l'IA, cela me parle profondément. Les discussions sur l'IA sur Twitter sont effectivement plus instantanées et plus approfondies que n'importe quel média traditionnel. Les chercheurs des laboratoires y publient directement les dernières avancées, et parfois se « disputent » entre eux — Gavin mentionne que l'équipe PyTorch de Meta et l'équipe Jax de Google ont eu un débat public sur X, au point que les responsables des deux laboratoires ont dû intervenir en déclarant : « Les membres de notre laboratoire n'ont pas le droit de dire du mal de l'autre laboratoire. »
***
## Scaling Laws : notre « moment Égypte ancienne »
Après la sortie de Gemini 3, beaucoup se sont intéressés à ce que cela implique pour les scaling laws (lois d'échelle). Gavin offre une perspective que je n'avais jamais entendue :
> Notre compréhension des scaling laws du pré-entraînement est probablement comparable à celle des anciens Égyptiens concernant le soleil. Ils pouvaient mesurer avec une précision extraordinaire — la Grande Pyramide dont l'axe est-ouest s'aligne parfaitement avec les équinoxes, tout comme Stonehenge. Des mesures parfaites. Mais ils ne comprenaient pas la mécanique orbitale. Ils ne savaient pas pourquoi le soleil se lève à l'est et se couche à l'ouest.
Cette analogie m'a fait longuement réfléchir. Nous pouvons effectivement prédire avec une grande précision : en multipliant par 10 la puissance de calcul d'un modèle, de combien ses performances s'amélioreront. Mais nous ne savons pas pourquoi. Ce n'est pas une « loi », mais une « observation empirique » — une observation que nous mesurons avec une extrême précision sans en comprendre le mécanisme sous-jacent.
Alors pourquoi Gemini 3 est-il important ? Parce qu'il prouve que cette « observation empirique » tient toujours. À un moment où les puces Blackwell étaient retardées et où tout le monde s'inquiétait de savoir si « les scaling laws avaient cessé de fonctionner », Gemini 3 a apporté une réponse claire : **elles fonctionnent toujours**.
Mais ce qui est encore plus intéressant, c'est ce que Gavin dit ensuite : sans l'apparition des « modèles de raisonnement » (reasoning models), le développement de l'IA entre 2024 et 2025 aurait dû stagner.
Pourquoi ? Parce qu'après que XAI a réussi à faire fonctionner 200 000 GPU Hopper de manière coordonnée, l'étape suivante nécessitait d'attendre les puces Blackwell. Il est impossible de maintenir plus de 200 000 Hopper en état de « cohérence » — autrement dit, de les faire fonctionner comme un tout unifié. Et Blackwell a pris du retard.
> Sans les modèles de raisonnement, du milieu de 2024 à aujourd'hui, l'IA n'aurait fait aucun progrès. Tout aurait stagné. Imaginez-vous ce que cela signifierait pour le marché ? Nous vivrions dans un environnement complètement différent. Les modèles de raisonnement ont en quelque sorte sauvé l'IA, car ils lui ont permis de continuer à progresser sans Blackwell.
C'est une perspective dont je n'avais pas conscience auparavant : les modèles de raisonnement (comme o1) ne sont pas simplement une nouvelle capacité — ils ont en réalité « sauvé » le rythme de développement de toute l'industrie de l'IA.
***
## La guerre des puces : Google « aspire l'oxygène »
En abordant la concurrence entre GPU et TPU, Gavin prononce une phrase qui m'a marqué :
> Google est actuellement le producteur de tokens au coût le plus bas. Ce qu'ils font depuis un moment, je dirais qu'ils « aspirent l'oxygène économique de l'écosystème IA » — c'est une stratégie extrêmement rationnelle de leur part.
En tant que producteur à bas coût, Google propose des services IA à prix cassés (voire à perte), rendant la vie très difficile aux concurrents. C'est une stratégie classique du secteur technologique, mais Gavin souligne un changement intéressant :
> L'IA est la première fois de ma carrière où être le « producteur à bas coût » compte vraiment dans le secteur technologique. Apple ne vaut pas des milliers de milliards parce qu'elle produit des téléphones à bas coût. Microsoft ne vaut pas des milliers de milliards parce qu'elle produit des logiciels à bas coût. NVIDIA ne vaut pas des milliers de milliards parce qu'elle produit des accélérateurs IA à bas coût. Cela n'a jamais compté.
Mais à l'ère de l'IA, quand l'énergie devient le facteur limitant, **le nombre de tokens produits par watt** devient crucial. Si vous pouvez produire 3 à 5 fois plus de tokens par watt, c'est 3 à 5 fois plus de revenus. Le prix du calcul devient hors sujet, car le goulot d'étranglement est l'énergie.
Ce paysage est sur le point de changer. Les puces Blackwell commencent enfin à être déployées, et Gavin prédit que le premier modèle Blackwell viendra de XAI :
> Selon Jensen, personne ne construit de centres de données plus vite qu'Elon. Jensen l'a dit publiquement.
Lorsque Blackwell et les puces Ruben suivantes seront déployées à grande échelle, l'avantage de Google en tant que producteur à bas coût disparaîtra. À ce moment-là, seront-ils toujours disposés à opérer leur activité IA avec une marge brute de -30 % ? Le calcul changera du tout au tout.
***
## Les centres de données spatiaux : une idée folle, mais logique du point de vue des premiers principes
Quand Patrick demande s'il existe « des idées folles dont on ne parle pas assez », Gavin commence à parler des centres de données spatiaux. Au départ, j'ai cru qu'il plaisantait, mais après avoir écouté son analyse, j'ai réalisé que c'est peut-être la partie la plus visionnaire de toute l'interview.
> Du point de vue des premiers principes, les centres de données spatiaux surpassent les centres de données terrestres dans toutes les dimensions.
Voici son raisonnement :
**1. Énergie** : dans l'espace, les satellites peuvent être exposés au soleil 24 heures sur 24, et l'intensité du rayonnement solaire est 6 fois supérieure à celle au sol. De plus, comme il y a toujours du soleil, vous n'avez pas besoin de batteries — qui représentent une part importante des coûts. L'énergie la moins chère du système solaire est donc le « solaire spatial ».
**2. Refroidissement** : sur Terre, une grande partie du coût et du poids des centres de données est consacrée au refroidissement. Mais dans l'espace ? Le refroidissement est gratuit. Placez les radiateurs du côté ombragé du satellite, où la température approche le zéro absolu.
**3. Réseau** : dans un centre de données, les racks sont connectés par fibre optique — essentiellement des lasers traversant des câbles. Qu'y a-t-il de plus rapide ? Des lasers traversant le vide. Donc si vous connectez des satellites dans l'espace par laser, le réseau est en fait plus rapide que dans un centre de données terrestre.
**4. Expérience utilisateur** : actuellement, quand vous posez une question à l'IA, le signal va de votre téléphone à une antenne-relais, puis par fibre optique jusqu'à un centre de données, le calcul est effectué, et tout le chemin est parcouru en sens inverse. Mais si un satellite pouvait communiquer directement avec votre téléphone (Starlink a déjà prouvé cette capacité de connexion directe), la chaîne serait beaucoup plus courte.
Bien sûr, cela nécessite des lancements massifs de Starship pour se concrétiser, probablement dans 5 à 6 ans. Mais Gavin souligne une convergence intéressante : Tesla, SpaceX et XAI sont en train de fusionner. XAI sera le « module intelligent » des robots Optimus, SpaceX construira des centres de données dans l'espace pour fournir de la puissance de calcul à l'IA — ces trois entreprises forment une boucle de renforcement mutuel de leurs avantages compétitifs.
***
## La « plateforme en feu » du SaaS
Si les sections précédentes vous ont enthousiasmé quant à l'avenir de l'IA, celle-ci pourrait vous inquiéter pour l'avenir de nombreuses entreprises existantes.
Gavin le dit sans détour : **les entreprises de SaaS applicatif commettent exactement la même erreur que les détaillants physiques face au e-commerce**.
Les détaillants physiques regardaient Amazon en pensant : « Le e-commerce est une activité à faible marge, comment pourrait-il être plus efficace que nous ? Actuellement, les clients viennent eux-mêmes en magasin et ramènent leurs achats chez eux. » Ils voyaient clairement la demande des clients, mais refusaient d'investir parce qu'ils n'aimaient pas la structure de marge du e-commerce. Résultat ? La marge bénéficiaire du commerce de détail nord-américain d'Amazon est aujourd'hui supérieure à celle de nombreux détaillants traditionnels.
Les entreprises SaaS font face à la même situation. Le logiciel traditionnel est écrit une fois puis distribué à l'infini, avec des marges brutes pouvant atteindre 80-90 %. Mais l'IA est différente — chaque utilisation nécessite un nouveau calcul, et une bonne entreprise IA n'aura peut-être qu'une marge brute de 40 %.
> Si vous voulez créer un agent IA et que vous n'êtes pas disposé à opérer avec une marge brute inférieure à 35 %, vous n'y arriverez jamais. Parce que les entreprises natives de l'IA opèrent précisément à ce niveau de marge. Si vous voulez protéger votre marge de 80 %, vous garantissez votre échec dans l'IA. Absolument garanti.
Gavin décrit cela comme une « décision de vie ou de mort », et **à part Microsoft, presque toutes les entreprises échouent**.
Il cite le fameux mémo de Nokia sur « la plateforme en feu » : votre plateforme est en flammes. Mais il y a en fait une excellente nouvelle plateforme juste à côté, vous pouvez y sauter, puis revenir éteindre le feu sur l'ancienne. Maintenant vous avez deux plateformes.
Salesforce, ServiceNow, HubSpot, GitLab, Atlassian — il estime que toutes ces entreprises peuvent et devraient exécuter cette stratégie : rendre publics vos revenus IA, rendre publique votre marge brute IA (une faible marge prouve justement que c'est de la « vraie IA »), puis pointer du doigt les concurrents soutenus par le capital-risque qui sont encore en perte et dire : « J'ai ce qu'ils n'ont pas : une activité qui génère du cash-flow. »
***
## L'histoire de la construction d'un investisseur
À la fin de l'interview, Patrick pose une question plus personnelle : comment présenteriez-vous votre métier à un jeune ?
La réponse de Gavin commence par « l'investissement est une quête de vérité », mais ce qui est vraiment fascinant, c'est son parcours de vie.
Son plan initial était : moniteur de ski en hiver, guide de rafting en été, escalade pendant les intersaisons, tout en essayant d'écrire des romans et de faire de la photographie animalière. C'était son « projet de vie » à l'université, et ses parents le soutenaient pleinement.
Mais ses parents ont eu une petite demande : pourrait-il faire un stage professionnel, juste un, n'importe lequel ?
Le seul stage qu'il a pu trouver était dans le département de gestion de patrimoine privé d'une maison de courtage. Le travail était simple : chaque fois que l'entreprise publiait un rapport de recherche, il devait vérifier quels clients détenaient cette action, puis leur envoyer le rapport.
Et puis il a commencé à lire ces rapports.
> Je me suis dit : « Mon Dieu, c'est la chose la plus fascinante que je puisse imaginer. »
Il décrit l'investissement comme un « jeu mêlant compétence et chance », un peu comme le poker. Vous pouvez perdre par malchance — par exemple, le siège de l'entreprise dans laquelle vous avez investi est frappé par une météorite — mais la plupart du temps, la compétence compte. Et le moyen d'acquérir un avantage est de posséder la connaissance historique la plus approfondie, combinée à la compréhension la plus juste du monde actuel, pour former un jugement différencié sur « ce qui va se passer ensuite ».
C'était son troisième jour de stage. Il est allé en librairie acheter le livre de Peter Lynch et l'a terminé en deux jours. Puis il a lu Buffett, lu « Market Wizards », lu les lettres de Buffett aux actionnaires — deux fois. Puis il a appris la comptabilité en autodidacte. De retour à l'université, il a changé sa spécialisation d'anglais et histoire à histoire et économie.
Il mentionne également une expérience en tant qu'agent d'entretien. Pendant qu'il travaillait à la station de ski d'Alta, il faisait le ménage des chambres. Un jour, en nettoyant une chambre, il a remarqué que le client lisait le même livre que lui. Il a dit : « C'est un bon livre, j'en suis à peu près au même endroit que vous. » L'autre l'a regardé comme s'il était un extraterrestre, puis, encore plus stupéfait, a demandé : « Vous lisez des livres ? »
> Cela a changé de manière permanente ma façon de traiter les autres.
***
## Épilogue : l'IA obtient ce dont elle a besoin
Vers la fin de l'interview, Gavin prononce ce que je considère comme la réflexion la plus fascinante :
> Ces deux dernières années, quoi que l'IA ait besoin pour continuer à progresser, elle l'a obtenu. Avez-vous déjà vu l'opinion publique américaine changer d'avis sur un sujet aussi vite que sur le nucléaire ? C'est arrivé comme ça. Et précisément au moment où l'IA en avait besoin. Maintenant que nous atteignons les limites énergétiques sur Terre, soudain la discussion sur les centres de données spatiaux apparaît. Chaque fois qu'un goulot d'étranglement risque de ralentir l'IA, tout s'accélère au contraire.
Cela rappelle le concept de « technium » (système technologique) proposé par Kevin Kelly dans « Ce que veut la technologie » : la technologie, prise dans son ensemble, semble avoir une sorte de volonté propre, aspirant à devenir toujours plus puissante.
Peut-être n'est-ce qu'une coïncidence. Peut-être n'est-ce que beaucoup de personnes brillantes qui résolvent des problèmes. Mais le schéma observé par Gavin — chaque obstacle rencontré par l'IA finit par être éliminé d'une manière ou d'une autre — mérite vraiment réflexion.
# Sélections commentées
# Curations – Sélections commentées
Vous trouverez ici mes réflexions et synthèses issues du visionnage de vidéos d'experts techniques et de la lecture de blogs de qualité.
Il ne s'agit pas de simples notes de lecture, mais de contenus enrichis par ma propre compréhension et mon expérience pratique.
## Sources
* Analyses de vidéos techniques
* Sélections d'articles de blog
* Synthèses de podcasts et d'interviews
## Derniers contenus
### Penser comme un Agent : la philosophie de conception d'outils de l'équipe Claude Code
Thariq, ingénieur chez Anthropic, partage son expérience de conception d'outils pour agents acquise lors de la construction de Claude Code — des trois itérations d'AskUserQuestion aux trois refactorisations de TodoWrite, du RAG à la divulgation progressive, chaque cas pointe vers la même méthodologie fondamentale : voir le monde comme un Agent.
[Lire l'article complet →](./claude-code-seeing-like-an-agent)
### 90 % des étudiants utilisent l'IA, mais personne ne connaît les règles
Quatre étudiants de Princeton, Berkeley, LSE ont discuté de la réalité de l'IA sur les campus — triche, confusion, polarisation, et cette question que personne n'ose poser : à quoi sert encore l'université ? Ce n'est pas un film promotionnel, ce sont de véritables interrogations et réflexions.
[Lire l'article complet →](./ai-on-campus-student-perspectives)
### De la guerre des puces aux centres de données spatiaux : la prochaine décennie de l'industrie de l'IA
De la guerre des puces aux centres de données spatiaux, du dilemme existentiel du SaaS à l'essence de l'investissement. Gavin Baker partage dans le podcast « Invest Like the Best » ses analyses les plus profondes sur l'industrie de l'IA : pourquoi vous devez utiliser la version payante de l'IA, pourquoi les Scaling Laws ressemblent à « la compréhension du soleil par les anciens Égyptiens », et comment les modèles de raisonnement ont « sauvé » le rythme de développement de toute l'industrie de l'IA.
[Lire l'article complet →](./gavin-baker-ai-economics)
# Le tweet de Karpathy a explosé et a atteint 62 000 étoiles : qu'a fait exactement Andrej-Karpathy-Skills ?
Le 27 janvier 2026, Andrej Karpathy a publié un très long tweet sur X – un essai de programmation de 11 sections, d'environ 1 400 mots, décrivant les pièges qu'il a rencontrés lors de la transition de « 80 % d'écriture manuscrite + 20 % d'agent » en novembre à « 80 % d'agent + 20 % de polissage » en décembre. Le tweet a finalement atteint **7,69 millions de vues, 39 000 likes et 36 000 favoris**.
Trois mois plus tard, un référentiel GitHub appelé `forrestchang/andrej-karpathy-skills` a été lancé, regroupant le tweet de Karpathy en quatre règles installables. En deux semaines, il a atteint **62,7 000 étoiles et 5,5 000 forks**, devenant ainsi le numéro 1 sur la liste hebdomadaire GitHub en avril 2026.
Ontologie d'entrepôt : **un fichier Markdown**.
***
## 1. De quoi se plaint Karpathy ?
Le long article de Karpathy énumère essentiellement quatre « conditions chroniques » codées en LLM.
**La première maladie : faire secrètement des hypothèses à votre place**
> "Le type d'erreur le plus courant est que les modèles font des hypothèses incorrectes pour vous et les appliquent ensuite sans les valider. Ils ne gèrent pas leur propre confusion, ils ne recherchent pas de clarification, ils ne montrent pas d'incohérences, ils ne présentent pas de compromis, ils ne réfutent pas quand il est temps de réfuter et ils sont un peu trop flatteurs."
Il s’agit d’une « erreur collusoire ». Vous dites "Ajoutez un identifiant pour moi", il ne demande pas quelle authentification utiliser, s'il faut mémoriser l'appareil ou comment gérer la session - il présente simplement une solution qu'il juge raisonnable. Au moment où vous aurez fini de réviser et constaterez qu’il est différent de ce que vous voulez, il aura déjà 500 lignes écrites.
**La deuxième maladie : la sur-ingénierie**
> "Ils aiment particulièrement compliquer à l'excès leur code et leurs API, gonfler les couches d'abstraction et ne pas nettoyer le code mort. Ils utiliseront 1 000 lignes de code pour implémenter une structure inefficace, gonflée et fragile, et vous devrez cajoler comme un enfant et dire : 'Eh bien, pourquoi ne faites-vous pas ça ?', avant de dire : 'Bien sûr !' puis réduisez-le immédiatement à 100 lignes.
Il s’agit de l’effet d’observateur le plus typique du codage LLM : tout en étant « généreusement » donné en contexte, il récompensera « généreusement » la complexité\*\*. Mode stratégie, mode usine, injection de dépendances : tout vous est proposé.
**La troisième maladie : changez quelque chose que vous n'avez pas demandé de changer**
> "Ils modifient ou suppriment parfois certains commentaires et certains codes parce qu'ils ne l'aiment pas ou ne le comprennent pas complètement - même si ces changements n'ont rien à voir avec la tâche en cours."
Vous lui demandez de corriger un bug, et il supprime commodément le commentaire TODO inachevé à côté de lui au motif qu'il "ne semble plus nécessaire".
**Maladie 4 : Même si vous écrivez les règles dans CLAUDE.md, cela ne fonctionnera toujours pas**
> "Le problème ci-dessus existe toujours même si j'ai effectué quelques simples tentatives de réparation dans CLAUDE.md."
C’est la phrase la plus déchirante de tout le tweet. Karpathy, ancien membre fondateur de l'équipe OpenAI et directeur de Tesla AI, n'a pas pu écrire CLAUDE.md qui maintiendrait Claude complètement en ligne.
***
## 2. Solution à `andrej-karpathy-skills`
`forrestchang` Systématisez les solutions à ces quatre maladies en quatre principes et regroupez-les dans un fichier CLAUDE.md.
| Principes | Maladies correspondantes | Actions de base |
| ------------------------------------ | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Réfléchissez avant de coder** | Assumer secrètement | Énoncez clairement vos hypothèses, énumérez plusieurs interprétations, arrêtez-vous et demandez si vous êtes confus et réfutez si nécessaire |
| **La simplicité d'abord** | Sur-ingénierie | Écrivez uniquement le code minimum requis ; n'écrivez pas de flexibilité spéculative, de gestion des erreurs, d'abstraction |
| **Changements chirurgicaux** | Modifications non autorisées | Ne touchez que ce qui est nécessaire ; ne refactorisez pas et ne modifiez pas le style ; signaler uniquement les autres codes morts mais ne les supprimez pas |
| **Exécution axée sur les objectifs** | Désalignement de la méthode | Donnez des critères de réussite + des tests vérifiables et laissez la boucle du modèle passer d'elle-même |
Le quatrième principe est une référence directe à une autre phrase célèbre du tweet de Karpathy – « Leverage » :
C'est le point d'appui de toute la méthodologie du projet : \*\*Les trois premiers principes empêchent LLM de déconner ; le quatrième principe vous indique comment véritablement exploiter ses atouts. \*\*
***
## 3. Comment installer et utiliser
**Méthode A : En tant que plug-in Claude Code** (il est recommandé de l'utiliser en premier)
```bash
/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills
```
Après installation, le dialogue Claude Code de tous les projets respectera automatiquement ces quatre principes. Il prend effet globalement et peut être désactivé à tout moment par `/plugin`.
**Méthode B : Copier manuellement CLAUDE.md**
Entrez le dépôt → ouvrez `CLAUDE.md` → copiez → collez dans `CLAUDE.md` dans le répertoire racine de votre projet. Uniquement efficace pour ce projet.
Le référentiel fournit également des adaptations supplémentaires `CURSOR.md` et `.cursor/rules/` - un ensemble de contenu couvrant les principaux IDE d'IA.
***
## 4. Pourquoi peut-il atteindre 62 000 étoiles ?
C’est un phénomène qui mérite d’être étudié. 62,7 000 étoiles est un nombre exagéré pour un « dépôt de fichier unique » - à titre de comparaison, le markitdown de Microsoft (9 000) et les compétences d'agent d'Addy Osmani (4,6 000) au cours de la même période combinées ne sont pas autant que cela.
Ventilé par poids d'impact :
**1. Karpathy IP Endorsement** - Le même contenu ne peut pas dépasser 10 000 s'il s'appelle `forrestchang-skills`. Karpathy est doté du capital culturel de "l'ancienne équipe fondatrice d'OpenAI + directeur Tesla AI + instructeur CS231n", et ses tweets sont accompagnés d'une étiquette "à lire absolument".
**2. Timing parfait** — L'Opus 4.7 est sorti le 16 avril et les plaintes pour ingénierie excessive étaient à leur paroxysme. Le repo est apparu juste au moment où tout le monde cherchait un antidote pour "empêcher Claude d'être si fou".
**3. Les problèmes sont universels** - Chaque utilisateur de Claude Code / Cursor a marché sur ces quatre pièges, et le taux d'empathie est proche de 100 %.
**4. Le seuil est extrêmement bas** - 1 fichier ou 2 lignes de commandes. Le coût d’une étoile est si bas qu’on peut l’ignorer. "Si vous ne l'installez pas, vous perdrez."
**5. Fort sentiment de vérifiabilité** – Les 4 principes sont clairs et faciles à retenir, et faciles à capturer et à transmettre. Contrairement au guide de projet d'invite de 1 000 lignes qui est prohibitif.
**6. README bilingue** ——`README.zh.md` consomme directement le trafic circulaire de l'IA chinoise, V2EX/instantanément/explose simultanément sur Weibo.
**7. Promotion croisée de l'auteur** - La phrase dans la colonne du haut *"Découvrez mon nouveau projet Multica"* dirige le trafic vers la propre plateforme d'agents commerciaux de l'auteur `multica-ai/multica`. \*\*Ce dépôt est essentiellement le sommet de l'entonnoir d'acquisition de clients de Multica. \*\*
**8. Meta fit** - Les « erreurs de codage LLM » dont il traite sont exactement ce que tous les lecteurs rencontrent lorsqu'ils codent avec LLM. La lecture et l'utilisation sont intégrées et le taux de conversion est extrêmement élevé.
En un mot : ce qu'il vend, ce n'est pas du code ou des outils, mais **emballer les émotions de Karpathy dans des règles installables** - c'est le cas le plus typique du « contenu est un produit » dans le cercle de la programmation de l'IA en 2026.
***
## 5. Mes suggestions d'utilisation
**Utilisez d'abord la méthode A pour installer globalement**. Voyez si cela améliore votre expérience lors de l'écriture d'outils et de scripts - en particulier lorsque vous laissez Claude modifier le code d'autres personnes, si cela réduit le problème des modifications aléatoires.
**Nous déciderons après une semaine ou deux s'il faut le fusionner dans le projet CLAUDE.md**. Le CLAUDE.md de chaque projet est déjà rempli de connaissances du domaine (système de conception, spécification des composants, processus de déploiement), tandis que l'ensemble de Karpathy est une méthodologie générale. Les deux ne s’opposent pas et peuvent se superposer – mais le timing doit attendre que vous soyez vraiment sûr de son utilité.
**Faites attention au coût** : Cela fera poser plus de questions à Claude, ce qui sera agaçant pour les gens habitués à « générer en une seule phrase » ; il risque de ne pas faire le léger nettoyage qui devrait être fait (trop strict) ; il sera contraint par des tâches d'exploration très vagues.
**Valeur plus profonde** : cela vous oblige à énoncer clairement vos exigences - ce qui se trouve être la condition préalable à toute ingénierie logicielle de haute qualité.
**Suivi à noter** : Forrestchang lui-même fait également la promotion de Multica, une « plateforme d'agents gérés open source » pour produire le mécanisme de compétences. Si cet ensemble de quatre principes devient finalement la norme de facto, Multica sera son véhicule utilitaire. Faites attention à cette ligne.
***
## Ressources de référence
# Indie Dev
Documentation de l'ensemble du processus de développement indépendant — de la préparation à la mise en ligne, étape par étape.
# Bark
# Bark
Un outil qui vous permet d'envoyer des notifications push personnalisées sur votre iPhone via de simples requêtes HTTP. Gratuit, open source et auto-hébergeable.
## Pourquoi le recommander
* **API minimaliste** - Une seule commande curl pour envoyer une notification push, sans configuration complexe
* **Open source et gratuit** - Licence MIT, entièrement open source, aucun frais
* **Respect de la vie privée** - Possibilité d'auto-héberger le serveur, les données de notification entièrement sous votre contrôle
## Cas d'utilisation
* **Notifications de scripts** - Recevoir une alerte lorsqu'une tâche longue est terminée, comme une sauvegarde de données
* **Surveillance de services** - Recevoir une alerte immédiate en cas d'anomalie du serveur
* **Intégration d'automatisation** - Résultats de build CI/CD, notifications de tâches planifiées
## Démarrage rapide
1. Téléchargez Bark depuis l'[App Store](https://apps.apple.com/app/bark-custom-notifications/id1403753865)
2. Ouvrez l'application et copiez votre adresse de push
3. Envoyez votre première notification :
```bash
curl https://api.day.app/YOUR_KEY/Hello
```
Notification avec titre :
```bash
curl https://api.day.app/YOUR_KEY/Titre/Contenu
```
## Informations du projet
* GitHub : [Finb/Bark](https://github.com/Finb/Bark)
* Stars : 7.2k+
* Licence : MIT
# Toolkit
# Toolkit
Vous trouverez ici les projets GitHub de qualité et les logiciels pratiques que j'ai découverts.
# Guide complet de Claude Agent Teams
## Introduction
Si vous avez utilisé les Subagents de Claude Code, vous avez probablement trouvé le développement parallèle déjà assez puissant. Mais les Subagents ont une limitation : ils ne peuvent que rapporter leurs résultats à l'Agent principal, sans pouvoir communiquer entre eux.
Agent Teams change complètement la donne. Imaginez : un Agent se charge de l'audit de sécurité, un autre de l'optimisation des performances, un troisième de la couverture de tests — non seulement ils travaillent en parallèle, mais ils peuvent aussi dialoguer directement, se challenger mutuellement et parvenir à un consensus. C'est la valeur fondamentale d'Agent Teams.
## Comprendre Agent Teams
L'architecture d'Agent Teams ressemble à une véritable équipe de développement :
```
┌─────────────────────────────────────────────────────────┐
│ Vous (utilisateur) │
└─────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
│ (Instance Claude principale, coordinateur) │
└─────────────────────────────────────────────────────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Teammate 1 │ │ Teammate 2 │ │ Teammate 3 │
│ Audit sécu. │◄─►│ Optim. perf.│◄─►│ Couv. tests │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
└──────────────┼──────────────┘
▼
┌─────────────────┐
│ Liste de tâches │
│ partagée │
└─────────────────┘
```
### Différences avec les Subagents
| Caractéristique | Subagent | Agent Teams |
| --------------------------- | ------------------------------------------------------------ | ---------------------------------------------------------- |
| **Contexte** | Contexte indépendant, résultats renvoyés à l'Agent principal | Contexte indépendant, exécution totalement autonome |
| **Communication** | Rapporte uniquement à l'Agent principal | Les Teammates communiquent directement entre eux |
| **Coordination des tâches** | L'Agent principal gère tout le travail | Liste de tâches partagée, coordination autonome |
| **Cas d'utilisation** | Tâches ciblées ne nécessitant que les résultats | Travaux complexes nécessitant discussion et collaboration |
| **Coût en tokens** | Plus faible : résumé renvoyé au contexte principal | Plus élevé : chaque Teammate est une instance indépendante |
En résumé : **les Subagents sont des sous-traitants envoyés exécuter des tâches, tandis qu'Agent Teams est une équipe projet collaborant dans la même pièce**.
### Pourquoi Agent Teams fonctionne
L'intuition fondamentale : **la spécialisation apporte la concentration**.
Lorsqu'un seul Agent traite des tâches complexes à plusieurs étapes, le contexte gonfle continuellement, nécessitant fréquemment un `/clear` pour réinitialiser. Agent Teams permet à chaque Teammate de conserver un domaine de spécialisation étroit, un contexte propre et des performances plus stables.
## Activer Agent Teams
Agent Teams est actuellement une fonctionnalité expérimentale, désactivée par défaut. Vous devez l'activer manuellement :
**Méthode 1 : Variable d'environnement**
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
**Méthode 2 : settings.json (recommandée, persistante)**
```json
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
```
## Utilisation de base
### Créer votre première Agent Team
Une fois activée, il suffit de demander à Claude en langage naturel de créer une équipe :
```
Crée une agent team pour examiner la PR #142.
Génère trois examinateurs :
- Un spécialisé dans les problèmes de sécurité
- Un qui vérifie l'impact sur les performances
- Un qui valide la couverture des tests
Qu'ils examinent chacun et rapportent leurs découvertes.
```
**Mots-clés** : Utilisez "create an agent team" ou "spawn an agent team". Si vous dites simplement "spawn agents", cela peut confondre Subagent et Agent Teams.
### Modes d'affichage
Agent Teams prend en charge deux modes d'affichage :
| Mode | Description | Prérequis |
| --------------- | --------------------------------------------------------- | ---------------------------- |
| **In-process** | Tous les Teammates s'exécutent dans le terminal principal | Aucune exigence particulière |
| **Split panes** | Chaque Teammate a son propre volet | Nécessite tmux ou iTerm2 |
Le mode par défaut est `auto` : si vous êtes dans tmux, il utilise les split panes, sinon le mode in-process.
**Configuration du mode d'affichage** :
```json
{
"teammateMode": "in-process"
}
```
**Spécification pour une seule session** :
```bash
claude --teammate-mode in-process
```
### Raccourcis clavier courants
| Action | Raccourci |
| ------------------------------------ | ------------ |
| Basculer entre les Teammates | `Shift+Down` |
| Revenir au Teammate précédent | `Shift+Up` |
| Afficher/masquer la liste des tâches | `Ctrl+T` |
| Interrompre le Teammate actuel | `Escape` |
| **Activer le Delegate Mode** | `Shift+Tab` |
| Entrer dans la session d'un Teammate | `Enter` |
### Delegate Mode (important)
Le Delegate Mode est l'une des fonctionnalités les plus importantes d'Agent Teams :
| Mode | Comportement du Lead |
| ----------------- | ----------------------------------------------------------------------- |
| **Mode normal** | Le Lead peut implémenter des tâches lui-même, écrire du code |
| **Delegate Mode** | Le Lead ne peut que coordonner, pas écrire de code ni exécuter de tests |
**Pourquoi le Delegate Mode est-il nécessaire** :
Sans cette restriction, le Lead a souvent tendance à « accaparer le travail » — alors que trois Teammates attendent de travailler, le Lead commence à coder lui-même. Avec le Delegate Mode activé, le Lead est contraint de jouer un rôle de pur chef de projet, se limitant à gérer les tâches, communiquer avec les Teammates et examiner les résultats.
```
# Après le lancement de l'équipe, appuyez immédiatement sur Shift+Tab pour activer
```
## Cas pratiques
### Cas 1 : Revue de code parallèle
Un seul examinateur a tendance à approfondir un seul type de problème. En séparant les dimensions d'examen en domaines indépendants, la sécurité, les performances et la couverture des tests reçoivent tous une attention égale :
```
Crée une agent team pour examiner cette PR. Génère trois examinateurs :
- Examinateur sécurité : vérifie l'authentification, l'autorisation, les vulnérabilités d'injection
- Examinateur performance : analyse la complexité algorithmique, les requêtes de base de données, les stratégies de cache
- Examinateur tests : valide la couverture des tests, les cas limites, la gestion d'erreurs
Qu'ils examinent chacun puis discutent des problèmes trouvés entre eux.
```
### Cas 2 : Débogage par hypothèses concurrentes
Lorsque la cause première est inconnue, un Agent unique a tendance à s'arrêter dès qu'il trouve une explication apparemment plausible. Laisser les Teammates se challenger mutuellement évite ce problème :
```
L'utilisateur signale que l'application se ferme après l'envoi d'un message,
au lieu de maintenir la connexion.
Génère 5 agent teammates pour investiguer différentes hypothèses.
Qu'ils discutent entre eux, tentent de réfuter les théories des autres,
comme un débat scientifique. Mettez à jour le rapport d'investigation
avec les conclusions ayant fait l'objet d'un consensus.
```
**Mécanisme clé** : La structure de débat. Plusieurs enquêteurs indépendants tentent activement de réfuter les théories des autres — les hypothèses qui survivent ont plus de chances d'être la véritable cause première.
### Cas 3 : Production de contenu en masse
Application typique pour des tâches non techniques — transformer une entrée en plusieurs sorties :
```
Crée une agent team pour transformer ce script vidéo en contenu pour quatre plateformes :
- Rédacteur d'article LinkedIn
- Rédacteur de thread Twitter
- Rédacteur de newsletter
- Rédacteur d'article de blog
Emplacement du script : /content/scripts/video-20.md
```
Chaque Teammate crée de manière indépendante tout en maintenant la cohérence du contenu.
### Cas 4 : Groupe d'assurance qualité
Un contrôle qualité de site de blog, déployant 5 Agents testant différents aspects en parallèle :
```
Crée une agent team pour un contrôle qualité complet du blog :
- Agent 1 : Test des pages principales (accueil, à propos, contact)
- Agent 2 : Test des pages d'articles (rendu, navigation, métadonnées SEO)
- Agent 3 : Vérification des liens (liens internes, liens externes, liens morts)
- Agent 4 : Validation SEO (titres, descriptions, données structurées)
- Agent 5 : Tests d'accessibilité (labels ARIA, contraste, navigation au clavier)
Générez un rapport de problèmes classés par priorité.
```
**Résultat** : Un contrôle complet réalisé en quelques minutes, là où il aurait fallu une exécution séquentielle manuelle. Chaque Agent se concentre sur son domaine, et les résultats sont consolidés en une liste de problèmes classés par priorité.
### Cas 5 : Mode de discussion multi-tours
Un pattern de prompt utile — faire discuter les Teammates comme en réunion :
```
Utilise Agent Teams pour créer 4 teammates qui discutent de [décision technique],
en 3 tours de discussion. Que les teammates échangent entre eux à chaque tour.
Un des teammates est spécifiquement chargé de la perspective Red Team,
pour formuler des critiques.
```
Ce mode est particulièrement adapté aux décisions d'architecture, aux choix technologiques et à d'autres scénarios nécessitant une évaluation multi-perspectives.
### Cas 6 : Projet de compilateur C
Anthropic a utilisé 16 Agents pour construire de zéro un compilateur C capable de compiler le noyau Linux :
| Indicateur | Données |
| --------------------- | ---------------------------------------------- |
| Nombre d'Agents | 16 instances parallèles |
| Sessions | \~2 000 sessions Claude Code |
| Coût | \~20 000 $ |
| Lignes de code | 100 000 lignes |
| Utilisation de tokens | 2 milliards en entrée + 140 millions en sortie |
Résultat final : un compilateur Rust capable de construire un Linux 6.9 amorçable sur x86, ARM et RISC-V.
## Gestion d'équipe
### Spécifier les Teammates et le modèle
Claude décide automatiquement du nombre de Teammates à générer selon la tâche, mais vous pouvez également le spécifier explicitement :
```
Crée 4 teammates pour refactoriser ces modules en parallèle.
Chaque teammate utilise le modèle Sonnet.
```
### Demander l'approbation du plan
Pour les tâches complexes ou à haut risque, vous pouvez demander aux Teammates de préparer un plan avant l'exécution :
```
Génère un teammate architecte pour refactoriser le module d'authentification.
Avant toute modification, exige l'approbation du plan.
```
Le Teammate, une fois son plan terminé, envoie une demande d'approbation au Lead. Le Lead peut alors approuver ou renvoyer avec des commentaires.
### Communiquer directement avec les Teammates
Chaque Teammate est une session Claude Code complète. Vous pouvez envoyer un message directement à n'importe quel Teammate :
* **Mode in-process** : Basculez avec `Shift+Down`, puis saisissez votre message
* **Mode split-pane** : Cliquez directement sur le volet correspondant
### Fermer des Teammates
```
Demande au teammate d'audit de sécurité de se fermer
```
Le Lead envoie une demande de fermeture ; le Teammate peut l'accepter ou la refuser (en expliquant pourquoi).
### Nettoyer l'équipe
Une fois terminé, demandez au Lead de nettoyer les ressources :
```
Nettoie l'équipe
```
**Important** : Passez toujours par le Lead pour le nettoyage. Ne laissez pas les Teammates effectuer le nettoyage, car cela pourrait entraîner des incohérences d'état des ressources.
## Bonnes pratiques
### Contrôle de la taille de l'équipe
Recommandations de taille :
| Taille de l'équipe | Cas d'utilisation |
| ------------------ | -------------------------------------------------------------- |
| 3 | Revue multi-perspectives simple |
| 4-5 | Développement de fonctionnalités ou refactorisation standard |
| 6+ | Migrations à grande échelle ou tâches d'architecture complexes |
**Règle générale** : Attribuer 5-6 tâches par Teammate est un bon équilibre. Si vous avez 15 tâches indépendantes, 3 Teammates constituent un bon point de départ.
### Granularité des tâches
* **Trop petite** : Les coûts de coordination dépassent les bénéfices
* **Trop grande** : Les Teammates travaillent trop longtemps sans point de contrôle, augmentant le risque de gaspillage
* **Idéale** : Des unités de travail indépendantes et complètes avec des résultats clairs (une fonction, un fichier de test, un rapport de revue)
### Éviter les conflits de fichiers
Deux Teammates modifiant le même fichier provoquent des écrasements. Lors de la répartition du travail, assurez-vous que chaque Teammate est responsable d'un ensemble de fichiers différent :
```
Teammate 1 : responsable du répertoire src/auth/
Teammate 2 : responsable du répertoire src/api/
Teammate 3 : responsable du répertoire src/utils/
```
### Surveillance et orientation
Vérifiez régulièrement la progression des Teammates et corrigez la direction si nécessaire. Laisser l'équipe fonctionner sans supervision trop longtemps augmente le risque de gaspillage.
Si le Lead commence à implémenter des tâches lui-même au lieu d'attendre les Teammates :
```
Attends que tes teammates aient terminé leurs tâches avant de continuer
```
### Fournir suffisamment de contexte
Les Teammates chargent automatiquement le contexte du projet (CLAUDE.md, serveurs MCP, Skills), mais n'héritent pas de l'historique de conversation du Lead. Fournissez suffisamment de détails lors de la génération :
```
Génère un teammate d'audit de sécurité avec ce prompt :
"Audite les vulnérabilités de sécurité du module d'authentification dans src/auth/.
Concentre-toi sur la gestion des tokens, la gestion des sessions, la validation des entrées.
L'application utilise des tokens JWT stockés dans des cookies httpOnly.
Rapporte les découvertes avec un niveau de gravité."
```
### Mode de vérification par auto-rapport
Incluez des critères de vérification clairs dans la description de la tâche pour que les Teammates s'auto-évaluent une fois terminé :
```
À la fin de la tâche, rapporte au Lead :
1. Quels fichiers tu as examinés
2. Quels problèmes tu as trouvés
3. Quelles modifications tu as apportées
4. Si les critères de validation (tests passés, lint sans avertissement, etc.) sont satisfaits
```
Ce mode réduit la charge de vérification du Lead tout en s'assurant que la tâche est réellement terminée et pas seulement « apparemment terminée ».
## Techniques avancées
### Utiliser les Hooks pour imposer des portes qualité
Imposez des règles via les Hooks lorsque les Teammates terminent leur travail :
```json
{
"hooks": {
"TeammateIdle": [
{
"command": "your-validation-script.sh"
}
],
"TaskCompleted": [
{
"command": "your-quality-check.sh"
}
]
}
}
```
* `TeammateIdle` : S'exécute quand un Teammate est sur le point de devenir inactif. Renvoyer le code de sortie 2 envoie un retour et fait continuer le Teammate
* `TaskCompleted` : S'exécute quand une tâche est marquée comme terminée. Renvoyer le code de sortie 2 bloque la complétion et envoie un retour
### Pré-approbation des permissions
Les demandes de permissions des Teammates remontent au Lead, ce qui peut causer des interruptions fréquentes. Pré-approuvez les opérations courantes avant la génération :
```json
{
"permissions": {
"allow": [
"Read(*)",
"Write(src/**)",
"Bash(npm test)"
]
}
}
```
### Combinaison avec Worktree
Agent Teams peut être utilisé conjointement avec Worktree, chaque Teammate travaillant dans son propre worktree :
```
Crée une agent team, chaque teammate travaille dans un worktree indépendant,
pour éviter les conflits de fichiers.
```
### Outils d'orchestration tiers
En plus d'Agent Teams natif, la communauté a développé des outils d'orchestration :
| Outil | Description |
| --------------- | ---------------------------------------------------------------- |
| **Gas Town** | Outil de gestion de sessions Claude parallèles multiples |
| **Multiclaude** | Exécution d'instances Claude dans plusieurs fenêtres de terminal |
Ces outils offrent des alternatives en dehors de la fonctionnalité expérimentale Agent Teams, mais nécessitent davantage de configuration manuelle. Si Agent Teams natif répond à vos besoins, il est recommandé de privilégier la fonctionnalité officielle.
## Limitations actuelles
Agent Teams est encore une fonctionnalité expérimentale ; il est important d'en connaître les limitations :
| Limitation | Description |
| ------------------------------------------------ | ---------------------------------------------------------------------------------------- |
| Impossible de restaurer les teammates in-process | `/resume` et `/rewind` ne restaurent pas les teammates in-process |
| L'état des tâches peut être en retard | Les Teammates oublient parfois de marquer les tâches comme terminées |
| La fermeture peut être lente | Les Teammates terminent leur requête en cours avant de se fermer |
| Une équipe par session | Le Lead ne peut gérer qu'une seule équipe à la fois |
| Pas d'équipes imbriquées | Les Teammates ne peuvent pas créer leurs propres équipes |
| Lead fixe | La session qui crée l'équipe est le Lead, non transférable |
| Split panes nécessitent tmux/iTerm2 | Non pris en charge dans les terminaux VS Code, Windows Terminal, Ghostty |
| Plan mode au niveau session | L'état Plan mode d'un Teammate est verrouillé à la génération, non modifiable en session |
## Considérations de coût
La consommation de tokens d'Agent Teams est significativement plus élevée qu'une session unique :
| Scénario | Consommation de tokens | Multiplicateur de coût |
| ----------------------------------- | ---------------------- | ---------------------- |
| Session Agent unique | \~200k tokens | 1x |
| 3 Teammates | \~800k tokens | \~4x |
| 5 Teammates | \~1,2M tokens | \~6x |
| 16 Teammates (cas du compilateur C) | 2 milliards de tokens | 20 000 $ / 2 semaines |
**Analyse des coûts** :
* Chaque Teammate est une instance Claude totalement indépendante, avec son propre contexte
* La communication entre Teammates consomme également des tokens
* Le Lead doit coordonner tous les Teammates, ce qui ajoute des frais supplémentaires
**Quand cela en vaut la peine** :
* Tâches de recherche nécessitant une exploration parallèle
* Revues multi-perspectives (sécurité, performances, tests)
* Décisions nécessitant discussion et consensus
* Pas pour les tâches routinières pouvant être effectuées séquentiellement
* Pas pour les tâches parallèles ne nécessitant pas de communication mutuelle (les Subagents sont plus économiques)
## Mon retour d'expérience
### Quand utiliser Agent Teams
Mes critères de décision :
1. **La tâche nécessite plusieurs perspectives** : Expertise de domaines différents (sécurité + performances + tests)
2. **Discussion et consensus nécessaires** : Hypothèses concurrentes, décisions d'architecture
3. **L'exploration parallèle apporte de la valeur** : Comparaison de plusieurs approches d'implémentation
Si vous avez uniquement besoin d'exécution parallèle sans communication mutuelle, les Subagents ou Worktree sont plus adaptés.
### Commencez par la recherche et la revue
Si vous débutez avec Agent Teams, commencez par des tâches qui ne nécessitent pas d'écriture de code : revue de PR, recherche de solutions techniques, investigation de bugs. Ces tâches ont des limites claires, démontrent la valeur de l'exploration parallèle et évitent les défis de coordination liés à l'implémentation parallèle.
### Combinaison avec d'autres fonctionnalités
| Combinaison | Effet |
| ---------------------- | ----------------------------------------------------- |
| Agent Teams + Worktree | Chaque Teammate travaille dans un environnement isolé |
| Agent Teams + Hooks | Automatisation des contrôles qualité et des retours |
| Agent Teams + Skills | Chaque Teammate dispose de compétences spécialisées |
## Conclusion
Agent Teams représente un nouveau paradigme du développement assisté par IA : du « un assistant IA » à « une équipe IA ».
Anthropic a utilisé 16 Agents pendant 2 semaines, pour 20 000 $, et a produit 100 000 lignes de code pour un compilateur C. L'enseignement clé de ce projet : **la qualité des tests est plus importante que tout**. Les Agents résolvent de manière autonome les problèmes que vous leur posez, donc les validateurs de tâches doivent être quasi parfaits, sinon les Agents résoudront le mauvais problème.
Retenez trois points essentiels :
| Point | Description |
| ----------------- | ------------------------------------------------------------------------------------ |
| **Collaboration** | Les Teammates peuvent communiquer directement, pas seulement rapporter des résultats |
| **Partage** | Coordination du travail via une liste de tâches partagée |
| **Supervision** | Vérification régulière de la progression, correction de direction en temps opportun |
Pour commencer, c'est très simple :
```
Crée une agent team pour [votre tâche]
```
***
**Lectures complémentaires** :
* [Guide complet de Claude Worktree](/fr/docs/notes/claude-worktree) — Comprendre la combinaison entre Worktree et Agent Teams
* [Guide complet de Claude Subagent](/fr/docs/notes/claude-subagent) — Comparer les cas d'utilisation de Subagent et Agent Teams
* [Guide de démarrage rapide Tmux](/fr/docs/notes/tmux-tutorial) — Utiliser Tmux pour gérer plusieurs sessions Agent
**Références** :
* [Documentation officielle Claude Code - Agent Teams](https://code.claude.com/docs/en/agent-teams)
* [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler)
* [Claude Code Agent Teams: The Complete Guide](https://claudefa.st/blog/guide/agents/agent-teams)
* [Multi-Agent Orchestration: Running 10+ Claude Instances in Parallel](https://dev.to/bredmond1019/multi-agent-orchestration-running-10-claude-instances-in-parallel-part-3-29da)
**Tutoriels vidéo** :
* [7 Unexpected Use Cases for Claude Code Agent Teams](https://www.youtube.com/watch?v=dlb_XgFVrHQ) — 7 cas d'utilisation non techniques
* [Claude Code Multi-Agent Workflows](https://www.youtube.com/watch?v=X8afcX2s2Mo) — Flux de travail multi-agents expliqué en détail
* [Agent Teams Orchestration Deep Dive](https://www.youtube.com/watch?v=Dyj9ShddyMw) — Analyse approfondie de l'orchestration d'équipe
* [Claude Code Agent Teams Setup Guide](https://www.youtube.com/watch?v=pSJiFqVxx3k) — Tutoriel de configuration complet
# Analyse complete de l'architecture du systeme Claude
## Introduction
En septembre 2025, Anthropic a realise une levee de fonds de 13 milliards de dollars avec une **valorisation de 183 milliards de dollars**, devenant ainsi la quatrieme plus grande entreprise privee au monde. Son produit phare, Claude Code, a attire **115 000** developpeurs actifs depuis son lancement en fevrier, traitant **195 millions de lignes** de code par semaine, avec une croissance des utilisateurs de **300 %**.
Plus fascinant encore, Dario Amodei, PDG d'Anthropic, a revele : **90 % du code de Claude Code est ecrit par lui-meme**.
**Comment est-ce possible ?**
Comment un assistant de programmation IA peut-il « s'ecrire lui-meme » ? Qu'est-ce qui rend son architecture si unique pour lui permettre d'assister — voire de remplacer — aussi efficacement les developpeurs humains ?
La reponse se trouve dans l'**architecture modulaire** de Claude : MCP fournit les outils, Skills enseigne comment les utiliser, Subagents execute les tâches en parallele, et Hooks assure un controle deterministe — ces composants collaborent pour donner a Claude la capacite de « travailler comme un programmeur ».
Cet article vous offre une **vue d'ensemble de cette architecture** — le positionnement de chaque composant, leurs synergies, et des exemples de configuration pour demarrer rapidement. Les articles suivants approfondiront les details de chaque composant.
## Vue d'ensemble de l'architecture
Le systeme Claude utilise une **architecture modulaire** ou les composants sont classes par fonction et fonctionnent comme des **collaborateurs complementaires** plutôt que des dependances hierarchiques :
**Point essentiel** : ces composants sont des extensions **complementaires de meme niveau**, et non des dependances hierarchiques — vous pouvez les combiner librement selon vos besoins :
| Vous souhaitez... | Utiliser... | Description en une phrase |
| ------------------------------------------------------- | ------------- | --------------------------------------------------------------------------------------------------------- |
| Connecter des sources de donnees et services externes | **MCP** | Donner a Claude des « mains et des pieds » pour acceder aux bases de donnees, API et systemes de fichiers |
| Enseigner a Claude des workflows specifiques | **Skills** | Permettre a Claude de « savoir » comment travailler dans un domaine particulier |
| Traiter des tâches complexes en parallele | **Subagents** | Decomposer les grandes tâches en sous-tâches, avec plusieurs Agents travaillant simultanement |
| Declencher rapidement des operations repetitives | **Commands** | Lancement en un clic des workflows courants, sans repeter les instructions |
| S'assurer que certaines operations s'executent toujours | **Hooks** | Quelle que soit la decision de Claude, cette etape doit s'executer |
***
## Environnement d'execution principal
### Agent SDK — Le moteur d'execution
Agent SDK est le **moteur d'execution central** de l'ensemble du systeme Claude Agent, fournissant :
* **Boucle principale (Main Loop)** : le cycle de travail central de l'Agent
* **Gestion du contexte** : budgets de tokens, compression automatique (declenchee a 92 % d'utilisation)
* **Repartition des outils** : decision de l'outil a utiliser et comment l'executer
* **Systeme de permissions** : controle des droits d'acces aux outils
Le mode de fonctionnement principal de l'Agent est une simple **boucle de retroaction** :
```
Collecter le contexte → Executer les actions → Verifier le travail → Repeter
```
***
### Built-in Tools — Outils natifs
Claude Agent integre plus de 20 outils natifs, repartis en trois categories :
| Categorie | Outils | Description |
| ------------- | ------------------- | ------------------------------------------------------------------- |
| **Lecture** | Read, Glob, Grep | Lecture de fichiers, correspondance de motifs, recherche de contenu |
| **Operation** | Write, Edit, Bash | Ecriture de fichiers, edition, execution de commandes |
| **Reseau** | WebSearch, WebFetch | Recherche web, recuperation de pages web |
Ces outils sont **disponibles par defaut**, sans configuration supplementaire. Claude interagit avec l'ordinateur a travers ces outils, exactement comme un programmeur utilise un IDE.
***
## Configuration et contexte
### CLAUDE.md — Contexte persistant
A chaque nouvelle conversation, vous deviez repeter le contexte du projet, les normes de codage, les conventions d'architecture... CLAUDE.md permet de **configurer une seule fois et charger automatiquement**.
CLAUDE.md est comme un **README pour l'IA** — il informe Claude du contexte du projet, des methodes de travail et des conventions.
#### Hierarchie de priorite
Claude charge les fichiers CLAUDE.md dans l'ordre suivant, les **fichiers les plus specifiques ayant la priorite la plus elevee** :
```
Enterprise (plus basse)
↓
User (~/.claude/CLAUDE.md)
↓
Project (./CLAUDE.md)
↓
Module (./src/module/CLAUDE.md) (plus haute)
```
#### Contenu recommande
CLAUDE.md devrait inclure les informations essentielles suivantes :
| Categorie | Exemple de contenu |
| ----------------------- | -------------------------------------------------------- |
| **Stack technique** | Next.js 14 + TypeScript, Tailwind CSS |
| **Commandes de build** | `npm run dev`, `npm run build`, `npm run test` |
| **Normes de code** | Conventions de nommage, configuration des outils de lint |
| **Structure du projet** | Description de l'utilite des repertoires cles |
**Principe cle** : restez concis. CLAUDE.md est **charge a chaque conversation** — un fichier trop long gaspille de precieux tokens.
***
## Packaging et distribution
### Plugins — Unites installables
Les configurations d'equipe sont dispersees, difficiles a partager et a standardiser. Chacun a son propre ensemble de Skills, Commands, Hooks... Comment unifier la gestion ?
Les Plugins regroupent **Skills + Commands + Subagents + Hooks + MCP** en **unites installables**, permettant une distribution en un clic et la standardisation au sein de l'equipe.
#### Structure de repertoire
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # Manifeste du plugin (requis)
├── commands/ # Slash Commands
├── agents/ # Subagents
├── skills/ # Skills
├── hooks/ # Configuration des Hooks
├── .mcp.json # Configuration du serveur MCP
└── README.md # Documentation
```
#### Exemple de configuration
```json
// plugin.json
{
"name": "frontend-toolkit",
"version": "1.0.0",
"description": "Boite a outils de developpement frontend",
"author": "Your Team"
}
```
```bash
# Methodes d'installation
claude plugin install github:your-org/your-plugin # Depuis GitHub
claude plugin install /path/to/plugin # Depuis un chemin local
```
**Ressources associees**
| Ressource | Description |
| -------------------------------------------------------------------------------- | ------------------------------------------------------------- |
| [claude-plugins-official](https://github.com/anthropics/claude-plugins-official) | Depot officiel des plugins Anthropic |
| [wshobson/agents](https://github.com/wshobson/agents) | 24.3k etoiles, collection de modeles d'Agent de haute qualite |
| [Claude Plugin Hub](https://www.claudepluginhub.com/) | Marketplace communautaire pour decouvrir des plugins |
| [awesome-claude-code](https://github.com/hesreallyhim/awesome-claude-code) | 19.3k etoiles, liste organisee de ressources Claude Code |
***
## Capacites d'extension (modules complementaires)
Les capacites d'extension du systeme Claude sont composees de plusieurs **modules complementaires**, chacun avec un role distinct, fonctionnant en synergie :
| Module | Role fonctionnel | Mode d'activation |
| ------------- | ----------------------------------------------------------------- | ------------------------------ |
| **MCP** | Connexion aux donnees et services externes (WHAT) | Disponible apres configuration |
| **Skills** | Connaissance procedurale — enseigner a Claude comment faire (HOW) | Correspondance automatique |
| **Subagents** | Contexte independant, delegation de tâches en parallele | Invocation explicite |
| **Commands** | Workflows repetitifs | `/cmd` manuel |
| **Hooks** | Controle deterministe, pilote par evenements | Declenchement automatique |
***
### MCP — Connectivite externe
#### Philosophie de conception
Dans l'approche traditionnelle, chaque source de donnees externe necessite une **integration sur mesure**, conduisant a un cauchemar d'integration N-fois-M. MCP fournit un protocole standardise pour **integrer une fois, utiliser partout**.
MCP (Model Context Protocol) est concu comme le **port USB-C des applications IA** :
| Caracteristique | Description |
| --------------------------- | ----------------------------------------------------------------------------------------------- |
| **Standard ouvert** | Publie en novembre 2024, donne a la Fondation Linux en decembre 2025 |
| **Adoption industrielle** | Adopte par OpenAI, Microsoft, Google, AWS et d'autres |
| **Echelle de l'ecosysteme** | Plus de 97 millions de telechargements mensuels du SDK, des milliers de serveurs communautaires |
#### Modele d'architecture
```
┌─────────────────┐
│ MCP Host │ (Claude Desktop, IDE, outils IA)
│ (Main Agent) │
└────────┬────────┘
│ 1:N
┌────┴────┬────────┬────────┐
│ Client │ Client │ Client │ (clients de protocole)
└────┬────┴────┬───┴────┬───┘
│ │ │
┌────┴────┐ ┌──┴───┐ ┌─┴────┐
│ Server │ │Server│ │Server│ (expose des capacites specifiques)
│ GitHub │ │ DB │ │ Slack│
└─────────┘ └──────┘ └──────┘
```
**Cas d'utilisation** : connexion a des bases de donnees, integration de services tiers (GitHub, Slack, Notion), acces a des API privees, traitement de flux de donnees en temps reel.
#### Exemple de configuration
Creez `.mcp.json` a la racine du projet :
```json
{
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
}
}
}
```
***
### Commands — Workflows manuels
Les Slash Commands fournissent des workflows repetitifs a **declenchement manuel**.
| Caracteristique | Description |
| ------------------------- | ---------------------------------------------- |
| **Mode de declenchement** | Saisir manuellement `/command-name` |
| **Emplacement** | `.claude/commands/` |
| **Utilite** | Workflows repetitifs, operations standardisees |
**Exemple** : creez `.claude/commands/review.md`
```markdown
Veuillez effectuer une revue de code sur les modifications en cours, en vous concentrant sur :
1. Le style et la coherence du code
2. Les problemes de performance potentiels
3. Les vulnerabilites de securite
4. La couverture des tests
```
Ensuite, saisissez `/review` pour le declencher.
***
### Hooks — Controle deterministe
Les Hooks sont au coeur du **controle deterministe** — certaines operations doivent s'executer et ne peuvent pas dependre du jugement du LLM.
| Categorie | Evenement | Moment de declenchement |
| ------------ | -------------------- | ------------------------------------------------- |
| **Outil** | `PreToolUse` | Avant l'execution de l'outil |
| | `PostToolUse` | Apres l'execution reussie de l'outil |
| | `PostToolUseFailure` | Apres l'echec de l'execution de l'outil |
| | `PermissionRequest` | Lors d'une demande de permission |
| **Session** | `SessionStart` | Au demarrage d'une session |
| | `SessionEnd` | A la fin d'une session |
| | `Stop` | Lorsque Claude termine une reponse |
| **Subagent** | `SubagentStart` | Au demarrage d'un subagent |
| | `SubagentStop` | A l'arret d'un subagent |
| **Autre** | `UserPromptSubmit` | Apres la soumission d'un prompt par l'utilisateur |
| | `Notification` | Evenements de notification |
| | `PreCompact` | Avant la compression du contexte |
#### Exemple de configuration
Formatage automatique des fichiers TypeScript :
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "jq -r '.tool_input.file_path' | { read f; [[ $f == *.ts ]] && npx prettier --write \"$f\"; }"
}
]
}
]
}
}
```
#### En pratique : boucle autonome
**Ralph Wiggum** est un plugin officiel d'Anthropic qui utilise le hook Stop pour implementer une boucle d'iteration autonome :
```bash
/ralph-loop "Implementer une API TODO avec CRUD et tests" --max-iterations 20
```
**Fonctionnement** : le hook Stop intercepte la sortie de Claude, reinjecte le prompt original et continue l'iteration jusqu'a ce que la tâche soit terminee ou que le nombre maximum d'iterations soit atteint.
**Cas d'utilisation ideaux** : tâches necessitant plusieurs iterations (passage de tests, refactoring de code) et tâches disposant de moyens de verification automatique.
***
### Subagents — Delegation de tâches et execution parallele
#### Philosophie de conception
Un Agent unique fait face a des defis : fenetre de contexte limitee, impossibilite de paralleliser, responsabilites floues. Les Subagents adoptent le modele d'architecture **Orchestrator-Worker** pour resoudre ces problemes :
```
Main Agent (Orchestrateur)
├── Analyser la demande utilisateur
├── Formuler un plan
├── Decomposer les tâches
└── Generer des subagents specialises
↓
┌────────┬────────┬────────┐
│ Sub 1 │ Sub 2 │ Sub 3 │ (Workers - execution parallele)
│ Code │ Tests │ Docs │
└────────┴────────┴────────┘
↓
Agreger les resultats → L'agent principal synthetise la sortie
```
#### Caracteristiques principales
| Caracteristique | Description |
| ------------------------------------- | --------------------------------------------------------------------------- |
| **Isolation du contexte** | Chaque Subagent dispose d'un contexte independant, evitant la contamination |
| **Specialisation des tâches** | Des prompts systeme personnalises definissent des roles dedies |
| **Controle des permissions d'outils** | Les Subagents peuvent etre limites a des outils specifiques |
| **Execution parallele** | Plusieurs Subagents travaillent simultanement |
**Donnees de performance** : les systemes multi-agents surpassent les agents uniques de 90,2 %, et la parallelisation peut reduire le temps de recherche de 90 % (la consommation de tokens est d'environ 15x, mais cela en vaut la peine pour les tâches complexes).
#### Exemple de configuration
Creez un fichier Markdown dans `.claude/agents/` :
```markdown
---
name: Code Reviewer
description: Un subagent specialise dans la revue de code
tools:
- Read
- Grep
- Glob
---
Vous etes un expert senior en revue de code. Concentrez-vous sur :
1. La qualite du code et sa maintenabilite
2. Les bugs potentiels et les cas limites
3. Les opportunites d'optimisation de performance
4. Les vulnerabilites de securite
```
***
### Skills — Connaissances procedurales
#### Philosophie de conception
Les Skills sont des **manuels de travail reutilisables pour l'IA** — des packages de connaissances modulaires que Claude peut charger dynamiquement a la demande. Le principe de conception central est la **divulgation progressive (Progressive Disclosure)** :
```
📚 Manuel de Skills
│
├─ 📋 Table des matieres ──── [Couche de metadonnees] Prechargee au demarrage (~30-50 tokens)
│ name: "weekly-report"
│ description: "Generer des rapports hebdomadaires standardises"
│
├─ 📖 Contenu principal ────── [Couche de documentation centrale] Chargee quand pertinent (~centaines a milliers de tokens)
│ # Weekly Report Generator
│ ## Instructions
│ Generer des rapports hebdomadaires selon cette structure...
│
└─ 📎 Annexe ────────────── [Couche de ressources de reference] Chargee a la demande
references/
├── template.xlsx
└── examples/
```
**MCP vs Skills** : MCP donne a Claude la capacite d'acceder aux outils (WHAT), tandis que Skills enseigne a Claude comment utiliser efficacement ces outils (HOW).
#### Avantages principaux
| Avantage | Description |
| -------------------------- | -------------------------------------------------------------------------------------------------------- |
| **Efficacite en tokens** | Les metadonnees ne prennent que 30-50 tokens ; des dizaines de Skills peuvent etre actives simultanement |
| **Activation automatique** | Correspondance automatique basee sur le contexte de la tâche, sans declenchement manuel |
| **Composable** | Plusieurs Skills collaborent automatiquement |
| **Portable** | Experience coherente a travers Claude.ai, Claude Code et l'API |
#### Exemple de configuration
Creez un repertoire dans `.claude/skills/` :
```
my-skill/
├── SKILL.md # Instructions principales (requis)
├── scripts/ # Scripts executables (optionnel)
└── references/ # Materiaux de reference (optionnel)
```
Structure principale de SKILL.md :
```yaml
---
name: code-review # Nom du Skill
description: Revue de code pour la qualite et la securite # Description concise (utilisee pour la correspondance automatique)
---
# Code Review Skill
## Instructions
[Instructions detaillees etape par etape...]
## Output Format
[Exigences de format de sortie...]
```
**Point cle** : la `description` dans le frontmatter est utilisee pour la correspondance automatique — restez concis et precis.
***
## References officielles
**Philosophie de conception**
| Ressource | Description |
| --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) | Article fondamental sur l'architecture des Agents |
| [Building agents with the Claude Agent SDK](https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk) | Pratiques d'ingenierie de l'Agent SDK |
| [Equipping Agents with Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Philosophie de conception des Skills |
| [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol) | Annonce de lancement de MCP |
**Documentation officielle**
| Ressource | Description |
| ----------------------------------------------------------------------------- | ----------------------------------------------- |
| [Skills explained](https://claude.com/blog/skills-explained) | Comparaison des Skills avec d'autres composants |
| [Using CLAUDE.md Files](https://claude.com/blog/using-claude-md-files) | Guide d'utilisation de CLAUDE.md |
| [Claude Code Subagents](https://code.claude.com/docs/en/sub-agents) | Documentation officielle des Subagents |
| [Claude Code Hooks](https://code.claude.com/docs/en/hooks) | Documentation officielle des Hooks |
| [MCP Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/mcp) | Documentation officielle de MCP |
| [MCP Specification](https://modelcontextprotocol.io/specification/2025-11-25) | Specification du protocole MCP |
**Analyses approfondies**
| Ressource | Description |
| --------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- |
| [Understanding Claude Code's Full Stack](https://alexop.dev/posts/understanding-claude-code-full-stack/) | Inclut des diagrammes d'architecture |
| [Claude Agent Skills: Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Analyse approfondie du fonctionnement interne des Skills |
| [How Claude Code is built](https://newsletter.pragmaticengineer.com/p/how-claude-code-is-built) | Les coulisses de la construction de Claude Code |
| [Claude Skills vs. MCP](https://intuitionlabs.ai/articles/claude-skills-vs-mcp) | Comparaison technique entre Skills et MCP |
***
## Lectures complementaires
Si vous souhaitez approfondir les concepts et pratiques des Skills, consultez :
* [Qu'est-ce que Claude Skills](/fr/docs/notes/claude-skills/concept) — Explication detaillee des principes fondamentaux des Skills
* [Guide pratique de Claude Skills](/fr/docs/notes/claude-skills/practice) — Creez votre premier Skill pas a pas
* [Guide complet de Claude Subagent](/fr/docs/notes/claude-subagent) — Utilisation et personnalisation des subagents
* [Analyse approfondie de GSD](/fr/docs/notes/gsd/concept) — Un systeme de programmation IA base sur l'ingenierie de contexte
* [Mes bonnes pratiques Claude Code](/fr/blog/claude-code-best-practices) — Conseils pour l'utilisation quotidienne de Claude Code
# Claude Code 真正厉害的,不是会调用工具
我最开始理解 Claude Code 的时候,也很容易把它想成一件简单的事:Claude 后面接了一组工具。模型觉得该读文件,就调 `Read`;觉得该改代码,就调 `Edit`;觉得该跑测试,就调 `Bash`。
这个理解没有错,但太浅了。
如果 Claude Code 只是“会调用工具”,它不会比一个接了 function calling 的聊天框强太多。真正让它可以进入真实项目的,不是工具数量,而是它把这些工具放进了一条可以持续推进、不断验证、遇到风险能被拦住的循环里。
这篇笔记想讲的不是“Claude Code 有哪些工具”,而是另一个更重要的问题:
> 当模型开始真正改文件、跑命令、连外部系统时,什么样的工具系统才配得上真实开发环境?
我的答案是:Claude Code 的重点不是 tool calling,而是 agent runtime。
## 工具调用的误解
我们平时说“工具调用”,很容易想到一个 API 过程:
```text
模型判断意图
-> 选择函数
-> 按 schema 填参数
-> 函数返回结果
-> 模型继续回答
```
这套模型适合解释天气查询、数据库问答、简单的业务 API 调用。但它不足以解释 Claude Code。
因为开发任务不是一次函数调用能完成的。
你让 Claude Code 处理一个登录测试失败,它不知道失败原因在哪里;它要先跑测试,看错误;再读测试文件;再搜相关实现;再改代码;再跑测试;如果失败继续读日志;如果改动影响了别的模块,还要扩大验证范围。
这不是“调用一个工具拿答案”,而是一条循环:
```text
观察当前状态
-> 选择下一步行动
-> 执行工具
-> 接收结果
-> 更新判断
-> 再选择下一步
```
这才是 Claude Code 和普通聊天模型的分界线。普通聊天模型主要生成文本,Claude Code 会在这个循环里持续改变工作环境:读文件、改文件、跑命令、查文档、调用外部系统。
但一旦模型能改变环境,问题就不再是“给它多少工具”,而是“怎样让它正确地使用工具”。
## 一次失败测试背后的真实链路
假设你打开 Claude Code,说:
```text
登录测试挂了,帮我修一下。
```
看起来这只是一个很小的任务。实际发生的事情会复杂得多。
Claude Code 第一反应通常不是直接改代码,因为它还不知道测试怎么挂,也不知道登录逻辑在哪,更不知道项目用的是哪套测试框架。它需要先看见现场。最直接的方式是跑测试:
```text
npm test -- login
```
如果测试输出里出现:
```text
Expected redirect to /dashboard
Received redirect to /login?next=/dashboard
```
下一步判断就变了。Claude 可能会读测试文件,看断言上下文;再用搜索找 `next=`、`redirect`、`dashboard`;如果项目有 middleware 或 auth guard,它还会继续读这些文件。
也就是说,`Bash` 返回的不是一个“答案”,而是下一轮推理的依据。`Read` 不是单纯读文件,而是在补齐当前任务的事实。`Grep` 不是搜索工具本身,而是在缩小问题空间。`Edit` 不是最终动作,后面还必须有验证。
这条链路大概长这样:
```text
Bash 跑测试
-> 得到失败现场
-> Read 读测试和实现
-> Grep 找调用点
-> Read 读相关代码
-> Edit 修改
-> Bash 再跑测试
-> 根据结果继续调整或收束
```
这就是我说它不是工具箱的原因。工具箱只是把锤子、螺丝刀、扳手放在一起。Claude Code 更像一个工作台:工具、当前材料、操作顺序、检查点和安全边界都被组织在同一个流程里。
## 工具不是越多越好
一旦接入 MCP,这个问题会变得更明显。
本地开发里,Claude Code 常用的内置工具并不算太多:找文件、搜代码、读文件、改文件、跑命令、查网页。它们覆盖了程序员的基本闭环。
但 MCP 会把 Claude Code 接到更大的世界:GitHub、Jira、数据库、监控平台、浏览器、内部 API。每接一个系统,都会多出一批工具。GitHub 有 issue、PR、comment、file、review;Sentry 有 event、issue、release;数据库有 schema、query、migration;公司内部工具又会继续膨胀。
工具越多,问题越不是“模型能不能调”,而是“模型能不能找到当前真正该用的那个工具”。
如果把所有工具定义一股脑塞进上下文,至少有两个坏处。
第一,工具定义会吃掉大量上下文。工具名、描述、参数 schema、返回格式,每一个都要 token。工具一多,上下文先被工具清单塞满,真正的任务材料反而变少。
第二,工具太多会降低选择质量。很多工具名字相近、参数相近、边界相近。模型一次看到太多能力,不一定更聪明,反而更容易选错。
这就是 `ToolSearch` 这类机制出现的位置。
`Grep` 搜的是代码,`WebSearch` 搜的是网页,`ToolSearch` 搜的是能力。它解决的不是“事实在哪里”,而是“当前有没有一个工具能让我做这件事”。
这个区别很关键。Claude Code 不应该在任务开始时背着全部工具定义工作,而应该在需要某类能力时,按需发现工具,再把少量相关工具定义带进上下文。
所以 MCP 和 ToolSearch 的关系可以这样理解:
```text
MCP 解决工具从哪里来
ToolSearch 解决工具太多后怎么找到
```
这对我们自己设计 agent 也很有启发。很多人给 agent 加工具时,只关心“能不能接上 API”。但真正影响效果的,往往是工具能不能被发现、边界是否清楚、返回结果是否高信号。
一个叫 `query` 的万能工具,看似灵活,实际很难被模型稳定使用。一个叫 `search_sentry_events` 的工具,在用户说“查一下最近登录失败的线上报错”时就清楚得多。
工具设计不是后端 API 封装,而是给模型设计行动空间。
## 能执行之前,先要被约束
工具调用请求只是意图,不应该等于执行。
这是 Claude Code 和普通 function calling 最大的差异之一。因为 Claude Code 的工具有真实副作用。
读源码和删文件不是一回事。跑测试和执行部署脚本不是一回事。查数据库和改生产数据库不是一回事。发 Slack、创建 PR、关闭 issue、改配置,这些都不是“生成文本”级别的风险。
所以 Claude Code 需要多层边界。
第一层是权限。哪些工具可以自动执行,哪些要询问用户,哪些应该直接拒绝。`Read` 某个源码文件通常可以自动放行,`npm test` 也可以预先允许;但 `rm -rf`、生产数据库写操作、外部消息发送,就应该被拦住或要求确认。
第二层是 hooks。权限回答“能不能执行”,hooks 回答“执行前后必须发生什么”。比如读取 `.env` 前拦截,改文件后自动格式化,任务结束前跑测试,压缩上下文前归档 transcript。这些规则不应该只写在提示词里,而应该变成运行时机制。
第三层是 sandbox。权限和 hooks 仍然在 Claude Code 层面判断,sandbox 则限制命令执行时到底能访问哪些文件、网络和系统资源。
第四层是 checkpoint。文件改坏了,能不能回退。这里要特别区分:checkpoint 能回滚本地文件修改,但不能回滚外部副作用。如果一个工具已经写了生产数据库、调用了远程 API、发出了消息,本地 checkpoint 并不能让世界恢复原状。
这几层合在一起,才让 Claude Code 从“敢调用工具”变成“可以被放进真实项目里用”。
```text
工具是否可见
-> 调用是否被允许
-> 执行前后是否触发 hook
-> 命令是否被 sandbox 限制
-> 文件修改是否可以 checkpoint 回退
-> 结果是否能被验证
```
如果没有这些边界,工具越多,agent 越危险。
## 工具结果不是日志垃圾桶
工具调用还有一个常被低估的问题:结果怎么回到模型。
如果跑测试只输出三行错误,直接放进上下文没问题。但如果命令输出几万行日志,或者数据库工具返回上万行记录,全部塞回上下文就会污染后续判断。
Agent 的上下文不是垃圾桶。工具结果应该服务下一步决策。
一个工具返回:
```json
{
"rows": [...]
}
```
看起来很完整,实际上可能把判断压力全部扔给模型。另一个工具返回:
```json
{
"summary": "过去 24 小时登录失败集中在 OAuth callback",
"top_errors": [...],
"suggested_next_queries": [...]
}
```
反而更适合 agent,因为它提供的是下一步行动所需的高信号信息。
这也是为什么“把现有 API 包一层给模型用”通常不够。人类 API 的设计目标是给工程师调用,agent 工具的设计目标是帮助模型判断下一步。两者不是一回事。
好的 agent 工具应该有几个特征:
* 名字明确,让模型知道什么时候该用。
* 参数少而清楚,不把业务判断藏在一个万能字符串里。
* 返回结果有摘要、有证据、有下一步提示。
* 对大结果做过滤、分页或聚合。
* 明确说明副作用和权限边界。
Claude Code 给我们的启发不是“多接工具”,而是“工具也要为推理链路设计”。
## 真正值得学的是收束能力
如果只看工具清单,Claude Code 并不神秘。
读文件、搜代码、改文件、跑命令、查网页、接 MCP,这些能力很多 agent 都有。真正差异在于:Claude Code 能不能让一次任务从混乱状态逐步收束到可验证结果。
这件事比工具数量更重要。
一个不会收束的 agent,会一直搜、一直读、一直改,最后把上下文塞满,把项目弄乱。一个能收束的 agent,会不断问:
* 当前最缺的事实是什么?
* 哪个工具能以最低成本拿到这个事实?
* 这一步有没有副作用?
* 结果是否足够支持下一步?
* 什么时候该停止,停止前要验证什么?
这才是 Claude Code 的工具系统最值得借鉴的地方。
它不只是回答“模型怎么调用函数”,而是在回答:
> 当工具越来越多、任务越来越长、副作用越来越真实的时候,怎样让模型仍然只看见该看的东西,只做被允许做的事,并且每一步都能被验证?
这也是我觉得很多 AI 编程工具会分化的地方。
早期大家比的是模型够不够强、能不能写出代码。接下来会越来越多地比 runtime:上下文怎么组织,工具怎么发现,权限怎么约束,结果怎么回流,任务怎么验证,失败怎么回退。
模型能力当然重要,但在真实工程里,模型不是单独工作的。它必须被放进一套能收束的系统里。
Claude Code 真正厉害的,不是会调用工具。
它厉害的是:它把工具调用变成了一个可以持续行动、可以被约束、可以被验证的工作流。
## 参考资料
* [Claude Code Tools reference](https://code.claude.com/docs/en/tools-reference)
* [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)
* [Agent SDK: How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop)
* [Scale to many tools with tool search](https://code.claude.com/docs/en/agent-sdk/tool-search)
* [Connect Claude Code to tools via MCP](https://code.claude.com/docs/en/mcp)
* [Claude Code Permission modes](https://code.claude.com/docs/en/permission-modes)
* [Agent SDK permissions](https://code.claude.com/docs/en/permissions)
* [Claude Code Hooks guide](https://code.claude.com/docs/en/hooks-guide)
* [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing)
* [Introducing advanced tool use on the Claude Developer Platform](https://www.anthropic.com/engineering/advanced-tool-use)
* [Writing effective tools for AI agents](https://www.anthropic.com/engineering/writing-tools-for-agents)
# Guide complet de Claude Worktree
## Introduction
Lorsque vous travaillez sur des tâches complexes avec Claude Code, vous avez peut-être rencontré ce dilemme : vous avez trois tâches indépendantes à traiter, mais exécuter plusieurs instances de Claude dans le même répertoire provoque des conflits de code — un Agent modifie des fichiers pendant qu'un autre touche aux mêmes, et le résultat lors de la fusion est un vrai désordre.
En février 2025, Anthropic a publié la commande `--worktree`, changeant complètement la donne. Désormais, vous pouvez exécuter `claude -w feature-1`, `claude -w feature-2` et `claude -w bugfix-1` dans trois terminaux distincts, chaque Agent travaillant dans son propre environnement isolé sans interférer avec les autres.
## Comprendre Worktree
Imaginez que vous êtes un architecte concevant trois pièces différentes simultanément. L'approche traditionnelle consiste à dessiner sur le même plan, ce qui devient confus avec les modifications constantes. Worktree vous donne trois plans distincts, chacun dédié à la conception d'une pièce, pour être ensuite fusionnés dans le plan principal.
Techniquement, Worktree est une fonctionnalité native de Git. La commande `--worktree` de Claude Code encapsule cette fonctionnalité pour la rendre beaucoup plus simple — une seule commande crée un environnement isolé, lance une instance de Claude et nettoie automatiquement une fois terminé.
### Pourquoi ne pas simplement cloner plusieurs fois
Vous pourriez demander : pourquoi ne pas simplement cloner le dépôt plusieurs fois ?
| Approche | Utilisation disque | Difficulté de synchronisation | Complexité de nettoyage |
| ------------------- | -------------------------------------------- | ---------------------------------- | ---------------------------------- |
| Clones multiples | Dépôt complet pour chaque copie | Pull/push manuels nécessaires | Suppression manuelle du répertoire |
| Git Worktree | Fichiers de travail uniquement, .git partagé | Historique partagé automatiquement | `git worktree remove` |
| Claude `--worktree` | Fichiers de travail uniquement, .git partagé | Historique partagé automatiquement | Nettoyage automatique à la sortie |
Les worktrees partagent la même base de données `.git` — tout l'historique des commits et les informations de branches sont partagés. Cela signifie qu'un commit créé dans un worktree est immédiatement visible dans tous les autres.
### Quand utiliser Worktree
Avant de vous lancer, évaluez si votre tâche se prête bien à l'utilisation d'un worktree.
Règle générale : **si une tâche nécessite plus de 30 minutes, envisagez d'utiliser un worktree**. Pour les tâches courtes, le worktree est un gaspillage de temps — créer l'environnement, installer les dépendances, puis fusionner peut prendre plus longtemps que la tâche elle-même. Mais pour les tâches nécessitant un travail approfondi, l'isolation du worktree devient très précieuse.
| Adapté au Worktree | Peu adapté |
| --------------------------------------------------------- | -------------------------------------------------- |
| Développement de fonctionnalités indépendantes | Petites modifications réalisables en 10 minutes |
| Refactorisation parallèle de modules différents | Tâches nécessitant des interactions fréquentes |
| Tâches de longue durée | Forte dépendance à d'autres modifications en cours |
| Modifications expérimentales nécessitant des tests isolés | Corrections de bugs simples |
### Prérequis
Avant d'utiliser worktree, assurez-vous que les conditions suivantes sont remplies :
| Condition | Description |
| --------------------------- | ------------------------------------------------------------------------------ |
| Git initialisé | Vous devez être dans un dépôt Git (avec un répertoire `.git`) |
| Au moins un commit | Les dépôts vides ne peuvent pas créer de worktrees |
| Branche distante disponible | Par défaut, le code est extrait depuis la branche distante (ex. `origin/main`) |
## Flux de travail complet : de la création au nettoyage
Voici le flux de travail complet, de la création d'un Worktree jusqu'au nettoyage final, dans l'ordre réel de développement.
### Étape 1 : Créer un Worktree
#### Création depuis la branche distante par défaut
Utilisez le paramètre `-w` ou `--worktree` pour lancer Claude :
```bash
# Créer un worktree nommé "feature-auth" et lancer Claude
claude -w feature-auth
# Générer automatiquement un nom aléatoire (ex. "bright-running-fox")
claude -w
```
Cette commande effectue en réalité quatre opérations :
1. Crée un nouveau répertoire de travail dans `/.claude/worktrees/feature-auth/`
2. Crée une nouvelle branche nommée `worktree-feature-auth`
3. Extrait le code depuis la branche distante par défaut (ex. `origin/main` ou `origin/master`) — **attention, ce n'est pas votre branche actuelle**
4. Lance Claude Code dans le nouveau répertoire
Tous les worktrees se trouvent dans le répertoire `.claude/worktrees/` :
```
votre-projet/
├── .claude/
│ └── worktrees/
│ ├── feature-auth/ ← Premier worktree
│ ├── bugfix-123/ ← Deuxième worktree
│ └── refactor-api/ ← Troisième worktree
├── src/
└── package.json
```
Il est recommandé d'ajouter ce chemin au `.gitignore` :
```bash
# .gitignore
.claude/worktrees/
```
#### Création depuis la branche actuelle/spécifique
`-w` extrait toujours depuis la branche distante par défaut et ne permet pas actuellement de spécifier une branche de base. Si vous souhaitez créer un worktree basé sur votre branche actuelle (ou une branche spécifique), trois approches sont possibles :
**Approche 1 : Création manuelle avec Git**
```bash
# Créer un worktree basé sur le HEAD actuel
git worktree add -b my-feature .claude/worktrees/my-feature HEAD
# Ou basé sur une branche spécifique
git worktree add -b hotfix .claude/worktrees/hotfix origin/release-v2
# Puis lancer Claude dans ce répertoire
cd .claude/worktrees/my-feature && claude
```
Cette approche vous donne un contrôle total — vous pouvez créer des worktrees à partir de n'importe quelle branche ou commit, et le répertoire de travail démarre sur la bonne branche dès le début. La documentation officielle recommande également : « Pour un contrôle accru sur la branche et l'emplacement, créez le worktree directement avec Git, puis exécutez Claude dans ce répertoire. »
**Approche 2 : Création dans une session (recommandée)**
Dans une session Claude existante, demandez simplement à Claude de créer un worktree :
```
> Crée un worktree depuis la branche actuelle
> start a worktree
```
Contrairement à la commande `-w`, les worktrees créés en session sont **automatiquement basés sur la branche actuelle**, et non sur la branche distante par défaut. Claude gère automatiquement la création du worktree et bascule dessus — aucune commande Git manuelle n'est nécessaire. Si vous travaillez déjà sur une branche feature, c'est l'approche la plus pratique — une simple phrase vous donne un environnement isolé basé sur votre branche actuelle.
**Approche 3 : Créer avec `-w` d'abord, changer de branche en session**
Créez d'abord un worktree avec `claude -w`, puis demandez à Claude de basculer vers la branche cible dans la session. L'inconvénient est qu'il extrait d'abord depuis la branche distante par défaut, puis bascule — une étape supplémentaire, moins propre que les deux premières approches. Et si la branche cible est déjà occupée par un autre worktree, vous rencontrerez un conflit de branche :
Comme le montre la capture d'écran, Claude détecte le conflit de branche et propose deux choix : revenir au répertoire principal, ou créer une nouvelle branche de travail basée sur la branche cible. Bien que cela fonctionne au final, le processus est moins direct que les approches 1 et 2.
**Avancé : Encapsulation avec Makefile pour des commandes en un clic**
Si vous avez fréquemment besoin de créer des worktrees depuis la branche actuelle, vous pouvez ajouter une commande raccourci dans le `Makefile` à la racine du projet, enchaînant création + ouverture de l'éditeur + lancement de Claude en un seul pipeline déterministe :
```makefile
# Créer un worktree depuis la branche actuelle et lancer l'environnement de développement
# Utilisation : make worktree name=my-feature
worktree:
@if [ -z "$(name)" ]; then \
echo "Utilisation : make worktree name="; \
echo "Exemple : make worktree name=fix-login-bug"; \
exit 1; \
fi
@echo "→ Création du worktree depuis $$(git branch --show-current) : $(name)"
git worktree add -b $(name) .claude/worktrees/$(name) HEAD
@echo "→ Initialisation de l'environnement"
cd .claude/worktrees/$(name) && npm install
@echo "→ Ouverture dans Zed"
zed .claude/worktrees/$(name)
@echo "→ Lancement de Claude"
cd .claude/worktrees/$(name) && claude
```
L'utilisation est très simple :
```bash
# Créer un worktree depuis la branche actuelle, ouvrir dans Zed, lancer Claude
make worktree name=fix-login-bug
# Lancer plusieurs tâches en parallèle
make worktree name=feature-search
make worktree name=refactor-api
```
Par rapport à la saisie manuelle de plusieurs commandes Git/cd/claude, `make worktree name=xxx` ne nécessite qu'une seule ligne, et le processus est exactement le même à chaque fois — pas d'oubli d'étape ni d'erreur de chemin. Notez que le Makefile utilisant `git worktree add` natif, le Hook `WorktreeCreate` de Claude Code ne se déclenchera pas (ce Hook ne s'active qu'avec `claude -w` ou lors de la création de worktree en session). Les étapes d'initialisation de l'environnement (installation des dépendances, copie du `.env`, etc.) doivent donc être écrites directement dans le Makefile, comme dans l'exemple `npm install` ci-dessus.
### Étape 2 : Initialiser l'environnement
Une fois le worktree créé, la première chose à faire est d'initialiser l'environnement de développement. Chaque nouveau worktree est un répertoire indépendant — `node_modules`, les environnements virtuels, les fichiers `.env`, etc. ne sont pas transférés automatiquement.
Claude Code fournit le Hook `WorktreeCreate` pour automatiser la configuration de l'environnement :
```json
{
"hooks": {
"WorktreeCreate": [
{
"command": "npm install && cp ../.env .env"
}
]
}
}
```
Ainsi, à chaque création de worktree, les dépendances sont automatiquement installées et le fichier de variables d'environnement est automatiquement copié. Étapes d'initialisation courantes :
| Type de projet | Commande d'initialisation |
| -------------- | -------------------------------------------------------------------------- |
| Node.js | `npm install` ou `yarn` |
| Python | `pip install -r requirements.txt` ou activation de l'environnement virtuel |
| Go | `go mod download` |
| Général | Copie des fichiers `.env`, configuration des variables d'environnement |
Si les Hooks ne sont pas configurés, vous pouvez également exécuter `/init` au début de chaque session worktree pour vous assurer que Claude comprend correctement le contexte du répertoire de travail actuel et relit la structure du projet et la configuration CLAUDE.md.
### Étape 3 : Commit et fusion
Une fois l'environnement prêt et le développement terminé, l'étape suivante est de fusionner les modifications vers la branche cible.
**Fusion vers la branche main**
Le cas le plus courant — le worktree a été créé depuis `origin/main`, et les modifications doivent être fusionnées vers `main`. Dans la session Claude du worktree, dites simplement :
```
> Commite toutes les modifications, pousse vers le dépôt distant, puis crée une PR vers main
```
Claude gère automatiquement l'ensemble du flux commit → push → `gh pr create`.
**Fusion vers une branche feature**
Si vous développez sur la branche `feature-x` et que les modifications du worktree doivent être fusionnées vers `feature-x` plutôt que `main` :
```
> Commite et pousse les modifications, puis crée une PR ciblant la branche feature-x
```
Claude exécutera `gh pr create --base feature-x`, créant directement une PR vers la branche feature.
Vous pouvez également quitter la session worktree (en choisissant de conserver le worktree), puis revenir au répertoire principal et lancer Claude :
```
> Fusionne les modifications de la branche worktree-my-task dans la branche actuelle
```
Si certains commits du worktree ne vous intéressent pas, vous pouvez les sélectionner avec cherry-pick :
```
> Montre-moi l'historique des commits de la branche worktree-my-task, puis cherry-pick les commits liés au module d'authentification vers la branche actuelle
```
> **Astuce** : Tous les worktrees partagent la même base de données `.git` — les commits créés dans un worktree sont immédiatement visibles dans le répertoire principal, sans nécessiter d'opérations push/pull supplémentaires.
### Étape 4 : Sortie et nettoyage
Une fois les modifications fusionnées, vous pouvez quitter la session worktree.
Lors de la sortie d'une session worktree, Claude gère automatiquement les choses selon l'état :
| État | Action |
| ------------------------------------- | -------------------------------------------------- |
| **Aucune modification** | Supprime automatiquement le worktree et la branche |
| **Modifications ou commits présents** | Vous invite à choisir entre conserver ou supprimer |
Les worktrees conservés persistent pour que vous puissiez continuer à travailler dessus ultérieurement.
Vous pouvez également configurer un Hook `WorktreeRemove` pour automatiser le nettoyage :
```json
{
"hooks": {
"WorktreeRemove": [
{
"command": "echo 'Worktree cleaned up'"
}
]
}
}
```
**Commandes de gestion manuelle**
Si vous devez gérer manuellement les worktrees, utilisez les commandes Git standard :
```bash
# Lister tous les worktrees
git worktree list
# Supprimer manuellement un worktree
git worktree remove .claude/worktrees/feature-auth
# Nettoyer les références de worktree obsolètes
git worktree prune
```
> **Attention** : Ne supprimez pas directement un répertoire worktree avec `rm -rf`. La bonne méthode est d'utiliser `git worktree remove`, ou si vous l'avez déjà supprimé par erreur, exécutez `git worktree prune` pour nettoyer les références résiduelles.
## Modes de développement parallèle
Maintenant que vous maîtrisez le flux de travail de base, voyons comment exploiter les worktrees pour le développement parallèle.
### Parallélisme multi-terminaux
L'utilisation la plus courante est l'exécution simultanée dans plusieurs onglets de terminal :
```bash
# Terminal 1 : Travailler sur l'authentification utilisateur
claude -w feature-auth
# Terminal 2 : Corriger un bug de paiement
claude -w bugfix-payment
# Terminal 3 : Refactoriser le module API
claude -w refactor-api
```
Chaque instance de Claude travaille dans son propre worktree sans interférer avec les autres. Vous pouvez :
* Laisser Claude développer une nouvelle fonctionnalité dans un terminal
* Laisser Claude corriger un bug dans un autre terminal
* Continuer votre propre revue de code dans un troisième terminal
### Implémentation compétitive
Une utilisation efficace consiste à faire implémenter la même fonctionnalité indépendamment par plusieurs Agents :
```bash
# Trois terminaux exécutés séparément
claude -w feature-search-v1
claude -w feature-search-v2
claude -w feature-search-v3
```
Donnez-leur les mêmes spécifications et laissez chacun implémenter de manière indépendante. Ensuite, comparez les trois solutions et fusionnez la meilleure. Cela exploite le non-déterminisme des LLM — la même entrée peut produire des résultats différents, et parfois la deuxième version est meilleure.
L'exploration de design UI est également très adaptée à ce mode. Supposons que vous souhaitiez repenser l'interface de votre application, mais que vous ne sachiez pas quel style convient le mieux :
```bash
# Trois Agents implémentent différents styles
claude -w ui-minimal # Style minimaliste
claude -w ui-colorful # Couleurs vives
claude -w ui-glassmorphism # Style glassmorphisme
```
Une fois terminé, exécutez simultanément les trois serveurs de développement (sur des ports différents), comparez côte à côte et fusionnez votre solution préférée dans la branche principale — bien plus efficace que le cycle traditionnel « construire une version, examiner, reconstruire ».
### Isolation de Subagent
Les worktrees ne sont pas réservés à l'instance Claude principale — ils fonctionnent également avec les Subagents. Ajoutez `isolation: worktree` dans le frontmatter de votre Subagent personnalisé :
```yaml
---
name: code-migrator
description: Handles large-scale code migrations. Use for batch refactoring.
tools: Read, Write, Edit, Bash, Grep, Glob
isolation: worktree
---
You are a code migration specialist...
```
Vous pouvez également le dire directement à Claude dans la conversation :
```
> utilise des worktrees pour tes agents
> use worktrees for your agents
```
Lorsqu'un Subagent est configuré avec l'isolation worktree :
```
Agent principal (répertoire principal)
│
├── Lancement Agent Migration 1 ──→ worktree-migration-1/
│ └── Traitement de src/auth/
│
├── Lancement Agent Migration 2 ──→ worktree-migration-2/
│ └── Traitement de src/api/
│
└── Lancement Agent Migration 3 ──→ worktree-migration-3/
└── Traitement de src/utils/
```
Chaque Subagent travaille indépendamment dans son propre worktree sans interférence. Une fois terminé, le worktree est automatiquement nettoyé (s'il n'y a pas de modifications non commitées).
### Intégration avec Tmux et l'IDE
Le paramètre `--tmux` lance automatiquement Claude dans une nouvelle session Tmux, pour qu'il continue de fonctionner même si vous fermez le terminal :
```bash
claude -w feature-auth --tmux
```
Si vous utilisez VS Code ou Cursor, le panneau de contrôle de code source reconnaît automatiquement tous les worktrees — le dépôt principal apparaît comme un repo, chaque worktree comme un repo indépendant, et vous pouvez basculer, commiter et pousser directement depuis l'IDE. Les worktrees peuvent également être combinés avec les [boucles Ralph](/fr/docs/notes/ralph-wiggum/concept) — chaque boucle Ralph s'exécute dans son propre worktree, de sorte que même une boucle échouée n'affecte pas la branche principale.
## Conseils et bonnes pratiques
### Pièges courants
1. **Confusion sur l'origine de la branche** : `-w` crée des worktrees depuis la **branche distante par défaut**, pas votre branche actuelle. Si vous exécutez `claude -w my-task` alors que vous êtes sur la branche `feature-x`, le code du nouveau worktree provient de `origin/main` et n'inclura pas les modifications de `feature-x`. Pour travailler depuis votre branche actuelle, consultez [Création depuis la branche actuelle/spécifique](#création-depuis-la-branche-actuellespécifique).
2. **Les modifications non commitées ne sont pas transférées** : Lors de la création d'un worktree, les modifications non indexées ou non commitées du répertoire principal n'apparaîtront pas dans le nouveau worktree. Les worktrees sont créés uniquement à partir de l'historique des commits, alors assurez-vous de commiter les modifications importantes d'abord.
3. **Une même branche ne peut pas être utilisée par plusieurs worktrees** : Git ne permet pas à deux worktrees d'extraire simultanément la même branche. Si votre répertoire principal est déjà sur `feature-x`, tenter d'extraire `feature-x` dans un worktree échouera. Chaque worktree doit être sur une branche différente.
4. **L'environnement nécessite une réinitialisation** : Chaque nouveau worktree n'inclut pas les dépendances d'exécution comme `node_modules`. Configurez un Hook `WorktreeCreate` pour automatiser cette étape (voir [Étape 2 : Initialiser l'environnement](#étape-2--initialiser-lenvironnement)).
### Recommandations d'utilisation
N'en abusez pas. Bien que techniquement vous puissiez ouvrir de nombreux worktrees, chaque instance de Claude consomme des crédits API, trop de tâches parallèles deviennent difficiles à suivre, et les conflits de fusion se complexifient.
**Conventions de nommage** : Adoptez de bonnes habitudes de nommage pour faciliter la gestion :
```bash
# Bon nommage
claude -w feature-user-auth
claude -w bugfix-payment-123
claude -w refactor-api-v2
# Mauvais nommage
claude -w test
claude -w temp
claude -w 1
```
## Systèmes de versionnement non-Git
Si vous utilisez SVN, Perforce ou Mercurial, vous pouvez obtenir un effet d'isolation similaire en configurant les Hooks `WorktreeCreate` et `WorktreeRemove`. Une fois configurés, l'utilisation de `--worktree` invoquera vos commandes personnalisées au lieu du comportement Git par défaut.
## Conclusion
Worktree est une fonctionnalité que l'équipe Claude Code utilise au quotidien — Boris Cherny la qualifie de « conseil de productivité numéro un ». La valeur fondamentale est simple : **permettre à plusieurs Agents de travailler en parallèle sans interférer les uns avec les autres**.
Pour commencer, c'est très simple :
```bash
claude -w your-task-name
```
***
**Lectures complémentaires** :
* [Guide complet de Claude Subagent](/fr/docs/notes/claude-subagent) — Comprendre l'intégration entre Subagent et Worktree
* [Analyse approfondie de Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept) — Une autre approche pour améliorer l'efficacité de la programmation IA
* [Architecture système de Claude expliquée](/fr/docs/notes/claude-architecture) — Comprendre la place de Worktree dans l'architecture globale
**Références** :
* [Documentation officielle Claude Code - Common Workflows](https://code.claude.com/docs/en/common-workflows)
* [Annonce Worktree de Boris Cherny](https://www.threads.com/@boris_cherny/post/DVAAnexgRUj)
* [Documentation officielle Git Worktree](https://git-scm.com/docs/git-worktree)
* [incident.io - Shipping faster with Claude Code and Git Worktrees](https://incident.io/blog/shipping-faster-with-claude-code-and-git-worktrees)
* [Dev.to - Git worktree + Claude Code: My Secret to 10x Developer Productivity](https://dev.to/kevinz103/git-worktree-claude-code-my-secret-to-10x-developer-productivity-520b)
**Tutoriels vidéo** :
* [I'm using claude --worktree for everything now](https://www.youtube.com/watch?v=yv8VZpov8bk) — Démonstration complète du flux de travail worktree
* [Git Worktrees: The secret sauce to Claude Code!](https://www.youtube.com/watch?v=up91rbPEdVc) — Méthodes de création manuelle de worktree
* [Native Worktrees Just Killed Traditional Claude Code Workflows](https://www.youtube.com/watch?v=lj_xZn-Yf18) — Présentation détaillée de la fonctionnalité worktree native
* [Claude Code Worktrees in 7 Minutes](https://www.youtube.com/watch?v=z_VI51k-tn0) — Tutoriel de prise en main rapide, incluant l'utilisation avec Subagent
* [Stop Using Claude Code on One Branch](https://www.youtube.com/watch?v=6nFRJftouI0) — Avantages du développement parallèle multi-worktree
# Documentation
Bienvenue dans le centre de documentation. Vous trouverez ici les documents techniques et tutoriels que j'ai rassemblés sur Claude Code.
# Guide de prise en main rapide de Tmux
## Introduction
Si vous avez deja utilise les Agent Teams de Claude Code ou si vous souhaitez executer plusieurs instances de Claude simultanement, Tmux est un outil pratiquement indispensable. Il vous permet d'executer plusieurs sessions dans une seule fenetre de terminal, les sessions continuent de fonctionner en arriere-plan meme si vous fermez le terminal, et Claude peut automatiquement creer et gerer plusieurs Agents dans Tmux.
Ce tutoriel est concu specialement pour les utilisateurs de Claude Code. Il couvre a la fois les bases de Tmux et les techniques d'integration avec Claude Code.
## Comprendre Tmux
Les trois concepts fondamentaux de Tmux :
```
┌─────────────────────────────────────────────────────────┐
│ Session (session) │
│ ┌─────────────────────────────────────────────────────┐│
│ │ Window (fenetre) ││
│ │ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ ││
│ │ │ Pane 1 │ │ Pane 2 │ │ Pane 3 │ ││
│ │ │ (Claude 1) │ │ (Claude 2) │ │ (Logs) │ ││
│ │ │ │ │ │ │ │ ││
│ │ │ │ │ │ │ │ ││
│ │ └──────────────┘ └──────────────┘ └───────────┘ ││
│ └─────────────────────────────────────────────────────┘│
│ Window 1: Development Window 2: Testing │
└─────────────────────────────────────────────────────────┘
```
| Concept | Analogie | Description |
| ----------- | -------------------- | --------------------------------------------------------------------------------- |
| **Session** | Espace de travail | Conteneur de niveau superieur, continue de fonctionner meme en cas de deconnexion |
| **Window** | Onglet de navigateur | Une session peut contenir plusieurs fenetres |
| **Pane** | Ecran divise | Une fenetre peut etre divisee en plusieurs volets |
## Installation et bases
### Installer Tmux
```bash
# macOS
brew install tmux
# Ubuntu/Debian/WSL
sudo apt-get install tmux
# CentOS/RHEL
sudo yum install tmux
```
Verifier l'installation :
```bash
tmux -V
# Affiche par exemple : tmux 3.6a
```
### La touche prefixe
Toutes les commandes Tmux commencent par une **touche prefixe**, qui est `Ctrl+B` par defaut.
Pour saisir une commande :
1. Appuyez sur `Ctrl+B` (maintenez enfonce)
2. Relâchez, puis appuyez sur la touche de commande
Par exemple, pour diviser la fenetre : `Ctrl+B` puis `%`
## Aide-memoire des commandes courantes
### Gestion des sessions
| Commande | Description |
| --------------------------- | -------------------------------------------------------- |
| `tmux` | Creer une nouvelle session |
| `tmux new -s name` | Creer une session nommee |
| `tmux ls` | Lister toutes les sessions |
| `tmux attach -t name` | Se connecter a une session |
| `tmux kill-session -t name` | Fermer une session |
| `Ctrl+B d` | Detacher la session courante (execution en arriere-plan) |
### Gestion des fenetres
| Raccourci | Description |
| ------------ | ---------------------------------- |
| `Ctrl+B c` | Creer une nouvelle fenetre |
| `Ctrl+B n` | Fenetre suivante |
| `Ctrl+B p` | Fenetre precedente |
| `Ctrl+B 0-9` | Basculer vers la fenetre specifiee |
| `Ctrl+B ,` | Renommer la fenetre courante |
| `Ctrl+B &` | Fermer la fenetre courante |
### Gestion des volets
| Raccourci | Description |
| -------------------------------- | ---------------------------------- |
| `Ctrl+B %` | Division verticale (gauche/droite) |
| `Ctrl+B "` | Division horizontale (haut/bas) |
| `Ctrl+B touches directionnelles` | Se deplacer entre les volets |
| `Ctrl+B x` | Fermer le volet courant |
| `Ctrl+B z` | Maximiser/restaurer un volet |
| `Ctrl+B {` | Deplacer le volet vers la gauche |
| `Ctrl+B }` | Deplacer le volet vers la droite |
### Autres commandes utiles
| Raccourci | Description |
| ---------- | ------------------------------------------ |
| `Ctrl+B [` | Entrer en mode copie (defilement possible) |
| `q` | Quitter le mode copie |
| `Ctrl+B ?` | Afficher tous les raccourcis |
## Integration avec Claude Code
### Pourquoi Claude Code a besoin de Tmux
1. **Mode split-pane des Agent Teams** : chaque Teammate s'affiche dans un volet independant
2. **Execution en arriere-plan** : les tâches continuent meme si vous fermez le terminal
3. **Persistance des sessions** : contexte complet restaure apres une reconnexion
4. **Gestion multi-instances** : executez plusieurs sessions Claude simultanement
### Utilisation de base : executer Claude en arriere-plan
```bash
# Lancer Claude dans tmux
tmux new -s claude-work
claude
# Detacher la session (Claude continue de fonctionner)
# Ctrl+B d
# Se reconnecter plus tard
tmux attach -t claude-work
```
### Utilisation du parametre --tmux
Claude Code prend nativement en charge l'integration Tmux :
```bash
# Lancer Claude dans une nouvelle session tmux
claude --tmux
# Utiliser avec worktree
claude -w feature-auth --tmux
```
Cela effectue automatiquement :
1. La creation d'une nouvelle session tmux
2. Le lancement de Claude Code dans cette session
3. Le nommage de la session sous la forme `claude-{ID-aleatoire}`
### Mode Tmux des Agent Teams
Les Agent Teams peuvent utiliser le mode d'affichage split-pane, avec chaque Teammate dans un volet independant :
```json
// settings.json
{
"teammateMode": "tmux"
}
```
Ou via la ligne de commande :
```bash
claude --teammate-mode tmux
```
Resultat :
```
┌─────────────────────────────────────────────────────────┐
│ Team Lead │
├─────────────────┬─────────────────┬─────────────────────┤
│ Teammate 1 │ Teammate 2 │ Teammate 3 │
│ Security │ Performance │ Testing │
│ │ │ │
└─────────────────┴─────────────────┴─────────────────────┘
```
## Configuration pratique
### Fichier \~/.tmux.conf recommande
Creez ou editez `~/.tmux.conf` :
```bash
# Utiliser Ctrl+A comme touche prefixe (plus accessible)
set -g prefix C-a
unbind C-b
bind C-a send-prefix
# Activer le support de la souris
set -g mouse on
# Augmenter le tampon d'historique (Claude produit beaucoup de sortie)
set -g history-limit 50000
# Navigation entre volets en style vim
bind h select-pane -L
bind j select-pane -D
bind k select-pane -U
bind l select-pane -R
# Raccourcis de division plus intuitifs
bind | split-window -h
bind - split-window -v
# Rechargement rapide de la configuration
bind r source-file ~/.tmux.conf \; display "Config reloaded!"
# Support 256 couleurs
set -g default-terminal "screen-256color"
set -ga terminal-overrides ",*256col*:Tc"
# Numerotation des fenetres a partir de 1 (0 est trop eloigne)
set -g base-index 1
setw -g pane-base-index 1
# Optimisation de la barre d'etat
set -g status-position bottom
set -g status-left-length 40
set -g status-right-length 60
```
Recharger la configuration :
```bash
tmux source-file ~/.tmux.conf
```
### Configuration dediee a Claude Code
Configuration optimisee pour Claude Code :
```bash
# Raccourci pour ouvrir Claude en popup
bind -r y run-shell '\
SESSION="claude-$(echo #{pane_current_path} | md5sum | cut -c1-8)"; \
tmux has-session -t "$SESSION" 2>/dev/null || \
tmux new-session -d -s "$SESSION" -c "#{pane_current_path}" "claude"; \
tmux display-popup -w80% -h80% -E "tmux attach-session -t $SESSION"'
```
Cette configuration produit l'effet suivant :
1. Appuyez sur `Ctrl+A y` pour ouvrir une fenetre popup Claude
2. Chaque repertoire dispose de sa propre session Claude
3. La session continue de fonctionner apres la fermeture de la popup
4. La conversation precedente est restauree a la reouverture
## Workflows courants
### Workflow 1 : projets en parallele
```bash
# Creer une session independante pour chaque projet
tmux new -s project-a
# Lancer Claude dedans
claude -w feature-x
# Detacher, puis creer une autre session
# Ctrl+B d
tmux new -s project-b
claude -w bugfix-y
# Basculer entre les sessions
tmux switch -t project-a
tmux switch -t project-b
# Ou lister toutes les sessions pour choisir
# Ctrl+B s
```
### Workflow 2 : tableau de bord de developpement
Creer un environnement de developpement multi-volets :
```bash
# Creer une session
tmux new -s dev
# Diviser en trois volets
# Ctrl+B % (division verticale)
# Ctrl+B " (division horizontale du côte droit)
# Disposition des volets :
# ┌───────────┬───────────┐
# │ Claude │ Logs │
# │ ├───────────┤
# │ │ Tests │
# └───────────┴───────────┘
# Executer Claude dans le premier volet
claude
# Basculer vers le deuxieme volet (Ctrl+B fleche droite)
tail -f logs/app.log
# Basculer vers le troisieme volet
npm test -- --watch
```
### Workflow 3 : developpement a distance
La fonctionnalite la plus puissante de Tmux est la persistance des sessions, particulierement adaptee au developpement a distance via SSH :
```bash
# Se connecter au serveur distant
ssh user@server
# Creer une session tmux
tmux new -s remote-claude
# Lancer Claude
claude
# Se deconnecter de SSH (Claude continue de fonctionner)
# Ctrl+B d
exit
# Se reconnecter plus tard
ssh user@server
tmux attach -t remote-claude
# La session Claude est entierement restauree
```
### Workflow 4 : supervision des Agent Teams
Utiliser tmux pour superviser tous les Teammates d'un Agent Teams :
```bash
# Lancer Claude en mode tmux
claude --teammate-mode tmux
# Creer un Agent Team
# "Creer un agent team pour auditer le code..."
# L'ecran se divise automatiquement, un volet par Teammate
# Vous pouvez cliquer sur differents volets pour communiquer directement avec le Teammate correspondant
```
## Depannage
### Problemes courants
| Probleme | Solution |
| --------------------------- | ---------------------------------------------------------------- |
| Couleurs mal affichees | Verifiez que `TERM=xterm-256color` |
| La souris ne fonctionne pas | Ajoutez `set -g mouse on` a la configuration |
| Problemes de copier-coller | Utilisez `Enter` pour copier en mode copie |
| Session disparue | Verifiez avec `tmux ls`, il peut s'agir d'un redemarrage systeme |
### Nettoyer les sessions orphelines
Claude Code peut parfois laisser des sessions tmux non nettoyees :
```bash
# Lister toutes les sessions
tmux ls
# Fermer une session specifique
tmux kill-session -t session-name
# Fermer toutes les sessions (attention !)
tmux kill-server
```
### Utilisateurs d'iTerm2
Si vous utilisez iTerm2 sous macOS, vous pouvez utiliser son integration native :
```bash
# Utiliser le mode d'integration tmux d'iTerm2
tmux -CC
# Ou dans Claude Code
claude --teammate-mode tmux
```
iTerm2 convertit automatiquement les volets tmux en onglets et ecrans divises natifs.
## Mon retour d'experience
### Quand utiliser Tmux
| Scenario | Tmux necessaire ? |
| --------------------------------------------- | -------------------- |
| Conversation simple et ponctuelle avec Claude | Non |
| Tâches de longue duree | Oui |
| Agent Teams | Fortement recommande |
| Developpement a distance | Indispensable |
| Projets en parallele | Recommande |
### Configuration minimale
Si vous ne souhaitez pas vous compliquer avec la configuration, retenez simplement ces commandes :
```bash
# Creer une session
tmux new -s work
# Detacher (execution en arriere-plan)
Ctrl+B d
# Se reconnecter
tmux attach -t work
# Diviser les volets
Ctrl+B % # division gauche/droite
Ctrl+B " # division haut/bas
# Changer de volet
Ctrl+B touches directionnelles
```
### Les meilleures combinaisons avec Claude Code
1. **Worktree + Tmux** : chaque worktree dans une session tmux independante
```bash
claude -w feature-auth --tmux
```
2. **Agent Teams + Tmux** : gestion visuelle de tous les Teammates
```bash
claude --teammate-mode tmux
```
3. **Tâches longues + detachement** : lancez puis detachez, revenez verifier plus tard
```bash
# Lancer
tmux new -s migration
claude
# "Executer la migration de base de donnees..."
# Ctrl+B d
# Quelques heures plus tard
tmux attach -t migration
```
## En conclusion
Tmux est un outil essentiel pour une utilisation efficace de Claude Code, en particulier dans les scenarios suivants :
| Point cle | Description |
| ----------------- | ------------------------------------------------------ |
| **Persistance** | Les sessions ne sont pas perdues en cas de deconnexion |
| **Parallelisme** | Gerez plusieurs instances Claude simultanement |
| **Visualisation** | Affichage split-pane des Agent Teams |
Trois commandes de base suffisent pour commencer :
* `tmux new -s name` pour creer une session
* `Ctrl+B d` pour detacher une session
* `tmux attach -t name` pour se reconnecter
***
**Lectures complementaires** :
* [Guide complet de Claude Agent Teams](/fr/docs/notes/claude-agent-teams) — Les Agent Teams necessitent Tmux pour le mode split-pane
* [Guide complet de Claude Worktree](/fr/docs/notes/claude-worktree) — Worktree peut fonctionner en arriere-plan avec Tmux
**Ressources** :
* [Tmux Wiki - Getting Started](https://github.com/tmux/tmux/wiki/Getting-Started)
* [A Quick and Easy Guide to tmux](https://hamvocke.com/blog/a-quick-and-easy-guide-to-tmux/)
* [Red Hat - A beginner's guide to tmux](https://www.redhat.com/en/blog/introduction-tmux-linux)
* [How to run Claude Code in a Tmux popup window](https://www.devas.life/how-to-run-claude-code-in-a-tmux-popup-window-with-persistent-sessions/)
* [Claude Code + tmux: The Ultimate Terminal Workflow](https://www.blle.co/blog/claude-code-tmux-beautiful-terminal)
**Tutoriels video** :
* [Tmux Basics Tutorial](https://www.youtube.com/watch?v=UgHbHqg_Wmo) — Introduction aux bases de Tmux
* [Tmux + Claude Code Workflow](https://www.youtube.com/watch?v=vtB1J_zCv8I) — Workflow d'integration avec Claude Code
* [Advanced Tmux Configuration](https://www.youtube.com/watch?v=yHQym4LF1yA) — Techniques de configuration avancee
# Sprint MVP : terminer les fonctionnalités essentielles en deux semaines
Ceci est un article de test.
# S'inscrire au programme Apple Developer
Pour publier votre application sur l'App Store, la première étape consiste à vous inscrire à l'Apple Developer Program (programme développeur Apple). Cela représente un investissement annuel de ¥688 (99 $), un passage obligé pour tout développeur iOS indépendant.
Cet article vous présente les différents types de comptes, les préparatifs nécessaires avant l'inscription ainsi que le processus d'inscription complet.
## Comparaison des types de comptes
L'Apple Developer Program propose trois types de comptes, adaptés à différents scénarios de développement :
| Caractéristique | Compte individuel | Compte organisation | Compte entreprise |
| ---------------------------- | --------------------------------------- | ----------------------------------------------------------- | -------------------------------------------- |
| Cotisation annuelle | ¥688 ($99) | ¥688 ($99) | ¥1 988 ($299) |
| Publication sur l'App Store | ✅ | ✅ | ❌ (distribution interne uniquement) |
| Nom du développeur affiché | Nom personnel | Nom de l'organisation/société | Nom de l'organisation |
| Gestion des membres d'équipe | ❌ | ✅ | ✅ |
| Numéro D-U-N-S | Non requis | Requis | Requis |
| Délai de vérification | Rapide (généralement sous 48 heures) | Plus long (vérification des informations de l'organisation) | Plus long |
| Public cible | Développeurs indépendants, particuliers | Sociétés, studios | Applications internes de grandes entreprises |
**Le choix pour un développeur indépendant** : si vous êtes un développeur individuel, optez directement pour le **compte individuel**. Le processus est le plus simple, la vérification la plus rapide, et les fonctionnalités sont amplement suffisantes. Le nom du développeur affiché sur l'App Store sera votre nom réel.
## Préparatifs avant l'inscription
### Conditions préalables
Avant de commencer l'inscription, assurez-vous d'avoir préparé les éléments suivants :
* **Apple ID** : si vous n'en avez pas encore, créez-en un sur [appleid.apple.com](https://appleid.apple.com). Il est recommandé d'utiliser votre adresse e-mail habituelle, car toutes les notifications liées au développement seront envoyées à cette adresse.
* **Authentification à deux facteurs** : l'authentification à deux facteurs (Two-Factor Authentication) doit être activée sur votre Apple ID. Sur iPhone, accédez à « Réglages → Apple ID → Connexion et sécurité → Authentification à deux facteurs » pour l'activer.
* **Appareil Apple** : le processus d'inscription nécessite une vérification d'identité sur un iPhone ou un iPad, via l'application Apple Developer.
### Exigences supplémentaires pour le compte organisation
Si vous vous inscrivez avec un compte organisation, vous aurez également besoin de :
* **Numéro D-U-N-S** : faites votre demande à l'avance sur le site officiel de Dun & Bradstreet ; le traitement prend 5 à 14 jours ouvrés.
* **Statut de représentant légal** : la personne effectuant l'inscription doit être le représentant légal de l'organisation ou un mandataire autorisé.
* **Informations de l'organisation** : adresse du siège social, nom du représentant légal, coordonnées, etc.
## Préparation du matériel de développement
L'inscription au compte développeur n'est que la première étape ; le développement iOS nécessite également certains outils matériels et logiciels.
### Matériel indispensable
* **Mac** — Xcode ne fonctionne que sous macOS, c'est une exigence incontournable. Un Mac équipé d'une puce Apple Silicon (série M) est recommandé pour sa rapidité de compilation et sa capacité à exécuter directement le simulateur iOS. Un MacBook Air série M suffit pour le développement indépendant ; si votre budget est limité, le Mac mini est une alternative.
* **iPhone / iPad (recommandé mais non obligatoire)** — Le simulateur couvre la plupart des scénarios de débogage, mais les tests sur appareil réel restent irremplaçables pour les performances, les capteurs (caméra/GPS/NFC), les notifications push, etc. Même sans compte développeur payant, vous pouvez déboguer sur un appareil réel avec un Apple ID gratuit (avec toutefois des limitations comme la re-signature tous les 7 jours ; voir la FAQ en fin d'article).
### Outils de développement
* **Xcode** — L'IDE officiel d'Apple, disponible gratuitement sur le Mac App Store. Son volume est conséquent (environ 12 Go+) ; la première installation demande un peu de patience.
* **Apple Developer App** — Pour l'inscription au compte, le visionnage des vidéos WWDC et la consultation de la documentation.
* **TestFlight** — Outil de distribution pour les tests bêta, le canal officiel pour inviter des utilisateurs à tester votre application.
### Points d'attention
* Les versions de macOS et Xcode doivent rester à jour ; Apple publie une nouvelle version de Xcode après chaque WWDC, nécessitant généralement l'une des 1 à 2 dernières versions majeures de macOS.
* Les mises à jour de Xcode sont fréquentes et volumineuses ; prévoyez un espace disque suffisant (au moins 50 Go).
* Si votre application utilise des fonctionnalités matérielles (caméra, Bluetooth, NFC, etc.), les tests sur appareil réel sont indispensables.
* Si vous ne disposez pas de Mac, les services de Mac dans le cloud (comme MacStadium, AWS EC2 Mac) constituent une alternative, mais l'expérience n'égale pas celle d'un appareil natif.
## Processus d'inscription
### Première étape : télécharger l'application Apple Developer
Sur votre iPhone ou iPad, ouvrez l'App Store, recherchez « Apple Developer » et téléchargez l'application.
### Deuxième étape : se connecter et commencer l'inscription
Ouvrez l'application Apple Developer et connectez-vous avec votre Apple ID. Appuyez sur l'onglet « Compte », puis sur « S'inscrire à l'Apple Developer Program ».
### Troisième étape : remplir les informations et vérifier votre identité
Remplissez vos informations personnelles selon les instructions :
1. **Confirmation des informations d'identité** : nom, adresse et autres informations de base.
2. **Vérification d'identité** : selon votre pays de résidence, l'application peut vous demander de photographier une pièce d'identité officielle (passeport, permis de conduire, etc.) ou de prendre un selfie pour la vérification.
3. **Acceptation des conditions** : lisez et acceptez le contrat de licence de l'Apple Developer Program.
> La vérification d'identité doit être effectuée dans un environnement bien éclairé pour garantir la netteté des photos. L'ensemble du processus d'inscription doit être complété sur le même appareil.
### Quatrième étape : payer la cotisation annuelle
Après avoir vérifié que toutes les informations sont correctes, procédez au paiement de la cotisation annuelle de ¥688 ($99). Les modes de paiement associés à votre Apple ID sont acceptés. Un e-mail de confirmation vous sera envoyé après le paiement.
### Cinquième étape : attendre la vérification
* **Compte individuel** : la vérification est généralement effectuée sous 48 heures. J'ai effectué le paiement le 14 mars et reçu l'e-mail de bienvenue le matin du 15 mars, soit moins de 24 heures.
* **Compte organisation** : Apple vérifie les informations de l'organisation et le numéro D-U-N-S, ce qui peut prendre plus de temps.
Une fois la vérification terminée, vous pourrez vous connecter au tableau de bord développeur sur [developer.apple.com](https://developer.apple.com) et accéder à toutes les ressources de développement.
## Gestion de l'abonnement et renouvellement
L'Apple Developer Program fonctionne sur un abonnement annuel, renouvelé automatiquement chaque année au tarif de ¥688.
### Renouvellement automatique
Le renouvellement automatique est activé par défaut ; le montant sera prélevé sur le mode de paiement associé à votre Apple ID avant la date d'expiration. Il est recommandé de conserver le renouvellement automatique pour éviter qu'une expiration du compte n'affecte vos applications publiées (voir les questions fréquentes ci-dessous).
### Annulation ou gestion de l'abonnement
Si vous devez modifier les paramètres de renouvellement, ouvrez « Réglages → Apple ID → Abonnements » sur votre iPhone et recherchez l'Apple Developer Program pour le gérer.
## Questions fréquentes
**Q : Que se passe-t-il si j'oublie de renouveler ?**
Après l'expiration de votre compte, vos applications seront retirées de l'App Store, mais elles ne seront pas supprimées. Une fois le paiement effectué à nouveau, vos applications seront restaurées. Cependant, les téléchargements et mises à jour des utilisateurs seront affectés entre-temps ; il est donc recommandé d'activer le renouvellement automatique.
**Q : Comment demander un numéro D-U-N-S ?**
Accédez au lien ci-dessous, remplissez les informations de votre entreprise et soumettez votre demande. Le traitement prend généralement 5 à 14 jours ouvrés. Les comptes individuels n'ont pas besoin de ce numéro.
**Q : Comment contacter Apple en cas de problème ?**
Rendez-vous sur le [support Apple Developer](https://developer.apple.com/contact/) ; vous pouvez les contacter par chat en ligne ou par téléphone. Le support en français est disponible et le temps de réponse est satisfaisant.
**Q : Peut-on commencer à développer avant de s'inscrire ?**
Vous pouvez commencer à explorer, mais il faut connaître les limitations du compte gratuit. Avec un Apple ID gratuit, vous pouvez écrire du code dans Xcode, déboguer avec le simulateur et même installer l'application sur votre propre appareil. C'est amplement suffisant pour apprendre Swift et valider des idées d'interface de base.
Toutefois, le compte gratuit comporte de nombreuses limitations : les applications installées sur un appareil réel doivent être recompilées et réinstallées tous les 7 jours, vous êtes limité à 3 appareils par plateforme, et les fonctionnalités telles que les notifications push, iCloud, TestFlight et les achats intégrés ne sont pas disponibles. Si votre application nécessite ces fonctionnalités, vous aurez besoin d'un abonnement payant dès la phase de développement — pas seulement pour la publication sur l'App Store.
Notre conseil : si vous souhaitez simplement débuter avec Swift et exécuter quelques démos, le compte gratuit suffit. Dès que vous commencez un projet sérieux, inscrivez-vous rapidement à l'abonnement payant pour éviter que les limitations ne retardent votre progression.
# Validation d'idée : d'une inspiration vague à une direction exécutable
Ceci est un article de test.
# Choix technologiques : pourquoi Next.js + Supabase
Ceci est un article de test.
# Agent 记忆系统学习笔记(五):从 Cognee / LightRAG 看懂知识图谱型长期记忆
## 写在前面:为什么最后看 Cognee / LightRAG
前四篇看的是 Agent 与用户交互中的记忆问题:
```text
Mem0:从对话中抽取可召回事实
Zep:把事实放进时间知识图谱
Letta:让 Agent 自主管理上下文层级
LangGraph:提供状态、checkpoint、store 等框架抽象
```
这一篇转向另一条路线:**知识图谱型长期记忆**。
代表项目是 Cognee 和 LightRAG。
这类系统的核心问题不是“用户喜欢什么”,而是:
```text
如何把大量文档、代码、网页、数据库和业务记录
变成 Agent 可以长期查询、持续更新、保留来源、理解关系的知识记忆?
```
这和简单向量 RAG 不一样。向量 RAG 更擅长找“语义相似片段”,但它不天然知道:
```text
哪些实体有关联
这个事实来自哪里
两个文档是否在讲同一个对象
关系是局部的还是全局的
一个回答需要跨多少跳关系
```
GraphRAG 型记忆试图补上这部分。
***
## 一、为什么 Agent 记忆不只是聊天记录
如果只做个人助理,一个 memory layer 可能主要存用户偏好:
```text
用户喜欢中文
用户是前端工程师
用户正在研究 Agent memory
用户不喜欢太长的回答
```
但很多真实 Agent 面对的是另一类问题:
```text
读一个公司代码库
理解一组产品文档
查询历史工单
分析客户会议记录
回答跨文档问题
在企业知识库中做多跳推理
```
这时,“记住用户说过什么”远远不够。系统需要记住的是一个外部世界:
```text
产品、模块、接口、人员、项目、客户、事件、文件、版本、依赖关系
```
这些东西不适合只放成一堆孤立 memory,也不适合只靠 embedding 相似度。
比如用户问:
```text
这个 API 改动会影响哪些下游模块?
```
这不是一个单纯语义相似问题,而是关系问题。
再比如用户问:
```text
A 客户最近提到的性能问题,和之前哪个版本的改动有关?
```
这又是实体、时间、来源、事件之间的组合问题。
所以 Cognee / LightRAG 这类系统会把“记忆”建成图。
***
## 二、Cognee:把数据变成可持续改进的 AI memory
Cognee 的定位很直接:它是一个开源 AI memory 平台。
它的基本目标是把各种数据源转成长期可查询、可改善、可遗忘的记忆。相比传统 RAG,Cognee 更强调 memory 这个概念,因为它不只是存 chunks,而是维护实体、关系、向量和来源。
Cognee 的一个重要特点是三层存储:
### Relational Store:管理来源和元数据
关系数据库负责管理数据对象本身:
```text
documents
datasets
chunks
users
permissions
pipeline 状态
来源 metadata
```
这层很容易被忽略,但在真实应用里非常重要。因为记忆不是一堆无主文本。你需要知道:
```text
这条记忆来自哪个文档
属于哪个用户或组织
是否可以被当前用户访问
什么时候写入
是否应该被删除
```
没有这层,长期记忆很快会变成无法治理的数据垃圾场。
### Vector Store:负责语义相似
向量库存 embeddings。
它负责回答:
```text
哪些文本片段和当前问题语义相近?
哪些 DataPoints 可能相关?
```
这是传统 RAG 的核心能力,Cognee 没有放弃它。区别是:向量检索只是 Cognee 记忆系统的一部分,不是全部。
### Graph Store:表达实体和关系
图数据库存实体和关系。
它负责回答:
```text
A 和 B 是什么关系?
某个实体连接到哪些事件?
这个模块依赖哪些服务?
某个客户的问题和哪些工单、版本、人员相关?
```
这层让记忆从“相似片段集合”变成“可导航的知识结构”。
***
## 三、Cognee 的 remember / recall:把底层 pipeline 包成记忆操作
Cognee v1.0 的抽象很有意思。它把用户面对的主要操作概括成四个:
### remember:写入记忆
`remember` 可以写永久记忆,也可以写带 session\_id 的会话记忆。
永久记忆通常会进入完整处理链路:
```text
数据规范化
切块
提取实体和关系
生成 DataPoints
写入关系库、向量库、图数据库
建立可检索结构
```
这相当于把原始数据变成长期知识记忆。
会话记忆则更像短期或中期上下文。它适合保存一次会话中的内容,并在之后按 session-aware 的方式召回。
### recall:召回记忆
`recall` 不只是做一次 embedding search。Cognee 会根据查询和参数自动路由到不同检索方式,例如:
```text
session memory
graph completion
temporal search
summary search
chunk search
```
这里的关键变化是:召回不再等于“top-k chunk”。
召回变成一个路由问题:
```text
当前问题需要最近会话?
需要某个实体的图邻居?
需要时间范围?
需要摘要?
还是只需要语义片段?
```
### improve:持续改进记忆
`improve` 的意义是:长期记忆不是一次性索引,而是可以持续改进。
一个系统可能先把会话写成 session memory,之后再把其中稳定、有价值的内容提炼到永久图记忆里。
这和人类记忆很像:不是每一句话都立刻进入长期记忆,而是经过整理、抽象、合并后才稳定下来。
### forget:可治理的遗忘
长期记忆一定需要遗忘。
原因很简单:
```text
数据可能过期
用户可能撤回授权
事实可能被更新
低质量记忆会污染召回
法规要求可删除
```
所以 GraphRAG 型记忆不能只考虑“怎么存更多”,还要考虑“怎么删除、怎么隔离、怎么审计”。
***
## 四、LightRAG:轻量图增强 RAG 引擎
LightRAG 的目标和 Cognee 接近,但气质不同。
Cognee 更像一个 AI memory platform,强调 remember / recall / improve / forget 这样的记忆操作。
LightRAG 更像一个轻量 GraphRAG 引擎,强调高效索引、实体关系抽取、增量更新,以及多种查询模式。
LightRAG 的典型流程是:
```text
1. 文档切分成 chunks
2. LLM 从 chunks 中抽取 entities 和 relations
3. 生成实体、关系、文本片段的 embedding
4. 写入图存储、向量存储、KV / 文档状态存储
5. 查询时根据模式组合局部图、全局图和向量片段
```
它的查询模式很有代表性:
```text
local:围绕具体实体做局部检索
global:查全局主题、社区和宏观关系
hybrid:合并 local 和 global
naive:传统向量 chunk 检索
mix:综合 local / global / naive
```
这说明一个问题:GraphRAG 并不是要替代向量检索,而是把向量检索放进更大的检索框架。
当问题是“某个实体附近发生了什么”,local graph 很有用。
当问题是“这个知识库整体上有哪些主题”,global graph 更有用。
当问题只是找一段相似文字,naive vector search 反而足够。
好的记忆系统应该能在这些模式之间切换。
***
## 五、图检索和向量检索到底差在哪
这一点非常关键。
向量检索回答的是:
```text
哪些文本和 query 在语义空间里接近?
```
图检索回答的是:
```text
哪些实体通过关系连接?
这条关系是什么类型?
可以沿关系走到哪里?
```
两者解决的问题不同。
### 向量检索的优势
向量检索适合:
```text
模糊语义匹配
找相似段落
召回没有明确实体的问题
处理自然语言表达差异
快速搭建 baseline RAG
```
它的弱点是:
```text
关系结构弱
多跳推理弱
来源和实体合并困难
容易召回语义相近但关系不对的片段
```
### 图检索的优势
图检索适合:
```text
实体关系查询
依赖分析
多跳推理
时间线和事件链
跨文档归并
```
它的弱点是:
```text
建图成本高
实体抽取和关系抽取会出错
schema 设计复杂
图过大后检索和排序也不简单
```
所以 Cognee / LightRAG 这类系统通常不会二选一,而是同时保留图和向量。
真正的问题不是“图好还是向量好”,而是:
```text
当前问题应该先走语义相似,还是先走实体关系?
召回结果如何合并、去重、排序、压缩?
```
***
## 六、GraphRAG 型记忆和前几篇方案的关系
现在可以把整个系列串起来。
Mem0、Zep、Letta、LangGraph 关注的主要是 Agent 与用户交互时的记忆。Cognee / LightRAG 则更偏 Agent 面对外部知识世界时的记忆。
### 和 Mem0 的区别
Mem0 更像“用户级长期偏好记忆”。
它适合记录:
```text
用户喜欢什么
用户是谁
用户过去说过哪些稳定事实
当前回答前应该召回哪些个人记忆
```
Cognee / LightRAG 更适合记录:
```text
一个组织的知识库
一个代码库的实体关系
产品文档里的概念依赖
跨文档、跨实体、跨事件的知识网络
```
### 和 Zep 的区别
Zep 也是图,但它的图更偏“对话中出现的事实、实体、关系、时间”。
Cognee / LightRAG 的图更偏“外部知识库里的实体、关系、文档、主题”。
粗略说:
```text
Zep:conversation graph memory
Cognee / LightRAG:knowledge graph memory
```
当然二者边界不是绝对的。Cognee 也支持 session memory,Zep 也能处理知识关系。但它们默认关注点不同。
### 和 Letta 的区别
Letta 关心 Agent 如何管理自己的上下文窗口:
```text
核心记忆放什么
外部记忆何时查
上下文满了如何换页
Agent 如何主动编辑 memory
```
Cognee / LightRAG 关心知识库本身如何组织:
```text
文档怎么切
实体怎么抽
关系怎么建
图和向量怎么共同召回
```
一个是 Agent 内部上下文管理,一个是外部知识记忆基础设施。
### 和 LangGraph 的区别
LangGraph 是编排框架。
你完全可以在 LangGraph 的某个 node 里调用 Cognee 或 LightRAG:
```text
用户问题进入 graph
node 先读 thread state
再调用 Cognee / LightRAG 检索知识记忆
把检索结果和短期 state 一起喂给 LLM
必要时把新信息写回 store 或外部 memory platform
```
也就是说,LangGraph 和 Cognee / LightRAG 不是替代关系,而是上下游关系。
***
## 七、什么时候该用 GraphRAG 型记忆
不是所有 Agent 都需要 GraphRAG。
如果你的应用只是:
```text
个人聊天机器人
简单客服 FAQ
少量文档问答
用户偏好记忆
```
那纯向量 RAG 或 Mem0 这类 memory layer 可能已经够了。
但如果你面对的是:
```text
大型文档库
代码库理解
企业知识库
跨文档事实合并
多实体关系推理
需要来源追踪和权限治理
知识持续更新
```
那 GraphRAG 型记忆就值得考虑。
它真正适合的问题有几个特征:
```text
实体很多
关系重要
上下文跨文档
问题经常需要多跳
来源和权限不可忽略
知识会持续增长和被修正
```
反过来,如果只是把十几篇文章做问答,强行上知识图谱可能是过度工程。
***
## 八、我的判断:GraphRAG 是长期知识记忆,不是万能记忆
看完 Cognee / LightRAG,我的判断是:GraphRAG 型系统解决的是“长期知识记忆”,不是所有 Agent 记忆问题。
它非常适合把外部数据世界变成可查询结构。
但它不直接解决这些问题:
```text
当前任务执行到哪一步
用户刚刚在这个 thread 里说了什么
Agent 应该如何修改自己的行为规则
哪些用户偏好应该立刻生效
会话上下文满了怎么办
```
这些仍然需要 LangGraph、Letta、Mem0、Zep 或你自己的状态系统来处理。
所以我更愿意把 Agent memory 分成两大类:
```text
交互记忆:围绕用户、对话、任务、偏好、行为改进
知识记忆:围绕文档、实体、关系、来源、跨文档推理
```
Cognee / LightRAG 是第二类的代表。
它们提醒我们:长期记忆不仅是“用户说过什么”,也可能是“世界是什么样的”。
***
## 九、整个系列的收束
写到这里,这个系列可以形成一个比较完整的地图:
```text
Mem0:事实抽取与多信号召回
Zep:带时间维度的对话知识图谱
Letta:Agent 自主管理上下文层级
LangGraph / LangMem:框架级状态、store 和记忆类型抽象
Cognee / LightRAG:知识图谱增强的长期知识记忆
```
如果把它们按问题来分:
| 问题 | 更接近的方案 |
| --------------------------- | ------------------- |
| 用户偏好和事实如何自动记住 | Mem0 |
| 对话事实如何带时间和关系 | Zep |
| Agent 如何自己管理有限上下文 | Letta / MemGPT |
| 复杂 Agent 应用如何持久化状态和长期 store | LangGraph / LangMem |
| 大规模外部知识如何变成长期可查询记忆 | Cognee / LightRAG |
这也说明,Agent 记忆不是单点能力,而是一组系统能力的组合。
一个成熟 Agent 可能同时需要:
```text
checkpointer 保存当前 thread state
store 保存用户长期偏好
memory extractor 抽取稳定事实
graph memory 管理实体关系和时间
knowledge memory 检索外部知识库
prompt optimizer 改进行为规则
forget / permission / provenance 做治理
```
这就是为什么“给 Agent 加记忆”听起来简单,真正做起来却非常复杂。
因为你不是在加一个数据库,而是在设计 Agent 与时间、用户、知识和自身行为之间的关系。
## 参考资料
* [Cognee Documentation](https://docs.cognee.ai/)
* [Cognee GitHub](https://github.com/topoteretes/cognee)
* [Cognee Core Concepts](https://docs.cognee.ai/core-concepts)
* [Cognee Architecture](https://docs.cognee.ai/architecture)
* [LightRAG GitHub](https://github.com/HKUDS/LightRAG)
* [LightRAG Paper](https://arxiv.org/abs/2410.05779)
EOF
# Agent 记忆系统学习笔记(四):从 LangGraph / LangMem 看懂框架级记忆抽象
## 写在前面:为什么第四篇看 LangGraph / LangMem
前三篇分别看了三种系统路线:
```text
Mem0:自动抽取 + 多信号召回
Zep:时间知识图谱 + Context Block
Letta:Agent 自主管理记忆 + 上下文层级
```
这一篇我想换一个角度,不再只看某个记忆产品,而是看**框架级记忆抽象**。
LangGraph / LangMem 的价值在于,它把记忆问题拆得很清楚:
```text
短期记忆:当前 thread 的 state 和 checkpoint
长期记忆:跨 thread 的 store 和 namespace
语义记忆:facts about user
情景记忆:past experiences / examples
程序记忆:instructions / prompts / behavior rules
```
如果说 Mem0、Zep、Letta 分别是不同的系统实现,那么 LangGraph 更像一套开发框架里的基本积木。它不替你决定所有记忆策略,而是告诉你:哪些东西属于状态,哪些东西属于 store,哪些东西应该同步写,哪些东西可以后台写。
***
## 一、LangGraph 的核心问题:记忆先是状态管理
很多 Agent 记忆讨论会直接跳到向量库、图数据库、长期偏好。但 LangGraph 的起点更工程化:**一个 Agent 工作流本身就是一个状态机。**
每次调用图,都会有 state。state 里可能有:
```text
messages
summary
retrieved documents
tool outputs
intermediate artifacts
human feedback
下一步要执行的节点
```
如果这些状态不能持久化,Agent 就无法跨轮继续,也无法在中断后恢复,更无法做 human-in-the-loop 或 time travel。
所以 LangGraph 的记忆首先不是“长期知识”,而是“工作流状态”。
LangGraph 把持久化拆成两套系统:
```text
Checkpointer:保存 thread-scoped graph state
Store:保存 cross-thread long-term memory
```
这个划分非常重要。它让我们把“当前会话连续性”和“跨会话长期记忆”分开看。
***
## 二、短期记忆:Checkpointer 保存 thread state
短期记忆在 LangGraph 里通常指 thread-scoped memory。
一个 thread 就像邮件里的一个会话串。用户在这个 thread 里持续和 Agent 对话,Agent 需要记住当前对话发生过什么。
LangGraph 用 checkpointer 保存 graph state 的 checkpoint。调用时传入:
```text
thread_id
```
框架就能找到这个 thread 对应的状态,并在下一次调用时继续使用。
短期记忆适合存:
```text
最近消息
当前任务状态
工具调用结果
中间产物
会话摘要
等待用户确认的 interrupt 状态
```
它的核心价值不只是聊天连续性,还包括:
```text
恢复中断任务
查看历史 state
time travel 到过去 checkpoint
容错和重试
human-in-the-loop 审批
```
这和 Mem0 / Zep 的长期召回不一样。Checkpointer 解决的是:“这个工作流刚刚进行到哪里?”
***
## 三、长期记忆:Store 保存跨会话数据
长期记忆在 LangGraph 里通过 store 实现。
Store 保存的是 application-defined data,也就是应用自己定义的 JSON 文档。每条数据通常放在一个 namespace 下面,再用 key 标识。
比如:
```text
namespace = (user_id, "memories")
key = memory_id
value = {"text": "用户喜欢短而直接的回答"}
```
Store 可以在 graph node 内部读写。也就是说,某个节点在调用模型前可以:
```text
根据当前用户输入搜索长期记忆
把搜到的记忆拼进 system prompt
必要时把新记忆写回 store
```
和 checkpointer 的区别是:
| 维度 | Checkpointer | Store |
| ---- | ----------------------- | ------------------------------ |
| 范围 | 单个 thread | 跨 thread |
| 内容 | graph state snapshots | 应用定义的长期数据 |
| 用途 | 会话连续性、恢复、中断、time travel | 用户偏好、事实、样例、共享知识 |
| 访问方式 | 通过 thread\_id 恢复 state | 通过 namespace / key / search 访问 |
所以,如果用户在 thread A 说“记住我叫 Bob”,然后在 thread B 问“我叫什么”,这就不应该只靠 checkpointer,而应该写进 store。
***
## 四、三类长期记忆:Semantic、Episodic、Procedural
LangGraph / LangMem 的另一个价值,是把长期记忆按类型拆开。
### Semantic Memory:事实和知识
Semantic memory 是事实记忆。
例如:
```text
用户喜欢 Python。
用户偏好简洁回答。
Alice 负责 ML 团队。
某个项目使用 Postgres。
```
这是大多数人想到“长期记忆”时最先想到的东西。Mem0 主要处理的也是这一类。
Semantic memory 可以有两种常见形态:
```text
Profile:一个持续更新的用户画像 JSON
Collection:一组独立 memory 文档
```
Profile 的好处是整体性强,坏处是越大越难安全更新。Collection 的好处是易追加、召回率高,坏处是关系和全局一致性更难维护。
### Episodic Memory:经历和案例
Episodic memory 记的是过去发生过的事件或经验。
在 Agent 里,它经常表现为 few-shot examples:
```text
之前遇到类似任务时,Agent 是怎么做的?
哪个工具调用序列成功解决了问题?
用户之前如何评价某种回答?
```
它回答的不是“事实是什么”,而是“过去怎样处理过”。
这类记忆对 coding agent、客服 agent、自动化 agent 特别重要。因为很多能力不是背事实,而是复用成功轨迹。
### Procedural Memory:规则和行为
Procedural memory 记的是“怎么做事”。
在 AI Agent 里,它通常不是真的改模型权重,而是改:
```text
system prompt
instructions
tool usage guidelines
workflow rules
response style
```
LangMem 的 prompt optimizer 就属于这个方向:从成功和失败轨迹中总结行为改进,再生成新的 prompt 规则。
这点对整个系列很重要。前面 Mem0 和 Zep 更多偏 semantic / episodic;Letta 的 persona、policies、scratchpad 则很接近 procedural / working memory。LangGraph 把这些类型统一到了一个分类框架里。
***
## 五、写入时机:hot path 还是 background
长期记忆还有一个关键问题:什么时候写?
LangGraph 文档把它分成两类:
### Hot path 写入
Hot path 是在用户请求链路里直接写记忆。
比如用户说:
```text
记住,我以后都想用中文回答。
```
Agent 立刻把它写入 store。下一轮马上生效。
优点:
```text
实时
透明
新记忆立刻可用
```
缺点:
```text
增加延迟
Agent 要分心判断是否写入
写入质量可能影响主任务
```
### Background 写入
Background 是在主流程之外异步整理记忆。
比如:
```text
对话结束后总结用户偏好
每天批量整理用户历史
后台从工具调用轨迹中抽取可复用经验
定期优化 system prompt
```
优点:
```text
不影响主链路延迟
可以用更复杂的抽取逻辑
适合批处理和人工审核
```
缺点:
```text
新记忆不会立刻生效
需要任务队列和触发策略
可能出现过期或重复记忆
```
这也是记忆系统真正复杂的地方。不是“要不要写入向量库”,而是要回答:
```text
谁来判断这件事值得记?
什么时候写?
写到哪一类记忆?
旧记忆冲突时怎么办?
是否需要用户确认?
```
***
## 六、召回链路:state 直接读,store 按需搜
LangGraph 的召回链路也很清楚:
```text
短期上下文来自 state
长期上下文来自 store
node 负责把二者组装进 prompt
```
一次典型调用可能是:
```text
1. Graph node 收到当前 state
2. 从 state 读取最近 messages 和当前任务状态
3. 根据 user_id 和当前 query 搜索 store
4. 得到相关长期记忆
5. 把短期状态 + 长期记忆组织进模型 prompt
6. 模型生成回答或工具调用
7. 根据结果更新 state,必要时写入 store
```
这和“把所有历史对话都塞进上下文”有本质区别。
LangGraph 的方式是分层的:
```text
当前 thread 内的连续性:state / checkpoint
跨 thread 的长期偏好:store / namespace
相关性筛选:search
最终上下文构造:node / prompt
```
也就是说,记忆不是一个单独 API,而是嵌在 graph execution 里的上下文工程。
***
## 七、LangMem:把记忆抽取和行为优化工具化
LangGraph 本身提供 state、checkpoint、store、node 这些基础设施。LangMem 则更进一步,提供围绕长期记忆的工具包。
它重点做几件事:
```text
从对话中抽取和更新长期记忆
管理 semantic / episodic / procedural memory
从反馈和轨迹中优化 agent 行为
和 LangGraph store 集成
```
如果 LangGraph 是“可以建记忆系统的框架”,LangMem 更像“常见记忆策略的 SDK”。
这两个东西放在一起,形成了一个很完整的抽象层:
```text
LangGraph:状态机、持久化、store、runtime
LangMem:记忆抽取、记忆更新、prompt 优化、行为学习
```
但它们仍然不是 Mem0 或 Zep 那种开箱即用的 memory service。你需要自己决定:
```text
记忆 schema 怎么设计
namespace 怎么划分
哪些 node 负责检索
哪些 node 负责写入
是否要向量搜索
是否要人工审核
如何处理冲突、遗忘和权限
```
这就是框架级方案和产品级方案的区别。
***
## 八、和 Mem0、Zep、Letta 的对应关系
把前四篇放在一起看,会更清楚:
### Mem0:记忆是可排序的事实集合
Mem0 更偏产品化记忆层。它帮你自动抽取 memory,再在召回时做语义、图关系、时间、实体等多信号排序。
你关心的是:
```text
怎么接入 memory API
怎么让系统自动记住用户偏好
怎么在回答前拿到相关 memories
```
### Zep:记忆是带时间的关系图
Zep 的重点是事实、实体、关系、时间和 invalidation。
它关心的是:
```text
事实什么时候成立
后来有没有被推翻
哪些实体和关系影响当前问题
如何组装成 context block
```
### Letta:记忆是 Agent 可管理的上下文层级
Letta / MemGPT 的核心是让 Agent 自己管理 memory hierarchy。
它关心的是:
```text
core memory 放什么
archival memory 怎么查
context 满了怎么换页
Agent 如何通过工具修改自己的记忆
```
### LangGraph:记忆是 Agent 应用里的可编排基础设施
LangGraph 不把记忆封装成一个黑盒服务,而是把它拆成:
```text
state
checkpoint
store
namespace
node
prompt assembly
background task
```
它关心的是:
```text
一个真实 Agent 应用如何可靠运行、恢复、检索、写入和演化
```
所以 LangGraph 更像 memory operating system 的开发框架,而不是单独的 memory app。
***
## 九、我的判断:LangGraph 适合把“记忆”放回工程系统里
看完 LangGraph / LangMem,我最大的感受是:它把“记忆”从一个模型能力问题,重新拉回到工程系统问题。
真实应用里,Agent 记忆至少包含四层:
```text
执行状态:当前任务跑到哪里了
会话上下文:这个 thread 里刚刚发生了什么
长期用户偏好:跨会话稳定存在的事实
行为改进:Agent 应该如何变得更会做事
```
不同层的生命周期完全不同。
```text
执行状态可能几分钟后就没用
会话上下文可能保留几天
用户偏好可能长期有效
行为规则需要持续评估和版本管理
```
如果用一个“记忆库”概念把它们全部混在一起,系统迟早会变得混乱。
LangGraph 的优点是:它没有假装所有问题都能靠一个 memory API 解决。它提供了足够底层、但又足够实用的抽象,让你按自己的应用边界来设计记忆。
它的缺点也来自这里:你需要自己承担更多设计工作。
如果你想快速给聊天机器人加长期用户偏好,Mem0 可能更快。
如果你想要时间图谱和事实失效,Zep 更直接。
如果你想研究 Agent 自主管理上下文,Letta 更有启发。
如果你要搭一个复杂、多节点、可恢复、可调度、可观测的 Agent 应用,那么 LangGraph / LangMem 会更自然。
***
## 十、这一篇的结论
LangGraph / LangMem 让我重新理解了 Agent 记忆系统的一个基本事实:
> 记忆不是一个单独模块,而是贯穿 Agent 执行、状态管理、检索、写入、提示词构造和行为优化的系统能力。
它的核心贡献不是提出一种新的记忆数据库,而是把记忆拆成了几个工程上可操作的概念:
```text
短期记忆:thread state + checkpointer
长期记忆:namespace + store
记忆类型:semantic / episodic / procedural
写入时机:hot path / background
召回方式:state read / store search / prompt assembly
```
这套抽象非常适合用来回看整个系列。
如果说:
```text
Mem0 解决“哪些事实值得记、怎么排”
Zep 解决“事实之间有什么时序关系”
Letta 解决“Agent 如何自己管理上下文”
```
那么 LangGraph 解决的是:
```text
这些记忆机制如何变成一个可运行、可恢复、可扩展的 Agent 应用架构。
```
下一篇我会继续看 Cognee / LightRAG 这条路线。它们把重点从“用户和 Agent 的交互记忆”推向“知识图谱增强的长期知识记忆”。这会帮助我们理解另一类问题:当 Agent 面对的不是一个用户偏好,而是一大堆文档、代码和业务知识时,记忆系统应该怎么建。
## 参考资料
* [LangChain Docs: Memory](https://docs.langchain.com/oss/python/concepts/memory)
* [LangGraph Docs: Add memory](https://langchain-ai.github.io/langgraph/how-tos/memory/add-memory/)
* [LangGraph Docs: Persistence](https://langchain-ai.github.io/langgraph/concepts/persistence/)
* [LangChain Blog: LangMem SDK](https://blog.langchain.com/langmem-sdk-launch/)
# Agent 记忆系统学习笔记(三):从 Letta / MemGPT 看懂 Agent 自主管理记忆
## 写在前面:第三篇为什么看 Letta
前两篇我们已经看了两种记忆系统路线。
Mem0 更像一个产品化的 memory ranking layer:把对话抽取成 memory,然后用语义、BM25、实体、时间和衰减信号把相关记忆召回。
Zep 更像 temporal graph + context assembly:把用户、业务数据和历史事件组织成 Context Graph,再把相关 facts、entities、episodes、observations 组装成 Context Block。
这一篇看第三种路线:**Agent 自己管理记忆。**
Letta 的前身和思想来源是 MemGPT。MemGPT 论文的核心比喻很有意思:LLM 的上下文窗口像计算机的主内存,很快、很贵、很有限;外部存储像磁盘,很大但需要主动读取。既然传统操作系统可以通过虚拟内存管理有限的物理内存,那么 Agent 能不能也管理自己的上下文?
Letta 继承了这个思路。它不是只问“如何自动召回 memory”,而是问:
> 如果 Agent 是一个有状态的程序,它能不能自己决定哪些内容常驻上下文,哪些内容放入外部长期记忆,什么时候搜索,什么时候写入,什么时候修改?
这就是 Letta 最值得学习的地方。
***
## 一、Letta 的核心问题:Agent 为什么需要管理自己的记忆
大多数记忆系统会把“召回”放在应用层或基础设施层处理。
典型流程是:
```text
用户输入 → 系统自动搜索记忆 → 把结果塞进 prompt → LLM 回答
```
这当然有效,但它有一个假设:外部系统比 Agent 更知道该召回什么。
Letta 的思路不同。它认为一个真正有状态的 Agent,不应该只是被动接收别人塞进来的上下文,而应该参与管理自己的状态。
在 Letta 里,一个 stateful agent 不只是一次 LLM 调用。它包含:
```text
system prompt
memory blocks
messages
reasoning
tool calls
tools
runs / steps
conversations
```
这些状态都会被持久化到数据库里。即使某些消息被压缩、从上下文窗口里移出去,底层记录也不会丢。
这个定位很关键。Letta 不是把 memory 当成一个旁路组件,而是把 memory 放进 Agent runtime 本身。
***
## 二、MemGPT 的操作系统比喻
理解 Letta,最好先理解 MemGPT 的比喻。
MemGPT 论文提出的是 **virtual context management**。它借鉴操作系统的分层内存:
```text
CPU 可以直接访问的内存很有限
磁盘容量大但访问慢
操作系统负责在内存和磁盘之间调度数据
给程序一种“我有更大内存”的错觉
```
放到 LLM Agent 里就是:
```text
LLM context window = 快速但有限的主内存
外部存储 / archival memory = 容量大但需要搜索的磁盘
Agent memory tools = 在两者之间搬运信息的系统调用
```
所以 Letta / MemGPT 路线的核心不是“检索算法”,而是“上下文管理”。
它要解决的问题是:
```text
什么必须放在上下文里?
什么可以放在外部,需要时再搜?
Agent 什么时候应该写入长期记忆?
Agent 什么时候应该修改自己的 core memory?
Agent 如何在多轮对话中保持状态连续?
```
这让 Letta 和 Mem0、Zep 的气质很不一样。Mem0 和 Zep 更像记忆基础设施,Letta 更像一个带记忆操作能力的 Agent OS。
***
## 三、Letta 的记忆分层
Letta 的文档里有一个很重要的 context hierarchy。不同类型的信息,根据重要性、大小和访问频率,应该进入不同层。
这里最核心的是两层:
```text
Memory Blocks:常驻上下文
Archival Memory:外部长期记忆,按需搜索
```
另外还有 files、external RAG、messages、shared blocks 等扩展形态。
我理解 Letta 的核心判断是:
> 不是所有记忆都应该被检索。最重要的记忆应该直接常驻上下文;低频的大量记忆才应该按需搜索。
这和 Mem0、Zep 很不一样。
Mem0 / Zep 更强调“检索出当前相关内容”。Letta 则先问:“这条信息是不是重要到不应该等检索?”
***
## 四、Core Memory:Memory Blocks 为什么重要
Letta 里最核心的抽象是 memory block。
Memory block 是一段结构化的上下文,会被放进 Agent 的 prompt 里。它可以有 label、description、value、limit。常见 block 有:
```text
persona:Agent 自己是谁,应该如何表现
human:用户是谁,有什么偏好和长期背景
scratchpad:当前任务的工作记忆
organization:组织级共享信息
policies:只读规则和约束
```
Memory block 的关键点是:**它不需要召回。**
只要 block attach 到某个 Agent,它就在上下文里,Agent 每次推理都能看见。
这适合存什么?
```text
用户姓名
非常稳定的偏好
Agent persona
关键约束
当前任务状态
必须长期遵守的政策
```
Letta 文档特别强调 description 字段。因为 Agent 会根据 block 的 label 和 description 判断怎么读写这块记忆。
这点很有启发:Memory block 不是一个随便写字符串的地方,它更像一个带说明书的上下文槽位。description 写得好不好,直接影响 Agent 是否知道该把什么写进去。
***
## 五、Archival Memory:外部长期记忆不是常驻上下文
Memory blocks 适合小而重要的信息。但如果信息很多,就不能全部放进上下文。
这时 Letta 用 archival memory。
Archival memory 是一个语义可搜索的外部长期存储。Agent 可以通过工具:
```text
archival_memory_insert
archival_memory_search
```
来写入和搜索。
它适合存:
```text
大量历史互动
研究笔记
文档摘要
技术参考
客服历史
用户调研
不必总是可见、但未来可能有用的信息
```
Letta 文档里有个区分很重要:
```text
Archival memory 是 intentional storage。
Conversation search 是 historical retrieval。
```
也就是说,archival memory 是 Agent 主动决定“这值得长期保存”;conversation search 则是搜索过去实际说过什么。
比如用户说:
```text
我偏好 Python 做数据科学项目。
```
Agent 可以把它写入 archival memory:
```text
User prefers Python for data science.
```
以后也可以通过 conversation search 找到原始对话。前者是抽取后的知识,后者是历史证据。
***
## 六、Letta 的关键差异:记忆管理是工具调用
Letta 最有意思的地方,是它让记忆操作变成 Agent 可调用的工具。
这和自动召回型系统很不一样。
在 Mem0 或 Zep 里,应用通常会在回答前自动拿到相关 memory 或 context block。Agent 不一定知道召回过程发生了什么。
在 Letta 里,Agent 可以在对话中决定:
```text
我需要更新 human block。
我需要把这条事实写进 archival memory。
我现在信息不够,需要搜索 archival memory。
我需要修改 scratchpad。
```
这是一种 agent-managed memory。
它的优点是灵活。Agent 可以根据任务需要动态管理上下文,不必完全依赖应用层预设的检索策略。
但它也有代价:
```text
Agent 可能忘记搜索
Agent 可能写入错误记忆
Agent 可能把不该常驻的信息放进 core memory
Agent 可能多次改写 block,产生冲突或覆盖
```
所以 Letta 的记忆路线更像“给 Agent 更大的自主权”,而不是“用基础设施替 Agent 做完所有判断”。
***
## 七、Context Hierarchy:什么时候用哪一层
Letta 的 context hierarchy 可以总结成一句话:
> 越重要、越小、越高频的信息,越应该靠近上下文;越大、越低频、越外部的信息,越应该通过工具检索。
我把它理解成四层:
### Memory Blocks:小而关键,必须常驻
适合:
```text
用户名字
Agent persona
核心规则
当前任务状态
高频偏好
```
它的风险是占用上下文,而且如果 block 被错误覆盖,影响会很大。
### Files:较大、只读、可部分打开和搜索
适合公司文档、指南、说明书。Agent 不需要一次看完,但可以打开片段、搜索内容。
### Archival Memory:长期、低频、按需语义搜索
适合 Agent 自己积累的知识和经历。
它不是每次都出现,但当 Agent 觉得需要时,可以搜索。
### External RAG / MCP:无限外部知识
适合超大规模知识库、业务数据库、外部系统。Letta 可以通过自定义工具或 MCP 让 Agent 访问。
这套层级最大的价值是:它把“记忆放在哪里”变成一个可判断的问题,而不是默认全都扔进向量库。
***
## 八、Shared Memory:多 Agent 协作里的共享状态
Letta 的 shared memory blocks 也很值得注意。
一个 block 可以被多个 Agent attach。这样多个 Agent 都能看到同一份信息。如果一个 Agent 更新 block,其他 Agent 也能看到更新。
这适合多 Agent 协作:
```text
task_queue:主管 Agent 写任务,Worker Agent 读任务并更新状态
organization:所有 Agent 共享组织背景
user_preferences:多个服务同一用户的 Agent 共享用户偏好
handoff_context:Agent A 交接给 Agent B 前写入上下文
system_config:只读全局配置
```
这个能力说明,memory block 不只是“个人记忆”,也可以是协作状态。
但它也带来并发问题。Letta 文档提醒,memory\_rethink 这种完整重写操作在多 Agent 同时编辑时并不安全,容易最后写入覆盖前面的修改。更安全的是 append 型的 memory\_insert,或者让某个 Agent 成为 block 的 owner。
这让我觉得 Letta 的 shared block 更像一个轻量共享状态系统,而不是传统数据库。它适合协调,但要小心并发写入。
***
## 九、和 Mem0、Zep 的关键差异
现在可以把三篇放在一起看。
| 维度 | Mem0 | Zep | Letta / MemGPT |
| -------- | ---------------- | ---------------------- | ---------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph | stateful agent + memory hierarchy |
| 主要问题 | 如何抽取和召回相关 memory | 如何表达关系和时间变化 | Agent 如何管理自己的上下文 |
| 常驻上下文 | 通常由应用注入搜索结果 | Context Block 注入 | Memory Blocks 常驻 |
| 外部记忆 | memory store | graph / context lake | archival memory / files / external tools |
| Agent 角色 | 使用被召回的 memory | 使用被组装的 context | 主动搜索、写入、修改记忆 |
| 适合场景 | 个性化记忆层 | 关系和状态演化 | 长期自主 Agent、复杂上下文管理 |
简单说:
```text
Mem0 关注怎么把记忆召回得更准。
Zep 关注怎么把记忆组织成时间图谱。
Letta 关注 Agent 如何自己管理记忆和上下文。
```
这三者不是互斥关系,而是三个不同抽象层。
一个复杂系统甚至可能同时用它们的思想:用 memory blocks 放核心状态,用图谱维护业务关系,用向量或多信号检索管理大规模长期记忆。
***
## 十、我对 Letta 的判断
Letta 最值得学习的地方,不是某个具体检索算法,而是它对 Agent 的定位。
它把 Agent 当成一个有状态的长期程序,而不是一个每次从 prompt 重建出来的无状态函数。
所以 Letta 的问题意识是:
```text
Agent 的状态如何持久化?
哪些状态应该常驻上下文?
哪些状态应该外置?
Agent 能不能自己编辑状态?
多个 Agent 如何共享状态?
上下文窗口不够时,如何分层管理?
```
这对长期陪伴型 Agent、个人助理、研究助理、多 Agent 工作流都很重要。
它适合:
```text
需要长期人格和用户关系的 Agent
需要 Agent 自己维护任务状态的系统
需要多 Agent 共享上下文或交接的工作流
需要让 Agent 主动整理记忆的研究型产品
```
它不一定适合:
```text
只想简单接一个自动记忆 API
不希望 Agent 自己改记忆
需要严格确定性的记忆写入和召回
记忆治理、审批、审计要求非常强的场景
```
我的判断是:
> Letta / MemGPT 的价值在于,它把“记忆召回”升级成“上下文管理”。它不是单纯帮 Agent 找记忆,而是给 Agent 一套管理自己状态的机制。
这也是为什么它应该放在 Mem0 和 Zep 后面分析。Mem0 让我们理解 memory pipeline,Zep 让我们理解 temporal graph,Letta 让我们理解 Agent-managed memory。
***
## 十一、这一篇之后,分析框架继续扩展
到这里,这个系列已经有三种视角:
```text
Mem0:记忆是可排序的事实集合
Zep:记忆是带时间的关系图
Letta:记忆是 Agent 可管理的上下文层级
```
所以以后分析任何 Agent memory 系统,我会继续加几个问题:
```text
它是否允许 Agent 主动管理记忆?
哪些记忆是常驻上下文,哪些记忆是按需搜索?
Agent 修改记忆时有没有工具边界?
多 Agent 是否能共享记忆?
记忆被错误修改后如何恢复?
记忆管理是系统自动完成,还是 Agent 自己参与?
```
下一篇我建议看 LangGraph / LangMem。它更像一个框架级抽象,可以把 short-term memory、long-term memory、semantic / episodic / procedural memory 这些概念统一起来。
***
## 参考资料
* [Letta Stateful Agents](https://docs.letta.com/guides/core-concepts/stateful-agents/)
* [Letta Memory Blocks](https://docs.letta.com/guides/core-concepts/memory/memory-blocks/)
* [Letta Shared Memory](https://docs.letta.com/guides/core-concepts/memory/shared-memory/)
* [Letta Archival Memory](https://docs.letta.com/guides/core-concepts/memory/archival-memory/)
* [Letta Context Hierarchy](https://docs.letta.com/guides/core-concepts/memory/context-hierarchy/)
* [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560)
# Agent 记忆系统学习笔记(一):从 Mem0 看懂记忆与召回链路
## 写在前面:这不是 Mem0 产品介绍,而是记忆系统入门
这个系列我想解决一个问题:**AI Agent 的记忆和召回到底是怎么设计的?**
如果只看表层,很多系统都在说自己有 memory:有的接一个向量库,有的保留聊天历史,有的做 GraphRAG,有的让 Agent 自己调用工具搜索记忆。但这些方案背后的关键问题其实是同一组:
```text
什么内容值得被记住?
它被存到哪里?
未来怎么找到它?
找到以后怎么排序?
什么时候应该忘掉、降权或保留历史?
最后怎么把记忆交给模型使用?
```
这一篇先从 Mem0 开始。原因很简单:Mem0 是一个比较典型的工程化 Agent Memory 案例,它把长期记忆拆成了写入、存储、实体链接、多信号召回、时间推理和排序等模块。学完它之后,再看 Zep、Letta、LangGraph、Cognee,会更容易建立统一的分析框架。
我对这篇笔记的定位是:**不是教你怎么接入 Mem0,而是用 Mem0 学会拆解一个记忆系统。**
***
## 一、先建立心智模型:记忆不是聊天记录
很多人第一次理解 Agent Memory,会把它等同于“保存历史聊天”。但这只是最低级的记忆。
聊天记录是原始材料,记忆是被提炼后的可复用事实。
比如用户说:
```text
我不太喜欢惊悚片,最近更想看科幻电影。
```
原始聊天记录只是这句话本身。但一个记忆系统真正应该沉淀的是:
```text
用户不喜欢惊悚片。
用户偏好科幻电影。
```
以后用户问“今晚看什么”,Agent 不需要用户重复偏好,而是能主动召回这两条记忆。这才是长期记忆的价值:**让每一次交互都能沉淀成未来可用的上下文。**
所以,记忆系统不是一个日志系统。日志系统记录“发生过什么”;记忆系统要判断“未来什么会有用”。
***
## 二、Mem0 的整体位置和写入链路
### Mem0 在 Agent 架构中的位置
Mem0 可以理解成应用和大模型之间的一层长期记忆基础设施。
它有两个最核心的动作:
```text
add :把新的对话或事实写入记忆系统
search :在未来根据 query 召回相关记忆
```
这两个 API 看起来简单,但背后分别对应一条完整链路。
***
### 记忆写入链路:从对话到可召回事实
先看写入。
Mem0 不是简单把整段对话塞进数据库,而是要先做一层“蒸馏”:从对话中抽取值得保留的事实。
举个例子:
```text
用户:我下个月要去东京,我比较喜欢住在涩谷附近,也不吃贝类。
助手:好的,我会记住你的旅行偏好。
```
比较理想的 memory extraction 结果是:
```text
用户下个月计划去东京。
用户偏好住在涩谷附近。
用户不吃贝类。
```
这一步是 Agent Memory 和普通 RAG 的核心区别之一。
传统 RAG 往往默认资料已经存在,重点在检索;Agent Memory 需要先判断什么内容值得被写成记忆。
***
### ADD-only:为什么 Mem0 不急着覆盖旧记忆?
Mem0 新版算法里一个重要取舍是 **ADD-only extraction**。也就是写入阶段主要做新增,而不是让 LLM 在每次写入时决定 ADD、UPDATE、DELETE。
这件事一开始看起来有点反直觉:如果只新增,记忆不是会越积越多吗?
会。但这是有意为之。
因为很多事实不是互相冲突,而是随时间演化。
```text
2025 年:用户住在纽约。
2026 年:用户搬到了旧金山。
```
如果系统把“用户住在纽约”直接更新成“用户住在旧金山”,那它就丢掉了历史。以后用户问:
```text
我搬家之前住在哪里?
```
系统可能就答不上来。
所以 Mem0 的思路是:
```text
写入时尽量保留事实
召回时再判断当前问题需要哪条事实
```
这个设计很关键。它把复杂度从写入阶段转移到了召回阶段。
写入阶段少做破坏性判断,可以降低误删、误合并、误覆盖的风险;召回阶段再根据语义、实体、时间和新旧程度决定哪条记忆更应该出现。
这也是我认为 Mem0 最值得学习的地方之一:**长期记忆不应该只保存“当前状态”,还应该保存“状态如何变化”。**
***
## 三、Mem0 的存储与 Graph Memory
### Mem0 的存储:向量库只是其中一层
很多人会把记忆系统想象成一个向量数据库。但 Mem0 的结构不止这一层。
从官方文档和评估说明看,它至少可以拆成三类存储:
Vector Store 负责“意思像不像”。Entity Store 负责“是不是和同一个人、地点、项目、概念有关”。History Store 则更偏系统内部的事件记录和上下文管理。
这个分层说明一件事:**一个好的记忆系统不能只靠 embedding。**
Embedding 适合语义相似,但它对专有名词、时间、实体关系、精确关键词并不总是稳定。记忆系统需要更多索引结构共同工作。
***
### Graph Memory:Mem0 的图更像“实体增强检索”
Mem0 现在的 Graph Memory 是内置的,不再要求使用 Neo4j、Memgraph、Kuzu 这类外部图数据库。
它的基本逻辑是:
之后用户问:
```text
Alice 现在和什么项目有关?
```
系统就不只看语义相似,也会看哪些 memory 和 Alice 这个实体连接。相关 memory 会获得排序加权。
但要注意:Mem0 的 Graph Memory 不是完整知识图谱。它不会严格记录:
```text
Alice -- 管理 --> Bob
Alice -- 就职于 --> Acme
```
它更像:
```text
Alice 这个实体连接了哪些 memories?
哪些 memories 共同提到了 Alice?
```
所以我会把 Mem0 的图理解成 **entity-aware retrieval**,也就是实体增强召回,而不是强 schema 的图推理系统。
这个判断很重要。因为如果你只是想让 Agent 更容易召回和某个人、项目、公司相关的历史,Mem0 的图已经很有用;但如果你要做复杂关系推理、路径查询、关系约束,那就需要看 Zep/Graphiti 或传统图数据库方案。
***
## 四、Mem0 的召回链路
### 召回链路总览:Mem0 怎么“想起来”?
记忆系统最难的不是存,而是召回。
因为长期记忆会越来越多,里面会有旧信息、相似信息、冲突信息、不同 session 的信息、不同用户的信息。真正的问题是:**当前这一刻,哪几条记忆最应该进入上下文?**
Mem0 的召回链路可以这样理解:
这条链路里,每一步都在减少错误召回的概率。
***
### 第一步:先确定搜索边界,而不是全局乱搜
记忆召回首先要解决隔离问题。
如果一个系统里有多个用户、多个 Agent、多个应用、多个 session,搜索时不能直接全局检索。否则很容易出现“串记忆”。
Mem0 用这些字段来划分记忆空间:
| 字段 | 作用 | 例子 |
| --------- | ------------ | -------------- |
| user\_id | 某个用户的长期记忆 | 用户偏好、历史行为 |
| agent\_id | 某个 Agent 的记忆 | 旅行助手、客服助手、学习助手 |
| app\_id | 某个应用或租户 | 白标应用、不同产品线 |
| run\_id | 某次任务或会话 | 一张客服工单、一次旅行规划 |
所以一个典型 search 应该像这样:
```python
client.search(
query="用户有什么饮食禁忌?",
filters={"user_id": "alice"}
)
```
这一层不是锦上添花,而是安全底线。
我建议以后拆所有 memory 系统时,都先问:
```text
它的记忆隔离模型是什么?
按用户隔离,还是按会话隔离?
多 Agent 共享记忆时怎么防止串上下文?
```
没有 scope 的记忆系统,很难进入生产。
***
### 第二步:语义检索,解决“意思相近”
语义检索是最基础的一层。
query 会被转成 embedding,然后和 memory embedding 做相似度匹配。
适合的问题是:
```text
用户喜欢什么类型的电影?
用户对旅行有什么偏好?
之前这个项目有哪些背景?
用户对远程办公怎么看?
```
这类问题不一定有精确关键词,但语义上可以匹配。
不过,语义检索不是万能的。它常见的问题是:
```text
专有名词可能召回不稳
订单号、会议名、项目名可能被弱化
时间问题容易混淆
实体关系不够清晰
相似但不相关的内容可能被召回
```
所以 Mem0 后面又叠了关键词、实体和时间信号。
***
### 第三步:BM25,解决“精确词命中”
有些查询不是语义问题,而是关键词问题。
比如:
```text
Q1 roadmap
Alice
invoice-38291
San Francisco
March 10
```
这些词对用户来说非常关键,但在向量空间里不一定有足够稳定的表达。BM25 的价值就是让这些关键词能影响排序。
所以 Mem0 的召回不是只看:
```text
这条 memory 和 query 意思像不像?
```
还要看:
```text
这条 memory 有没有命中 query 里的关键字?
```
不过 Mem0 文档里有一个细节值得记住:在 OSS v3 的说明中,BM25 更像是 boost signal,不一定是独立扩展候选池。也就是说,它主要提高相关候选的排序,而不是完全替代向量召回。
这说明 Mem0 的多信号召回更像:
```text
先通过语义找到候选
再用关键词和实体信号调整排序
```
***
### 第四步:实体匹配,解决“和谁有关”
实体匹配是 Mem0 召回链路中最值得关注的一层。
用户问:
```text
Alice 相关的项目有哪些?
```
系统会从 query 中识别实体:
```text
Alice
```
然后去 Entity Store 中找 Alice,找到和 Alice 连接的 memories,并给它们加权。
流程大概是:
这类能力对长期记忆很重要。因为真实问题经常围绕实体展开:某个人、某家公司、某个项目、某个地点、某个概念。
向量检索能找到语义类似的内容,但实体匹配能帮助系统找到“围绕同一对象分散在不同对话里的记忆”。
***
### 第五步:时间推理,解决“什么时候为真”
长期记忆系统一定会遇到时间问题。
用户的偏好会变,住址会变,工作会变,项目状态会变。
```text
以前住在纽约,现在住在旧金山。
以前喜欢咖啡,现在戒咖啡了。
上个月在做 A 项目,现在转到 B 项目。
```
所以召回时不能只问:
```text
哪条 memory 最像 query?
```
还要问:
```text
这条 memory 在时间上是否适合当前问题?
```
Mem0 Platform v3 的 Temporal Reasoning 处理的就是这类问题:
```text
last week
next month
right now
currently
as of March 2025
since when
upcoming
```
比如:
```text
用户问:我现在住在哪里?
```
系统应该更偏向:
```text
用户 2026 年搬到了旧金山。
```
而不是:
```text
用户 2025 年住在纽约。
```
但如果用户问:
```text
我搬家之前住在哪里?
```
旧记忆又应该被召回。
这就是为什么 ADD-only 和时间推理要放在一起理解:ADD-only 保留历史,Temporal Reasoning 决定在当前问题下哪段历史更相关。
***
### 第六步:Memory Decay,解决旧记忆污染
长期记忆还有一个问题:旧记忆会越来越多。
有些旧记忆仍然重要,比如过敏信息;有些旧记忆只是临时信息,比如上季度的项目名称。如果所有记忆永远平等,旧信息就会污染召回。
Mem0 的 Memory Decay 可以理解成一种搜索时的轻量排序偏置:
```text
刚被召回过的记忆 → 得到轻微增强
长期没被触达的记忆 → 得到轻微降权
```
注意,它不是删除,也不是硬过滤。
这很像人类记忆:我们不是彻底忘掉过去,而是最近频繁使用的信息更容易浮现。
不过 decay 不能太强。因为低频信息不代表不重要,比如医疗过敏、法律约束、安全偏好。一个好的 memory decay 应该是“温和影响排序”,而不是决定生死。
***
### 第七步:Criteria Retrieval,让“相关性”变成业务定义
Mem0 还有一个高级能力叫 Criteria Retrieval。它的核心思想是:有些应用里,“相关”不只是语义相似。
比如一个心理健康助手,可能更想优先召回带有情绪信号的记忆;一个学习助手可能更关心“好奇心”和“挫败感”;一个客服 Agent 可能更关心“紧急程度”和“风险”。
也就是说,普通检索问的是:
```text
这条 memory 和 query 有多像?
```
Criteria Retrieval 问的是:
```text
这条 memory 是否符合我这个应用定义的重要性?
```
可以把它理解成业务层的 relevance shaping。
这说明记忆召回最后一定会走向业务定制。因为不同 Agent 对“重要记忆”的定义不一样。
***
### 第八步:Rerank,解决 top 结果顺序
前面的语义、关键词、实体、时间、衰减都可以产生一个综合分数。但在一些高风险或高体验要求的场景里,还需要 rerank。
Rerank 的作用是把候选结果重新排序,让最应该进入上下文的 memory 排在前面。
适合开启 rerank 的场景:
```text
用户只能看到少量结果
第 1 条结果的质量非常关键
医疗、客服、付费用户体验等高精度场景
复杂 query 需要更强语义判断
```
不适合无脑开启的原因也很简单:它会增加延迟。
所以 rerank 应该是一种按场景打开的精度增强,而不是默认万能药。
***
### 最终:记忆如何进入 LLM 上下文
召回不是终点。召回出来的 memories 最终要进入 LLM 的上下文。
一个常见结构是:
```text
系统指令:
你是一个有长期记忆的助手。以下是和当前用户相关的历史记忆。
用户记忆:
- 用户不喜欢惊悚片。
- 用户偏好科幻电影。
- 用户下个月计划去东京。
- 用户不吃贝类。
当前用户输入:
今晚有什么推荐?
```
这一步看起来简单,但也有几个问题:
```text
召回多少条?
按什么顺序放?
是否区分事实、偏好、计划和历史事件?
是否告诉模型这些记忆可能过时?
是否允许模型基于记忆主动提问确认?
```
记忆注入的质量会直接影响回答质量。召回太少,模型缺上下文;召回太多,模型会被噪音干扰。
所以长期记忆的目标不是“召回尽可能多”,而是“召回刚好够用”。
***
## 五、用一张总图理解 Mem0
把写入和召回放在一起,Mem0 的整体链路可以画成这样:
这张图里最重要的是:Mem0 不是单一检索器,而是一条闭环。
```text
对话产生记忆
记忆影响未来对话
未来对话继续产生新记忆
```
这就是 Agent 长期记忆的基本循环。
***
## 六、我对 Mem0 的初步判断
我现在对 Mem0 的理解是:它不是一个“记忆数据库”,而是一个**记忆排序系统**。
它真正解决的是:
```text
在大量、分散、跨时间、跨实体、可能过时的记忆里,
为当前 query 找出最应该进入上下文的几条。
```
它的优点是工程化程度很高,几乎把生产环境里会遇到的问题都拆成了对应模块:
```text
多用户隔离 → Entity-Scoped Memory
语义召回 → Vector Search
关键词命中 → BM25
实体相关 → Graph Memory / Entity Linking
时间变化 → Temporal Reasoning
旧记忆污染 → Memory Decay
业务重要性 → Criteria Retrieval
高精度排序 → Rerank
```
它的局限也很清楚:
第一,Graph Memory 更像实体增强检索,不是完整知识图谱。
第二,ADD-only 会让 memory 增长,后续必须依赖排序、衰减和清理策略。
第三,官方 benchmark 很有参考价值,但真正上线前仍然要在自己的业务数据上评估。
第四,隐私、同意、可见性、错误记忆纠正,这些治理问题不是接入 Mem0 就自动解决,应用层仍然要认真设计。
所以我对它的定位是:
> Mem0 很适合作为学习 Agent Memory 工程化的第一个样本。它展示了现代长期记忆系统应该有哪些模块,但它不是所有记忆问题的最终答案。
***
## 七、以后拆其他记忆系统,可以沿用这套问题
这篇是系列第一篇。后面不管研究 Zep、Letta、LangGraph 还是 Cognee,我都会尽量用同一套问题拆:
```text
1. 它怎么写入记忆?
原始对话、摘要、事实抽取,还是结构化事件?
2. 它怎么存储记忆?
向量库、全文索引、图数据库、SQL、对象存储?
3. 它怎么划分记忆边界?
user、agent、session、app、organization?
4. 它怎么召回?
向量、BM25、图遍历、实体匹配、时间推理、工具调用?
5. 它怎么排序?
score fusion、rerank、decay、recency、业务规则?
6. 它怎么处理变化?
update、delete、ADD-only、时间版本、事实失效?
7. 它怎么把记忆交给模型?
自动注入、工具调用、上下文块、prompt 模板?
8. 它有什么治理能力?
可见、可删、可审计、可纠错、权限隔离?
```
如果用这套框架看 Mem0,可以得到一句总结:
> Mem0 的核心是:写入时把对话蒸馏成事实,存储时建立语义和实体索引,召回时融合语义、关键词、实体、时间、衰减和业务信号,再把最相关的记忆注入上下文。
这也是我理解 Agent Memory 的第一块拼图。
***
## 参考资料
* [Mem0 Memory Types](https://docs.mem0.ai/core-concepts/memory-types)
* [Mem0 Add Memory](https://docs.mem0.ai/core-concepts/memory-operations/add)
* [Mem0 Search Memory](https://docs.mem0.ai/core-concepts/memory-operations/search)
* [Mem0 Graph Memory](https://docs.mem0.ai/platform/features/graph-memory)
* [Mem0 Entity-Scoped Memory](https://docs.mem0.ai/platform/features/entity-scoped-memory)
* [Mem0 Temporal Reasoning](https://docs.mem0.ai/platform/features/temporal-reasoning)
* [Mem0 Memory Decay](https://docs.mem0.ai/platform/features/memory-decay)
* [Mem0 Criteria Retrieval](https://docs.mem0.ai/platform/features/criteria-retrieval)
* [Mem0 Memory Evaluation](https://docs.mem0.ai/core-concepts/memory-evaluation)
* [Mem0 arXiv Paper](https://arxiv.org/abs/2504.19413)
# Agent 记忆系统学习笔记(二):从 Zep 看懂时间图谱记忆
## 写在前面:为什么第二篇看 Zep
上一篇我用 Mem0 拆了一遍 Agent 记忆系统的基本链路:对话如何被抽取成 memory,memory 如何被存储,未来又如何通过语义、关键词、实体、时间、衰减和 rerank 找回来。
这一篇换一个视角:**如果记忆不是一条条事实,而是一张会随时间变化的图,会发生什么?**
这就是 Zep 最值得分析的地方。它的核心不是“给 memory 加一个图增强”,而是把用户、业务数据和 Agent 工作记录组织成一张 **temporal Context Graph**。在这张图里,实体是节点,事实和关系是边,原始数据是 episode,边上还带着事实何时开始有效、何时失效的时间信息。
所以这篇不是 Zep 接入教程,而是继续这个系列的学习目标:用 Zep 作为样本,理解**时间图谱记忆**这条技术路线。
***
## 一、Zep 的核心问题:长期记忆为什么需要图
在 Mem0 那篇里,我把 Mem0 理解成一个“记忆排序系统”。它的基本单元是 memory fact:一条被抽取出来、未来可以召回的事实。
但有些问题用单条 fact 很难表达清楚。
比如:
```text
用户原来在 A 公司。
用户后来加入 B 公司。
用户和 Alice 一起负责 Q1 路线图。
Alice 后来转去负责定价项目。
Q1 路线图因为预算问题延期。
```
如果只是把这些都存成独立 memory,当然也能搜。但真正的难点是:这些事实之间有关系,而且关系会随时间变化。
你问:
```text
用户现在和 Alice 还在同一个项目吗?
```
这不是单纯的语义相似问题。系统需要知道:
```text
用户是谁
Alice 是谁
他们之前有什么关系
这个关系从什么时候开始
是否已经结束
Q1 路线图和定价项目分别是什么
哪个事实代表当前状态
```
Zep 的答案是:不要只存一堆 memory,而是构建一张带时间的 Context Graph。
从这个角度看,Zep 不是单纯的 memory store,而是一个把交互数据、业务数据和历史状态持续编织成图的系统。
***
## 二、Zep 的核心抽象:Context Graph
Zep 文档里反复提到一个概念:**Context Graph**。
可以把它理解成某个用户、对象或业务主体的一张时间知识图谱。
这张图里主要有三类东西:
### Entity Nodes:实体节点
节点代表实体,比如用户、公司、地点、项目、产品、账号、概念。
和普通图谱不同的是,Zep 的实体节点不只是一个名字。它还会维护围绕这个实体的叙事摘要,比如这个用户是谁、这个项目发生了什么、这个账号有哪些历史状态。
### Entity Edges / Facts:事实边
边代表两个实体之间的关系,而 fact 存在边上。
比如:
```text
Emily 的账号因为支付失败被暂停。
```
这里可能有两个实体:
```text
Emily 的账号
支付失败
```
关系边上保存事实:
```text
User account Emily0e62 has a suspended status due to payment failure.
```
更重要的是,这条 fact 不是一个静态字符串,它带时间属性。
### Episodes:原始证据
Episode 是开发者交给 Zep 的原始数据:聊天消息、文本片段、JSON 记录、邮件、工单、CRM 数据等。
这点很重要。Zep 并不是抽取完 fact 就把原始数据扔掉。Episode 会被保留,作为事实背后的证据。以后如果 Agent 需要引用原话、找上下文、追溯来源,就可以回到 episode。
所以 Zep 的结构不是:
```text
对话 → memory
```
而更像:
```text
episode → entities + facts + summaries + observations → context block
```
***
## 三、写入链路:episode 如何变成时间图谱
Zep 支持几种数据入口。
一种是聊天消息:
```text
thread.add_messages
```
另一种是业务数据:
```text
graph.add(message / text / JSON)
```
这说明 Zep 的目标不只是“记住聊天”,还包括把业务系统里的动态数据也放进同一张图。
这里最关键的概念是 episode。
每次你往 Zep 里加一段消息、文本或 JSON,它都会先成为一个 episode。然后系统从 episode 里抽取实体、关系和事实,并把它们写进 Context Graph。
举个例子,输入一条业务 JSON:
```json
{
"event": "payment_failed",
"account": "Emily0e62",
"reason": "Card expired",
"amount": 99.99,
"created_at": "2024-09-15T00:00:00Z"
}
```
Zep 可能从里面抽出:
```text
账号 Emily0e62 发生了一笔 99.99 的失败交易。
失败原因是 Card expired。
这个事件发生在 2024-09-15。
```
这些不是简单存在一个列表里,而会进入图:账号、交易、失败原因可能成为实体或关系的一部分,事实被挂在边上,时间字段也进入事实生命周期。
这就是 Zep 和普通 memory fact 系统的第一个大差异:**Zep 从源头上把记忆当成图结构来生成,而不是事后给 memory 做实体连接。**
***
## 四、Zep 的时间模型:事实不是覆盖,而是失效
Zep 最值得学习的设计,是它对事实变化的处理。
很多系统遇到状态变化时会做 update:
```text
旧事实:用户喜欢 Adidas
新事实:用户喜欢 Puma
```
然后旧事实被覆盖。
但 Zep 的思路不是简单覆盖,而是让事实带生命周期。
Zep 的 fact 边上有四个时间字段:
| 时间字段 | 含义 |
| ----------- | ---------------- |
| created\_at | Zep 什么时候知道这个事实 |
| valid\_at | 这个事实在现实中什么时候开始为真 |
| invalid\_at | 这个事实在现实中什么时候不再为真 |
| expired\_at | Zep 什么时候知道这个事实失效 |
这个设计非常重要,因为它区分了两个时间:
```text
现实中什么时候发生
系统什么时候知道
```
比如用户 6 月 1 日搬家,但 6 月 10 日才告诉 Agent。那么:
```text
valid_at = 6 月 1 日
created_at = 6 月 10 日
```
如果 9 月 1 日又搬走,但 9 月 5 日才告诉 Agent:
```text
invalid_at = 9 月 1 日
expired_at = 9 月 5 日
```
这比“最新事实覆盖旧事实”细得多。它让系统可以回答:
```text
用户现在住在哪里?
用户 8 月时住在哪里?
系统是什么时候知道用户搬家的?
```
这也是 Zep / Graphiti 路线最有价值的地方:它不是只记住事实,而是记住事实的时间边界。
***
## 五、召回链路:Zep 怎么把图谱变成上下文
Zep 的召回和 Mem0 很不一样。
Mem0 更像:
```text
query → 搜 memory → 返回 top memories → 应用注入 prompt
```
Zep 更像:
```text
query 或最近消息 → 搜 Context Graph → 选择多种 context type → 组装 Context Block → 直接放进 prompt
```
Zep 有两种常见使用方式。
第一种是高层方法:
```text
thread.get_user_context()
```
它会根据当前 thread 最近几条消息,在整个 User Graph 里找相关上下文,然后返回一个 prompt-ready 的 Context Block。
第二种是低层方法:
```text
graph.search()
```
你可以指定要搜 facts、entities、episodes、observations、thread summaries,或者使用 `scope="auto"` 让 Zep 自己跨类型选择并组装上下文。
这背后用到几类信号:
```text
semantic similarity:语义相似
BM25 full-text:关键词精确匹配
Breadth-first search:从指定节点出发做图连接相关性
RRF / MMR / cross encoder / node distance 等 reranker
时间过滤:created_at、valid_at、invalid_at、expired_at
```
所以 Zep 的召回不是简单“图遍历”,而是混合检索:语义、关键词、图距离、时间过滤、跨类型 rerank 一起工作。
***
## 六、Context Types:Zep 召回的不只是事实
Zep 还有一个很值得学习的点:它把上下文拆成多种形态。
这比“只返回 memory list”更细。
### Facts:精确事实
适合回答具体问题,比如:
```text
用户账号为什么被暂停?
用户什么时候开始使用某个工具?
某个项目的决定是什么?
```
Facts 的优势是精确,而且有时间范围。
### Entities:实体摘要
适合回答围绕某个实体的问题,比如:
```text
我们知道 Alice 的哪些信息?
这个项目过去发生过什么?
这个账号的整体状态是什么?
```
实体摘要不是某一条 fact,而是围绕一个节点的叙事。
### Episodes:原始证据
适合需要原话和来源的时候。
比如 Agent 不能只说“用户投诉过登录问题”,它还需要引用当时用户怎么说的,就应该回到 episode。
### Thread Summaries:会话摘要
适合恢复某个历史会话。
它回答的是:这条 thread 里发生过什么,而不是全局用户画像。
### Observations:跨事实形成的稳定模式
这是 Zep 比较有意思的一层。Observation 不是单条 fact,而是从多个 facts 和 entities 中总结出的稳定模式、承诺、决策或状态变化。
比如:
```text
用户经常在账单失败后需要人工支持。
用户对支付可靠性非常敏感。
某个项目持续被预算问题影响。
```
这种信息如果只看单条 fact,会显得很碎;但作为 observation,就能表达“长期模式”。
### User Summary:用户基线画像
Zep 的默认 Context Block 会包含 user summary。它提供一个用户的长期基线画像,让 Agent 不用每次都从零理解用户是谁。
这套 context type 的拆分,说明 Zep 的目标不是简单召回,而是做 context assembly:根据当前问题选择合适形态的上下文。
***
## 七、Context Block:Zep 不只是返回检索结果
Zep 的一个关键产品化设计是 Context Block。
它不是让你拿到一堆搜索结果后自己拼 prompt,而是直接返回一个可以塞进系统提示词的字符串。
典型结构类似:
```text
用户是谁、长期偏好是什么、当前账户状态如何……
- 某条相关事实。(Date range: from - to)
- 某条相关事实。(Date range: from - to)
```
这件事很重要,因为很多 memory 系统只解决“搜到什么”,但没有解决“怎么把搜到的东西喂给模型”。
而 Zep 直接把召回和 prompt 组装连在一起:
```text
检索结果 → context selection → 格式化 → 字符预算控制 → Context Block
```
这对应用开发者很友好,但也有取舍:默认 Context Block 更方便,定制性则需要通过 context template 或 advanced context construction 来做。
我理解它的定位是:
> Zep 不只是 graph retrieval,而是面向 Agent 的上下文工程系统。
***
## 八、和 Mem0 的关键差异
看完 Zep,再回头看 Mem0,两者差异会很清楚。
| 维度 | Mem0 | Zep / Graphiti |
| ---- | ------------------------ | --------------------------------------------- |
| 核心抽象 | memory fact | temporal Context Graph |
| 图的角色 | 实体连接,用于 ranking boost | 图是主要记忆结构 |
| 时间处理 | retrieval ranking 中的时间信号 | fact 边自带 valid\_at / invalid\_at |
| 原始证据 | memory 和历史记录 | episode 是一等对象 |
| 召回结果 | memories | Context Block / facts / entities / episodes 等 |
| 更适合 | 个性化偏好、长期事实、轻量 memory 层 | 关系变化、状态演化、多源业务数据、时间查询 |
简单说:
```text
Mem0 更像一个产品化 memory ranking layer。
Zep 更像一个 temporal graph + context assembly layer。
```
这不是谁更好,而是抽象层不同。
如果你的应用主要是记用户偏好、历史任务、简单事实,Mem0 的 mental model 更直接。
如果你的应用里有大量实体关系、业务事件、状态变化、跨系统数据,Zep 的图谱模型会更自然。
***
## 九、我对 Zep 的判断
我觉得 Zep 最值得学习的不是“用了知识图谱”,而是它把长期记忆里的几个难题都放进了图模型:
第一,事实不是孤立的,而是实体之间的关系。
第二,事实不是永远为真,而是有时间边界。
第三,原始数据不能丢,episode 要作为证据保留。
第四,召回不是单一结果类型,而是按任务选择 facts、entities、episodes、observations、summary。
第五,最终目标不是返回数据库记录,而是组装出 LLM 能直接使用的 Context Block。
它适合的场景是:
```text
客户支持:账号、工单、支付、历史问题持续变化
销售 / CRM:人、公司、机会、会议、承诺之间关系复杂
医疗 / 教育:用户状态和偏好随时间演化
企业 Agent:业务系统 JSON、文档、聊天记录要汇入同一上下文
长周期项目管理:项目状态、负责人、决策和风险不断变化
```
它不一定适合的场景是:
```text
只需要轻量用户偏好记忆
不需要关系建模和时间有效性
希望完全自托管且不想接入托管服务
需要自己控制每个图数据库细节
```
另外要注意一点:Zep 是商业托管产品,Graphiti 是它背后的开源 temporal knowledge graph 框架。研究技术路线时可以把 Zep 当成“产品化图谱记忆”,把 Graphiti 当成“开源图谱引擎”。如果后面要自建,Graphiti 会是更值得深入看的部分。
***
## 十、这一篇之后,我的分析框架更新了
上一篇 Mem0 让我看到一个 memory layer 需要回答:
```text
怎么抽取 memory?
怎么多信号召回?
怎么排序?
怎么注入上下文?
```
这一篇 Zep 让我加上几个新问题:
```text
这个系统有没有显式实体和关系?
事实是否有时间生命周期?
原始证据是否被保留?
它返回的是 memory list,还是 prompt-ready context?
它能否表达跨事实的长期模式?
```
所以,到目前为止,这个系列里可以形成一个初步判断:
> 轻量长期记忆看 memory ranking;复杂长期记忆看 temporal graph;真正面向 Agent 的系统,最后都要走向 context assembly。
下一篇我会看 Letta / MemGPT。它和 Mem0、Zep 又不一样:它关注的是 Agent 如何主动管理自己的 core memory 和 archival memory。
***
## 参考资料
* [Zep Key Concepts](https://help.getzep.com/concepts)
* [Zep Graph Overview](https://help.getzep.com/graph-overview)
* [Zep Facts](https://help.getzep.com/facts)
* [Zep Adding Messages](https://help.getzep.com/adding-messages)
* [Zep Adding Business Data](https://help.getzep.com/adding-business-data)
* [Zep Episodes](https://help.getzep.com/episodes)
* [Zep Searching the Graph](https://help.getzep.com/searching-the-graph)
* [Zep Retrieving Context](https://help.getzep.com/retrieving-context)
* [Zep Context Types](https://help.getzep.com/context-types)
* [Zep Observations](https://help.getzep.com/observations)
# Introduction au Concept
## Introduction
Lorsque vous avez configuré un workflow parfait dans Claude Code — commandes personnalisées, hooks de revue de code, Skills dédiés — vous pourriez vous demander : est-il possible d'empaqueter tout cela et de le partager avec votre équipe ou la communauté ?
C'est exactement le problème que Plugin résout.
Si les Skills sont des « manuels d'instructions » pour l'IA, alors un Plugin est une « boîte à outils » : il regroupe les Skills, les Commands, les Hooks, les serveurs MCP et toutes les autres configurations pour que vous puissiez tout installer et distribuer en une seule commande.
## Comprendre Plugin
Imaginez que vous êtes un artisan expérimenté qui a accumulé au fil des années un ensemble d'outils de confiance : marteaux, scies, règles, divers tournevis. Chaque fois que vous changez d'établi, vous devez transporter chaque outil un par un et tout réorganiser. Un Plugin est comme une boîte à outils bien conçue : elle contient non seulement tous vos outils, mais les maintient aussi classés par catégorie, prêts à l'emploi où que vous alliez.
D'un point de vue technique, Plugin est le mécanisme d'empaquetage d'extensions de Claude Code. Un Plugin peut contenir :
| Composant | Fonction | Emplacement du fichier |
| ------------------- | -------------------------------------------------- | ---------------------- |
| **Commandes slash** | Points d'entrée pour actions rapides | `commands/` |
| **Subagents** | Sous-agents spécialisés | `agents/` |
| **Skills** | Paquets de connaissances pour l'IA | `skills/` |
| **Hooks** | Scripts d'automatisation déclenchés par événements | `hooks/` |
| **Serveurs MCP** | Connexions aux systèmes externes | `.mcp.json` |
| **Serveurs LSP** | Configuration des serveurs de langage | `.lsp.json` |
Ces composants fonctionnent ensemble pour former une solution de workflow complète.
## Plugin vs Configuration autonome
Dans Claude Code, vous pouvez placer les configurations dans le répertoire `.claude/` du projet ou les empaqueter en tant que Plugin. La différence fondamentale réside dans la **méthode de distribution** et l'**espace de noms** :
| Aspect | Configuration autonome (`.claude/`) | Plugin |
| ---------------------- | -------------------------------------------------------- | -------------------------------------------- |
| Nom de la commande | `/hello` | `/plugin-name:hello` |
| Cas d'utilisation | Workflows personnels, configuration spécifique au projet | Partage d'équipe, distribution communautaire |
| Gestion des versions | Gérée avec le code du projet | Supporte le Semantic Versioning |
| Méthode de mise à jour | Synchronisation manuelle | Supporte les mises à jour automatiques |
| Gestion des conflits | Peut entrer en conflit avec d'autres configurations | Isolation par espace de noms |
**Quand choisir Plugin** :
* Vous devez partager des configurations de workflow avec les membres de votre équipe
* Vous souhaitez réutiliser le même ensemble d'outils dans plusieurs projets
* Vous prévoyez de distribuer des configurations à la communauté
* Vous avez besoin du contrôle de version et des mises à jour automatiques
**Quand utiliser la configuration autonome** :
* Expérimentations rapides à usage personnel
* Configurations spécifiques au projet qui n'ont pas besoin d'être réutilisées
* Commandes simples à usage unique
## Structure de répertoires du Plugin
La structure standard d'un Plugin se présente comme suit :
```
my-plugin/
├── .claude-plugin/ # 元数据目录
│ └── plugin.json # 必需:插件清单
├── commands/ # 斜杠命令
│ ├── review.md
│ └── deploy.md
├── agents/ # 子代理
│ └── code-reviewer.md
├── skills/ # Agent Skills
│ └── code-review/
│ └── SKILL.md
├── hooks/ # 事件钩子
│ └── hooks.json
├── scripts/ # 辅助脚本
│ └── format-code.sh
├── .mcp.json # MCP 服务器配置
└── .lsp.json # LSP 服务器配置
```
**Points importants** :
* `plugin.json` doit être placé dans le répertoire `.claude-plugin/`
* Les autres répertoires (commands, agents, skills, etc.) se trouvent à la racine du plugin
* Ne placez pas les répertoires fonctionnels dans `.claude-plugin/`
## Fichier de configuration principal
Le cœur d'un Plugin est `.claude-plugin/plugin.json`, qui définit les métadonnées et les chemins des composants du plugin :
```json
{
"name": "my-awesome-plugin",
"version": "1.0.0",
"description": "一个示例插件",
"author": {
"name": "Your Name",
"email": "you@example.com"
},
"keywords": ["example", "demo"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
| Champ | Requis | Description |
| ------------- | ------ | --------------------------------------------------------------------------- |
| `name` | Oui | Identifiant unique du plugin, utilisez des lettres minuscules et des tirets |
| `version` | Non | Numéro de version sémantique |
| `description` | Non | Description courte du plugin |
| `author` | Non | Informations sur l'auteur |
| `keywords` | Non | Tags pour la découvrabilité |
| `commands` | Non | Chemin du fichier ou répertoire de commandes |
| `agents` | Non | Chemin du fichier ou répertoire d'agents |
| `skills` | Non | Chemin du répertoire Skills |
| `hooks` | Non | Chemin de configuration des hooks |
| `mcpServers` | Non | Chemin de configuration MCP |
## Portées d'installation
Plugin supporte quatre portées d'installation pour s'adapter aux différents cas d'utilisation :
| Portée | Fichier de configuration | Objectif |
| --------- | ----------------------------- | ----------------------------------------------------- |
| `user` | `~/.claude/settings.json` | Plugins personnels, disponibles dans tous les projets |
| `project` | `.claude/settings.json` | Plugins d'équipe, partagés via le contrôle de version |
| `local` | `.claude/settings.local.json` | Spécifique au projet, dans le gitignore |
| `managed` | `managed-settings.json` | Gestion d'entreprise (lecture seule) |
La portée d'installation par défaut est `user`. Si vous souhaitez enregistrer la configuration du plugin dans Git pour une utilisation en équipe, choisissez la portée `project`.
## Avantages principaux
### Isolation des espaces de noms
Les commandes du Plugin portent un préfixe d'espace de noms (par exemple, `/my-plugin:review`), évitant les conflits de noms avec d'autres plugins ou configurations de projet. Cela est particulièrement important dans la collaboration d'équipe : les plugins développés par différentes équipes peuvent coexister harmonieusement.
### Gestion des versions
Plugin supporte le Semantic Versioning, vous permettant de :
* Suivre l'historique des modifications du plugin
* Revenir à des versions antérieures si nécessaire
* Recevoir automatiquement les mises à jour compatibles
### Distribution facilitée
Via le Plugin Marketplace, vous pouvez :
* Héberger vos plugins sur GitHub
* Permettre aux utilisateurs d'installer avec une simple commande
* Gérer automatiquement les dépendances et les mises à jour
### Collaboration d'équipe
Plugin est particulièrement adapté aux scénarios d'équipe :
* Unifier la chaîne d'outils de développement de l'équipe
* Les nouveaux membres obtiennent tous les outils en une seule commande
* La gestion centralisée des configurations réduit le travail en double
## Écosystème Plugin
L'écosystème Plugin de Claude Code se développe rapidement. Début 2025, l'écosystème a atteint une échelle considérable :
* **229+ plugins** actifs dans l'écosystème
* **239 Agent Skills** distribués sur le marketplace
* **200+ serveurs MCP** préconstruits dans le toolkit Docker
**Ressources officielles** :
| Ressource | Lien | Description |
| ---------------------------------- | ---------------------------------------------------------------------------------------- | ------------------------------- |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Dépôt officiel de Skills |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Répertoire officiel de plugins |
| Docker MCP Toolkit | [Site web](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | 200+ serveurs MCP préconstruits |
**Sélections de la communauté** :
| Ressource | Lien | Description |
| ---------------------- | ----------------------------------------------------------------- | ------------------------------------- |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | 243 plugins collectés automatiquement |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Compilation des meilleures pratiques |
| claude-plugins.dev | [Site web](https://claude-plugins.dev/) | Registre communautaire et CLI |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agents + 15 orchestrateurs |
| Compound Engineering | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | 17 agents spécialisés |
## Relation avec les autres fonctionnalités
Plugin est un concept de « conteneur » qui peut inclure d'autres fonctionnalités de l'écosystème Claude Code :
```
Plugin (Conteneur)
├── Skills (Paquets de connaissances)
├── Commands (Commandes rapides)
├── Agents (Sous-agents)
├── Hooks (Hooks d'événements)
└── MCP/LSP (Connexions externes)
```
Comprendre cette hiérarchie est important :
* **Skills** enseignent à Claude comment faire quelque chose
* **Commands** fournissent des points d'entrée rapides
* **Agents** gèrent des tâches indépendantes et spécialisées
* **Hooks** permettent l'automatisation pilotée par événements
* **Plugin** regroupe tout cela pour faciliter la distribution et la gestion
### Skills vs Plugins
La différence entre Skills et Plugins peut prêter à confusion au début. Selon l'analyse de [Young Leaders Tech](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) :
| Caractéristique | Skills | Plugins |
| ---------------- | --------------------------------------------- | -------------------------------------- |
| **Portée** | Tous les produits Claude (Web, API, Code) | Claude Code uniquement |
| **Contenu** | Guides Markdown + scripts optionnels | Commands, Agents, Hooks, MCP, Skills |
| **Activation** | Automatique (le modèle décide quand utiliser) | Variable (dépend du type de composant) |
| **Idéal pour** | Enseigner une expertise de domaine à Claude | Étendre l'environnement Claude Code |
| **Distribution** | Dépôts GitHub, système de fichiers | Marketplace décentralisé |
**Point clé** : Les Skills sont déclenchés automatiquement par le modèle sans invocation manuelle ; les Plugins sont un mécanisme d'empaquetage qui résout le défi du partage distribué. Les deux peuvent être utilisés ensemble : un Plugin peut contenir des Skills.
## Résumé
Claude Code Plugin est essentiellement un **mécanisme d'empaquetage et de distribution de workflows**. Il résout les problèmes de réutilisation des configurations et de collaboration d'équipe, vous permettant de partager votre chaîne d'outils soigneusement élaborée avec davantage de personnes.
Retenez trois mots-clés :
| Mot-clé | Signification |
| ---------------- | ---------------------------------------------------------------- |
| **Empaquetage** | Intègre plusieurs composants de configuration en une seule unité |
| **Isolation** | Les espaces de noms évitent les conflits |
| **Distribution** | Partage facile via le Marketplace |
Maintenant que vous comprenez les concepts, le prochain article, [Guide pratique de Claude Code Plugin](/fr/docs/notes/claude-plugin/practice), vous guidera dans la pratique : créer un Plugin de zéro, publier sur le Marketplace et appliquer les meilleures pratiques de collaboration d'équipe.
Si vous n'êtes pas encore familier avec les composants qu'un Plugin peut contenir, nous vous recommandons de lire [Qu'est-ce que Claude Skills](/fr/docs/notes/claude-skills/concept) pour comprendre les concepts fondamentaux des Skills.
# Guide Pratique
## Rappel rapide
Dans l'article précédent, nous avons exploré les concepts fondamentaux des Plugins : il s'agit du mécanisme de Claude Code pour empaqueter et distribuer des workflows, regroupant Commands, Skills, Agents, Hooks et d'autres composants en une seule unité pour le partage en équipe et la distribution communautaire. Cet article adopte une approche pratique et vous guide à travers le processus complet, de la création à la publication.
## Créer votre premier Plugin
### Etape 1 : Créer la structure de répertoires
```bash
mkdir my-first-plugin
mkdir my-first-plugin/.claude-plugin
mkdir my-first-plugin/commands
```
### Etape 2 : Créer le manifeste du plugin
Définissez les métadonnées de votre plugin dans `.claude-plugin/plugin.json` :
```json
{
"name": "my-first-plugin",
"description": "我的第一个 Claude Code 插件",
"version": "1.0.0",
"author": {
"name": "Your Name"
}
}
```
### Etape 3 : Ajouter des commandes slash
Créez des fichiers Markdown dans le répertoire `commands/`. Chaque fichier correspond à une commande :
`commands/hello.md` :
```markdown
---
description: 向用户发送友好的问候
---
# Hello 命令
请热情地问候用户,并询问今天可以帮助他们做什么。
```
### Etape 4 : Tester le plugin
Utilisez le drapeau `--plugin-dir` pour charger votre plugin local et le tester :
```bash
claude --plugin-dir ./my-first-plugin
```
Exécutez la commande dans Claude Code :
```
/my-first-plugin:hello
```
### Etape 5 : Ajouter des arguments de commande
Les commandes prennent en charge les arguments fournis par l'utilisateur. Mettez à jour `hello.md` :
```markdown
---
description: 向指定用户发送个性化问候
---
# Hello 命令
请热情地问候名为 "$ARGUMENTS" 的用户,并询问今天可以帮助他们做什么。
如果用户没有提供名字,就使用"朋友"作为称呼。
```
Testez la commande avec des arguments :
```
/my-first-plugin:hello 小明
```
**Marqueurs d'arguments pris en charge** :
* `$ARGUMENTS` - Toute la saisie de l'utilisateur
* `$1`, `$2`, `$3` - Arguments individuels
## Ajouter d'autres composants
### Ajouter des Skills
Créez un répertoire `skills/`. Chaque Skill est un dossier contenant un fichier `SKILL.md` :
```
my-first-plugin/
├── skills/
│ └── code-review/
│ └── SKILL.md
```
`skills/code-review/SKILL.md` :
```yaml
---
name: code-review
description: 审查代码质量、安全性和可维护性
---
当审查代码时,请检查以下方面:
1. **代码组织**:结构是否清晰
2. **错误处理**:异常是否被妥善处理
3. **安全隐患**:是否存在安全漏洞
4. **测试覆盖**:关键逻辑是否有测试
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Ajouter des Subagents
Créez un répertoire `agents/` :
`agents/code-reviewer.md` :
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家。
当被调用时:
1. 运行 git diff 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(如注入、敏感信息泄露)
- 性能优化机会
```
### Ajouter des Hooks
Les Hooks vous permettent d'exécuter automatiquement des scripts lorsque des événements spécifiques se produisent. Créez `hooks/hooks.json` :
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Write|Edit",
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PLUGIN_ROOT}/scripts/format-code.sh"
}
]
}
]
}
}
```
**Important** : Utilisez la variable d'environnement `${CLAUDE_PLUGIN_ROOT}` pour référencer les fichiers au sein du répertoire de votre plugin, garantissant que les chemins se résolvent correctement quel que soit l'emplacement d'installation du plugin.
Créez le script correspondant `scripts/format-code.sh` :
```bash
#!/bin/bash
# 格式化代码
cd "$CLAUDE_PROJECT_DIR" && make format 2>/dev/null || true
```
N'oubliez pas de définir les permissions d'exécution :
```bash
chmod +x scripts/format-code.sh
```
### Ajouter des serveurs MCP
Si votre plugin a besoin de se connecter à des systèmes externes, créez `.mcp.json` :
```json
{
"mcpServers": {
"plugin-database": {
"command": "${CLAUDE_PLUGIN_ROOT}/servers/db-server",
"args": ["--config", "${CLAUDE_PLUGIN_ROOT}/config.json"],
"env": {
"DB_PATH": "${CLAUDE_PLUGIN_ROOT}/data"
}
}
}
}
```
## Structure complète du Plugin
Un Plugin entièrement fonctionnel pourrait ressembler à ceci :
```
my-plugin/
├── .claude-plugin/
│ └── plugin.json # 插件清单
├── commands/
│ ├── review.md # 代码审查命令
│ ├── deploy.md # 部署命令
│ └── test.md # 测试命令
├── agents/
│ ├── code-reviewer.md # 代码审查代理
│ └── debugger.md # 调试代理
├── skills/
│ └── code-standards/
│ └── SKILL.md # 代码规范知识
├── hooks/
│ └── hooks.json # 事件钩子配置
├── scripts/
│ ├── format-code.sh # 格式化脚本
│ └── run-tests.sh # 测试脚本
├── .mcp.json # MCP 配置
├── LICENSE
├── README.md
└── CHANGELOG.md
```
Le `plugin.json` correspondant :
```json
{
"name": "dev-toolkit",
"version": "1.0.0",
"description": "开发者工具箱:代码审查、测试、部署一站式解决方案",
"author": {
"name": "Your Team",
"email": "team@example.com"
},
"homepage": "https://github.com/you/dev-toolkit",
"repository": "https://github.com/you/dev-toolkit",
"license": "MIT",
"keywords": ["development", "code-review", "deployment"],
"commands": "./commands/",
"agents": "./agents/",
"skills": "./skills/",
"hooks": "./hooks/hooks.json",
"mcpServers": "./.mcp.json"
}
```
## Publier sur le Marketplace
### Qu'est-ce que le Marketplace
Le Marketplace est le centre de distribution des Plugins. Vous pouvez le considérer comme une "boutique de plugins" : les utilisateurs peuvent installer vos plugins publiés avec une simple commande.
### Créer la configuration du Marketplace
Créez `.claude-plugin/marketplace.json` dans votre dépôt GitHub :
```json
{
"name": "my-marketplace",
"owner": {
"name": "Your Name",
"email": "you@example.com"
},
"plugins": [
{
"name": "dev-toolkit",
"source": "./plugins/dev-toolkit",
"description": "开发者工具箱",
"version": "1.0.0"
},
{
"name": "doc-generator",
"source": {
"source": "github",
"repo": "you/doc-generator-plugin"
},
"description": "文档生成工具"
}
]
}
```
### Types de sources de plugins
Le Marketplace prend en charge plusieurs types de sources :
**Chemin relatif** (plugins au sein du même dépôt) :
```json
{
"name": "my-plugin",
"source": "./plugins/my-plugin"
}
```
**Dépôt GitHub** :
```json
{
"name": "github-plugin",
"source": {
"source": "github",
"repo": "owner/plugin-repo"
}
}
```
**N'importe quel dépôt Git** :
```json
{
"name": "git-plugin",
"source": {
"source": "url",
"url": "https://gitlab.com/team/plugin.git"
}
}
```
### Processus de publication
1. **Créer un dépôt GitHub**
2. **Pousser votre code** :
```bash
git init
git add .
git commit -m "Initial release"
git push origin main
```
3. **Les utilisateurs ajoutent votre Marketplace** :
```bash
/plugin marketplace add your-username/your-repo
```
4. **Les utilisateurs installent les plugins** :
```bash
/plugin install dev-toolkit@your-marketplace
```
## Installer et gérer les Plugins
### Via le menu interactif
```bash
/plugin
```
Cela ouvre une interface interactive dans laquelle vous pouvez parcourir, installer, activer et désactiver des plugins.
### Via la ligne de commande
**Ajouter un Marketplace** :
```bash
/plugin marketplace add owner/repo # GitHub
/plugin marketplace add https://example.com/marketplace.json # URL
/plugin marketplace add ./local-marketplace # 本地
```
**Installer des plugins** :
```bash
# 安装到用户范围(默认)
/plugin install formatter@my-marketplace
# 安装到项目范围(团队共享)
/plugin install formatter@my-marketplace --scope project
# 安装到本地范围(gitignored)
/plugin install formatter@my-marketplace --scope local
```
**Autres commandes de gestion** :
```bash
/plugin enable # 启用插件
/plugin disable # 禁用插件
/plugin uninstall # 卸载插件
/plugin update # 更新插件
```
### Valider les plugins
Vérifiez que la configuration de votre plugin est correcte avant de publier :
```bash
claude plugin validate .
```
Ou dans Claude Code :
```
/plugin validate .
```
## Configuration pour la collaboration d'équipe
### Partager la configuration des plugins dans un projet
Validez la configuration des plugins dans le contrôle de version pour que les membres de l'équipe l'obtiennent automatiquement :
`.claude/settings.json` :
```json
{
"extraKnownMarketplaces": {
"company-tools": {
"source": {
"source": "github",
"repo": "your-org/claude-plugins"
}
}
},
"enabledPlugins": {
"code-formatter@company-tools": true,
"deployment-tools@company-tools": true
}
}
```
Après le clonage du projet par les membres de l'équipe, ces plugins seront automatiquement disponibles.
### Restrictions Marketplace pour les entreprises
Pour les environnements d'entreprise nécessitant un contrôle strict, vous pouvez restreindre les Marketplaces autorisés dans les paramètres administrés :
```json
{
"strictKnownMarketplaces": [
{
"source": "github",
"repo": "company/approved-plugins"
}
]
}
```
Le définir comme un tableau vide `[]` désactive entièrement les plugins externes.
## Référence des commandes CLI
| Commande | Description |
| -------------------------------------- | ------------------------------------------------------------- |
| `/plugin` | Ouvrir l'interface de gestion interactive |
| `/plugin install @` | Installer un plugin |
| `/plugin uninstall ` | Désinstaller un plugin |
| `/plugin enable ` | Activer un plugin |
| `/plugin disable ` | Désactiver un plugin |
| `/plugin update ` | Mettre à jour un plugin |
| `/plugin validate .` | Valider la configuration du plugin dans le répertoire courant |
| `/plugin marketplace add ` | Ajouter un Marketplace |
| `/plugin marketplace list` | Lister les Marketplaces ajoutés |
| `/plugin marketplace update` | Mettre à jour le cache du Marketplace |
| `/plugin marketplace remove ` | Supprimer un Marketplace |
## Bonnes pratiques
### Bonnes pratiques de développement
1. **Gardez les Skills ciblés** : Chaque Skill doit exceller dans une seule tâche ; évitez les conceptions fourre-tout
2. **Rédigez des descriptions claires** : Aidez Claude à comprendre quand utiliser vos composants
3. **Testez d'abord avec votre équipe** : Validez en interne avant de distribuer à la communauté
4. **Documentez les changements de version** : Consignez les modifications de chaque version dans CHANGELOG.md
### Bonnes pratiques de structure de répertoires
* Placez `commands/`, `agents/`, `skills/` dans le répertoire racine du plugin
* Ne mettez que `plugin.json` dans le répertoire `.claude-plugin/`
* Utilisez `${CLAUDE_PLUGIN_ROOT}` pour référencer les fichiers au sein du plugin
* N'utilisez jamais `../` pour accéder à des fichiers en dehors du plugin
### Bonnes pratiques pour les Hooks
1. Les scripts doivent être exécutables : `chmod +x script.sh`
2. Utilisez un shebang pour déclarer l'interpréteur : `#!/bin/bash`
3. Utilisez la variable `${CLAUDE_PLUGIN_ROOT}` pour garantir l'exactitude des chemins
4. Testez les scripts indépendamment avant de les intégrer aux Hooks
### Bonnes pratiques de gestion des versions
Suivez le Semantic Versioning :
* **MAJOR** (1.0.0 → 2.0.0) : Changements incompatibles
* **MINOR** (1.0.0 → 1.1.0) : Nouvelles fonctionnalités (rétrocompatible)
* **PATCH** (1.0.0 → 1.0.1) : Corrections de bugs (rétrocompatible)
## Résolution des problèmes courants
| Problème | Cause possible | Solution |
| ------------------------------- | -------------------------------------- | ------------------------------------------------------------------------ |
| Le plugin ne se charge pas | plugin.json mal formaté | Validez avec `claude plugin validate` |
| La commande n'apparaît pas | Structure de répertoires incorrecte | Assurez-vous que `commands/` est à la racine, pas dans `.claude-plugin/` |
| Les Hooks ne se déclenchent pas | Le script n'est pas exécutable | Exécutez `chmod +x script.sh` |
| Chemin introuvable | Utilisation de chemins relatifs | Utilisez `${CLAUDE_PLUGIN_ROOT}` |
| Le serveur MCP échoue | Variables d'environnement non définies | Vérifiez la configuration des chemins dans `.mcp.json` |
## Migrer depuis une configuration existante
Si vous disposez déjà de configurations dans le répertoire `.claude/`, suivez ces étapes pour les migrer vers un Plugin :
1. **Créer la structure du Plugin** :
```bash
mkdir my-plugin/.claude-plugin
```
2. **Créer plugin.json** :
```json
{
"name": "my-plugin",
"description": "从现有配置迁移的插件",
"version": "1.0.0"
}
```
3. **Copier les fichiers existants** :
```bash
cp -r .claude/commands my-plugin/
cp -r .claude/agents my-plugin/
cp -r .claude/skills my-plugin/
```
4. **Migrer les Hooks** :
Copiez la configuration `hooks` de `.claude/settings.json` vers `hooks/hooks.json`
5. **Tester** :
```bash
claude --plugin-dir ./my-plugin
```
## Ressources d'apprentissage
### Documentation officielle
| Ressource | Lien | Description |
| --------------------- | ----------------------------------------------------------------------------------------- | -------------------------------------- |
| Plugins Reference | [code.claude.com/docs](https://code.claude.com/docs/en/plugins-reference) | Documentation de référence des plugins |
| Create Plugins | [code.claude.com/docs](https://code.claude.com/docs/en/plugins) | Guide de création de plugins |
| Best Practices | [Anthropic Engineering](https://www.anthropic.com/engineering/claude-code-best-practices) | Bonnes pratiques officielles |
| Agent Skills Standard | [agentskills.io](https://agentskills.io) | Spécification du standard ouvert |
### Dépôts officiels
| Ressource | Lien | Description |
| ---------------------------------- | ---------------------------------------------------------------------------------------- | ------------------------------ |
| anthropics/skills | [GitHub](https://github.com/anthropics/skills) | Dépôt officiel de Skills |
| anthropics/claude-plugins-official | [GitHub](https://github.com/anthropics/claude-plugins-official) | Catalogue officiel de plugins |
| Docker MCP Toolkit | [Site web](https://www.docker.com/blog/add-mcp-servers-to-claude-code-with-mcp-toolkit/) | Plus de 200 MCPs préconstruits |
### Ressources communautaires
| Ressource | Lien | Description |
| ------------------------ | ---------------------------------------------------------------------------- | -------------------------------------------- |
| claude-plugins.dev | [Site web](https://claude-plugins.dev/) | Registre communautaire et CLI |
| awesome-claude-plugins | [GitHub](https://github.com/quemsah/awesome-claude-plugins) | Collection de 243 plugins |
| awesome-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Compilation des bonnes pratiques |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agents + 15 orchestrateurs |
| Tutoriel jeremylongshore | [GitHub](https://github.com/jeremylongshore/claude-code-plugins-plus-skills) | Des centaines de plugins + tutoriels Jupyter |
### Lectures recommandées
| Article | Source |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders Tech |
| [Improving your coding workflow with Claude Code Plugins](https://composio.dev/blog/claude-code-plugin) | Composio |
| [Building My First Claude Code Plugin](https://alexop.dev/posts/building-my-first-claude-code-plugin/) | Alexander Opalic |
## Perspectives d'avenir
Le système de Plugins représente un bond qualitatif dans les capacités d'extension de Claude Code. Avec la croissance de la communauté, nous pouvons nous attendre à :
* **Un écosystème de plugins plus riche** : Couvrant un large éventail de scénarios de développement et de workflows
* **Des fonctionnalités de niveau entreprise** : Une gestion des permissions et des capacités d'audit plus complètes
* **Une compatibilité multiplateforme** : Le standard ouvert Skills a déjà été adopté par plusieurs fournisseurs
C'est le moment idéal pour vous lancer. Vous pouvez commencer par de simples commandes, ajouter progressivement des Skills et des Hooks, et finalement construire une solution complète de workflow.
Si vous souhaitez en savoir plus sur le composant Subagent que les Plugins peuvent inclure, consultez [Qu'est-ce que les Claude Code Subagents](/fr/docs/notes/claude-subagent/concept).
# Introduction au concept de sous-agent Claude Code
\##Présentation
Lorsque vous utilisez Claude Code pour gérer des tâches complexes, vous avez peut-être été confronté à un tel dilemme : le contexte de la conversation principale devient de plus en plus long, l'IA commence à « oublier » les informations importantes précédentes et la qualité de la réponse diminue progressivement.
Subagent est né pour résoudre ce problème.
Si Skills est le « manuel de travail » de Claude, alors le sous-agent est « l'employé à temps plein » que vous embauchez : ils ont leur propre poste de travail indépendant (contexte), se concentrent sur un type de travail spécifique et vous rapportent les résultats une fois terminé.
## Comprendre le sous-agent
Imaginez que vous êtes le PDG d'une entreprise. Lorsque l’entreprise est petite, vous gérez tout vous-même. Mais à mesure que votre entreprise se développe, vous commencez à embaucher des employés à temps plein : des comptables pour les finances, des RH pour le recrutement et des ingénieurs pour le développement. Chaque employé travaille à son propre poste et vous rend compte une fois la tâche terminée.
Le sous-agent joue exactement ce rôle dans Claude Code.
D'un point de vue technique, les sous-agents sont des assistants IA spécialisés qui présentent les caractéristiques suivantes :
| Caractéristiques | Descriptif |
| -------------------------- | ------------------------------------------------------------------- |
| **Contexte indépendant** | Chaque sous-agent s'exécute dans sa propre fenêtre contextuelle |
| **Capacités spécialisées** | Optimisé pour des types de tâches spécifiques |
| **Outils configurables** | Ne peut accéder qu'à un ensemble d'outils spécifié |
| **Invites personnalisées** | Il existe des invites système spéciales pour guider le comportement |
## Pourquoi avons-nous besoin d'un contexte indépendant ?
Il s’agit du concept de conception de base de Subagent et mérite une compréhension approfondie.
Dans une conversation normale, toutes les informations sont empilées dans le même contexte. Lorsque Claude recherche la base de code, analyse les fichiers, puis apporte des modifications, tous ces traitements intermédiaires occupent de l'espace contextuel. À mesure que la conversation progresse et que le contexte devient plus chargé, Claude peut commencer à « oublier » des informations importantes antérieures.
Le sous-agent modifie les éléments suivants :
```
主对话(专注于高层目标)
│
├── 用户:帮我优化这个模块的性能
│
├── Claude:我来分析一下...
│ │
│ └── [调用 Explore Subagent]
│ │ 在独立上下文中:
│ ├── 搜索相关文件
│ ├── 分析代码结构
│ ├── 识别性能瓶颈
│ └── 返回:发现 3 个优化点...
│
└── Claude:根据分析,我发现 3 个优化点...
```
Le processus d'analyse du sous-agent ne pollue pas la conversation principale. Le dialogue principal n'a reçu que des résultats raffinés, tout en conservant clarté et concentration.
## Type de sous-agent intégré
Claude Code fournit trois sous-agents intégrés puissants, couvrant les scénarios d'utilisation les plus courants :
### Explorer le sous-agent
**Ciblage** : exploration rapide et en lecture seule de la base de code.
**Caractéristiques** :
* Utiliser le modèle Haiku (rapide, faible latence)
* Strictement en lecture seule - les fichiers ne peuvent pas être créés, modifiés ou supprimés
* Outils disponibles : Glob, Grep, Read, Bash (opérations en lecture seule)
**Quand utiliser** :
Lorsque vous posez des questions exploratoires telles que « Où cette fonction est-elle implémentée ? et "Comment les erreurs sont-elles gérées ?" Claude appellera automatiquement le sous-agent Explore.
**Niveau de détail** :
| Niveau | Descriptif | Scénarios applicables |
| -------------- | ---------------------------------------------- | --------------------------------------------------------------------- |
| Rapide | Recherche rapide avec une exploration minimale | Requêtes simples et ciblées |
| Moyen | Exploration modérée | Équilibrer vitesse et exhaustivité |
| Très minutieux | Analyse complète | Des questions complexes qui nécessitent une compréhension approfondie |
### Sous-agent du plan
**Positionnement** : Étudiez la base de code et préparez un plan de mise en œuvre.
**Caractéristiques** :
* Utiliser le modèle Sonnet (capacités d'inférence plus fortes)
* Uniquement les outils d'exploration : Read, Glob, Grep, Bash
* Appelé automatiquement en mode planification
**Quand utiliser** :
Lorsque vous entrez en mode planification et que Claude doit effectuer des recherches avant de proposer un plan, le sous-agent Plan collectera automatiquement des informations, puis proposera des suggestions de plan basées sur les résultats de la recherche.
### Sous-agent à usage général
**Positionnement** : gérez des tâches complexes en plusieurs étapes.
**Caractéristiques** :
* Utiliser le modèle Sonnet
* Accès à tous les outils (y compris la lecture et l'écriture)
* Convient aux tâches complexes nécessitant une exploration et une modification
**Quand utiliser** :
Lorsque la tâche implique plusieurs étapes, nécessite une recherche avant de la modifier, ou la recherche initiale peut échouer et plusieurs stratégies doivent être essayées.
## Ma compréhension et ma pratique
Si vous observez attentivement les trois sous-agents intégrés officiels, vous découvrirez une chose en commun : **Ce sont tous des tâches de recherche et de planification**. Explore est responsable de l'exploration de la base de code, Plan est responsable de l'élaboration des plans et même General-Purpose est principalement utilisé pour la recherche et l'analyse. Aucun d’entre eux n’est spécifiquement conçu pour écrire du code.
Cela confirme ma compréhension de Subagent : **La valeur fondamentale de Subagent n'est pas un "contexte propre", mais de permettre à l'agent principal de se concentrer sur ses tâches**.
### Mode de division du travail
Ma méthode d'utilisation est très simple : le sous-agent est responsable du travail de « collecte d'informations » tel que la recherche, la planification et l'examen, et l'agent principal est responsable de l'exécution réelle.
```
Subagent(调研员) 主 Agent(执行者)
│ │
├── 探索代码库结构 │
├── 分析依赖关系 │
├── 制定实施计划 │
└── 返回精炼的上下文 ──────────→ 基于上下文执行任务
│
├── 编写代码
├── 修改文件
└── 运行测试
```
### Pourquoi ne pas laisser le sous-agent écrire du code ?
Certaines personnes aiment que l'agent principal planifie plusieurs sous-agents pour écrire du code. Je pense que ce n'est pas fiable. La raison est simple : **Le contexte manque cruellement**.
Le contexte du sous-agent est indépendant. Il ne sait pas ce qui a été discuté, quelles décisions ont été prises et quelles contraintes ont été imposées lors de la conversation principale. Lui demander d'écrire du code, c'est comme demander à un nouvel employé d'accomplir une tâche sans aucune information de base : le code produit ne correspondra probablement pas à vos attentes.
Au contraire, il est bien plus raisonnable de positionner Subagent en « chercheur » :
* La tâche de recherche elle-même ne nécessite pas beaucoup de contexte
* Les informations sont renvoyées à la place du code, qui peut être utilisé par l'agent principal en fonction du contexte complet
* Même si les résultats de l'enquête sont biaisés, l'agent principal peut les corriger
### Mon utilisation quotidienne
1. **Avant de démarrer une nouvelle tâche** : laissez l'agent Explore comprendre rapidement la structure du code concerné
2. **Planification de tâches complexes** : laissez l'agent du plan analyser les exigences et formuler les étapes de mise en œuvre
3. **Code Review** : laissez l'agent Review vérifier la qualité du code et les problèmes de sécurité.
4. **Codage réel** : l'agent principal écrit du code en fonction du contexte collecté
Voici la liste des agents que j'utilise actuellement :
L'avantage est que la fenêtre contextuelle de l'agent principal reste propre, avec uniquement « les informations dont j'ai besoin » au lieu de « un ensemble de résultats intermédiaires générés lors du processus de recherche du sous-agent ».
## Comparaison avec d'autres fonctions
### Subagent vs Skills
C'est la confusion la plus courante. Différence fondamentale : **Les compétences injectent des connaissances à Claude ; Le sous-agent crée des travailleurs indépendants**.
| Dimensions | Compétences | Sous-agent |
| ------------------------------- | ---------------------------------------------------- | -------------------------------------------------------- |
| **Fonctionnalités principales** | Fournir une expertise et des instructions | Agents qui effectuent des tâches de manière indépendante |
| **Contexte** | Partager le contexte de la conversation principale | Avoir un contexte indépendant |
| **Méthode de déclenchement** | Correspondance automatique basée sur la description | Délégation automatique ou appel manuel |
| **Scénarios applicables** | Rendre Claude meilleur dans certains types de tâches | Tâches indépendantes complexes et en plusieurs étapes |
Pour le dire au sens figuré : les compétences sont comme du matériel de formation, permettant à Claude d'apprendre à faire quelque chose ; Le sous-agent est comme un employé à temps plein, qui accomplit la tâche de manière indépendante sur son poste de travail et rend compte des résultats.
Les deux peuvent être combinés : un sous-agent de révision de code peut charger la compétence de spécification de code pour obtenir l'effet combiné de « connaissances expertes et professionnelles ».
### Commande sous-agent vs slash
| Dimensions | Sous-agent | Commande barre oblique |
| ------------------------ | ----------------------------------------- | -------------------------------- |
| **Méthode d'activation** | Délégation automatique ou appel explicite | Saisie du manuel d'utilisation |
| **Contexte** | Contexte autonome | Conversation principale partagée |
| **Complexité** | Convient aux tâches complexes | Convient aux opérations simples |
La commande slash est une touche de raccourci et vous entrez `/review` pour déclencher une opération prédéfinie ; Le sous-agent est un travailleur indépendant qui peut effectuer des tâches complexes en plusieurs étapes de manière autonome.
### Subagent vs Plugin
Le plugin est un concept de « conteneur », qui peut contenir des sous-agents :
```
Plugin(容器)
├── Commands(快捷命令)
├── Skills(知识包)
├── Agents(子代理)← 这就是 Subagent
└── Hooks(事件钩子)
```
Vous pouvez définir un sous-agent dans le répertoire `agents/` du plugin et le distribuer avec le plugin.
## Modèle de conception agent
Anthropic résume six modèles de conception agents de base dans sa documentation officielle. Comprendre ces modèles peut aider à mieux concevoir les systèmes de sous-agents :
| Modèle | Idée de base | Application de sous-agent |
| -------------------------- | -------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| **Chaînage rapide** | Décomposer des tâches complexes en plusieurs étapes séquentielles | Appel en chaîne de plusieurs sous-agents |
| **Routage** | Distribué à des processeurs spécialisés en fonction du type d'entrée | Différents types de tâches sont délégués à des sous-agents spécialisés |
| **Parallélisation** | Exécuter plusieurs sous-tâches indépendantes en même temps | Démarrer plusieurs sous-agents en parallèle |
| **Orchestre-travailleurs** | Le coordinateur central attribue les tâches aux travailleurs | Claude comme coordonnateur, Sous-agent comme ouvrier |
| **Évaluateur-Optimiseur** | Sortie du générateur, optimisation de l'évaluateur | Générer un sous-agent + examiner un sous-agent |
| **Agents** | Agents indépendants qui prennent des décisions autonomes | Chaque sous-agent fonctionne indépendamment |
Ces modes peuvent être utilisés en combinaison. Par exemple, un système de qualité de code peut utiliser à la fois :
* **Parallélisation** : exécutez simultanément des analyses de sécurité et des analyses de performances.
* **Orchestre-Ouvriers** : Maître Claude coordonne plusieurs sous-agents spécialisés
* **Evaluator-Optimizer** : vérifiez le code immédiatement après la génération
## Avantages principaux
### Protection du contexte
La plus grande valeur du sous-agent réside dans la protection du contexte de la conversation principale. Les processus intermédiaires tels que la recherche de code et l'analyse de fichiers ne seront pas accumulés dans le dialogue principal, permettant au dialogue principal de toujours se concentrer sur des objectifs de haut niveau.
### Capacités de spécialisation
Vous pouvez créer des sous-agents spécialisés pour des domaines spécifiques, configurés avec des instructions détaillées et des outils appropriés. Un sous-agent spécialisé est plus performant dans une tâche spécifique qu'un Claude polyvalent.
### Contrôle des autorisations flexible
Chaque sous-agent peut avoir des droits d'accès aux outils différents. Par exemple, la classe d'exploration Subagent n'accorde que des autorisations en lecture seule et la classe de modification Subagent n'accorde que des autorisations en écriture. Ce contrôle précis améliore la sécurité.
### Réutilisabilité
Une fois créé, le sous-agent peut être réutilisé dans plusieurs projets ou partagé avec les équipes via des plugins.
## Quand utiliser le sous-agent
**Scénarios appropriés pour l'utilisation du sous-agent** :
* Nécessite un contexte indépendant pour exécuter des tâches
* Les tâches sont des flux de travail complexes en plusieurs étapes
* Nécessite un ensemble d'outils différent de celui de la conversation principale
* Les tâches peuvent prendre beaucoup de temps à s'exécuter
**Scénarios non adaptés à l'utilisation du sous-agent** :
* Requête unique et simple
* Nécessite une interaction étroite avec le dialogue principal
* Les missions peuvent être accomplies rapidement
## Scénarios d'application typiques
### Révision du code
```
代码审查 Subagent
├── 专门的审查提示
├── 只读工具(Read, Grep, Glob)
└── 输出:结构化的审查报告
```
Lorsque vous complétez un morceau de code, vous pouvez demander au sous-agent Code Review de l'examiner dans un contexte indépendant sans interférer avec votre travail de développement principal.
### Analyse de débogage
```
调试 Subagent
├── 错误分析专家提示
├── 读写工具
└── 输出:根因分析 + 修复方案
```
Lorsqu'une erreur est rencontrée, le sous-agent de débogage peut analyser en profondeur la cause de l'erreur, essayer diverses hypothèses et enfin fournir des suggestions de réparation.
### Exploration de la base de code
```
探索 Subagent
├── Haiku 模型(快速)
├── 只读工具
└── 输出:代码结构概览
```
Lorsque vous débutez dans un nouveau projet, Explore Subagent peut rapidement cartographier votre base de code sans encombrer votre conversation principale avec des tonnes de résultats de recherche.
## Ressources d'apprentissage
### Ressources officielles
| Ressources | Liens | Instructions |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| Documentation du code Claude | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Entrée de la documentation officielle |
| Guide des sous-agents | [Documents Claude Code](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Documentation officielle du sous-agent |
| Recherche sur les systèmes multi-agents | [Ingénierie anthropique](https://www.anthropic.com/engineering/multi-agent-research-system) | Détails de la recherche sur une amélioration des performances de 90,2 % |
| Modèle de conception agent | [Documents anthropiques](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Explication détaillée de six modèles de conception de base |
### Ressources communautaires
| Ressources | Liens | Instructions |
| ------------------- | ----------------------------------------------------------------- | ---------------------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 Agents + 15 Orchestrateurs |
| Ingénierie composée | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Plugins pour 17 agents spécialisés |
| génial-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Résumé des meilleures pratiques |
## Résumé
Claude Code Subagent est essentiellement un assistant d'IA spécialisé et indépendant du contexte. Il résout le problème de la surcharge d'informations dans les tâches complexes grâce à l'isolation du contexte, gardant la conversation principale claire et ciblée à tout moment.
Retenez trois mots clés :
| Mots-clés | Signification |
| --------------- | ----------------------------------------------------------------------------- |
| **Indépendant** | Chaque sous-agent possède sa propre fenêtre contextuelle |
| **Spécialisé** | Optimisé pour des types de tâches spécifiques |
| **Délégation** | Claude peut déléguer des tâches au sous-agent automatiquement ou manuellement |
Après avoir compris le concept, le prochain article "[Claude Code Subagent Practical Guide](/fr/docs/notes/claude-subagent/practice)" vous amènera à la pratique : création d'un sous-agent personnalisé, configuration des autorisations de l'outil et bonnes pratiques dans les projets réels.
Si vous souhaitez connaître les compétences que le sous-agent peut charger, veuillez lire "[Que sont les compétences de Claude](/fr/docs/notes/claude-skills/concept)". Si vous souhaitez packager Subagent pour la distribution, veuillez lire "[Qu'est-ce que le plugin Claude Code](/fr/docs/notes/claude-plugin/concept)".
# Guide pratique du sous-agent Claude Code
## Examen rapide
Dans l'article précédent, nous avons découvert le concept de base de Subagent : il s'agit d'un assistant d'IA spécialisé indépendant du contexte qui résout le problème de la surcharge d'informations dans les tâches complexes grâce à l'isolation du contexte. Claude Code dispose de trois sous-agents intégrés : Explorer, Planifier et Général. Cet article vous emmènera d'un point de vue pratique pour créer un sous-agent personnalisé et maîtriser son utilisation avancée.
## Gérer le sous-agent
### Via la commande /agents
Le plus simple est d'utiliser l'interface interactive :
```bash
/agents
```
Cela ouvrira un menu dans lequel vous pourrez :
* Afficher tous les sous-agents (intégrés + personnalisés)
-Créer un nouveau sous-agent
* Modifier la configuration et les autorisations des outils pour le sous-agent existant
* Supprimer le sous-agent inutile
* Voir quel sous-agent est actif en cas de conflit de nom
### Grâce à la gestion des fichiers
Le sous-agent est stocké sous forme de fichier Markdown. Vous pouvez également créer et modifier des fichiers directement.
**Emplacement de stockage** :
| emplacement | chemin | portée |
| ------------------ | ------------------------------ | -------------------------------------------------- |
| Niveau du projet | `.claude/agents/` | Dédié au projet en cours et peut être soumis à Git |
| Niveau utilisateur | `~/.claude/agents/` | Disponible pour tous les projets |
| Plugin | Répertoire `agents/` du plugin | Installé avec le plugin |
**Priorité** : Niveau projet > Niveau utilisateur > Niveau plugin
Lorsqu'un sous-agent portant le même nom existe à plusieurs emplacements, celui avec la priorité la plus élevée écrasera celui avec la priorité la plus faible.
## Créez votre premier sous-agent
### Étape 1 : Créer un répertoire
```bash
mkdir -p .claude/agents
```
### Étape 2 : Créer un fichier Markdown
`.claude/agents/code-reviewer.md`:
```markdown
---
name: code-reviewer
description: 专业的代码审查代理,用于代码质量检查。在完成代码编写后主动使用。
tools: Read, Grep, Glob, Bash
---
你是一位资深的代码审查专家,专注于确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 分析修改的文件
3. 提供结构化的审查反馈
审查清单:
- 代码可读性和命名规范
- 错误处理和边界条件
- 安全漏洞(注入、敏感信息泄露)
- 性能优化机会
- 测试覆盖情况
输出格式:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
```
### Étape 3 : tester le sous-agent
Dans Claude Code :
```
> 用 code-reviewer 代理审查我最近的修改
```
Ou laissez Claude choisir automatiquement :
```
> 帮我审查一下代码质量
```
Si `description` est écrit suffisamment clairement, Claude reconnaîtra et appellera automatiquement votre sous-agent.
## Explication détaillée des champs de configuration
Le fichier de configuration du sous-agent se compose de deux parties : le contenu YAML et le corps Markdown.
### YAML Frontmatter
```yaml
---
name: your-agent-name
description: 描述这个代理做什么,以及何时应该被使用
tools: Tool1, Tool2, Tool3
model: sonnet
permissionMode: default
skills: skill1, skill2
---
```
| Champ | Obligatoire | Descriptif |
| ---------------- | ----------- | ------------------------------------------------------------------------------------------------------ |
| `name` | est | un identifiant unique, utilisez des lettres minuscules et des traits d'union |
| `description` | Oui | Description en langage naturel (Claude l'utilise pour déterminer quand appeler) |
| `tools` | Non | Liste d'outils séparés par des virgules. En cas d'omission, tous les outils sont hérités |
| `model` | Non | Sélection de modèle : `sonnet`, `opus`, `haiku` ou `inherit` |
| `permissionMode` | Non | Mode d'autorisation (voir ci-dessous) |
| `skills` | Non | Compétences chargées automatiquement (le sous-agent n'hérite pas des compétences de la session parent) |
### Mode d'autorisation
| Mode | Descriptif |
| ------------------- | ------------------------------------------------- |
| `default` | Vérification normale des autorisations |
| `acceptEdits` | Accepter automatiquement les opérations d'édition |
| `bypassPermissions` | Ignorer toutes les vérifications d'autorisation |
| `plan` | Proposez seulement un plan, pas l'exécutez |
| `ignore` | Ignorer ce sous-agent |
### Texte de démarque
Le texte est l'invite système du sous-agent. Plus vous écrivez de manière détaillée, meilleures sont les performances du sous-agent.
Une bonne invite système doit inclure :
* Définition claire des rôles
* Étapes de travail spécifiques
* Liste de contrôle clé
* Exigences relatives au format de sortie
## Mécanisme de déclenchement
### Délégation automatique
Claude décidera automatiquement s'il doit déléguer en fonction du contenu de la tâche et du `description` du sous-agent.
**Conseil pour encourager l'utilisation automatique** : utilisez des mots déclencheurs dans `description` :
```yaml
description: Use PROACTIVELY after writing code for quality checks
```
ou :
```yaml
description: MUST BE USED when encountering errors or test failures
```
### Appel explicite
Dites directement à Claude quel sous-agent utiliser :
```
> 使用 code-reviewer 代理检查我的代码
> 让 debugger 代理分析这个错误
> 调用 test-runner 代理运行测试
```
\##Configuration des outils
### Liste des outils couramment utilisés
| Outils | Instructions |
| ----------- | ----------------------------------- |
| `Read` | Lire le contenu du fichier |
| `Write` | Écrire dans un fichier |
| `Edit` | Modifier le fichier |
| `Glob` | Correspondance de modèle de fichier |
| `Grep` | Recherche d'expressions régulières |
| `Bash` | Exécuter la commande shell |
| `WebFetch` | Obtenir du contenu Web |
| `WebSearch` | Rechercher sur le Web |
### Stratégie de configuration des outils
**Lecture seule Sous-agent** (exploration, analyse) :
```yaml
tools: Read, Grep, Glob, Bash
```
REMARQUE : même si Bash est inclus, Subagent ne doit être utilisé que pour les commandes en lecture seule (ls, git status, git log, etc.).
**Lecture et écriture du sous-agent** (réparation, refactorisation) :
```yaml
tools: Read, Edit, Write, Bash, Grep, Glob
```
**Principe du moindre privilège** : accordez uniquement les outils nécessaires pour éviter les opérations accidentelles.
## Modèle de sous-agent pratique
### Réviseur de code
```markdown
---
name: code-reviewer
description: Expert code review. Use PROACTIVELY after writing or modifying code.
tools: Read, Grep, Glob, Bash
model: inherit
---
你是一位资深的代码审查专家,确保代码质量和安全性。
当被调用时:
1. 运行 `git diff` 查看最近的更改
2. 聚焦于修改的文件
3. 立即开始审查
审查清单:
- 代码清晰可读
- 函数和变量命名规范
- 无重复代码
- 正确的错误处理
- 无暴露的密钥或 API 密码
- 输入验证完整
- 测试覆盖充分
- 性能考虑到位
按优先级组织反馈:
- 严重问题(必须修复)
- 警告(建议修复)
- 建议(可以改进)
包含具体的修复示例。
```
### Expert en débogage
```markdown
---
name: debugger
description: Debugging specialist. Use PROACTIVELY when encountering errors or test failures.
tools: Read, Edit, Bash, Grep, Glob
---
你是一位调试专家,专注于根因分析。
当被调用时:
1. 捕获错误信息和堆栈跟踪
2. 识别复现步骤
3. 定位失败位置
4. 实施最小修复
5. 验证解决方案有效
调试流程:
- 分析错误信息和日志
- 检查最近的代码更改
- 形成并测试假设
- 添加战略性的调试日志
- 检查变量状态
对每个问题提供:
- 根因解释
- 支持诊断的证据
- 具体的代码修复
- 测试方法
- 预防建议
专注于修复根本问题,而非表面症状。
```
### Testeur
```markdown
---
name: test-runner
description: Test automation expert. Use PROACTIVELY to run tests and fix failures.
tools: Read, Edit, Bash, Grep, Glob
permissionMode: acceptEdits
---
你是一位测试自动化专家。
当你看到代码更改时:
1. 识别相关的测试文件
2. 运行适当的测试
3. 如果测试失败,分析原因并修复
测试策略:
- 优先运行与更改相关的测试
- 分析失败的测试输出
- 区分代码问题和测试问题
- 修复后重新运行验证
对于新功能:
- 确认测试覆盖关键路径
- 检查边界条件测试
- 验证错误处理测试
```
### Générateur de documents
```markdown
---
name: doc-generator
description: Documentation specialist. Use when creating or updating documentation.
tools: Read, Write, Grep, Glob
---
你是一位技术文档专家。
当被调用时:
1. 分析代码结构和注释
2. 识别公共 API 和关键功能
3. 生成清晰的文档
文档风格:
- 简洁明了
- 包含代码示例
- 解释为什么,而不只是是什么
- 考虑读者背景
输出格式:
- API 参考用 Markdown
- 使用恰当的标题层级
- 包含目录(如果文档较长)
```
### Scanner de sécurité
```markdown
---
name: security-scanner
description: Security specialist. Use PROACTIVELY when reviewing code for security issues.
tools: Read, Grep, Glob, Bash
---
你是一位安全专家,专注于发现代码中的安全漏洞。
扫描范围:
- 注入漏洞(SQL、命令、XSS)
- 认证和授权问题
- 敏感数据暴露
- 安全配置错误
- 依赖漏洞
检查清单:
- 用户输入是否经过验证和转义
- 敏感数据是否加密存储
- API 密钥是否硬编码
- 是否使用安全的默认配置
- 依赖是否有已知漏洞
输出格式:
- 严重(立即修复)
- 高危(尽快修复)
- 中危(计划修复)
- 低危(考虑修复)
每个问题包含:
- 漏洞描述
- 风险说明
- 修复建议
- 参考资料
```
## Utilisation avancée
### Modèles de conception au niveau de la production
Dans les environnements de production, il existe plusieurs modèles éprouvés de collaboration multi-agents :
#### 3 Mode Amigos
Un modèle de collaboration composé de trois rôles : produit, architecture et mise en œuvre :
```
PM Agent → Architect Agent → Claude Code
(产品定义) (技术设计) (代码实现)
```
| Rôles | Responsabilités | Configuration des outils |
| ------------------ | ------------------------------------------- | -------------------------- |
| Agent PM | Définition des fonctions, tri des exigences | Lire, recherche sur le Web |
| Agent d'architecte | Conception de solutions techniques | Lire, Glob, Grep |
| Claude Code | Mise en œuvre du code | Tous les outils |
#### Pipeline en trois étapes
Décomposez les tâches complexes en trois étapes claires :
```
规格制定 → 架构评审 → 实现测试
(Spec) (Review) (Implement)
```
Chaque étape est responsable d'un sous-agent dédié et la sortie sert d'entrée à l'étape suivante.
#### Stratégie d'orchestration de modèles
Différents modèles sont utilisés à différentes étapes pour optimiser les coûts et les effets :
| Scène | Modèle recommandé | Raison |
| ---------------------- | ----------------- | ------------------------------ |
| Phase de planification | Sonnet | Raisonnement approfondi requis |
| Phase d'exécution | Haïku | Rapide et peu coûteux |
| Étape de révision | Sonnet | Un jugement global requis |
Exemple de configuration :
```yaml
---
name: quick-executor
model: haiku
---
```
### Lien sous-agent
Pour les flux de travail complexes, plusieurs sous-agents peuvent être liés :
```
> 首先用 code-analyzer 代理找出性能问题,
> 然后用 optimizer 代理修复它们
```
### Exécution avec reprise
L'exécution du sous-agent peut être suspendue et reprise, en conservant le contexte précédent complet :
**Appel initial** :
```
> 用 code-analyzer 代理开始分析认证模块
[Agent 完成初始分析并返回 agentId: "abc123"]
```
**Agent de récupération** :
```
> 恢复代理 abc123,继续分析授权逻辑
[Agent 继续,保持之前的完整上下文]
```
**Scénario d'utilisation** :
* Études de longue durée, réalisées en plusieurs sessions
* Améliorations itératives, en gardant le contexte
* Flux de travail en plusieurs étapes pour traiter les tâches associées en séquence
### Configurer les compétences du sous-agent
Le sous-agent n'hérite pas automatiquement des compétences de la session parent. Si nécessaire, déclarez-le explicitement :
```yaml
---
name: code-reviewer
skills: code-standards, security-checklist
---
```
### Définition dynamique CLI
Pas besoin de sauvegarder le fichier, définissez le sous-agent temporaire directement sur la ligne de commande :
```bash
claude --agents '{
"quick-reviewer": {
"description": "Quick code review",
"prompt": "You are a code reviewer...",
"tools": ["Read", "Grep", "Glob"],
"model": "haiku"
}
}'
```
Convient pour des tests rapides ou une utilisation unique.
\## meilleures pratiques
### 1. Restez concentré
```markdown
✅ 好:单一责任
---
name: code-reviewer
description: Expert code review for quality and security
---
❌ 差:试图做太多
---
name: super-agent
description: Does everything - reviews, tests, deploys, documents...
---
```
Un sous-agent qui fait bien une chose vaut mieux qu’un sous-agent qui fait plusieurs choses.
### 2. Rédigez une description claire
Claude utilise `description` pour décider quand utiliser le sous-agent. Une bonne description doit répondre :
1. \*\*Que fait ce sous-agent ? \*\* Répertoriez les capacités spécifiques
2. \*\*Quand doit-il être utilisé ? \*\* Contient des mots déclencheurs
```markdown
✅ 好:具体和清晰
description: 从 PDF 文件提取文本和表格,填写表单,合并文档。
当处理 PDF 文件或用户提到 PDF、表单、文档提取时使用。
❌ 差:太模糊
description: 处理文档
```
### 3. Restreindre l'accès aux outils
Accordez uniquement les outils dont vous avez besoin :
```yaml
---
name: code-reviewer
tools: Read, Grep, Glob, Bash
---
```
Cela empêche Subagent de modifier accidentellement des fichiers et lui permet de se concentrer davantage sur son travail de révision.
### 4. Écrivez des invites système détaillées
Plus les invites du système sont détaillées, meilleures sont les performances du sous-agent :
* Définition claire des rôles
* Étapes de travail spécifiques
* Liste de contrôle clé
* Exigences relatives au format de sortie
### 5. Contrôle des versions
Validez le sous-agent au niveau du projet dans Git :
```bash
git add .claude/agents/
git commit -m "Add code-reviewer subagent"
```
Les membres de l'équipe obtiennent automatiquement le même sous-agent après le clonage du projet.
## Dépannage des problèmes courants
| Problème | Cause possible | Solutions |
| ---------------------------------- | --------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Le sous-agent n'est pas appelé | la description n'est pas assez claire | Ajoutez des mots déclencheurs pour le rendre plus spécifique |
| Le sous-agent n'est pas appelé | Emplacement du fichier incorrect | Assurez-vous que le fichier est dans `.claude/agents/` ou `~/.claude/agents/` |
| Les outils ne sont pas disponibles | Erreur de configuration du champ Outils | Vérifiez l'orthographe des noms d'outils et assurez-vous qu'ils sont séparés par des virgules |
| La sortie est instable | L'invite du système est trop vague | Ajouter des étapes spécifiques et des exigences de format de sortie |
| Contexte perdu | Séance terminée | Utilisation de l'exécution avec reprise |
| Conflit de nom | Sous-agent portant le même nom à plusieurs endroits | Utilisez `/agents` pour voir lequel est actif |
## Partager avec l'équipe
### Méthode 1 : via Git
Placez Subagent dans le répertoire `.claude/agents/` et soumettez-le au référentiel du projet. Les membres de l'équipe sont automatiquement obtenus après le clonage.
### Méthode 2 : via un plugin
Placez le sous-agent dans le répertoire `agents/` du plugin et distribuez-le via le mécanisme du plugin.
### Méthode 3 : partage au niveau de l'utilisateur
Placez les sous-agents couramment utilisés dans `~/.claude/agents/` pour les rendre disponibles dans tous les projets. La synchronisation entre plusieurs machines peut être gérée à l'aide de dotfiles.
## Ressources d'apprentissage
### Documentation officielle
| Ressources | Liens | Instructions |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------- |
| Documentation du code Claude | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) | Entrée de la documentation officielle |
| Guide des sous-agents | [Documents Claude Code](https://docs.anthropic.com/en/docs/claude-code/sub-agents) | Détails de configuration du sous-agent |
| Modèles de conception agent | [Documents anthropiques](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/agentic-design-patterns) | Six modèles de conception de base |
| Etude multi-agents | [Ingénierie anthropique](https://www.anthropic.com/engineering/multi-agent-research-system) | Détails de l'étude sur l'amélioration des performances de 90,2 % |
### Ressources communautaires
| Ressources | Liens | Instructions |
| ------------------- | ----------------------------------------------------------------- | --------------------------------------- |
| wshobson/agents | [GitHub](https://github.com/wshobson/agents) | 99 agents + 15 modèles d'orchestrateur |
| Ingénierie composée | [GitHub](https://github.com/EveryInc/compound-engineering-plugin) | Plugins pour 17 agents spécialisés |
| génial-claude-code | [GitHub](https://github.com/hesreallyhim/awesome-claude-code) | Résumé des bonnes pratiques Claude Code |
### Lecture recommandée
| Article | Source | Sujet |
| -------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | --------------------------------------- |
| Construire des agents efficaces | [Anthropique](https://www.anthropic.com/research/building-effective-agents) | Principes de conception des agents |
| Comment nous avons construit notre système de recherche multi-agents | [Anthropique](https://www.anthropic.com/engineering/multi-agent-research-system) | Pratique de l'architecture multi-agents |
| Claude Compétences, commandes, sous-agents et plugins | [Technologie des jeunes leaders](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Analyse de comparaison de fonctions |
## Résumé
Claude Code Subagent est un outil puissant pour améliorer l'efficacité de la programmation de l'IA. Il rend les tâches complexes gérables via des contextes indépendants et des configurations spécialisées.
Démarrage rapide :
1. Exécutez `/agents` pour ouvrir l'interface de gestion
2. Créez un sous-agent simple (tel qu'un réviseur de code)
3. Testez la délégation automatique et les appels explicites
4. Ajustez la configuration selon vos besoins
Au fur et à mesure que vous approfondissez votre utilisation, vous pouvez progressivement :
* Créez un sous-agent exclusif pour votre équipe
* Configurer les liens de sous-agents pour gérer des flux de travail complexes
* Gérer des tâches à long terme avec une exécution pouvant être reprise
Si vous souhaitez packager et distribuer Subagent avec d'autres configurations, veuillez lire "[Claude Code Plugin Practical Guide](/fr/docs/notes/claude-plugin/practice)".
# Conseils avancés
## Notifications terminal : alertes à la fin des tâches
Vous souhaitez recevoir une notification lorsque Claude termine une tâche ?
```bash
claude config set --global preferredNotifChannel terminal_bell
```
Combinez cela avec les notifications d'iTerm2, ou utilisez `terminal-notifier` pour des notifications personnalisées (consultez la configuration des Hooks dans les [Meilleures pratiques](/fr/blog/claude-code-best-practices)).
## Utilisation avancée des Hooks
Les Hooks ne se limitent pas à l'exécution de commandes shell. Il existe en réalité quatre types :
1. **command**:Shell 命令(最常见)
2. **http**:POST JSON 到 URL(支持自定义 headers 和环境变量展开)
3. **prompt**:发给 Claude 评估(比如「所有任务都完成了吗?」)
4. **agent**:启动一个有工具访问权限的子代理来验证
Quelques événements Hook avancés à connaître :
* `PostCompact`:压缩完成后触发,适合注入提醒让 Claude 重新读取关键文件
* `SessionStart`:写入 `$CLAUDE_ENV_FILE` 可以给整个会话持久化环境变量
* `PreToolUse`:可以修改工具输入(`updatedInput`),甚至自动批准或拒绝操作
## Écosystème de plugins
Utilisez `/plugin` pour parcourir et installer des plugins communautaires. Quelques plugins notables :
* **dx**(by ykdojo):提供 `/handoff`(自动写交接文档)、`/clone`(克隆对话)、`/half-clone`(只克隆最近的对话减少上下文)
* **mine**(by anipotts):把所有 Claude Code 会话数据导入 SQLite,支持成本追踪、缓存分析、错误记忆等查询
## Agent Teams : collaboration multi-agent
Définissez une variable d'environnement pour activer la fonctionnalité expérimentale Agent Teams :
```bash
export CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
Une fois activée, une session peut agir en tant que Team Lead, coordonnant plusieurs agents travaillant simultanément via git worktree. Chaque agent s'exécute de manière indépendante dans sa propre fenêtre de contexte, ce qui est idéal pour le développement parallèle de grands projets.
Notez toutefois que la consommation de tokens augmente de 4 à 15 fois, à utiliser avec discernement.
## Stratégies de prompts
Les conseils suivants proviennent du fil Twitter de Boris Cherny sur les pratiques d'équipe — essentiellement les meilleures pratiques d'« ingénierie de prompts » appliquées à Claude Code.
### Utilisez Claude comme votre réviseur de code
Ne vous contentez pas de demander à Claude d'écrire du code — demandez-lui aussi de réviser le vôtre :
```
Grill me on these changes and don't make a PR until I pass your test.
```
Ou demandez-lui de prouver que le code fonctionne :
```
Prove to me this works. Diff behavior between main and my feature branch.
```
### Ne reformulez pas quand vous n'êtes pas satisfait
Le conseil n°6 de Boris : si Claude donne une réponse médiocre, ne reformulez pas votre question. Dites plutôt « Cette solution n'est pas assez bonne, dis-moi précisément ce qui peut être amélioré ». Itérer sur la réponse existante fonctionne mieux que de repartir de zéro.
### Laissez Claude mettre à jour son propre CLAUDE.md
Après avoir corrigé une erreur, ajoutez :
```
Update your CLAUDE.md so you don't make that mistake again.
```
Boris affirme que Claude est étonnamment doué pour écrire ses propres règles. Avec le temps, CLAUDE.md devient de plus en plus précis et la qualité des conversations s'améliore continuellement.
### Dites simplement « fix »
Avec Slack MCP activé, collez un rapport de bug depuis Slack et dites un seul mot : **fix**. Zéro changement de contexte.
Ou quand le CI échoue, dites simplement :
```
Go fix the failing CI tests.
```
Pas besoin d'analyser manuellement les logs ni d'expliquer le problème — laissez Claude consulter les logs, diagnostiquer le problème et le corriger.
## En conclusion
Claude Code évolue très rapidement, et ces conseils sont en constante évolution. Nous vous recommandons de suivre le Changelog officiel pour rester à jour.
Si vous n'avez pas encore lu mes articles précédents, je vous suggère de commencer par les flux de travail fondamentaux :
### Lectures complémentaires
* [Mes meilleures pratiques avec Claude Code](/fr/blog/claude-code-best-practices) — Conseils essentiels sur les flux de travail et guide des commandes slash
* [Contrôle qualité en programmation IA : 5 lignes de défense](/fr/blog/claude-code-quality-control) — Système d'assurance qualité pour la programmation avec Claude Code
* [Architecture du système Claude expliquée](/fr/docs/notes/claude-architecture) — Comprendre MCP, Skills, Subagents, Hooks et plus encore
# Commandes Pratiques et Automatisation
## `/diff` : Visualiseur Interactif de Diff
Tapez `/diff` pour ouvrir une vue interactive de diff :
* **Flèches gauche/droite** : Basculez entre le git diff complet (toutes les modifications) et les modifications par tour de Claude
* **Flèches haut/bas** : Parcourez les différents fichiers
Bien plus agréable que d'exécuter `git diff` dans le terminal, surtout lorsque les modifications concernent plusieurs fichiers.
## `/simplify` : Revue de Code Multi-Agents
L'exécution de `/simplify` lance 3 agents de revue en parallèle :
* Agent de **réutilisation du code** : Recherche les motifs dupliqués
* Agent de **qualité du code** : Vérifie la lisibilité et la structure
* Agent d'**efficacité** : Analyse les surcharges de performance inutiles
Les trois agents travaillent indépendamment, puis agrègent les résultats — corrigeant automatiquement les problèmes valides et ignorant les faux positifs.
## `/security-review` : Scan de Sécurité
Effectue un audit de sécurité sur les modifications de la branche courante, vérifiant les injections SQL, XSS, les failles d'authentification, les problèmes de traitement des données et les vulnérabilités des dépendances. Chaque découverte passe par une validation adversariale pour réduire les faux positifs.
## Fonctionnalités Cachées de `/copy`
`/copy` ne se contente pas de copier la dernière réponse. Lorsque la réponse contient des blocs de code, un sélecteur interactif apparaît vous permettant de choisir un bloc de code spécifique au lieu de copier la réponse entière. Vous pouvez également passer un numéro pour copier des réponses antérieures : `/copy 2` copie l'avant-dernière, `/copy 3` copie l'antépénultième — sans avoir à faire défiler et sélectionner manuellement.
## `/batch` : Refactoring Parallèle à Grande Échelle
```
/batch 把 src/ 下所有组件从 Class 组件迁移到函数组件
```
C'est une fonctionnalité majeure. `/batch` analyse la base de code, décompose la tâche en 5 à 30 unités indépendantes, lance un agent dédié pour chaque unité travaillant dans un git worktree isolé, puis chaque agent effectue un commit et ouvre un PR.
Idéal pour les migrations à grande échelle, l'ajout en masse d'annotations de types, les renommages globaux et autres scénarios similaires.
## `/loop` : Tâches Planifiées
```
/loop 5m 检查部署是否完成
/loop 1h /review-pr 1234
```
Crée une tâche planifiée au sein de la session qui se répète à l'intervalle spécifié. Utile pour surveiller l'état d'un déploiement, vérifier périodiquement des PRs, etc. C'est au niveau de la session (disparaît à la fermeture), limité à 50 tâches, avec une expiration automatique après 3 jours.
## Entrée par Pipe : Transmettez N'importe Quoi à Claude
```bash
# 让 Claude 分析错误日志
cat error.log | claude -p "分析这个错误日志,找出根本原因"
# 让 Claude 总结最近的改动
git diff HEAD~3 | claude -p "总结这三次提交的改动"
# 让 Claude 解读命令输出
kubectl get pods | claude -p "哪些 pod 状态异常?"
```
`-p` est le mode headless (non interactif), idéal pour une utilisation dans les scripts et les pipelines CI/CD.
## Paramètres Cachés du Mode Headless
Le mode `-p` dispose de paramètres extrêmement puissants mais peu connus :
```bash
# 设置花费上限(超过就停)
claude -p --max-budget-usd 5.00 "重构认证模块"
# 限制对话轮数
claude -p --max-turns 3 "修复这个测试"
# 输出 JSON 格式(方便程序解析)
claude -p --output-format json "分析这个项目"
# 要求输出符合特定 JSON Schema
claude -p --json-schema '{"type":"object","properties":{"summary":{"type":"string"}}}' "总结项目"
# 多轮 headless 对话(用 session-id 保持上下文)
claude -p --session-id my-task "第一步:分析代码"
claude -p --session-id my-task "第二步:生成测试"
# 指定备用模型(主模型过载时自动切换)
claude -p --fallback-model sonnet "复杂分析"
# 限制可用工具
claude -p --tools "Read,Grep,Glob" "只读分析,不要改代码"
# 完全替换系统提示词
claude -p --system-prompt "你是一个 Python 专家" "优化这段代码"
```
# Configuration et diagnostics
## `/statusline` : Barre d'état personnalisée
Utilisez `/statusline` pour personnaliser les informations affichées dans la barre d'état inférieure en utilisant des descriptions en langage naturel. Vous pouvez également créer manuellement un script `~/.claude/statusline.sh`.
Les informations affichables incluent : modèle actuel, branche git, nombre de fichiers non commités, progression d'utilisation du contexte, coût de la session, etc. Lorsque vous avez plusieurs fenêtres Claude ouvertes pour différentes tâches, la barre d'état vous aide à identifier rapidement ce que fait chaque fenêtre.
## Autocomplétion de settings.json
Ajoutez `$schema` au début de votre settings.json, et VS Code / Cursor fournira l'autocomplétion et la validation des options de configuration :
```json
{
"$schema": "https://json.schemastore.org/claude-code-settings.json"
}
```
## Quelques paramètres cachés utiles
```json
{
"showTurnDuration": true,
"env": {
"DISABLE_AUTOUPDATER": "1"
}
}
```
* `showTurnDuration` : Affiche la durée de chaque tour de conversation
* `DISABLE_AUTOUPDATER` : Désactive la vérification automatique des mises à jour, réduisant la surcharge de contexte
## `/stats` et `/insights` : Analyse d'utilisation
* `/stats` : Visualise l'utilisation quotidienne, l'historique des sessions, les séries d'utilisation et les préférences de modèle, avec filtrage par plage de dates
* `/insights` : Analyse tout votre historique Claude Code, vous indique quels flux de travail sont efficaces, où se trouvent les goulots d'étranglement, et génère des suggestions d'optimisation
## history.jsonl : Historique des prompts
Claude enregistre chaque prompt que vous envoyez dans `~/.claude/history.jsonl`. Vous pouvez demander à Claude d'analyser ce fichier pour identifier les modèles de prompts et les opportunités d'optimisation.
## `/doctor` : Bilan de santé
Lorsque vous rencontrez des problèmes étranges, exécutez `/doctor` (ou `claude doctor` dans le terminal). Il vérifie l'état de l'installation, la version, l'état de l'authentification et les dépendances système pour vous aider à identifier rapidement les problèmes.
## Outils communautaires
L'outil communautaire `ccusage` permet de suivre l'utilisation des tokens :
```bash
npx ccusage daily
```
Si vous avez utilisé `--dangerously-skip-permissions` ou approuvé de nombreuses commandes, vous pouvez utiliser `cc-safe` pour rechercher des risques :
```bash
npx cc-safe .
```
Il vérifie `.claude/settings.json` à la recherche de commandes à haut risque comme `sudo`, `rm -rf`, `chmod 777`, `git reset --hard`, etc.
## Plus de commandes slash à connaître
| Commande | Fonction |
| --------------------- | ---------------------------------------------------------------------------------------------------------- |
| `/export [filename]` | Exporter la conversation en texte brut |
| `/pr-comments [PR]` | Récupérer les commentaires de PR (détecte automatiquement la branche actuelle) |
| `/release-notes` | Voir le journal des modifications de la version actuelle |
| `/plugin` | Parcourir et installer les plugins communautaires |
| `/fast` | Basculer le mode rapide |
| `claude --debug` | Activer les logs de débogage au démarrage (supporte le filtrage par catégorie, ex., `--debug "api,hooks"`) |
| `/install-github-app` | Installer GitHub App pour les revues automatiques de PR |
# Gestion du contexte
## `/compact` accepte des arguments
Beaucoup savent que `/compact` peut compresser le contexte, mais peu savent qu'il accepte des arguments pour spécifier ce qu'il faut conserver :
```
/compact 保留所有关于数据库 schema 的讨论,以及当前的重构方案
```
Ainsi, la compression donnera la priorité au contenu que vous avez spécifié, évitant la perte de contexte critique.
## Écrivez des instructions de survie à la compaction dans CLAUDE.md
Ajoutez une section `## Compact Instructions` dans votre CLAUDE.md pour indiquer à Claude ce qui doit être préservé lors de la compaction :
```markdown
## Compact Instructions
When summarizing, preserve all TypeScript type changes, error patterns encountered, and the current refactoring plan.
```
Ainsi, même la compaction automatique ne perdra pas d'informations critiques.
## Empêchez Claude d'abandonner prématurément à cause du budget de tokens
Ajoutez ceci dans votre CLAUDE.md :
```markdown
Your context window will be automatically compacted as it approaches its limit.
Never stop tasks early due to token budget concerns.
Always complete tasks fully, even if the end of your budget is approaching.
```
Parfois, Claude s'arrête de manière proactive lorsque le contexte est presque plein, en disant « le contexte est presque plein ». Ajouter ceci l'empêche d'abandonner prématurément.
## Protocole Handoff : passation de session
Lorsque le contexte est presque plein mais que la tâche n'est pas terminée, demandez à Claude d'écrire un document de passation :
```
把剩余的计划写到 HANDOFF.md 里,说明你尝试了什么、什么有效、什么没效。
```
Ensuite, ouvrez une nouvelle session et utilisez simplement `@HANDOFF.md` pour restaurer le contexte complet. Cela compresse plus de 10K tokens de contexte en moins de 2K, bien plus précis que `/compact`.
## Compactez proactivement à 70-80%
Un point facile à négliger : lorsque le contexte approche de sa limite, Claude déclenche automatiquement la compaction. Mais quand la compaction automatique survient en pleine tâche, elle peut perdre des informations critiques et dégrader la qualité des réponses suivantes.
Une meilleure approche est la **gestion proactive** : exécutez manuellement `/compact` lorsque le contexte atteint 70-80% — c'est bien plus efficace que d'attendre la compaction automatique. Exécutez `/clear` immédiatement après avoir terminé une tâche ; ne laissez pas le contexte gonfler indéfiniment.
Vous pouvez également déclencher la compaction automatique plus tôt via une variable d'environnement :
```json
{
"env": {
"CLAUDE_AUTOCOMPACT_PCT_OVERRIDE": "50"
}
}
```
## `/context` : diagnostic du contexte
Vous ne savez pas combien d'espace reste dans la fenêtre de contexte ? `/context` vous le dira :
* Quels outils ou services MCP consomment le plus de contexte
* Le pourcentage d'utilisation actuel de la capacité
* Des suggestions d'optimisation ciblées
J'ai constaté que parfois, le simple fait d'avoir certains services MCP enregistrés (même sans les utiliser) peut consommer plus de 30% de la fenêtre de contexte. Utilisez `/context` pour vérifier ; nettoyer les MCP inutilisés peut libérer un espace considérable.
## Chargement différé automatique des outils MCP
Lorsque les définitions d'outils MCP dépassent 10% du contexte, Claude Code active automatiquement Tool Search — chargeant un index de recherche léger au lieu des définitions complètes d'outils. Cela réduit la consommation de contexte MCP de plus de 85% (par exemple, de 77K tokens à 8.7K). Cette fonctionnalité est **activée par défaut** et ne nécessite aucune configuration manuelle.
À noter : Tool Search ne prend en charge que les modèles Sonnet 4+ et Opus 4+, pas Haiku. Si votre `ANTHROPIC_BASE_URL` pointe vers un proxy non officiel, Tool Search sera automatiquement désactivé (car la plupart des proxies ne transmettent pas les blocs `tool_reference`).
Pour personnaliser le comportement, configurez-le dans settings.json :
```json
{
"env": {
"ENABLE_TOOL_SEARCH": "auto:5"
}
}
```
Valeurs de configuration prises en charge :
* **Non défini** : Activé par défaut
* **`true`** : Activation forcée (y compris les scénarios de proxy non officiel)
* **`auto`** : S'active lorsque le contexte dépasse 10% (équivalent au comportement par défaut)
* **`auto:`** : Seuil personnalisé, par exemple `auto:5` signifie activation au-delà de 5%
* **`false`** : Désactivé, tous les outils MCP sont préchargés
# Astuces cachées de Claude Code
Astuces pratiques pour Claude Code rassemblées à partir des tweets du fondateur Boris Cherny, de la communauté et du changelog. Nombre d'entre elles sont de véritables pépites cachées -- raccourcis, fonctionnalités méconnues, astuces en ligne de commande et bien plus -- qu'il est difficile d'abandonner une fois que vous les avez adoptées.
## Sommaire
* [Raccourcis clavier](./shortcuts) -- Changement de mode avec Shift+Tab, rappel d'historique avec Esc+Esc, Ctrl+S pour la sauvegarde temporaire, et plus
* [Saisie et interaction](./input-interaction) -- Commandes terminal avec `!`, injection de fichiers avec `@`, collage d'URL, /btw pour les interruptions, mode Vim
* [Réflexion et contrôle du modèle](./thinking-model) -- Mots-clés think/ultrathink, /effort, subagents, opusplan
* [Gestion des sessions](./session-management) -- /rename, /branch, /color, contrôle à distance
* [Gestion du contexte](./context-management) -- Paramètres de /compact, directives de compression, protocole Handoff, chargement différé MCP
* [Commandes et automatisation](./commands-automation) -- /diff, /simplify, /batch, /loop, mode Headless
* [Configuration et diagnostics](./config-diagnostics) -- statusline, settings.json, /stats, /doctor
* [Avancé](./advanced) -- Hooks en profondeur, écosystème de plugins, Agent Teams, philosophie des prompts
### Pour aller plus loin
* [Architecture du système Claude en détail](/fr/docs/notes/claude-architecture) -- Comprendre MCP, Skills, Subagents, Hooks et les autres composants
* [Guide complet des Claude Subagents](/fr/docs/notes/claude-subagent) -- Concepts et mise en pratique des sous-agents
# Saisie et interaction
## `!`:直接运行终端命令
在输入框以 `!` 开头,可以直接在 Claude Code 内执行终端命令,不需要切换到另一个终端窗口:
```
! git status
! npm run build
! docker ps
```
输入 `!` 加命令前缀后按 Tab 还能自动补全历史命令。
## `@` + 文件路径:注入文件上下文
在输入时用 `@` 加文件路径,可以把文件内容直接注入到上下文中:
```
帮我看看 @src/auth/login.ts 和 @src/auth/middleware.ts 之间的逻辑有没有问题
```
支持 Tab 键自动补全路径,不需要手动输入完整路径。比让 Claude 自己去读文件更快,因为省去了工具调用的开销。
## 直接粘贴 URL
直接把 URL 粘贴到输入中,Claude 会自动抓取网页内容作为上下文:
```
参考这个 API 文档 https://docs.example.com/api/v2 来写客户端代码
```
## 喂 `/llms-full.txt` 让 Claude 自己查文档
很多开源项目的文档站点会提供 `/llms-full.txt` 文件(LLM 友好的完整文档)。遇到某个库的问题时,把这个文件的 URL 粘贴给 Claude,它能自己查文档解决绝大部分问题:
```
参考 https://docs.astro.build/llms-full.txt 帮我解决这个路由问题
```
## `/btw`:在 Claude 工作时插嘴
这是 2026 年 3 月刚加的新功能。当 Claude 正在执行任务时,你可以用 `/btw` 发起一个旁路对话——问问它在想什么、给它补充信息,而不需要打断当前任务。
正如 Anthropic 工程师 @trq212 在推特上说的:「没人会用 Ctrl+C 打断同事,你只需要说一句 'btw',他们就会抬头看你。」
## Vim 模式
输入 `/vim` 开启 Vim 模式,支持:
* 模式切换(Normal/Insert)
* 导航(h/j/k/l, w/b/e, 0/$)
* 编辑操作(d, c, y, p)
* 文本对象(iw, aw, i", a())
如果你是 Vim 用户,这比默认的输入体验好太多。用 `/config` 可以设置为永久开启。
## 语音模式
输入 `/voice` 激活语音模式,长按空格键说话,松开发送。适合不想打字但又需要给 Claude 交代任务的时候。按键可以在 `keybindings.json` 中自定义。
# Gestion des sessions
## `/rename` : Nommer la session
```
/rename my-auth-refactor
```
Donnez un nom à votre session actuelle. L'avantage est que dans le sélecteur interactif de sessions (`claude --resume`), les sessions nommées peuvent être sélectionnées et restaurées directement sans avoir à appuyer sur Entrée pour confirmer ; vous pouvez aussi lancer directement depuis le terminal avec `claude --resume my-auth-refactor`.
Dans le sélecteur, il suffit de taper du texte pour rechercher et filtrer. Il prend aussi en charge ces raccourcis clavier : `Ctrl+V` pour prévisualiser la session, `Ctrl+R` pour renommer, `Ctrl+A` pour basculer l'affichage de tous les projets et `Ctrl+B` pour filtrer par branche.
## `/branch` : Bifurquer la conversation
Comme les branches git, `/branch` crée une bifurcation au point actuel de la conversation. Vous pouvez essayer différentes approches dans la bifurcation sans affecter la conversation d'origine. Si le résultat ne vous convient pas, revenez simplement à la branche d'origine et continuez.
## `/color` : Colorer la fenêtre
Définissez une couleur pour la barre de prompt de la session actuelle. Prend en charge red, blue, green, yellow, purple, orange, pink et cyan.
## Boîte à outils complète de gestion des sessions en ligne de commande
```bash
# 恢复当前目录最近的会话
claude --continue
# 打开会话选择器,或按名称恢复
claude --resume
claude --resume my-auth-refactor
# 启动时直接命名会话
claude -n "auth-refactor"
# Fork 上一次会话(保留上下文,创建新分支)
claude -c --fork-session
# 恢复与特定 PR 关联的会话
claude --from-pr 123
# 在隔离的 git worktree 中启动
claude -w
```
Les sessions de Claude Code sont sauvegardées automatiquement (pas besoin de Ctrl+S), vous pouvez donc reprendre votre dernier travail avec `--continue` à chaque ouverture du terminal.
## `claude --remote` : Continuer sur un autre appareil
```bash
claude --remote "your task description"
```
Lancez une session web que vous pouvez poursuivre sur claude.ai ou l'application mobile.
## `/remote-control` : Contrôler Claude local depuis votre téléphone
Tapez `/remote-control` dans Claude Code sur votre ordinateur, et un code de connexion sera généré. Entrez ensuite ce code dans l'application Claude sur votre téléphone pour contrôler à distance la session Claude Code locale — donnez des instructions à Claude sur votre ordinateur depuis votre téléphone.
# Raccourcis Clavier
Le systeme de raccourcis de Claude Code est bien plus riche que ce que la plupart des utilisateurs imaginent -- appuyez sur `?` pour afficher tous les raccourcis disponibles dans votre contexte actuel.
## Shift+Tab : Basculement cyclique des modes
Il s'agit probablement du raccourci le plus important. Appuyer sur `Shift+Tab` permet de basculer entre trois modes :
**Normal Mode → Auto-Accept Mode → Plan Mode → Normal Mode**
Inutile de saisir `/plan` ou `/auto-accept` manuellement -- une seule touche suffit. Mon habitude : lorsque je recois une nouvelle tache, j'appuie deux fois pour passer en Plan Mode, je valide l'approche, puis j'appuie une fois de plus pour passer en Auto-Accept et laisser Claude executer de maniere autonome.
## Esc + Esc : La machine a remonter le temps
Appuyez deux fois sur `Esc` et un menu de retour arriere (Rewind) apparait :
* **Restaurer le code et la conversation** : vous revenez a un point de controle anterieur, les fichiers et l'historique de conversation sont tous deux restaures
* **Conversation uniquement** : les messages sont annules mais vos modifications de code actuelles sont conservees
* **Code uniquement** : les modifications de fichiers sont annulees mais l'historique de conversation est conserve
Claude suit automatiquement chaque modification de fichier comme point de controle. Cette approche est bien plus fine que `git checkout .`, car vous pouvez revenir a n'importe quelle modification individuelle, et pas seulement au dernier commit.
Point important : seuls les fichiers modifies directement par Claude via ses outils sont suivis. Les fichiers que vous avez modifies manuellement, les `git push` ou autres operations externes ne peuvent pas etre annules.
## Ctrl+S : Stockage temporaire de prompts (Prompt Stash)
Vous etes en train d'ecrire un prompt et vous devez gerer autre chose en priorite ? Appuyez sur `Ctrl+S` et votre saisie actuelle est mise de cote :
Vous pouvez ensuite saisir une autre commande ou question. Une fois ce message envoye, le contenu stocke se **restaure automatiquement** dans le champ de saisie pour que vous puissiez reprendre la ou vous en etiez.
Considerez cela comme un `git stash` applique aux prompts. Exemple concret : vous redigez une longue description de refactoring, puis vous realisez que vous souhaitez d'abord que Claude verifie un fichier -- appuyez sur `Ctrl+S` pour stocker votre description, posez votre question sur le fichier, et une fois la reponse obtenue, votre description revient automatiquement.
## Ctrl+B : Envoyer des taches en arriere-plan
Claude traite une tache longue (comme un refactoring de grande envergure) et vous souhaitez travailler sur autre chose ? Appuyez sur `Ctrl+B` pour envoyer la tache en cours en arriere-plan -- votre terminal est immediatement disponible pour de nouvelles saisies.
Utilisez `Ctrl+T` pour consulter la liste des taches en arriere-plan, et appuyez deux fois sur `Ctrl+F` pour arreter tous les agents en arriere-plan.
> Utilisateurs de tmux : la touche de prefixe par defaut de tmux est egalement `Ctrl+B`. Vous devrez donc appuyer deux fois pour declencher la fonctionnalite d'arriere-plan de Claude.
## Ctrl+G : Rediger de longs prompts dans votre editeur
Il arrive que vous ayez besoin de fournir a Claude un ensemble d'instructions detaillees, et la saisie dans le terminal est peu pratique. Appuyez sur `Ctrl+G` pour ouvrir votre `$EDITOR` par defaut (VS Code, Vim, etc.), redigez votre prompt dans l'editeur, et celui-ci est automatiquement envoye a Claude lorsque vous enregistrez et fermez le fichier.
Pour modifier l'editeur par defaut, configurez-le dans votre fichier shell (`~/.zshrc` ou `~/.bashrc`) :
```bash
# VS Code
export EDITOR="code --wait"
# Zed
export EDITOR="zed --wait"
# Vim
export EDITOR="vim"
```
Le parametre `--wait` est essentiel -- il indique a l'editeur d'attendre que vous fermiez le fichier avant de rendre le controle. Sans cela, Claude recoit immediatement un contenu vide. Les editeurs en terminal comme Vim bloquent naturellement, ce parametre n'est donc pas necessaire pour eux.
Particulierement utile pour les descriptions d'exigences sur plusieurs paragraphes ou le collage de materiaux de reference volumineux. En Plan Mode, vous pouvez meme utiliser `Ctrl+G` pour modifier directement dans votre editeur le plan genere par Claude.
## Cmd+T : Activer la reflexion etendue
Le raccourci par defaut est `Cmd+T` (ou `Meta+T` sous Windows/Linux) et il active ou desactive le mode de reflexion etendue (Extended Thinking). Lorsqu'il est active, Claude raisonne plus en profondeur avant de repondre -- ideal pour les decisions d'architecture complexes ou la recherche de bugs difficiles.
Attention toutefois : la plupart des terminaux (iTerm2, Terminal.app, Warp, etc.) interceptent `Cmd+T` pour ouvrir un nouvel onglet, ce qui rend ce raccourci souvent inutilisable en pratique. Deux solutions : utilisez `/keybindings` pour le reassigner a une touche sans conflit, ou utilisez simplement la commande `/effort` pour modifier la profondeur de reflexion (meme effet, avec un controle plus fin du niveau).
## Raccourcis Readline
Le champ de saisie de Claude Code prend en charge les raccourcis Readline standard -- les habitues du terminal s'y retrouveront immediatement :
| Raccourci | Fonction |
| ----------------- | --------------------------------------- |
| Ctrl+A | Aller au debut de la ligne |
| Ctrl+E | Aller a la fin de la ligne |
| Ctrl+W | Supprimer le mot precedent |
| Ctrl+U | Supprimer jusqu'au debut de la ligne |
| Ctrl+K | Supprimer jusqu'a la fin de la ligne |
| Ctrl+Y | Coller le dernier texte supprime |
| Alt+Y | Parcourir l'historique des suppressions |
| Option+Left/Right | Sauter par mot (Mac) |
## Raccourcis d'approbation : `y/n/d/e`
Lorsque Claude propose une modification de fichier et attend votre confirmation, quatre raccourcis a une seule touche controlent le processus :
* `y` : Accepter
* `n` : Refuser
* `d` : Afficher le diff complet
* **`e` : Modifier avant d'accepter**
`e` est le plus meconnu et pourtant le plus utile -- il vous permet d'ajuster les modifications de Claude avant qu'elles ne soient appliquees. Quelques lignes ne vous conviennent pas ? Inutile de tout refuser et recommencer, appuyez simplement sur `e` et corrigez.
## Aide-memoire
| Raccourci | Fonction |
| ----------- | --------------------------------------------------------------------------------------------------------------------- |
| Shift+Tab | Basculer entre les modes : Normal → Auto-Accept → Plan |
| Esc+Esc | Ouvrir le menu de retour arriere |
| Ctrl+S | Stocker la saisie actuelle, restauration automatique apres le prochain envoi |
| Ctrl+B | Envoyer la tache en cours en arriere-plan |
| Ctrl+T | Consulter la liste des taches en arriere-plan |
| Ctrl+F (x2) | Arreter tous les agents en arriere-plan |
| Ctrl+G | Rediger un prompt dans un editeur externe |
| Ctrl+O | Activer l'affichage detaille des outils |
| Cmd+T | Activer la reflexion etendue (peut etre intercepte par le terminal ; envisagez de reassigner ou d'utiliser `/effort`) |
| `\` + Enter | Saisie multiligne (aucune configuration requise) |
| Shift+Enter | Saisie multiligne (necessite d'executer `/terminal-setup` au prealable) |
| Up / Down | Parcourir l'historique des saisies |
| Ctrl+R | Rechercher dans l'historique des commandes |
| Ctrl+L | Effacer l'ecran (historique conserve) |
| Ctrl+C | Annuler la generation en cours |
| Ctrl+D | Quitter Claude Code |
| `?` | Afficher tous les raccourcis disponibles |
## Raccourcis personnalises
Si les raccourcis par defaut ne correspondent pas a vos habitudes, utilisez `/keybindings` pour ouvrir `~/.claude/keybindings.json` et les personnaliser. Les modifications prennent effet immediatement -- aucun redemarrage n'est necessaire.
La syntaxe prend en charge les combinaisons de touches (par exemple, `ctrl+shift+c`) et le mode Chord (par exemple, `ctrl+k ctrl+s` -- appuyez sur Ctrl+K, relacher, puis appuyez sur Ctrl+S). Il existe 16 contextes de liaison differents (Chat, Autocomplete, Confirmation, DiffDialog, etc.), chacun avec son propre ensemble d'actions assignables.
# Réflexion et contrôle des modèles
## Contrôler la profondeur de réflexion avec des mots-clés
L'ajout de mots-clés spécifiques dans vos prompts déclenche différents niveaux de budget de réflexion. Il s'agit d'une fonctionnalité exclusive à Claude Code (non disponible sur claude.ai web) :
| Mot-clé | Budget de réflexion | Cas d'utilisation |
| ----------------------------- | ------------------- | -------------------------------------------- |
| `think` | \~4 000 tokens | Questions de programmation courantes |
| `think hard` / `megathink` | \~10 000 tokens | Logique complexe, dépendances multi-fichiers |
| `think harder` / `ultrathink` | \~31 999 tokens | Conception d'architecture, bugs difficiles |
En pratique, j'ajoute généralement `think hard` lorsque Claude donne une réponse superficielle, puis je repose la question. Pour les problèmes particulièrement complexes (comme le débogage à travers plusieurs services), j'utilise directement `ultrathink`.
## `/effort` : Contrôler la profondeur de réflexion
En plus des mots-clés (think / ultrathink), vous pouvez utiliser `/effort` pour définir directement la profondeur de réflexion :
```
/effort low # 简单任务,跳过深度思考,更快更省
/effort high # 复杂任务,深度推理
/effort max # 最大思考预算(仅 Opus)
/effort auto # 让 Claude 自己判断
```
Le paramètre reste actif pendant toute la session. Utilisez `low` pour les modifications de fichiers simples et `max` pour la conception d'architectures complexes — vous économisez ainsi de l'argent sans sacrifier la qualité.
## Le mot-clé `use subagents`
Ajoutez `use subagents` à la fin de n'importe quelle requête, et Claude décomposera la tâche en plusieurs sous-agents qui s'exécutent en parallèle. Cela est non seulement plus rapide, mais permet également de garder la fenêtre de contexte de l'agent principal propre.
Boris a spécifiquement mentionné ce point sur Twitter : déléguer les tâches individuelles aux sous-agents pour maintenir le contexte de l'agent principal focalisé.
## `opusplan` : La meilleure stratégie de modèles en termes de rapport qualité-prix
Résumé en une phrase : **Opus réfléchit, Sonnet exécute**.
## `/model` : Changer de modèle
Utilisez `/model` pour changer de modèle à tout moment pendant une session. Par exemple, utilisez Sonnet au quotidien, basculez temporairement sur Opus pour les problèmes complexes, puis revenez une fois terminé.
## Contrôle du style de sortie
Dans `/config`, sélectionnez « Output style » — il existe deux modes peu courants mais très utiles :
* **Mode Explanatory** : Claude insère des « points de connaissance » entre les tâches, expliquant les frameworks et les patterns de code pertinents — idéal pour découvrir un nouveau projet
* **Mode Learning** : Un mode d'apprentissage collaboratif où Claude ajoute des marqueurs `TODO(human)` dans le code pour que vous les implémentiez vous-même, au lieu de vous donner directement la réponse
Vous pouvez également créer des fichiers de style de sortie personnalisés (au format Markdown) dans `~/.claude/output-styles/` pour modifier directement le prompt système. Attention : les styles de sortie personnalisés **remplaceront entièrement** le prompt système de programmation par défaut, sauf si vous définissez `keep-coding-instructions: true`.
# Introduction conceptuelle
## Introduction
En octobre 2025, Anthropic a discrètement lancé une nouvelle fonctionnalité appelée Claude Skills. Malgré un lancement modeste, le blogueur technologique renommé Simon Willison l'a qualifiée de « peut-être plus importante que MCP » et a prédit qu'elle déclencherait une « explosion cambrienne » dans l'écosystème des outils IA.
Cette évaluation est loin d'être exagérée. Si vous utilisez régulièrement des assistants IA, vous avez probablement rencontré une frustration familière : chaque nouvelle conversation vous oblige à ressaisir les mêmes instructions de flux de travail ; vous parvenez enfin à ajuster l'IA à votre satisfaction, pour tout recommencer depuis le début dans une nouvelle fenêtre de discussion. Skills a été conçu précisément pour résoudre ce point de friction.
## Comprendre Claude Skills
Imaginez que vous êtes le dirigeant d'une entreprise et que vous remettez à chaque nouvel employé un manuel opérationnel détaillant les flux de travail, les directives de marque et les procédures standard pour traiter les problèmes courants. Claude Skills est essentiellement ce « manuel opérationnel » pour un assistant IA — lui permettant d'accomplir des tâches spécifiques de manière reproductible et standardisée.
D'un point de vue technique, les Skills sont des dossiers contenant des instructions, des scripts et des ressources que Claude peut charger dynamiquement à la demande. Chaque Skill enseigne à Claude comment gérer un type de tâche particulier de manière cohérente, et ces connaissances persistent d'une conversation à l'autre. Cela signifie que vous n'avez besoin de le « former » qu'une seule fois — par la suite, Claude se souviendra de la marche à suivre à chaque utilisation.
### Les trois composants essentiels
Un Skill complet se compose de trois parties :
| Composant | Fonction | Obligatoire |
| -------------------------- | -------------------------------------------------------------------------------------------- | ----------- |
| **SKILL.md** | Document d'instructions principal contenant les métadonnées et les instructions détaillées | Obligatoire |
| **Documents de référence** | Directives de marque, documents de politique, modèles et autres informations complémentaires | Facultatif |
| **Scripts** | Code Python/JavaScript pour les calculs complexes ou les opérations sur les fichiers | Facultatif |
SKILL.md est « l'âme » de l'ensemble du Skill. Sa structure de base est la suivante :
```yaml
---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---
# Your Skill Name
## Instructions
Provide clear, step-by-step guidance for Claude.
## Examples
Show concrete examples of using this Skill.
## Guidelines
- Guideline 1
- Guideline 2
```
Le frontmatter YAML en début de fichier contient deux champs clés : `name` est l'identifiant du Skill, limité à 64 caractères ; `description` indique à Claude ce que fait le Skill et quand l'utiliser, limité à 200 caractères. C'est sur cette description que Claude s'appuie pour déterminer quand invoquer un Skill donné — plus elle est claire et précise, plus la probabilité que le Skill soit correctement déclenché est élevée.
### Cas d'utilisation
Les Skills couvrent un large éventail d'applications dans les tâches répétitives du quotidien :
**Génération de documents** : Création par lots de feuilles de calcul Excel, de présentations PowerPoint, de documents Word et de rapports PDF. Anthropic fournit même un ensemble officiel de Skills documentaires prêts à l'emploi.
**Conformité de marque** : Intégrez les couleurs de votre marque, les règles d'utilisation du logo, les directives d'espacement et le ton rédactionnel dans un Skill, garantissant que tout le contenu généré par l'IA respecte les standards de la marque.
**Comptes rendus de réunions** : Résumez automatiquement les enregistrements de réunions, extrayez les points d'action, assignez les responsables et générez les e-mails de suivi.
**Analyse de données** : Exécutez des flux d'analyse standardisés, tels que la veille concurrentielle (extraction structurée des mises à jour produits, des changements de prix et des commentaires d'analystes) ou l'analyse financière (analyse des rapports de résultats et construction de modèles financiers).
**Gestion de projet** : Élaborez des plans de projet à partir d'objectifs, suggérez des jalons et générez des rapports hebdomadaires ou des synthèses pour les investisseurs.
## Architecture de divulgation progressive
L'aspect le plus élégant des Skills réside dans la manière dont l'information est chargée. Les descriptions d'outils MCP traditionnels peuvent consommer des milliers, voire des dizaines de milliers de tokens, alors que les métadonnées des Skills n'occupent que quelques dizaines de tokens. Cela signifie que vous pouvez activer simultanément un grand nombre de Skills sans craindre que les descriptions d'outils saturent la fenêtre de contexte.
Cette efficacité provient d'un patron architectural appelé **divulgation progressive** (Progressive Disclosure). Les Skills utilisent une structure d'information à trois couches qui charge le contenu à la demande — à l'image d'un manuel avec une table des matières :
```
📚 Manuel des Skills
│
├─ 📋 Table des matières ───────────────── [Couche métadonnées] Préchargée au démarrage
│ │
│ │ name: "weekly-report"
│ │ description: "Générer des rapports hebdomadaires standardisés"
│ │
│ │ ✓ Seulement 30-50 tokens
│ │ ✓ Toutes les métadonnées des Skills visibles simultanément
│ │
│
├─ 📖 Chapitres ────────────────────────── [Couche document principal] Chargée quand pertinent
│ │
│ │ # Générateur de rapport hebdomadaire
│ │
│ │ ## Instructions
│ │ Générer les rapports hebdomadaires selon cette structure...
│ │
│ │ ## Exemples
│ │ Entrée : Fonctionnalité de connexion terminée cette semaine...
│ │ Sortie : ### Réalisations de la semaine ...
│ │
│ │ ⚡ Développé uniquement quand Claude le juge nécessaire
│ │ 📊 Consomme des centaines à des milliers de tokens
│ │
│
└─ 📎 Annexes ──────────────────────────── [Couche ressources de référence] Chargée au besoin
│
│ references/
│ ├── brand-guide.md Directives de marque
│ ├── template.xlsx Modèle de rapport
│ └── examples/ Rapports précédents
│
│ 🔍 Chargée uniquement quand explicitement nécessaire
│ 📦 Peut contenir d'abondantes ressources de référence
```
Vous consultez d'abord la table des matières pour voir quels chapitres sont disponibles (couche métadonnées), puis vous ouvrez le chapitre pertinent pour le lire (couche document principal), et enfin vous consultez les annexes si vous avez besoin de plus de détails (couche ressources de référence).
| Couche | Contenu | Moment de chargement | Coût en tokens |
| ---------------------------------- | ------------------------------------ | ----------------------- | -------------------- |
| **Couche métadonnées** | name + description | Préchargée au démarrage | 30-50 |
| **Couche document principal** | Contenu complet de SKILL.md | Chargée quand pertinent | Centaines à milliers |
| **Couche ressources de référence** | Fichiers de référence, modèles, etc. | Chargée au besoin | À la demande |
Cela s'aligne parfaitement avec la nature fondamentale des grands modèles de langage — « fournir du texte au modèle pour qu'il comprenne ». Les Skills n'introduisent pas de protocoles complexes ni d'appels API ; ils utilisent des structures textuelles soigneusement organisées pour aider l'IA à acquérir et appliquer efficacement les connaissances. Simon Willison a qualifié cette conception d'« absurdement élégante », précisément parce qu'elle résout un problème complexe de la manière la plus directe possible.
## Avantages clés
### Efficacité en tokens
L'architecture de divulgation progressive des Skills offre une efficacité exceptionnelle en tokens. Une comparaison simple illustre la différence :
| Approche | Tokens au démarrage | Total pour 100 Skills |
| ----------------------------------- | ------------------------------- | ------------------------------------ |
| Traditionnelle (chargement complet) | Milliers à dizaines de milliers | Peut dépasser la fenêtre de contexte |
| Skills (divulgation progressive) | 30-50 | 3 000-5 000 |
Puisque les métadonnées de chaque Skill n'occupent que quelques dizaines de tokens, vous pouvez activer simultanément des dizaines, voire des centaines de Skills, le contenu complet étant chargé à la demande — sans jamais gaspiller le précieux espace de contexte.
### Composabilité
Plusieurs Skills peuvent collaborer automatiquement. Lorsque vous présentez une tâche complexe, Claude identifie intelligemment les Skills à invoquer et les coordonne pour accomplir la tâche.
Par exemple, si vous dites « Générez un rapport trimestriel à partir de ces données de ventes », Claude pourrait :
1. Invoquer un Skill d'analyse de données pour traiter les données brutes
2. Invoquer un Skill de génération de graphiques pour créer des visualisations
3. Invoquer un Skill documentaire pour produire le rapport final
Tout au long de ce processus, vous n'avez jamais besoin de spécifier manuellement quel Skill utiliser — Claude les sélectionne et les combine automatiquement en fonction des exigences de la tâche.
### Portabilité
Le même Skill fonctionne sur toutes les plateformes de l'écosystème Anthropic :
| Plateforme | Description |
| ----------- | ------------------------------------------------------------ |
| Claude.ai | Interface web, adaptée aux utilisateurs généraux |
| Claude Code | Outil en ligne de commande, adapté aux développeurs |
| API | Intégration programmatique, adaptée au développement système |
Un Skill de rédaction de marque que vous créez pour votre équipe se comportera de manière cohérente sur toutes ces plateformes, réalisant véritablement le principe **construire une fois, utiliser partout**.
> **Stratégies des autres plateformes IA** : Skills est actuellement une fonctionnalité exclusive à Anthropic. OpenAI utilise une approche à double voie avec Custom GPTs + Assistants API (deux systèmes non unifiés) ; Microsoft Copilot et Google Gemini se concentrent sur l'intégration approfondie au sein de leurs écosystèmes respectifs plutôt que sur des modules de compétences réutilisables. Claude Skills est considéré comme un facteur de différenciation significatif.
### Données d'efficacité
Selon les benchmarks internes d'Anthropic, les équipes utilisant Skills ont **réduit de 73 % le temps consacré à l'ingénierie de prompts répétitifs**. Il ne s'agit pas seulement de gains d'efficacité — plus important encore, cela représente la standardisation et la réutilisabilité des flux de travail. Les membres de l'équipe n'ont plus besoin de maintenir chacun leur propre ensemble de prompts ; ils partagent plutôt un ensemble unique de Skills vérifiés.
## Résumé
Claude Skills est, fondamentalement, un ensemble de **manuels réutilisables pour les assistants IA**. Grâce à l'architecture de divulgation progressive, ils atteignent une efficacité exceptionnelle en tokens, permettant à l'IA de maîtriser de vastes connaissances spécialisées sans consommer le précieux espace de contexte.
Retenez trois concepts clés, et vous aurez saisi l'essence des Skills :
| Concept | Signification |
| -------------- | ---------------------------------------------------------------------------------------------- |
| **Efficace** | Les métadonnées n'occupent que quelques dizaines de tokens ; le contenu se charge à la demande |
| **Composable** | Plusieurs Skills collaborent automatiquement |
| **Portable** | Expérience cohérente sur toutes les plateformes |
Maintenant que vous comprenez les concepts, le prochain article, [Guide pratique de Claude Skills](/fr/docs/notes/claude-skills/practice), vous guidera dans la pratique : comment activer et installer des Skills, créer votre premier Skill personnalisé et éviter les pièges courants.
Si vous souhaitez formaliser davantage vos flux de travail, consultez [Qu'est-ce que le développement piloté par les spécifications](/fr/docs/notes/speckit/concept) pour apprendre à faire passer la programmation IA de « l'intuition » à « l'ingénierie ».
# Guide pratique
## Rappel rapide
Dans l'[article précédent](/fr/docs/notes/claude-skills/concept), nous avons abordé les concepts fondamentaux des Skills : des manuels réutilisables pour les assistants IA qui atteignent une efficacité exceptionnelle en tokens grâce à une architecture de divulgation progressive, avec trois atouts clés — efficacité, composabilité et portabilité. Cet article adopte une approche pratique pour vous aider à comprendre comment les Skills se distinguent des autres fonctionnalités, à apprendre à activer, installer et créer des Skills, et à maîtriser les bonnes pratiques tout en évitant les pièges courants.
## Comparaison des fonctionnalités
L'écosystème Claude propose plusieurs fonctionnalités, et il peut être difficile de les distinguer au premier abord. Le tableau ci-dessous offre un aperçu rapide :
| Fonctionnalité | Ce que c'est | Idéal pour | Persistance |
| -------------- | ------------------------- | --------------------------------------------------- | ---------------------------------------- |
| **Skills** | Paquets d'expertise | Tâches répétitives, flux de travail standardisés | Persistant entre les conversations |
| **Prompts** | Instructions instantanées | Requêtes ponctuelles | Conversation en cours uniquement |
| **Projects** | Bases de connaissances | Informations contextuelles, documentation de projet | Au sein de l'espace de travail du projet |
| **MCP** | Connecteurs | Données externes, appels API | Connexion continue |
| **Subagents** | Sous-agents | Délégation de tâches, traitement parallèle | Entre les sessions |
### Skills vs MCP
C'est la source de confusion la plus fréquente. La distinction fondamentale : **MCP connecte Claude aux données ; Skills enseigne à Claude comment traiter les données**. Ils se complètent plutôt qu'ils ne se remplacent.
| Dimension | Skills | MCP |
| -------------------------- | ------------------------------------------------------------ | ------------------------------------------------------- |
| **Fonction principale** | Enseigne à Claude comment exécuter des tâches | Connecte Claude aux systèmes externes |
| **Consommation de tokens** | Très faible (dizaines de tokens) | Plus élevée (milliers à dizaines de milliers de tokens) |
| **Complexité technique** | Simple (Markdown + YAML) | Complexe (spécification de protocole complète) |
| **Cas d'usage typiques** | Rédaction de marque, génération de rapports, flux de travail | Requêtes en base de données, appels API, services cloud |
| **Portabilité** | Claude.ai/Code/API | Adopté par plusieurs fournisseurs de modèles |
Une fois cette distinction comprise, vous saurez quand utiliser chacun. Utilisez MCP lorsque vous avez besoin d'interroger des bases de données, d'appeler des API ou d'accéder à des services cloud. Utilisez Skills lorsque vous devez suivre un style rédactionnel spécifique, exécuter des flux de travail standardisés ou réutiliser une expertise métier.
La bonne pratique consiste à combiner les deux : utilisez MCP pour vous connecter à votre système CRM et extraire les données clients, puis utilisez Skills pour définir comment analyser ces données et générer des rapports.
### Skills vs Subagents
La distinction fondamentale : **Skills rend Claude plus compétent pour certains types de tâches ; Subagents permet à Claude de déléguer des tâches à des « spécialistes indépendants ».**
| Dimension | Skills | Subagents |
| ----------------------- | ----------------------------------------------------------------- | -------------------------------------------------------- |
| **Fonction principale** | Fournit expertise et instructions | Sous-agents indépendants qui exécutent des tâches |
| **Contexte** | Injecté dans le contexte de la conversation principale | Possède sa propre fenêtre de contexte indépendante |
| **Cas d'usage** | Rendre Claude plus performant sur des types de tâches spécifiques | Tâches complexes et multi-étapes indépendantes |
| **Activation** | Correspondance automatique basée sur la description | Invocation manuelle ou délégation automatique par Claude |
| **Portabilité** | Claude.ai/Code/API | Claude Code et Agent SDK uniquement |
Voyez les choses ainsi : les Skills sont comme des supports de formation — ils enseignent à Claude comment faire quelque chose. Les [Subagents](/fr/docs/notes/claude-subagent) sont comme des employés dédiés — ils disposent de leur propre espace de travail (contexte) et de leurs propres autorisations (outils), accomplissent les tâches de manière indépendante et rendent compte des résultats.
Les deux peuvent être combinés : par exemple, un sous-agent de revue de code peut charger des Skills de bonnes pratiques spécifiques à un langage, réalisant une combinaison « spécialiste + expertise métier ». Selon une [recherche d'Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system), les systèmes multi-agents (Claude Opus 4 comme orchestrateur + Claude Sonnet 4 en sous-agents) ont obtenu des scores supérieurs de 90,2 % par rapport aux configurations mono-agent dans les évaluations internes.
### Skills vs commandes slash
Si vous avez utilisé Claude Code, vous connaissez les [commandes slash](/fr/blog/claude-code-best-practices) comme `/commit` et `/review`. La distinction fondamentale : **Skills s'active automatiquement selon le contexte ; les commandes slash nécessitent une saisie manuelle pour être déclenchées.**
| Dimension | Skills | Commandes slash |
| ------------------------------ | -------------------------------------------------- | ---------------------------------------------- |
| **Activation** | Automatique (correspondance par contexte) | Saisie manuelle (ex. : `/commit`) |
| **Condition de déclenchement** | Claude évalue la pertinence d'après la description | L'utilisateur saisit explicitement la commande |
| **Cas d'usage** | Renforcement de capacités « toujours actif » | Opérations explicites et reproductibles |
| **Perception utilisateur** | Invisible, s'applique automatiquement | Nécessite de mémoriser les noms de commandes |
Exemple : lorsque vous tapez `/commit`, Claude exécute un flux de commit prédéfini — c'est une commande slash. Lorsque vous dites « aidez-moi à rédiger un rapport hebdomadaire », Claude identifie et charge automatiquement le Skill de génération de rapport hebdomadaire sans aucune commande — ce sont les Skills.
Pour retenir facilement : les commandes slash sont des raccourcis clavier que vous déclenchez manuellement ; les Skills sont des connaissances de fond que Claude utilise à sa propre initiative.
### Skills vs Plugins
Les Plugins constituent le mécanisme de paquets d'extension de Claude Code. La distinction fondamentale : **Skills sont des extensions de capacités à activation automatique ; les Plugins sont des configurations de flux de travail complètes, empaquetées et distribuables.**
| Dimension | Skills | Plugins |
| ----------------------- | ---------------------------------------- | ------------------------------------------ |
| **Fonction principale** | Extension de capacités | Distribution de flux de travail empaquetés |
| **Activation** | Activation automatique selon le contexte | Composants fusionnés après installation |
| **Portée** | Multi-plateforme (Claude.ai/Code/API) | Claude Code uniquement |
| **Contenu** | Instructions + scripts + ressources | Commandes slash + hooks + skills |
| **Distribution** | Dossier autonome | Installation via marketplace |
Le point clé : les Plugins peuvent contenir des Skills (dans leur répertoire `skills/`) et représentent une unité d'empaquetage plus large. Lorsque vous installez un Plugin, ses Skills sont automatiquement activés, les commandes slash apparaissent dans l'auto-complétion et les hooks sont fusionnés avec votre configuration existante.
En résumé : utilisez les Skills pour étendre les capacités de Claude ; utilisez les Plugins pour distribuer des configurations de flux de travail standardisées au sein de votre équipe.
## Tutoriel pratique
### Option 1 : Activer les Skills intégrés
C'est le moyen le plus simple pour commencer. Anthropic fournit un ensemble de Skills documentaires pratiques :
| Skill | Fonctionnalité |
| --------------------- | ---------------------------------------------------------------------------------------- |
| **Excel (xlsx)** | Créer des feuilles de calcul, analyser des données, générer des rapports avec graphiques |
| **PowerPoint (pptx)** | Créer des présentations, modifier des diapositives, analyser le contenu de présentations |
| **Word (docx)** | Créer des documents, modifier du contenu, mettre en forme du texte |
| **PDF (pdf)** | Générer des documents et rapports PDF formatés |
**Étapes d'activation** :
1. Connectez-vous à [Claude.ai](https://claude.ai)
2. Cliquez sur votre avatar en haut à droite et accédez aux **Settings**
3. Trouvez l'option **Capabilities**
4. Activez les Skills dont vous avez besoin
Une fois activés, testez immédiatement : « Créez une feuille de calcul Excel pour le budget des ventes du T3 avec des ventilations mensuelles et des totaux. »
> **Remarque** : Nécessite un plan Pro, Max, Team ou Enterprise, et la fonctionnalité d'exécution de code doit être activée.
### Option 2 : Installer des Skills communautaires
Si vous utilisez Claude Code, vous pouvez installer des Skills contribués par la communauté via des commandes.
**Installation depuis le marketplace de plugins** :
```bash
# Ajouter le dépôt officiel de Skills
/plugin marketplace add anthropics/skills
# Installer le paquet de Skills documentaires
/plugin install document-skills@anthropic-agent-skills
# Installer le paquet de Skills d'exemple
/plugin install example-skills@anthropic-agent-skills
```
**Emplacements de stockage des Skills** :
| Emplacement | Chemin | Description |
| ----------------- | ------------------- | ------------------------------------------- |
| Skills personnels | `~/.claude/skills/` | Disponibles uniquement pour vous |
| Skills de projet | `.claude/skills/` | Versionnés avec git, partagés avec l'équipe |
### Option 3 : Créer un Skill personnalisé
C'est ici que les Skills révèlent toute leur puissance — créer des flux de travail adaptés à vos besoins.
**Étape 1 : Créer la structure de dossiers**
```bash
mkdir -p ~/.claude/skills/weekly-report
cd ~/.claude/skills/weekly-report
```
Un dossier de Skill complet pourrait ressembler à ceci :
```
weekly-report/
├── SKILL.md # Instructions principales (obligatoire)
├── template.md # Modèle de rapport (facultatif)
└── examples/ # Exemples de rapports (facultatif)
├── good-example.md
└── bad-example.md
```
**Étape 2 : Rédiger le SKILL.md**
SKILL.md est le cœur de l'ensemble du Skill. Il se compose de deux parties : le frontmatter YAML (métadonnées) et le corps en Markdown (instructions détaillées).
**Métadonnées obligatoires** :
| Champ | Exigence | Description |
| ------------- | ------------------- | ----------------------------------------------------------- |
| `name` | 64 caractères max. | Identifiant unique du Skill |
| `description` | 200 caractères max. | Indique à Claude quand utiliser ce Skill (très important !) |
**Métadonnées facultatives** :
| Champ | Description |
| --------------- | -------------------------------------------------- |
| `dependencies` | Paquets requis, ex. : `python>=3.8, pandas>=1.5.0` |
| `allowed-tools` | Liste des outils autorisés |
| `model` | Remplacement optionnel du modèle |
Un exemple complet de Skill de génération de rapport hebdomadaire :
```yaml
---
name: weekly-report
description: 根据本周工作内容生成标准化的周报,包含进展、问题和下周计划
---
# 周报生成助手
## 使用场景
当用户需要生成周报、工作总结或进度汇报时,使用此技能。
## 输出格式
请按以下结构生成周报:
### 本周完成
- 列出已完成的主要工作项
- 每项包含简短说明和成果
### 进行中
- 列出正在进行的工作
- 标注当前进度和预期完成时间
### 遇到的问题
- 列出阻碍进展的问题
- 如果有,说明需要的支持
### 下周计划
- 列出下周的主要任务
- 按优先级排序
## 风格要求
- 使用简洁的表达
- 避免过于技术化的术语
- 突出成果和影响
## 示例
**输入**:这周完成了用户登录功能,修复了 3 个 bug,参加了产品评审。
**输出**:
### 本周完成
- 用户登录功能开发:完成前后端联调,支持邮箱和手机号登录
- Bug 修复:解决了 3 个高优先级问题,提升系统稳定性
### 进行中
- (无)
### 遇到的问题
- (无)
### 下周计划
- 开始用户注册功能开发
- 编写单元测试用例
```
**Étape 3 : Tester**
Testez dans Claude : « Aidez-moi à générer le rapport de cette semaine. Cette semaine, j'ai terminé la fonctionnalité de connexion utilisateur, corrigé 3 bugs et assisté à deux réunions de revue produit. »
### Utiliser le Skill Creator
Si vous préférez ne pas rédiger le SKILL.md de zéro, Claude dispose d'un Skill intégré skill-creator qui vous guide de manière interactive :
```
Help me create a skill for [your workflow]
```
Claude vous posera une série de questions pour clarifier vos besoins, puis générera un brouillon de SKILL.md.
## Principes techniques
### Les Skills comme système de méta-outils
Fondamentalement, les Skills constituent un **système de méta-outils** — ils n'exécutent pas directement du code mais injectent des instructions spécialisées dans le contexte de la conversation, modifiant la manière dont Claude raisonne.
Lorsque vous déclenchez un Skill, deux choses se produisent :
1. **Message de métadonnées** : Un indicateur d'état visible montrant quel Skill est en cours de chargement
2. **Prompt du Skill** : Les instructions complètes de SKILL.md sont envoyées à Claude, masquées à l'utilisateur
### Mécanisme de découverte et de sélection
Comment Claude sait-il quel Skill invoquer ? La réponse : **il s'appuie entièrement sur la compréhension du langage.**
Le nom et la description de tous les Skills activés sont formatés en une liste dynamique intégrée au prompt système. Lorsque vous envoyez un message, Claude utilise ses capacités natives de compréhension du langage pour faire correspondre votre intention et décider d'invoquer ou non un Skill donné.
C'est pourquoi le champ `description` est si crucial — c'est la seule base de décision de Claude. Il n'y a pas de routage algorithmique complexe ; la décision se fait entièrement dans le processus de raisonnement de Claude.
## Bonnes pratiques
Grâce à une utilisation extensive en conditions réelles, la communauté a distillé quatre règles d'or pour la création de Skills :
**1. Rester concentré**
Un Skill doit faire une seule chose bien. Plusieurs Skills ciblés sont bien plus utiles qu'un seul Skill fourre-tout — ils sont plus faciles à maintenir et à composer.
**2. Rédiger des descriptions claires**
Le champ description détermine quand Claude invoque votre Skill, soyez donc précis sur les scénarios applicables. « Générer des rapports d'analyse trimestriels à partir de données de ventes » est une bonne description ; « traiter des données » est trop vague.
**3. Fournir des exemples**
Inclure des exemples d'entrée/sortie dans votre SKILL.md améliore significativement la cohérence des résultats, particulièrement pour les tâches ayant des exigences de formatage spécifiques.
**4. Commencer simplement**
Commencez par des instructions en Markdown pur, validez les résultats, puis envisagez d'ajouter des scripts. Augmentez la complexité progressivement.
### Résolution des problèmes courants
| Problème | Cause probable | Solution |
| ---------------------------- | ----------------------------- | ------------------------------------------------------------------- |
| Le Skill ne se déclenche pas | Description pas assez précise | Réécrire avec des descriptions de scénarios plus spécifiques |
| Le Skill ne se déclenche pas | Skill mal installé | Vérifier les chemins de fichiers et le nommage |
| Résultats incohérents | Exemples manquants | Ajouter davantage d'exemples d'entrée/sortie |
| Résultats incohérents | Instructions trop vagues | Ajouter des contraintes et des exigences de formatage |
| Chargement lent | Fichier trop volumineux | Déplacer les fichiers volumineux vers un sous-répertoire references |
### Considérations de sécurité
Les Skills peuvent exécuter du code, la sécurité est donc importante :
* **Vérifier la source** : N'utilisez que des Skills provenant de canaux de confiance
* **Examiner les scripts** : Inspectez le code des scripts dans les Skills avant installation
* **Protéger les informations sensibles** : Ne codez jamais en dur les clés API ou les mots de passe dans les Skills
* **Gérer les autorisations** : En utilisation par une équipe, soyez attentif au périmètre de partage des Skills
## Limitations actuelles
En tant que fonctionnalité émergente, Skills présente actuellement certaines limitations :
| Limitation | Détails |
| --------------------------------------- | -------------------------------------------------------------------------------------------------- |
| ~~**Écosystème Anthropic uniquement**~~ | Résolu — voir ci-dessous |
| **Pas de mécanisme de revue** | Pas encore de flux de revue ou d'audit intégré |
| **Courbe d'apprentissage** | Les équipes doivent adapter leurs flux de travail et établir des processus de gestion des versions |
| **Stade précoce** | L'écosystème est encore en évolution |
> **Mise à jour majeure (18 décembre 2025)** : Anthropic a officiellement publié Agent Skills en tant que [standard ouvert](https://agentskills.io). La spécification et le SDK de référence sont disponibles sur [agentskills.io](https://agentskills.io).
>
> **Entreprises/produits ayant adopté le standard** :
> 
>
> * **Microsoft** : Intégré dans VS Code et GitHub
> * **OpenAI** : ChatGPT et Codex CLI utilisent la même architecture
> * **Outils de développement** : Cursor, Goose, Amp, OpenCode
> * **Skills partenaires** : Atlassian, Figma, Canva, Stripe, Notion, Zapier
>
> De plus, Anthropic, OpenAI et Block ont cofondé l'[Agentic AI Foundation](https://www.linuxfoundation.org/) (hébergée par la Linux Foundation), avec Google, Microsoft et AWS qui ont également rejoint l'initiative. Cela signifie que Skills évolue d'une fonctionnalité propriétaire vers un standard industriel — les Skills écrits pour Claude Code peuvent interopérer avec OpenAI Codex CLI.
>
> Références :
>
> * [Anthropic makes agent Skills an open standard - SiliconANGLE](https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/)
> * [OpenAI Adds 'Skills' Framework - WinBuzzer](https://winbuzzer.com/2025/12/13/openai-adds-skills-framework-to-chatgpt-and-codex-cli-mirroring-anthropics-agent-standard-xcxwbn/)
## Ressources d'apprentissage
### Ressources officielles
| Ressource | Lien | Description |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------- | ----------------------------------------- |
| Dépôt GitHub Skills | [anthropics/skills](https://github.com/anthropics/skills) | Exemples officiels, 22k+ Stars |
| Documentation Claude Code | [code.claude.com/docs](https://code.claude.com/docs/en/skills) | Guide d'utilisation des Skills |
| Centre d'aide | [support.claude.com](https://support.claude.com) | FAQ |
| Blog technique | [Anthropic Engineering](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) | Analyse technique approfondie |
| Démarrage rapide API | [docs.claude.com](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/quickstart) | Guide d'intégration pour les développeurs |
| Standard ouvert Agent Skills | [agentskills.io](https://agentskills.io) | Spécification officielle et SDK |
### Sélections de la communauté
| Ressource | Lien | Description |
| --------------------- | ------------------------------------------------------------------------------------- | --------------------------------------------------------- |
| awesome-claude-skills | [VoltAgent/awesome-claude-skills](https://github.com/VoltAgent/awesome-claude-skills) | Collection de Skills sélectionnés |
| Claude Command Suite | [qdhenry/Claude-Command-Suite](https://github.com/qdhenry/Claude-Command-Suite) | 148+ commandes slash, 54 agents IA |
| Office Skills | [tfriedel/claude-office-skills](https://github.com/tfriedel/claude-office-skills) | Skills de création et d'édition de documents bureautiques |
### Lectures recommandées
| Article | Auteur |
| ------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| [Claude Skills are awesome, maybe a bigger deal than MCP](https://simonwillison.net/2025/Oct/16/claude-skills/) | Simon Willison |
| [Claude Agent Skills: A First Principles Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/) | Lee Han Chung |
| [Skills explained: How Skills compares to prompts, Projects, MCP, and subagents](https://claude.com/blog/skills-explained) | Anthropic (officiel) |
| [Understanding Claude Code: Skills vs Commands vs Subagents vs Plugins](https://www.youngleaders.tech/p/claude-skills-commands-subagents-plugins) | Young Leaders |
## Perspectives
L'émergence des Skills représente une direction importante dans l'évolution des outils IA — permettre à l'IA non seulement d'exécuter des tâches, mais aussi d'apprendre et de mémoriser des méthodes de travail spécifiques. Simon Willison a prédit que les Skills déclencheraient une « explosion cambrienne » dans l'espace des outils IA, et cette prédiction n'est pas exagérée.
À mesure que de plus en plus de développeurs et d'équipes créent et partagent des Skills, nous pouvons nous attendre à voir :
* **Des marketplaces de Skills spécialisés** : Des experts de chaque secteur empaqueté leurs connaissances en Skills réutilisables
* **Une intégration approfondie Skills + MCP** : Formant des flux de travail de bout en bout complets
* **Des plateformes de Skills de niveau entreprise** : Collaboration d'équipe, gestion des versions et contrôle d'accès
C'est le moment idéal pour commencer. Voici ce que vous pouvez faire immédiatement : connectez-vous à Claude.ai et activez les Skills documentaires. Cette semaine, essayez d'installer un Skill communautaire et de créer votre premier Skill simple. À long terme, identifier le travail répétitif au sein de votre équipe et construire progressivement une bibliothèque de Skills dédiée sera un moyen efficace d'améliorer la productivité.
### Lectures complémentaires
* [Guide complet de l'architecture système de Claude](/fr/docs/notes/claude-architecture) — La place des Skills dans l'architecture globale de Claude
* [Le guide complet des Claude Subagents](/fr/docs/notes/claude-subagent) — Plongée approfondie dans le mécanisme des Subagents
* [Mes bonnes pratiques Claude Code](/fr/blog/claude-code-best-practices) — Astuces et conseils pour l'utilisation quotidienne de Claude Code
# Analyse approfondie de Skill-Creator : utilisez les données pour stimuler le développement de vos compétences
\##Présentation
Cet article est basé sur des informations de mars 2026 et correspond à Claude Code v2.1+.
Si vous avez lu [Concept](/fr/docs/notes/claude-skills/concept) et [Practice](/fr/docs/notes/claude-skills/practice), vous devriez déjà savoir comment écrire manuellement un fichier SKILL.md - définir le frontmatter, écrire une commande, l'enregistrer dans le répertoire `.claude/skills/`, et le tour est joué.
Mais voici une question fondamentale : \*\*Comment savez-vous que vos compétences sont vraiment utiles ? \*\*
Vous avez peut-être modifié le libellé d'un paragraphe et avez l'impression qu'il fonctionne mieux, mais ce n'est que votre sentiment subjectif. Peut-être qu'avec un mot d'invite différent, la nouvelle version serait pire. Peut-être que vos compétences ne sont pas du tout améliorées par rapport au fait d'être non qualifié - Claude peut tout aussi bien réussir tout seul.
Dans les chapitres conceptuels et pratiques, le processus de développement des compétences est le suivant : **Écrit → Essayé → Se sentir bien → En ligne**. L'ensemble du processus repose sur l'intuition, il n'y a pas de quantification et il n'y a aucun moyen de répondre « Dans quelle mesure cette compétence est-elle meilleure que l'absence de compétence du tout ? » Et Skill-Creator a transformé cette chose en ingénierie : **Écrit → Tests parallèles avec/sans compétences → Comparaison A/B de tests aveugles → Notation quantitative → Itération de feedback → Vérification des données**.
C'est pourquoi Skill-Creator existe. Il vous aide non seulement à "générer un SKILL.md", mais fournit un ensemble complet de boucles créer → tester → évaluer → optimiser, vous permettant de laisser vos données parler d'elles-mêmes.
## Qu'est-ce que Skill-Creator
Skill-Creator lui-même est également une compétence - un fichier SKILL.md de 33 Ko ainsi que des fichiers de guidage de sous-agents, des scripts Python et des visionneuses HTML. Sa structure de répertoires ressemble à ceci :
```
skill-creator/
├── SKILL.md # 主指令文件(486 行)
├── agents/ # 子代理指导
│ ├── grader.md # 评分代理
│ ├── comparator.md # 盲测对比代理
│ └── analyzer.md # 分析代理
├── eval-viewer/ # 评估结果查看器
│ ├── generate_review.py
│ └── viewer.html
├── assets/
│ └── eval_review.html # 触发评估审查界面
├── scripts/ # Python 工具脚本
│ ├── run_eval.py # 运行触发评估
│ ├── run_loop.py # 优化循环
│ ├── improve_description.py # 描述优化
│ ├── aggregate_benchmark.py # 聚合基准测试
│ ├── package_skill.py # 打包为 .skill 文件
│ └── quick_validate.py # 快速校验
└── references/
└── schemas.md # JSON Schema 定义
```
L'installation est également très simple :
```bash
# 通过 Claude Code 插件市场
/plugins # 然后搜索 skill-creator 安装
# 或通过 skills.sh
npx skills add anthropics/skills -- skill skill-creator
```
## Suivez à nouveau ceci : évaluer et optimiser une compétence existante
Passons en revue le processus complet de Skill-Creator en utilisant les compétences que j'utilise réellement. Je gère un marché de plug-ins Claude Code [yux-claude-hub](https://github.com/wuyuxiangX/yux-claude-hub), dans lequel la compétence `yux-video-summary` est utilisée pour convertir les sous-titres vidéo en résumés structurés - prenant en charge la détection des langues chinoise et anglaise, les deux modes de sortie DUAL\_FILE/SINGLE\_FILE, le nettoyage des mots de remplissage, etc. Le SKILL.md de la compétence ressemble à ceci :
```yaml
---
name: yux-video-summary
description: Transform a video transcript file into a structured,
organized summary with key points, timeline, and cleaned transcript.
Use when the user has a transcript file and wants it summarized.
allowed-tools: Read, Write, Glob, Grep
---
```
La compétence a été écrite, mais comment savoir si elle est vraiment utile ? \*\* C'est là que Skill-Creator entre en scène.
> Il y a un principe d'écriture important dans le code source de Skill-Creator : *"Essayez d'expliquer le **pourquoi** derrière tout. Si vous vous retrouvez à écrire TOUJOURS ou JAMAIS en majuscules, c'est un drapeau jaune - recadrez et expliquez le raisonnement."* Signification : Une bonne compétence doit **expliquer pourquoi**, plutôt que d'empiler des règles rigides.
### Étape 1 : Créez des cas de test et exécutez l'évaluation
Question centrale : \*\*Cette compétence est-elle vraiment meilleure que pas de compétence du tout ? \*\*
Ouvrez Claude Code et saisissez directement :
```
Use the skill creator to create evals for the yux-video-summary skill
```
Skill-Creator lira d'abord les définitions et les schémas des compétences, puis générera automatiquement des cas de test et des assertions quantitatives. Mon exécution a généré 3 cas de test et 39 assertions :
Notez qu'il ne compile pas les cas de test avec désinvolture : il comprend les deux modes de sortie DUAL\_FILE et SINGLE\_FILE définis dans la compétence, et conçoit spécifiquement des scénarios qui couvrent différents types de vidéos (tutoriels, interviews de podcast, partage de technologie) et combinaisons de langues. La conception d'Assertions est également très particulière, de la détection de la langue à la sélection du mode de sortie en passant par la qualité du contenu et le nettoyage des mots de remplissage chinois et anglais, elle est bien plus complète que je souhaite tester les dimensions moi-même.
Ensuite, le système démarre simultanément deux sous-agents indépendants pour chaque scénario de test : with\_skill (chargement des compétences) et **without\_skill** (ligne de base, aucune compétence n'est chargée). **6 agents parallèles** (3 scénarios de test × 2 versions) ont été démarrés en même temps, chacun s'exécutant dans un **arbre de travail indépendant** sans interférer les uns avec les autres.
> La compétence PDF d'Anthropic rencontrait auparavant des problèmes pour gérer les formulaires non remplissables : Claude devait placer le texte à des coordonnées précises sans définir de champs. Le point de défaillance a été isolé grâce à Eval, et l'équipe a ensuite corrigé la logique de positionnement. C'est la valeur d'Eval : transformer « quelque chose ne va pas » en « ce qui ne va pas exactement ici ».
### Étape 2 : notation du relais de trois sous-agents
Une fois toutes les opérations terminées, les trois sous-agents professionnels apparaissent **automatiquement** dans l'ordre :
**Grader** Vérifie les assertions une par une. Il vérifiera si le résumé de la version with\_skill contient la table de présentation, si le mode DUAL\_FILE est correctement sélectionné, si le mot de remplissage a été nettoyé, puis enregistrera la réussite/l'échec et la preuve de chaque élément, générant `grading.json` :
```json
{
"expectations": [
{ "text": "摘要包含 Overview 表格", "passed": true, "evidence": "Found overview table with Type, Duration, Language fields" },
{ "text": "正确选择 DUAL_FILE 模式", "passed": true, "evidence": "Generated separate summary and transcript files" },
{ "text": "filler 词已清理", "passed": false, "evidence": "Found 'you know' in transcript line 42" }
],
"summary": { "passed": 2, "failed": 1, "total": 3, "pass_rate": 0.67 }
}
```
**Comparator** effectue une comparaison A/B aveugle : il reçoit deux résumés, mais **ne sait pas quelle est la version de compétence et quelle est la version de base**. Il ne voit que le « Sortie A » et la « Sortie B » et les juge indépendamment en fonction de ses propres normes de qualité pour déterminer le gagnant.
**Analyzer** combine les résultats ci-dessus pour établir un diagnostic : quelles assertions ont réussi, quelles que soient les compétences ou non (indiquant que cette assertion n'a aucune différenciation et doit être remplacée par une meilleure assertion), quels résultats ont une variance élevée (le test est instable) et quel est le compromis entre le temps et le jeton. Enfin, des suggestions d'améliorations sont données.
### Étape 3 : Examiner les résultats dans Eval Viewer
Une fois la notation terminée, Skill-Creator ouvrira automatiquement une visionneuse HTML dans votre navigateur.
**Onglet Sorties** Vous pouvez afficher la sortie de chaque scénario de test un par un. Il y a une zone de texte de commentaires en bas – notez ce que vous pensez n'est pas assez bon, comme « le résumé n'a pas de chronologie » et « le mot de remplissage n'est pas nettoyé ». Après avoir lu tous les cas d'utilisation, cliquez sur **Soumettre tous les avis** et les commentaires seront enregistrés dans `feedback.json`.
**Onglet Résultats du benchmark** Vous pouvez voir la comparaison quantitative : le taux de réussite, la consommation de temps, la consommation de jetons de with\_skill et without\_skill, ainsi que la comparaison élément par élément de chaque assertion.
### Étape 4 : Itérer et améliorer jusqu'à ce que vous soyez satisfait
Retournez voir Claude Code et dites-lui que vous avez fini de donner votre avis. Skill-Creator lira `feedback.json` et donnera des suggestions d'analyse et d'amélioration basées sur les données de référence :
Ma compétence a bien fonctionné avec un taux de réussite de 97 %. Skill-Creator a identifié avec précision un petit problème : la vidéo de l'interview manquait de paragraphes de citations notables et a fait des suggestions pour le réparer.
La clé est qu'il ne corrige pas les cas de test individuels - il généralise vos commentaires, comprend les exigences qui les sous-tendent et ajuste la structure globale de la compétence, puis réécrit SKILL.md, réexécute tous les tests dans le répertoire `iteration-2/` et ouvre une nouvelle visionneuse d'évaluation afin que vous puissiez comparer le résultat des deux tours. Ce cycle continue jusqu'à ce que vous soyez satisfait.
> Une philosophie d'amélioration remarquable dans le code source de Skill-Creator : *"Nous essayons de créer des compétences qui peuvent être utilisées un million de fois dans de nombreuses invites différentes. Plutôt que d'introduire des changements fastidieux et excessifs, ou des MUST trop contraignants, s'il y a un problème persistant, essayez de créer des branches et d'utiliser différentes métaphores."* Idée de base : **Évitez le surajustement** pour tester des cas et poursuivez les capacités de généralisation.
### Étape 5 (facultatif) : Optimisez la description pour que la compétence se déclenche au bon moment
La qualité de la compétence est vérifiée, mais il y a un autre problème qui passe facilement inaperçu : le champ `description` de la compétence détermine quand Claude l'appellera.
Entrée :
```
Use the skill creator to optimize the description for yux-video-summary
```
Skill-Creator génère automatiquement environ 20 requêtes d'évaluation (la moitié doit être déclenchée, l'autre moitié ne doit pas être déclenchée), et l'interface de révision s'ouvre dans le navigateur :
Notez que ces requêtes sont disponibles en chinois et en anglais, couvrant une variété d'expressions réelles. Les requêtes "Ne devraient pas déclencher" ne devraient pas être trop scandaleuses - un bon contre-exemple est "Aidez-moi à résumer le procès-verbal de cette réunion", qui partage le mot-clé "résumé" avec le résumé vidéo, mais nécessite en réalité des compétences en traitement de documents plutôt qu'en résumé vidéo.
Vous pouvez modifier le texte de la requête directement sur la page, cliquer sur **+ Ajouter une requête** pour en ajouter de nouvelles, utiliser le bouton Supprimer pour supprimer celles inappropriées et vous pouvez également activer le commutateur Devrait déclencher pour chaque requête. Après avoir confirmé qu'il est correct, cliquez sur **Export Eval Set** pour exporter le fichier JSON. Retournez sur Claude Code et dites-lui que vous l'avez exporté. Le système exécutera automatiquement la boucle d'optimisation en arrière-plan :
L'ensemble du processus est entièrement automatisé : divisez la requête en un ensemble de formation et un ensemble de test 60/40, optimisez de manière itérative la description sur l'ensemble de formation (jusqu'à 5 tours) et utilisez les résultats de l'ensemble de test pour sélectionner la meilleure version afin d'éviter le surajustement. Après l'exécution, la comparaison des descriptions avant et après l'optimisation sera affichée :
La description optimisée devient plus spécifique : elle clarifie les types de fichiers pris en charge (.vtt/.srt), met l'accent sur les fonctionnalités du pipeline (nettoyage des remplissages, logique DUAL/SINGLE\_FILE) et utilise MUST USE pour exclure les scénarios qui ne doivent pas être déclenchés. Anthropic a utilisé en interne cet ensemble d'optimiseurs pour exécuter ses propres compétences en matière de création de documents. En conséquence, la précision de déclenchement de 5 compétences publiques sur 6 a été améliorée.
### Utilisation avancée : injection de contexte dynamique
Si vous souhaitez que la compétence injecte automatiquement du contexte lors du chargement, vous pouvez intégrer une commande shell dans SKILL.md en utilisant la syntaxe `!` de Skills 2.0 :
```markdown
## Project Context
File tree: !`find . -type f -not -path '*/node_modules/*' | head -50`
Package info: !`cat package.json 2>/dev/null || echo "No package.json"`
Recent commits: !`git log --oneline -10`
```
Ces commandes sont exécutées avant que Claude ne voie la compétence, et les données sont intégrées directement dans l'invite. Par rapport au fait de laisser Claude explorer les fichiers un par un, cela permet d'économiser beaucoup de temps et de jetons.
## Deux types de compétences : laquelle créer ?
Avant d'utiliser Skill-Creator, il est nécessaire de comprendre les deux types de compétences définis par Anthropic :
**Type d'amélioration des capacités** - Laissez le modèle faire des choses qu'il ne peut pas ou ne peut pas bien faire auparavant. Par exemple :
* Compétences en génération d'images : Claude ne peut pas générer d'images de manière native, mais cela peut être réalisé en appelant des outils tels que nanobanner via des compétences.
* Compétences en conception frontale : les conceptions d'IA par défaut sont souvent très "à saveur d'IA", et de bonnes compétences en conception peuvent grandement améliorer la qualité.
**Préférence de codage** - consolidez votre flux de travail spécifique. Le modèle possède déjà des capacités individuelles, mais vous avez besoin d'un ordre d'exécution précis. Par exemple :
* Compétences en matière d'examen des relations publiques : vérifier la sécurité du code selon des procédures fixes et produire des rapports sur le niveau de risque
* Compétences en matière de résumé vidéo : sortie selon une structure de modèle spécifique, détection automatique de la langue et nettoyage des mots de remplissage
Les raisons pour lesquelles ces deux types de compétences doivent être testées sont différentes : le **type d'amélioration des capacités** peut devenir inutile à mesure que le modèle évolue - si la ligne de base (without\_skill) peut également réussir toutes les assertions, cela signifie que le modèle est suffisamment natif et que cette compétence peut être retirée ; Le **type de codage** est plus durable, mais vous devez vérifier s'il est vraiment fidèle à votre workflow.
Les capacités d'évaluation de Skill-Creator vous permettent de vérifier en permanence si une compétence est toujours utile, plutôt que d'utiliser aveuglément une compétence qui pourrait être obsolète.
## Ce que dit la communauté
La mise à jour de Skill-Creator a suscité de nombreuses discussions, de X/Twitter à Reddit en passant par les blogs indépendants, et les vrais retours sont plus précieux que la documentation officielle.
### Est-ce vraiment utile ? Les données parlent
La question la plus directe : est-il vraiment préférable d’ajouter des compétences que de ne pas en ajouter ? \*\* Plusieurs mesures réelles donnent une réponse claire.
Reddit u/hashpanak a effectué une évaluation sur la compétence de génération de titre et a obtenu un taux de réussite de 100 % avec\_skill et seulement 60 % sans\_skill. Lorsqu'on lui a demandé si le coût du jeton en valait la peine, il a répondu : "Absolument. Après optimisation, les tâches répétées peuvent être converties en scripts, ce qui permettra d'économiser des jetons." u/spences10 est encore plus extrême : il a effectué 250 évaluations sandbox et augmenté le taux d'activation des compétences de 84 % à 100 %. La section des commentaires u/Manfluencer10kultra a déclaré : "**Cela devrait devenir une pratique standard.**"
Le blogueur [Nathan Onn](https://www.nathanonn.com/claude-code-skill-creator-guide/) a évalué les compétences en matière de sécurité de WordPress : les 21 assertions ont été réussies (la ligne de base n'était que de 90,5 %) et la vitesse était 9,9 % plus rapide. Son résumé : **"Les compétences étaient autrefois de l'art, maintenant elles sont de l'ingénierie."**
@0zhuxiaofeng a donné des chiffres plus précis du point de vue du flux de travail réel : "Après l'avoir utilisé pendant un mois, le plus grand changement est que run\_eval permet aux compétences de se noter elles-mêmes. L'agent que j'exécute pour les opérations de contenu évalue désormais automatiquement l'effet après chaque version, et les compétences médiocres sont directement éliminées et réécrites. **L'intervention manuelle a été réduite de 3 heures par jour à une demi-heure**."
### Angle mort négligé : déclencheur ≠ qualité
Le blogueur [Mager](https://www.mager.co/blog/2026-03-08-claude-code-eval-loop/) a souligné un point mort que personne n'a mentionné : **Les compétences peuvent réussir l'évaluation de la qualité mais échouer lors de l'évaluation du déclencheur** - la qualité du résultat est très bonne, mais elle ne sera jamais appelée. Après trois tours d'optimisation `run_loop.py`, il a déclenché l'évaluation vers 13/13. Aperçu principal : "La description d'une compétence n'est pas une métadonnée, mais un paramètre apprenable : vous devez optimiser le comportement de routage réel."
Cela coïncide avec la suggestion de @ DrWang5257 : "Ne réécrivez pas le tout d'un coup. Divisez-le d'abord en trois sections : les conditions de déclenchement, les modèles d'entrée et le repli en cas d'échec, et répétez étape par étape. De cette façon, la vitesse de mise à jour est rapide et le taux de roulement est faible. "
### De vrais problèmes
Bien que l'effet soit bon, il existe également de nombreux pièges :
* **La consommation de jetons est énorme**. @konghao10 a dit sans ambages que "la consommation de jetons est énorme" - exécuter 6 agents parallèles en même temps n'est vraiment pas bon marché. Reddit u/munkymead a également déclaré que "faire un test sérieux coûte cher".
* **Si vous avez trop de compétences, vous vous battrez**. \[Le blogueur RoboRhythms Noah Albert] ([https://www.roborhythms.com/best-claude-code-skills-2026/](https://www.roborhythms.com/best-claude-code-skills-2026/)) a découvert que **commence à avoir des problèmes lorsque les compétences atteignent 8 à 10** : Claude remettra en question le résultat, générera des préfaces plus verbeuses et aura parfois des conflits de commandes entre les compétences. Cependant, Reddit u/Specialist\_Solid523 a répliqué : "Les compétences mal écrites ne mangent que le contexte. **Les compétences bien écrites rendent presque toujours l'utilisation de vos jetons plus efficace.**"
* **SKILL.md s'allonge avec plus d'itérations**. Reddit u/IulianHI a souligné une contradiction : avec des améliorations itératives, les fichiers de compétences continuent de s'étendre, \*\* mais évincent la fenêtre contextuelle pour réellement faire les choses \*\*. Les cas de test qui ne couvrent que le chemin heureux manquent les 5 % critiques.
* **La gestion des versions est manquante**. @fengqve se plaint "Pourquoi Skill \*\*n'a-t-il **pas le concept de version** ? Il a été mis à jour tellement de fois qu'il est difficile de décrire de quelle mise à jour il s'agit." Ceci est particulièrement douloureux après plusieurs séries d’itérations.
* **Le mode sans tête a un bug**. Il y a un problème clé sur GitHub : la compétence n'est jamais déclenchée en mode `claude -p`, ce qui fait que le rappel décrivant la boucle d'optimisation est toujours à 0 % ([#36570](https://github.com/anthropics/claude-code/issues/36570)).
### Penser plus loin : l'auto-amélioration récursive
@vista8 a partagé un article connexe \[Memento-Skills: Let Agents Design Agents] ([https://github.com/Memento-Teams/Memento-Skills](https://github.com/Memento-Teams/Memento-Skills)), et quelqu'un dans la zone de commentaires l'a résumé avec précision : « Le principal goulot d'étranglement de Skill est l'itération - il est facile d'écrire la première version, mais il est difficile de l'améliorer et de mieux l'utiliser dans des scénarios réels. Si vous pouvez automatiser ce cycle « utiliser → évaluer → améliorer », cela équivaut à installer un moteur d'auto-évolution pour Agent.
Un fil de discussion de type 104 sur Reddit r/ClaudeAI discute également de cette direction. Mais le commentaire principal a jeté de l'eau froide dessus - u/Tatrions a déclaré : "La boucle récursive fonctionne, mais le plus difficile est de savoir quand faire confiance aux améliorations. Nous avons constaté que nous devons faire du contrôle des preuves - ne validez pas les modifications à moins qu'un échec ne se produise au moins deux fois. Sinon, chaque cycle "répare" quelque chose qui n'est pas cassé en premier lieu, et cela finit par être pire. "
## Installation et Ecologie
Skill-Creator, en tant que l'une des compétences officiellement maintenues par Anthropic, est incluse dans l'entrepôt [anthropics/skills](https://github.com/anthropics/skills), qui contient plus de 17 compétences de niveau production.
L'écosystème de compétences au sens large connaît également une croissance rapide : [skills.sh](https://skills.sh) Le marché offre une expérience de découverte et d'installation pratique, et la communauté a maintenu plus de 1 234 compétences d'agent.
## Écrivez à la fin
Le problème principal résolu par Skill-Creator est le suivant : \*\*Comment savez-vous que vos compétences sont vraiment efficaces ? \*\*
En l’absence de cela, le développement des compétences repose sur « écrire → essayer → se sentir bien ». Avec Skill-Creator vous pouvez :
* Testez les effets qualifiés et non qualifiés avec **Parallel Agent**
* Éliminez les biais d'évaluation grâce à la **Comparaison A/B aveugle**
* Visualisez les résultats et laissez des commentaires avec **Eval Viewer**
* Utilisez **Description Optimizer** pour contrôler avec précision le moment de déclenchement des compétences
* Utilisez des **boucles itératives** pour vous améliorer continuellement jusqu'à ce que vous soyez satisfait
Ceci est conforme au concept de développement piloté par les tests en génie logiciel : il ne s'agit pas simplement de "simplement écrire le code et de penser qu'il peut s'exécuter", mais "d'utiliser des tests pour prouver qu'il fonctionne réellement comme prévu".
Anthropic a avancé une perspective intéressante dans le blog officiel : à mesure que les capacités du modèle s'améliorent, SKILL.md peut évoluer du "plan de mise en œuvre" (indiquant à Claude **comment**) à une "description de la spécification" (indiquant à Claude **quoi** et laissant le modèle le comprendre par lui-même). Le framework Eval est le premier pas dans cette direction – Eval décrit « quoi faire ». Si un jour cette description suffit à elle-même pour devenir une compétence, alors le système de tests établi par Skill-Creator deviendra encore plus important.
Si vous utilisez déjà Skills, essayez d'utiliser `/skill-creator` pour évaluer vos compétences les plus utilisées. Vous pourriez être surpris de constater que certaines compétences ne valent pas mieux que l’absence de compétences du tout, et c’est là que commence l’optimisation.
Lecture connexe :
* [Que sont les Compétences Claude](/fr/docs/notes/claude-skills/concept) — Comprendre les principes fondamentaux des Compétences
* [Guide pratique](/fr/docs/notes/claude-skills/practice) — Créez votre première compétence
# Introduction au concept
## Introduction
Dans l'[analyse approfondie de Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept), nous avons decouvert un probleme fondamental : le **Context Rot** — a mesure que la conversation s'allonge, la fenetre de contexte de Claude se remplit de code echoue, de discussions obsoletes et d'informations non pertinentes, ce qui degrade continuellement la qualite des resultats.
La solution de Ralph est de "tout redemarrer" : une boucle infinie en bash lance a chaque fois une nouvelle instance de Claude, en transmettant l'etat via le systeme de fichiers. Simple, efficace, mais avec des limites evidentes — ce n'est qu'une methodologie, sans comprehension du projet, sans planification par phases, sans verification de la qualite. Vous devez rediger les specs vous-meme, orchestrer les taches vous-meme, et juger vous-meme si "c'est termine ou non".
Comme Chase AI l'a resume avec precision dans sa video : **Ralph Loop est une arme extremement puissante, mais la plupart des gens n'ont pas besoin d'une arme — ils ont besoin d'un arsenal complet.** La boucle Ralph depend entierement de la preparation en amont : votre PRD est-il suffisamment bon ? Les definitions fonctionnelles sont-elles assez precises ? Savez-vous a quoi ressemble "termine" ? Si les reponses a ces questions ne sont pas precises, peu importe le nombre d'iterations de la boucle, ce sera du garbage in, garbage out.
Et si vous vouliez un systeme qui **ne se contente pas de faire tourner Claude en boucle, mais qui comprend reellement votre projet et livre du code de maniere fiable** ?
C'est exactement ce que **GSD (Get Shit Done)** se propose de faire.
## Qu'est-ce que GSD
Le createur de GSD est **TÂCHES** (GitHub : glittercowboy), un developpeur independant. Sa motivation est directe :
> "Je ne suis pas une entreprise de 50 personnes. Je ne veux pas jouer au theatre d'entreprise. Je suis juste un creatif qui veut faire de bonnes choses."
Lors de son stream en direct, TÂCHES a demontre un fait saisissant : il **n'ecrit jamais de code a la main**. Il a utilise GSD pour construire en 4 heures, a partir de zero, une application native macOS complete de generation musicale (Sample Digger), entierement sans code ecrit manuellement. Il ne se definit pas comme un programmeur, mais comme un "chef de projet de haut niveau" — decrire la vision, prendre les decisions cles, valider les resultats. GSD rend ce mode de travail possible.
> "This has like 100x'd my ability to make cool shit with Claude Code because it's just created this systematization."
>
> — TÂCHES
D'autres outils de developpement pilote par les specifications — BMAD, SpecKit — ont chacun leur valeur, mais ils tendent a introduire des workflows d'entreprise complexes : ceremonies de sprint, story points, synchronisations avec les parties prenantes. Pour les developpeurs independants ou les petites equipes, ces processus sont eux-memes un fardeau. Comme Chase AI l'a souligne : "It's not enterprise theater. We understand that you're just one person, you just want some sort of scaffolding around Claude Code to make sure it executes the tasks it says it's going to execute in an effective way."
La philosophie de conception de GSD est de **cacher la complexite dans le systeme**. L'utilisateur n'a besoin que de quelques commandes simples, tandis que le systeme gere en arriere-plan toute la gestion du contexte, l'orchestration des taches et la verification de la qualite. En un mois apres sa publication, le projet a obtenu pres de 3 000 stars sur GitHub et 14 000 installations npm, TÂCHES le mettant a jour presque 15 a 20 fois par jour.
### La place de GSD dans l'ecosysteme des outils
| Dimension | Ralph Wiggum | SpecKit | BMAD | **GSD** |
| ----------------------------- | --------------------------------- | ------------------------------------- | ----------------------------------- | ---------------------------------------------- |
| Positionnement | Technique d'execution (bash loop) | Boite a outils de generation de specs | Framework entreprise | **Ingenierie de contexte + spec-driven** |
| Capacite de planification | Aucune (spec a fournir soi-meme) | Forte (spec→plan→taches) | Forte (processus agile complet) | **Forte (recherche→discussion→planification)** |
| Autonomie d'execution | La plus elevee (mode AFK) | Declenchement manuel a chaque etape | Declenchement manuel a chaque etape | **Declenchement manuel a chaque etape** |
| Mode de participation humaine | Human on the Loop | Human in the Loop | Human in the Loop | **Human in the Loop** |
| Gestion du Context Rot | Redemarrage nouvelle session | Aucune solution integree | Aucune solution integree | **Contexte frais via sous-agents** |
| Verification qualite | Depend des tests externes | Verification de build | Processus QA integre | **Verification auto + UAT** |
| Complexite utilisateur | La plus basse | Moyenne | Assez elevee | **Basse** |
| Complexite systeme | La plus basse | Moyenne | Assez elevee | **Elevee** |
Ce tableau revele un compromis cle : **Ralph obtient l'autonomie d'execution la plus elevee avec la complexite systeme la plus basse** — une fois lance, vous pouvez aller dormir ; tandis que **GSD echange une complexite systeme elevee contre la qualite de planification et la validation humaine** — vous avez l'opportunite d'intervenir a chaque phase. SpecKit et BMAD se situent entre les deux, offrant des capacites de planification mais sans l'ingenierie de contexte de GSD ni l'execution autonome de Ralph.
GSD et Ralph ne sont pas contradictoires. GSD herite des principes fondamentaux de Ralph — contexte frais, fichiers comme source de verite — mais construit dessus un systeme complet de comprehension et d'execution de projet. Si Ralph c'est "donner une tache a l'IA et la laisser reessayer en boucle", GSD c'est "comprendre ce que vous voulez, rechercher comment le faire, planifier les etapes, executer et verifier".
Le resume de Chase AI est tres juste : **La boucle Ralph suppose que vous arrivez avec un plan complet — GSD vous aide a construire ce plan.** GSD prend votre idee a moitie formee, pose des questions approfondies, effectue des recherches pour vous, genere un PRD complet, le decompose en taches atomiques, puis livre le projet de bout en bout. Et lors de l'execution du code, il utilise precisement les principes fondamentaux qui rendent la boucle Ralph puissante : le contexte frais des sous-agents, et des taches aussi petites et precises que possible.
## Workflow principal
Le workflow de GSD est un cycle **discussion → planification → execution → verification**, avec des entrees et sorties clairement definies a chaque phase.
### 1. Initialisation du projet
```text
/gsd:new-project
```
Une seule commande lance l'ensemble du processus. Le systeme va :
1. **Questionner** — Poser des questions de maniere continue jusqu'a comprendre pleinement votre idee (objectifs, contraintes, preferences techniques, cas limites)
2. **Rechercher** — Envoyer des agents paralleles pour etudier les domaines pertinents (optionnel mais recommande)
3. **Extraire les exigences** — Distinguer le contenu v1, v2 et hors perimetre
4. **Feuille de route** — Creer une planification par phases correspondant aux exigences
Vous approuvez la feuille de route, puis la construction commence. L'experience de TÂCHES montre que : plus la description initiale est detaillee, moins le systeme pose de questions ; plus elle est vague, plus il en pose. Il recommande de preparer un document de vision sommaire avant de demarrer — pas besoin de connaitre la stack technique ou les details d'implementation, il suffit de decrire ce que vous voulez.
**Fichiers produits** : `PROJECT.md`, `REQUIREMENTS.md`, `ROADMAP.md`, `STATE.md`
> Vous avez deja une base de code ? Executez d'abord `/gsd:map-codebase`, le systeme enverra des agents paralleles pour analyser votre stack technique, architecture, conventions et problemes potentiels. Ensuite, `/gsd:new-project` pourra planifier en se basant sur la base de code existante.
### 2. Phase de discussion
```text
/gsd:discuss-phase 1
```
Chaque phase de la feuille de route n'a qu'une ou deux phrases de description, ce qui est insuffisant pour construire ce que vous voulez. La phase de discussion sert a **capturer vos preferences d'implementation** avant la recherche et la planification.
Le systeme analyse la phase courante et identifie les "zones grises" — ces points de decision ou plusieurs implementations sont raisonnables :
* Fonctionnalites visuelles → mise en page, interactions, gestion des etats vides
* API/CLI → format de reponse, gestion des erreurs, niveau de detail
* Systeme de contenu → structure, ton, profondeur, flux
Chaque decision que vous prenez ici affecte directement la qualite de la recherche et de la planification ulterieures. Vous pouvez sauter cette etape (le systeme utilisera des valeurs par defaut raisonnables), mais une discussion approfondie permet au systeme de construire quelque chose de plus conforme a vos attentes.
**Fichier produit** : `{phase}-CONTEXT.md`
### 3. Phase de planification
```text
/gsd:plan-phase 1
```
Le systeme va :
1. **Rechercher** — Etudier comment implementer la phase courante, guide par les decisions de la phase de discussion
2. **Planifier** — Creer 2 a 3 plans de taches atomiques, en utilisant un format structure XML
3. **Verifier** — Controler que le plan satisfait les exigences, iterer et corriger jusqu'a validation
Un principe de conception important est le **Goal-Backward Planning** (planification par retro-objectif). Au lieu de partir de "que devrions-nous construire", on demande "pour atteindre l'objectif, quelles conditions doivent etre remplies ?" — puis on deduit le plan et les taches en sens inverse. TÂCHES affirme que cette approche "a considerablement ameliore la qualite des resultats", car chaque tache comprend sa relation avec les autres, au lieu d'etre simplement un element d'une liste a faire.
Chaque plan est suffisamment petit pour etre execute dans une fenetre de contexte entierement nouvelle. C'est la cle — **il n'y aura pas de degradation de la qualite**.
**Fichiers produits** : `{phase}-RESEARCH.md`, `{phase}-{N}-PLAN.md`
### 4. Phase d'execution
```text
/gsd:execute-phase 1
```
Le systeme va :
1. **Execution par vagues** — Les taches independantes sont executees en parallele, celles avec des dependances en sequence
2. **Contexte frais** — Chaque plan est execute dans un nouveau contexte de 200k tokens, zero accumulation de dechets
3. **Commits atomiques** — Chaque tache fait l'objet d'un git commit independant
4. **Verification des objectifs** — Verifier que la base de code implemente bien les fonctionnalites promises par la phase
Lors du stream en direct de TÂCHES, il a complete le developpement de 3 phases entieres, **la fenetre de contexte principale restant constamment a 24%**. Le sous-agent GSD Executor n'a besoin de charger que moins de 1 000 lignes de contexte pour achever une phase complete — vous pouvez executer 10 plans consecutifs, le contexte restant en dessous de 50%. C'est une experience completement differente du travail direct dans Claude Code : vous ne "jouez plus a la roulette russe en pariant sur le moment ou vous allez heurter le mur de la fenetre de contexte".
**Fichiers produits** : `{phase}-{N}-SUMMARY.md`, `{phase}-VERIFICATION.md`
### 5. Phase de verification
```text
/gsd:verify-work 1
```
La verification automatisee peut controler si le code existe et si les tests passent. Mais la fonctionnalite **fonctionne-t-elle comme vous l'attendez** ? Cela necessite votre confirmation.
Le systeme va :
1. **Extraire les livrables testables** — Lister les choses que vous devriez pouvoir faire maintenant
2. **Guider la verification une par une** — "Pouvez-vous vous connecter par email ?" Oui/Non, ou decrire le probleme
3. **Diagnostiquer automatiquement les echecs** — Envoyer un agent de debogage pour trouver la cause racine
4. **Creer un plan de correction** — Un plan de correction directement executable
Si tout passe, on continue a la phase suivante. S'il y a des problemes, executez a nouveau `/gsd:execute-phase` pour lancer le plan de correction.
C'est la plus grande difference ideologique entre GSD et la boucle Ralph : **Ralph est hands-off — vous lancez et vous le laissez tourner ; GSD a une etape de validation humaine a la fin de chaque phase.** Chase AI souligne que la boucle Ralph est du type "partez a la conquete" — elle tourne toute seule sans se retourner ; GSD s'assure que vous pouvez intervenir et corriger le cap a chaque point critique, evitant que les erreurs s'accumulent couche apres couche sans supervision.
De plus, GSD fournit un processus de debogage dedie. Lorsque la verification trouve un probleme, `/gsd:debug` lance un **sous-agent de debogage isole**, avec son propre workflow hypothese-preuves-solution, creant une documentation de debogage independante qui trace l'ensemble du processus d'investigation, sans polluer le contexte principal.
**Fichier produit** : `{phase}-UAT.md`
### Repeter le cycle
```text
/gsd:discuss-phase 2 → /gsd:plan-phase 2 → /gsd:execute-phase 2 → /gsd:verify-work 2
...
/gsd:complete-milestone → /gsd:new-milestone
```
Chaque phase passe par le cycle complet **discussion → planification → execution → verification**. Le contexte reste frais, la qualite reste constante.
Une fois toutes les phases terminees, `/gsd:complete-milestone` archive le jalon et marque la version. Puis `/gsd:new-milestone` demarre la construction de la version suivante.
## Pourquoi cela fonctionne : principes techniques
La fiabilite de GSD n'est pas le fruit du hasard, elle repose sur quatre piliers techniques cles.
### Context Engineering
Claude Code est extremement puissant lorsqu'il recoit le bon contexte. La plupart des gens ne savent pas comment lui fournir le bon contexte. GSD gere cela pour vous.
| Fichier | Role |
| ----------------- | --------------------------------------------------------------------------------- |
| `PROJECT.md` | Vision du projet, toujours charge |
| `research/` | Connaissances ecosysteme (stack technique, fonctionnalites, architecture, pieges) |
| `REQUIREMENTS.md` | Exigences par version, avec tracabilite des phases |
| `ROADMAP.md` | Direction et progression |
| `STATE.md` | Decisions, blocages, position — memoire inter-sessions |
| `PLAN.md` | Taches atomiques + structure XML + etapes de verification |
| `SUMMARY.md` | Journal d'execution, commite dans l'historique |
Chaque fichier a une **limite de taille** basee sur le seuil de degradation de qualite de Claude. Rester sous ces limites garantit une qualite de sortie elevee et constante. La fenetre de contexte principale reste a 30-40%, le travail reel etant effectue dans le contexte frais de 200k des sous-agents.
Chase AI a une explication intuitive du context rot : **quelle que soit la taille de la fenetre de contexte — Sonnet, Opus, meme une fenetre d'un million de tokens — les tokens de la premiere moitie sont plus efficaces que ceux de la seconde moitie.** Ce n'est pas un bug, c'est une propriete inherente des LLM. Le mecanisme autocompact integre de Claude Code ne peut que partiellement attenuer le probleme. La solution de GSD est plus radicale : chaque tache atomique est executee dans un nouveau sous-agent, garantissant que chaque tache obtient le meilleur de Claude.
Les propres donnees de TÂCHES le confirment : sur le plan Max a $200/mois, il consomme environ $30 000 de tokens Opus par mois. Cela semble beaucoup, mais comme chaque tache est executee dans un contexte frais, les reprises sont rares, et l'efficacite reelle est bien superieure a celle d'un travail de rafistolage repete dans un contexte degrade.
### XML Prompt Formatting
Chaque plan est un XML structure optimise pour Claude :
```xml
Create login endpointsrc/app/api/auth/login/route.ts
Use jose for JWT (not jsonwebtoken - CommonJS issues).
Validate credentials against users table.
Return httpOnly cookie on success.
curl -X POST localhost:3000/api/auth/login returns 200 + Set-CookieValid credentials return cookie, invalid return 401
```
Des instructions precises, aucune devinette necessaire, la verification est integree dans chaque tache.
### Multi-Agent Orchestration
Chaque phase utilise le meme schema : un orchestrateur leger envoie des agents specialises, collecte les resultats et les achemine vers l'etape suivante.
| Phase | Ce que fait l'orchestrateur | Ce que font les agents |
| ------------- | ------------------------------------------------- | --------------------------------------------------------------------------------------- |
| Recherche | Coordonne, presente les decouvertes | 4 chercheurs paralleles etudient stack technique, fonctionnalites, architecture, pieges |
| Planification | Valide, gere les iterations | Le planificateur cree le plan, le verificateur valide, boucle jusqu'a validation |
| Execution | Regroupe en vagues, suit la progression | Les executants implementent en parallele, chacun avec un contexte frais de 200k |
| Verification | Presente les resultats, achemine l'etape suivante | Le verificateur controle la base de code, le debogueur diagnostique les echecs |
L'orchestrateur ne fait jamais le gros du travail. Il envoie des agents, attend, integre les resultats. Le resultat : vous pouvez executer une phase entiere — recherche approfondie, creation et validation de multiples plans, des milliers de lignes de code ecrites en parallele, verification automatique — **tandis que votre fenetre de contexte principale reste a 30-40%**.
### Atomic Git Commits
Chaque tache est commitee independamment des qu'elle est terminee :
```text
abc123f docs(08-02): complete user registration plan
def456g feat(08-02): add email confirmation flow
hij789k feat(08-02): implement password hashing
lmn012o feat(08-02): create registration endpoint
```
Avantages : `git bisect` peut localiser la tache exacte qui a echoue, chaque tache peut etre annulee independamment, l'historique clair aide Claude a comprendre l'evolution du code dans les sessions futures.
## Les limites de GSD
GSD est puissant, mais comprendre ce qu'il **ne peut pas faire** est tout aussi important.
### GSD est un workflow guide par l'humain, pas un agent autonome
GSD ne peut pas tourner de maniere persistante. Chaque frontiere de phase — de `discuss` a `plan` a `execute` a `verify` — necessite que vous saisissiez manuellement une commande. Vous ne pouvez pas dire "fais-moi une app" et aller dormir.
Cela contraste fortement avec le mode AFK de Ralph. Ralph est concu pour "lancer et aller dormir" — la boucle infinie en bash continue de tourner jusqu'a ce que la tache soit terminee ou echoue. GSD exige que vous soyez present a chaque point critique : approuver la feuille de route, repondre aux questions de discussion, declencher la planification, lancer l'execution, confirmer les resultats de verification.
Lors de son stream en direct de 4 heures, TÂCHES a saisi des commandes en continu : `new-project`, `discuss-phase 1`, `plan-phase 1`, `execute-phase 1`, `verify-work 1`, `discuss-phase 2`... Chaque transition necessite qu'il appuie sur Entree. Ce n'est pas accidentel — c'est un choix de conception delibere.
### Un compromis de conception delibere
Ralph sacrifie la capacite de planification au profit de l'autonomie d'execution ; GSD sacrifie l'autonomie d'execution au profit de la qualite de planification et de la validation humaine. **C'est un compromis de conception, pas un defaut.**
* **Avantage de Ralph** : Vous pouvez le laisser tourner pendant que vous dormez et il completera une fonctionnalite entiere. Mais si les specs ne sont pas assez bonnes, il foncera a toute allure dans la mauvaise direction.
* **Avantage de GSD** : Vous pouvez corriger le cap a la fin de chaque phase. Mais vous devez etre present tout au long du processus, sans pouvoir vous absenter.
Quel serait l'etat ideal ? Si la discussion, la planification, l'execution et la verification de GSD pouvaient etre enchaines en un cycle automatique — similaire a la boucle bash de Ralph, mais avec la planification structuree et la verification qualite de GSD — ce serait le meilleur des deux mondes. Mais un tel outil n'existe pas encore. C'est peut-etre la prochaine direction a explorer.
## Ressources video
Les videos suivantes peuvent vous aider a mieux comprendre de maniere visuelle l'utilisation et l'efficacite de GSD.
## Conclusion
GSD represente une direction dans l'evolution des outils de programmation IA : passer de "faire ecrire du code par l'IA" a "faire livrer des projets de maniere fiable par l'IA".
Ralph Wiggum a prouve une intuition cle — un contexte frais est plus precieux qu'un contexte accumule. GSD s'appuie sur cette base en ajoutant la comprehension du projet (new-project), la capture des decisions (discuss), la planification structuree (plan), l'execution parallele (execute) et la verification qualite (verify), formant ainsi un cycle complet en boucle fermee.
Pour les developpeurs independants et les petites equipes, la valeur de GSD reside dans le fait qu'il encapsule des pratiques d'ingenierie complexes en quelques commandes simples. Vous n'avez pas besoin de comprendre l'orchestration de sous-agents ou l'ingenierie de prompts XML — il vous suffit de decrire ce que vous voulez, puis de laisser le systeme s'en occuper.
Chase AI le dit bien : GSD est fait pour ceux qui "ne viennent pas d'un background technique, mais qui veulent tout de meme construire des projets de bout en bout dans Claude Code de maniere durable et reproductible". Et le stream en direct de TÂCHES le prouve — quelqu'un qui se decrit comme "capable tout au plus d'ecrire une page HTML Hello World tout seul" a construit une application de bureau native complete avec GSD.
Ce n'est pas de la magie. C'est **mettre la bonne complexite au bon endroit** — le systeme assume la complexite d'orchestration, l'humain se concentre sur la creativite et les decisions. Et ses limites meritent egalement d'etre respectees : GSD a choisi de garder l'humain toujours present, ce qui est a la fois sa limitation et la source de sa fiabilite.
Vous voulez passer a la pratique ? Poursuivez avec le [guide pratique GSD](/fr/docs/notes/gsd/practice) — couvrant la reference complete des commandes, les details de configuration, les demonstrations de workflows reels et les questions frequentes.
***
**Lectures connexes** :
* [Analyse approfondie de Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept) — Analyse complete du probleme du Context Rot et de la methodologie Ralph
* [Qu'est-ce que le developpement pilote par les specifications](/fr/docs/notes/speckit/concept) — De la programmation intuitive au developpement pilote par les specifications
* [Guide complet des sous-agents Claude](/fr/docs/notes/claude-subagent) — Une autre maniere de maintenir un contexte propre
* [Architecture complete du systeme Claude](/fr/docs/notes/claude-architecture) — Architecture globale des composants Hooks, Subagent, etc.
* [Mes meilleures pratiques Claude Code](/fr/blog/claude-code-best-practices) — Conseils d'utilisation quotidienne de Claude Code
# Guide pratique
## Introduction
Dans l'[article précédent](/fr/docs/notes/gsd/concept), nous avons exploré en profondeur les principes fondamentaux de GSD — l'ingénierie de contexte, l'orchestration de sous-agents, la planification par objectifs inversés et les commits atomiques. Ces concepts semblent élégants, mais entre « comprendre la théorie » et « mener un projet à bien », il reste de nombreux détails opérationnels à maîtriser.
Dans cet article, nous passons à la pratique. Vous apprendrez le système complet de commandes de GSD, les options de configuration, la structure des fichiers produits, ainsi que la manière de l'utiliser pour livrer une fonctionnalité complète de A à Z.
## Installation et configuration
### Installation
```bash
npx get-shit-done-cc@latest
```
L'installateur vous demandera de choisir :
1. **Environnement d'exécution** — Claude Code, OpenCode, Gemini CLI ou tous
2. **Portée** — Globale (tous les projets) ou locale (projet actuel)
Après l'installation, saisissez `/gsd:help` dans votre environnement d'exécution pour vérifier que l'installation a réussi.
### Recommandé : mode sans permissions
GSD est conçu pour une automatisation sans friction. Il est recommandé de lancer Claude Code de la manière suivante :
```bash
claude --dangerously-skip-permissions
```
Si vous préférez ne pas utiliser ce drapeau, vous pouvez configurer des permissions granulaires dans `.claude/settings.json`.
### Mise à jour
```text
/gsd:update
```
GSD est mis à jour très fréquemment (TÂCHES pousse environ 15 à 20 mises à jour par jour). Il est recommandé d'exécuter cette commande régulièrement pour rester à jour.
## Référence complète des commandes
Toutes les interactions avec GSD se font via des commandes préfixées par `/gsd:`. Voici la liste complète, classée par fonctionnalité.
### Commandes du flux de travail principal
Ces cinq commandes constituent la boucle principale de GSD. Elles s'utilisent dans l'ordre.
| Commande | Description |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:new-project` | Initialise un projet. Le système pose des questions jusqu'à comprendre votre vision, puis effectue des recherches, extrait les exigences et crée une feuille de route |
| `/gsd:discuss-phase [N]` | Discute des zones grises de la phase N. Capture vos préférences d'implémentation pour orienter la planification |
| `/gsd:plan-phase [N]` | Crée un plan de tâches atomiques pour la phase N. Comprend trois sous-étapes : recherche, planification, vérification |
| `/gsd:execute-phase ` | Exécute la phase N. Les sous-agents implémentent les tâches en parallèle, chaque tâche faisant l'objet d'un commit indépendant |
| `/gsd:verify-work [N]` | Vérifie les livrables de la phase N. Vous guide pour confirmer chaque élément et diagnostique automatiquement les problèmes |
> `[N]` indique un paramètre optionnel — en cas d'omission, le système détecte automatiquement la phase en cours. `` indique un paramètre obligatoire.
### Gestion des jalons
| Commande | Description |
| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:audit-milestone` | Audite la progression du jalon actuel — vérifie l'état de toutes les phases et identifie les éléments non terminés |
| `/gsd:complete-milestone` | Archive le jalon actuel, marque la version et prépare le passage au cycle suivant |
| `/gsd:new-milestone [name]` | Crée un nouveau jalon. Vous pouvez optionnellement fournir un nom ; le système planifiera la suite en se basant sur le travail déjà accompli |
### Gestion des phases
| Commande | Description |
| --------------------------------- | -------------------------------------------------------------------------------------------------------- |
| `/gsd:add-phase` | Ajoute une nouvelle phase à la fin de la feuille de route |
| `/gsd:insert-phase [N]` | Insère une phase urgente à la position indiquée ; les phases suivantes sont automatiquement renumérotées |
| `/gsd:remove-phase [N]` | Supprime la phase indiquée et supprime en cascade tous les fichiers produits associés |
| `/gsd:list-phase-assumptions [N]` | Liste toutes les hypothèses et dépendances de la phase indiquée, pour identifier les risques potentiels |
### Quick Mode et outils
| Commande | Description |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/gsd:quick [--full]` | Mode rapide — passe la recherche, la vérification du plan et la validation. Adapté aux petites tâches. `--full` active toutes les garanties |
| `/gsd:debug [desc]` | Lance un sous-agent de débogage isolé. Vous pouvez optionnellement décrire le problème ; le système émet des hypothèses → collecte des preuves → résout |
| `/gsd:add-todo [desc]` | Enregistre une idée dans la liste de tâches, sans modifier la feuille de route |
| `/gsd:check-todos` | Affiche la liste de tâches en cours |
| `/gsd:map-codebase` | Analyse la base de code existante — pile technique, architecture, conventions, problèmes potentiels |
### Gestion de session et configuration
| Commande | Description |
| ------------------ | ------------------------------------------------------------------------------------------------------------- |
| `/gsd:pause-work` | Met le travail en pause. Sauvegarde l'état actuel dans STATE.md, pour faciliter la reprise ultérieure |
| `/gsd:resume-work` | Reprend le travail. Lit l'état précédent depuis STATE.md et continue là où vous vous étiez arrêté |
| `/gsd:progress` | Affiche la progression globale du projet — nombre de phases terminées, position actuelle, éléments en attente |
| `/gsd:help` | Affiche toutes les commandes disponibles avec une brève description |
| `/gsd:settings` | Affiche et modifie la configuration de GSD |
| `/gsd:set-profile` | Change le profil de modèle (quality / balanced / budget) |
| `/gsd:update` | Met à jour GSD vers la dernière version |
## Configuration détaillée
### Profils de modèle
GSD propose trois profils de modèle, que vous pouvez changer via `/gsd:set-profile` :
| Profil | Planification | Exécution | Vérification | Cas d'usage |
| --------------------- | ------------- | --------- | ------------ | -------------------------------------------------------------------------- |
| quality | Opus | Opus | Sonnet | Projets complexes, fonctionnalités critiques, première utilisation |
| balanced (par défaut) | Opus | Sonnet | Sonnet | Développement quotidien, meilleur compromis pour la plupart des situations |
| budget | Sonnet | Sonnet | Haiku | Fonctionnalités simples, budget limité, itérations rapides |
### Paramètres principaux
Affichez et modifiez les paramètres suivants via `/gsd:settings` :
| Paramètre | Valeur par défaut | Description |
| ------------------------ | ----------------- | ----------------------------------------------------------------------------------------- |
| `mode` | `balanced` | Choix du profil de modèle |
| `depth` | `standard` | Profondeur de recherche : `quick` (rapide) / `standard` (standard) / `deep` (approfondie) |
| `git.branching_strategy` | `feature` | Stratégie de branches Git : `feature` (par fonctionnalité) / `phase` (par phase) / `none` |
### Interrupteurs de flux de travail
Les agents suivants peuvent être activés ou désactivés individuellement, pour ajuster le compromis entre vitesse et qualité :
| Interrupteur | Par défaut | Description |
| -------------- | ---------- | ----------------------------------------------------------- |
| `research` | Activé | Effectuer une recherche automatique avant la planification |
| `plan_check` | Activé | Vérifier automatiquement le plan après sa création |
| `verifier` | Activé | Vérifier automatiquement après l'exécution |
| `auto_advance` | Désactivé | Passer automatiquement à la phase suivante après achèvement |
> Désactiver `research` et `plan_check` peut accélérer considérablement le processus, mais risque de réduire la qualité de la planification. Il est recommandé de ne les désactiver qu'après avoir acquis une bonne connaissance du projet.
## Structure des fichiers produits
Tous les états et fichiers produits par GSD sont enregistrés dans le répertoire `.planning/`. Comprendre cette structure vous aidera pour le débogage et les interventions manuelles.
### Fichiers au niveau du projet
| Fichier | Rôle | Moment de création |
| ----------------- | ----------------------------------------------------------- | ------------------------------------ |
| `PROJECT.md` | Vision et périmètre du projet | `new-project` |
| `REQUIREMENTS.md` | Document d'exigences versionné, avec traçabilité des phases | `new-project` |
| `ROADMAP.md` | Planification des phases et progression | `new-project` |
| `STATE.md` | État actuel — décisions, blocages, position | `new-project`, mis à jour en continu |
### Fichiers au niveau de la phase
Chaque phase produit les fichiers suivants (exemple avec la phase 1) :
| Fichier | Rôle | Moment de création |
| -------------------- | --------------------------------------------------------- | ------------------ |
| `01-CONTEXT.md` | Enregistrement des décisions prises lors de la discussion | `discuss-phase 1` |
| `01-RESEARCH.md` | Résultats de recherche et investigation technique | `plan-phase 1` |
| `01-01-PLAN.md` | Plan de la première tâche atomique | `plan-phase 1` |
| `01-02-PLAN.md` | Plan de la deuxième tâche atomique | `plan-phase 1` |
| `01-01-SUMMARY.md` | Compte-rendu d'exécution de la première tâche | `execute-phase 1` |
| `01-02-SUMMARY.md` | Compte-rendu d'exécution de la deuxième tâche | `execute-phase 1` |
| `01-VERIFICATION.md` | Résultats de la vérification automatique | `execute-phase 1` |
| `01-UAT.md` | Enregistrement des tests de recette utilisateur | `verify-work 1` |
### Exemple de structure de répertoire
```text
.planning/
├── PROJECT.md
├── REQUIREMENTS.md
├── ROADMAP.md
├── STATE.md
├── research/
│ ├── tech-stack.md
│ ├── features.md
│ ├── architecture.md
│ └── pitfalls.md
├── 01-CONTEXT.md
├── 01-RESEARCH.md
├── 01-01-PLAN.md
├── 01-02-PLAN.md
├── 01-01-SUMMARY.md
├── 01-02-SUMMARY.md
├── 01-VERIFICATION.md
├── 01-UAT.md
├── 02-CONTEXT.md
├── 02-RESEARCH.md
├── 02-01-PLAN.md
│ ...
└── todos.md
```
## Démonstration d'un flux de travail concret
L'exemple suivant illustre le processus complet, de l'initialisation à la livraison, en prenant comme cas « l'ajout d'une fonctionnalité de commentaires à un système de blog ».
### Étape 1 : Initialiser le projet
```text
/gsd:new-project
```
Le système commence à poser des questions successives :
```
> Que souhaitez-vous construire ?
"Je veux ajouter une fonctionnalité de commentaires à mon blog Next.js.
Avec support des commentaires anonymes et authentifiés,
rendu Markdown et panneau d'administration.
Pile technique : Prisma + PostgreSQL."
```
Plus votre description est détaillée, moins le système aura besoin de poser de questions complémentaires. TÂCHES recommande de préparer un document de vision sommaire décrivant ce que vous voulez — pas besoin de connaître les détails techniques.
Une fois terminé, le système produit quatre fichiers et vous demande d'approuver la feuille de route. Après approbation, la phase de construction commence.
> **Vous avez déjà une base de code ?** Exécutez d'abord `/gsd:map-codebase` : le système analysera votre architecture et vos conventions existantes, et `new-project` pourra ensuite planifier en s'appuyant sur le code déjà en place.
### Étape 2 : Discuter de la phase
```text
/gsd:discuss-phase 1
```
Le système identifie les zones grises et les aborde une par une :
```
> Imbrication des commentaires : supporter plusieurs niveaux ou seulement un niveau de réponse ?
> Commentaires anonymes : avec captcha ou soumission directe ?
> Panneau d'administration : opérations en lot ou modération individuelle ?
```
Chaque décision que vous prenez ici influence directement la qualité de la planification ultérieure. En cas de doute, vous pouvez laisser le système utiliser les valeurs par défaut — mais une discussion approfondie réduit considérablement les reprises lors de la phase d'exécution.
### Étape 3 : Planifier la phase
```text
/gsd:plan-phase 1
```
Le système va :
1. Rechercher comment implémenter un système de commentaires avec Prisma + PostgreSQL
2. Créer 2 à 3 plans de tâches atomiques (par exemple : modèle de données, routes API, composants front-end)
3. Vérifier automatiquement que le plan couvre toutes les exigences
Chaque plan est suffisamment petit pour être réalisé dans une fenêtre de contexte entièrement nouvelle.
### Étape 4 : Exécuter la phase
```text
/gsd:execute-phase 1
```
Le système lance l'exécution par vagues :
* **Vague 1** (sans dépendances) : schéma de base de données, modèles Prisma — exécution en parallèle
* **Vague 2** (dépend de la vague 1) : routes API, CRUD des commentaires — exécution en parallèle
* **Vague 3** (dépend de la vague 2) : composant front-end de commentaires — exécution indépendante
Chaque tâche s'exécute dans un contexte neuf de 200k tokens et fait l'objet d'un commit Git indépendant à la fin.
### Étape 5 : Vérifier les livrables
```text
/gsd:verify-work 1
```
Le système vous guide pour confirmer chaque élément :
```
> ✅ Les tables de base de données ont été créées
> ✅ Les routes API retournent les codes de statut corrects
> ❓ Voyez-vous la zone de saisie de commentaire sous l'article de blog ? [oui/non/décrire le problème]
> ❓ La page se met-elle à jour en temps réel après la soumission d'un commentaire ? [oui/non/décrire le problème]
```
En cas d'échec, le système diagnostique automatiquement le problème et crée un plan de correction. Il vous suffit de relancer `/gsd:execute-phase 1` pour appliquer les correctifs.
### Scénarios courants
**Insérer une phase urgente** : un changement de besoin nécessite l'insertion d'un nouveau travail avant la phase actuelle.
```text
/gsd:insert-phase 2
```
Les phases suivantes sont automatiquement renumérotées (l'ancienne phase 2 devient la phase 3, et ainsi de suite).
**Mettre en pause et reprendre** : vous devez interrompre le travail pour traiter autre chose.
```text
/gsd:pause-work # Sauvegarde l'état actuel
# ... traiter autre chose ...
/gsd:resume-work # Reprendre là où vous vous étiez arrêté
```
**Annuler un résultat insatisfaisant** :
```bash
git reset --hard HEAD~3 # Revenir à l'état précédant l'exécution
```
```text
/gsd:remove-phase 2 # Supprimer en cascade tous les fichiers produits de cette phase
```
TÂCHES a démontré cette opération plusieurs fois en direct — quand le résultat ne lui convient pas, il annule. Net et sans bavure.
## Flux de travail de débogage
Lorsque la vérification révèle un problème, ou que vous rencontrez un bug en cours de développement, GSD fournit un processus de débogage dédié.
```text
/gsd:debug La page ne se met pas à jour en temps réel après la soumission d'un commentaire
```
Le système lance un **sous-agent de débogage isolé** dont le processus est le suivant :
1. **Hypothèses** — Génère plusieurs hypothèses de causes racines à partir de la description du problème
2. **Collecte de preuves** — Vérifie chaque hypothèse en examinant le code, les journaux et les requêtes réseau
3. **Résolution** — Une fois la cause racine identifiée, crée un plan de correction
Caractéristiques clés :
* **Isolation du contexte** : l'agent de débogage dispose de sa propre fenêtre de contexte et ne pollue pas le contexte de développement principal
* **Traçabilité documentée** : crée un document de débogage indépendant retraçant l'ensemble de l'investigation
* **Plan de correction** : une fois le diagnostic terminé, produit un plan de correction prêt à être exécuté
Cette approche est bien plus efficace que de déboguer directement dans le contexte principal — les informations de débogage ne s'accumulent pas dans votre fenêtre principale.
## Retours d'expérience
En combinant les démonstrations en direct de TÂCHES et l'expérience d'utilisation de Chase AI, voici quelques recommandations pratiques.
### Ralentir pour aller plus vite
TÂCHES admet que lorsqu'il a commencé à utiliser GSD, il avait un état d'esprit « vite, vite, vite ». Mais il a découvert que **passer plus de temps en phase de recherche et de discussion réduit les reprises en phase d'exécution**. La nouvelle version de GSD a ajouté les étapes `research-project` et `define-requirements` précisément pour s'assurer de bien cadrer la direction avant de passer à l'action.
> "I was definitely when I first started doing things with GSD, I was very much like move fast, move fast, move fast. But I'm just finding that taking these extra steps... I think you're going to really really love to see the results."
>
> —— TÂCHES
### Purger le contexte entre les phases
L'habitude de TÂCHES est d'**exécuter `clear` entre chaque phase**, pour garder le contexte principal épuré. Il utilise le terminal Warp, chaque fenêtre en plein écran (Command+Shift+Enter), exécutant la phase en cours dans une fenêtre tout en préparant la phase suivante dans une autre.
### Le compromis du coût en tokens
L'approche par sous-agents de GSD consomme effectivement plus de tokens qu'une utilisation directe de Claude Code. Mais Chase AI avance un argument convaincant : **« plan twice, prompt once » (planifier deux fois, prompter une fois) est plus économique à long terme que « prompter une fois, puis corriger encore et encore ».** Car réussir du premier coup dans un contexte neuf est bien plus efficace que de corriger indéfiniment dans un contexte dégradé.
### Gérer les résultats insatisfaisants
Si vous n'êtes pas satisfait du résultat d'une phase, vous pouvez exécuter `git reset --hard` puis utiliser `/gsd:remove-phase` pour supprimer en cascade tous les fichiers produits de cette phase. TÂCHES a démontré cette opération en direct — un rendu visuel ne lui convenait pas, il est simplement revenu à l'état précédent satisfaisant. Net et sans bavure.
### Le système de tâches
`/gsd:add-todo` vous permet d'enregistrer des idées dans une liste de tâches à tout moment, sans modifier la feuille de route. Ces idées peuvent être reprises lors de `/gsd:discuss-milestone` comme données d'entrée pour le prochain jalon. La stratégie de TÂCHES est « d'abord les fonctionnalités, le peaufinage de l'interface au milestone 2 ».
## Questions fréquentes et bonnes pratiques
### Bonnes pratiques
**Fournissez une description initiale détaillée.** La qualité de `/gsd:new-project` dépend de la qualité de vos données d'entrée. Préparez un document de vision sommaire — décrivez les objectifs, les utilisateurs, les fonctionnalités clés et les contraintes connues. Plus la description est précise, moins le système posera de questions et plus la planification sera pertinente.
**Nettoyez le contexte entre les phases.** Après chaque phase, exécutez `clear` ou `/compact` pour garder la fenêtre de contexte principale épurée. L'habitude de TÂCHES est de maintenir le contexte principal à 30-40 %.
**Testez d'abord avec le Quick Mode.** Pour les petites fonctionnalités dont vous n'êtes pas certain, testez d'abord avec `/gsd:quick`. Si le résultat est concluant, intégrez-le dans la feuille de route officielle.
**Exécutez map-codebase sur un projet existant.** Avant d'utiliser GSD sur une base de code existante, exécutez `/gsd:map-codebase`. Le système analysera la pile technique, l'architecture et les conventions, et la planification ultérieure sera mieux adaptée au code existant.
### FAQ
**Q : Quels environnements d'exécution GSD prend-il en charge ?**
R : Claude Code, OpenCode et Gemini CLI. Vous pouvez en choisir un ou tous lors de l'installation.
**Q : Quelle est la différence entre le Quick Mode et le mode complet ?**
R : Le Quick Mode fournit les garanties de base de GSD (commits atomiques, suivi d'état), mais ignore les étapes de recherche, de vérification du plan et de validation. Il est adapté aux corrections de bugs, aux petites fonctionnalités et aux modifications de configuration qui ne nécessitent pas une planification complète.
**Q : Peut-on mettre en pause pendant l'exécution ?**
R : Oui. `/gsd:pause-work` sauvegarde l'état actuel dans STATE.md. Lors du prochain `/gsd:resume-work`, le système reprendra là où vous vous étiez arrêté.
**Q : Comment contrôler le coût en tokens ?**
R : Trois approches — (1) Passer au profil `budget` : `/gsd:set-profile budget` ; (2) Désactiver les agents `research` ou `plan_check` ; (3) Utiliser `/gsd:quick` pour les tâches simples.
**Q : Peut-on utiliser GSD avec Ralph ?**
R : Oui. GSD et Ralph répondent à des problématiques différentes — GSD gère la planification et l'exécution structurée, Ralph gère les boucles d'exécution autonomes. Vous pouvez utiliser `new-project` et `plan-phase` de GSD pour générer un plan complet, puis utiliser les boucles Ralph pour exécuter les phases qui ne nécessitent pas d'intervention humaine.
**Q : Comment gérer la collaboration en équipe ?**
R : Le répertoire `.planning/` peut être versionné dans Git. Plusieurs personnes peuvent exécuter différentes phases et fusionner les résultats via Git. Toutefois, il est recommandé d'éviter d'exécuter la même phase simultanément.
## Conclusion
La valeur fondamentale de GSD réside dans le fait de **cacher la complexité dans le système et de laisser la simplicité à l'utilisateur**. Vous n'avez besoin que de quelques commandes — `new-project`, `discuss-phase`, `plan-phase`, `execute-phase`, `verify-work` — et le système gère en coulisses toute la gestion de contexte, l'orchestration des sous-agents et la vérification de la qualité.
De l'installation à la livraison, GSD offre un chemin clair : décrire ce que vous voulez → discuter des détails d'implémentation → générer des plans atomiques → exécuter en parallèle → vérifier les livrables. À chaque étape, vous avez la possibilité d'intervenir, et chaque étape est documentée.
Ce n'est pas de la magie « un bouton et c'est fait ». C'est un système qui requiert votre participation tout en prenant en charge la majeure partie de la charge cognitive. Comme le dit TÂCHES : vous êtes le chef de projet, GSD est votre équipe d'exécution.
***
**Pour aller plus loin** :
* [Analyse approfondie de GSD](/fr/docs/notes/gsd/concept) — Principes fondamentaux, flux de travail et architecture technique
* [Analyse approfondie de Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept) — Context Rot et la méthodologie Ralph
* [Guide pratique de snarktank/ralph](/fr/docs/notes/ralph-wiggum/snarktank) — Installation de Ralph, rédaction de PRD et mise en pratique
* [Qu'est-ce que le développement piloté par les spécifications](/fr/docs/notes/speckit/concept) — Du Vibe Coding au développement piloté par les spécifications
* [Guide pratique de Speckit](/fr/docs/notes/speckit/practice) — Détail des commandes Speckit et cas d'usage complet
# gstack : Quand le PDG de YC met son expérience entrepreneuriale dans Claude Code
\##Présentation
Dans les notes précédentes, nous avons exploré diverses « solutions d'amélioration » dans l'écosystème Claude Code, depuis la boucle infinie de [Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept) jusqu'au développement axé sur les spécifications de [GSD](/fr/docs/notes/gsd/concept). Ils essaient tous de répondre à la même question : \*\*Comment faire passer la programmation de l'IA de « l'adaptation » à une « livraison fiable » ? \*\*
La réponse de Ralph est "tout redémarrer" - utilisez un nouveau processus à chaque fois pour éviter la pourriture du contexte. La réponse de GSD est « Spécification Driven » : garantir la qualité grâce à des cycles structurés de planification et de validation des phases. Mais que se passe-t-il si vous souhaitez non seulement un système d’exécution, mais une équipe d’ingénierie virtuelle complète ? Le PDG prend les décisions relatives aux produits, le responsable de l'ingénierie examine l'architecture, le concepteur contrôle l'expérience, le QA exécute de vrais tests de navigateur et l'ingénieur de publication gère le lancement... tout cela est joué par l'IA et est commandé par vous.
C'est l'idée centrale de gstack.
## Qu'est-ce que Gstack
**Garry Tan**, le créateur de gstack, possède une riche expérience technique et entrepreneuriale : il a commencé à écrire du code à l'âge de 14 ans, est diplômé de Stanford Computer Engineering, est le 10e employé de Palantir, a cofondé Posterous (acquis plus tard par Twitter) et est président-directeur général de Y Combinator depuis 2023.
Il a utilisé gstack pour publier plus de 600 000 lignes de code de production (35 % de tests) en 60 jours, soit en moyenne plus de 10 000 lignes par jour, tout en continuant à exécuter YC à plein temps. L'un des projets, garylist.org, a été lancé en 21 jours, avec 150 000 lignes de code et 35 % de couverture de tests. Selon ses propres mots, la qualité du code dépasse le précédent projet entrepreneurial pour lequel il a dépensé 5 millions de dollars, deux ans et 10 ingénieurs.
Depuis que le projet est devenu open source le 11 mars 2026, il est passé de la v0 à la v0.15.1.0 en 3 semaines et GitHub a reçu plus de 60 500 étoiles. Licence MIT, entièrement open source.
## La place de gstack dans l'écosystème des outils
| Dimensions | Code Claude natif | Ralph Wiggum | GSD | Kit de spécifications | Superpouvoirs | **gstack** |
| ---------------------- | ------------------------------------- | ------------------------------- | ----------------------------------------------------- | ------------------------------------- | --------------------------------- | ------------------------------------------------------ |
| Positionnement de base | Assistant de codage universel de l'IA | Itération de boucle infinie | Ingénierie contextuelle + axée sur les spécifications | Exigences → Spécifications → Tâches | Discipline des processus + TDD | **Équipe virtuelle basée sur les rôles** |
| Modèle de base | Programmation conversationnelle | Boucle Bash + Nouveau processus | Feuille de route par phases | Spécification → Plan → Tâches | Pipeline de développement strict | **Processus de sprint en sept étapes** |
| Implication humaine | Conversations en direct | Intervention (AFK) | Vérification par étape | Approbation des spécifications | Validation par étape | **Révision des rôles par étape** |
| Capacités uniques | Codage de base | Itération illimitée | Contexte Gestion de la pourriture | Suivi des exigences | TDD forcé | **Automatisation du navigateur + révision multi-rôle** |
| Convient aux scénarios | Tâches simples | Itération continue | Gestion de projets à grande échelle | Des projets aux exigences rigoureuses | Assurance qualité de l'ingénierie | **Développement de produits complet** |
Un modèle clé peut être vu dans le tableau : \*\*Ces outils ne se font pas concurrence, mais résolvent des problèmes de programmation d'IA dans différentes dimensions. \*\*
Superpowers utilise la **discipline des processus** pour garantir la qualité du code (TDD obligatoire, dialogue structuré, plan de mise en œuvre) ; GSD utilise l'**ingénierie de contexte** pour gérer des projets complexes (planification des phases, nouveau contexte du sous-agent, état du système de fichiers) ; gstack utilise la **décomposition des rôles** pour améliorer la qualité de la prise de décision (le point de vue du PDG examine les produits, les responsables de l'ingénierie examinent l'architecture, le contrôle qualité exécute de vrais navigateurs).
Pour faire simple, Superpowers est basé sur des garde-fous de processus, et gstack est basé sur la conception de rôles : le premier est adapté à la mise en œuvre de projets de 1 à N, et le second est adapté à la construction de produits de 0 à 1. \*\* Les deux sont des produits complémentaires plutôt que concurrents. \*\*
## Workflow de base : les sept étapes du Sprint
gstack organise l'ensemble du processus de développement en un cycle de **Réfléchir → Planifier → Construire → Révision → Test → Expédier → Réfléchir**, appelé "Le Sprint" - pas un Sprint agile, mais un rythme de développement de "les rôles apparaissent en séquence".
### 1. Réfléchissez – Clinique de produits
```text
/office-hours
```
C'est la compétence la plus distinctive de Gstack. L’inspiration vient directement des heures de bureau de YC : les entrepreneurs vont à la rencontre des partenaires de YC et se livrent à une introspection. L'IA vous posera **6 questions forçantes** :
1. Qui en a spécifiquement besoin ?
2. Et s’ils ne l’ont pas aujourd’hui ?
3. Pourquoi cette question est-elle urgente maintenant ?
4. Comment savez-vous que cela fonctionne ?
5. Que se passe-t-il si vous ne faites rien ?
6. Quelle est la plus petite version que vous puissiez publier ?
Le but n'est pas de vous aider à écrire du code, mais de **réexaminer le problème lui-même** avant d'écrire du code.
### 2. Plan — Examen multi-rôles
```text
/plan-ceo-review # CEO 视角:寻找 10 星级产品
/plan-eng-review # 工程经理:锁定架构和边界
/plan-design-review # 设计师:评分 0-10,说明如何做到 10 分
/autoplan # 自动依次运行三个审查
```
L'examen du PDG est essentiellement un « mode fondateur » : au lieu d'exécuter les exigences littéralement, vous prenez du recul et demandez « Quel est le véritable objectif de ce produit ? Il prend en charge quatre modes : étendre la portée, développer sélectivement, maintenir la portée et réduire la portée.
### 3. Build — implémentation du codage
Commencez à coder selon le plan approuvé. Cette étape utilise les fonctionnalités standard de Claude Code.
### 4. Examen — Examen parallèle par des experts
```text
/review
```
Cette compétence envoie **7 sous-agents parallèles** en même temps pour examiner le code sous 7 perspectives : tests, maintenabilité, sécurité, performances, migration de données, contrat API et attaque de l'équipe rouge. Les problèmes évidents seront automatiquement résolus.
### 5. Test — Contrôle qualité du vrai navigateur
```text
/qa
```
Pas un test pratique. La compétence QA lance un **véritable navigateur Chromium sans tête**, ouvre votre application, clique sur les boutons, remplit les formulaires et prend des captures d'écran - tout comme le ferait un vrai testeur. Corrigez automatiquement les bogues, générez des tests de régression et revérifiez une fois les bogues découverts.
### 6. Expédier – publication en un clic
```text
/ship
```
Synchronisez automatiquement la branche principale, exécutez des tests, examinez les différences, mettez à jour les numéros de version et CHANGELOG, validez, poussez, créez des PR. Si le projet ne dispose pas d'un cadre de test, il en créera même un en premier.
### 7. Réfléchir – réviser et apprendre
```text
/retro
```
Rapport hebdomadaire de style responsable de l'ingénierie : analysez l'historique des validations, le taux de tests et les tendances en matière de qualité du code. Soutenez l'analyse d'équipes composées de plusieurs personnes et suivez des indicateurs tels que le « nombre de jours de sortie consécutifs ».
## Pourquoi ça marche : principes techniques
### Browse Daemon : mettez les yeux sur l'IA
La contribution technique la plus unique de gstack est le Browse Daemon - une instance Chromium persistante sans tête qui communique via HTTP localhost. Le premier appel lance le navigateur (\~ 3 secondes) et chaque commande suivante ne prend que 100 à 200 ms. Cela signifie que l'IA peut réellement voir votre application, plutôt que de deviner la structure du DOM.
Il introduit également le **Ref System** (référence d'élément `@e1`, `@e2`) pour localiser les éléments via l'arborescence d'accessibilité sans écrire de sélecteurs CSS. Il s'agit d'une « contribution véritablement technique » généralement reconnue par la communauté (y compris les critiques).
### Répartition des rôles : pas un agent, mais une équipe
Ce que fait gstack, c'est désassembler tous les rôles en fichiers d'invite indépendants, permettant à Claude Code de basculer vers les perspectives de différents rôles à différentes étapes pour réviser le code. Il s’agit essentiellement d’une ingénierie d’invite raffinée.
L'idée principale est la suivante : \*\*La planification n'est pas égale à la révision, la révision n'est pas égale à la publication, et le goût du fondateur et la rigueur de l'ingénierie sont des modes de pensée complètement différents. \*\* Au lieu de laisser un agent général faire tout, changez de « mode cérébral » si nécessaire : réflexion du fondateur, rigueur technique, révision paranoïaque, exécution rapide.
### Trois grandes philosophies
ETHOS.md de gstack enregistre trois concepts fondamentaux :
1. **Boil the Lake** : lorsque l'IA ramène le coût marginal de l'exhaustivité à zéro, choisissez toujours une implémentation complète : couverture de test à 100 %, tous les cas extrêmes, tous les chemins d'erreur. Les « raccourcis de version » sont une pensée ancienne.
2. **Rechercher avant de construire** : Trois niveaux de connaissances : modèles éprouvés, solutions nouvelles et populaires et premiers principes. Commencez par comprendre ce que chacun fait, remettez en question ses hypothèses et découvrez pourquoi les solutions habituelles sont fausses.
3. **Souveraineté de l'utilisateur** : recommandation d'IA, prise de décision humaine. Même si deux modèles d’IA parviennent à un consensus, le jugement de l’utilisateur prime toujours, car l’utilisateur possède une connaissance du domaine, une perspective stratégique et des goûts.
## Les limites et controverses de gstack
La réaction de la communauté à gstack est probablement l’outil de programmation d’IA le plus polarisant.
**Le bon côté** : les fondateurs et les constructeurs non techniques conviennent généralement que les compétences de « réflexion produit » telles que `/office-hours` et `/plan-ceo-review` ont aidé de nombreux développeurs indépendants à réexaminer l'orientation du produit avant de commencer à coder. La revue technique (`/review`) peut en effet découvrir certaines vulnérabilités de sécurité cachées. Ce modèle d’examen parallèle multi-angles a une valeur pratique.
Le **côté questionnement** est également très direct :
* **L'indicateur LOC est peu significatif** : 600 000 lignes de code en 60 jours. Le nombre de lignes de code n'est jamais un indicateur de qualité. Une grande quantité de code peut n’être qu’un échafaudage et un passe-partout.
* **Essentiellement un modèle d'invite** : chaque compétence est un fichier SKILL.md et le seuil technique n'est pas élevé. La vraie valeur ne réside pas dans le fichier lui-même, mais dans la qualité de la conception de l'invite.
* **Limitations du code d'auto-révision de l'IA** : `/review` Laisser l'IA réviser le code écrit par l'IA équivaut à corriger vos propres devoirs. Le parallélisme multirôle peut atténuer ce problème, mais il s’agit toujours du même modèle.
* **Bonus effet célébrité** : Si le fondateur n'est pas le PDG de YC, il y a de fortes chances que ce projet ne reçoive pas une telle attention.
**Mon avis** : Mis à part les controverses, les éléments vraiment précieux de gstack sont au nombre de deux : la technologie d'automatisation du navigateur de Browse Daemon et le modèle de conception de décomposition des rôles. Rien de tout cela ne dépend de qui est Garry Tan. L'importance fondamentale de la roleisation ne se situe pas au niveau technique, mais au niveau comportemental : elle vous aide à organiser votre flux de travail d'IA de manière plus consciente, plutôt que de tout confier à un agent général.
gstack convient au forking et à la personnalisation. Vous pouvez acquérir les compétences dont vous avez besoin et modifier les invites souhaitées, plutôt que de toutes les copier.
## Ressources vidéo
## Écrivez à la fin
gstack représente une direction intéressante pour les outils de programmation de l'IA : non pas rendre l'IA plus autonome (voie de Ralph), ni rendre le processus plus rigide (voie des Superpuissances), mais laisser l'IA jouer différents rôles pour améliorer la qualité des décisions. Sa controverse illustre simplement la richesse de l’écosystème de programmation de l’IA : aucune solution ne convient à tout le monde.
Si gstack vous intéresse, l'étape suivante consiste à lire le [Chapitre pratique](/fr/docs/notes/gstack/practice) - un tutoriel étape par étape depuis l'installation jusqu'à l'exécution du workflow complet.
***
**Lecture connexe** :
* [Introduction aux concepts GSD](/fr/docs/notes/gsd/concept) — Une autre solution de programmation IA structurée
* [Analyse approfondie de Ralph Wiggum](/fr/docs/notes/ralph-wiggum/concept) — Comprendre le point de départ de l'itération en boucle infinie
* [Claude Skills Concept](/fr/docs/notes/claude-skills/concept) — Comprendre le mécanisme sous-jacent des compétences
# Panorama des compétences frontales de Gstack : flux de travail de l'IA, de la conception au lancement
\##Présentation
Dans les notes précédentes, nous avons parlé de [Qu'est-ce que gstack](/fr/docs/notes/gstack/concept), [Comment exécuter le workflow](/fr/docs/notes/gstack/practice) et [La structure d'ingénierie des compétences](/fr/docs/notes/gstack/skill-architecture). Mais il y a une question qui n'a pas été abordée : parmi les plus de 60 compétences acquises après l'installation de gstack, lesquelles sont liées à la conception front-end/UI ? Dans quel ordre ? \*\*
Cette note fait deux choses : d'abord, classer les \~27 compétences liées au front-end par fonction, puis utiliser un petit projet intéressant - la page anniversaire du compte à rebours - pour parcourir le flux de travail complet du début à la fin, avec des captures d'écran à chaque étape, afin que vous puissiez voir l'effet réel.
## Panorama du kit de compétences front-end
Les compétences frontales de gstack peuvent être divisées en 6 couches fonctionnelles, de la fondation au toit, chaque couche résolvant des problèmes à différentes étapes.
### Infrastructure de conception
Un réglage unique au niveau du projet pour déterminer le langage de conception, et toutes les compétences ultérieures feront référence à ces références.
| Compétence | Que faire | Quand utiliser |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| `/design-consultation` | Consultation complète du système de conception, correspondance des couleurs de sortie, polices, espacement, direction de la texture | Au démarrage d'un nouveau projet, ou si vous souhaitez redéfinir le style visuel |
| `/teach-impeccable` | Collectez les préférences de conception en une seule fois et écrivez-les dans le fichier de configuration AI | Exécutez une fois après avoir installé Gstack pour permettre à l'IA de se souvenir de votre esthétique |
| `/brand-guidelines` | Appliquer la correspondance des couleurs de la marque et les spécifications de police existantes | Postulez directement lorsqu'il existe un manuel de marque existant |
> Si le projet a déjà `DESIGN.md`, ce niveau peut être ignoré.
### Exploration de la conception
Lorsque vous n’êtes pas sûr de la direction, comparez rapidement plusieurs options.
| Compétence | Que faire | Quand utiliser |
| ------------------ | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------ |
| `/design-shotgun` | Générez 3 à 5 solutions visuelles, ouvrez le panneau de comparaison | Vous ne savez pas quel style vous souhaitez, vous voulez voir les possibilités |
| `/frontend-design` | Générer un code d'interface frontale reconnaissable au niveau de la production | Travaillez directement après que la direction soit claire |
| `/canvas-design` | Générer des affiches, des arts visuels (PNG/PDF) | Nécessite une conception visuelle statique plutôt que des composants Web |
### Implémentation de la conception
Transformez le plan en code véritablement exécutable et gérez la composition, la mise en page et la réactivité.
| Compétence | Que faire | Quand utiliser |
| ------------------------ | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| `/design-html` | Convertir le brouillon de conception confirmé en HTML/CSS de qualité production | J'ai des maquettes que je souhaite implémenter directement |
| `/mobile-responsiveness` | Mise en page réactive et interaction tactile adaptées aux mobiles | Adaptation mobile à partir de zéro |
| `/adapt` | Adaptation des points d'arrêt selon les appareils et les tailles d'écran | Il existe une version de bureau et doit être adaptée aux téléphones mobiles/tablettes |
| `/typeset` | Sélection des polices, niveau, taille, épaisseur et optimisation de la lisibilité | La disposition du texte semble « presque dénuée de sens » |
| `/arrange` | Espacement de mise en page, rythme visuel, réparation d'alignement | Espacement incohérent, la mise en page semble encombrée ou dispersée |
### Améliorations de la conception
Sur la base de l'achèvement fonctionnel, injectez des effets dynamiques, des détails de personnalité et d'émotion.
| Compétence | Que faire | Quand utiliser |
| ------------ | ------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
| `/animate` | Ajoutez des micro-interactions et des animations ciblées | Fonctionnalité de la page OK mais semble « rigide » |
| `/delight` | Ajoutez des détails surprises et des touches personnalisées | Vous voulez que les utilisateurs se souviennent de cette page |
| `/bolder` | Amplifier l'impact visuel | Le design est trop simple et trop sûr |
| `/colorize` | Ajoutez une couleur stratégique à l'interface monotone | La page est trop grise, trop simple et manque de chaleur |
| `/overdrive` | Effets techniques de niveau explosion - shader, physique du ressort, animation du lecteur de défilement | Une certaine zone veut un effet wow |
| `/onboard` | Nouveau processus de guidage des utilisateurs, conception d'état vide | Expérience utilisateur pour la première fois |
Ces quatre compétences d'amélioration sont dans une **relation progressive** : `animate` est l'effet dynamique de base, `delight` est émotionnel, `bolder` est l'amplification et `overdrive` est l'explosion. Empilez-les étape par étape en fonction des besoins du projet, il n'est pas nécessaire de tous les utiliser.
### Optimisation de la conception
Convergence et raffinement – éliminer les excès, aligner les écarts et polir les aspérités.
| Compétence | Que faire | Quand utiliser |
| ------------ | ----------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| `/polish` | Polissage de qualité finale : alignement, espacement, cohérence | Un dernier passage avant la sortie |
| `/quieter` | Réduire l'intensité de la stimulation visuelle | Le design est trop chic et bruyant |
| `/distill` | Minimiser et supprimer la complexité inutile | Il y a trop d'éléments sur la page et je souhaite les réduire |
| `/normalize` | Aligner les normes du système de conception (jeton, espacement, couleur) | Le style s'écarte des spécifications DESIGN.md |
| `/clarify` | Améliorer la rédaction UX, les messages d'erreur et le libellé des étiquettes | La rédaction prête à confusion, les messages d'erreur ne sont pas conviviaux |
### Revue et vérification de la conception
Inspection systématique avant de se connecter, trouver les problèmes, les noter et les résoudre.
| Compétence | Que faire | Quand utiliser |
| --------------------- | --------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
| `/plan-design-review` | Examen du plan de conception avant la mise en œuvre (score de 0 à 10) | Je veux que l'IA examine le plan du point de vue d'un concepteur |
| `/design-review` | Contrôle qualité visuel après implémentation, comparaison et réparation automatiques des captures d'écran | Une fois le code écrit, vérifiez le degré de restauration visuelle |
| `/critique` | Évaluation UX : hiérarchie visuelle, charge cognitive, résonance émotionnelle | Vous souhaitez un rapport de révision de conception structuré |
| `/audit` | Bilan technique : accessibilité, performances, thèmes, réactivité | Vérifications systématiques avant mise en ligne |
| `/benchmark` | Tests de référence de performance, comparaison avant/après | Vous souhaitez quantifier l'impact des changements sur les performances |
## Démonstration pratique : utilisez la page anniversaire du compte à rebours pour parcourir l'ensemble du processus
Le simple fait de regarder le tableau de classification est trop abstrait. Nous utilisons un petit projet pour rassembler les compétences ci-dessus : créer une seule page de **compte à rebours/anniversaire** : choisissez une date significative et créez un affichage de compte à rebours avec une animation numérique et des effets d'arrière-plan.
Ce projet est petit mais complet, juste ce qu'il faut pour couvrir la plupart des 6 niveaux de compétences. Le processus complet est :
```
基建 → 探索 → 构建 → 增强 → 调优 → 审查 → 发布
```
> Vous n'êtes pas obligé d'exécuter les 7 étapes à chaque fois. Une fois que vous êtes devenu compétent, les liens couramment utilisés ne comportent que `/frontend-design → /animate → /polish → /ship` quatre étapes. Afin de montrer toute la capacité, chaque étape est franchie ici.
### Étape 1 : Infrastructure – Déterminer le langage de conception
**Skill**:`/design-consultation` + `/teach-impeccable`
Cela ne doit être fait qu’une seule fois au début du projet. Sorties `DESIGN.md`, qui permet à l'IA de mémoriser vos préférences de conception. Si le projet a déjà `DESIGN.md`, ignorez-le directement.
```text
> /design-consultation
"我要做一个倒计时纪念日页面,风格偏好是深色背景、
数字大而醒目、整体氛围温暖但克制"
```
{/* À FAIRE : Capture d'écran – Fragment DESIGN.md produit par design-consultation */}
### Étape 2 : Exploration - Comparaison de plusieurs options
**Skill**:`/design-shotgun`
En cas de doute sur la direction, laissez l'IA générer 3 à 5 solutions visuelles et ouvrez le panneau de comparaison pour choisir.
```text
> /design-shotgun
"做一个倒计时/纪念日单页,展示距离某个日期的天/时/分/秒,
要有数字翻牌动画,背景有微妙的粒子或渐变效果,
整体氛围要温暖但不俗气。给我 3 个方向。"
```
{/* À FAIRE : Capture d'écran – 3 panneaux de comparaison de solutions générés par design-shotgun */}
Choisissez une direction parmi 3 options. Si vous savez exactement ce que vous voulez, sautez cette étape et passez directement à l’étape 3.
### Étape 3 : Build – produire du code au niveau de la production
**Skill**:`/frontend-design` + `/adapt`
lien central. code tout en garantissant une réactivité dès le départ.
```text
> /frontend-design
"基于方案 2 实现倒计时页面。要求:
- 实时倒计时(天/时/分/秒)
- 响应式布局(桌面 + 手机)
- 日期可配置"
```
{/* À FAIRE : Capture d'écran – L'effet de la page de bureau une fois la construction terminée */}
{/* À FAIRE : Capture d'écran – Effet de version mobile (après adaptation /adapt) */}
### Étape 4 : Améliorer – Injecter du mouvement et de la personnalité
**Compétence** : `/animate` → `/delight` (sur demande `/overdrive`)
Ces trois éléments sont dans une relation progressive : `animate` est l'effet dynamique de base, `delight` est le détail émotionnel et `overdrive` est l'effet explosif. Ajoutez des calques si nécessaire.
```text
> /animate
"给倒计时数字加翻牌动画效果,页面加载时有入场动画"
> /delight
"给页面加一些有趣的小细节,比如到达整点时的微庆祝效果"
> /overdrive (可选,想要 wow effect 的话)
"给背景加粒子效果或流体动画,60fps"
```
Notez les contraintes d'animation référencées à `DESIGN.md` - si le système de conception n'autorise qu'une transition en survol de 150 ms, `/overdrive` ne s'appliquera pas. C'est un bon exercice de jugement.
{/* À FAIRE : Capture d'écran ou GIF – avant et après l'amélioration du mouvement */}
### Étape 5 : Réglage - Convergence et polissage
**Compétence** : `/typeset` + `/polish` (sur demande `/distill`, `/normalize`)
Alignement des espacements, hiérarchie des polices, rythme visuel. Si vous constatez que vous en avez trop ajouté, utilisez `/distill` pour la soustraction.
```text
> /typeset
"检查数字字体的大小、粗细层级是否合理"
> /polish
"最终打磨,检查对齐、间距、暗色模式支持"
```
{/* À FAIRE : Capture d'écran – comparaison détaillée du vernis avant et après le vernis */}
### Étape 6 : Révision – Contrôle systématique
**Skill**:`/design-review` + `/audit`
Assurance qualité visuelle + revue technique. `/design-review` prendra automatiquement des captures d'écran pour comparer et résoudre les problèmes, `/audit` vérifiera l'accessibilité et les performances.
```text
> /design-review
"视觉 QA,截图对比桌面和手机端"
> /audit
"检查可访问性和性能"
```
{/* À FAIRE : Capture d'écran – le rapport de notation produit par l'audit */}
### Étape 7 : Libération
**Skill**:`/ship`
Processus de publication standard de Gstack : tests, examen des différences, création de relations publiques.
```text
> /ship
```
***
**Résultats attendus** : une page de compte à rebours visuellement exquise, utilisant 8 à 10 compétences frontales dans le processus. Ce qui est plus important est d'établir l'intuition de « quelle compétence utiliser à quel stade ».
## Aide-mémoire quotidien
Ce qui précède est le processus complet. Si vous rencontrez des problèmes spécifiques dans le développement quotidien, consultez simplement ce tableau :
| Ma question actuelle | Quoi utiliser |
| ----------------------------------------------------------------------------- | ------------------------------------------------ |
| Je ne sais pas quel style je veux | `/design-shotgun` |
| La fonction page est bonne mais elle semble "presque inutile" | `/polish` |
| J'ai l'impression que quelque chose ne va pas mais je ne peux pas l'expliquer | `/design-review` |
| La police/la mise en page semble maladroite | `/typeset` |
| Espacement désordonné et disposition encombrée | `/arrange` |
| Le style s'écarte du système de conception | `/normalize` |
| Vous souhaitez extraire des composants publics | `/extract` |
| La page est trop complexe et je souhaite soustraire | `/distill` |
| Le design est trop simple et trop sûr | `/bolder` ou `/colorize` |
| Le design est trop chic et bruyant | `/quieter` |
| Le texte du message d'erreur n'est pas convivial | `/clarify` |
| Il y a un problème avec l'affichage sur le téléphone mobile | `/adapt` |
| Vous souhaitez ajouter des effets d'animation | `/animate` (de base) ou `/overdrive` (explosion) |
| Contrôle systématique avant la mise en ligne | `/audit` |
| debug | `/investigate` |
## Résumé
Cette note fait deux choses :
1. **Panorama** - les 27 compétences front-end de gstack sont classées en 6 couches (infrastructure→exploration→implémentation→amélioration→optimisation→révision)
2. **Démonstration pratique** - Utilisez une page d'anniversaire de compte à rebours pour parcourir le flux de travail complet, montrant quelles compétences sont utilisées à chaque étape et pourquoi
À retenir : l'utilisation la plus puissante de ces compétences n'est pas de les appeler individuellement, mais de les combiner dans un pipeline : explorer les directions, créer des implémentations, améliorer le peaufinage et réviser les versions, avec une sélection claire des compétences à chaque étape.
Mais ne vous laissez pas limiter par le processus : une fois que vous maîtrisez le sujet, `/frontend-design → /animate → /polish → /ship` quatre étapes suffisent la plupart du temps.
***
**Lecture connexe** :
* [gstack Concepts](/fr/docs/notes/gstack/concept) — Qu'est-ce que gstack et quels problèmes résout-il ?
* [Gstack Practical Chapter](/fr/docs/notes/gstack/practice) — Workflow complet de l'installation à l'exécution
* [gstack Skill Architecture Teardown](/fr/docs/notes/gstack/skill-architecture) — Que peuvent apprendre les développeurs de Skills ?
* [Claude Skills Concept](/fr/docs/notes/claude-skills/concept) — Comprendre le mécanisme sous-jacent des compétences
# Pratique Gstack : flux de travail complet, de l'installation à l'exécution
\##Présentation
Dans [Concept](/fr/docs/notes/gstack/concept), nous avons découvert le positionnement central de gstack - un ensemble de compétences basées sur les rôles qui transforme Claude Code en une équipe d'ingénierie virtuelle, et son positionnement différencié dans l'écosystème des outils de programmation d'IA par rapport à GSD, Superpowers, Ralph et d'autres solutions.
Cet article pratique se concentre sur **comment utiliser** : de l'installation et de la configuration à l'exécution du flux de travail complet, vous aidant à démarrer avec gstack en 30 minutes.
\##Installation et configuration
### Conditions préalables
* **Claude Code** est installé et disponible
* **Git** installé
* **Bun v1.0+** installé (gstack est construit sur Bun)
* Les utilisateurs Windows ont également besoin de Node.js
### Installation globale (recommandée, réalisée en 30 secondes)
```bash
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup
```
Le script d'installation fait trois choses :
1. Ajoutez les informations sur les compétences de gstack à votre fichier `CLAUDE.md`
2. Placez tous les fichiers de compétences dans le répertoire des compétences
3. Installez Playwright et le navigateur Chromium correspondant (pour `/browse` et `/qa`)
### Installation au niveau du projet (partage en équipe)
Si vous souhaitez que les membres de l'équipe obtiennent automatiquement gstack après le clonage du référentiel :
```bash
cp -Rf ~/.claude/skills/gstack .claude/skills/gstack
rm -rf .claude/skills/gstack/.git
cd .claude/skills/gstack && ./setup
```
\###Support multi-agents
gstack ne se limite pas à Claude Code et prend actuellement en charge **10 agents de programmation IA**. `./setup` détecte automatiquement les hôtes installés par défaut :
```bash
./setup --host codex # OpenAI Codex CLI
./setup --host opencode # OpenCode
./setup --host cursor # Cursor
./setup --host factory # Factory Droid
./setup --host slate # Slate
./setup --host kiro # Kiro
./setup --host hermes # Hermes
./setup --host gbrain # GBrain(修改版)
./setup --host openclaw # OpenClaw(通过 ACP 派发 Claude Code 会话)
```
Le chemin d'installation des compétences de chaque hôte a la forme de `~/./skills/gstack-*/` et n'interfère pas les uns avec les autres.
> 💡 **Options supplémentaires pour les utilisateurs d'OpenClaw** : En plus d'appeler via ACP, OpenClaw peut également installer directement 4 compétences méthodologiques natives (`gstack-openclaw-office-hours`, `gstack-openclaw-ceo-review`, `gstack-openclaw-investigate`, `gstack-openclaw-retro`) via ClawHub, qui peuvent être utilisées en conversation sans session Claude Code.
### Mode équipe (partage en équipe + mises à jour automatiques, recommandé)
La v1.x introduit le mode équipe : chaque développeur installe gstack globalement, et l'entrepôt enregistre uniquement "nous utilisons gstack", et les mises à jour se produisent automatiquement :
```bash
(cd ~/.claude/skills/gstack && ./setup --team) && \
~/.claude/skills/gstack/bin/gstack-team-init required && \
git add .claude/ CLAUDE.md && git commit -m "require gstack for AI-assisted work"
```
Remplacer `required` par `optional` est un « rappel doux » plutôt qu’obligatoire. Chaque fois que vous démarrez Claude Code, il exécutera automatiquement une vérification de mise à jour (limitation une fois par heure, sûre et silencieuse en cas de panne du réseau). Il n’y a aucun fichier vendu dans l’entrepôt et il n’y a pas de dérive de version.
### Mise à jour
```bash
cd ~/.claude/skills/gstack && git pull && ./setup
```
Ou utilisez `/gstack-upgrade` directement dans Claude Code.
## Référence complète des commandes
### Processus de sprint
| Commande | Rôle | Descriptif |
| --------------------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/office-hours` | Heures de bureau de YC | 6 questions forçantes pour reconstruire l'orientation du produit et générer des documents de conception |
| `/plan-ceo-review` | PDG / Fondateur | Vous recherchez des produits 10 étoiles, disponibles en quatre modèles de gamme |
| `/plan-eng-review` | Responsable de l'ingénierie | Architecture de verrouillage, flux de données, cas extrêmes, matrice de test |
| `/plan-design-review` | Concepteur senior | Concevez un score de 0 à 10, expliquez comment obtenir 10 points |
| `/plan-devex-review` | Leader de l'expérience développeur | Explorez les portraits des développeurs, évaluez TTHW et concevez des moments magiques ; trois modes (DX EXPANSION / POLISH / TRIAGE), 20-45 questions de forçage |
| `/autoplan` | Révision du pipeline | Exécutez automatiquement la revue PDG → Conception → Ingénierie → DX en séquence, décidez automatiquement en fonction des principes de prise de décision de codage et ne vous lancez que des « décisions de goût » |
### Conception
| Commande | Descriptif |
| ---------------------- | ------------------------------------------------------------------------------------------------------ |
| `/design-consultation` | Créez un système de conception complet à partir de zéro et générez DESIGN.md |
| `/design-shotgun` | Générez plusieurs variantes de conception IA et comparez les sélections dans le navigateur |
| `/design-html` | Générez du HTML/CSS de qualité production, prenez en charge la détection du framework React/Svelte/Vue |
### Examen et sécurité
| Commande | Rôle | Descriptif |
| ---------------- | ----------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/review` | Ingénieur d'état-major | Trouvez les bogues qui peuvent passer le CI mais qui exploseront en production, résoudront automatiquement les problèmes évidents et marqueront les lacunes d'intégrité |
| `/investigate` | Expert en débogage | Débogage systématique des causes profondes. Règle de fer : ne corrigez pas le bug tant que vous n’avez pas trouvé la cause première ; arrêter après 3 échecs de correctifs |
| `/design-review` | Concepteur capable d'écrire du code | Audit visuel + réparation automatique, soumission atomique, captures d'écran de comparaison avant et après |
| `/devex-review` | Testeur DX | Exécutez réellement l'intégration : parcourez les documents, exécutez le processus de saisie, chronométrez le TTHW, les erreurs de capture d'écran, comparez avec le score `/plan-devex-review` |
| `/cso` | Agent de sécurité | OWASP Top 10 + Modélisation des menaces STRIDE, 17 règles d'exclusion des faux positifs, seuil de confiance de 8/10, chaque résultat est accompagné de scénarios d'utilisation spécifiques |
### Tests et assurance qualité
| Commande | Descriptif |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/qa` | Ouvrez le vrai test du navigateur et trouvez le bug → Correction de la validation atomique → Générer un test de régression → Re-vérifier |
| `/qa-only` | Comme ci-dessus mais uniquement pour le reporting, aucune modification du code |
| `/benchmark` | Test de performances de base : chargement des pages, Core Web Vitals, taille des ressources, prise en charge avant et après comparaison |
| `/browse` | Commandes de navigateur de niveau \~ 100 ms, Chromium réel, captures d'écran, remplissage de formulaire, clics sur les éléments |
| `/open-gstack-browser` | Démarrez le navigateur GStack : contrôle visible de l'IA Chromium, livré avec une extension de barre latérale, furtivité anti-exploration, routage automatique du modèle (opération Sonnet/analyse Opus), prend en charge l'importation de cookies en un clic |
| `/setup-browser-cookies` | Importez des cookies depuis de vrais navigateurs (Chrome/Arc/Brave/Edge) vers des sessions sans tête pour tester les pages requises pour la connexion |
| `/pair-agent` | Couplage de navigateur d'agent Cross-AI : partagez le même navigateur GStack avec OpenClaw / Hermes / Codex / Cursor, etc., chaque agent a un onglet indépendant, est livré avec un tunnel ngrok pour prendre en charge les agents distants, jeton de portée + isolation d'onglet + limite de débit + attribution de comportement |
### Libération, exploitation et maintenance
| Commande | Descriptif |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/ship` | Synchroniser la branche principale → Exécuter des tests → Couverture d'audit → Mettre à jour la version → Soumettre le push → Créer un PR ; Bootstrap automatique lorsque le projet ne dispose pas de framework de test |
| `/land-and-deploy` | Fusionner PR → Attendre CI → Déployer → Vérifier l'état de l'environnement de production |
| `/canary` | Surveillance Canary post-déploiement : erreurs de console, régressions de performances, échecs de pages |
| `/setup-deploy` | `/land-and-deploy` Configuration unique : plateforme de détection automatique (Fly.io/Render/Vercel/Netlify/Heroku/GitHub Actions/custom) + URL de production + commande de déploiement |
| `/setup-gbrain` | Démarrez avec la base de données GBrain en un clic (en 5 minutes) : PGLite local, URL existante de Supabase, ou créez automatiquement un nouveau projet Supabase via l'API de gestion ; Enregistrement MCP + autorisations de lecture-écriture/lecture seule/refus au niveau de l'entrepôt |
### Réviser et apprendre
| Commande | Descriptif |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `/retro` | Rapport hebdomadaire sur la perception de l'équipe : démontage par habitant, statistiques des séquences de victoires, tendances de santé des tests, opportunités de croissance ; `/retro global` sur tous les projets + outils d'IA (Claude Code / Codex / Gemini) |
| `/document-release` | Mettre à jour automatiquement la documentation du projet pour qu'elle corresponde au code publié (README / ARCHITECTURE / CONTRIBUTING / CLAUDE.md / TODOS) ; `/ship` est désormais automatiquement appelé |
| `/learn` | Gérer les mémoires d'apprentissage inter-sessions : afficher, rechercher, élaguer, exporter, accumuler par projet |
| `/context-save` `/context-restore` | Package en mode point de contrôle continu : validation WIP automatique pour enregistrer le contexte, utilisez `/context-restore` pour reconstruire la session après un crash/un changement |
### Protection de sécurité
| Commande | Descriptif |
| ----------------------- | -------------------------------------------------------------------------------- |
| `/careful` | Avertissement de fonctionnement dangereux : rm -rf, DROP TABLE, force-push, etc. |
| `/freeze` / `/unfreeze` | Verrouiller/déverrouiller la portée de l'édition sur un répertoire spécifique |
| `/guard` | Combinaison `/careful` + `/freeze`, mode de sécurité le plus élevé |
| `/checkpoint` | Enregistrer/restaurer l'instantané de l'état de fonctionnement |
### Intégration des outils
| Commande | Descriptif |
| -------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/codex` | Intégration OpenAI Codex CLI : révision indépendante du code (porte réussite/échec), mode confrontation, mode consultation ; une analyse de chevauchement entre modèles sera donnée après l'exécution avec `/review` |
| `/health` | Tableau de bord de qualité du code : tsc + biome + knip + shellcheck + tests → note globale 0-10 |
| `/skillify` | Consolider le flux de travail actuel dans une compétence réutilisable |
| `/scrape` | Flux de travail de scraping Web |
| `/landing-report` | Rapport sur les performances et l'expérience de la page de destination |
| `/make-pdf` | Générer un document PDF |
| `/benchmark-models` `/model-overlays` `/plan-tune` | Comparaison inter-modèles, superposition de couverture, optimisation des forfaits |
### Standalone CLI(v0.19+)
En plus de la commande slash, gstack est également livré avec un ensemble de CLI autonomes (non exécutées dans la ses