← 提示词库 Anthropic/claude-code/skills/claude-api/shared/cost-optimization.md 原文 md
🌐 中英双语对照

Cost Optimization - Cutting Spend per Completed Task / 成本优化——降低每个已完成任务的开销

If you arrived via /claude-api cost-optimize: this is the right file. Execute the steps below in order rather than summarizing the guide back to the user - presenting the profile, the ranked plan, and the findings IS part of the execution. Start with Step 0 (establish scope, quality bar, and baseline), and finish with Step 4's two deliverables: the cost profile and the changes.

**如果你是通过 /claude-api cost-optimize 进入本文件的:**来对地方了。请按顺序执行以下步骤,而不是把本指南总结一遍复述给用户——呈现画像、排序列表与结论本身就是执行的一部分。从步骤 0(确定范围、质量标准与基线)开始,以步骤 4 的两项交付物(成本画像与变更)收尾。

API spend is optimized in units of cost per completed task, not cost per token. A model with a higher sticker price can be the cheaper option if it finishes the job in fewer turns, and a cheaper model that fails still bills its tokens, then the retry, then whatever the failure costs downstream. Every judgment below reads cost and quality together.

API 开销以每个已完成任务的成本、而非每 token 的成本为单位进行优化。标价更高的模型如果能在更少的轮次内完成工作,反而可能是更便宜的选择;而一个失败的廉价模型照样要为它的 token 计费,随后还有重试的费用,以及失败在下游造成的一切代价。下文的每项判断都将成本与质量放在一起衡量。

【评论】以"每完成任务成本"而非"每 token 成本"为优化单位,是本指南的核心立场:它把重试与失败的下游代价一并计入,避免只看 token 单价造成误判。

The levers divide into two kinds, and the order of the steps is load-bearing:

这些杠杆分为两类,而步骤的先后顺序是有实际效力的:

Where this workflow sits: the prompt-audit subcommand (shared/prompt-audit.md) audits the prompt surface (prompts, skills, tool descriptions) alone; this workflow is the holistic cost pass - request shape, caching, loop structure, output, batching, effort, model - and runs that audit as one sub-lever of input hygiene (§ 2.2) rather than restating its patterns; and once the project has an eval, the levers become a hillclimb - one change at a time against the eval, keep or revert (Step 3).

本工作流的定位:prompt-audit 子命令(shared/prompt-audit.md)只审计提示词表面(提示词、技能、工具描述);本工作流是整体性的成本处理——请求形态、缓存、循环结构、输出、批处理、effort、模型——并把该审计作为输入卫生的一个子杠杆(§ 2.2)来执行,而不是复述其模式;一旦项目有了评测(eval),这些杠杆就变成一次爬山——针对评测一次只改一处,保留或回退(步骤 3)。

Measured expectations below give the direction and rough size of each effect from Anthropic's published runs (sources at the end); the per-model figures live on those pages and change with each model release, so most are not restated here. They are directional, not guarantees - the validation loop in Step 3 is what makes a number true for this project - and fetching the Pricing and Cost Optimization pages is a required step, not background (Step 0 -> Fetch before you size). For measured figures, wherever a fetched page differs from what is quoted here, the page wins. For which history edits are valid, the Preserved thinking page governs (§ 2.1); the cost guide's and the cookbook's client-side prune recipes do not account for it.

下文的实测预期给出了 Anthropic 公开实验中每项效应的方向和大致量级(来源见文末);分模型的数字位于那些页面上,且随每次模型发布而变化,因此此处大多不予复述。它们是方向性的,不是保证——步骤 3 的验证回路才是让数字对本项目成立的东西——而抓取 Pricing 与 Cost Optimization 页面是必经步骤,不是背景阅读(步骤 0 -> 定规模前先抓取)。就实测数字而言,凡抓取到的页面与此处引述不一致,以页面为准。至于哪些历史编辑是合法的,以 Preserved thinking 页面为准(§ 2.1);成本指南与 cookbook 的客户端剪枝方案并未将其考虑在内。


Step 0: Establish scope, quality bar, and baseline / 步骤 0:确定范围、质量标准与基线

First, establish three things - from the request and the repository where they answer it, and from the user where they don't. Unlike the prompt audit, this workflow is interactive by design: when context for a lever is missing, or a step would spend real money, work through it with the user rather than assuming. It is not expected to one-shot the audit. State all three at the top of the report (the baseline value itself may read "pending Step 1" at first).

**首先确定三件事——能从请求和仓库得到答案的从那里取,得不到的向用户问。**与提示词审计不同,本工作流在设计上就是交互式的:当某个杠杆缺少上下文,或某一步要花费真金白银时,与用户一起解决,而不是自行假设。不要求一次跑完整个审计。在报告开头陈述全部三项(基线值本身起初可以写"待步骤 1")。

  1. Scope. If the request names files or directories, that is the scope. Otherwise it is every place the project calls the Claude API - request builders, agent loops, batch jobs. Note distinct traffic classes (an interactive path and a nightly job are different workloads even on one key): the profile, the ranking, and every validation later run per class, and "cost per task" means nothing blended across classes. Also establish which platform the code targets (first-party Anthropic API, Claude Platform on AWS, Bedrock, Vertex, or Foundry) - feature availability varies, and it filters which levers are even on the table.
    **范围。**如果请求点名了文件或目录,那就是范围。否则就是项目中调用 Claude API 的每一处——请求构造器、代理循环、批处理作业。记下不同的流量类别(交互路径与夜间作业即便在同一把密钥上也是不同的工作负载):画像、排序以及后续每项验证都按类别分别运行,"每任务成本"跨类别混合后毫无意义。同时确定代码面向哪个平台(Anthropic 一方 API、AWS 上的 Claude Platform、Bedrock、Vertex 还是 Foundry)——功能可用性各不相同,它会过滤掉哪些杠杆根本不在考虑之列。

  2. Quality bar. Find the project's eval, test suite, or outcome checks for its LLM calls. If none exists, say so prominently in the report: without one, savings cannot be told apart from regressions. Do not stop - free wins are safe to propose regardless - but mark every tradeoff lever "needs an eval before applying", and ask the user what outcome check they can provide. An eval only validates the traffic class it covers: mark levers on uncovered paths the same way. If the only check is the user's own manual review, it gates free wins - it never clears a tradeoff. The full no-eval endgame - including a minimal eval recipe that unblocks tradeoffs - is in Step 3.
    **质量标准。**找到项目针对其 LLM 调用的评测、测试套件或结果校验。如果不存在,在报告中显著说明:没有它,节省与退化就无法区分。不要因此停下——免费收益无论如何都可以放心提出——但把每个权衡类杠杆标注为"应用前需要评测",并询问用户能提供什么样的结果校验。评测只对它覆盖的流量类别有效:对未覆盖路径上的杠杆照此标注。如果唯一的校验是用户本人的人工审查,它可以约束免费收益——但永远不能放行权衡项。完整的无评测应对方案——包括一份能解锁权衡项的最小评测配方——在步骤 3。

  3. Baseline cost per task. The baseline is whatever honest number is cheapest to obtain, in this order:
    **每任务基线成本。**基线就是以最廉价方式能得到的诚实数字,按以下优先顺序:

    • From history, free: with Admin API access, pull Step 1's usage and cost reports forward and compute the baseline from them - the reports supply the dollars, but the per-task denominator must come from the user or the application's own logs; or roll up the application's own logged usage objects per task, not per request - four token counts, each at its own rate: regular input, cache writes (1.25x input for the 5-minute duration, 2x for 1-hour), cache reads (a fraction of base input that differs by model), and output - multiplier structure as published on the pricing page; take the values from it when you fetch the rates, not from memory.
      从历史记录,免费:如有 Admin API 访问权限,把步骤 1 的用量与成本报告提前拉取过来并据此计算基线——报告给出金额,但每任务分母必须来自用户或应用自身的日志;或者按任务(而非按请求)汇总应用自己记录的 usage 对象——四项 token 计数,各按各的费率:普通输入、缓存写入(5 分钟时长为输入的 1.25 倍,1 小时为 2 倍)、缓存读取(基础输入的一个比例,因模型而异)、输出——倍率结构以定价页面公布者为准;抓取费率时以页面取值,不要凭记忆。
    • From a baseline run, paid: run the project's eval (or, with no eval, replay a representative sample of real requests) and roll up the same way. This spends real API money: state the expected cost - from Step 1's token estimates and live pricing, and "estimated - pending Step 1" is an acceptable first answer - and get the user's approval before running it. If the user declines the spend, estimate the baseline from the code and any bill figure they can read off the Console, label it an estimate, and continue.
      从基线运行,付费:运行项目的评测(没有评测时,重放一份有代表性的真实请求样本),并按同样方式汇总。这会花费真实的 API 资金:说明预期成本——依据步骤 1 的 token 估算与实时定价,"估算值——待步骤 1"是可以接受的第一版答案——**并在运行前取得用户批准。**如果用户拒绝这笔支出,就依据代码以及用户能从控制台读到的任何账单数字来估算基线,标注为估算值,然后继续。

    Fetch before you size - this gates every rate and every measured figure in the audit. Before any dollar amount, multiplier, or published figure is written into the report or into code, WebFetch two rows from shared/live-sources.md: Pricing (per-model rates and multipliers) and Cost Optimization (current measured expectations and the starting-model recommendation). Fetch both in one turn. When a lever that edits conversation history, system, or tools (§ 2.2, § 2.3, a model switch) enters the shortlist, also fetch the Preserved thinking page (URL in § 2.1) before proposing its diff, for which models and accounts run the check. Record in the report which pages were fetched and when, so a reader can see it happened, and cite the fetched page beside each figure taken from it. Remembered rates or figures, for any model, are not a source. If a fetch fails, say so in the report. For rates: with Admin API access, derive effective realized rates by dividing cost-report amounts by the usage report's matching token counts (same model, same token type); otherwise ask the user for the current rates. For measured expectations: size them in relative buckets with no figures. Only if none of these is available and the user still wants a number may one appear, and then labeled "unverified - from memory" every place it appears.

    **定规模前先抓取——这一步约束着审计中的每一条费率和每一个实测数字。**在任何金额、倍率或公开数字写入报告或代码之前,先用 WebFetch 抓取 shared/live-sources.md 中的两行:Pricing(分模型费率与倍率)和 Cost Optimization(当前实测预期与起始模型建议)。在同一轮中把两者都抓取。当某个会编辑对话历史、system 或 tools 的杠杆(§ 2.2、§ 2.3、模型切换)进入候选清单时,在提出其 diff 之前还要抓取 Preserved thinking 页面(URL 见 § 2.1),以确认哪些模型与账户会执行该检查。在报告中记录抓取了哪些页面、何时抓取,让读者能看到这一步确实发生过,并在每个取自页面的数字旁注明出处页面。凭记忆记下的任何模型的费率或数字都不是来源。如果抓取失败,在报告中说明。费率方面:如有 Admin API 访问权限,用成本报告金额除以用量报告中对应的 token 计数(同模型、同 token 类型)推导实际实现费率;否则向用户询问当前费率。实测预期方面:以不带数字的相对档位来定规模。只有当以上途径都不可用而用户仍想要一个数字时才可以给出,并且在其出现的每一处都标注"未经验证——凭记忆"。

【评论】反复要求实时抓取官方页面、明令禁止凭记忆引用费率,是一种防幻觉设计:防止模型用训练时记错的旧价格推算节省额。

For counting tokens in prompts and files, see shared/token-counting.md (count_tokens returns the count without running inference). Sanity-check an estimated baseline against any known monthly bill: divergence usually means multi-turn history growth the single-turn estimate missed.

统计提示词与文件中的 token 数见 shared/token-counting.md(count_tokens 无需运行推理即返回计数)。把估算的基线与任何已知的月账单做一致性核对:偏差通常意味着单轮估算遗漏了多轮历史增长。

Step 1: Profile where the tokens go / 步骤 1:画像 token 的去向

The profile can be measured or estimated. Measure when the organization's access allows it; fall back to reading the code. Either way, the levers that pay are decided by the workload's shape, not by the list of what exists.

画像可以是实测的,也可以是估算的。组织权限允许时就实测;否则退回到读代码。无论哪种方式,值得动的杠杆由工作负载的形态决定,而不是由现存事物的清单决定。

Measure it - the Usage and Cost Admin API (preferred) / 实测——用量与成本 Admin API(首选)

If the user has an Admin API key (sk-ant-admin01-... - a different key type from the standard API key; not available for individual accounts - creation and scopes are covered in the Admin API docs, reachable from the Usage and Cost Admin API URL in shared/live-sources.md), pull the real numbers instead of estimating. These are report reads, not model calls - they consume no tokens. Full parameters and response schemas: the Usage and Cost Admin API URL in shared/live-sources.md.

如果用户有 Admin API 密钥(sk-ant-admin01-...——与标准 API 密钥不同的另一种密钥类型;个人账户不可用——创建方式与权限范围见 Admin API 文档,可从 shared/live-sources.md 中的 Usage and Cost Admin API URL 访问),就拉取真实数字而不是估算。这些是报告读取,不是模型调用——不消耗 token。完整参数与响应模式:shared/live-sources.md 中的 Usage and Cost Admin API URL。

The measured profile answers directly: the real cache hit rate (cache_read_input_tokens against uncached input), how much traffic already rides the batch tier, the input/output balance, and where spend concentrates by model, key, and workspace. Check that the measured footprint plausibly matches the audited code (same models, a believable order of magnitude): the report covers the whole organization, and a key shared across projects blends their traffic - making per-project reads, including Step 3's post-cutover confirmation, unattributable. On a mismatch, reconcile against the code estimate, scope usage-report queries by api_key_ids[] / workspace_ids[] where the separation exists (the cost report takes neither filter - it segments only by workspace, via group_by), and recommend per-project keys or workspaces as a measurement prerequisite where it doesn't. Optimization effort follows the audited scope's spend, not the org blend.

实测画像直接回答:真实缓存命中率(cache_read_input_tokens 相对未缓存输入)、已有多少流量走批处理档、输入/输出平衡,以及支出按模型、密钥、工作区集中在哪里。检查实测足迹与被审计的代码是否合理吻合(相同模型、可信的数量级):报告覆盖整个组织,跨项目共享的密钥会把它们的流量混在一起——使按项目读取(包括步骤 3 切换后的确认)无法归因。若有不符,与代码估算核对;在存在隔离之处用 api_key_ids[] / workspace_ids[] 限定用量报告查询(成本报告不接受这两种过滤——只能通过 group_by 按工作区分段);在没有隔离之处,建议把按项目的密钥或工作区作为度量的前置条件。优化精力跟随被审计范围的支出,而不是组织整体的混合值。

Estimate it from the code / 从代码估算

Without Admin API access (no Admin key, a Claude Enterprise organization, Claude Platform on AWS - whose feature availability shared/claude-platform-on-aws.md covers - or another cloud platform these reports do not cover) - and even with it, for the structural facts no usage report can show - read the request-building code:

在没有 Admin API 访问权限时(没有 Admin 密钥、Claude Enterprise 组织、AWS 上的 Claude Platform——其功能可用性见 shared/claude-platform-on-aws.md——或这些报告不覆盖的其他云平台)——即便有权限,对于那些任何用量报告都显示不了的结构性事实——去读构造请求的代码:

Per-model defaults, parameter support, and per-platform feature availability change across releases. For any "what happens when thinking/effort is omitted", "does this model accept effort", "what levels does it support", or "is this feature available on Bedrock/Vertex/Foundry" question, read the answer from SKILL.md -> Thinking & Effort, shared/models.md, or shared/platform-availability.md (or the live Models API) - never assume, and never encode the answer in this guide.

**分模型的默认值、参数支持与分平台功能可用性随版本而变。**任何"省略 thinking/effort 会怎样"、"这个模型是否接受 effort"、"支持哪些档位"或"该功能在 Bedrock/Vertex/Foundry 上是否可用"的问题,都从 SKILL.md -> Thinking & Effort、shared/models.md 或 shared/platform-availability.md(或实时 Models API)读取答案——绝不臆测,也绝不把答案硬编码进本指南。

Ask for the app's own usage logs first / 先索要应用自己的用量日志

Before ranking on estimates, ask the user whether the application already logs response.usage per request - and if so, to paste a representative day's worth. That turns cache hit rate, the input/output split, and thinking-token spend from guesses into measurements at zero API cost, and it decides which tier of the ranking table below applies. Two fields sharpen it where the response carries them: usage.output_tokens_details.thinking_tokens meters thinking spend directly, and a request that runs a server-side step such as compaction itemizes that step's tokens under usage.iterations - sum the iterations rather than reading the top-level counts alone. If the app doesn't log usage yet, note that adding it is itself a free-win diff (Step 3) and proceed on the code estimate.

在基于估算排序之前,询问用户应用是否已按请求记录 response.usage——如果有,请其粘贴有代表性的一天量。这会把缓存命中率、输入/输出拆分与思考 token 花费从猜测变成零 API 成本的实测,并决定下文排序表适用哪一档。响应携带时有两个字段能让它更精确:usage.output_tokens_details.thinking_tokens 直接计量思考花费;而运行服务端步骤(如压缩)的请求会在 usage.iterations 下分项列出该步骤的 token——把各次迭代求和,不要只读顶层计数。如果应用尚未记录用量,注明添加它本身就是一份免费收益 diff(步骤 3),然后按代码估算继续。

Estimating cache hit rate without usage data. If the app logs request timestamps, simulate the TTL walk: sort timestamps, count a hit whenever the gap to the previous request is <= TTL (reads refresh the entry), and run it for each cache TTL the platform offers (see shared/prompt-caching.md) - the difference between durations is the longer-TTL lever's ceiling on the user's real traffic. If only aggregate volume is known, approximate with Poisson arrivals: hit rate ~ 1 - e^(-lambda·TTL) where lambda is requests per second. Either beats comparing average gap to TTL, which ignores burstiness.

**在没有用量数据时估算缓存命中率。**如果应用记录请求时间戳,就模拟 TTL 行走:对时间戳排序,凡与前一请求的间隔 <= TTL 就计一次命中(读取会刷新条目),并对平台提供的每种缓存 TTL 各跑一遍(见 shared/prompt-caching.md)——两种时长之差就是更长 TTL 杠杆在用户真实流量上的上限。如果只知道总量,用泊松到达近似:命中率 ~ 1 - e^(-lambda·TTL),其中 lambda 是每秒请求数。这两种方法都优于把平均间隔与 TTL 直接比较——后者忽略了突发性。

Rank the levers / 对杠杆排序

Before touching code, size each lever the profile makes applicable so the shortlist can be ordered. How you quote the size depends on what data you have - an estimate and a measurement must not look the same in the report:

在改动代码之前,先给画像判定适用的每个杠杆定出规模,让候选清单可以排序。引用规模的方式取决于你手上有什么数据——估算与实测在报告中绝不能呈现得一样:

Data available Quote each ceiling as
Admin API usage/cost report Dollar range, labeled measured
App-side usage logs, or a user-reported bill total only % of current bill, with dollars only as a parenthetical "(~ $Y at your reported $X/mo)" - the % is the claim; the $ is the user's own arithmetic
Neither (pure code read) Relative buckets - "largest / medium / small", or an order-of-magnitude band - no specific figures
可用数据 上限的引用方式
Admin API 用量/成本报告 美元区间,标注为 measured
仅有应用侧 usage 日志,或用户报出的账单总额 当前账单的百分比,美元只作括注"(~ $Y at your reported $X/mo)"——百分比是主张;美元是用户自己的算术
两者皆无(纯读代码) 相对档位——"最大/中等/较小"或数量级区间——不给具体数字

Before sizing, drop any lever the target platform doesn't support (shared/platform-availability.md is the single source of truth - do not assume 1P availability carries to Bedrock, Vertex, Foundry, or Claude Platform on AWS). A lever that can't ship on the user's platform isn't worth ranking; list it under "skipped" with the availability reason instead.

定规模之前,先剔除目标平台不支持的杠杆(shared/platform-availability.md 是唯一权威来源——不要假定一方 API 的可用性会延伸到 Bedrock、Vertex、Foundry 或 AWS 上的 Claude Platform)。在用户平台上无法落地的杠杆不值得排序;把它连同可用性原因列在"已跳过"之下。

Within whichever unit applies, size each lever from the measured (or estimated) spend components and the measured expectations described in Step 2 and quantified on the Cost Optimization page - use the copy fetched in Step 0 for any per-model figure (if that fetch failed, carry the ceiling as a relative bucket) - for example:

在适用的单位体系内,从实测(或估算)的支出分量、步骤 2 所述并在 Cost Optimization 页面上量化的实测预期出发,给每个杠杆定规模——任何分模型数字都用步骤 0 抓取的副本(若那次抓取失败,则把上限按相对档位携带)——例如:

Ceilings that claim the same tokens (caching an inlined document versus deleting it) are mutually exclusive: compute each ceiling unconditionally, rank, then deflate each for its overlap with the levers above it, so the shortlist can never sum past the bill.

对同一批 token 提出主张的上限(缓存一份内联文档与删掉它互斥):先无条件计算每个上限,排序,然后按其与更上方杠杆的重叠逐级折减,使候选清单之和永远不会超过账单。

Present the ranked shortlist with the profile evidence behind each number - labeled as ranked by savings ceiling, not application order (Step 2's § 2.x numbering decides the sequence) - and say where the list stops: a lever whose ceiling is a small fraction of the bill - or would not repay the approved runs and effort needed to validate it - does not earn an eval cycle, and most levers will not earn a place on any given workload (the "Workload shape -> lever" table near the end of this file is the map for matching profile to levers). On a small bill the honest shortlist may be empty: "nothing here is worth changing" is a successful finding, not a failure - report it plainly. Expected savings are planning numbers, not results - Step 3's measurements are the results.

呈现排好序的候选清单时,附上每个数字背后的画像证据——标注为按节省上限排序、而非应用顺序(步骤 2 的 § 2.x 编号决定顺序)——并说明清单到哪里为止:一个上限只占账单很小比例——或收回不了验证它所需的已批准运行与精力——的杠杆不配得到一轮评测,而且大多数杠杆在任一给定工作负载上都得不到位置(本文件末尾附近的"工作负载形态 -> 杠杆"表就是把画像对应到杠杆的地图)。账单很小时,诚实的候选清单可能是空的:"这里没有值得改的"是一个成功结论,不是失败——如实报告。预期节省是规划数字,不是结果——步骤 3 的测量才是结果。

Step 2: Work the levers in order / 步骤 2:按顺序运用杠杆

Free wins may be applied directly when the request asked for edits (a bare subcommand invocation has not asked - propose). Tradeoff levers (2.6 onward) are always presented with their measured quality cost and applied only on the user's explicit acceptance - never trade accuracy for cost silently. And every run that exercises the model - the baseline, each lever's validation pass - spends real API money: get explicit approval before each one, with the expected cost, or once as a Step 3 measurement budget that covers them.

当请求本身要求改动时,免费收益可以直接应用(单纯调用子命令不算要求——只提出建议)。权衡类杠杆(2.6 起)总是连同其实测质量代价一起呈现,并且只在用户明确接受后才应用——绝不在暗中用准确率换成本。而每一次真正驱动模型的运行——基线、每个杠杆的验证轮——都花费真实 API 资金:每次之前取得明确批准并说明预期成本,或者一次性取得覆盖它们的步骤 3 测量预算。

【评论】本节把"花钱必须先获用户明确批准"设为硬性门槛,将成本控制与权限控制绑定,防止代理自行产生大额 API 支出。

Pricing multipliers quoted below (cache write rates, batch discount) are current as of writing - confirm against the Pricing URL in shared/live-sources.md before computing any ceiling. The cache-read rate differs by model and is deliberately not quoted here.

下文引用的定价倍率(缓存写入费率、批处理折扣)以撰写时为准——计算任何上限之前先对照 shared/live-sources.md 中的 Pricing URL 确认。缓存读取费率因模型而异,此处刻意不予引述。

2.1 Prompt caching - first, and it stays on / 2.1 提示词缓存——第一顺位,且永久开启

Every turn of an agentic task resends the entire growing conversation - system prompt, tool definitions, every prior turn - so a 40-turn task sends its first turn 40 times and task cost grows with roughly the square of turn count. Caching does not stop the resending; it reprices everything already cached to the model's cache-read rate, a small fraction of base input.

代理型任务的每一轮都重发整段不断增长的对话——系统提示词、工具定义、此前每一轮——因此一个 40 轮的任务会把第 1 轮发送 40 次,任务成本大致随轮数的平方增长。缓存并不阻止重发;它把已缓存的一切按模型的缓存读取费率重新计价,那是基础输入的一个很小比例。

For design and placement - the prefix-match invariant, classifying inputs by stability, breakpoint patterns, the anti-pattern table - read shared/prompt-caching.md and follow its workflow; do not improvise cache_control markers. Points that matter specifically for cost:

设计与摆放——前缀匹配不变量、按稳定性分类输入、断点模式、反模式表——读 shared/prompt-caching.md 并遵循其工作流;不要即兴发挥 cache_control 标记。专门与成本相关的要点:

2.2 Input tokens - progressive disclosure / 2.2 输入 token——渐进式披露

Send the model what the task needs, let it fetch the rest. Each sub-lever has a skip-when; the caveat at the end of this section governs all of them.

把任务需要的发给模型,其余让它自己去取。每个子杠杆都有"何时跳过";本节末尾的告示约束全部子杠杆。

Caveat for the whole section: a smaller prefix is not automatically a cheaper task. Deferring context means the model may spend discovery turns fetching what it previously read inline. Validate against the eval - on the cookbook's workload, wrapping the manual in a tool matched the explicit-breakpoint config on cost and gave back accuracy. These changes edit system or tools, so apply them to new conversations only where the harness can, and otherwise say in the report what a conversation already in flight pays when they deploy: a restarted cache and, on models that run the preserved-thinking check, invalid thinking blocks in its history (§ 2.1). Tools declared up front with defer_loading are the form that stays valid.

本节整体告示:更小的前缀不自动等于更便宜的任务。延后上下文意味着模型可能花发现轮去取此前内联可读的内容。对照评测验证——在 cookbook 的工作负载上,把手册包进工具在成本上与显式断点配置打平,但回吐了一些准确率。这些变更会编辑 system 或 tools,因此在 harness 能做到之处只应用于新会话;做不到时,在报告中说明已在进行中的会话在其部署时要付出什么:缓存重启,以及在运行 preserved-thinking 检查的模型上历史中失效的思考块(§ 2.1)。用 defer_loading 预先声明的工具是保持有效的形式。

2.3 Agent-loop hygiene - keep long loops from compounding / 2.3 代理循环卫生——防止长循环复利放大

Only relevant when the profile shows deep loops with bulky accumulating results; short loops never trigger these and the added machinery is pure overhead.

只在画像显示存在带庞大累积结果的深循环时才相关;短循环永远不会触发这些,新增机制纯属开销。

2.4 Output tokens / 2.4 输出 token

2.5 Batch processing / 2.5 批处理

50% off every token in the request, including cache reads and writes - the discounts stack. The second-largest free lever after caching for unattended agent work - evaluation runs, backfills, scheduled jobs.

请求中的每个 token 都打五折,包括缓存读与写——折扣可叠加。对于无人值守的代理工作——评测运行、回填、定时作业——这是仅次于缓存的第二大免费杠杆。

2.6 Effort and budgets - the first tradeoffs / 2.6 effort 与预算——最初的权衡项

From here down, every lever trades capability for cost. Sweep on the eval, one change at a time.

从这里往下,每根杠杆都在用能力换成本。在评测上扫描,一次只改一处。

2.7 Model selection - last, deliberately / 2.7 模型选择——刻意放在最后

Model choice constrains the intelligence ceiling, which is why it comes after every lever that doesn't. (The exception is when the project has an eval and you are running the hillclimb loop: there shared/evals/cost-hillclimb.md walks model x effort early, because an eval can detect the case where a stronger model at lower effort is the cheaper cell. Without an eval that case is invisible and a model swap stays the riskiest change - keep it last.)

模型选择约束智能上限,这正是它排在所有不约束智能上限的杠杆之后的原因。(例外是项目有评测且你在跑爬山回路时:那里 shared/evals/cost-hillclimb.md 会提前走模型 x effort,因为评测能发现"更强模型配更低 effort 才是更便宜单元格"的情形。没有评测时该情形不可见,换模型仍是最危险的变更——留在最后。)

Step 3: Apply, measure, keep or revert - one lever at a time / 步骤 3:应用、测量、保留或回退——一次一根杠杆

Workload shape -> lever / 工作负载形态 -> 杠杆

Adapted from the cookbook's takeaways table, for mapping a profile to levers (row 1's watch-out is extended):

改编自 cookbook 的要点表,用于把画像对应到杠杆(第 1 行的注意事项有所扩展):

Where the cost is Reach for Skip it or watch out when
Same system prompt and tools re-billed on every call Prompt caching with auto first, then an explicit breakpoint on the static prefix when many independent conversations share it or prefix layers change at different rates, and the 1-hour TTL or a keep-alive where gaps run past five minutes (§ 2.1) Anything dynamic sits above the breakpoint - move that content into the user turn. And a cache that already reads well needs nothing: concurrent-batch misses (§ 2.5) aren't breakers, and a 1-hour TTL doesn't reach calls that are hours apart
Large reference document in every prompt Move it behind a tool or skill Each call needs most of the document rather than a section, or the eval shows misses on cases that hinge on rules the model has to go looking for
Many or heavy tool schemas Tool search with defer_loading Under roughly 10K schema tokens, where the search step is overhead
Images, PDFs, or large files in context Downscale images to what the task needs, and use the Files API plus code execution for tables and PDFs There is nothing to extract or compute so the sandbox only adds tokens
Unbounded user-supplied input Token counting as an ingestion gate
Bulky results piling up across a long loop Context editing or compaction server-side; a client-side prune at natural boundaries, gated on models that run the preserved-thinking check (§ 2.3) Loops are short or the cleared content is still needed, and note that every edit breaks the cache from that point - and a client-side edit of earlier turns also invalidates later thinking on models that run the preserved-thinking check
One self-contained step with bulky intermediates Subagent, optionally on a cheaper model The deciding model needs that intermediate context to judge well
Long visible responses Specify the output shape with an example, with max_tokens as a backstop and a stop-sequence sentinel for early exits
Thinking and tool calls dominate, and the eval has headroom Lower effort first; check for a newer model in the tier (§ 2.7), then drop a model tier and re-sweep effort Always a direct capability trade, so step down one notch at a time against the eval
Mostly routine cases with a few hard ones Advisor tool on a cheaper driver There is no cheap signal to gate the consult, leaving the driver to spot hard cases itself
No one is waiting on the response Batch API, flattening a tool loop into one request by pre-fetching its inputs if you have to A user is waiting, or when flattening changes how the model reasons
成本在哪里 用什么 何时跳过或当心
每次调用都重复计费同样的系统提示词与工具 先自动提示词缓存;当许多独立会话共享静态前缀或前缀各层变化速率不同时,在静态前缀上加显式断点;间隔常超过五分钟处用 1 小时 TTL 或保活(§ 2.1) 断点上方有任何动态内容——把那些内容挪进用户轮。已经读得很好的缓存什么都不需要:并发批处理未命中(§ 2.5)不是破坏者,而 1 小时 TTL 也够不着相隔数小时的调用
每个提示词都带大型参考文档 把它挪到工具或技能之后 每次调用需要文档的大部分而非一节,或评测显示在依赖模型须自行寻找规则的用例上出现失误
工具模式多或重 用 defer_loading 做工具搜索 模式 token 约在 10K 以下,此时搜索步骤是开销
上下文中的图片、PDF 或大文件 图片降采样到任务所需;表格与 PDF 用 Files API 加代码执行 没有可提取或可计算的内容,沙箱只会增加 token
无上界的用户提供输入 用 token 计数作摄取闸门
庞大结果在长循环中不断堆积 服务端上下文编辑或压缩;自然边界上的客户端剪枝,在运行 preserved-thinking 检查的模型上受限制(§ 2.3) 循环短或被清除内容仍有用;注意每次编辑都从该点起破坏缓存——且在运行 preserved-thinking 检查的模型上,对较早轮次的客户端编辑还会使其后的思考失效
单个自成体系步骤带庞大中间结果 子代理,可选更便宜的模型 决策模型需要那段中间上下文才能判断好
可见回复过长 用示例规定输出形状,max_tokens 作后备,停止序列哨兵作提前退出
思考与工具调用占主导且评测有余量 先降 effort;查层级内有无更新模型(§ 2.7),再降一档模型并重扫 effort 永远是直接的能力交换,因此对照评测一次只降一档
常规案例居多、少数困难 更廉价驱动模型上配顾问工具 没有廉价信号为咨询设闸门,只能靠驱动模型自己发现困难案例
没有人在等响应 Batch API,必要时预先抓取输入把工具循环压平成一个请求 有用户在等,或压平会改变模型推理方式时

Step 4: Deliverables / 步骤 4:交付物

  1. The cost profile and plan: the Step 0 assumptions (scope, quality bar, baseline), the Step 1 token profile, and the levers chosen with the measured expectation each one carries - plus the levers deliberately skipped and why, so the next person doesn't re-litigate them. Label the shortlist table as ranked by savings ceiling, not application order, so it can't be misread as the diff sequence.
    成本画像与计划:步骤 0 的假设(范围、质量标准、基线)、步骤 1 的 token 画像,以及所选杠杆及各自携带的实测预期——外加刻意跳过的杠杆及原因,免得下一个人重新争论一遍。给候选清单表标注"按节省上限排序、非应用顺序",以免被误读为 diff 顺序。
  2. The changes: one diff per lever so effects attribute - applied and measured (expected versus measured cost per task, pass rate held or not) where the user approved the runs; left as proposals carrying their expected savings and published quality cost where they didn't, or where a tradeoff lever still needs an eval. When nothing cleared the ranking floor, this deliverable is "no changes recommended" - a successful outcome; say it plainly rather than manufacturing a lever.
    变更:一根杠杆一个 diff,让效果可归因——用户批准运行之处,已应用并测量(预期与实测的每任务成本对比、通过率是否守住);未批准之处、或权衡杠杆仍需评测之处,保留为携带预期节省与公开质量代价的建议。当没有任何东西过得了排名底线时,这项交付物就是"不建议变更"——这是一个成功结果;如实说出,而不是硬造一根杠杆。

Report skeleton (section order and required columns - keep the rest flexible):

报告骨架(小节顺序与必需列为固定项——其余保持灵活):

Sources and live references / 来源与实时引用

The measured results above come from two published Anthropic sources (and the Admin API facts in Step 1 from a third); Step 0 requires the first and the Pricing page before anything is sized; the cookbook and Admin API docs are fetched when the user needs the full write-ups or schemas:

上文的实测结果来自两个 Anthropic 公开来源(步骤 1 的 Admin API 事实来自第三个);在任何东西定规模之前,步骤 0 要求抓取第一个与 Pricing 页面;当用户需要完整论述或模式时才抓取 cookbook 与 Admin API 文档: