Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ Click the menu-bar item to open a three-tab panel:
- **Composition** (构成) — where the money went: spend broken down by model and by project.
- **Insights** (洞察) — code output, tool-use mix, golden-hours heatmap, savings tips, and a quota-depletion forecast.

> **How the numbers are computed.** *Spend* is an **estimate at pay-as-you-go API prices** (`Pricing.swift`), **not your subscription bill** — on a Max / ChatGPT plan, read it as "equivalent API value," not money actually charged. *Code output* is **approximate git attribution**: all non-merge commits in a session's working directory within the time window — it can't tell hand-written from AI commits and excludes uncommitted work. *Codex* token totals are de-duplicated from each session's cumulative counter (`total_token_usage`), so they no longer double-count replayed events.
> **How the numbers are computed.** *Spend* is an **estimate at standard pay-as-you-go API prices** (`Pricing.swift`), **not your subscription bill** — on a Max / ChatGPT plan, read it as "equivalent API value," not money actually charged. Models without a published rate use a visibly approximate estimate. *Code output* is **approximate git attribution**: all non-merge commits in a session's working directory within the time window — it can't tell hand-written from AI commits and excludes uncommitted work. *Codex* uses per-response token usage in recent logs; older logs use cumulative-counter deltas without charging the prior session's starting balance.

## Privacy

Expand Down
2 changes: 1 addition & 1 deletion README_ZH.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,7 @@ make package # 产出 dist/CodingBar.app
- **构成**(Composition)— 钱花在哪:按模型和按项目拆解花费。
- **洞察**(Insights)— 代码产出、工具使用占比、黄金时段热力图、省钱提示、额度燃尽预测。

> **数字是怎么算的。** *花费*是按 **API 现付价估算**(`Pricing.swift`),**不是你的订阅账单**——包月 Max / ChatGPT 用户应把它读作「等价 API 价值」,而非实际扣款。*代码产出*是**近似的 git 归因**:会话工作目录在时间窗内的全部非 merge 提交——无法区分手写与 AI 提交,且不含未提交改动。*Codex* 的 token 总量按每个会话的累计计数器(`total_token_usage`)做差分去重,不再把重复事件算两遍。
> **数字是怎么算的。** *花费*是按 **标准 API 现付价估算**(`Pricing.swift`),**不是你的订阅账单**——包月 Max / ChatGPT 用户应把它读作「等价 API 价值」,而非实际扣款。未公布价格的模型使用带近似标记的估价。*代码产出*是**近似的 git 归因**:会话工作目录在时间窗内的全部非 merge 提交——无法区分手写与 AI 提交,且不含未提交改动。*Codex* 的新日志使用逐请求 token 用量;旧日志从累计计数器做差分,不把前一会话的初始累计值算作本次用量。

## 隐私

Expand Down
48 changes: 30 additions & 18 deletions Sources/CodingBar/SelfTest.swift
Original file line number Diff line number Diff line change
Expand Up @@ -30,10 +30,16 @@ enum SelfTest {
let september = Date(timeIntervalSince1970: 1_788_220_800)
check("Fable 5 1h cache pricing", abs(Pricing.cost(model: "claude-fable-5", tokens: millionTokens,
at: july, cacheWrite1h: 1_000_000) - 81) < 0.000_001)
check("Sonnet 5 intro pricing", abs(Pricing.cost(model: "claude-sonnet-5", tokens: millionTokens,
at: july, cacheWrite1h: 1_000_000) - 16.2) < 0.000_001)
check("Sonnet 5 standard pricing", abs(Pricing.cost(model: "claude-sonnet-5", tokens: millionTokens,
at: september, cacheWrite1h: 1_000_000) - 24.3) < 0.000_001)
check("Sonnet 5 permanent price", abs(Pricing.cost(model: "claude-sonnet-5", tokens: millionTokens,
at: july, cacheWrite1h: 1_000_000) - 16.2) < 0.000_001
&& abs(Pricing.cost(model: "claude-sonnet-5", tokens: millionTokens,
at: september, cacheWrite1h: 1_000_000) - 16.2) < 0.000_001)
check("latest Claude tiers and cache reads",
abs(Pricing.cost(model: "claude-opus-5-5", tokens: millionTokens,
at: september, cacheWrite1h: 1_000_000) - 32.2) < 0.000_001
&& abs(Pricing.cost(model: "claude-fable-5-1", tokens: millionTokens,
at: september, cacheWrite1h: 1_000_000) - 80.25) < 0.000_001
&& Pricing.priceIsExact(model: "claude-mythos-5-1"))

let openAIBaseTokens = TokenBreakdown(input: 100_000, output: 100_000,
cacheRead: 100_000, cacheWrite: 100_000)
Expand All @@ -44,10 +50,18 @@ enum SelfTest {
&& Pricing.normalize(model: "gpt-5.6-luna") == "openai/gpt-5.6-luna"
&& Pricing.priceIsExact(model: "gpt-5.6"))
check("GPT-5.6 Sol base and long-context pricing",
abs(Pricing.cost(model: "gpt-5.6-sol", tokens: openAIBaseTokens, at: july,
billingInputTokens: 272_000) - 4.175) < 0.000_001
&& abs(Pricing.cost(model: "gpt-5.6-sol", tokens: openAIBaseTokens, at: july,
billingInputTokens: 272_001) - 6.85) < 0.000_001)
abs(Pricing.cost(model: "gpt-5.6-sol", tokens: openAIBaseTokens, at: september,
billingInputTokens: 272_000) - 2.94) < 0.000_001
&& abs(Pricing.cost(model: "gpt-5.6-sol", tokens: openAIBaseTokens, at: september,
billingInputTokens: 272_001) - 4.88) < 0.000_001)
check("GPT-6 exact rates and long context",
Pricing.priceIsExact(model: "gpt-6-astra")
&& Pricing.priceIsExact(model: "gpt-6-sol")
&& Pricing.priceIsExact(model: "gpt-6-luna")
&& abs(Pricing.cost(model: "gpt-6-astra", tokens: openAIBaseTokens,
at: september, billingInputTokens: 272_001) - 12.2) < 0.000_001
&& abs(Pricing.cost(model: "gpt-6-luna", tokens: openAIBaseTokens,
at: september, billingInputTokens: 272_001) - 0.122) < 0.000_001)
check("GPT prices cover current and historical IDs",
abs(Pricing.cost(model: "gpt-5.4-mini", tokens: TokenBreakdown(output: 1_000_000),
at: july) - 4.5) < 0.000_001
Expand All @@ -63,22 +77,20 @@ enum SelfTest {

// Regression: the family-keyword fallback used to funnel every Opus into 4.8, so a
// real `claude-opus-5` record was renamed and merged into the 4.8 row. Each tier
// must resolve to itself, and an unrecognized version to the newest — not a pinned
// older one, which is how this broke in the first place.
// must resolve to itself; an unrecognized version keeps its own approximate ID.
check("Opus 5 keeps its own identity",
Pricing.normalize(model: "claude-opus-5") == "anthropic/claude-opus-5"
&& Pricing.displayName(forCanonicalKey: "anthropic/claude-opus-5") == "Opus 5")
check("older Opus tiers still resolve to themselves",
Pricing.normalize(model: "claude-opus-4-8") == "anthropic/claude-opus-4-8"
&& Pricing.normalize(model: "claude-opus-4-6") == "anthropic/claude-opus-4-6")
check("dated Opus 5 variant resolves via the fallback",
check("dated Opus 5 variant resolves exactly",
Pricing.normalize(model: "claude-opus-5-20260315") == "anthropic/claude-opus-5")
check("unknown Opus/Sonnet versions resolve to the newest, not a pinned tier",
Pricing.normalize(model: "claude-opus-9") == "anthropic/claude-opus-5"
&& Pricing.normalize(model: "claude-sonnet-9") == "anthropic/claude-sonnet-5")
// "Unknown → newest" is only safe while every *known* tier is enumerated: Opus 4.1
// costs 3x the 4.5+ tiers, so falling through to Opus 5 would bill it at a third of
// its real rate. Mythos 5 is Fable-tier and would otherwise hit the $3/$15 fallback.
check("unknown Claude versions keep their identity and approximate marker",
Pricing.normalize(model: "claude-opus-9") == "claude-opus-9"
&& !Pricing.priceIsExact(model: "claude-opus-9")
&& Pricing.normalize(model: "claude-sonnet-9") == "claude-sonnet-9")
// Opus 4.1 costs 3x the 4.5+ tiers, so explicit historical rows matter.
check("off-tier Opus versions resolve to themselves, not the newest",
Pricing.normalize(model: "claude-opus-4-1-20250805") == "anthropic/claude-opus-4-1"
&& Pricing.normalize(model: "claude-opus-4-5-20251101") == "anthropic/claude-opus-4-5")
Expand All @@ -89,7 +101,7 @@ enum SelfTest {
&& abs(Pricing.cost(model: "claude-mythos-5", tokens: millionTokens, at: july)
- Pricing.cost(model: "claude-fable-5", tokens: millionTokens, at: july)) < 0.000_001)
check("bare family selectors mean the current model",
Pricing.normalize(model: "opus") == "anthropic/claude-opus-5"
Pricing.normalize(model: "opus") == "anthropic/claude-opus-5-5"
&& Pricing.normalize(model: "sonnet") == "anthropic/claude-sonnet-5")
// 1M each of input/output/cacheRead/cacheWrite at $5 / $25 / $0.5 / $6.25 = $36.75.
check("Opus 5 priced at the Opus tier, not the generic fallback",
Expand Down
12 changes: 7 additions & 5 deletions Sources/CodingBarCore/Aggregator.swift
Original file line number Diff line number Diff line change
Expand Up @@ -144,18 +144,20 @@ public enum Aggregator {
// contributing record priced via a family guess / fallback rate.
let home = FileManager.default.homeDirectoryForCurrentUser.path
func breakdown(from records: [RawRecord]) -> (models: [ModelStat], projects: [ProjectStat]) {
var modelMap: [String: (tokens: TokenBreakdown, cost: Double, exact: Bool)] = [:]
var modelMap: [String: (model: String, provider: Provider, tokens: TokenBreakdown, cost: Double, exact: Bool)] = [:]
for r in records {
let key = Pricing.normalize(model: r.model)
var entry = modelMap[key] ?? (tokens: TokenBreakdown(), cost: 0, exact: true)
let groupKey = r.provider.rawValue + "·" + key
var entry = modelMap[groupKey] ?? (model: key, provider: r.provider,
tokens: TokenBreakdown(), cost: 0, exact: true)
entry.tokens += r.tokens
entry.cost += recordCost(r)
entry.exact = entry.exact && Pricing.priceIsExact(model: r.model)
modelMap[key] = entry
modelMap[groupKey] = entry
}
let models: [ModelStat] = modelMap
.map { key, entry in
ModelStat(model: key, provider: Pricing.provider(forCanonicalKey: key),
.map { _, entry in
ModelStat(model: entry.model, provider: entry.provider,
tokens: entry.tokens, cost: entry.cost, pricedExact: entry.exact)
}
.sorted { $0.cost > $1.cost }
Expand Down
22 changes: 16 additions & 6 deletions Sources/CodingBarCore/ClaudeScanner.swift
Original file line number Diff line number Diff line change
Expand Up @@ -17,21 +17,31 @@ public enum ClaudeScanner {
return ([], [])
}

var seenIds = Set<String>()
var allRecords: [RawRecord] = []

let records = scanner.scan(directory: projectsDir) { fileURL in
parseFile(fileURL)
}
return deduplicate(records)
}

// Dedup by message.id across all files
/// Streamed assistant messages repeat the same id. The later record carries
/// the final output-token count and often the completed tool-use content.
static func deduplicate(_ records: [RawRecord]) -> (records: [RawRecord], seenIds: Set<String>) {
var seenIds = Set<String>()
var indexByID: [String: Int] = [:]
var allRecords: [RawRecord] = []
for record in records {
if let mid = record.messageId {
guard seenIds.insert(mid).inserted else { continue }
if let index = indexByID[mid] {
if record.tokens.output >= allRecords[index].tokens.output {
allRecords[index] = record
}
continue
}
indexByID[mid] = allRecords.count
seenIds.insert(mid)
}
allRecords.append(record)
}

return (allRecords, seenIds)
}

Expand Down
49 changes: 8 additions & 41 deletions Sources/CodingBarCore/Coach.swift
Original file line number Diff line number Diff line change
Expand Up @@ -2,20 +2,6 @@ import Foundation

enum Coach {

// Canonical keys for Opus and Haiku pricing families. Every Opus tier belongs here:
// a missing one doesn't degrade the tip, it silently excludes that model's turns from
// the count entirely, so the advice goes quiet exactly when a new Opus becomes the
// model people actually run.
private static let opusKeys: Set<String> = [
"anthropic/claude-opus-5",
"anthropic/claude-opus-4-8",
"anthropic/claude-opus-4-7",
"anthropic/claude-opus-4-6",
]
private static let haikuKeys: Set<String> = [
"anthropic/claude-haiku-4-5",
]

// A "simple" turn has zero or one tool call and fewer than 300 output tokens.
private static func isSimpleTurn(_ record: RawRecord) -> Bool {
record.toolNames.count <= 1 && record.tokens.output < 300
Expand All @@ -24,43 +10,24 @@ enum Coach {
static func opusOnSimpleTip(from todayRecords: [RawRecord], language: AppLanguage) -> Insight? {
let claudeToday = todayRecords.filter { $0.provider == .claude }

// Cost delta: only count the non-cached input (cache tokens are already cheap
// regardless of model — switching models won't help much there).
var totalSimpleNetInput = 0 // non-cached input tokens only
var totalSimpleCacheRead = 0 // cache-read tokens (priced differently)
var totalSimpleOutput = 0
// Compare each turn at its actual Opus rate; cache writes are excluded
// because switching models would recreate the cache rather than reuse it.
var totalSaved = 0.0
var count = 0

for r in claudeToday {
let key = Pricing.normalize(model: r.model)
guard opusKeys.contains(key) else { continue }
guard key.hasPrefix("anthropic/claude-opus-"), Pricing.priceIsExact(model: r.model) else { continue }
guard isSimpleTurn(r) else { continue }
totalSimpleNetInput += r.tokens.input
totalSimpleCacheRead += r.tokens.cacheRead
totalSimpleOutput += r.tokens.output
let compared = TokenBreakdown(input: r.tokens.input, output: r.tokens.output,
cacheRead: r.tokens.cacheRead)
totalSaved += Pricing.cost(model: r.model, tokens: compared, at: r.timestamp)
- Pricing.cost(model: "claude-haiku-4-5", tokens: compared, at: r.timestamp)
count += 1
}

guard count >= 3 else { return nil } // not enough to matter

// Price the delta off the current Opus, not a pinned older one. Identical numbers
// today (both tiers are $5/$25), but this is what keeps the saving honest the next
// time the tiers diverge.
let opusKey = "anthropic/claude-opus-5"
let haikuKey = "anthropic/claude-haiku-4-5"
let opusInputPrice = Pricing.inputPrice(forCanonicalKey: opusKey)
let haikuInputPrice = Pricing.inputPrice(forCanonicalKey: haikuKey)
// Cache read price delta is small; include it for completeness
let opusCacheReadPrice = Pricing.cacheReadPrice(forCanonicalKey: opusKey)
let haikuCacheReadPrice = Pricing.cacheReadPrice(forCanonicalKey: haikuKey)
let opusOutputPricePerM = 25.0 // USD/1M
let haikuOutputPricePerM = 5.0 // USD/1M

let savedInput = Double(totalSimpleNetInput) * (opusInputPrice - haikuInputPrice) / 1_000_000
let savedCacheRead = Double(totalSimpleCacheRead) * (opusCacheReadPrice - haikuCacheReadPrice) / 1_000_000
let savedOutput = Double(totalSimpleOutput) * (opusOutputPricePerM - haikuOutputPricePerM) / 1_000_000
let totalSaved = savedInput + savedCacheRead + savedOutput

guard totalSaved >= 0.2 else { return nil }

let saved = String(format: "%.2f", totalSaved)
Expand Down
Loading
Loading