Optimization Roadmap
The loop: Hypothesis → Implement → Observe → Validate → Score ↑. Every captured win — validated here or measured from a change you already made — locks in savings and unlocks the next.
Spend timeline
Daily provider spend from your telemetry (uncalibrated cost basis), in USD — the currency your provider bills in. The dashed marker is the measured migration; the guide lines are the average daily spend before and after it.
The annotated run-rate is engine-derived business value in EUR: USD→EUR at 0.92 (fixed reference rate, effective 2026-07-06).
Measured wins — annualized run-rate
gpt-4-turbo → claude-sonnet · since 2026-06-06
Moved gpt-4-turbo→claude-sonnet ~24d ago; repriced at gpt-4-turbo rates = $22,071.51 vs $6,884.78 actual (USD) → run-rate ≈ €244,361/yr EUR (USD→EUR at 0.92 (fixed reference rate, effective 2026-07-06); medium confidence, 24d). Calibrated to the monthly spend you declared (×1.15).
Detected from telemetry via a repricing counterfactual (your actual recent tokens priced at the old model vs the new) — a run-rate from a short window, cost-only, quality not measured.
Before you switched gpt-4-turbo → claude-sonnet, Bayeto predicted €0.203/call. You measured €0.221/call.
Compared per call, so the score reflects the estimate and not traffic growth. Annualized for context: €206,904/yr at the pre-switch run-rate vs €244,361/yr measured after — volume moved -2%.
accuracy = 1 − |0.203 − 0.221| / 0.221 = 91.9% · signed error €-0.0179/call (under-predicted)
Denominator: the measured per-call saving; over- and under-prediction count symmetrically (absolute error). Graded population: client-reporting, compliance-assistant, credit-memo, research-copilot · 65,408 post-switch calls · formula pvm-v1.
- Unlockable
Finish the gpt-4-turbo → claude-sonnet migration: Meeting Summaries missed the rollout
unlocks €40,940/yr · open hypothesis
75 → 77score would +2 - 77 → 77score would ±0
- 77 → 77score would ±0
- Unlockable
Route small-context Portfolio Analytics traffic off claude-sonnet to GLM-5.2
unlocks €567/yr · open hypothesis
77 → 77score would ±0 - 77 → 77score would ±0
- 77 → 77score would ±0
- Unlockable
Route small-context Python Quant traffic off claude-sonnet to GLM-5.2
unlocks €320/yr · open hypothesis
77 → 77score would ±0 - 77 → 77score would ±0
Target 77 — every step above implemented. New telemetry surfaces new hypotheses; the target rises with them.
Held — not offered as wins
Declaring these surfaces would release €18,883/yr. They carry an unmeasured quality risk on a surface that is declared critical, or is not yet cleared. Declaring a surface low-stakes releases its move; declaring it critical keeps it blocked. Where a move competes with one already leading, the figure shown is what you would gain by swapping — the two are never additive.
Right-size Research Copilot: move low-output claude-sonnet traffic to claude-haiku
You declared this surface critical and the quality impact is unmeasured — Bayeto won't recommend cheapening it.
Declaring this surface would release €6,965/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Right-size Credit Memo: move low-output claude-sonnet traffic to claude-haiku
Carries an unmeasured quality risk on a surface whose stakes aren't cleared. Declare it low-stakes to release it.
Declaring this surface would release €6,084/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Right-size Compliance Assistant: move low-output claude-sonnet traffic to claude-haiku
You declared this surface critical and the quality impact is unmeasured — Bayeto won't recommend cheapening it.
Declaring this surface would release €3,426/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Compress the un-cached context Research Copilot re-sends every call
You declared this surface critical and the quality impact is unmeasured — Bayeto won't recommend cheapening it.
Declaring this surface would release €633/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Compress the un-cached context Credit Memo re-sends every call
Carries an unmeasured quality risk on a surface whose stakes aren't cleared. Declare it low-stakes to release it.
Declaring this surface would release €1,483/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Compress the un-cached context Compliance Assistant re-sends every call
You declared this surface critical and the quality impact is unmeasured — Bayeto won't recommend cheapening it.
Declaring this surface would release €291/yr — the net gain over the move already leading this application's spend, which it replaces rather than adds to.
Right-size Client Reporting: move low-output claude-sonnet traffic to claude-haiku
You declared this surface critical and the quality impact is unmeasured — Bayeto won't recommend cheapening it.
Clearing this releases nothing: a larger move already leads this application’s spend, and the two are mutually exclusive.
ledger version 22c35e346df765c3 — identical on every page, export, and decision record of this workspace state · how to verify it