Objective lens — the root object, switchable
The same telemetry under a different objective yields different decisions. The default view evaluates each application under its own declared objective; a lens re-evaluates every application under one. Ranking always recomputes; a move is only rejected where a constraint can bind on your data — so the lens states which of its constraints can, and which cannot.
Active constraints: latency <= 1700ms · quality >= 0.75 · 3 recommendations survive this objective · 14 moves rejected as non-Pareto — e.g. “p95 latency would rise to 2910ms > cap 1700ms”.
- latency <= 1700 — can reject: highest observed p95 2912ms; worst projected move -5% → 2766ms
- quality >= 0.75 — cannot bind here: lowest measured quality 1; worst projected move 0% → 1
Recommendations
3 generated optimization hypotheses, ranked by business value under each application’s objective. Nothing here is seeded — the engine produces them from telemetry.
2 of 3 shown · filters active · Clear filters
Outstanding — ranked by value × confidence × objective alignment
Route small-context Portfolio Analytics traffic off claude-sonnet to GLM-5.2
Pin Portfolio Analytics serving to the caller’s region
Ruled out by your objective · 14 generated, then refused by a constraint — not in the count above · €128,653/yr what these would have saved, if your constraints allowed them
These moves would save money and your objective forbids them. Each says which constraint bound and by how much — change the constraint and they become available.
Right-size Client Reporting: move low-output claude-sonnet traffic to claude-haiku
would save €1,457/yr- p95 latency would rise to 2910ms > cap 1700ms
Extend prompt caching across Client Reporting
would save €708/yr- p95 latency would rise to 2765ms > cap 1700ms
Compress the un-cached context Compliance Assistant re-sends every call
would save €1,604/yr- p95 latency would rise to 2025ms > cap 1700ms
Right-size Compliance Assistant: move low-output claude-sonnet traffic to claude-haiku
would save €5,967/yr- p95 latency would rise to 2132ms > cap 1700ms
Extend prompt caching across Compliance Assistant
would save €1,313/yr- p95 latency would rise to 2025ms > cap 1700ms
Compress the un-cached context Credit Memo re-sends every call
would save €2,112/yr- p95 latency would rise to 2395ms > cap 1700ms
Right-size Credit Memo: move low-output claude-sonnet traffic to claude-haiku
would save €8,255/yr- p95 latency would rise to 2521ms > cap 1700ms
Extend prompt caching across Credit Memo
would save €629/yr- p95 latency would rise to 2395ms > cap 1700ms
Compress the un-cached context Meeting Summaries re-sends every call
would save €6,952/yr- p95 latency would rise to 2766ms > cap 1700ms
Right-size Meeting Summaries: move low-output gpt-4-turbo traffic to claude-sonnet
would save €40,940/yr- p95 latency would rise to 2912ms > cap 1700ms
Finish the gpt-4-turbo → claude-sonnet migration: Meeting Summaries missed the rollout
would save €40,940/yr- p95 latency would rise to 2912ms > cap 1700ms
Compress the un-cached context Research Copilot re-sends every call
would save €3,240/yr- p95 latency would rise to 2393ms > cap 1700ms
Right-size Research Copilot: move low-output claude-sonnet traffic to claude-haiku
would save €11,930/yr- p95 latency would rise to 2519ms > cap 1700ms
Extend prompt caching across Research Copilot
would save €2,606/yr- p95 latency would rise to 2393ms > cap 1700ms
Considered & not surfaced · 19 examined and never generated — not in the count above · €0/yr if the telemetry had supported them
Candidates the engine examined and declined — it shows its work on the moves it didn’t make, not only the ones it did.
Batch the offline Client Reporting jobs
would save €0/yrThe pinned catalog carries no batch tariff for Client Reporting's models — every schedule in it is synchronous, on-demand. Batch pricing is published per provider; Bayeto has not captured it, so there is no discount to quote and this move is not priced.
Cache economics for Client Reporting (claude-haiku)
would save €0/yrExamined Client Reporting's claude-haiku cache: reuse 12.4× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Client Reporting (claude-sonnet)
would save €0/yrExamined Client Reporting's claude-sonnet cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Response length for Client Reporting (claude-sonnet)
would save €0/yrExamined: output is 55.7% of this segment's cost, but responses average only 798 tokens — already short, so there is no padding to cap.
Cache economics for Compliance Assistant (claude-haiku)
would save €0/yrExamined Compliance Assistant's claude-haiku cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Compliance Assistant (claude-sonnet)
would save €0/yrExamined Compliance Assistant's claude-sonnet cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Credit Memo (claude-sonnet)
would save €0/yrExamined Credit Memo's claude-sonnet cache: reuse 12.3× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Credit Memo (gpt-4o-mini)
would save €0/yrCredit Memo's gpt-4o-mini bills cache-writes at or below the input rate ($0.15/1M vs $0.15/1M) — churn carries no premium, so there's nothing to fix.
Prompt caching for Portfolio Analytics
would save €0/yrPortfolio Analytics's prefix (~490 tokens) is below the provider's 1024-token minimum cacheable length — caching cannot engage.
Cache economics for Portfolio Analytics (claude-sonnet)
would save €0/yrExamined Portfolio Analytics's claude-sonnet cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Portfolio Analytics (gpt-4o-mini)
would save €0/yrPortfolio Analytics's gpt-4o-mini bills cache-writes at or below the input rate ($0.15/1M vs $0.15/1M) — churn carries no premium, so there's nothing to fix.
Response length for Portfolio Analytics (claude-sonnet)
would save €0/yrExamined: output is 65.3% of this segment's cost, but responses average only 190 tokens — already short, so there is no padding to cap.
Batch the offline Python Quant jobs
would save €0/yrThe pinned catalog carries no batch tariff for Python Quant's models — every schedule in it is synchronous, on-demand. Batch pricing is published per provider; Bayeto has not captured it, so there is no discount to quote and this move is not priced.
Prompt caching for Python Quant
would save €0/yrPython Quant's prefix (~289 tokens) is below the provider's 1024-token minimum cacheable length — caching cannot engage.
Cache economics for Python Quant (claude-sonnet)
would save €0/yrExamined Python Quant's claude-sonnet cache: reuse 8× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Python Quant (gpt-4o-mini)
would save €0/yrPython Quant's gpt-4o-mini bills cache-writes at or below the input rate ($0.15/1M vs $0.15/1M) — churn carries no premium, so there's nothing to fix.
Response length for Python Quant (claude-sonnet)
would save €0/yrExamined: output is 66.1% of this segment's cost, but responses average only 160 tokens — already short, so there is no padding to cap.
Cache economics for Research Copilot (claude-haiku)
would save €0/yrExamined Research Copilot's claude-haiku cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
Cache economics for Research Copilot (claude-sonnet)
would save €0/yrExamined Research Copilot's claude-sonnet cache: reuse 12.5× vs break-even 0.28× — the cache nets in your favor. Healthy; nothing to change.
ledger version 1b1a8afa88e54575 — identical on every page, export, and decision record of this workspace state · how to verify it