Skip to content
BayetoThornbury Fixed Income (demo)
Demo — fictional firm,
real engine output
Analyze your own system
All recommendations
Heldobjective: lowest costFrontier model on routine workcomplexity Low

Right-size Credit Memo: move low-output claude-sonnet traffic to claude-haiku

Route the low-output claude-sonnet traffic in Credit Memo down one tier to claude-haiku

Not led with — review needed — stakes not declared. This carries an unmeasured quality risk and you haven't told us how critical this surface is. On telemetry alone Bayeto flags it for review rather than leading with it. Declare the app critical and it's held; declare it low-stakes and it becomes a normal recommendation.

Verify before implementing — this recommendation rests on evidence the engine does not have:

  • No quality evaluation on this segment — the downgrade's quality impact is a risk estimate, not a measurement.
  • Application criticality not declared — treated conservatively.
  • No independent quality evaluation — the quality impact is unmeasured.

Implement only if these checks pass. The full self-critique is in the reasoning trail below.

Annual value€8,255direct bill savings
Payback period10d≈€220 to implement · Low effort
Confidence60%evidence Low
Sample size16,320requests observed

Payback divides by an implementation cost of €220 under the agent-assisted-v1-2026-08 cost model: 2h of human review at €100/h plus a €20 agent-run allowance — an AI agent implements, your team reviews and verifies. Implemented by hand instead, the same move is estimated at 2 engineer-days (≈€1,600). Every figure is a prior, not a measurement of your team — and the model is versioned so a changed assumption can never move a figure silently.

Context completeness

Confidence 60% is how sure the engine is on the evidence it has. This is how much of the evidence that could change the conclusion the engine could establish for this application — a different question. The declare-context page counts what you have answered workspace-wide, which can legitimately read differently: an “I don't know” is an answer there and establishes nothing here.

3/6
dimensions present
  • Telemetry
  • Provider bill— calibrated from the monthly spend declared for this workspace, not a reconciled invoice
  • Prior experiments
  • Per-app criticality
  • Codebase
  • Quality evaluation

We cannot detect Per-app criticality, Codebase, Quality evaluation — you can tell us, and confidence is re-derived on the fuller evidence.

Observation → Evidence → Inference → Recommendation

every fact, judgment, self-critique and failure condition behind this decision
1

Observations (facts)

  • [avg_input_tokens] Credit Memo sends 29254 input tokens on average vs 1051 output. (N=16,320)
2

Evidence (objective-relevant)

  • 99.4% of Credit Memo spend runs on claude-sonnet, but the prompt side (input + any cached prefix) dominates cost (output only 15.3%; input 84.5% · cache-read 0.1% · cache-write 0.1% · output 15.3%) — the output:input ratio is NOT a complexity signal here. The case rests on claude-haiku being cheaper for the SAME token mix, net of any cache re-warming. (strength 0.35)
3

Inference (judgment)

The prompt side dominates this segment's cost, so the output:input ratio is not a complexity signal — but a full-segment downgrade to claude-haiku is a clean, cache-aware cost win; quality risk is unmeasured.

4

Recommendation (move)

Route the low-output claude-sonnet traffic in Credit Memo down one tier to claude-haiku

Credit Memo: all traffic on the current premium modelCredit Memo: Route the low-output claude-sonnet traffic in Credit Memo down one tier to claude-haiku
5

Critical Review (self-critique)

Assumptions

  • Assumes claude-haiku preserves answer quality for this workload — quality is UNMEASURED here.
  • Does NOT use the output:input ratio as a complexity signal — the prompt side dominates cost (output 15.3%); the case rests only on claude-haiku being cheaper for the same token mix.
  • Assumes the segment is cleanly switchable in full; the €8255 is net of a cache re-warm amortized over the observed reuse (12.3×).

Alternative explanations

  • The large prompt may be an intentional cached system prefix (architecture), not model over-provisioning — the ratio alone can't tell these apart.
  • The cost is spread across token types; a caching change would save little here.

Missing evidence

  • No quality evaluation on this segment — the downgrade's quality impact is a risk estimate, not a measurement.
  • Application criticality not declared — treated conservatively.
  • No independent quality evaluation — the quality impact is unmeasured.

This is wrong if…

  • Wrong if this workload needs claude-sonnet-tier reasoning for correctness (e.g. a compliance-critical answer).
  • Wrong if the segment can't be switched as a whole (per-turn routing) — cache-thrash would erase the saving.
  • Wrong if this workload's real quality bar is higher than assumed and the change degrades answers.

Download the canonical decision record (JSON) — the deterministic decision block (byte-pinned by CI), the run context (engine + catalog versions, measured-window boundaries, exact pricing rows, fixed FX with effective date, formula inputs, the overlap group, and the ledger version), and the LLM-phrased narrative as a visibly separate section.

Business case & technical detail

the phrased explanation — the decision above is computed without it

Business case

Credit Memo's claude-sonnet spend is dominated by the prompt side (input + any cached prefix), not output — so this is not "over-provisioned for complexity". But claude-haiku is cheaper across this exact token mix, and a full-segment switch nets a saving even after any cache re-warming on the new model. Estimated €8,255/yr at the current run-rate, payback 10 days, confidence 60%.

Technical detail

Route this segment to claude-haiku. Expected quality change is UNKNOWN — not measured; validate with a canary / Simulate on a traffic slice before full rollout, and roll back if graded quality drops. Prompt caches are per-model, so switch the WHOLE segment at once (never per-turn) to avoid cache-thrash. Affects ~18,743 req/mo; expected cost -49.6%, latency +0%, quality unmeasured (medium risk).

Benchmark context & Optimization Memory

your value against the cohort, and what prior outcomes taught the engine

Optimization Memory

No prior implementations yet — confidence derived from evidence only.

Why NOT the alternatives?

each rejected route, with the constraint that rejected it
  • Cut cache-write cost instead (keep claude-sonnet)
    cost+0%latency+0%quality+0%Cache-writes are a small share here, so a caching change saves little.
  • Keep claude-sonnet (no change)
    cost+0%latency+0%quality+0%Quality-safe but forgoes the saving; justified only if this traffic genuinely needs top-tier reasoning.

Implementation packet

advisory_onlyrisk: model-substitution

Reviewed by Bayeto engineering on 2026-07-26 · rule version v1 · 94% of the €56,350/yr open on this workspace is carried by a packet an engineer can execute (5 of 8 open moves); 3 of 14 specs meet that standard. Missing mandatory for this risk class: guardrails, stopIf, rollback, qualityValidation.

Review attests that these steps are sound and reversible as written. It does not guarantee an outcome: whether the change saves money is measured by the validation loop afterwards, never promised by the reviewer.

Not reviewed to executable standard

Guardrails — watch these, Stop if, Rollback, Quality validation are all unwritten for this move, for one reason: Not reviewed to executable standard. The move is inferred from cost and token shape, and Bayeto has no view of task OUTCOMES, so it cannot state what quality bar the cheaper model must clear — nor assemble a comparison, since it stores aggregates rather than request bodies.

What you would need to decide

Bayeto withholds steps here because these are your calls to make, not ours. Answer them and this move can be specified.

  • What quality bar must the cheaper model clear for this task, expressed as something you can measure?

    Bayeto infers that the model is overpowered from cost and token shape. It has no view of task outcomes, so it cannot tell you what 'still good enough' means here.

  • Do you hold a held-out set of real requests to compare the two models on?

    Bayeto stores aggregates, not request bodies, so it cannot assemble the comparison sample — and a comparison on invented inputs would grade the wrong thing.

  • Preconditions1 specified instruction (mandatory for this risk class)
  • Steps1 specified instruction (mandatory for this risk class)
  • Guardrails — watch thesenone specified (mandatory for this risk class)withheld — for the reason stated above
  • Stop ifnone specified (mandatory for this risk class)withheld — for the reason stated above
  • Quality validationnone specified (mandatory for this risk class)withheld — for the reason stated above
  • Rollbacknone specified (mandatory for this risk class)withheld — for the reason stated above

Instruction text is withheld here: the public demo shows complete packets only for the moves reviewed to the executable standard — 3 of 14 specs — as its worked examples; the rest summarize what exists and what was deliberately not specified.

Simulate — Before You Build

Projected annual savings€7,259scenario realization (sim-v1 assumption) at 60% confidence
Downside / base / upside€5,779 · €7,259 · €8,255realization = 0.7 + 0.3 × 60% = 0.879 · base €8,255 · sim-v1
cost-49.6%latency (prior)+0%qualityunmeasured (medium risk)

The cost delta derives from this workspace’s telemetry. The latency and quality deltas are heuristic priors carried by the rule: no quality evaluation is connected, and no rule measures the latency effect of its own move — the figures are per-rule constants, so nothing here observes either outcome.

  • · Applies to ~18,743 requests/month at current mix.
  • · Scenario realization assumption (sim-v1, not a statistical expected value): 0.7 + 0.3 × confidence = 0.879. The 0.7 floor encodes implementation shortfall — even a certain move rarely captures 100% at rollout; rationale published at bayeto.ai/methodology.
  • · Per-request cost/latency/quality deltas hold at current traffic shape.
© 2026 Bayeto