Skip to content
BayetoThornbury Fixed Income (demo)
Demo — fictional firm,
real engine output
Analyze your own system
All recommendations
Heldobjective: highest qualityFrontier model on routine workcomplexity Low

Right-size Client Reporting: move low-output claude-sonnet traffic to claude-haiku

Route the low-output claude-sonnet traffic in Client Reporting down one tier to claude-haiku

Not led with — declared-critical surface. You've declared this surface critical and the quality impact is unmeasured — so Bayeto won't recommend downgrading it. Declaring a quality evaluation for this candidate (Declared context) is what changes that.

Verify before implementing — this recommendation rests on evidence the engine does not have:

  • No quality evaluation on this segment — the downgrade's quality impact is a risk estimate, not a measurement.
  • No independent quality evaluation — the quality impact is unmeasured.

Implement only if these checks pass. The full self-critique is in the reasoning trail below.

Annual value€1,457direct bill savings
Payback period56d≈€220 to implement · Low effort
Confidence65%evidence Medium
Sample size14,400requests observed

Payback divides by an implementation cost of €220 under the agent-assisted-v1-2026-08 cost model: 2h of human review at €100/h plus a €20 agent-run allowance — an AI agent implements, your team reviews and verifies. Implemented by hand instead, the same move is estimated at 2 engineer-days (≈€1,600). Every figure is a prior, not a measurement of your team — and the model is versioned so a changed assumption can never move a figure silently.

Context completeness

Confidence 65% is how sure the engine is on the evidence it has. This is how much of the evidence that could change the conclusion the engine could establish for this application — a different question. The declare-context page counts what you have answered workspace-wide, which can legitimately read differently: an “I don't know” is an answer there and establishes nothing here.

4/6
dimensions present
  • Telemetry
  • Provider bill— calibrated from the monthly spend declared for this workspace, not a reconciled invoice
  • Per-app criticality
  • Prior experiments
  • Codebase
  • Quality evaluation

We cannot detect Codebase, Quality evaluation — you can tell us, and confidence is re-derived on the fuller evidence.

Observation → Evidence → Inference → Recommendation

every fact, judgment, self-critique and failure condition behind this decision
1

Observations (facts)

  • [avg_input_tokens] Client Reporting sends 3578 input tokens on average vs 797 output. (N=14,400)
2

Evidence (objective-relevant)

  • 94.5% of Client Reporting spend runs on claude-sonnet; output is 55.7% of this segment's cost (input 42.9% · cache-read 0.7% · cache-write 0.7% · output 55.7%) — a material share, so the small answers signal over-provisioning. (strength 0.5)
3

Inference (judgment)

Output is a material share of this segment's cost and the answers are tiny — the premium tier is over-provisioned; a one-tier downgrade is cost-positive after netting cache re-warming, with unmeasured quality risk.

4

Recommendation (move)

Route the low-output claude-sonnet traffic in Client Reporting down one tier to claude-haiku

Client Reporting: all traffic on the current premium modelClient Reporting: Route the low-output claude-sonnet traffic in Client Reporting down one tier to claude-haiku
5

Critical Review (self-critique)

Assumptions

  • Assumes claude-haiku preserves answer quality for this workload — quality is UNMEASURED here.
  • Assumes the small answers reflect task simplicity; output is 55.7% of this segment's cost (a material share), so the tier premium is genuinely being paid on cheap work.
  • Assumes the segment is cleanly switchable in full; the €1457 is net of a cache re-warm amortized over the observed reuse (12.5×).

Alternative explanations

  • The large prompt may be an intentional cached system prefix (architecture), not model over-provisioning — the ratio alone can't tell these apart.
  • The cost is spread across token types; a caching change would save little here.

Missing evidence

  • No quality evaluation on this segment — the downgrade's quality impact is a risk estimate, not a measurement.
  • No independent quality evaluation — the quality impact is unmeasured.

This is wrong if…

  • Wrong if this workload needs claude-sonnet-tier reasoning for correctness (e.g. a compliance-critical answer).
  • Wrong if the segment can't be switched as a whole (per-turn routing) — cache-thrash would erase the saving.
  • Wrong if this workload's real quality bar is higher than assumed and the change degrades answers.

Download the canonical decision record (JSON) — the deterministic decision block (byte-pinned by CI), the run context (engine + catalog versions, measured-window boundaries, exact pricing rows, fixed FX with effective date, formula inputs, the overlap group, and the ledger version), and the LLM-phrased narrative as a visibly separate section.

Business case & technical detail

the phrased explanation — the decision above is computed without it

Business case

Client Reporting pays claude-sonnet rates to produce tiny answers where output is a material 55.7% of cost — the hallmark of an over-provisioned model. Right-sizing to claude-haiku keeps the same work at a lower tier. Estimated €1,457/yr at the current run-rate, payback 56 days, confidence 65%.

Technical detail

Route this segment to claude-haiku. Expected quality change is UNKNOWN — not measured; validate with a canary / Simulate on a traffic slice before full rollout, and roll back if graded quality drops. Prompt caches are per-model, so switch the WHOLE segment at once (never per-turn) to avoid cache-thrash. Affects ~16,140 req/mo; expected cost -46.6%, latency +0%, quality unmeasured (medium risk).

Benchmark context & Optimization Memory

your value against the cohort, and what prior outcomes taught the engine

Optimization Memory

No prior implementations yet — confidence derived from evidence only.

Why NOT the alternatives?

each rejected route, with the constraint that rejected it
  • Cut cache-write cost instead (keep claude-sonnet)
    cost-0.3%latency+0%quality+0%Cache-writes are a small share here, so a caching change saves little.
  • Keep claude-sonnet (no change)
    cost+0%latency+0%quality+0%Quality-safe but forgoes the saving; justified only if this traffic genuinely needs top-tier reasoning.

Implementation packet

advisory_onlyrisk: model-substitution

Reviewed by Bayeto engineering on 2026-07-26 · rule version v1 · 94% of the €56,350/yr open on this workspace is carried by a packet an engineer can execute (5 of 8 open moves); 3 of 14 specs meet that standard. Missing mandatory for this risk class: guardrails, stopIf, rollback, qualityValidation.

Review attests that these steps are sound and reversible as written. It does not guarantee an outcome: whether the change saves money is measured by the validation loop afterwards, never promised by the reviewer.

Not reviewed to executable standard

Guardrails — watch these, Stop if, Rollback, Quality validation are all unwritten for this move, for one reason: Not reviewed to executable standard. The move is inferred from cost and token shape, and Bayeto has no view of task OUTCOMES, so it cannot state what quality bar the cheaper model must clear — nor assemble a comparison, since it stores aggregates rather than request bodies.

What you would need to decide

Bayeto withholds steps here because these are your calls to make, not ours. Answer them and this move can be specified.

  • What quality bar must the cheaper model clear for this task, expressed as something you can measure?

    Bayeto infers that the model is overpowered from cost and token shape. It has no view of task outcomes, so it cannot tell you what 'still good enough' means here.

  • Do you hold a held-out set of real requests to compare the two models on?

    Bayeto stores aggregates, not request bodies, so it cannot assemble the comparison sample — and a comparison on invented inputs would grade the wrong thing.

  • Preconditions1 specified instruction (mandatory for this risk class)
  • Steps1 specified instruction (mandatory for this risk class)
  • Guardrails — watch thesenone specified (mandatory for this risk class)withheld — for the reason stated above
  • Stop ifnone specified (mandatory for this risk class)withheld — for the reason stated above
  • Quality validationnone specified (mandatory for this risk class)withheld — for the reason stated above
  • Rollbacknone specified (mandatory for this risk class)withheld — for the reason stated above

Instruction text is withheld here: the public demo shows complete packets only for the moves reviewed to the executable standard — 3 of 14 specs — as its worked examples; the rest summarize what exists and what was deliberately not specified.

Simulate — Before You Build

Projected annual savings€1,304scenario realization (sim-v1 assumption) at 65% confidence
Downside / base / upside€1,020 · €1,304 · €1,457realization = 0.7 + 0.3 × 65% = 0.895 · base €1,457 · sim-v1
cost-46.6%latency (prior)+0%qualityunmeasured (medium risk)

The cost delta derives from this workspace’s telemetry. The latency and quality deltas are heuristic priors carried by the rule: no quality evaluation is connected, and no rule measures the latency effect of its own move — the figures are per-rule constants, so nothing here observes either outcome.

  • · Applies to ~16,140 requests/month at current mix.
  • · Scenario realization assumption (sim-v1, not a statistical expected value): 0.7 + 0.3 × confidence = 0.895. The 0.7 floor encodes implementation shortfall — even a certain move rarely captures 100% at rollout; rationale published at bayeto.ai/methodology.
  • · Per-request cost/latency/quality deltas hold at current traffic shape.
© 2026 Bayeto