Skip to content
BayetoThornbury Fixed Income (demo)
Demo — fictional firm,
real engine output
Analyze your own system
All recommendations
Outstandingobjective: lowest costUnfinished migrationcomplexity Low

Finish the gpt-4-turbo → claude-sonnet migration: Meeting Summaries missed the rollout

Route Meeting Summaries's gpt-4-turbo traffic to claude-sonnet

Verify before implementing — this recommendation rests on evidence the engine does not have:

  • No per-app quality evaluation on this surface; sibling history is the (strong) proxy.
  • No independent quality evaluation — the quality impact is unmeasured.

Implement only if these checks pass. The full self-critique is in the reasoning trail below.

Annual value€40,940direct bill savings
Payback period2d≈€220 to implement · Low effort
Confidence72%evidence High
Sample size276,000requests observed

Payback divides by an implementation cost of €220 under the agent-assisted-v1-2026-08 cost model: 2h of human review at €100/h plus a €20 agent-run allowance — an AI agent implements, your team reviews and verifies. Implemented by hand instead, the same move is estimated at 2 engineer-days (≈€1,600). Every figure is a prior, not a measurement of your team — and the model is versioned so a changed assumption can never move a figure silently.

Context completeness

Confidence 72% is how sure the engine is on the evidence it has. This is how much of the evidence that could change the conclusion the engine could establish for this application — a different question. The declare-context page counts what you have answered workspace-wide, which can legitimately read differently: an “I don't know” is an answer there and establishes nothing here.

4/6
dimensions present
  • Telemetry
  • Provider bill— calibrated from the monthly spend declared for this workspace, not a reconciled invoice
  • Per-app criticality
  • Prior experiments
  • Codebase
  • Quality evaluation

We cannot detect Codebase, Quality evaluation — you can tell us, and confidence is re-derived on the fuller evidence.

Observation → Evidence → Inference → Recommendation

every fact, judgment, self-critique and failure condition behind this decision
1

Observations (facts)

  • [measured_win] Measured impact — you moved gpt-4-turbo→claude-sonnet ~24d ago. Repricing the last 24d of ACTUAL work at gpt-4-turbo rates = $22,071.51 vs $6,884.78 actual on claude-sonnet (both USD) → realized RUN-RATE ≈ €212,488/yr EUR (USD→EUR at 0.92 (fixed reference rate, effective 2026-07-06)), based on 24d post-switch (medium confidence — tightens as data accrues; cost-only, quality not measured). Period per-call context (confoundable, USD): $0.1634 → $0.0552. Calibrated to the reconciled bill (×1.15) ⇒ €244,361/yr. (N=276,000)
  • [monthly_spend_usd] Meeting Summaries spends about $4,051/mo across 1 models. (N=24,000)
2

Evidence (objective-relevant)

  • Meeting Summaries still runs 100% of its spend on gpt-4-turbo — the family the rest of the fleet left. The identical gpt-4-turbo → claude-sonnet move was MEASURED across 6 sibling applications since 2026-06-06 (24 days, medium confidence): this surface missed the rollout. (strength 0.705)
3

Inference (judgment)

Meeting Summaries missed a migration the fleet already validated; finishing it captures measured — not assumed — economics.

4

Recommendation (move)

Route Meeting Summaries's gpt-4-turbo traffic to claude-sonnet

Meeting Summaries: all traffic on the current premium modelMeeting Summaries: Route Meeting Summaries's gpt-4-turbo traffic to claude-sonnet
5

Critical Review (self-critique)

Assumptions

  • Assumes Meeting Summaries's workload is comparable to the 6 migrated siblings — the evidence is on siblings, not this app.

Alternative explanations

  • The app may have been deliberately exempted from the rollout (a known dependency on the prior family) — confirm with the owning team.

Missing evidence

  • No per-app quality evaluation on this surface; sibling history is the (strong) proxy.
  • No independent quality evaluation — the quality impact is unmeasured.

This is wrong if…

  • Wrong if intentionally exempted; wrong if its prompt mix differs materially from the migrated fleet.
  • Wrong if this workload's real quality bar is higher than assumed and the change degrades answers.

Download the canonical decision record (JSON) — the deterministic decision block (byte-pinned by CI), the run context (engine + catalog versions, measured-window boundaries, exact pricing rows, fixed FX with effective date, formula inputs, the overlap group, and the ledger version), and the LLM-phrased narrative as a visibly separate section.

Business case & technical detail

the phrased explanation — the decision above is computed without it

Business case

You already made this exact move fleet-wide and Bayeto measured it since 2026-06-06 (24 days of post-switch data, medium confidence). Meeting Summaries still runs 100% of its spend on gpt-4-turbo. Completing the migration you already validated captures ≈ €40,940/yr — the quality risk is bounded by your own post-switch history, not an assumption. Estimated €40,940/yr at the current run-rate, payback 2 days, confidence 72%.

Technical detail

Switch the whole segment at once (per-model caches — a partial switch thrashes the prefix cache). Reuse the acceptance signals from the migrated siblings; the rollback path is already known from the fleet migration. Affects ~30,000 req/mo; expected cost -78.5%, latency +0%, quality unmeasured (low risk).

Benchmark context & Optimization Memory

your value against the cohort, and what prior outcomes taught the engine

Optimization Memory

No prior implementations yet — confidence derived from evidence only.

Why NOT the alternatives?

each rejected route, with the constraint that rejected it
  • Drop two tiers instead
    cost-50%latency+0%qualityunmeasured (medium risk)A deeper cut, but NOT what was validated — the measured evidence covers gpt-4-turbo → claude-sonnet, nothing further.
  • Keep gpt-4-turbo here
    cost+0%latency+0%quality+0%Forgoes a saving your own fleet already validated. Only right if this app was deliberately exempted from the rollout — confirm with the owning team.
  • Migrate per-request instead of the whole segment
    cost-39.3%latency+0%qualityunmeasured (low risk)A partial switch splits the per-model prompt cache — every call re-writes the prefix instead of reading it. Whole-segment is the clean move (and what the siblings did).

Implementation packet

executablerisk: model-substitution

Reviewed by Bayeto engineering on 2026-08-15 · rule version v1 · 94% of the €56,350/yr open on this workspace is carried by a packet an engineer can execute (5 of 8 open moves); 3 of 14 specs meet that standard.

Review attests that these steps are sound and reversible as written. It does not guarantee an outcome: whether the change saves money is measured by the validation loop afterwards, never promised by the reviewer.

Complete packet, fictional system

Shown in full on the public demo because this move is reviewed to the executable standard — the one worked example a buyer can evaluate. Every figure derives from the fictional Thornbury dataset and its pinned scenario catalog: synthetic, for judging the packet's form, not a customer result.

Finish a migration this tenant has already performed and measured. The rest of the fleet moved off the prior family; this application did not, and it is still paying the prior family's rates for the same work. The quality argument is NOT that the two models are equivalent — it is that this customer already accepted the switch elsewhere and their own post-switch telemetry is the evidence. What that evidence cannot do is speak for THIS workload, which is why every step below is a canary before a cutover.

Preconditions

  • MEASURED ON YOUR FLEET, NOT ON THIS WORKLOAD. The identical gpt-4-turbo to claude-sonnet switch has been running on 6 sibling applications (client-reporting, compliance-assistant, credit-memo, portfolio-analytics, python-quant, research-copilot) since 2026-06-06 — 24 days and 124,800 calls of post-switch telemetry, from which your own data measured 244,361 EUR/yr. That evidence is why this move is specified for execution while other model substitutions are not. It was NOT measured on this application: no call from this workload took part in it, which is the precondition that identified the workload as a laggard in the first place.

    calculatedsnapshot.realizedWins[] (fromFamily/toFamily, byApplication, windowDays)

  • This application runs dominantly on the prior family, which is what makes a clean switch isolable rather than a partial rewrite.

    calculatedrollup model slices vs MIN_LAGGARD_SHARE

  • The target family is priced in the pinned catalog, so the saving is a reprice at published rates rather than an estimate.

    directcatalog.ratesFor(win.toFamily)

  • VERIFY BEFORE MIGRATING; Bayeto does not observe prompt content. Confirm this workload's prompts, tool definitions and output parsers are not pinned to the prior family's quirks — a response format, a tokenizer assumption or a system-prompt idiom that the siblings did not depend on will fail here and the sibling evidence will not have caught it.

    human-authoredreviewer instruction - prompt content is never sent to Bayeto

  • VERIFY: no contractual, residency or certification commitment pins this workload to the prior family. The siblings' migration is not evidence about this workload's obligations, which Bayeto cannot see.

    human-authoredreviewer instruction

  • VERIFY: the workload has a quality metric you can run before and after on a held-out sample. Without one there is no abort condition below that can fire on quality, and this packet should not be executed.

    human-authoredreviewer instruction

Steps

  • Route a SMALL share of this application's traffic to the target family, leaving the rest on the prior one. A canary, not a cutover — the measurement you are relying on was taken on other applications.

    human-authoredreviewer instruction

  • Run the workload's own quality metric over the canary output and the held-out baseline, and compare. Do this before widening, not after.

    human-authoredcustomer-supplied quality metric

  • Compare realised cost PER CALL between the canary and the baseline, not monthly totals — traffic volume moves between windows and would mask the effect either way.

    calculatedtelemetry.costUsd / call counts

  • Widen in stages once the canary holds, keeping the prior family reachable at every stage.

    human-authoredreviewer instruction

  • Expect the first requests on the target family to pay cache re-warming: any prompt-prefix cache warmed on the prior family is cold on the new one. The recommendation's own figure is already net of this, and it is a one-off, not a regression.

    calculatedcacheWarmingPenaltyUsd (netted into the recommendation)

Guardrails — watch these

  • The workload's own quality metric, canary vs baseline. This is the guardrail the sibling evidence cannot substitute for.

    human-authoredcustomer-supplied quality metric

  • Request error rate, before vs after.

    directtelemetry.status

  • Realised cost per call, before vs after — the saving must appear in your own telemetry, not only in the reprice.

    calculatedtelemetry.costUsd / call counts

  • OUTPUT LENGTH DISTRIBUTION. A different model can be systematically more verbose at the same prompt, which turns a cheaper per-token rate into a more expensive call. The reprice assumes this workload's observed token mix carries over; this is the guardrail that catches it when it does not.

    calculatedtelemetry.outputTokens distribution

  • End-to-end latency distribution.

    directtelemetry.latencyMs

  • Retry and fallback rate. A model that fails a schema or a parser more often can hold quality and error rate flat while costing more in retries.

    directtelemetry.status / retry markers

Stop if

  • The workload's quality metric falls below the threshold you set before starting.

    human-authoredcustomer-supplied quality metric

  • Request error rate rises, or the retry rate rises.

    directtelemetry.status

  • Realised cost per call does not fall in the canary. The sibling applications' measured saving is evidence about them; if it does not reproduce here, the reprice's assumption about this workload's token mix is what was wrong, and widening will not fix it.

    calculatedtelemetry.costUsd / call counts vs the measured sibling saving

  • Output length rises materially against the baseline, even if quality and error rate hold.

    calculatedtelemetry.outputTokens distribution

Quality validation

  • Run the workload's existing quality metric over a held-out sample on THIS application, before and after, and compare. The measured outcome on sibling applications is the reason this move is offered; it is not a substitute for this step.

    human-authoredcustomer-supplied quality metric

  • Do NOT accept the sibling migration as validation for this workload. It is evidence that the switch is survivable in this tenant's environment, on other traffic. The engine's own strength ceiling for this rule encodes the same limit, and treating fleet evidence as workload evidence is the specific error this packet exists to prevent.

    human-authoredreviewer instruction — capped at 0.75, measured on comparable moves

  • Compare on the SAME held-out sample for both arms. A canary judged against fresh traffic confounds the model change with whatever else moved that week.

    human-authoredreviewer instruction

Rollback

  • Route this application's traffic back to the prior family — a routing-and-configuration revert, no data migration and no prompt change, because the canary never removed the prior path. Two costs are real and both are one-off: any prompt-prefix cache warmed on the target family is discarded, and the prior family's cache may itself have gone cold during the canary, so the first requests after reverting pay re-warm writes.

    human-authoredreviewer instruction - canary keeps the prior path reachable

Simulate — Before You Build

Projected annual savings€37,522scenario realization (sim-v1 assumption) at 72% confidence
Downside / base / upside€28,658 · €37,522 · €40,940realization = 0.7 + 0.3 × 72% = 0.917 · base €40,940 · sim-v1
cost-78.5%latency (prior)+0%qualityunmeasured (low risk)

The cost delta derives from this workspace’s telemetry. The latency and quality deltas are heuristic priors carried by the rule: no quality evaluation is connected, and no rule measures the latency effect of its own move — the figures are per-rule constants, so nothing here observes either outcome.

  • · Applies to ~30,000 requests/month at current mix.
  • · Scenario realization assumption (sim-v1, not a statistical expected value): 0.7 + 0.3 × confidence = 0.917. The 0.7 floor encodes implementation shortfall — even a certain move rarely captures 100% at rollout; rationale published at bayeto.ai/methodology.
  • · Per-request cost/latency/quality deltas hold at current traffic shape.
© 2026 Bayeto