Skip to content
BayetoThornbury Fixed Income (demo)
Demo — fictional firm,
real engine output
Analyze your own system

Pilot report — predicted vs measured

Thornbury Fixed Income (demo) · 2 committed moves · outcomes measured from 124,800 logged calls over 24 days. Cost basis calibrated to the monthly provider spend you declared.

Every outcome below is re-derived from your telemetry on each run — never a stored figure, never self-reported. A move that cannot yet support a verdict says so.

Predicted — committed moves
€49,195/yr

What the engine predicted for the 2 moves you committed to, captured at the moment each was marked.

Measured — net regression
€73/yr added

The measured net across concluded moves is a cost increase. Reported as found.

€43,261/yr measured and NOT counted

1 concluded move landed in a window where something else also changed, so the movement cannot be attributed to the move. The figures are real measurements of the window and are shown below in full — they feed no memory, enter no fee, and are excluded from the total above.

Cost validatednot attributablerule: overpowered_modelcredit-memo

Right-size Credit Memo: move low-output claude-sonnet traffic to claude-haiku

€8,255
predicted / yr
Implemented
2026-06-10
Owner
Named owner recorded
Marked from
Not recorded — this mark predates the lane record, so we cannot say whether Bayeto was leading with this move at the time.
Measured outcome
€43,261/yr (20d / 13,600 calls observed)
Per call
€0.2326 €0.0583 · predicted €0.0367/call · accuracy 21% (impl-v2)
Verdict
Implemented 2026-06-10; per-call cost on this application moved €0.2326 → €0.0583 (13,600 calls over 20 calendar days, 20 active, post-implementation) → measured ≈ €43,261.26/yr, extrapolated ×18.3 from that window. Cost-only; quality not measured; simultaneous changes on this application would confound the attribution. Mean prompt size on this application rose 10% across the boundary (26595.93 → 29277.84 tokens per call, uncached plus cache-reads). THIS AXIS CANNOT SEPARATE a workload change from a move that itself changes prompt size — where the move is one of those, part of this number is the move working — so it is a figure to read beside the saving, never a verdict about it. It suppresses nothing. ATTRIBUTION CONTAMINATED: the application's model mix shifted across the window boundary (claude-sonnet 3%→99%, gpt-4-turbo 97%→0%) — the measured movement includes changes this mark does not explain. This figure is a window measurement, not an attributable saving; it feeds no memory and may never enter a fee. If you know the window is clean, correcting or re-marking the implementation date re-derives it.
Regressedrule: migration_laggardmeeting-summaries

Finish the gpt-4-turbo → claude-sonnet migration: Meeting Summaries missed the rollout

€40,940
predicted / yr
Implemented
2026-06-18
Owner
Named owner recorded
Marked from
Not recorded — this mark predates the lane record, so we cannot say whether Bayeto was leading with this move at the time.
Measured outcome
€73/yr (12d / 12,000 calls observed)
Per call
€0.1241 €0.1243
Verdict
Implemented 2026-06-18; per-call cost on this application ROSE €0.1241 → €0.1243 (12,000 calls over 12 calendar days, 12 active, post-implementation) → a measured regression of ≈ €73/yr (extrapolated ×30.4 from that window), not a saving. Cost-only; simultaneous changes would confound the attribution. Mean prompt size on this application rose 0% across the boundary (11989.59 → 12014.7 tokens per call, uncached plus cache-reads). THIS AXIS CANNOT SEPARATE a workload change from a move that itself changes prompt size — where the move is one of those, part of this number is the move working — so it is a figure to read beside the saving, never a verdict about it. It suppresses nothing.

Track record on your own data

Before the switch Bayeto predicted €0.203/call from moving gpt-4-turbo→claude-sonnet; the measured saving is €0.221/call — 92% accurate. Compared per call, so the figure reflects the engine's prediction quality and not traffic growth (volume -2% vs the prediction's basis).

What was deliberately not pursued

Held for safety
7 moves worth €34,566/yr were found and withheld: 5 on surfaces you declared critical, 2 on surfaces whose stakes you have not declared — declaring them is what lets Bayeto judge the risk rather than withhold. Not offered — not a matter of price.
Declined
0 recommendations worth €0/yr were reviewed and declined.
Still on the table
8 moves worth €47,104/yr remain available to pursue.

How these figures were produced

Recommendations come from a deterministic decision engine: the same telemetry and the same objective always yield the same recommendation, and no language model participates in the decision. Outcomes are measured by re-aggregating this workspace's own rows before and after each implementation date — cost only; quality is not measured, and simultaneous changes on an application would confound the attribution. Every figure on this page reconciles to the same canonical ledger as every other page and export of this workspace state.

ledger version 22c35e346df765c3 — identical on every page, export, and decision record of this workspace state · how to verify it

© 2026 Bayeto