Route small-context Portfolio Analytics traffic off claude-sonnet to GLM-5.2
Route the 85.7% of Portfolio Analytics traffic that is small-context from claude-sonnet to Zhipu GLM-5.2
Verify before implementing — this recommendation rests on evidence the engine does not have:
- No measured quality of GLM-5.2 vs the premium model on THIS app's small-context traffic.
- No independent quality evaluation — the quality impact is unmeasured.
Implement only if these checks pass. The full self-critique is in the reasoning trail below.
Payback divides by an implementation cost of ≈€220 under the agent-assisted-v1-2026-08 cost model: 2h of human review at €100/h plus a €20 agent-run allowance — an AI agent implements, your team reviews and verifies. Implemented by hand instead, the same move is estimated at 2 engineer-days (≈€1,600). Every figure is a prior, not a measurement of your team — and the model is versioned so a changed assumption can never move a figure silently.
Context completeness
Confidence 97% is how sure the engine is on the evidence it has. This is how much of the evidence that could change the conclusion the engine could establish for this application — a different question. The declare-context page counts what you have answered workspace-wide, which can legitimately read differently: an “I don't know” is an answer there and establishes nothing here.
- Telemetry
- Provider bill— calibrated from the monthly spend declared for this workspace, not a reconciled invoice
- Per-app criticality
- Prior experiments
- Codebase
- Quality evaluation
We cannot detect Codebase, Quality evaluation — you can tell us, and confidence is re-derived on the fuller evidence.
Observation → Evidence → Inference → Recommendation
— every fact, judgment, self-critique and failure condition behind this decision
Observations (facts)
- [small_context_share] 85.7% of Portfolio Analytics traffic is small-context on a premium model. (N=17,280)
- [model_small_context_requests] 14804 of 15180 claude-sonnet requests in Portfolio Analytics are small-context. (N=15,180)
Evidence (objective-relevant)
- 85.7% of Portfolio Analytics traffic (14804 sampled requests) is small-context on a premium model. (strength 0.95)
Inference (judgment)
Small-context traffic on claude-sonnet is over-specified for the work; a capable low-cost model is Pareto-superior under a cost objective.
Recommendation (move)
Route the 85.7% of Portfolio Analytics traffic that is small-context from claude-sonnet to Zhipu GLM-5.2
Critical Review (self-critique)
Assumptions
- Assumes SMALL input ⇒ SIMPLE task — that a low-cost model handles short-context requests within quality tolerance.
Alternative explanations
- A short prompt can still pose a hard question; context length is not task difficulty.
Missing evidence
- No measured quality of GLM-5.2 vs the premium model on THIS app's small-context traffic.
- No independent quality evaluation — the quality impact is unmeasured.
This is wrong if…
- Wrong if the small-context slice includes high-stakes questions where the cheaper model degrades answers.
- Wrong if this workload's real quality bar is higher than assumed and the change degrades answers.
Download the canonical decision record (JSON) — the deterministic decision block (byte-pinned by CI), the run context (engine + catalog versions, measured-window boundaries, exact pricing rows, fixed FX with effective date, formula inputs, the overlap group, and the ledger version), and the LLM-phrased narrative as a visibly separate section.
Business case & technical detail
— the phrased explanation — the decision above is computed without it
Business case
Portfolio Analytics pays premium claude-sonnet rates on 85.7% of traffic a capable low-cost model can serve. The quality effect is unmeasured — validate on your own workload before a full cutover. Estimated €567/yr at the current run-rate, payback 142 days, confidence 97%.
Technical detail
Add a router branch: when input ≤ 1000 tokens, serve from GLM-5.2 instead of claude-sonnet. No prompt changes required. Affects ~18,505 req/mo; expected cost -79.2%, latency -5%, quality unmeasured (low risk).
Benchmark context & Optimization Memory
— your value against the cohort, and what prior outcomes taught the engine
Benchmark context
Share of premium-model traffic that is small-context (lower is leaner). (illustrative)
Optimization Memory
Learned from 4 prior implementations, 4 succeeded (N=172,500).
Why NOT the alternatives?
— each rejected route, with the constraint that rejected it
- Route to Claude Haikucost-86%latency-50%qualityunmeasured (medium risk)Cheaper still, but a cross-tier drop carries a higher unmeasured quality risk on harder tickets (rework/escalation cost). Neither candidate's quality effect on this workload has been measured.
- Route to Claude Opuscost+600%latency+70%qualityunmeasured (low risk)Cost +600% and latency +70% for an UNMEASURED quality change → strictly dominated.
- Route to GPT-4ocost+210%latency-2%qualityunmeasured (low risk)Cost +210% for an UNMEASURED quality change → not Pareto-optimal under a cost objective.
Implementation packet
advisory_onlyrisk: model-substitutionReviewed by Bayeto engineering on 2026-07-26 · rule version v1 · 94% of the €56,350/yr open on this workspace is carried by a packet an engineer can execute (5 of 8 open moves); 3 of 14 specs meet that standard. Missing mandatory for this risk class: guardrails, stopIf, rollback, qualityValidation.
Review attests that these steps are sound and reversible as written. It does not guarantee an outcome: whether the change saves money is measured by the validation loop afterwards, never promised by the reviewer.
Not reviewed to executable standard
Guardrails — watch these, Stop if, Rollback, Quality validation are all unwritten for this move, for one reason: Not reviewed to executable standard. "Small context" is a proxy for "easy", and Bayeto measures only the first — it sees token counts, never task difficulty — so which of these requests the cheaper model can actually serve is a judgement about your workload rather than an arithmetic result.
What you would need to decide
Bayeto withholds steps here because these are your calls to make, not ours. Answer them and this move can be specified.
Which of these small-context requests are simple ENOUGH for the cheaper model, as opposed to merely short?
Bayeto sees token counts, never task difficulty. A short prompt can carry a hard question, and the correlation between length and difficulty is an assumption about your workload that only you can check.
What is the longest request the cheaper model must still handle without truncating?
The rollup reports an average, not the tail. A request above the cheaper tier fails outright rather than costing more, and how often that is acceptable is your call.
- Preconditions — 2 specified instructions (mandatory for this risk class)
- Steps — 1 specified instruction (mandatory for this risk class)
- Guardrails — watch these — none specified (mandatory for this risk class)withheld — for the reason stated above
- Stop if — none specified (mandatory for this risk class)withheld — for the reason stated above
- Quality validation — none specified (mandatory for this risk class)withheld — for the reason stated above
- Rollback — none specified (mandatory for this risk class)withheld — for the reason stated above
Instruction text is withheld here: the public demo shows complete packets only for the moves reviewed to the executable standard — 3 of 14 specs — as its worked examples; the rest summarize what exists and what was deliberately not specified.
Simulate — Before You Build
The cost delta derives from this workspace’s telemetry. The latency and quality deltas are heuristic priors carried by the rule: no quality evaluation is connected, and no rule measures the latency effect of its own move — the figures are per-rule constants, so nothing here observes either outcome.
- · Applies to ~18,505 requests/month at current mix.
- · Scenario realization assumption (sim-v1, not a statistical expected value): 0.7 + 0.3 × confidence = 0.991. The 0.7 floor encodes implementation shortfall — even a certain move rarely captures 100% at rollout; rationale published at bayeto.ai/methodology.
- · Per-request cost/latency/quality deltas hold at current traffic shape.