PPT App Testing: Cost Anomaly Across Model Multipliers (Noticed Along the Way)
While testing the PPT App today (single-page PPT generation), noticed some inconsistencies between token cost and the model multiplier numbers β flagging it here.
These werenβt dedicated cost-measurement tests β just observations picked up during regular PPT testing. But since each run used the same prompt and same variables (search on, same template choice or none), with only the model/multiplier changed, the data still seems worth sharing.
Gemini Γ3
- 1-1: no template β 221 tokens / 10min
- 1-2: no template β 255.28 tokens / 7min
Gemini Γ6
- 2-1: no template β 299 tokens / 10min
- 2-2: template β 602 tokens / 14min
- 2-2: template β 602 tokens / 14min
- 2-5: youming template β 523.455 tokens / 20min18s
Gemini Γ18
- 3: no template β 2717.623 tokens / 13min
GPT Γ4
- 4-1: template β 177.629 tokens / 9min
Kimi Γ6.78
- 5: no template β 585.105 tokens / 20min29s
Qwen Γ3.31
- 6: no template β 1108.927 tokens / 15min35s
GPT Γ20
- 7: no template β 378 tokens / 6min
GPT Γ28
- 8: no template β 409.358 tokens / 10min
Claude Γ30
- 9: no template β 2235.757 tokens β cost is roughly in the same range as Gemini Γ18 (2717.623), despite the multiplier being nearly double
The issue
Not a dedicated cost test, but since scenario (single-page PPT), prompt, and variables were all consistent β only the model/multiplier changed β cost should scale somewhat predictably with the multiplier. It doesnβt:
- Gemini 6β18 (3x multiplier increase) β actual cost jumped ~5β9x (vs. normal Γ6 baseline of ~300β600 tokens)
- Claude Γ30 costs about the same as Gemini Γ18, despite the multiplier being nearly double
Just something noticed along the way, not a rigorous cost comparison experiment β but the numbers diverge enough that itβs worth the team taking a look at whether thereβs a billing logic issue.