GPT-5.4 mini vs Gemini 3 Flash Preview: cost compared

Last verified

Gemini 3 Flash Preview is 33% cheaper on input and 33% cheaper on output than GPT-5.4 mini. Whether that makes it the right choice depends entirely on whether it does your task well enough — this page gives you the cost side of that trade, precisely.

Rates side by side

$ per million tokensGPT-5.4 miniGemini 3 Flash Preview
Input$0.75$0.50
Output$4.50$3.00
Cached input$0.07$0.05
Output : input ratio6.0×6.0×

What the difference is worth

Monthly cost for the same workload on each, without caching:

WorkloadGPT-5.4 miniGemini 3 Flash PreviewDifference
Support chatbot
1,000 conversations/month · 2,000 in + 500 out each
$3.75$2.50$1.25 saved
RAG search
10,000 queries/month · 8,000 in + 400 out each
$78.00$52.00$26.00 saved
Bulk extraction
10,000 documents · 3,000 in + 300 out each
$36.00$24.00$12.00 saved
Coding agent
100 runs/month · 400,000 in + 20,000 out each
$39.00$26.00$13.00 saved

Which to use

Gemini 3 Flash Preview if the task is well-specified and mechanical — extraction, classification, formatting, routine transformation. The saving is real and compounds with volume.

GPT-5.4 mini if the task involves genuine reasoning, long-horizon planning, or work where a wrong answer is expensive to catch. Paying 1.5× more on input is trivial compared to the cost of shipping a bad result.

These are different providers, so switching means a different SDK and a different tokenizer. The same text is a different number of tokens on each, so the price difference above is approximate at the margin. Measure token counts on your real inputs before committing.

Don't skip caching

Cached input costs $0.07 on GPT-5.4 mini and $0.05 on Gemini 3 Flash Preview. If your prompts share a stable prefix, caching on the more expensive model can beat switching to the cheaper one outright — worth checking before you migrate anything.

Sources

OpenAI rates verified 2026-08-05 (source). Google rates verified 2026-08-05 (source).