What happened?
Google has released Gemini 3.6 Flash with improved coding, knowledge-work, multimodal and token efficiency, alongside 3.5 Flash-Lite for high-throughput workloads.
The update makes workload-specific model routing more important than choosing one flagship for everything. Teams could reserve 3.6 Flash for complex planning and use Flash-Lite for repetitive extraction or classification, but should measure total cost with their own prompt lengths and retry rates.
Why does it matter?
Agent operating cost depends not only on per-call pricing but also on output tokens, tool calls and throughput, so both models can directly change the economics of large-scale automation.
Who should care?
Related AI
AIZIGOO view
Speed and benchmark figures come from Google and the cited evaluators; real-world results may vary by workload.