Google has introduced Gemini 3.1 Flash-Lite, a new model designed to deliver fast performance and lower costs for developers building large-scale AI applications.
The company describes Gemini 3.1 Flash-Lite as the fastest and most cost-efficient model in the Gemini 3 series. It is specifically designed for high-volume developer workloads where speed, responsiveness and affordability are critical.
Starting today, the model is rolling out in preview to developers through the Gemini API in Google AI Studio and to enterprises through Vertex AI.
Gemini 3.1 Flash-Lite Designed for High-Volume Agentic AI Tasks
Google says Gemini 3.1 Flash-Lite was built to support high-frequency workloads where latency and efficiency matter most.
The model is aimed at Gemini Flash Lite high-volume agentic AI tasks such as translation at scale, content moderation, and other workflows that require rapid responses across large datasets.
According to Google, the model delivers faster responses compared with earlier generations. Benchmark results from Artificial Analysis show it achieves a 2.5x faster Time to First Answer Token and a 45% increase in output speed compared with the earlier 2.5 Flash model while maintaining similar or better quality.


The model also achieved an Elo score of 1432 on the Arena.ai leaderboard and performed strongly on reasoning and multimodal benchmarks. Google reported results including 86.9% on GPQA Diamond and 76.8% on MMMU Pro, surpassing some earlier Gemini models such as 2.5 Flash.

This combination of speed and capability makes the model suitable for real-time AI systems that must handle large volumes of requests without sacrificing responsiveness.
Gemini 3.1 Flash-Lite Pricing API Cost Developers Significantly Less
One of the main selling points behind Gemini 3.1 Flash-Lite is its pricing structure.
Google said the model is priced at $0.25 per one million input tokens and $1.50 per one million output tokens. This Gemini 3.1 Flash-Lite pricing API cost developers significantly less compared with larger AI models while still delivering strong performance.
The company also introduced configurable “thinking levels” available in AI Studio and Vertex AI. These controls allow developers to determine how much reasoning the model should apply to a given task, helping balance cost, speed and accuracy depending on the workload.
In addition to handling high-volume tasks, the model can also perform more complex functions such as generating user interfaces and dashboards, building simulations, or following multi-step instructions.
Several early-access developers and companies — including Latitude, Cartwheel and Whering — are already testing Gemini 3.1 Flash-Lite through AI Studio and Vertex AI.
According to Google, early testers highlighted the model’s efficiency and reasoning capabilities, noting that it can process complex inputs while maintaining instruction accuracy similar to larger models.
With Gemini 3.1 Flash-Lite, Google is targeting developers who need scalable AI infrastructure that can power large workloads without dramatically increasing operational costs.

Don’t miss out on our latest news—follow us for the latest AI news, breakthroughs, and insights that matter.