# gemini-25-flash-lite-stable-ga-google-2025-07-22

## Veille

Gemini 2.5 Flash-Lite - Google - Stable GA - Cost-efficient - Fastest model - Developer Blog

## Titre Article

Gemini 2.5 Flash-Lite is now stable and generally available - Google Developers Blog

## Date

2025-07-22

## URL

https://developers.googleblog.com/en/gemini-25-flash-lite-is-now-stable-and-generally-available/

## Keywords

Gemini 2.5 Flash-Lite, AI, machine learning, Google AI Studio, Vertex AI, large language model, cost-efficient, speed, latency, multimodal, context window, Grounding with Google Search, Code Execution, URL Context, translation, classification

## Authors

Logan Kilpatrick (Group Product Manager), Zach Gleicher (Product Manager)

## Ton

**Profile:** Product developer | Institutional developer relations | Informative-technical | Expert

Google's product managers adopt the voice of the Google Developers Blog aimed at a technical audience. The focus on model specifications (cost-efficiency, speed, latency, context window) addresses developer priorities. Accessible technical language (multimodal, grounding, code execution) demonstrates the depth of capabilities. Professional and confident tone, typical of Google's developer communications. Feature-centered structure with availability information facilitating adoption. Typical of developer-oriented product announcements (in the style of AWS What's New, Azure Updates), providing technical details to the developer community making integration decisions.

## Pense-betes

- **Fastest and most cost-efficient model** in the Gemini 2.5 family
- **Pricing**: **$0.10 per 1M input tokens**, **$0.40 per 1M output tokens**
- **40% reduction in audio input pricing** since the preview launch
- **Lower latency** than 2.0 Flash-Lite and 2.0 Flash across a wide range of prompts
- **1 million token context window**
- **Native reasoning capabilities**
- **Advanced tools**: Grounding with Google Search, Code Execution, URL Context
- **Optimized for latency-sensitive applications**: translation, classification
- **Benchmarks**: outperforms 2.0 Flash-Lite in code, math, science, reasoning, multimodal understanding
- **Real-world deployments**: Satlyt (**45% latency reduction**, **30% lower power consumption**), HeyGen (video planning automation, **180+ languages**), DocsHound (long video processing, thousands of screenshots extracted), Evertune (fast synthesis of large volumes)
- **Stable version**: specify "gemini-2.5-flash-lite", preview alias removed on **August 25**

## RésuméDe400mots

Google announced the **stable general availability** of Gemini 2.5 Flash-Lite, marking a significant advance in the Gemini 2.5 model family. Positioned as **the fastest and most cost-efficient model**, priced at **$0.10 per million input tokens and $0.40 per million output tokens**, Flash-Lite is designed to deliver exceptional intelligence per dollar. This release builds on the success of 2.5 Pro and 2.5 Flash, completing the 2.5 model lineup ready for large-scale production use.

**Performance and latency optimization**

Gemini 2.5 Flash-Lite is particularly optimized for **latency-sensitive applications** such as translation and classification, where speed and cost efficiency take priority without compromising quality. It shows **lower latency** than its predecessors, 2.0 Flash-Lite and 2.0 Flash, across a wide range of prompts. In addition, Google has **cut audio input pricing by 40%** since the preview launch, reinforcing its financial accessibility.

**Benchmark quality**

Despite its cost-efficient positioning, the model demonstrates **high quality across various benchmarks**, including code, math, science, reasoning, and multimodal understanding, **outperforming 2.0 Flash-Lite**. Developers building with 2.5 Flash-Lite have access to a robust feature set, including a **1 million token context window**, controllable thinking budgets, and native tool support. These tools include **Grounding with Google Search, Code Execution, and URL Context**, enabling more sophisticated and integrated AI applications.

**Validated real-world applications**

Practical applications of Gemini 2.5 Flash-Lite are already visible across several successful deployments. **Satlyt** leverages its speed to achieve a **45% reduction in latency** for critical onboard diagnostics and a **30% drop in power consumption** for its decentralized space computing platform. **HeyGen** relies on the model to automate video planning, optimize content, and translate videos into **more than 180 languages**, enabling global and personalized user experiences. **DocsHound** turns product demos into comprehensive documentation by quickly processing long videos and extracting thousands of screenshots. **Evertune** benefits from Flash-Lite's speed to rapidly scan and synthesize large volumes of model outputs, providing dynamic and timely analysis for brands monitoring their representation across AI models.

**Accessibility and deployment**

Developers can start using the **stable version** of Gemini 2.5 Flash-Lite by specifying **"gemini-2.5-flash-lite"** in their code, with the preview alias being removed on **August 25**. The model is available in **Google AI Studio and Vertex AI**, inviting developers to explore its capabilities for building innovative AI solutions. This strategic positioning of Flash-Lite as the go-to choice for cost-efficient, latency-sensitive applications illustrates Google's commitment to offering a multi-tier model lineup addressing developers' varied needs, from ultra-fast inference to deep reasoning, while maintaining competitive pricing to democratize access to advanced AI.

## GrapheDeConnaissance

- Google —publie→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.99)
- Logan Kilpatrick —publie→ Gemini 2.5 Flash-Lite stable GA (EVENEMENT, 0.98)
- Zach Gleicher —publie→ Gemini 2.5 Flash-Lite stable GA (EVENEMENT, 0.98)
- Gemini 2.5 Flash-Lite —fait_partie_de→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.99)
- Gemini 2.5 Flash-Lite —mesure→ $0.10 / 1M tokens input, $0.40 / 1M tokens output (MESURE, 0.99)
- Gemini 2.5 Flash-Lite —surpasse→ Gemini 2.0 Flash-Lite (TECHNOLOGIE, 0.95)
- Gemini 2.5 Flash-Lite —utilise→ Grounding with Google Search (TECHNOLOGIE, 0.97)
- Gemini 2.5 Flash-Lite —observé_dans→ Google AI Studio (TECHNOLOGIE, 0.98)
- Gemini 2.5 Flash-Lite —observé_dans→ Vertex AI (TECHNOLOGIE, 0.98)
- Satlyt —utilise→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.97)
- Satlyt —mesure→ réduction de 45% de la latence des diagnostics onboard (MESURE, 0.96)
- HeyGen —utilise→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.97)
- DocsHound —utilise→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.97)
- Evertune —utilise→ Gemini 2.5 Flash-Lite (TECHNOLOGIE, 0.97)
- Gemini 2.5 Flash-Lite —mesure→ fenêtre de contexte de 1 million de tokens (MESURE, 0.99)

---
Canonical: https://www.thekb.eu/en/fiches/gemini-25-flash-lite-stable-ga-google-2025-07-22/
