A Concise Guide to Gemini Models (2026-05-20)
Listen to this article
AI-generated voice10:55
Goal: Help beginners and architects quickly understand
gemini-3.5-flashand other cutting-edge Gemini models, clarifying which model to choose for different scenarios to achieve the optimal balance between performance and cost.
One-Sentence Conclusion
Currently, the latest mainline of the Google Gemini API has evolved to Gemini 3.5. The current stable flagship model is Gemini 3.5 Flash, with the API identifier (Model Code):
gemini-3.5-flash
Quick Selection Cheat Sheet
| Core Requirement | Recommended Model | Reason for Choice |
|---|---|---|
| Complex Coding, Agents, long-context tasks, multi-step tool use | gemini-3.5-flash | The latest stable flagship model. Extremely fast, strong reasoning, designed specifically for real-world Agent workflows. |
| Strongest Pro-level reasoning, complex software engineering, highly reliable tool use | gemini-3.1-pro-preview | Pro-level preview model, ideal for highly difficult, deep reasoning tasks. |
| Ultra-low cost, massive simple tasks, classification, extraction, translation | gemini-3.1-flash-lite | Ultimate cost-efficiency, extremely fast response, perfect for high-frequency lightweight tasks. |
| Requires legacy Gemini 3 Flash behavior or Computer Use | gemini-3-flash-preview | Preview model supporting Computer Use (browser/desktop automation); not the preferred choice for general tasks. |
| Real-time bidirectional voice conversation | gemini-3.1-flash-live-preview | Dedicated Live API model supporting low-latency audio interaction. |
| Image generation and editing | gemini-3.1-flash-image-preview<br>gemini-3-pro-image-preview | Dedicated image generation models, not for general text tasks. |
1. Flagship Workhorse: Gemini 3.5 Flash
π Basic Info
| Item | Description |
|---|---|
| API Model Code | gemini-3.5-flash |
| Model Status | Stable |
| Latest Update | 2026-05 |
| Knowledge Cutoff | 2025-01 |
| Input Modalities | Text, Image, Video, Audio, PDF |
| Output Modalities | Text |
| Context Window (Input) | 1,048,576 tokens (1M) |
| Max Output Tokens | 65,536 tokens (64K) |
π― Best Use Cases
gemini-3.5-flash is currently the most recommended general-purpose model to integrate first. It can serve as the default starting point for almost all new projects:
- Coding Agents: Code writing, auto-completion, and small-to-medium refactoring.
- Multi-step Tool Use: Executing stable Function Calling in complex Agent planning.
- Long-Context Tasks: Reading ultra-long documents, multi-file correlation analysis, execution, and verification.
- Enterprise Workflows: Intelligent assistants in business systems like ERP / CRM.
- Advanced Feature Support: Native support for Search Grounding (Google Search enhancement) and Structured Output (strict JSON Schema enforcement).
β Non-Use Cases
- Cannot directly generate or edit images.
- Does not support ultra-low latency bidirectional voice chat via the Live API.
- Does not support Computer Use (use
gemini-3-flash-previewif you need to test this feature).
π° Pricing (Paid Tier, Standard)
| Metric | Price |
|---|---|
| Input Tokens | USD 1.50 / 1M tokens |
| Output Tokens (including Thinking) | USD 9.00 / 1M tokens |
| Context Caching | USD 0.15 / 1M tokens |
β οΈ Note: If Google Search / Maps Grounding is enabled, additional charges will apply per actual Search Query once the free tier is exceeded.
2. Deep Reasoning: Gemini 3.1 Pro Preview
π Basic Info
| Item | Description |
|---|---|
| API Model Code | gemini-3.1-pro-preview |
| Custom Tools Endpoint | gemini-3.1-pro-preview-customtools |
| Model Status | Preview |
| Latest Update | 2026-02 |
| Knowledge Cutoff | 2025-01 |
| Input/Output Limits | Input 1M tokens / Output 64K tokens |
π― Best Use Cases
Suitable for "heavy reasoning" tasks that demand extreme logical rigor, where you are willing to trade off some speed and cost:
- Large-scale Code Refactoring Planning: Complex architectural design across modules and repositories.
- Deep Debugging: Analyzing ultra-long logs to locate elusive concurrency or memory bugs.
- Complex Software Engineering Decisions: Evaluating the pros and cons of multiple technical solutions.
- Highly Reliable Agents: If your Agent utilizes a large number of custom tools (e.g.,
view_file,run_shell), using the specially optimizedgemini-3.1-pro-preview-customtoolsendpoint is highly recommended.
π° Pricing (Paid Tier, Standard)
| Context Length | Input Price (per 1M) | Output Price (per 1M) |
|---|---|---|
| <= 200k tokens | USD 2.00 | USD 12.00 |
| > 200k tokens | USD 4.00 | USD 18.00 |
3. Ultimate Cost-Efficiency: Gemini 3.1 Flash-Lite
π Basic Info
| Item | Description |
|---|---|
| API Model Code | gemini-3.1-flash-lite |
| Model Status | Stable |
| Latest Update | 2026-05 |
| Input/Output Limits | Input 1M tokens / Output 64K tokens |
π― Best Use Cases
Designed to handle high-concurrency, lightweight tasks with extremely low latency and cost:
- Large-scale Text Processing: Batch translation, text classification, sentiment analysis.
- Structured Data Extraction: Extracting simple JSON fields from unstructured text.
- Intelligent Routing (Router): Serving as a front-end gateway to determine user intentβanswering simple queries directly and routing complex ones to Flash or Pro.
β Non-Use Cases
- Complex programming (Coding) and multi-step logical reasoning.
- Complex Agent planning and high-precision tool use.
π° Pricing (Paid Tier, Standard)
| Metric | Price |
|---|---|
| Input (Text / Image / Video) | USD 0.25 / 1M tokens |
| Input (Audio) | USD 0.50 / 1M tokens |
| Output (including Thinking) | USD 1.50 / 1M tokens |
4. Special Transition: Gemini 3 Flash Preview
π Basic Info
| Item | Description |
|---|---|
| API Model Code | gemini-3-flash-preview |
| Model Status | Preview |
| Latest Update | 2025-12 |
π― Best Use Cases
Following the release of gemini-3.5-flash, this model is no longer the preferred general-purpose choice, but it still holds value in specific scenarios:
- Computer Use Testing: Simulating human interaction with a browser or desktop (currently supported in this preview version).
- Legacy System Compatibility: Replicating or maintaining specific generation behaviors of the older Gemini 3 Flash.
π° Pricing (Paid Tier, Standard)
| Metric | Price |
|---|---|
| Input (Text / Image / Video) | USD 0.50 / 1M tokens |
| Input (Audio) | USD 1.00 / 1M tokens |
| Output (including Thinking) | USD 3.00 / 1M tokens |
5. Specialized Models at a Glance
These models are not intended for general text conversations but are specialized endpoints optimized for specific multimodal scenarios:
| Model Name | API Model Code | Core Purpose |
|---|---|---|
| Gemini 3.1 Flash Live Preview | gemini-3.1-flash-live-preview | Real-time voice conversation (Audio-to-Audio) with ultra-low latency. |
| Gemini 3.1 Flash TTS Preview | gemini-3.1-flash-tts-preview | High-quality Text-to-Speech conversion. |
| Gemini 3.1 Flash Image Preview | gemini-3.1-flash-image-preview | Fast image generation and inpainting/editing. |
| Gemini 3 Pro Image Preview | gemini-3-pro-image-preview | Image generation balancing high visual quality with complex prompt understanding. |
6. Legacy Retention: Gemini 2.5 Series
If your legacy projects are still using the 2.5 series, they remain active and available. However, it is highly recommended to evaluate and migrate new projects directly to the Gemini 3 / 3.5 series.
gemini-2.5-pro: Legacy high-precision reasoning model.gemini-2.5-flash: Legacy balanced low-latency model.gemini-2.5-flash-lite: Legacy low-cost model.
7. Recommended Selection Logic
π‘ General Decision Flow
- Default Choice: Start prototyping directly with
gemini-3.5-flash. - Cost Optimization: If the tasks are simple, highly concurrent, and budget-sensitive, downgrade to
gemini-3.1-flash-lite. - Capability Upgrade: If you encounter complex code refactoring or multi-step logical bottlenecks, upgrade to
gemini-3.1-pro-preview. - Multimodal Specialties: Use
Livefor voice conversations andImagemodels for image generation.
π’ ERP / AI Agent Implementation Recommendations
| Business Scenario | Recommended Model | Implementation Reason |
|---|---|---|
| Default Chatbox Agent | gemini-3.5-flash | Fast response, strong bilingual capabilities, smooth user experience. |
| SQL Generation / ERP Table Relationship Classification | gemini-3.1-flash-lite | Structured tasks; extremely low cost for batch processing. |
| Complex Debugging / Cross-module Refactoring | gemini-3.1-pro-preview | Highly logical, low hallucination rate. |
| Massive Log Summarization / Intent Routing | gemini-3.1-flash-lite | High throughput, highly cost-effective. |
| Multi-table Queries + Web Search (Grounding) | gemini-3.5-flash | Perfect support for Function Calling and Grounding. |
8. Developer Pitfalls to Avoid
π¨ Pitfall 1: Do not use includes("mini") to match lightweight models
In many open-source frameworks or custom codebases, developers like to identify lightweight models by checking if the model ID contains "mini".
However, in Gemini, the word "gemini" itself contains the substring "mini"!
If you use modelId.includes("mini"), all Gemini models will be misclassified as Lite-tier, causing configuration chaos.
Correct Approach (using regex boundary matching):
// Recommended matching pattern
const LITE_TIER_PATTERN = /\b(?:flash-lite|flash|mini|haiku|nano)\b/i;
const isLiteTier = LITE_TIER_PATTERN.test(modelId);
π¨ Pitfall 2: Deploying Preview models to production
Models with the -preview suffix (such as gemini-3.1-pro-preview) are preview versions. Google may adjust their backend behavior at any time, and rate limits are typically much stricter.
Golden Rule: For core production workflows, always prioritize Stable models without the
-previewsuffix (e.g.,gemini-3.5-flash).
π¨ Pitfall 3: Ignoring the cost of Thinking Tokens
In Gemini's billing design, "Thinking Tokens" generated by the model before outputting the final answer are billed as Output Tokens. For complex reasoning tasks, the actual bill may be higher than estimated pure text outputs. Architects should leave some budget buffer.
π¨ Pitfall 4: Hidden bills from Grounding
When Google Search or Google Maps Grounding is enabled, the API automatically retrieves real-world information based on the user's prompt. Please note that each search query is billed separately once the free quota is exceeded, which requires strict quota management in high-concurrency scenarios.
9. Minimalist Decision Tree
Start Selection
ββ Task Type?
β ββ Simple Tasks (Classification/Translation/Simple Extraction/Routing)
β β ββ π Choose gemini-3.1-flash-lite (Cost-saving, extremely fast)
β β
β ββ Core Tasks (General Coding/Agents/ERP Chat/Multi-step Tool Use)
β β ββ π Choose gemini-3.5-flash (All-rounder, stable, default choice)
β β
β ββ Highly Complex Tasks (Deep Debugging/Architecture Design/Long-context Reasoning)
β ββ π Choose gemini-3.1-pro-preview (High precision, strong reasoning)
β
ββ Special Modality Requirements?
ββ Real-time Bidirectional Voice Chat ββπ Choose gemini-3.1-flash-live-preview
ββ Image Generation & Editing ββββπ Choose gemini-3.1-flash-image-preview / gemini-3-pro-image-preview
