← Back to all posts

9 min readUpdated 20 May 2026
GeminiLLMGoogle CloudAPIAI Agent+1 more

A Concise Guide to Gemini Models (2026-05-20)

Yap Wei JunSoftware Engineer Β· Singapore

Listen to this article

AI-generated voice10:55

Download MP3

Goal: Help beginners and architects quickly understand gemini-3.5-flash and other cutting-edge Gemini models, clarifying which model to choose for different scenarios to achieve the optimal balance between performance and cost.


One-Sentence Conclusion

Currently, the latest mainline of the Google Gemini API has evolved to Gemini 3.5. The current stable flagship model is Gemini 3.5 Flash, with the API identifier (Model Code):

gemini-3.5-flash

Quick Selection Cheat Sheet

Core RequirementRecommended ModelReason for Choice
Complex Coding, Agents, long-context tasks, multi-step tool usegemini-3.5-flashThe latest stable flagship model. Extremely fast, strong reasoning, designed specifically for real-world Agent workflows.
Strongest Pro-level reasoning, complex software engineering, highly reliable tool usegemini-3.1-pro-previewPro-level preview model, ideal for highly difficult, deep reasoning tasks.
Ultra-low cost, massive simple tasks, classification, extraction, translationgemini-3.1-flash-liteUltimate cost-efficiency, extremely fast response, perfect for high-frequency lightweight tasks.
Requires legacy Gemini 3 Flash behavior or Computer Usegemini-3-flash-previewPreview model supporting Computer Use (browser/desktop automation); not the preferred choice for general tasks.
Real-time bidirectional voice conversationgemini-3.1-flash-live-previewDedicated Live API model supporting low-latency audio interaction.
Image generation and editinggemini-3.1-flash-image-preview<br>gemini-3-pro-image-previewDedicated image generation models, not for general text tasks.

1. Flagship Workhorse: Gemini 3.5 Flash

πŸ“Š Basic Info

ItemDescription
API Model Codegemini-3.5-flash
Model StatusStable
Latest Update2026-05
Knowledge Cutoff2025-01
Input ModalitiesText, Image, Video, Audio, PDF
Output ModalitiesText
Context Window (Input)1,048,576 tokens (1M)
Max Output Tokens65,536 tokens (64K)

🎯 Best Use Cases

gemini-3.5-flash is currently the most recommended general-purpose model to integrate first. It can serve as the default starting point for almost all new projects:

  • Coding Agents: Code writing, auto-completion, and small-to-medium refactoring.
  • Multi-step Tool Use: Executing stable Function Calling in complex Agent planning.
  • Long-Context Tasks: Reading ultra-long documents, multi-file correlation analysis, execution, and verification.
  • Enterprise Workflows: Intelligent assistants in business systems like ERP / CRM.
  • Advanced Feature Support: Native support for Search Grounding (Google Search enhancement) and Structured Output (strict JSON Schema enforcement).

❌ Non-Use Cases

  • Cannot directly generate or edit images.
  • Does not support ultra-low latency bidirectional voice chat via the Live API.
  • Does not support Computer Use (use gemini-3-flash-preview if you need to test this feature).

πŸ’° Pricing (Paid Tier, Standard)

MetricPrice
Input TokensUSD 1.50 / 1M tokens
Output Tokens (including Thinking)USD 9.00 / 1M tokens
Context CachingUSD 0.15 / 1M tokens

⚠️ Note: If Google Search / Maps Grounding is enabled, additional charges will apply per actual Search Query once the free tier is exceeded.


2. Deep Reasoning: Gemini 3.1 Pro Preview

πŸ“Š Basic Info

ItemDescription
API Model Codegemini-3.1-pro-preview
Custom Tools Endpointgemini-3.1-pro-preview-customtools
Model StatusPreview
Latest Update2026-02
Knowledge Cutoff2025-01
Input/Output LimitsInput 1M tokens / Output 64K tokens

🎯 Best Use Cases

Suitable for "heavy reasoning" tasks that demand extreme logical rigor, where you are willing to trade off some speed and cost:

  • Large-scale Code Refactoring Planning: Complex architectural design across modules and repositories.
  • Deep Debugging: Analyzing ultra-long logs to locate elusive concurrency or memory bugs.
  • Complex Software Engineering Decisions: Evaluating the pros and cons of multiple technical solutions.
  • Highly Reliable Agents: If your Agent utilizes a large number of custom tools (e.g., view_file, run_shell), using the specially optimized gemini-3.1-pro-preview-customtools endpoint is highly recommended.

πŸ’° Pricing (Paid Tier, Standard)

Context LengthInput Price (per 1M)Output Price (per 1M)
<= 200k tokensUSD 2.00USD 12.00
> 200k tokensUSD 4.00USD 18.00

3. Ultimate Cost-Efficiency: Gemini 3.1 Flash-Lite

πŸ“Š Basic Info

ItemDescription
API Model Codegemini-3.1-flash-lite
Model StatusStable
Latest Update2026-05
Input/Output LimitsInput 1M tokens / Output 64K tokens

🎯 Best Use Cases

Designed to handle high-concurrency, lightweight tasks with extremely low latency and cost:

  • Large-scale Text Processing: Batch translation, text classification, sentiment analysis.
  • Structured Data Extraction: Extracting simple JSON fields from unstructured text.
  • Intelligent Routing (Router): Serving as a front-end gateway to determine user intentβ€”answering simple queries directly and routing complex ones to Flash or Pro.

❌ Non-Use Cases

  • Complex programming (Coding) and multi-step logical reasoning.
  • Complex Agent planning and high-precision tool use.

πŸ’° Pricing (Paid Tier, Standard)

MetricPrice
Input (Text / Image / Video)USD 0.25 / 1M tokens
Input (Audio)USD 0.50 / 1M tokens
Output (including Thinking)USD 1.50 / 1M tokens

4. Special Transition: Gemini 3 Flash Preview

πŸ“Š Basic Info

ItemDescription
API Model Codegemini-3-flash-preview
Model StatusPreview
Latest Update2025-12

🎯 Best Use Cases

Following the release of gemini-3.5-flash, this model is no longer the preferred general-purpose choice, but it still holds value in specific scenarios:

  • Computer Use Testing: Simulating human interaction with a browser or desktop (currently supported in this preview version).
  • Legacy System Compatibility: Replicating or maintaining specific generation behaviors of the older Gemini 3 Flash.

πŸ’° Pricing (Paid Tier, Standard)

MetricPrice
Input (Text / Image / Video)USD 0.50 / 1M tokens
Input (Audio)USD 1.00 / 1M tokens
Output (including Thinking)USD 3.00 / 1M tokens

5. Specialized Models at a Glance

These models are not intended for general text conversations but are specialized endpoints optimized for specific multimodal scenarios:

Model NameAPI Model CodeCore Purpose
Gemini 3.1 Flash Live Previewgemini-3.1-flash-live-previewReal-time voice conversation (Audio-to-Audio) with ultra-low latency.
Gemini 3.1 Flash TTS Previewgemini-3.1-flash-tts-previewHigh-quality Text-to-Speech conversion.
Gemini 3.1 Flash Image Previewgemini-3.1-flash-image-previewFast image generation and inpainting/editing.
Gemini 3 Pro Image Previewgemini-3-pro-image-previewImage generation balancing high visual quality with complex prompt understanding.

6. Legacy Retention: Gemini 2.5 Series

If your legacy projects are still using the 2.5 series, they remain active and available. However, it is highly recommended to evaluate and migrate new projects directly to the Gemini 3 / 3.5 series.

  • gemini-2.5-pro: Legacy high-precision reasoning model.
  • gemini-2.5-flash: Legacy balanced low-latency model.
  • gemini-2.5-flash-lite: Legacy low-cost model.

πŸ’‘ General Decision Flow

  1. Default Choice: Start prototyping directly with gemini-3.5-flash.
  2. Cost Optimization: If the tasks are simple, highly concurrent, and budget-sensitive, downgrade to gemini-3.1-flash-lite.
  3. Capability Upgrade: If you encounter complex code refactoring or multi-step logical bottlenecks, upgrade to gemini-3.1-pro-preview.
  4. Multimodal Specialties: Use Live for voice conversations and Image models for image generation.

🏒 ERP / AI Agent Implementation Recommendations

Business ScenarioRecommended ModelImplementation Reason
Default Chatbox Agentgemini-3.5-flashFast response, strong bilingual capabilities, smooth user experience.
SQL Generation / ERP Table Relationship Classificationgemini-3.1-flash-liteStructured tasks; extremely low cost for batch processing.
Complex Debugging / Cross-module Refactoringgemini-3.1-pro-previewHighly logical, low hallucination rate.
Massive Log Summarization / Intent Routinggemini-3.1-flash-liteHigh throughput, highly cost-effective.
Multi-table Queries + Web Search (Grounding)gemini-3.5-flashPerfect support for Function Calling and Grounding.

8. Developer Pitfalls to Avoid

🚨 Pitfall 1: Do not use includes("mini") to match lightweight models

In many open-source frameworks or custom codebases, developers like to identify lightweight models by checking if the model ID contains "mini". However, in Gemini, the word "gemini" itself contains the substring "mini"! If you use modelId.includes("mini"), all Gemini models will be misclassified as Lite-tier, causing configuration chaos.

Correct Approach (using regex boundary matching):

// Recommended matching pattern
const LITE_TIER_PATTERN = /\b(?:flash-lite|flash|mini|haiku|nano)\b/i;
const isLiteTier = LITE_TIER_PATTERN.test(modelId);

🚨 Pitfall 2: Deploying Preview models to production

Models with the -preview suffix (such as gemini-3.1-pro-preview) are preview versions. Google may adjust their backend behavior at any time, and rate limits are typically much stricter.

Golden Rule: For core production workflows, always prioritize Stable models without the -preview suffix (e.g., gemini-3.5-flash).

🚨 Pitfall 3: Ignoring the cost of Thinking Tokens

In Gemini's billing design, "Thinking Tokens" generated by the model before outputting the final answer are billed as Output Tokens. For complex reasoning tasks, the actual bill may be higher than estimated pure text outputs. Architects should leave some budget buffer.

🚨 Pitfall 4: Hidden bills from Grounding

When Google Search or Google Maps Grounding is enabled, the API automatically retrieves real-world information based on the user's prompt. Please note that each search query is billed separately once the free quota is exceeded, which requires strict quota management in high-concurrency scenarios.


9. Minimalist Decision Tree

Start Selection
 β”œβ”€ Task Type?
 β”‚   β”œβ”€ Simple Tasks (Classification/Translation/Simple Extraction/Routing)
 β”‚   β”‚   └─ πŸ‘‰ Choose gemini-3.1-flash-lite (Cost-saving, extremely fast)
 β”‚   β”‚
 β”‚   β”œβ”€ Core Tasks (General Coding/Agents/ERP Chat/Multi-step Tool Use)
 β”‚   β”‚   └─ πŸ‘‰ Choose gemini-3.5-flash (All-rounder, stable, default choice)
 β”‚   β”‚
 β”‚   └─ Highly Complex Tasks (Deep Debugging/Architecture Design/Long-context Reasoning)
 β”‚       └─ πŸ‘‰ Choose gemini-3.1-pro-preview (High precision, strong reasoning)
 β”‚
 └─ Special Modality Requirements?
     β”œβ”€ Real-time Bidirectional Voice Chat β”€β”€πŸ‘‰ Choose gemini-3.1-flash-live-preview
     └─ Image Generation & Editing β”€β”€β”€β”€πŸ‘‰ Choose gemini-3.1-flash-image-preview / gemini-3-pro-image-preview

← Back to all posts