Google announced the general availability of Gemini 3.8 Flash on September 2, 2026, under the API model identifier gemini-3.8-flash. The model provides a 1,048,576-token input context window and a maximum output limit of 65,536 tokens. It incorporates selectable reasoning levels—low, medium, and high—with medium enabled as the default setting and minimal thinking unsupported.

The general model is accessible via the Gemini API, Google AI Studio, and Gemini Enterprise, as well as to Google AI Pro and Ultra subscribers in the Gemini app, Search AI Mode, and Google Sheets.

Pricing and workload trade-offs

Google introduced a two-stage pricing schedule for the API:

  • Introductory rates (through December 31, 2026): $0.75 per million input tokens and $3.75 per million output tokens.
  • Standard rates (beginning January 1, 2027): $1.50 per million input tokens and $7.50 per million output tokens.

Although introductory token rates match those of Gemini 3.7 Flash, Google cautions that Gemini 3.8 Flash is designed to consume more tokens on complex tasks. Total workload costs will depend on prompt volume, output length, thinking level, and tool usage. Google recommends dialing back reasoning effort or remaining on Gemini 3.7 Flash for efficiency-critical workflows. Teams planning sustained deployments also need to factor in the scheduled doubling of base token rates on January 1, 2027.

Restricted access for Flash Cyber

Google separately introduced Gemini 3.8 Flash Cyber, a specialized variant built with more permissive cybersecurity safeguards than the standard model.

Flash Cyber is not part of the public Gemini API. Google distributes it exclusively through its Fairwind program for vetted defenders, prioritizing government agencies, critical infrastructure entities, and software maintainers. Google did not announce a broader public API release date for the variant.