Gemini 2.5 retirement on Vertex AI: what breaks and how to prepare (with our real inventory)

Gemini 2.5 retires from Vertex AI on 20 October 2026. Official dates, our real call inventory and a migration checklist you can run this week.

Carlos Betancur Gálvez

By Carlos Betancur Gálvez

AI, digital marketing and medical marketing consultant · btodigital

20 October 2026 is the retirement date for Gemini 2.5 Pro, Flash and Flash-Lite on Vertex AI, according to the official Google Cloud documentation I checked on 4 October. When I audited the projects we run at btodigital, a digital agency in Medellín, Colombia, I found seven with active calls to those models, three that will be blocked, and code still pointing at models that no longer exist.

This piece is for you if you are a CTO, run an agency, or your company has something in production that calls Gemini through Google Cloud: a WhatsApp agent, a report generator, an email classifier. You get the verified dates, the real inventory we built (with numbers, without client names) and a migration checklist you can start this week.

Which dates are official and which are not?

There is a trap here, and I walked right into it.

On 15 September Google Cloud emailed us this timeline: on 20 October, projects that had not called those models between 22 July and 20 October would be blocked; active projects would keep working; Flash-Lite would shut down on 28 January 2027, and Flash and Pro on 31 March 2027. I built our inventory on that timeline and wrote it down in my notes.

The public model versions and lifecycle page, last updated on 2 October 2026, says something different: a single retirement date, 20 October 2026, for all three models. No January, no March. And back in April, the release notes had announced 16 October. The date has already moved once.

My position: plan for 20 October. The documentation is what Google commits to in writing for everyone, and the same page states that timelines may be extended but never brought forward. The extra time described in the email may exist for active projects. I would not bet a client’s operation on it.

ModelVertex AI retirement (official docs)Google’s suggested replacementGemini API (AI Studio)
gemini-2.5-pro20 Oct 2026gemini-3.8-flash or gemini-3.5-flashNo shutdown date; existing users only
gemini-2.5-flash20 Oct 2026gemini-3.8-flash, 3.5-flash-lite or 3.1-flash-liteNo shutdown date; existing users only
gemini-2.5-flash-lite20 Oct 2026gemini-3.8-flash, 3.1-flash-lite or Gemma 4No shutdown date; existing users only
gemini-2.5-flash-image15 Mar 2027gemini-3.1-flash-lite-imageShut down on 2 Oct 2026
gemini-2.0-flashRetired on 1 Jun 2026gemini-3.1-flash-liteShut down on 1 Jun 2026

Sources: Google Cloud “Model versions and lifecycle” (updated 2 Oct 2026) and the Gemini API “Gemini deprecations” page (updated 1 Oct 2026), both checked on 4 October 2026. Vertex AI now appears in that documentation as Gemini Enterprise Agent Platform: same service, new name.

Look at the last column. The same model runs on two different calendars depending on how you call it. Through Vertex AI (Google Cloud credentials), Gemini 2.5 Flash retires on 20 October. Through the Gemini API with an AI Studio key, it has no shutdown date, although Google has already limited access to people who were using it. The image model goes the other way: already gone in AI Studio, alive on Vertex AI until March. If your company uses both routes, you need two inventories.

What did our own inventory show?

I measured it on 3 October 2026 in Cloud Monitoring, counting model invocations per project since 22 July, which is the window Google uses to decide whether a project is active.

Project (anonymised)ModelMeasured callsWhat it means
A WhatsApp agent SaaS product2.5 Flash~40,000 since 22 JulMigrate first: it serves customers live
A digital news outlet2.5 Flash~15,000 since 22 JulUses thinkingBudget: 0, which changes in Gemini 3
A clinic’s sales dashboard and bot2.5 Flash~10,000 since 22 JulHas code calling 2.5 Pro, unused since July
A medical directory2.5 Flash~3,500 since 22 JulStraightforward migration
A client’s staging environment2.5 Flash-Lite~30,000 per monthWas the most urgent under the email timeline
An industrial quoting tool2.5 FlashActive (not quantified)Migrate with regression tests
An internal test project2.5 Flash20Decide whether it should still exist
Three projects with no calls since 31 Jul2.5 Flash0Blocked on 20 October

Three things surprised me.

A staging environment busier than several production systems. Thirty thousand calls a month in an environment nobody treats as critical. If it breaks, nobody notices until someone tries to test something and it fails.

Code calling a model the project will no longer be allowed to use. Gemini 2.5 Pro had exactly one call in the whole period (25 August, in the WhatsApp product). Two other projects had zero calls since 22 July, yet one of them contains a function that calls it by name. That function will fail on the day somebody needs it.

Dead code pointing at models that are already gone. In one project’s retrieval engine I found references to gemini-2.5-pro-preview-05-06 and gemini-2.0-flash-001. Both are out of service. Nobody had noticed because that path never runs, which is exactly the kind of path that runs on the worst possible day.

What actually breaks when you switch models?

Changing the model name is one line. What breaks is everything around it. This is what Google’s official migration guide says, and what it means in practice:

  • Reasoning control changes shape. In Gemini 3 the numeric thinking_budget is deprecated in favour of thinking_level, which takes levels (minimal, low, medium, high). There is no exact zero. Two of our projects deliberately switch reasoning off with thinkingBudget: 0; after migrating we will have to measure how much the lowest level still thinks.
  • The default is not the minimum. In Gemini 3.5 Flash the default reasoning effort is medium. If you leave it unset, you pay for reasoning your task may not need.
  • Temperature, top_p and top_k are deprecated across the 3.x family. If your system relied on temperature 0 for stable answers, Google recommends getting there with explicit system instructions instead.
  • Thought signatures become mandatory in multi-turn conversations: if they are missing, the model returns an error instead of a warning. The clinic dashboard already forwards complete response parts, so no risk there.
  • Function calling requires strict matching. Every function response must include the id of the call it answers. If not, the model tends to return empty responses without raising an error.
  • Token counts may go up. Google warns that response schemas and function calls are now counted in full, where they used to be undercounted. Your bill can grow without a single code change.

The reasoning point is not a footnote. In one of our products I measured 1,955 thinking tokens against 505 tokens of useful output per call. That thinking is billed as output and never shows up anywhere in the response.

What we learned

Reasoning can eat the output limit. If a call has reasoning on and a tight output-token limit, the model spends it thinking before writing the answer. An empty JSON, or one cut in half, comes back, the code cannot parse it and falls back to a canned message or default document. No errors in the logs: the output looks like a normal answer, just a wrong one. The lesson for this migration: check every call, not just the main one. In each, confirm how reasoning is set and whether the output limit is enough.

References to retired models only show up if you look for them. During the inventory I saw that gemini-2.0-flash, shut down on 1 June, was still referenced in five projects, and that gemini-2.5-flash-image, shut down in AI Studio on 2 October, appeared in four. We found them by doing the inventory by hand, not through anything that was watching.

My own notes had the wrong date. I recorded the email timeline and treated it as official. Had I not gone back to the documentation before writing this, I would have told more than one client they had until March.

How do you prepare this week?

  1. Call inventory per project. In Cloud Monitoring, query the model invocation metric grouped by model and project, starting 22 July. Any project with zero calls is a blocking candidate.
  2. Code inventory. Search for the string gemini- across code, environment variables, configuration stored in databases and one-off scripts. Include the code you believe is dead.
  3. Split Vertex AI from the Gemini API. Tag each call by route: Google Cloud credentials or AI Studio key. They follow different calendars.
  4. Pick the replacement per task, not per project. A short classification and a customer conversation do not need the same model, even if they live in the same service.
  5. Translate the parameters. thinking_budget becomes thinking_level; drop temperature, top_p and top_k; check thought signatures and the id in function responses.
  6. Regression tests with real cases. Build a set of 30 to 50 real, anonymised inputs per task, with the output you consider correct today. Compare valid JSON, length limits, accents and special characters, empty responses, finish reason and latency.
  7. Measure reasoning before going live. Read usageMetadata.thoughtsTokenCount on test calls with the old and new model. For formatting tasks, start at the lowest level.
  8. Watch cost for 72 hours. In the billing console, group by SKU and project and review daily spend after the switch. Token counts can rise even if nothing else changes.
  9. Model name in configuration, not in code. Rolling back should mean changing a variable, not shipping another deploy.
  10. Decide what to do with inactive projects. If nobody uses them, shut them down. If someone will need them, migrate straight to Gemini 3.

The order I would follow: first, whatever serves people live, like WhatsApp agents; then anything with high volume even if it looks secondary, like that staging environment; next, code that calls models with no recent usage; and last, inactive projects, which are a business decision more than a technical one.

If your WhatsApp agent runs on Gemini, this guide to WhatsApp AI agents explains how they work and what to decide before putting one in front of your customers. And if you want to see the rest of my daily stack, it is in the AI tools I use every day.

Frequently asked questions

When does Gemini 2.5 Flash retire on Vertex AI? According to the official Google Cloud documentation, checked on 4 October 2026, on 20 October 2026. The same date applies to Gemini 2.5 Pro and 2.5 Flash-Lite. Some customers received an email mentioning shutdowns in January and March 2027 for active projects, but that extension does not appear on the public page.

Will my project stop working on 20 October? If it did not call Gemini 2.5 between 22 July and 20 October, Google’s notice says it will be blocked. If it was active, the email says it keeps working for a while longer, but I would plan as if the date were 20 October, which is what the documentation states.

Does this affect the Gemini API with an AI Studio key the same way? No. On the Gemini API, 2.5 Flash has no announced shutdown date, although access is now limited to existing users. Same model, different calendars, which is why you should inventory them separately.

Which model should I migrate to? Google suggests gemini-3.8-flash or gemini-3.5-flash to replace 2.5 Pro and 2.5 Flash, and gemini-3.1-flash-lite for Flash-Lite. My advice is to decide per task and test on your own real cases, not from the table.

What happens to thinkingBudget: 0 in Gemini 3? The numeric parameter is deprecated and replaced by reasoning levels. There is no exact zero, so measure how many thinking tokens the lowest level uses on your task before you go live.

How long does a migration take? It depends on how many calls you have and how scattered they are across the code. The slow part is not switching the model, it is building regression tests from real cases. That is why I recommend starting there.

If you would rather have us run the inventory and migration with you, this is how I work as an AI consultant.

Share
Related posts