●CLI 0.62.0 — Gemini CLI v0.62.0 (Sep 29) adds Gemini 3.8 Flash and 3.5 Flash-Lite support, along with a fix that keeps the OAuth refresh token●10/02 — gemini-2.5-flash-image reaches its shutdown date today. The place to move is gemini-3.1-flash-image, and sweeping for leftover references is this week's job●TOOLS — An issue reports that Gemini CLI returns a 400 once MCP tools exceed 128. How you trim the tool count is the real question●NEW — The Gemini features that landed in Docs, Sheets, Slides and Drive in March: what I kept and what I dropped six months later●CLAUDE — Handing Claude Code's Japanese-text work to Gemini 3.8 is a pairing getting read on Zenn. The real question is which jobs to hand over and which to keep●12/31 — 90 days left until the introductory pricing ends for 3.8 / 3.7 / 3.6 Flash and robotics-er-2. 3.5 Flash stays as is●CLI 0.62.0 — Gemini CLI v0.62.0 (Sep 29) adds Gemini 3.8 Flash and 3.5 Flash-Lite support, along with a fix that keeps the OAuth refresh token●10/02 — gemini-2.5-flash-image reaches its shutdown date today. The place to move is gemini-3.1-flash-image, and sweeping for leftover references is this week's job●TOOLS — An issue reports that Gemini CLI returns a 400 once MCP tools exceed 128. How you trim the tool count is the real question●NEW — The Gemini features that landed in Docs, Sheets, Slides and Drive in March: what I kept and what I dropped six months later●CLAUDE — Handing Claude Code's Japanese-text work to Gemini 3.8 is a pairing getting read on Zenn. The real question is which jobs to hand over and which to keep●12/31 — 90 days left until the introductory pricing ends for 3.8 / 3.7 / 3.6 Flash and robotics-er-2. 3.5 Flash stays as is
Keeping Gemini API's Default-Model Shift From Becoming an Incident — Pinning Model IDs and Detecting Silent Upgrades in Production
When the default model quietly moves up, your output length, reasoning behavior, and cost change with zero code edits. This guide shows how to pin model IDs in a single source of truth and verify the effective model from the response to detect default changes.
One morning I was scanning the nightly batch logs and noticed the output was about 20% shorter than usual. I hadn't changed a single line of code. Only the cost graph had crept up against the previous day. Tracing it back, I found that the path calling the model through an alias — not an explicit ID — had started receiving responses from a different model. The default had quietly shifted.
On June 8, 2026, Gemini Enterprise locked its default model to 3.5 Flash and removed the disable toggle. The same exposure exists on the API side: any automation that relies on aliases or "no model specified" can, on some random day, start getting answered by a different model. The problem isn't whether the model is better. The incident is the change happening without you knowing about it.
Here's the design that actually held up across several apps where I run the Gemini API as an indie developer. It's the mechanism I built so that the late night I spent chasing this never has to happen twice.
Why an alias reference becomes a silent incident
An alias like gemini-flash-latest, or letting the SDK pick its default, is convenient the day you write it — you always get the newest model. But that "auto-upgrade" property wears two faces in production.
The first is behavioral change. Across generations, the same prompt yields different output length, formatting, and thinking depth. Any downstream step running regex or a JSON schema breaks quietly right here.
The second is cost change. When the responding model changes, the unit price changes. For a batch firing 100,000 calls a day, even a few tens of percent of price movement swings the monthly bill hard.
At minimum, pin these five things: the model ID, generation parameters (temperature, max_output_tokens), thinking settings, safety settings, and the "model generation you expect." That last one is verification metadata — the baseline for the guard below.
Verify the effective model with model_version
This is the heart of the article. A Gemini API response carries model_version, which tells you the model that actually answered — not what you requested, but what the server responded with. If you compare it against your expectation in a startup smoke call, you catch a default change immediately.
from google import genaifrom google.genai import typesclient = genai.Client(api_key="YOUR_GEMINI_API_KEY")# Single source of truth (this is the only thing you switch per environment)EXPECTED_MODEL = "gemini-2.5-pro" # explicit ID, never an aliasEXPECTED_VERSION_PREFIX = "gemini-2.5-pro" # expected model_version prefixdef assert_pinned_model() -> str: """Call once at startup. Fail fast if the effective model differs.""" resp = client.models.generate_content( model=EXPECTED_MODEL, contents="ping", config=types.GenerateContentConfig(max_output_tokens=8), ) actual = resp.model_version or "" if not actual.startswith(EXPECTED_VERSION_PREFIX): raise RuntimeError( f"model drift detected: expected '{EXPECTED_VERSION_PREFIX}*', " f"got '{actual}'. Abort the deploy." ) return actualif __name__ == "__main__": print("pinned model OK:", assert_pinned_model())
Calling assert_pinned_model() at app startup or the head of a batch is enough to prevent the worst case: production running on for hours while answered by an unexpected model. Failing loudly is the point. A hard stop is safer than quietly continuing.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦A startup guard that verifies the model that actually answered using response.model_version
✦Why alias references silently break, and the 5 settings to codify in a single source of truth
✦A 7-day migration playbook, plus a MAD-based behavioral baseline that catches drift even when the model ID never changes
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Beyond the startup check, recording model_version on every production response makes post-incident analysis far easier. Because every response is tied to its effective model, you can later say exactly when behavior changed.
import logginglogger = logging.getLogger("gemini.model_guard")def generate_with_guard(prompt: str): resp = client.models.generate_content( model=EXPECTED_MODEL, contents=prompt, config=types.GenerateContentConfig( temperature=0.4, max_output_tokens=2048, ), ) actual = resp.model_version or "unknown" um = resp.usage_metadata logger.info( "model=%s in_tok=%s out_tok=%s", actual, getattr(um, "prompt_token_count", None), getattr(um, "candidates_token_count", None), ) if not actual.startswith(EXPECTED_VERSION_PREFIX): # Not worth crashing, but it must reach a channel you'll notice logger.error("MODEL DRIFT at runtime: got %s", actual) notify_ops(f"Gemini model drift: {actual}") # to Slack, etc. return resp
Logging token counts alongside is the practical trick. Because token consumption shifts when the model changes, correlating a model_version change with a consumption change lets you explain a cost spike on the spot.
Snapshot the model registry in CI to gate diffs
Manual review always misses something. So collect your pinned model settings into one file and gate diffs in CI — to stop unintended changes (someone hurriedly reverting to an alias, say) before they merge.
The steps are:
Write per-environment model IDs and parameters into model_registry.json (the single source of truth).
The app reads this registry at startup and refuses to boot if any entry contains an alias like -latest.
CI runs tests for "no aliases" and "matches the expected prefix."
Any change must pass review. Make the diff itself visible.
import json, re, sysFORBIDDEN = re.compile(r"(latest|preview|exp)$")def check_registry(path="model_registry.json") -> int: reg = json.load(open(path, encoding="utf-8")) errors = [] for env, cfg in reg.items(): model = cfg.get("model", "") if FORBIDDEN.search(model): errors.append(f"{env}: alias not allowed -> {model}") if "expected_version_prefix" not in cfg: errors.append(f"{env}: expected_version_prefix is undefined") for e in errors: print("NG:", e) return 1 if errors else 0if __name__ == "__main__": sys.exit(check_registry())
This gate is a tiny test, but its leverage is large: it puts your production model selection under code review.
A 7-day playbook to turn a default change into an adoption
Detecting and stopping a default change is defense. When the new default is genuinely better, switch to offense and adopt it deliberately. I recommend this order.
Days 1–2: evaluate the new model offline with your real production prompts. On ~100 representative inputs, line up output length, JSON validity, and latency against the old model. Day 3: run a regression test against your golden dataset to confirm downstream regex and schemas don't break. Days 4–5: route about 5% of traffic to the new model and watch error rate and token consumption broken down by model_version. Day 6: if clean, update the registry's model and expected_version_prefix to the new ID and pass the CI gate. Day 7: cut over fully, keeping instant rollback to the old model available for 24 hours.
This is where the model_version logs you've been collecting pay off. You compare the same metrics before and after with real data, not guesswork. Not "it feels better" but "median output length dropped 18%, JSON validity unchanged at 99.6%." That precision is what makes a production switch calm.
The ID can match while the behavior quietly moves
Everything so far watches which model answered. Which means it only catches drift in the ID.
The same gemini-2.5-pro can keep answering while the server-side default for thinking allocation or a safety threshold gets retuned, and the output changes anyway. That's the failure that caught me later. The guard stayed silent while the JSON parse failure rate downstream went from 0.2% to 1.4% over a weekend.
So alongside the ID check I started keeping a behavioral baseline. A few dozen frozen inputs run once a day, with the output statistics recorded. The point isn't the absolute numbers — it's how far today sits from the recent distribution.
import json, statistics, datetime, pathlibBASELINE_PATH = pathlib.Path("behavior_baseline.jsonl")PROBES = [...] # 20-50 representative production inputs, frozen verbatimdef run_probes() -> dict: lengths, valid, thinking_ratio = [], 0, [] for p in PROBES: resp = generate_with_guard(p) text = resp.text or "" lengths.append(len(text)) try: json.loads(text) valid += 1 except json.JSONDecodeError: pass um = resp.usage_metadata think = getattr(um, "thoughts_token_count", 0) or 0 out = getattr(um, "candidates_token_count", 0) or 1 thinking_ratio.append(think / out) return { "date": datetime.date.today().isoformat(), "model": EXPECTED_MODEL, "len_median": statistics.median(lengths), "json_valid_rate": valid / len(PROBES), "thinking_ratio_median": statistics.median(thinking_ratio), }
Freezing the probe inputs is the precondition. Swap them out whenever it's convenient and the ground shifts under the comparison, which strips the baseline of any meaning.
The check uses median and MAD (median absolute deviation) rather than mean and standard deviation. Probe runs occasionally include one wildly off result, and a mean lets that single run drag the baseline with it.
def mad(xs): m = statistics.median(xs) return statistics.median([abs(x - m) for x in xs]) or 1e-9def check_drift(today: dict, history: list[dict], k: float = 4.0) -> list[str]: alerts = [] for key in ("len_median", "json_valid_rate", "thinking_ratio_median"): past = [h[key] for h in history[-14:]] if len(past) < 7: continue # while history is thin, record only base, spread = statistics.median(past), mad(past) if abs(today[key] - base) > k * spread: alerts.append( f"{key}: today={today[key]:.4f} base={base:.4f} mad={spread:.4f}" ) return alerts
k=4.0 is where false positives settled below one a week against fourteen days of history. At 3.0 it fired on ordinary variance, and I picked up the habit of ignoring the alert — worse than having no alert at all. Your threshold will differ, so record for two weeks and look at the spread before committing to a number.
With the baseline in place, changes that happen while model_version never moves finally get a name. Not "output feels shorter lately" but "median output length is 23% below the 14-day base, thinking ratio flat." Isolating a cause starts from that sentence.
Where pinning doesn't reach — caches, batches, and taking inventory
Even with the registry as a single source of truth, some paths let the model spec slip past. Cutover day is exactly where that bit me.
Context caching is the first. A cache is bound to the model ID it was created with, so the moment the registry moves to a new ID, existing caches are no longer reachable from the new model. While the TTL lingers, one path grabs an old-ID cache and another generates on the new ID — same feature, two flavors of response.
Batch is the second. Submission and completion are separated in time, so any job submitted mid-cutover completes under the old settings. Writing the registry revision into the job metadata at submission time lets you sort the results out afterward.
Path
When the model is resolved
What cutover requires
Plain generate_content
At request time
Updating the registry is enough
Context cache
At cache creation
Rebuild the cache on the new ID and explicitly delete the old one
Batch job
At submission
Pause submissions around cutover and drain in-flight jobs
Tuned model
At tuning time
Track the base model's retirement date separately
Above all, don't keep "where do we specify the model" in human memory. This script runs in CI and flags any literal that bypasses the registry.
import re, pathlib, sysMODEL_LITERAL = re.compile(r"""["'](gemini-[a-z0-9.\-]+)["']""")ALLOWED = {"model_registry.json", "model_registry.py"}def audit(root: str = "src") -> int: hits = [] for path in pathlib.Path(root).rglob("*.py"): if path.name in ALLOWED: continue for i, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1): m = MODEL_LITERAL.search(line) if m: hits.append(f"{path}:{i}: {m.group(1)}") for h in hits: print("hardcoded model:", h) return 1 if hits else 0if __name__ == "__main__": sys.exit(audit())
The first run surfaced model IDs in places I'd forgotten about: a small verification script, a docs-generation tool, and a migration script I'd written as a one-off. None of them serve production traffic. All of them were more than capable of quietly skewing the numbers I was comparing against.
Let the machine take inventory instead of your memory. Adding that one job to CI made the whole cutover far easier to see through.
Pitfalls and how to avoid them
In practice, a few things trip you up.
model_version won't always match your requested ID exactly. A minor suffix can be appended, so write the guard as a prefix match, not an exact match. Exact matching would crash production on a harmless patch update.
It's tempting to keep aliases "because they're handy in dev," but when the effective model diverges between dev and production, you breed bugs that only reproduce in production. Using the same explicit ID in dev, and gating new-model trials behind a separate verification flag, turned out safer in the end.
Finally, the smoke call costs money too, of course. Keep max_output_tokens minimal and limit it to startup, and it stays negligible. Skimping a few cents on safety to miss a silent incident costs far more.
The default moving up is an unstoppable trend. Rather than bracing against the unstoppable, build the state where you always notice when it moved — first. That small preparation lets you finish the late-night investigation exactly once. I hope it helps anyone working on the same problem.
Share
Thank You for Reading
Gemini Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.