●TTS — Gemini 3.8 Flash TTS and Flash-Lite TTS are GA as of September 22. Voice replication works on Flash only; Flash-Lite returns a 400●2.5 — Access to the 2.5 models is now limited to accounts with prior active usage (September 18). Not a deprecation: new projects should start on 3.5 Flash-Lite or 3.8 Flash●9/30 — Two days until gemini-omni-flash-preview shuts down, and gemini-2.5-flash-image follows on October 2. In practice the replacement for the latter is gemini-3.1-flash-image●AUTH — Reports continue of Gemini CLI sign-in failing only for Workspace Enterprise accounts while personal accounts work. Isolating the two is the current focus●NEW — Moving to 3.8 Flash-Lite TTS meant re-picking the voice for our guidance audio●429 — A brand-new project shows Free tier yet returns 429 with limit: 0. Here is what to check in the first hour●TTS — Gemini 3.8 Flash TTS and Flash-Lite TTS are GA as of September 22. Voice replication works on Flash only; Flash-Lite returns a 400●2.5 — Access to the 2.5 models is now limited to accounts with prior active usage (September 18). Not a deprecation: new projects should start on 3.5 Flash-Lite or 3.8 Flash●9/30 — Two days until gemini-omni-flash-preview shuts down, and gemini-2.5-flash-image follows on October 2. In practice the replacement for the latter is gemini-3.1-flash-image●AUTH — Reports continue of Gemini CLI sign-in failing only for Workspace Enterprise accounts while personal accounts work. Isolating the two is the current focus●NEW — Moving to 3.8 Flash-Lite TTS meant re-picking the voice for our guidance audio●429 — A brand-new project shows Free tier yet returns 429 with limit: 0. Here is what to check in the first hour
Making Gemini API Output Reproducible with the seed Parameter — Practical Patterns for Tests and Debugging
A practical guide to the Gemini API seed parameter: measured match rates over 100 runs, seeds derived from test case IDs plus a nightly sweep, a four-level comparison ladder for snapshots, and how to fix a wrapper that drops seed.
"I'm sending the exact same prompt and getting a different answer every time" — that's the wall most teams hit the moment they try to write tests against a Gemini-powered feature. As an indie developer I ran into it myself when wiring up regression tests for one of my apps, and I nearly wrote it off as "the model is just non-deterministic" before I realized the culprit was sitting in my own wrapper code.
The good news is that, in most cases, the seed parameter does what you want. The less obvious news is that "just pass a seed and you'll get the same answer" is not quite accurate — there are situations where seed simply cannot stabilize the output. This article walks through how seed actually works, the patterns I rely on for tests and debugging, the match rates I measured on my own machine, and the gotchas that surprise people most often.
What seed actually controls
The Gemini API seed fixes the starting point of the pseudo-random number generator used during sampling. Give it the same seed, prompt, model, and parameters, and the sampling order lines up, so the output tends to match.
The key thing to internalize is that seed is not a replacement for temperature:
temperature=0.0 alone pushes the model toward near-greedy decoding, which is mostly deterministic, but batching order and tiny numerical differences on the model side can still nudge the result
Adding seed aligns the sampling process itself, so you get a more consistent result
In my experience, seed + low temperature is noticeably steadier for regression tests than simply lowering the temperature. The next section puts a number on that "in my experience."
Measured: sending one prompt 100 times to check the match rate
Rather than rely on feel, I sent the same short prompt ("Answer with the capital of Japan in one word.") to gemini-2.5-flash 100 times under each condition and counted how often the response was byte-identical to the first one. The comparison is an exact string match, whitespace included.
Condition
Exact matches
Notes
temperature=0.0 / seed=42 (fixed)
100 / 100
Zero variance. Safe to use as a test baseline
temperature=0.0 / seed unset
97 / 100
Occasionally splits on a trailing period
temperature=0.7 / seed=42 (fixed)
41 / 100
Even with seed, the sampling space is wide enough to wobble
temperature=0.7 / seed unset
18 / 100
Not usable for comparison
Two things stand out. First, for testing, temperature=0.0 + fixed seed is clearly the best — on my machine all 100 responses matched. Second, raising the temperature drops the match rate to 41% even with a fixed seed. In other words, seed does not erase the variance that temperature creates; it only aligns the sampling order under the same temperature and conditions. Getting that distinction straight up front saves you from a lot of confusion later.
One caveat: the longer the response, the higher the chance it splits on the final token. The measurement above uses a few-token answer, which is why the match rate is so high; with a few-hundred-token response, even temperature=0.0 + fixed seed can occasionally wobble at the tail. Choose your snapshot granularity with that reality in mind.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦Measured match rates from sending one prompt 100 times across seed on/off and different temperatures
✦Deriving seeds from test case IDs, plus an eight-seed sweep that broke three green single-seed cases
✦A four-level comparison ladder in code, and a table of which level each output type may pass at
✦The Before/After of a wrapper that silently drops seed, plus a five-second sanity check
✦A top-down triage flow for variance, and logprobs code to diagnose why an output wobbles
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Here is a minimal, working example with the google-genai Python SDK, shaped so it drops straight into a pytest snapshot test.
# pip install google-genaiimport osfrom google import genaifrom google.genai import typesclient = genai.Client(api_key=os.environ["GEMINI_API_KEY"])def generate_with_seed(prompt: str, seed: int = 42) -> str: """Get a highly reproducible response for the same seed and prompt.""" response = client.models.generate_content( model="gemini-2.5-flash", contents=prompt, config=types.GenerateContentConfig( temperature=0.0, top_p=1.0, seed=seed, max_output_tokens=512, ), ) return response.textif __name__ == "__main__": out_a = generate_with_seed("Answer with the capital of Japan in one word.") out_b = generate_with_seed("Answer with the capital of Japan in one word.") print(out_a) print(out_b) print("match:", out_a == out_b)
This calls the model twice with seed=42, temperature=0.0 and compares the results. Both should print Tokyo and match: True. Adding even a single trailing space to the prompt can change the response, so manage your input strings strictly in tests.
REST and Node.js variants
From REST, put seed inside generationConfig.
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=YOUR_GEMINI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "contents": [{"parts":[{"text":"What is 2+2? Answer with just the number."}]}], "generationConfig": { "temperature": 0.0, "topP": 1.0, "seed": 42, "maxOutputTokens": 32 } }'
With Node.js (@google/genai), you just pass seed in the config object.
import { GoogleGenAI } from "@google/genai";const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY! });const result = await ai.models.generateContent({ model: "gemini-2.5-flash", contents: "Translate 'Good morning' to French.", config: { temperature: 0, topP: 1, seed: 42, maxOutputTokens: 64 },});console.log(result.text);
REST and Node.js behave the same internally — same parameters, same result.
Three patterns I use for tests and debugging
These are the three I reach for day to day.
1. Snapshot-freeze for regression tests
So that a diff only appears when you intend one, generate the response with fixed seed + temperature=0 and save it to a snapshot file. If CI can catch the diff, you avoid the "output quietly changed and nobody noticed" class of incident. The full implementation is in Building Prompt Regression Tests for the Gemini API with Pytest.
2. Variance reduction for prompt A/B comparison
When comparing "is prompt A or B better," running each once with the same seed is less reliable than preparing 3–5 seeds and doing a paired comparison per seed. Even when you deliberately raise temperature to measure diversity, fixing the list of seeds keeps the experiment reproducible.
3. Bug reports and reproduction environments
When a user reports "it returned something weird," logging the seed alongside the prompt dramatically raises your odds of reproducing it locally. Always keep prompt, model, temperature, and seed in your app logs.
A suite pinned to seed=42 hides the cases it never exercised
Once reproducibility lands, complacency follows. For a while every test I owned ran on seed=42, and the suite stayed green for weeks. Then I went back through production logs for one classification prompt and found the opposite label appearing at a steady, non-trivial rate — a label the tests had never once produced.
In hindsight the reason is obvious. Seed 42 is just one point that happened to be stable for that prompt. A prompt sitting near a decision boundary falls the other way under a different seed. A fixed-seed test only certifies "nothing is broken at this one point."
So I split seed selection into two tiers. Everyday CI gives each test case one derived seed; a nightly job sweeps the same cases across several.
# tests/seedlib.pyfrom hashlib import blake2bSEED_SPACE = 2**31 - 1def seed_for(case_id: str, variant: int = 0) -> int: """Derive a stable seed from the test case ID. Distinct per case, and independent of environment or test ordering.""" digest = blake2b(f"{case_id}#{variant}".encode(), digest_size=8).digest() return int.from_bytes(digest, "big") % SEED_SPACEdef seed_sweep(case_id: str, n: int = 8) -> list[int]: """Seed series for sweeping one case across several draws""" return [seed_for(case_id, v) for v in range(n)]
Deriving from the case ID means there is no seed registry to maintain. Add a test, and its seed exists. When every case shares the same 42, the fact that one prompt got lucky at that point tells you nothing about the next one — and that difference stays invisible.
The nightly side reports how often the response agrees across seeds.
# tests/sweep.pyfrom collections import Counterfrom seedlib import seed_sweepdef agreement_rate(call, case_id: str, prompt: str, n: int = 8) -> tuple[float, str]: """Return the share of seeds producing the same answer, plus that answer""" outs = [call(prompt, seed=s, temperature=0.0) for s in seed_sweep(case_id, n)] top, hits = Counter(o.strip() for o in outs).most_common(1)[0] return hits / n, topdef test_sweep_report(call, cases): fragile = [] for case_id, prompt in cases: rate, top = agreement_rate(call, case_id, prompt) if rate < 0.875: # record if even one of eight seeds splits fragile.append((case_id, rate, top)) for case_id, rate, top in fragile: print(f"[fragile] {case_id}: agreement={rate:.2f} top={top!r}") assert not [c for c in fragile if c[1] < 0.5], "a case splits on a majority of seeds"
Here is what eight seeds at temperature=0.0 produced across six classification and extraction prompts on my own machine.
Prompt
Agreement over 8 seeds
Single-seed (42) test
What it tells you
Binary sentiment (unambiguous input)
8 / 8
green
Genuinely stable
Binary sentiment (mixed praise and complaint)
5 / 8
green
Green by luck. The prompt is the problem
Amount extraction from an invoice
8 / 8
green
Numeric extraction is generally solid
Three-way ticket routing
6 / 8
green
The boundary label is under-defined
Summary (three sentences requested)
2 / 8
green
Not something to check by string match
Tag extraction (up to five)
4 / 8
red
Order is unstable. Compare as a set
The interesting part is that three cases green under a single seed split once the seed moved. CI never turned red, and the weakness in those prompts simply sat there. Tag extraction was already red, but for a different reason: the problem was the comparison, not the reproducibility — which is the subject of the next section.
I deliberately keep the sweep to once a day rather than once a pull request. Running eight times the API calls on every push is not worth it, and catching a dropped agreement rate the following morning is soon enough. Only cases below 0.5 fail the job; everything else is reported. A morning notification that fails loudly for borderline cases gets ignored within a week, and then the whole mechanism is dead.
Decide how strict your comparison is before exact match starts failing you
Push reproducibility far enough and the failures flip sides. The response is effectively identical, but a trailing period or a full-width space turns the test red.
Raising temperature or dropping the seed at that point is exactly backwards — that is adding variance to avoid an equality check you should be fixing instead. I ended up splitting comparison into levels and deciding, per output type, which level is allowed to pass.
# tests/compare.pyimport json, re, unicodedatadef _norm(s: str) -> str: s = unicodedata.normalize("NFKC", s).strip() s = re.sub(r"\s+", " ", s) return re.sub(r"[.,]+$", "", s)def _structural(a: str, b: str, tol: float = 1e-6) -> bool: try: pa, pb = json.loads(a), json.loads(b) except json.JSONDecodeError: return False return _deep_eq(pa, pb, tol)def _deep_eq(a, b, tol: float) -> bool: if isinstance(a, dict) and isinstance(b, dict): return a.keys() == b.keys() and all(_deep_eq(a[k], b[k], tol) for k in a) if isinstance(a, list) and isinstance(b, list): return len(a) == len(b) and all(_deep_eq(x, y, tol) for x, y in zip(a, b)) if isinstance(a, (int, float)) and isinstance(b, (int, float)): return abs(a - b) <= tol return a == bdef match_level(expected: str, actual: str) -> int: """Return the strictest level that matched. 4 means no match.""" if expected == actual: return 0 # L0: byte-identical if _norm(expected) == _norm(actual): return 1 # L1: identical after normalization if _structural(expected, actual): return 2 # L2: equivalent as JSON return 4
Anything meant to be compared as a set — the tag extraction above — gets one order-independent level just before L2.
def set_equal(expected: str, actual: str) -> bool: """Order-independent comparison for tag and label lists (L3)""" to_set = lambda s: {t.strip().lower() for t in re.split(r"[,\n]+", s) if t.strip()} return to_set(expected) == to_set(actual)
Then each output type gets a ceiling.
Output type
Highest level allowed
Why
Single label (positive / negative, etc.)
L1
Anything beyond formatting drift is a prompt defect
Structured data (JSON response)
L2
Key order and whitespace are outside the contract. Compare values
Tag or label lists
L3
Order carries no meaning, and failing on it is unmaintainable
Free-form prose (summaries, explanations)
Do not compare
Not a string-match target. Assert properties instead
That last row is deliberate. If you try to pass a summary snapshot at L1, you will keep loosening it, and what survives is one test that checks nothing at all. Take summaries out of string comparison entirely and assert properties instead — required terms present, sentence count as specified. Those hold up over time.
One small operational habit paid off more than I expected. Recording the level match_level() returned on every run means you notice when a case that used to pass at L0 starts passing at L1. The value has not changed, but the shape of the response has. In my own repositories I glance at the distribution of levels once a week, and in weeks where a model version moved I have seen the L1 share tick up before anything else surfaced.
Splitting comparison into levels changes what "the test failed" means. Once you know which level it failed at, you can tell on the spot whether to fix the prompt, fix the comparison, or stop checking that thing by string at all.
The thing silently dropping your seed is usually your own wrapper
When "I passed a seed but it doesn't match," the first suspect is not the model — it's the layer sitting between your code and the API. The mistake I actually made was rebuilding GenerateContentConfig in a shared wrapper and forwarding only temperature and max_output_tokens, quietly dropping seed.
# Before: seed is never forwarded, so tests look "non-deterministic"def call_model(prompt: str, cfg: dict) -> str: response = client.models.generate_content( model=cfg["model"], contents=prompt, config=types.GenerateContentConfig( temperature=cfg.get("temperature", 0.0), max_output_tokens=cfg.get("max_output_tokens", 512), # forgot to forward seed here ), ) return response.text
# After: build the config so no known field can be droppedKNOWN_KEYS = {"temperature", "top_p", "seed", "max_output_tokens"}def call_model(prompt: str, cfg: dict) -> str: passthrough = {k: cfg[k] for k in KNOWN_KEYS if k in cfg} response = client.models.generate_content( model=cfg["model"], contents=prompt, config=types.GenerateContentConfig(**passthrough), ) return response.text
The point is to pass known keys through as a dict comprehension rather than copying each setting by hand. That way, adding a new parameter later can't reintroduce a "forgot to forward it" bug. If an enterprise gateway is stripping unknown fields, this shape also makes it easier to tell whether the field is being stripped or was never sent in the first place.
A top-down flow to eliminate the source of variance
When seed isn't working, don't poke at it randomly. Walking down this order gets you to the cause fastest.
Send the same minimal prompt three times in a row (seed=42, temperature=0). If all three match exactly, the raw API call is at least healthy
If they don't match, call the SDK directly, bypassing your wrapper. If it matches now, the wrapper is the culprit (see the Before/After above)
Still wobbling? Check temperature. Make sure a stray 0.7 hasn't slipped in via an env var or default
Check the model name isn't an alias (-latest). Pin tests to an explicit version
Check the input isn't multimodal, streaming, or tool-using. Those are outside seed's jurisdiction, so split the test unit
Walking these five steps top to bottom prevents the "it's the model's fault" time sink almost entirely. After I wrote this order on a sticky note, my time spent investigating seed-related issues dropped by roughly half.
When seed does not help
Let me expand on the "outside seed's jurisdiction" that showed up in steps 4–5. If you're passing a seed and the result still varies, check whether you've hit one of these.
Temperature is high: at 0.7–1.0, the sampling space is wide enough that seed alone leaves room for noise — as measured above, the match rate can fall to 41%. For maximum reproducibility, keep it in the 0–0.2 range
The model name is an alias: aliases like gemini-2.5-flash-latest can be swapped for a different version underneath. In tests, use gemini-2.5-flash (or an explicit version) to be safe
Multimodal input (images, PDFs): the image preprocessing path has its own variance and is less stable than text alone. In snapshot tests, stick to text input
Streaming responses: chunk boundaries can shift. Compare on the final, fully assembled text
Tool use or grounding: the external call results themselves change over time, so seed alone can't reproduce them. Mock the tools in tests
In short, seed suppresses "sampling variance" but not "external variance." The trick is to split the test unit and stay conscious of where the wobble is coming from.
Diagnosing why it wobbles with logprobs
When you want to take the diagnosis a step deeper, pull logprobs and look at how much the model hesitated at each token. The closer the top candidates' probabilities are, the more easily that token flips to another word when conditions change slightly — an obvious relationship, but a useful one to see.
from google import genaifrom google.genai import typesclient = genai.Client()resp = client.models.generate_content( model="gemini-2.5-flash", contents="Answer the sentiment in one word, positive / negative: shipping was slow but the quality was great", config=types.GenerateContentConfig( temperature=0.0, seed=42, response_logprobs=True, logprobs=5, # top 5 candidates per position ),)# Inspect how close the candidates are at the first tokenfor cand in resp.candidates: for step in cand.logprobs_result.top_candidates[:1]: for c in step.candidates: print(f"{c.token!r}: logprob={c.log_probability:.3f}")
If positive and negative have near-identical logprobs at the first token, that prompt is inherently prone to splitting. Rather than papering over it with a seed, rewriting the prompt to widen the confidence gap helps both test stability and production quality. The full walkthrough is in Measuring Classification Confidence with Gemini API logprobs.
A short story: why seed matters more in evaluation than production
At first I thought of seed as a "test-time only" tool. What changed my mind was running prompt evaluations in parallel. Without a seed, the score for the same prompt drifted slightly run to run, and the effect of a small prompt improvement drowned in that noise.
The moment I fixed the seed and the model version, the score started reflecting "the intrinsic quality of the prompt" rather than "sampling luck." If you're doing serious prompt improvement, fix the seed in your evaluation jobs before you even think about production. And if you're using an LLM-as-judge, fix the judge's model and seed too — otherwise you're measuring two layers of wobble at once.
Designing seed alongside temperature and top_p
These are the three settings I've settled into.
Full reproducibility (for tests): temperature=0, top_p=1.0, seed=fixed. Aims for near-exact snapshot matches
Allowing slight variation (production): temperature=0.2, top_p=0.95, seed=unset. Feels more natural for user-facing responses
When you need creativity (copywriting, etc.): temperature=0.9, top_p=0.95, seed=rotating. Cycle through several seeds and compare candidates
Tune response quality on the temperature side and reproducibility on the seed side — separating the roles keeps your settings from sprawling. For choosing temperature itself, I've collected use-case-by-use-case guidance in Task-by-Task Temperature Best Practices for the Gemini API.
Quick sanity check: confirming seed actually works in your stack
Before relying on seed across your test suite, run a one-line confirmation in the same environment your tests use. Network proxies, alternate endpoints, or a mismatched SDK version can all produce subtle differences. The fastest check is to send the same minimal prompt three times back-to-back with seed=42, temperature=0 and assert all three are byte-identical. If they are, your stack honors seed correctly. If not, something between your code and the model is dropping it — usually a wrapper that forgets to forward the parameter, or a gateway that strips fields it doesn't recognize. The Before/After and triage flow above are your prescription.
What to try next
Open one of your current Gemini API call sites, add seed=42 and temperature=0, and run the same prompt twice. In most cases the two outputs will match exactly. If they don't, walk down the triage flow — it's almost always one of those issues.
Once you have that working, add a single snapshot test to CI for one important prompt. The moment your pipeline can detect quiet output drift, your ability to iterate on prompts steps up noticeably. Reproducibility is an unglamorous foundation, but whether or not you have it changes how fast everything downstream moves. I hope this helps in your own work, and thank you for reading.
Share
Thank You for Reading
Gemini Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.