◉GEMINI LABJP
●VEO 3.1 — The three veo-3.1 preview models shut down Oct 22, 14 days left. Move to gemini-omni-1.1-flash●OMNI — gemini-omni-flash-preview also shuts down Oct 22, not Sep 30. Check the deprecations table●NANO 2.1 — gemini-nano-banana-2.1 is GA. gemini-3.1-flash-image is deprecated with no shutdown date yet●CLI 0.63 — Gemini CLI stable is v0.63.0, with clearer retry progress on reconnect●TTS/LIVE — 2.5-era audio and TTS previews shut down Nov 17, 40 days left●NEW — Veo 3.1 previews stop Oct 22: take inventory before you swap●VEO 3.1 — The three veo-3.1 preview models shut down Oct 22, 14 days left. Move to gemini-omni-1.1-flash●OMNI — gemini-omni-flash-preview also shuts down Oct 22, not Sep 30. Check the deprecations table●NANO 2.1 — gemini-nano-banana-2.1 is GA. gemini-3.1-flash-image is deprecated with no shutdown date yet●CLI 0.63 — Gemini CLI stable is v0.63.0, with clearer retry progress on reconnect●TTS/LIVE — 2.5-era audio and TTS previews shut down Nov 17, 40 days left●NEW — Veo 3.1 previews stop Oct 22: take inventory before you swap
Articles/Gemini Basics
◉ Gemini Basics/2026-10-08Beginner

Free Tier Shows Up, but You Get a 429 With `limit: 0` — What to Check in the First Hour of a New Gemini API Project

A new Gemini API project says Free tier, yet the very first call returns 429 with limit: 0. Here is the order I check things in, how to read the error body, and a 40-line script that tests models one by one.

Gemini API247429 errorfree tier3RESOURCE_EXHAUSTEDtroubleshooting85beginner14

You create a new project, generate an API key, and the console clearly says Free tier. You paste the sample code, run it, and the very first call comes back with a 429. I see this question again and again on developer forums.

Most people's first thought is "I must have sent too many requests," so they wait. But a 429 that appears on the first call does not go away with waiting. The cause is not "you used it up" but "there was never any quota to begin with."

The first thing I would like to share is how to read limit: 0 in the error body. Once that is clear, the rest of the checking becomes almost mechanical.

limit: 0 Means "No Quota Assigned," Not "Too Many Requests"

An ordinary 429 happens when you exceed a fixed allowance — the limit is 10 and you sent an 11th. A limit: 0, on the other hand, says that this combination of model and project has no allocation at all.

So what you need to fix is not how often you call. It is which model you call, from which project, under which conditions. As the official rate limits page explains, limits are set per model and per project. That is why the same key can succeed with one model and fail with another.

The "Free tier" label in the console describes the project's tier. It is not a promise that every model is free. Mixing those two up is how people end up circling through settings screens for an hour.

The Order I Check Things in (Five Steps)

Before touching settings at random, decide the order. Mine goes top to bottom, because the upper steps settle the question fastest and the lower ones take the longest.

StepWhat to checkHow to settle it
1The model name you callTry one call with a different model. If it works, that model simply has no quota for you
2Key-to-project matchConfirm the project you are looking at in the console owns the key your code uses
3`quotaMetric` and `quotaValue` in the error bodyRead which metric is zero (the next section automates this)
4Supported region and billingCompare against the official [billing page](https://ai.google.dev/gemini-api/docs/billing) and the free-tier region notes
5Propagation delay after a changeIf you just linked billing, wait a little and rerun the same call

Steps 1 and 2 take minutes. Step 3 often narrows the cause down, and steps 4 and 5 are only worth doing after you have seen what step 3 says. Just keeping this order cuts down a lot of settings-screen wandering.

A 40-Line Script to Test Models One by One

Trying models by hand is tedious, and the error body is long and hard to read. So here is a script that lists the models your key can see, sends one minimal request to a few of them, and prints only quotaMetric and quotaValue from the error. It uses the standard library only.

import json, os, re, sys, urllib.request, urllib.error
 
KEY = os.environ["GEMINI_API_KEY"]  # keep the real key out of your code
BASE = "https://generativelanguage.googleapis.com/v1beta"
 
def call(path, body=None):
    data = json.dumps(body).encode() if body else None
    req = urllib.request.Request(
        f"{BASE}/{path}", data=data,
        headers={"x-goog-api-key": KEY, "Content-Type": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=30) as r:
            return r.status, json.load(r)
    except urllib.error.HTTPError as e:
        return e.code, json.loads(e.read() or b"{}")
 
# 1) List the models this key can see (does not consume quota)
_, listing = call("models")
names = [m["name"] for m in listing.get("models", [])
         if "generateContent" in m.get("supportedGenerationMethods", [])]
targets = sys.argv[1:] or [n for n in names if "flash" in n][:3]
print("visible models:", len(names))
 
# 2) Send one minimal request to each target
for name in targets:
    name = name if name.startswith("models/") else f"models/{name}"
    code, res = call(f"{name}:generateContent",
        {"contents": [{"parts": [{"text": "ping"}]}],
         "generationConfig": {"maxOutputTokens": 8}})
    if code == 200:
        print(f"OK    {name}")
        continue
    err = res.get("error", {})
    zero = bool(re.search(r"limit:\s*0", err.get("message", "")))
    for d in err.get("details", []):
        for v in d.get("violations", []):
            zero = zero or str(v.get("quotaValue")) == "0"
            print(f"  metric={v.get('quotaMetric')} value={v.get('quotaValue')}")
    kind = "zero quota" if zero else "used up, or another cause"
    print(f"{code}  {name}  -> {kind}")

The listing call (models) does not generate anything, so it does not consume quota. If no models show up at all, that points you at the key or the project, which is a useful branch in itself.

The request is minimal on purpose. I want to confirm in one shot that the failure comes from the quota, not from the content. With a small maxOutputTokens, a successful call costs next to nothing.

To use it, put your key in an environment variable and run python3 probe.py. Pass the model names you actually use as arguments to test only those.

Reading the Output

The result usually falls into one of three shapes.

OutputMeaningNext move
Some models OK, others "zero quota"This project has no quota for those modelsSwitch to a model that works and keep going; check billing if you need the others
Every model "zero quota"Project, region, or billing is the causeGo to steps 2 and 4
"used up, or another cause"Quota exists, so it is probably an ordinary rate limitSpace out retries; the design is covered in [Don't Retry Every Gemini 429 — Telling Rate Limits Apart From Spend Cap Exhaustion](/en/articles/gemini-api/gemini-api-429-retryable-vs-spend-cap-exhaustion-retry-design)

The first shape is almost anticlimactic. Learning that "that model just wasn't available" and getting unstuck on the spot is quite possible.

Why I Prefer a Probe Over Guessing

It is tempting to open the console and click through every settings page until something looks wrong. I have done it, and it rarely pays off, because most of those pages show the project's state, not the state of the particular model you are calling.

A probe flips that around. It asks the API itself, model by model, and the answer comes back in the same words the API will use later in production. That makes the result easy to paste into a note, compare next week, and hand to someone else.

It also keeps the problem small. One request per model is enough to separate "this model has no quota here" from "something is wrong with the whole project," and those two lead to very different next steps.

If It Still Doesn't Work, Gather These Before Asking

If you cannot narrow it down by now, you will end up asking on a forum or to support. Having these five things ready makes the reply faster and more accurate.

  1. The model name you called
  2. quotaMetric and quotaValue from the error body
  3. When it happened, with the time zone
  4. The project identifier (never paste the key itself)
  5. Whether you set up billing, and when

The same symptom is discussed in a forum thread, and my impression is that posts with these details get to a cause much sooner.

Closing

If I had to put today's point in one sentence:

A 429 on the very first call is something to read, not something to wait out.

The next step is small and fits in today. Run the script above once, and write down which models show "zero quota" for your project. Having that list in hand makes the next encounter with the same screen feel very different.

I now put this check first whenever I set up a new environment.

Share

Thank You for Reading

Gemini Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

If you found this article helpful, a small tip ($1.50) would mean a lot to us. Your support helps keep this site ad-free and covers server and hosting costs.

Related Articles

◉ Gemini Basics2026-08-28
What decides whether your Gemini API data trains Google's models, and the one exception on the paid tier
Whether your Gemini API prompts feed model training is not settled by the paid tier alone. Here is how AI Studio and the API define paid differently, how dataset sharing reverses the protection, and what changes by region.
◉ Gemini Basics2026-09-13
I edited the Drive doc, asked Gemini again, and it quoted the old sentence back
When Gemini answers from an outdated version of a Drive file, the file usually isn't the problem — the conversation is. Here's how sources stay attached for the whole conversation, why you can't remove one mid-thread, and what to do about partial reads.
◈ API / SDK2026-05-15
Three Places Gemini API Embedding Broke on Me — and What Actually Fixed Them
Notes from wiring Gemini API Embedding into an auto-categorization pipeline: INVALID_ARGUMENT from dimension mismatches, 429 rate limits that lost 400 items of work, and weak retrieval quality — plus normalization, checkpointing, and a recall@5 harness that made improvements measurable.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links