You create a new project, generate an API key, and the console clearly says Free tier. You paste the sample code, run it, and the very first call comes back with a 429. I see this question again and again on developer forums.
Most people's first thought is "I must have sent too many requests," so they wait. But a 429 that appears on the first call does not go away with waiting. The cause is not "you used it up" but "there was never any quota to begin with."
The first thing I would like to share is how to read limit: 0 in the error body. Once that is clear, the rest of the checking becomes almost mechanical.
limit: 0 Means "No Quota Assigned," Not "Too Many Requests"
An ordinary 429 happens when you exceed a fixed allowance — the limit is 10 and you sent an 11th. A limit: 0, on the other hand, says that this combination of model and project has no allocation at all.
So what you need to fix is not how often you call. It is which model you call, from which project, under which conditions. As the official rate limits page explains, limits are set per model and per project. That is why the same key can succeed with one model and fail with another.
The "Free tier" label in the console describes the project's tier. It is not a promise that every model is free. Mixing those two up is how people end up circling through settings screens for an hour.
The Order I Check Things in (Five Steps)
Before touching settings at random, decide the order. Mine goes top to bottom, because the upper steps settle the question fastest and the lower ones take the longest.
| Step | What to check | How to settle it |
|---|---|---|
| 1 | The model name you call | Try one call with a different model. If it works, that model simply has no quota for you |
| 2 | Key-to-project match | Confirm the project you are looking at in the console owns the key your code uses |
| 3 | `quotaMetric` and `quotaValue` in the error body | Read which metric is zero (the next section automates this) |
| 4 | Supported region and billing | Compare against the official [billing page](https://ai.google.dev/gemini-api/docs/billing) and the free-tier region notes |
| 5 | Propagation delay after a change | If you just linked billing, wait a little and rerun the same call |
Steps 1 and 2 take minutes. Step 3 often narrows the cause down, and steps 4 and 5 are only worth doing after you have seen what step 3 says. Just keeping this order cuts down a lot of settings-screen wandering.
A 40-Line Script to Test Models One by One
Trying models by hand is tedious, and the error body is long and hard to read. So here is a script that lists the models your key can see, sends one minimal request to a few of them, and prints only quotaMetric and quotaValue from the error. It uses the standard library only.
import json, os, re, sys, urllib.request, urllib.error
KEY = os.environ["GEMINI_API_KEY"] # keep the real key out of your code
BASE = "https://generativelanguage.googleapis.com/v1beta"
def call(path, body=None):
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(
f"{BASE}/{path}", data=data,
headers={"x-goog-api-key": KEY, "Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=30) as r:
return r.status, json.load(r)
except urllib.error.HTTPError as e:
return e.code, json.loads(e.read() or b"{}")
# 1) List the models this key can see (does not consume quota)
_, listing = call("models")
names = [m["name"] for m in listing.get("models", [])
if "generateContent" in m.get("supportedGenerationMethods", [])]
targets = sys.argv[1:] or [n for n in names if "flash" in n][:3]
print("visible models:", len(names))
# 2) Send one minimal request to each target
for name in targets:
name = name if name.startswith("models/") else f"models/{name}"
code, res = call(f"{name}:generateContent",
{"contents": [{"parts": [{"text": "ping"}]}],
"generationConfig": {"maxOutputTokens": 8}})
if code == 200:
print(f"OK {name}")
continue
err = res.get("error", {})
zero = bool(re.search(r"limit:\s*0", err.get("message", "")))
for d in err.get("details", []):
for v in d.get("violations", []):
zero = zero or str(v.get("quotaValue")) == "0"
print(f" metric={v.get('quotaMetric')} value={v.get('quotaValue')}")
kind = "zero quota" if zero else "used up, or another cause"
print(f"{code} {name} -> {kind}")The listing call (models) does not generate anything, so it does not consume quota. If no models show up at all, that points you at the key or the project, which is a useful branch in itself.
The request is minimal on purpose. I want to confirm in one shot that the failure comes from the quota, not from the content. With a small maxOutputTokens, a successful call costs next to nothing.
To use it, put your key in an environment variable and run python3 probe.py. Pass the model names you actually use as arguments to test only those.
Reading the Output
The result usually falls into one of three shapes.
| Output | Meaning | Next move |
|---|---|---|
| Some models OK, others "zero quota" | This project has no quota for those models | Switch to a model that works and keep going; check billing if you need the others |
| Every model "zero quota" | Project, region, or billing is the cause | Go to steps 2 and 4 |
| "used up, or another cause" | Quota exists, so it is probably an ordinary rate limit | Space out retries; the design is covered in [Don't Retry Every Gemini 429 — Telling Rate Limits Apart From Spend Cap Exhaustion](/en/articles/gemini-api/gemini-api-429-retryable-vs-spend-cap-exhaustion-retry-design) |
The first shape is almost anticlimactic. Learning that "that model just wasn't available" and getting unstuck on the spot is quite possible.
Why I Prefer a Probe Over Guessing
It is tempting to open the console and click through every settings page until something looks wrong. I have done it, and it rarely pays off, because most of those pages show the project's state, not the state of the particular model you are calling.
A probe flips that around. It asks the API itself, model by model, and the answer comes back in the same words the API will use later in production. That makes the result easy to paste into a note, compare next week, and hand to someone else.
It also keeps the problem small. One request per model is enough to separate "this model has no quota here" from "something is wrong with the whole project," and those two lead to very different next steps.
If It Still Doesn't Work, Gather These Before Asking
If you cannot narrow it down by now, you will end up asking on a forum or to support. Having these five things ready makes the reply faster and more accurate.
- The model name you called
quotaMetricandquotaValuefrom the error body- When it happened, with the time zone
- The project identifier (never paste the key itself)
- Whether you set up billing, and when
The same symptom is discussed in a forum thread, and my impression is that posts with these details get to a cause much sooner.
Closing
If I had to put today's point in one sentence:
A 429 on the very first call is something to read, not something to wait out.
The next step is small and fits in today. Run the script above once, and write down which models show "zero quota" for your project. Having that list in hand makes the next encounter with the same screen feel very different.
I now put this check first whenever I set up a new environment.