◉GEMINI LABJP
●CLI 0.63.0 — Gemini CLI 0.63.0 (Oct 6) shows retry progress on reconnect, tells a missing MCP config from broken JSON, and caps tool output in long agent loops●10/29 — 22 days until the image model gemini-3.1-flash-image shuts down. Move to gemini-nano-banana-2.1, added on Oct 6●10/22 — 15 days until the three Veo 3.1 models shut down. The place to move is gemini-omni-1.1-flash●NOTICE — A Zenn post reports that a plugin's admin screen said nothing when an AI model was retired. How to avoid missing such notices is the real question●NEW — Three Veo 3.1 previews stop on 10/22. Before swapping them, here is how to inventory every place that calls them●CLAUDE — Zenn keeps getting posts on handing Codex or Claude Code writing to Gemini. The real question is which jobs to hand over and which to keep●CLI 0.63.0 — Gemini CLI 0.63.0 (Oct 6) shows retry progress on reconnect, tells a missing MCP config from broken JSON, and caps tool output in long agent loops●10/29 — 22 days until the image model gemini-3.1-flash-image shuts down. Move to gemini-nano-banana-2.1, added on Oct 6●10/22 — 15 days until the three Veo 3.1 models shut down. The place to move is gemini-omni-1.1-flash●NOTICE — A Zenn post reports that a plugin's admin screen said nothing when an AI model was retired. How to avoid missing such notices is the real question●NEW — Three Veo 3.1 previews stop on 10/22. Before swapping them, here is how to inventory every place that calls them●CLAUDE — Zenn keeps getting posts on handing Codex or Claude Code writing to Gemini. The real question is which jobs to hand over and which to keep
Articles/API / SDK
◈ API / SDK/2026-07-04Advanced

Catching the Rows That Quietly Failed Overnight: A Per-Row Retry Ledger for the Gemini Batch API

A SUCCEEDED batch job is not the same as all-rows-succeeded. From running nightly batches as a solo developer, here is a per-row result ledger, a transient-vs-permanent failure classifier, selective retries, and a guard against retrying permanent failures forever, with a working SQLite state machine.

gemini-api284batch-api3retry-designidempotency5indie-dev47operations19

✦ Premium Article

In the morning, the batch job status read SUCCEEDED. I wrote the results back with a clear conscience and moved on to other work for the day.

A few days later, while scanning the classifications, I noticed something. One cluster of reviews had landed in Firestore with an empty category. A few dozen rows. The job as a whole had "succeeded," yet some of the rows inside it had quietly failed.

As an indie developer running several apps and four sites, the nightly batch becomes a "start it and go to sleep" tool. In my own case, classifying the reviews that pile up in App Store Connect and Google Play Console, I built the next stage of processing on the assumption that everything would be present by morning. That is exactly why a partial failure like this leaks downstream and contaminates later steps.

This article is about dropping the habit of reading a batch completion as "all rows succeeded," and instead recording results per row in a ledger and picking up only the rows that fell through. It focuses not on the first implementation of nightly processing, but on the question that always follows it: how do you recover the few dozen rows that failed?

"Completed" and "all succeeded" are different

The state of a batch job and the success or failure of each individual request inside it live on different layers. The job can finish cleanly while the output JSONL mixes successful responses and errors line by line.

A single output line usually takes one of the two shapes below. Key names shift a little between SDK versions, so I recommend peeking at your actual output with head before you build against it.

{"key": "review-000512", "response": {"candidates": [ ... ]}}
{"key": "review-000513", "error": {"code": 400, "message": "..."}}

In other words, even when the job is SUCCEEDED, rows carrying an error are perfectly normal. What I dropped were an aggressive review body caught by the safety filter, and a body made up entirely of emoji that failed schema extraction.

Here is one idea worth holding onto: record the job's outcome and each row's outcome separately. Mix the two layers into "the job succeeded, so insert everything," and a hole will always open up.

Keep a per-row ledger

So we prepare a ledger that holds a state for every request we submit. The key is the custom_id (here, key). I narrowed the states down to four.

StateMeaningWhat to do next
pendingNo result received yetInclude in the next batch
succeededA valid structured output was obtainedNothing (finalized)
retryableTransient failure (429 / 503, etc.)Resubmit up to the attempt cap
permanentInput- or safety-driven; a retry will not fix itDo not resubmit; route to a human

I chose SQLite because, for a solo developer's nightly batch, "one portable file that survives a mid-run crash" matters more than anything. The ledger schema and initialization are just this.

import sqlite3
import time
 
def open_ledger(path: str = "batch_ledger.db") -> sqlite3.Connection:
    conn = sqlite3.connect(path)
    conn.execute(
        """
        CREATE TABLE IF NOT EXISTS rows (
            key           TEXT PRIMARY KEY,
            payload       TEXT NOT NULL,   -- input request (JSON string)
            status        TEXT NOT NULL DEFAULT 'pending',
            attempts      INTEGER NOT NULL DEFAULT 0,
            last_error    TEXT,
            result        TEXT,            -- extracted result on success (JSON string)
            updated_at    REAL NOT NULL DEFAULT 0
        )
        """
    )
    conn.commit()
    return conn
 
 
def enroll(conn: sqlite3.Connection, key: str, payload: str) -> None:
    """Register as pending on first submission. Do not touch on resubmit."""
    conn.execute(
        "INSERT OR IGNORE INTO rows(key, payload, updated_at) VALUES (?, ?, ?)",
        (key, payload, time.time()),
    )
    conn.commit()

The INSERT OR IGNORE is the crucial part. Resubmitting the same key on night two must not roll a succeeded row back to pending. The ledger exists to protect a success once it is finalized.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦Build a SQLite ledger with four states (pending / succeeded / retryable / permanent) that reconciles the Batch output JSONL by custom_id, in copy-paste-ready code
✦Learn to separate transient failures (429, 503) from permanent ones (safety blocks, invalid input) and set a stop condition so permanent rows are never retried forever
✦See how to resubmit only the failed rows on night two and three, and how the ledger absorbs the cost-accounting drift that spans retry attempts
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Gemini Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

◈ API / SDK2026-08-23
Moving app AI work from runtime calls to a pre-ship batch pass
Where you put a Gemini call decides whether your request count scales with users or with assets. Here is the decision rule I used to move classification into a pre-ship batch pass, plus a resumable implementation.
◈ API / SDK2026-08-22
Once You Pass Twenty Mediation Groups, How Do You Find the Setting That Went Missing?
As ad mediation groups multiply, missing sources and type drift accumulate quietly. Here is the split I settled on: normalize the settings into one matrix, let code confirm the gaps, and send Gemini only the cells that need judgment.
◈ API / SDK2026-06-23
Integrating Gemini 3.2 Pro Function Calling into iOS/Android Apps: Production Design Patterns
A practical guide to integrating Gemini 3.2 Pro Function Calling into iOS and Android apps. Includes working SwiftUI, Kotlin, and Python code, plus production patterns proven in a real indie wallpaper app — cost, latency, staged rollout, and regression testing.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links