GEMINI LABJP
LYRIA — Lyria 3.5 entered public preview on September 3. It generates full-length songs rather than loops, with fine-grained control over duration and structureINPUT — The model ID is lyria-3.5, and it accepts images alongside text. Being able to say make it feel like this photo changes how you approach the promptSPEC — Output is 44.1kHz stereo, the input ceiling is 131,072 tokens, and it can write lyrics. Worth carrying those units whenever you quote the numbersVIDEO — The agentic video understanding released September 1 lets the model walk a timeline itself, pulling transcripts or frames on demand. Up to 88% fewer tokens on long contentQUIET — Nothing new has landed in the API release notes since September 3. A quiet stretch is still worth recording rather than passing over in silenceDICTATION — On the app side, holding the Fn key now dictates into your active window, dropping cleaned-up text straight at the cursorLYRIA — Lyria 3.5 entered public preview on September 3. It generates full-length songs rather than loops, with fine-grained control over duration and structureINPUT — The model ID is lyria-3.5, and it accepts images alongside text. Being able to say make it feel like this photo changes how you approach the promptSPEC — Output is 44.1kHz stereo, the input ceiling is 131,072 tokens, and it can write lyrics. Worth carrying those units whenever you quote the numbersVIDEO — The agentic video understanding released September 1 lets the model walk a timeline itself, pulling transcripts or frames on demand. Up to 88% fewer tokens on long contentQUIET — Nothing new has landed in the API release notes since September 3. A quiet stretch is still worth recording rather than passing over in silenceDICTATION — On the app side, holding the Fn key now dictates into your active window, dropping cleaned-up text straight at the cursor
Articles/Gemini Basics
Gemini Basics/2026-09-09Intermediate

Two dictations on one Fn key — where I let Gemini type, and where I don't

On a Mac, the Fn key now hosts two different dictations: hold it and Gemini drops polished text at your cursor, tap it twice and macOS transcribes you as-is. Here is the boundary I drew before letting either one into my day, and the order I check when nothing happens.

Gemini for Macdictationvoice inputmacOSshortcuts

Late one afternoon, with a pile of app review replies waiting, I pressed the Fn key inside a scratch text file and started talking. What came up was the familiar macOS microphone, not Gemini. One small difference in how I pressed the key had sent me somewhere else entirely.

Gemini on macOS listens while you hold Fn, and drops cleaned-up text at your cursor the moment you let go. The built-in dictation answers to a double tap of the same key. Two dictations with different temperaments now live on one key.

Before thinking about convenience, I decided something else first: which text fields I would never press that key in. What gets inserted does not always land somewhere forgiving.

Two dictations, one key

It helps to fix the mapping before anything else. Get this wrong and every later settings check goes to the wrong place.

What you doWhat answersWhat the text looks likeWhere the setting lives
Hold Fn while speakingGemini dictationFillers removed, shaped into paragraphs or listsGemini app settings (Speak to Window)
Tap Fn twiceBuilt-in macOS dictationClose to a raw transcript of what you saidSystem Settings, Keyboard, Dictation

With Gemini, holding the key is the input and releasing it is the commit. It drops the "ums," takes your mid-sentence correction over the thing you first said, shapes the result into paragraphs, and places it wherever your cursor happens to be in the active window.

The built-in one does the opposite on purpose. It gives you back roughly what you said. The presence or absence of that shaping is the real difference between them — not quality, but purpose.

What you have selected before you speak also changes the outcome. With text highlighted, the same gesture stops being an insertion and becomes an instruction applied to that selection. Adding and rewriting sit on one key — worth remembering before your thumb goes down.

One more thing worth knowing before you start troubleshooting: the hold-to-talk dictation is rolling out in English first. If nothing happens in another language, the answer may simply be that it has not reached you yet.

Fields I speak into, and fields I don't

For a while I treated this as a general-purpose input method that works anywhere. That did not go well. Dictation inserts text; it has no opinion about whether the destination is safe.

Three questions decide it for me now. Can I undo this? Is there a send button nearby? Does this need to be exactly right? If a field trips any of the three, I type it with my fingers.

VerdictFieldWhy
SpeakScratch notes, rough article drafts, first passes at store descriptions, draft replies to app reviewsEverything here can be rewritten as many times as I want
TypeAn email body with the recipient already filled in, a chat composerThe send action sits right there, and it does not come back
TypeTerminal, code inside an editor, config filesShaping introduces line breaks and punctuation I did not ask for
TypeFinal copy before publishing, password and verification code fieldsA misheard word costs far too much

Because the same key can rewrite a selection rather than add to it, I have started glancing at what is highlighted before I speak. A half-finished paragraph left selected turns a small addition into a replacement, and the paragraph I meant to extend is simply gone.

The terminal is the one I nearly learned the hard way. Picturing a nicely formatted paragraph landing on a command line is enough to make my stomach turn.

The rule I hold myself to is one line long.

Dictation gets me to a draft. Sending and running stay on my fingers.

Since drawing that line I have stopped hesitating over the feature itself. When the safe places are already decided, there is nothing to weigh each time.

The order I check when nothing happens

When the key does nothing, the cause is usually in one of four places. Working down the list saves a lot of wandering.

  1. Whether the feature has reached you at all. The English-first rollout means waiting is a real possibility.
  2. Whether dictation is enabled in the Gemini app settings, where the shortcut assignment can also be changed.
  3. What System Settings, Keyboard, Dictation currently has as its shortcut, and whether two Fn-based shortcuts are colliding.
  4. Whether a background utility is remapping the key. If something like Karabiner-Elements touches Fn or F5, the press never reaches the app at all.

There is an order behind that list. The first two are about whether the feature exists for you at all, and the last two are about whether your keypress ever arrives. Checking availability after spending twenty minutes in a settings panel is the kind of detour I would rather not repeat.

The fourth one is easy to miss. I had left a key remapper running for years and had entirely forgotten about it. Before deciding an app is broken and pacing between settings screens, I would look for a utility quietly intercepting the key.

Settings split between the app and the OS is a recurring shape in Gemini. An earlier piece of mine, deleting the chat wasn't enough, there was a second place the memory was coming from, was the same kind of stumble in a different room.

If it shapes your words, it also drops some of them

The value of Gemini's dictation is that it turns hesitant speech into something readable. The flip side is that part of what you said is quietly discarded along the way.

Numbers and proper nouns are where this caught me. Say a model name or a version number out loud and the shaping may nudge it toward a tidier spelling. Perfectly readable, and no longer the value I meant to publish.

So I now pause and reread immediately after dictating. I check three things in particular: the digits in any number, the spelling of proper nouns, and whether my negations survived. Negations tend to appear exactly where I corrected myself mid-sentence, and if the shaping absorbs one, the meaning flips.

There is a smaller habit underneath that one. When a sentence matters, I say it once and stop, rather than talking through three versions and trusting the shaping to pick the right one. The more corrections I stack inside a single hold, the less predictable the result becomes.

Numbers and names are faster to type by hand from the start — it took me a while to admit that. What dictation is genuinely good at is getting a half-formed thought out of my head and onto the page.

Try one day, in a scratch file only

Mixing dictation into everything at once makes it impossible to tell where it helped and where it got in the way. I kept it to scratch notes for a day, then widened from there.

If there is one thing to do today, open a throwaway text file, hold Fn, and say what you are thinking. Read back what appears and see how much of your own phrasing survived — that is the honest measure of how far you can trust it.

Thank you for reading this far. I am still working out how much of my drafting I want to hand over.

Share

Thank You for Reading

Gemini Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • Copy-paste ready implementation code
  • New advanced guides published daily
  • $5/mo or $15 for lifetime access
View Membership →

If you found this article helpful, a small tip ($1.50) would mean a lot to us. Your support helps keep this site ad-free and covers server and hosting costs.

Related Articles

Gemini Basics2026-09-04
Deleting the Chat Wasn't Enough. There Was a Second Place the Memory Was Coming From
How Gemini's memory of past chats actually works, why deleting a conversation often isn't enough, and how the fact that Gems and Gemini Live ignore that memory can be turned into a simple way to keep separate projects from bleeding into each other.
Gemini Basics2026-08-28
What decides whether your Gemini API data trains Google's models, and the one exception on the paid tier
Whether your Gemini API prompts feed model training is not settled by the paid tier alone. Here is how AI Studio and the API define paid differently, how dataset sharing reverses the protection, and what changes by region.
Gemini Basics2026-08-16
The Strongest Gemini Is in Preview and the Cheap One Is GA. That Ordering Should Drive Your Model Choice
Sort the Gemini lineup by release stage instead of capability and the order inverts: the strongest reasoning model sits in preview while the fast, inexpensive Flash models are GA. Here is how a solo developer handles that, plus a 30-line script that counts how many preview models your code already depends on.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links