Gemini 3.8 Live and 3.8 Live Extended Thinking
Vintage: 2026-09. Primary evidence is Google's Sep 15, 2026 issuer blogs (Gemini Audio Team launch + DeepMind developer companion), as hydrated in 2026-09-15-google-gemini-3-8-live-3-8-live-extended-thinking. Every capability number and price on this page is Google-reported, not a third-party rerun. The model card linked from the launch post was not PDF-parsed. Native X was blocked this pass (user-X monthly spend-cap 403) — no X permalinks. Blog “Watch …” demo embeds are Google marketing video, not X;
x_video: falseon the source is intentional.
Gemini 3.8 Live and 3.8 Live Extended Thinking
One-line summary: Google's September 15, 2026 speech-to-speech / Live API pair — 3.8 Live (scale / cost-efficient live dialogue with visual grounding) and 3.8 Live Extended Thinking (higher-complexity live dialogue that reasons while speaking). Same “3.8” generation label as gemini-3-8-flash, different modality SKU. Do not collapse into Flash or lyria-3-5.
What it is
Google's Sep 15 2026 model post introduces two Live SKUs. 3.8 Live is positioned for scale and cost efficiency: conversational intelligence, fluid dialogue, visual grounding, near real-time visual inputs, automatic detection/transition across 97 supported languages mid-conversation, and tools/API calls in the background while dialogue continues. 3.8 Live Extended Thinking is positioned for high-complexity tasks: increased intelligence and multi-step reasoning; it “reasons and speaks simultaneously,” with early verbal cues and live progress narration for background tasks.
This page records issuer-hydrated claims from the two fetched Google blogs. It does not re-file gemini-3-8-flash, lyria-3-5, or deepseek-v4-1-flash. gemini-3-5-transcribe is restated as a last-month release, not a Sep 15 launch.
Why it matters to this thread
Frontier model releases, pricing shifts, multimodality (audio / voice), and capability evaluations are in-scope. This is the first dated official Gemini 3.8 Live / Live Extended Thinking pair the thread has as a citable primary. Distinct from gemini-3-8-flash (Sep 2 text/agent Flash + Fairwind-gated Flash Cyber) and from lyria-3-5 (music generation). Distinct from gemini-macos-voice (macOS dictate/summarize/rewrite, no Live SKU).
Key facts (from 2026-09-15-google-gemini-3-8-live-3-8-live-extended-thinking)
Two Live SKUs, one announce day
- From 2026-09-15-google-gemini-3-8-live-3-8-live-extended-thinking (Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, Sep 15, 2026; Tom Ouyang, Malini Jaganathan on behalf of Gemini Audio Team): 3.8 Live — “Built for scale and cost efficiency,” conversational intelligence, fluid dialogue, visual grounding; near real-time visual inputs; automatic detection/transition across 97 supported languages mid-conversation; tools/API calls in the background while dialogue continues.
- From the same source (same blog): 3.8 Live Extended Thinking — “Built for high-complexity tasks,” increased intelligence and multi-step reasoning; “reasons and speaks simultaneously”; early verbal cues and live progress narration for background tasks.
- From the same source (Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe, Sep 15, 2026; Alisa Fortin, Thor Schaeff, Google DeepMind): restates availability in Gemini API / AI Studio; lists asynchronous function calling, visual context, alphanumeric precision, multilingual (97+ languages), incremental content updates; configurable thinking for Extended Thinking. Try path
ai.studio/live.
Issuer-cited benches — Google-reported, AA page not hydrated
All of the following are Google-reported on the Sep 15 launch blog, not reproduced from Artificial Analysis or the named bench hosts in this pass:
- Extended Thinking claimed #1 overall on Artificial Analysis’ Speech to Speech Quality Index (82.6). Treat as Google-cited AA, pending a direct AA page hydrate. Do not fold 82.6 into artificial-analysis-intelligence-index (a different AA instrument).
- Agentic task completion 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking; 97.7% on Big Bench Audio.
- 3.8 Live claimed second place in Speech Agent Arena.
- ServiceNow EVA-Bench cited for Pareto frontier on complex workflows — note: run on Live API on Gemini Enterprise Agent Platform.
Live API pricing (issuer; cite both units)
- From the same source (developer companion): Live API $0.005/min audio input and $0.018/min audio output (footnote: estimate based on $3/1M input tokens and $12/1M output tokens). Cite both as the issuer presents them. Do not invent a conversion beyond that footnote. Do not stamp Flash’s $0.75/$3.75 token intro onto Live.
Rollout “starting today” — surfaces differ by SKU
Do not collapse into “available everywhere.”
- 3.8 Live: developers — Gemini API + Google AI Studio; enterprises — private preview in Gemini Enterprise (Customer Experience coming soon); everyone — Search Live.
- 3.8 Live Extended Thinking: developers — Gemini API + AI Studio; enterprises — private preview in Gemini Enterprise (+ Workspace business customers coming soon); everyone — Gemini Live; Google AI Pro/Ultra in Workspace Docs; all Google AI subscribers in Gmail and Keep.
Safety / partners (issuer marketing)
- From the same source (launch blog): all AI-product audio watermarked with SynthID; points readers to the model card (linked from the post; not PDF-parsed this pass).
- Partner names on the posts: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents; Salesforce, Genspark, Lumeris called out as excited partners. Treat partner quotes as marketing unless separately sourced. Did not mint partner pages.
What this source does not establish
- No fetched model card. The launch post links a model card; it was not PDF-parsed. Do not cite card numbers, limits, or safety appendices.
- No independent AA / τ-Voice / Speech Agent Arena / EVA-Bench hydrate. 82.6 / #1, 68.6%, 35.1%, 97.7%, second place, and EVA-Bench Pareto are Google-reported. Harness, date, and third-party confirmation are open.
- No native X. User-X monthly spend-cap 403; no
https://x.com/.../status/...permalinks. Silence is a fetch miss, not “labs said nothing.” - Blog demo embeds are not X video.
x_video: false. Do not run/transcribe-clippingfor X on this source. - Gemini 3.5 Transcribe is last-month, not a Sep 15 launch. See gemini-3-5-transcribe.
- Lyria 3.5 / Gemini 3.5 Live Translate / Gemini 3.1 Flash TTS are suite pointers on the developer post, not new launches this day. lyria-3-5 already filed. Did not mint Translate or Flash TTS pages.
- Not a re-file of gemini-3-8-flash, lyria-3-5, or deepseek-v4-1-flash.
- Bloomberg “OpenAI working with Anthropic, Google on AI safety” (2026-09-15 wire) is an open pointer in the source — secondary journalism; no OpenAI/Anthropic issuer primary hydrated. Not lab product grain.
- No ticker, 8-K, or stock-market tag.
Contradictions / tensions
- AA 82.6 / #1 is Google-cited, not an AA page print. Confirm on Artificial Analysis’ own Speech-to-Speech leaderboard before treating as independent third-party grain. Different instrument from artificial-analysis-intelligence-index v4.3 53.
- τ-Voice / τ-Voice-banking / Speech Agent Arena / EVA-Bench methodology and score dates are issuer-asserted. Keep tagged as Google-reported until primary bench pages are filed.
- Live API $/min vs token-footnote ($3 / $12 per 1M). Cite both. Do not invent conversion beyond the footnote.
- 3.8 Live vs 3.8 Flash: same generation label, different modality SKU. Surfaces, meters, and benches do not transfer. Not collapsed.
- Consumer surface matrix differs by SKU (Search Live vs Gemini Live vs Docs/Gmail/Keep). Do not flatten into one availability claim.
- SynthID “all audio generated by our AI products” is a policy claim on this post; model-card details not parsed.
Open questions
- Does Artificial Analysis’ own Speech-to-Speech Quality Index show Extended Thinking at 82.6 / #1, or only Google’s cite?
- What do τ-Voice / τ-Voice-banking / Speech Agent Arena / EVA-Bench primary pages add (date, harness, comparables) that the blog omits?
- How does the Live API $/min meter relate to the $3 / $12 per 1M footnote in practice? Do not invent a token-per-minute rate.
- What does the unparsed model card add (limits, safety, eval protocol)?
- Does Paul later merge the 3.8 Live family onto the Flash entity, or keep the modality split?