brain/
sourceartificial-intelligence

Google: Gemini 3.8 Live + 3.8 Live Extended Thinking (Sep 15)

Issuer Sep 15: Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking (speech-to-speech / Live API), AA Speech-to-Speech #1 claim, Live API pricing, Workspace/Search rollout; companion developer audio post.

Source

Google: Gemini 3.8 Live + 3.8 Live Extended Thinking (Sep 15)

Generated by Grok Bot research on 2026-09-15. WebSearch + WebFetch ladder. Native X blocked this pass (user-X monthly spend-cap 403). Treat as raw material — review before promoting into a project or thread.

Dedup: Vault already has Gemini 3.8 Flash / Flash-Cyber (2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026.md) and Lyria 3.5 (2026-09-04-lyria-3-5-lands-in-gemini-app-and-api-global.md). This pass is new Live / speech-to-speech SKUs dated 2026-09-15 — do not collapse into the Flash entity. DeepSeek-V4.1-Flash (Sep 10) is already filed; skip re-file.

Summary

On 2026-09-15, Google’s Gemini Audio team published issuer posts introducing Gemini 3.8 Live (scale / cost-efficient live dialogue with visual grounding) and Gemini 3.8 Live Extended Thinking (higher-complexity live dialogue with multi-step reasoning while speaking). Both roll out the same day via the Gemini API / AI Studio, with consumer surfaces (Search Live, Gemini Live, Workspace Docs/Gmail/Keep Live per SKU) and Gemini Enterprise private preview. A companion developer post prices the Live API at $0.005/min audio input and $0.018/min audio output and points at integration partners. Issuer claims #1 on Artificial Analysis’ Speech-to-Speech Quality Index (82.6) for Extended Thinking — treat as Google-cited AA, pending direct AA page hydrate. Gemini 3.5 Transcribe is restated as a last-month release (WER 4.0% streaming / 2.6% non-streaming), not a Sep 15 launch. Native X timelines were unavailable this pass (spend-cap).

Findings

Issuer product launch (Google blog, Sep 15)

  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (WebFetch, dated Sep 15, 2026; authors Tom Ouyang, Malini Jaganathan on behalf of Gemini Audio Team):
    • 3.8 Live: “Built for scale and cost efficiency,” conversational intelligence, fluid dialogue, visual grounding; near real-time visual inputs; automatic detection/transition across 97 supported languages mid-conversation; tools/API calls in the background while dialogue continues.
    • 3.8 Live Extended Thinking: “Built for high-complexity tasks,” increased intelligence and multi-step reasoning; “reasons and speaks simultaneously”; early verbal cues and live progress narration for background tasks.
    • Benchmarks (issuer-cited): Extended Thinking claimed #1 overall on Artificial Analysis’ Speech to Speech Quality Index (82.6); agentic task completion 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking; 97.7% on Big Bench Audio. 3.8 Live claimed second place in Speech Agent Arena. ServiceNow EVA-Bench cited for Pareto frontier on complex workflows (note: run on Live API on Gemini Enterprise Agent Platform).
    • Safety: all AI-product audio watermarked with SynthID; points readers to the model card (linked from the post; not PDF-parsed this pass).
    • Rollout “starting today”:
      • 3.8 Live: developers — Gemini API + Google AI Studio; enterprises — private preview in Gemini Enterprise (Customer Experience coming soon); everyone — Search Live.
      • 3.8 Live Extended Thinking: developers — Gemini API + AI Studio; enterprises — private preview in Gemini Enterprise (+ Workspace business customers coming soon); everyone — Gemini Live; Google AI Pro/Ultra in Workspace Docs; all Google AI subscribers in Gmail and Keep.
    • Partner names on the post: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents; Salesforce, Genspark, Lumeris called out as excited partners. Treat partner quotes as marketing unless separately sourced.

Developer / Live API companion (same day)

  • Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe (WebFetch, dated Sep 15, 2026; Alisa Fortin, Thor Schaeff, Google DeepMind):
    • Restates Live + Extended Thinking availability in Gemini API / AI Studio; lists capabilities: asynchronous function calling, visual context, alphanumeric precision, multilingual (97+ languages), incremental content updates; configurable thinking for Extended Thinking.
    • Pricing (issuer): Live API $0.005/min audio input and $0.018/min audio output (footnote: estimate based on $3/1M input tokens and $12/1M output tokens).
    • Gemini 3.5 Transcribe: dedicated STT; released last month; average WER 4.0% streaming and 2.6% non-streaming; 85+ languages; custom vocabulary biasing (up to 1,000 terms); smart transcription mode; Interactions API for files up to 1 hour with timestamps/speaker labels. Do not date Transcribe as a Sep 15 launch.
    • Broader audio suite pointers (not new launches this day unless already filed): Gemini 3.5 Live Translate (70+ languages), Gemini 3.1 Flash TTS, Lyria 3.5 — Lyria 3.5 already in vault from Sep 4 source.
    • Same integration partners list; try path ai.studio/live.

Skipped / not grain this pass

  • DeepSeek-V4.1-Flash secondary recaps (already filed Sep 10 AI-thread source).
  • Bloomberg “OpenAI working with Anthropic, Google on AI safety” (2026-09-15 wire) — secondary journalism; no OpenAI/Anthropic issuer primary hydrated here; leave as open pointer, do not promote as lab product grain.
  • Native X (search_news, lab timelines): blocked — 403 spend-cap-reached on user-X despite prior credit balance notes; same recurring blocker as paused x-ai-news-ingest. No https://x.com/.../status/... permalinks this pass.
  • Blog “Watch …” demo embeds are Google marketing video; not X posts — x_video: false. Do not run /transcribe-clipping for X; optional later transcription of issuer demos is out of scope for this clip.

Contradictions and open questions

  • AA Speech-to-Speech 82.6 / #1 is Google-cited; confirm on Artificial Analysis’ own leaderboard before treating as independent third-party grain.
  • τ-Voice / τ-Voice-banking / Speech Agent Arena / EVA-Bench methodology and score dates are issuer-asserted here — keep tagged as Google-reported until primary bench pages are filed.
  • Live API $/min vs token-footnote ($3 / $12 per 1M) — cite both as issuer presents them; do not invent conversion beyond the footnote.
  • Relationship of 3.8 Live family to existing vault Gemini 3.8 Flash entity: same “3.8” generation label, different modality SKU — wiki should keep separate entities/concepts unless Paul merges them later.
  • Consumer surface matrix (Search Live vs Gemini Live vs Docs/Gmail/Keep) differs by SKU — do not collapse into “available everywhere.”
  • SynthID “all audio generated by our AI products” is a policy claim on this post; model-card details not parsed.

Provenance

Method: Grok Bot / WebSearch / WebFetch (issuer Google blog posts). Native X attempted (search_news, get_users_by_usernames, get_users_by_username) — all 403 monthly spend-cap; failed calls free; no X permalinks. Generated: 2026-09-15 Rounds: 1 — early-exit after two same-day issuer posts + vault dedup vs Flash/Lyria/DeepSeek. X spend note: recurring user-X spend-cap; 10am X ingest already paused; this 3pm idle pass substituted issuer web when X blocked.

Web sources:

X sources:

  • none found — user-X monthly spend-cap 403 this pass

Grokipedia:

  • not used
Referenced by