brain/
← all entities
entitygenericartificial-intelligence

Gemini 3.7 Flash

Notes

Gemini 3.7 Flash

Vintage: 2026-08. Primary evidence is official Google X posts dated 2026-08-13/14 plus independent and third-party X posts dated 2026-08-17/19. Capability claims are a snapshot of that week, not a model card.

One-line summary: Google DeepMind's August 2026 Flash-tier model, launched on X as a coding / knowledge-work / agents workhorse; independent SkillsBench print is #2 by 0.1pt; Google comms (@NewsFromGoogle, not DeepMind) claims #1 on AA-AnalystAgent.

What it is

A Gemini Flash-generation model announced on X by @GoogleDeepMind on 2026-08-13. Official posts position it as stronger than the prior Flash for coding, knowledge work, and web development, and as the default workhorse inside Gemini Spark and (a day later) Gemini chat for Pro/Ultra. This page records X-post claims only. The launch thread mentions a blog URL (https://goo.gle/3RNm2tZ) that was not fetched for this ingest.

Why it matters to this thread

Frontier model releases and capability evaluations are in-scope. This is the first dated official Gemini 3.7 Flash launch the thread has as a citable source. It also supplies an independent SkillsBench print and two non-official X claims (a Google VP product note; a researcher robotics bench) that must stay attributed to those accounts.

Key facts (from 2026-08-18-x-ai-news-18-aug-2026-gemini-3-7-flash-plus-lab-concentration)

All bullets are X posts (fetch_method: x-mcp). Permalinks live on the source page.

  • Official launch (X post, @GoogleDeepMind, 2026-08-13): "Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development."
  • Official launch details (X post, same thread, 2026-08-13): the follow-up post claims strong gains over 3.6 Flash in debugging and issue resolution; better functional web layouts/apps with fewer prompts; improved reasoning/accuracy on real-world business workflows. Surfaces named: try in @Antigravity; API in @GoogleAIStudio and @AndroidStudio; Google AI Pro/Ultra can use 3.7 Flash in Gemini Spark in @GeminiApp.
  • Gemini Spark (X post, @GeminiApp, 2026-08-13): Spark (Google AI Pro and Ultra, 160+ countries) uses 3.7 Flash starting that day. The post calls it the "most intelligent workhorse model for coding and agents."
  • Gemini chat (X post, @GeminiApp, 2026-08-14): available to all Pro and Ultra users in Gemini chat; claims improved reasoning/accuracy for multi-step tasks across files and emails.
  • Independent SkillsBench (X post, @ValsAI, 2026-08-18): "it ranks #2 on SkillsBench by leaning on skill files. It scores 65.9, 0.1pts behind Grok 4.5 (66.0), a gap inside the error bars." Not a Google announcement. See ai-coding-benchmarks.
  • Google product VP note tweet (X post, @joshwoodward, 2026-08-18): Josh Woodward, a Google product VP, circling back on X — not a model card. Claims: revamped Workspace tools in 1–2 weeks; 3.7 Flash showed improvements in tool calling, more coming; new Projects design done, implementing; 49 connectors (and counting) supported; several items marked done, including biggest over-triggering bugs.

Key facts (from 2026-08-20-x-ai-news-20-aug-2026-openai-private-safety-claude-connectors)

  • Google comms, AA-AnalystAgent (X post, @NewsFromGoogle, 2026-08-19; has_media: photo): "Gemini 3.7 Flash isn't just fast ⚡ It's now ranked #1 on @ArtificialAnlys' AA-AnalystAgent leaderboard for complex, real-world data analysis." Same post: "In a test of 80 real-world tasks across 14 business and scientific domains, Gemini 3.7 Flash combined reasoning and speed to deliver the highest overall accuracy, completing tasks 60% to 90% faster than other top-performing models and more than twice as fast (2.4x) as its closest accuracy rival."
  • Qualify: Google communications account, not @GoogleDeepMind and not a model card. Do not add a 60% / 53.8% / 50% per-model accuracy table — those numbers are not in this post (they were in an excluded Wes Roth post). The "60% to 90% faster" and "2.4x" figures are in the @NewsFromGoogle text. This is a data-analysis bench, not a coding bench — do not fold it into ai-coding-benchmarks.

What this source does not establish

  • No official pricing. Unofficial $0.75 / $3.75 tweets were excluded from the clipping and are not cited.
  • No fetched blog / model card. The goo.gle/3RNm2tZ URL is mentioned in the DeepMind X post only.
  • Robotics 92% / 32% is not Google's number. A researcher X post by @chooi_jeq (Jay Chooi, 2026-08-17) claims 3.7 Flash saturated "one of our physical tool-use benchmarks at 92%" versus 3.6 Flash at 32% three weeks earlier. Filed on ai-robotics-embodiment-wave as a researcher claim.
  • @karpathy had no original posts in the fetch window.
  • No unofficial AnalystAgent score table. Do not write Gemini 3.7 Flash 60% / Opus 5 53.8% / GPT-5.5 50%. Those scores are not in the @NewsFromGoogle post.

Adjacent product-account color (not a Flash claim)

  • From 2026-08-26-x-ai-news-26-aug-2026-openai-jalapeno-business-premium (@GeminiApp, 2026-08-25 — official Gemini product account, not @GoogleDeepMind): "Dictate, summarize files, and rewrite copy directly into any window with your voice in the Gemini app for macOS." No SKU, pricing, or model name in the post. Not a Gemini 3.7 Flash claim. Full treatment: gemini-macos-voice. @GoogleDeepMind had no original posts in that fetch window.
  • From 2026-08-27-significant-ai-developments-last-30-days (@GoogleDeepMind, 2026-08-26 — official DeepMind account): Gemini 3.5 Transcribe pointer (post text not fetched in that pass). Not a Gemini 3.7 Flash claim.
  • From 2026-08-27-x-ai-news-27-aug-2026-deepmind-double-blind-evals (official @GoogleDeepMind, 2026-08-26 — primary X; text/photos): Transcribe capability list (phone numbers / postal codes / order IDs in noise; filler-word removal; custom vocabulary; 85+ languages; Gemini app macOS + Gboard Android). Not a Gemini 3.7 Flash claim. Full treatment: gemini-3-5-transcribe. Video announce held for /transcribe-clipping. Do not invent SKU or pricing.
  • From 2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026 (Google blog + official X, Sep 2, 2026): Gemini 3.8 Flash / Flash Cyber launched as a later pair. Not a Gemini 3.7 Flash claim and not a rewrite of this page. Full treatment: gemini-3-8-flash. The 2026-09-02 overnight clip's unverified 3.8 cluster item is resolved there as a primary.

Strengths

  • Official launch + product-account posts give a dated availability path (Antigravity, API surfaces, Spark, chat).
  • Independent SkillsBench print is explicit that the 0.1pt gap to Grok 4.5 is inside the error bars.

Weaknesses

  • Entire official record here is X posts, not a model card.
  • SkillsBench methodology is unnamed beyond "leaning on skill files."
  • VP follow-up is a personal note tweet.
  • Robotics figures are a single researcher X post.

Open questions

  • What does the official blog / model card add (pricing, eval tables, safety) that the X thread omits?
  • Is SkillsBench 65.9 vs 66.0 a durable ranking or noise, as @ValsAI itself flags?
  • Does the researcher physical tool-use bench replicate, and on whose robot stack?

Sources

Related

Referenced by