brain/
← all entities
entitygenericartificial-intelligence

Gemini 3.8 Flash and 3.8 Flash Cyber

Notes

Vintage: 2026-09. Primary evidence is Google's Sep 2, 2026 model blog plus same-day official @GoogleDeepMind / @GeminiApp X, as hydrated in 2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026. Every capability number on this page is Google-reported, not a third-party rerun. The DeepMind model card URL was not fetched (WebFetch timeout). The root launch X post has an animated GIF; claims here are from post text and fetched official pages only — media was not described.

Gemini 3.8 Flash and 3.8 Flash Cyber

One-line summary: Google's September 2, 2026 Gemini 3.8 pair — two variants sharing a foundation, different mitigations, not the same deployed model. Flash is generally available (developer / enterprise / Pro/Ultra); Flash Cyber is gated via fairwind.

What it is

Google's Sep 2 2026 model post says Gemini 3.8 has two variants that share a foundation. Flash is the generally available line: Google says it improves software engineering, agents, and multi-step reasoning, and that it is available via developer, enterprise, and Gemini Pro/Ultra consumer surfaces. Flash Cyber is the cyber-specialized line, gated via fairwind. The same post says Flash and Cyber have different mitigations. Do not treat them as the same deployed model.

This page records issuer-hydrated claims from the fetched Google blog, the fetched Fairwind program page, and official X text. It resolves the 2026-09-02 overnight clip's unverified Gemini 3.8 Flash news-cluster item (research-clip://x-ai-overnight-lab-drops-2026-09-02 / 2026-09-02-x-ai-overnight-fable-5-1-astra-critical-reward-seeker): that pass treated timing as chatter and did not cite it. Official blog + DeepMind/GeminiApp X now confirm a Sep 2 launch. Do not keep "unverified" on the launch itself. Do not re-file Fable/Mythos 5.1, reward-seeker / Hacker-Opus, Astra Critical, or DeepMind agentic video from that overnight source.

Why it matters to this thread

Frontier model releases, pricing shifts, and capability evaluations are in-scope. This is the first dated official Gemini 3.8 Flash / Flash Cyber pair the thread has as a citable primary. Distinct from gemini-3-7-flash (August 2026 Flash-tier X launch) and from gemini-agentic-video-understanding (Sep 1 agentic video; not re-filed here). Distinct from claude-fable-5-1's same-model / different-safeguards pair and from openai-astra-critical-cyber — do not rank Google-reported Cyber benches against Astra / ExploitBench as if they share a harness.

Key facts (from 2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026)

Two variants, one foundation

  • From 2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026 (Google blog, Sep 2, 2026): Gemini 3.8 has two variants that share a foundation. Flash is generally available (developer, enterprise, Pro/Ultra). Flash Cyber is gated via Fairwind. Flash and Cyber have different mitigations.
  • From the same source (@GoogleDeepMind X, 2026-09-02, 2095175498967949359): launch post repeats the split in text — 3.8 Flash improvements across SWE, agentic tasks, and multi-step reasoning; Cyber described as Google's most capable cyber variant. That post has an animated GIF; claims here are from post text and fetched official pages only. x_video: false on the source is intentional.

Flash: availability, Google-claimed capability, introductory price

  • From the same source (Google blog): 54.9% HLE-Verified for Flash. That is a Google-reported score, not a third-party rerun in this pass. The same post says Flash may use more tokens.
  • From the same source (@GeminiApp X, 2026-09-02, 2095176192307691835): Flash is available to Pro/Ultra and claims more reliable comprehensive responses for advice, text analysis, and complex coding.
  • From the same source (@GoogleDeepMind X, 2095175502298243094): Flash is rolling out via Antigravity / API, with Cyber gated via Fairwind.
  • From the same source (Google blog, introductory API pricing): $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026, then $1.50 / $7.50. The cheap intro rate expires; do not stamp $0.75/$3.75 as the durable list price. Token-heavier generations (Google's own "may use more tokens" note) can eat the intro discount in practice. Consumer Pro/Ultra availability is a surface, not evidence that the same meter applies there.

Flash Cyber: Google-reported benches, not an independent bake-off

All of the following are Google-reported on the Sep 2 blog, not reproduced in the clipping:

  • >70% success on an internal 20-language vulnerability benchmark.
  • CWE-Bench pass@1 47.2% versus a leading frontier model at 47.8%. Google does not name that frontier model in the fetched post. The gap is 0.6 points; treat it as a near-tie on Google's own comparison, not as a demonstrated lead.
  • Chrome tests: 2.6× more correct patches than larger commercial models (unnamed).

Official X in the same window restates overlapping but not identical claims. Keep them as two Google claims, not one merged result:

  • From the same source (2095196704769237137): expert vulnerability detection and autonomous patching.
  • From the same source (2095196708254658820): deployable fixes in minutes inside an org cloud. "Fixes in minutes" is an X claim, not a numbered result on the fetched blog.
  • From the same source (2095196714479071566): CyberGym lead and 2.6× more valid fixes in Chrome tests.

Cross-model harness, effort level, tool access, and patch-validity criteria are not independently checked. Do not compare these numbers to other labs' cyber benches (including OpenAI's Astra / ExploitBench material on openai-astra-critical-cyber) as if they share a harness.

Fairwind gating

Flash Cyber is gated via fairwind. Full program rules live on that page. From the same source: vetted governments, healthcare, telecom, and core platforms get access to Flash Cyber and CodeMender; Google says over 650 partners; use restricted to defensive / academic dual-use tasks. Those are program rules and Google statements, not an audit.

What this source does not establish

Contradictions / tensions

  • Yesterday vs today: the 2026-09-02 overnight clip left Gemini 3.8 Flash timing as unverified cluster chatter. Official blog + DeepMind/GeminiApp X now confirm a Sep 2 launch. Cluster item resolved as a primary; do not keep "unverified" on the launch itself.
  • CWE-Bench near-tie vs "most capable cyber variant": Google reports 47.2% pass@1 vs 47.8% for an unnamed leading frontier model on the blog, while the launch X post calls Cyber the most capable cyber variant. Those two Google lines do not automatically agree. Not reconciled.
  • Harness comparability: CWE-Bench, the internal 20-language set, Chrome tests, CyberGym (X only), and other labs' cyber evals are not shown to share protocol, tools, or patch-accept criteria. Stamp Google-reported, not cross-lab rank.
  • 2.6× wording split: blog = 2.6× more correct patches than larger commercial models in Chrome tests; X = CyberGym lead and 2.6× more valid fixes in Chrome tests. Same-day Google, not proven identical metrics.
  • Introductory pricing expiry: $0.75/$3.75 holds only through 31 December 2026, then $1.50/$7.50 per the blog. Pair with "may use more tokens."
  • Fairwind 650 partners / zero retention / dual-use-only: Google program claims; no partner list or independent retention audit in this pass. See fairwind.

Open questions

  • What does the unfetched model card add (limits, safety appendices, eval protocol) that the blog omits?
  • Does the unnamed CWE-Bench "leading frontier model" at 47.8% stay unnamed, and does a named third-party rerun move the 0.6-point gap?
  • Are blog "correct patches" and X "valid fixes" / CyberGym the same Chrome-test metric?
  • Does Pro/Ultra consumer availability share the API intro meter, or is that a different surface?

Sources

Related

Referenced by