brain/
← all entities
entitygenericartificial-intelligence

Unsloth

Notes

Vintage: 2026-09. Primary evidence is verified @UnslothAI plus fetched issuer docs unsloth.ai/docs/models/glm-5.3-flash, as hydrated in 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash. Specs here are Unsloth docs, not a Z.ai model card.

Unsloth

One-line summary: Local-inference tooling issuer. On 2026-09-04 it published a GLM-5.3-Flash / Ox Alpha GGUF speedup — banner “3.3x faster inference,” docs 1.6–3.4× with optimized decoding and MTP.

What it is

The issuer of the Sep 4 local-inference optimizations for glm-5-3-flash. From 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash (verified @UnslothAI, 2026-09-04 12:32 UTC): GLM-5.3-Flash now runs “3.3x faster locally,” with “Local GGUF inference… 1.6–3.4× faster with optimized decoding and bonus multi-token prediction,” pointing at Unsloth docs and Hugging Face GGUFs (huggingface.co/unsloth/GLM-5.3-Flash-GGUF).

Fetched docs match: banner “Sep 4: GLM-5.3-Flash now runs with 3.3x faster inference”; model described as also known as ox-alpha; 320B total / 18B active multimodal open model; 1M context; hardware table (1-bit ~100 GB through BF16 ~650 GB); tok/s tables (e.g. tg32 @ 65536: 20.66 → 48.99 tok/s before MTP n=2 examples).

Why it matters to this thread

Local GGUF speedups on an already-filed Chinese open-weight Flash SKU are in-scope (chinese-open-weight-frontier-parity, edge-inference-shift). This is Unsloth-claimed decode/MTP, not a new Z.ai release and not a rename of overnight omen-alpha.

Key facts (from 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash)

  • Sep 4 banner: 3.3× faster inference (docs).
  • X text: 1.6–3.4× local GGUF with optimized decoding + MTP.
  • Unsloth-described shape: ox-alpha alias; 320B / 18B active; 1M context; 1-bit ~100 GB through BF16 ~650 GB.
  • Example tok/s in docs: tg32 @ 65536: 20.66 → 48.99 before MTP n=2 examples.
  • GGUF hub: unsloth/GLM-5.3-Flash-GGUF.

What this source does not establish

  • Not a Z.ai model card. 320B/18B and the hardware/tok/s tables are Unsloth docs.
  • Not an Omen identity. Unsloth’s page is about GLM-5.3-Flash / Ox Alpha. Do not attach those specs to omen-alpha.
  • @BAI_AGI “#1 on B.AI” / “2.41 trillion” cumulative tokens is platform marketing (video not transcribed) — not this issuer’s card. Filed as a flag on glm-5-3-flash.

Open questions

  • Do Z.ai weights or an official card independently state 320B / 18B / 1M, or does that remain Unsloth-described?

Sources

Related

Referenced by