Unsloth
Vintage: 2026-09. Primary evidence is verified @UnslothAI plus fetched issuer docs unsloth.ai/docs/models/glm-5.3-flash, as hydrated in 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash. Specs here are Unsloth docs, not a Z.ai model card.
Unsloth
One-line summary: Local-inference tooling issuer. On 2026-09-04 it published a GLM-5.3-Flash / Ox Alpha GGUF speedup — banner “3.3x faster inference,” docs 1.6–3.4× with optimized decoding and MTP.
What it is
The issuer of the Sep 4 local-inference optimizations for glm-5-3-flash. From 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash (verified @UnslothAI, 2026-09-04 12:32 UTC): GLM-5.3-Flash now runs “3.3x faster locally,” with “Local GGUF inference… 1.6–3.4× faster with optimized decoding and bonus multi-token prediction,” pointing at Unsloth docs and Hugging Face GGUFs (huggingface.co/unsloth/GLM-5.3-Flash-GGUF).
Fetched docs match: banner “Sep 4: GLM-5.3-Flash now runs with 3.3x faster inference”; model described as also known as ox-alpha; 320B total / 18B active multimodal open model; 1M context; hardware table (1-bit ~100 GB through BF16 ~650 GB); tok/s tables (e.g. tg32 @ 65536: 20.66 → 48.99 tok/s before MTP n=2 examples).
Why it matters to this thread
Local GGUF speedups on an already-filed Chinese open-weight Flash SKU are in-scope (chinese-open-weight-frontier-parity, edge-inference-shift). This is Unsloth-claimed decode/MTP, not a new Z.ai release and not a rename of overnight omen-alpha.
Key facts (from 2026-09-04-x-overnight-opencode-omen-alpha-unsloth-glm-5-3-flash)
- Sep 4 banner: 3.3× faster inference (docs).
- X text: 1.6–3.4× local GGUF with optimized decoding + MTP.
- Unsloth-described shape:
ox-alphaalias; 320B / 18B active; 1M context; 1-bit ~100 GB through BF16 ~650 GB. - Example tok/s in docs: tg32 @ 65536: 20.66 → 48.99 before MTP n=2 examples.
- GGUF hub: unsloth/GLM-5.3-Flash-GGUF.
What this source does not establish
- Not a Z.ai model card. 320B/18B and the hardware/tok/s tables are Unsloth docs.
- Not an Omen identity. Unsloth’s page is about GLM-5.3-Flash / Ox Alpha. Do not attach those specs to omen-alpha.
- @BAI_AGI “#1 on B.AI” / “2.41 trillion” cumulative tokens is platform marketing (video not transcribed) — not this issuer’s card. Filed as a flag on glm-5-3-flash.
Open questions
- Do Z.ai weights or an official card independently state 320B / 18B / 1M, or does that remain Unsloth-described?