Mostik.ai
Vintage: 2026-09. Primary evidence is the founder X post (2026-09-02) plus a same-day company confirmation, as hydrated in 2026-09-03-x-ai-pass-3-sep-2026-open-model-fleets-agentic-coding-price (
fetch_method: x-mcp). A same-evening recap states different performance and compute figures. Do not smooth the discrepancy. Do not carry either headline number as a wiki fact. Themostik.ai/read-morewriteup and the named WIRED piece are pointers, not fetched pages. ARC-AGI first-place is an unresolved claim, not a result.
Mostik.ai
One-line summary: Four-month-old startup (General Catalyst–backed) claiming a latent-space "bridge" that passes hidden states from a 753B frontier model into a 4B edge model with no text in between and neither model fine-tuned.
What it is
Sasha Malysheva announced Mostik.ai on 2 Sep 2026; the company confirmed the account. The mechanism claim, as written: models communicate in latent space; "hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched."
The stated configuration: "a 753B model reads the problem, and a 4B edge-class model writes the answer." Team/backing as stated: 15 people, 12 PhDs and a Fields medalist, four months, backed by General Catalyst.
No person-entity page for the founder — this source is not speaker-aware.
Why it matters to this thread
The load-bearing framing (not the unstable numbers) attacks a premise the thread already tracks: "everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning?" If a latent bridge worked between untouched models from different families, "which model is best" would partially decompose into who supplies reasoning vs who supplies tokens — inference economics, not weights, as the contested surface. Full treatment: latent-space-model-handoff.
Mostik also says it is "committed to preventing frontier model lock-in" and is "partnering with inference providers to accelerate open-weight adoption" — a commercial interest, name it.
Key facts (from 2026-09-03-x-ai-pass-3-sep-2026-open-model-fleets-agentic-coding-price)
- Founder announcement: https://x.com/aimalysheva/status/2095232794792255848
- Company confirmation: https://x.com/mostik_ai/status/2095253666018058384
- Mechanism claim as quoted above (latent hidden-state handoff; no text; neither model fine-tuned; different families).
- Configuration named: 753B reads, 4B writes.
- Backing / team as stated in the founder post.
- Pointers named, not fetched: WIRED piece;
mostik.ai/read-more.
What this source does not establish
- Headline performance is not stable. Founder: 80% as accurate as the frontier model at 20x faster. Same-evening recap (@kimmonismus, commentary, not the lab): the bridge "lets the 4B model close half the performance gap to the 753B model" and the bridged pair "uses 2.5x less compute than a score-matched mid-sized model." Those are different quantities against different baselines; 20x latency ≠ 2.5x compute. Do not carry either pair into this page as a result.
- ARC-AGI "first place" is unverifiable by design. The founder says it cannot be discussed while the competition runs. Hold as an unresolved claim.
- No fetched writeup, paper, or WIRED piece.
- 753B is a size class, not a named model in these posts. Do not equate it to GLM 5.2's 753B figure on chinese-open-weight-frontier-parity without a fetched card.
Open questions
- What does
mostik.ai/read-moreactually report for accuracy, latency, and compute? - Which 753B and which 4B, and from which families?
- Does "neither model is fine-tuned" survive a methods section?