brain/
conceptartificial-intelligence

Anthropic illicit-distillation attributions (September 2026)

Notes

Vintage: 2026-09. Primary evidence is the issuer September 2026 threat intelligence report as synthesized in 2026-09-11-anthropic-sep-2026-threat-intelligence-autonomous-cyber (method: grok-bot; x_video: false; curl-retrieved 2026-09-11). All volumes and lab attributions are Anthropic unilateral assessments. Named labs’ public replies (denial, partial admit, silence) were not retrieved in this pass. Secondary press restates issuer figures — discovery only, not independent measurement. Do not merge with deepseek-v4-1-flash product grain. Distinct from the training-loop concept distillation-and-iterated-amplification.

Anthropic illicit-distillation attributions (September 2026)

One-line summary: Anthropic’s Sep 10 threat-intelligence report attributes large-scale illicit Claude distillation to China-based labs — largest Alibaba / Qwen (>151M exchanges May–Jul), then Moonshot / Kimi (>23M) and DeepSeek (>12.1M in 14 July days) — plus Zhipu / Xiaomi / SenseTime / MiniMax cases; issuer-alleged until primary replies land.

The insight

This is a dated issuer accusation with per-lab counts, not a rewrite of ITAD-as-training-loop and not a DeepSeek product card. The report’s distillation section names fraudulent-account pools, silent customer-request relays (Claude replies displayed as Kimi / DeepSeek), and CoT / SFT / “thinking signature” extraction. Prefer per-lab issuer figures over TechCrunch’s rolled-up “nearly 200 million exchanges / five campaigns” framing.

Keep lab distillation numbers tagged as Anthropic assessments. Do not invent lab replies.

Evidence

All figures below are from 2026-09-11-anthropic-sep-2026-threat-intelligence-autonomous-cyber citing the issuer report. Secondary (TechCrunch, CNBC, Reuters) restates them — not independent measurement.

  • Alibaba / Qwen (issuer largest): >151 million exchanges May–Jul 2026; peak nearly 3M/day from >3,500 fraudulent accounts; CoT/SFT transcripts said to advance Qwen 3.5 / 3.6 / 3.7; also Claude used for RL env / architecture R&D per Anthropic. Two fraudulent-account pools (~5,000 in first pool with residential proxies / disposable emails / virtual cards).
  • Moonshot / Kimi (GTG-16002): silently forwarded customer requests to Claude and displayed Claude replies as Kimi; one 10-day window ~300,000 requests via 5,380 fraudulent accounts (mostly Singapore/Japan appearance); May–Jul scale >23 million exchanges; CoT extraction via cross-session “thinking signature” replay; issuer flags sensitive customer traffic including assessed PLA-affiliated CCTV analysis in Chengdu and SOE engineer credentials. Whether customers were notified: issuer says unknown.
  • DeepSeek (GTG-16001): similar silent relay + CoT extraction; >12.1 million exchanges over 14 days in July 2026; targeted Opus reasoning traces; harness-string tagging (Claude Code / Agent SDK / OpenCode) to select traffic to relay. Do not merge with yesterday’s deepseek-v4-1-flash launch.
  • Zhipu / Z.ai (GTG-16006): CoT cleaner pipeline; 770,609 cleaner exchanges in 10 June days via 273 fraudulent accounts; >3.4M attributed over 17 days Jun–Jul; attempted Fable cyber distillation then switched to Opus 4.6 / another US lab after Fable safeguards degraded attacks (issuer).
  • Xiaomi (GTG-16008): >400k exchanges over 20 days Mar–Apr via >1,500 accounts; replayed MiMo sessions through Claude for SFT/RL rather than serving Claude to users (issuer assessment).
  • SenseTime / MiniMax: third-party reseller / shell-proxy ecosystem harvesting transcripts; MiniMax shell proxy said to offer only Anthropic/OpenAI models (issuer inference of harvest intent).
  • Defenses named (issuer): metadata attribution of proxy nets; adversarial-extraction classifiers (strengthened with Fable 5); summarized thinking; Fable 5.1 preserved/encrypted thinking; identity verification for abuse signals.

What this source does not establish

  • Named labs have not been independently confirmed here. Account clustering, shared prompts, and proxy geography are Anthropic’s method as described in the clip — not a third-party audit.
  • No lab replies in this pass. Do not invent denial / admit / silence.
  • Moonshot “PLA-affiliated” user and sensitive-customer examples are Anthropic investigative narrative; chain-of-custody to named institutions is not independently verified here.
  • TechCrunch “nearly 200 million / five campaigns” is secondary roll-up — prefer per-lab issuer totals.
  • Not a rewrite of distillation-and-iterated-amplification (ITAD as shared training loop). This page is a Sep 2026 issuer accusation with volumes.
  • Not a rewrite of deepseek-v4-1-flash or qwen-3-8-max-0902. Qwen 3.5 / 3.6 / 3.7 are the SKUs Anthropic names; 3.8-Max-0902 is a different card.
  • Did not create Alibaba / Xiaomi / Zhipu / SenseTime / MiniMax entity pages — one-source accusation grain stays here.
  • Did not re-file DeepSeek-V4.1-Flash, OpenAI Data agent, ChatGPT FS, Defense Factory / AISI, or the Sep 9 alignment assessment.

Contradictions / tensions

  • Unilateral volumes vs no lab reply. Keep tagged issuer-alleged.
  • Secondary roll-up vs per-lab issuer table. Prefer the table.
  • DeepSeek product launch (Sep 10 10am) vs DeepSeek distillation allegation (this Sep 10 report). Adjacent same-week lab, different claims. Do not merge performance benches with threat-intel accusations.
  • June 2026 Moonshots Alibaba sock-puppet color on distillation-and-iterated-amplification (Diamandis relaying 28.8M / 25k accounts) is an earlier, smaller accusation — chronological prior, not the same print.

Open questions

  • Do Alibaba, Moonshot, DeepSeek, Zhipu, Xiaomi, SenseTime, or MiniMax issue primary replies?
  • Were Moonshot / DeepSeek customers notified of the silent-relay claim (issuer: unknown)?
  • Do the Qwen 3.5 / 3.6 / 3.7 “advanced by CoT/SFT” claims show up in anything other than Anthropic’s assessment?

Sources

Related

Referenced by